Information processing device, information processing program and information processing method
The system seamlessly switches between manual and automatic avatar control modes using operator movement capture, addressing the challenge of user discomfort in telexistence mode transitions, ensuring operator flexibility and user comfort.
Patent Information
- Application Number
- JP2024060665
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-04
- Publication Date
- 2025-10-17
AI Technical Summary
Existing communication systems cannot smoothly switch from telexistence mode to chatbot mode, leading to potential user discomfort when operators perform other tasks while interacting with users.
An information processing device and method that includes a camera to capture operator movements, allowing seamless switching between manual and automatic modes for controlling avatar actions based on predefined conditions, ensuring the operator remains unaware of the mode changes.
Enables smooth transitions between manual and automatic avatar control modes, allowing operators to perform other tasks without user awareness, enhancing user comfort and operational efficiency.
Smart Images

Figure 2025158279000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an information processing program, and an information processing method, and more particularly to an information processing device, a control program, and a control method for moving an avatar in accordance with the action of an operator or automatically, for example. [Background technology]
[0002] An example of this type of conventional information processing device is disclosed in Patent Document 1. In the communication system disclosed in Patent Document 1, when a user stops in front of a response terminal, response processing is performed in chatbot mode. When chatbot mode is set, a response pattern for responding to the user's inquiry is selected from an extraction / classification pattern storage unit using machine learning or a rule base. The response pattern includes coordinate data for controlling the avatar's facial expressions (including blinking) and gestures (including hand and arm movements and facial direction) as well as response voice data. An avatar image that changes according to this response pattern is generated, and lip-sync processing is performed to move the avatar's lips while the response voice is being emitted in synchronization with the response voice data included in the response pattern.
[0003] When a user makes an inquiry that cannot be resolved by chatbot mode, the user calls an operator. In response, the response mode of the response terminal is switched from chatbot mode to telexistence mode and a connection request is sent to the operator terminal. With telexistence mode set, the operator terminal converts the operator's facial expressions and gestures into coordinate data and sends it to the response terminal along with response voice data. The response terminal generates an avatar based on the coordinate data sent from the operator terminal, thereby generating character response information in which the operator's facial expressions and gestures are reflected in the avatar's facial expressions and behavior, and displays it to the user. In this case, lip-sync processing is performed to move the avatar's lips in synchronization with the operator's response voice data while the response voice is being generated. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2021-56940 Summary of the Invention [Problem to be solved by the invention]
[0005] The communication system of Patent Document 1 mentioned above can be switched from chatbot mode to telexistence mode, but it is not anticipated that it will be possible to switch from telexistence mode to chatbot mode.
[0006] In telexistence mode, there may be cases where the operator wants to temporarily perform other tasks while attending to the user, but the avatar may also be moved by the operator's actions while performing other tasks, which may cause the user to feel uncomfortable. Therefore, there is room for improvement.
[0007] Therefore, a primary object of the present invention is to provide a novel information processing device, an information processing program, and an information processing method.
[0008] Another object of the present invention is to provide an information processing device, an information processing program, and an information processing method that can smoothly switch between a manual mode, in which an avatar performs actions that reflect the operator's movements, and an automatic mode, in which the avatar operates automatically, without the operator being aware of the operation. [Means for solving the problem]
[0009] A first invention is an information processing device used by an operator who interacts with a user, comprising: a camera that captures an image of the operator; an avatar action control means that, in manual mode, controls the action of an avatar of the operator displayed on a display device of a user terminal used by the user based on an image captured by the camera; and a mode switching means that, when the operator's action satisfies a first condition based on the image captured by the camera, switches from the manual mode to an automatic mode that automatically controls the action of the avatar, and, in the automatic mode, switches from the automatic mode to the manual mode when the operator's action satisfies a second condition corresponding to the first condition.
[0010] The second invention is dependent on the first invention, and the first condition is that the operator has performed an action other than an action to respond to the user, and the second condition is that the operator has returned to an action or state to respond to the user.
[0011] A third invention is dependent on the second invention, and has a plurality of first conditions and second conditions corresponding to each of the plurality of first conditions.
[0012] In a fourth aspect of the present invention, the avatar movement control means further includes transmission means for generating movement data for controlling the movement of the avatar based on the image captured by the camera, and transmitting the movement data to the user terminal.
[0013] In a fifth aspect of the present invention, the avatar movement control means further includes transmitting means for generating an avatar image reflecting a movement of the operator based on an image captured by the camera, and transmitting the avatar image to the user terminal.
[0014] A sixth invention is an information processing program executed on an information processing device used by an operator who interacts with a user and equipped with a camera that photographs the operator, wherein the processor of the information processing device executes an avatar action control step in a manual mode to control the action of an avatar of the operator displayed on a display device of a user terminal used by the user based on an image captured by the camera, and a mode switching step to switch from the manual mode to an automatic mode that automatically controls the action of the avatar when the operator's action satisfies a first condition based on the image captured by the camera, and to switch from the automatic mode to the manual mode when the operator's action satisfies a second condition corresponding to the first condition in the automatic mode.
[0015] A seventh invention is an information processing method for an information processing device used by an operator who interacts with a user and equipped with a camera that photographs the operator, wherein the processor of the information processing device, in manual mode, controls the behavior of an avatar of the operator displayed on a display device of a user terminal used by the user based on an image captured by the camera, and when the operator's behavior satisfies a first condition based on the image captured by the camera, switches from manual mode to automatic mode in which the avatar's behavior is automatically controlled, and when, in automatic mode, the operator's behavior satisfies a second condition corresponding to the first condition, switches from automatic mode to manual mode. [Effects of the Invention]
[0016] According to this invention, it is possible to smoothly switch between a manual mode in which the avatar performs actions that reflect the operator's movements and an automatic mode in which the avatar operates automatically, without the operator being aware of it.
[0017] The above and other objects, features and advantages of the present invention will become more apparent from the following detailed description of the preferred embodiments with reference to the drawings. [Brief explanation of the drawings]
[0018] [Figure 1] FIG. 1 is a diagram showing an information processing system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing the electrical configuration of the user terminal shown in FIG. [Figure 3] FIG. 3 is a block diagram showing the electrical configuration of the operator terminal shown in FIG. [Figure 4] FIG. 4(A) is a diagram showing an example of a user's response screen displayed on the display device of the user terminal, and FIG. 4(B) is a diagram showing an example of an image captured by the operator terminal. [Figure 5] Figure 5(A) is a diagram showing another example of a user's response screen displayed on the display device of the user terminal, and Figure 5(B) is a diagram showing another example of an image captured by the operator terminal. [Figure 6] FIG. 6 is a diagram showing an example of the operator's response screen displayed on the display device of the operator terminal. [Figure 7] FIG. 7 is a diagram showing an example of a memory map of the RAM of the user terminal shown in FIG. [Figure 8] FIG. 8 is a diagram showing an example of a memory map of the RAM of the operator terminal shown in FIG. [Figure 9] FIG. 9 is a diagram showing an example of specific contents of the data storage area shown in FIG. [Figure 10] FIG. 10 is a flowchart showing an example of an interaction process of the CPU of the user terminal shown in FIG. [Figure 11] FIG. 11 is a flowchart showing a part of an example of an interaction process of the CPU of the operator terminal shown in FIG. [Figure 12] FIG. 12 is a flowchart showing another part of an example of the interaction processing by the CPU of the operator terminal shown in FIG. 3, and follows FIG. [Figure 13] FIG. 13 is a flowchart showing a first part of an example of the mode switching process of the CPU of the operator terminal shown in FIG. [Figure 14] FIG. 14 is a flowchart showing a second part of an example of the mode switching process of the CPU of the operator terminal shown in FIG. 3, and follows FIG. [Figure 15]FIG. 15 is a flowchart showing a third part of an example of the mode switching process of the CPU of the operator terminal shown in FIG. 3, and follows FIGS. 13 and 14. In FIG. [Figure 16] FIG. 16 is a flowchart showing a fourth part of an example of the mode switching process of the CPU of the operator terminal shown in FIG. 3, and follows FIG. DETAILED DESCRIPTION OF THE INVENTION
[0019] Referring to FIG. 1, an information processing system 10 of this embodiment includes a user terminal 12, which is communicatively connected to an operator terminal 16 via a network .
[0020] As an example, Web Real-Time Communication (WebRTC) technology allows the user terminal 12 and the operator terminal 16 to perform P2P (Peer to Peer) communication and send and receive video and / or audio in real time via a web browser.
[0021] Although not shown in the figure, WebRTC can be used in combination with multiple servers. Specifically, the multiple servers required to use WebRTC include a signaling server, a STUN server, and a TURN server. Although not described in detail, the signaling server is a server used to obtain information about the other party communicating via WebRTC. Furthermore, the STUN server and TURN server are servers used to traverse so-called NAT (Network Address Translation) when the other party is located on a different network.
[0022] In this embodiment, one user terminal 12 and one operator terminal 16 are shown, but in reality, multiple user terminals 12 and multiple operator terminals 16 are connected to the network 14, and a connection state is established for one user terminal 12 and one operator terminal 16 that are set to conduct a video call or web conference.
[0023] In addition, if the operator terminal 16 to be connected to the user terminal 12 is determined in advance, the user terminal 12 has connection information for the operator terminal 16, and the operator terminal 16 has connection information for the user terminal 12, a connection state between the user terminal 12 and the operator terminal 16 can be established without using a server such as the one described above.
[0024] The user terminal 12 is used by a user who uses a service provided by a website implemented by a server on the network 14. The operator terminal 16 is used by a person (hereinafter referred to as the "operator") who responds to inquiries about the services provided by the website. In this embodiment, an avatar image (hereinafter referred to as the "avatar image") corresponding to the operator's avatar is displayed on the display device 30 of the user terminal 12, and the user and the operator interact through the avatar. The operator interacts with the user by operating the avatar, i.e., by controlling the avatar's movements and speech. In other words, the user is also a person who uses an interactive service using an avatar provided by the operator of the operator terminal 16.
[0025] It should be noted that a server that realizes a website that provides a predetermined service may function as the signaling server described above.
[0026] The user terminal 12 is an information processing device, and is a notebook PC or a desktop PC, but other general-purpose terminals such as a smartphone or a tablet PC can also be used.
[0027] The network 14 is composed of an IP network (or IP network) including the Internet, and an access network (or access network) for accessing this IP network. The access network may be a public telephone network, a mobile phone network, a wired LAN, a wireless LAN, or a CATV (Cable Television) network.
[0028] The operator terminal 16 is an information processing device different from the user terminal 12, and is, for example, a notebook PC or a desktop PC, but other terminals such as a smartphone or a tablet PC can also be used.
[0029] Fig. 2 is a block diagram showing the electrical configuration of the user terminal 12 shown in Fig. 1. As shown in Fig. 2, the user terminal 12 includes a CPU 20, which is connected to a RAM 22, a communication interface (hereinafter referred to as "communication I / F") 24, and an input / output interface (hereinafter referred to as "input / output I / F") 26 via an internal bus.
[0030] The CPU 20 is responsible for overall control of the user terminal 12. However, instead of the CPU 20, a SoC (System-on-a-chip) including multiple functions such as a CPU function and a GPU (Graphics Processing Unit) function may be provided.
[0031] The RAM 22 is a main storage device and is used as a work area or buffer area for the CPU 20. Although not shown in the figure, the user terminal 12 is provided with a HDD and a ROM as auxiliary storage devices. However, a non-volatile memory such as an SSD may be used instead of or in addition to the HDD.
[0032] The communication I / F 24 has a wired interface for transmitting and receiving control signals and data to and from an external computer such as the operator terminal 16 via the network 14 under the control of the CPU 20. However, a wireless interface such as a wireless LAN or Bluetooth (registered trademark) can also be used as the communication I / F 24.
[0033] An input device 28, a display device 30, a microphone 32, a speaker 34, and a camera 36 are connected to the input / output I / F 26.
[0034] The input device 28 is a keyboard and a computer mouse. In addition, a touch panel may be provided in some cases. However, if a smartphone or a tablet PC is used as the user terminal 12, the input device 28 is a touch panel and hardware buttons.
[0035] The display device 30 is an LCD or an organic EL display. The touch panel may be provided on the display surface of the display device 30, or a touch display in which the touch panel is integrally formed with the display device 30 may be provided. This also applies to the operator terminal 16 described later. The microphone 62 is a sound-collecting microphone. The speaker 64 is a stereo speaker. The camera 66 is a CCD camera or a CMOS camera.
[0036] The input / output I / F 26 outputs operation data (or operation information) input from the input device 28 to the CPU 20, and also outputs image data generated by the CPU 20 to the display device 30, causing a screen corresponding to the image data to be displayed on the display device 30. However, there are also cases where image data received from an external computer (for example, the operator terminal 16) is output by the CPU 20.
[0037] The input / output I / F 26 converts the user's voice detected by the microphone 32 into digital voice data and outputs it to the CPU 20, converts voice data output by the CPU 20 into an analog voice signal and outputs it from the speaker 34, and outputs video data of a video including the user captured (detected) by the camera 36 (hereinafter referred to as "user video data") to the CPU 20. However, in this embodiment, the voice data output from the CPU 20 is operator voice data received from the operator terminal 16 or voice data obtained by converting the operator's voice into avatar voice. Furthermore, the user's video is a moving image.
[0038] 2 is an example and is not intended to be limiting. In another example, the user terminal 12 may include a camera.
[0039] Furthermore, if the user terminal 12 is a smartphone, it is equipped with a call circuit for making calls via a mobile phone network or via a mobile phone network and a public telephone network, but in this embodiment, such calls are not made, so they are not shown in the figure. This is also true when the operator terminal 16, which will be described later, is a smartphone.
[0040] Fig. 3 is a block diagram showing the electrical configuration of the operator terminal 16 shown in Fig. 1. As shown in Fig. 3, the operator terminal 16 includes a CPU 50, which is connected to a RAM 52, a communication I / F 54, and an input / output I / F 56 via an internal bus.
[0041] The CPU 50 is responsible for overall control of the operator terminal 16. However, instead of the CPU 50, an SoC including multiple functions such as a CPU function and a GPU function may be provided.
[0042] The RAM 52 is a main storage device and is used as a work area or buffer area for the CPU 50. Although not shown, the operator terminal 16 is provided with a HDD and a ROM as auxiliary storage devices. However, a non-volatile memory such as an SSD may be used instead of or in addition to the HDD.
[0043] The communication I / F 54 is a wired interface for transmitting and receiving control signals and data to and from an external computer such as the operator terminal 16 via the network 14 under the control of the CPU 50. However, a wireless interface such as a wireless LAN or Bluetooth (registered trademark) can also be used as the communication I / F 54.
[0044] An input device 58, a display device 60, a microphone 62, a speaker 64, and a camera 66 are connected to the input / output I / F 56.
[0045] The input device 58 is a keyboard and a computer mouse. In addition, a touch panel may be provided in some cases. However, when a smartphone or a tablet PC is used as the operator terminal 16, the input device 58 is a touch panel and hardware buttons.
[0046] The display device 60 is an LCD or an organic EL display. The microphone 62 is a sound-collecting microphone. The speaker 64 is a stereo speaker. The camera 66 is a CCD camera or a CMOS camera.
[0047] The input / output I / F 56 outputs operation data (or operation information) input from the input device 58 to the CPU 50, and also outputs image data generated by the CPU 50 to the display device 60, causing a screen corresponding to the image data to be displayed on the display device 60. However, image data received from an external computer (for example, the user terminal 12) may also be output by the CPU 50. The image data received from the user terminal 12 is user video data.
[0048] Furthermore, the input / output I / F 56 converts the operator's voice detected by the microphone 62 into digital voice data and outputs it to the CPU 50, converts the voice data output by the CPU 50 into an analog voice signal and outputs it from the speaker 64, and outputs video data of a video including the operator captured (detected) by the camera 66 (hereinafter referred to as "operator video data") to the CPU 50. However, in this embodiment, the voice data output from the CPU 50 is user voice data received from the user terminal 12. Furthermore, the video of the operator is a moving image.
[0049] The electrical configuration of the operator terminal 16 shown in FIG. 3 is an example and does not need to be limited to this.
[0050] In such an information processing system 10, as described above, when a user is using his / her own user terminal 12 to view the web screen of a specified service, the user can interact with an operator and make an inquiry about the specified service.
[0051] 4 shows an example of user-side response screen 100 displayed on display device 30 of user terminal 12 when a user and an operator are interacting. User-side response screen 100 may be displayed over a web screen of a predetermined service, as part of the web screen, or in place of the web screen of the predetermined service. However, when displayed as part of the web screen, only avatar image 102 may be displayed.
[0052] The user's response screen 100 includes an avatar image 102 representing the operator, and the avatar image 102 is an image of a character that resembles a human female. In the example shown in Fig. 4, the avatar image 102 on the user's response screen 100 is an image of the character's upper body. As an example, the background is filled with a predetermined color.
[0053] The avatar image 102 may also be an image of an animal or robot character, an anime character, a game character, or the like. The avatar image 102 may also be an image of the entire body of the character. However, in this embodiment, the avatar must have a skeletal structure similar to that of a human, since the avatar may move in accordance with the operator's movements.
[0054] Avatar image 102 or user-side response screen 100 is displayed on display device 30 when certain activation conditions are met. The certain activation conditions are that the user instructs the operator to be called, that the user does not perform any operation for a first predetermined time period (30 seconds in this embodiment) or longer, that the user instructs the same location or a similar location (a nearby location) on the web screen, and that the user returns to the same web screen multiple times (e.g., three times). However, except when the user instructs the operator to be called, when the certain activation conditions are met, user terminal 12 automatically calls the operator.
[0055] When the user or the user terminal 12 calls the operator, a connection state between the user terminal 12 and the operator terminal 16 is established by a signaling server or the like.
[0056] When the operator is responding to the user and the user and the operator are engaged in a conversation, the user terminal 12 detects the user's voice through the microphone 32 and transmits voice data corresponding to the detected voice (hereinafter referred to as "user voice data") to the operator terminal 16. The operator terminal 16 outputs the received user voice data to the speaker 64. Therefore, the operator can hear the user's voice.
[0057] Furthermore, when the user and the operator are engaged in a conversation, the camera 36 of the user terminal 12 captures an image of the user, and then transmits image data corresponding to the captured image of the user (moving image in this embodiment), i.e., user video data, to the operator terminal 16. The operator terminal 16 outputs the received user video data to the display device 60. Therefore, the operator can view the user's video. However, it is also possible to prevent the user video data from being transmitted to the operator terminal 16 by turning off the function of the camera 36.
[0058] On the other hand, when the operator terminal 16 detects the operator's voice through the microphone 62, it transmits voice data corresponding to the detected voice to the user terminal 12. The user terminal 12 outputs the received voice data to the speaker 34. Therefore, the user can hear the operator's voice.
[0059] If the avatar is a non-human character such as an animal or robot character, an anime character, or a game character, the voice can be converted into the voice of that character and output.
[0060] In addition, in this embodiment, the operator terminal 16 has two operating modes: a mode in which the avatar's movements are controlled by the operator's movements (hereinafter referred to as the "manual mode"), and a mode in which the avatar's movements are controlled automatically without relying on the operator's movements (hereinafter referred to as the "automatic mode").
[0061] In this embodiment, for ease of explanation, the mode in which the movement of the avatar is controlled by the movement of the operator is called "manual mode" as opposed to "automatic mode." However, in manual mode, the movement of the avatar is controlled based on a captured image of the operator (in this embodiment, a moving image), so the operator does not operate the input device 58 to move the avatar.
[0062] In manual mode, the avatar moves in accordance with the operator's movements. That is, the operator's movements are reflected in the avatar's movements. Therefore, when the operator responds to a user and the user and the operator are conversing, the avatar moves in accordance with the operator's movements. Furthermore, when the operator speaks, the avatar moves in accordance with the operator's movements and the operator's voice is output. In this case, avatar image 102 is lip-synchronized. That is, the lips of avatar image 102 move in accordance with the output of the voice uttered by the operator. Therefore, the avatar appears to be actually speaking.
[0063] In this embodiment, avatar movement includes moving the head, neck, and hands of avatar image 102, as well as changing the facial expression, eyelids, and lips of avatar image 102.
[0064] However, the facial expression of avatar 102 may be fixed to a predetermined expression (for example, a smile) and not changed.
[0065] In the manual mode, user terminal 12 controls the movement of the avatar using movement data generated based on a captured image of the operator, and generates and displays (or updates) image data for user-side response screen 100. A method for generating image data for avatar image 102 will be described later.
[0066] 4, the avatar faces forward, with its right hand slightly raised and its left hand lowered. In this avatar image 102, the avatar's facial expression is one of joy (here, a smile).
[0067] Fig. 5 shows an example of a captured image 150 captured by the camera 68 of the operator terminal 16 when the user and the operator are interacting. In the example shown in Fig. 5, the captured image 150 includes an image 152 of the subject, i.e., the operator, who is a human female. However, the image 152 is a moving image. For simplicity, the background image has been omitted from the captured image 150 shown in Fig. 5.
[0068] 5, in the photographed image 150, the operator faces forward, with his right hand slightly raised and his left hand lowered. In this photographed image 150, the operator's facial expression is one of joy (smiling).
[0069] 4 and 5, the movements and facial expression of the operator in captured image 150 are reflected in the movements and facial expression of avatar image 102. Furthermore, although it is difficult to see in FIGS. 4 and 5, changes (movements) in the eyelids and lips of the operator are also reflected in changes (movements) in the eyelids and lips of avatar image 102.
[0070] In the manual mode, motion data for controlling the motion of the avatar is generated in the operator terminal 16 based on the motion of the operator. To distinguish it from motion data in the automatic mode (automatic motion data), which will be described later, the motion data in the manual mode is sometimes called manual motion data. However, when there is no need to distinguish between automatic motion data and manual motion data, it is simply called "motion data."
[0071] The movement data is information (data) about the position and orientation of each part of the face for controlling the avatar's gestures, facial orientation, and facial expressions. Specifically, the movement data is information about the positions (coordinate data) of each part of the face, namely the eyebrows, eyes (upper eyelids, lower eyelids), nose, and mouth (upper lip, lower lip), the positions (coordinate data) of the head, hands, arms, shoulders, waist, knees, and ankles, and the orientation of the neck (or face), wrists, and ankles. However, if the avatar's abdomen and below is not displayed, information about the positions and orientations of the waist, knees, and ankles does not need to be included in the movement data.
[0072] The motion data for controlling the motion of the avatar based on the motion of the operator can be generated using an image processing library such as MediaPipe Holistic. In another example, the method disclosed in Japanese Patent Application Laid-Open No. 2021-56940 in the background art can be used.
[0073] In the manual mode, voice data (hereinafter referred to as "operator voice data") for controlling the speech (voice output) of the avatar based on the voice of the operator is generated in the operator terminal 16. The operator voice data may be data of the operator's voice, or may be data obtained by converting the operator's voice into the voice of the avatar.
[0074] The generated manual operation data and operator voice data are transmitted to user terminal 12. User terminal 12 operates the avatar in accordance with the received manual operation data, and generates and displays (or updates) avatar image 102. Thus, the avatar is operated. User terminal 12 also outputs the received operator voice data to speaker 34. At this time, avatar image 102 is lip-synchronized.
[0075] On the other hand, in the automatic mode, one piece of automatic action data is selected in a predetermined manner from among multiple pieces of pre-generated action data, i.e., automatic action data, regardless of the operator's action. The selected piece of automatic action data is transmitted to user terminal 12. Even in the automatic mode, user terminal 12 displays (or updates) user-side response screen 100 including avatar image 102 on display device 30. Therefore, the avatar operates independently of the operator's action, and a predetermined facial expression is reflected. Furthermore, when the operator speaks, the operator's voice data is also transmitted to user terminal 12, and avatar image 102 is lip-synchronized.
[0076] In this embodiment, when the operator speaks, the avatar's movement, i.e., automatic operation data, is selected (or determined) according to what the operator says. As a method for selecting (or determining) the avatar's movement according to what the operator says, the method disclosed in Japanese Patent Application Laid-Open No. 2020-6482 can be applied. However, in Japanese Patent Application Laid-Open No. 2020-6482, in order to determine the gestures (i.e., arm movements) of Android (registered trademark), the same method is appropriately modified to determine the avatar's movement.
[0077] In addition, the automatic operation data can be selected according to the emotion of the operator or user. The emotion of the operator or user can be estimated using a large-scale language model. A large-scale language model is a natural language processing model trained by machine learning using a large amount of text data, and can be used for various natural language processing such as text classification, sentiment analysis, information extraction, text summarization, text generation, and question answering. The history of the dialogue between the operator and the user and an estimation of the current emotion of the operator or user based on this history can be input to the large-scale language model as a prompt, and an estimation result of the emotion of the operator or user can be obtained from the large-scale language model.
[0078] However, if a facial image can be acquired from a video of the operator or user, the emotion of the operator or user can be estimated based on the facial image of the operator or user, and automatic operation data can be selected according to the estimated emotion. Since a method for estimating the emotion of a human being, such as an operator, using a facial image is already known, a description of this method will be omitted.
[0079] Known techniques can be used to estimate human emotions from facial images, such as those disclosed in "Hiroshi Kobayashi and Fumio Hara: Recognition of Basic Human Facial Expressions Using Neural Networks, Transactions of the Society of Instrument and Control Engineers, Vol. 29, No. 1, 112 / 118 (1993)," "Yosuke Koyanaka, Tsuneyasu Homma, Masao Sakai, and Kenichi Abe: Facial Expression Recognition Using Neural Networks, Bulletin of the Tohoku University School of Medicine and Health Sciences, 13(1):23-32, 2004," and "Daiki Nishime, Satoshi Endo, Yoshiaki Touma, Koji Yamada, and Yuhei Akamine: Acquisition of Facial Expressions and Analysis of Facial Features Using Convolutional Neural Networks, Transactions of the Japanese Society for Artificial Intelligence, Vol. 32, No. 5, FZ (2017)."
[0080] Another known technique for estimating human emotions based on feature points extracted from a facial image is the technique disclosed in Japanese Patent Application Laid-Open No. 2020-163660.
[0081] The emotions of the operator or user can also be estimated based on the operator's or user's voice. Known technologies can be used to estimate human emotions from voice. For example, the technologies disclosed in JP 2021-12285 A and "Mori, Daiki: Understanding Emotions and Attitudes from Speech," Journal of the Institute of Electronics, Information and Communication Engineers, Vol. 101, No. 9, 2018, can be used.
[0082] In other words, the automatic action data is action data about the avatar's actions determined in accordance with the content or emotion of speech.
[0083] As described above, regardless of whether the operation mode is manual or automatic, user terminal 12 controls the movement and speech of avatar image 102 (avatar) displayed on display device 30 in accordance with data from operator terminal 16. However, the data from operator terminal 16 may be movement data only, or movement data and operator voice data.
[0084] In addition, in this embodiment, operation data is transmitted from the operator terminal 16 to the user terminal 12, and the avatar image 102 is generated and displayed on the user terminal 12. However, it is also possible to generate image data for the avatar image 102 on the operator terminal 16 and transmit it to the user terminal 12, and display the user's response screen 100 including the avatar image 102 on the user terminal 12.
[0085] During a conversation, the operation mode of the avatar is set to manual mode or automatic mode on the operator terminal 16. In this embodiment, when the user and the operator start a conversation, the operation mode is set to manual mode. The operator may want to perform other tasks while attending to the user, and in such cases, the operation mode is temporarily set from manual mode to automatic mode. Examples of other tasks include creating, sending, or receiving e-mail, turning one's attention elsewhere, or leaving the desk. Furthermore, when the operator becomes available to attend to the user, the operation mode is set from automatic mode to manual mode.
[0086] In this embodiment, when the operation mode is the manual mode, if a first condition for setting the operation mode to the automatic mode is satisfied, the operation mode is set to the automatic mode. That is, the operation mode is switched from the manual mode to the automatic mode. Also, when the operation mode is the automatic mode, if a second condition for setting the operation mode to the manual mode is satisfied, the operation mode is set to the manual mode. That is, the operation mode is switched from the automatic mode to the manual mode.
[0087] In this embodiment, three first conditions are set, and a second condition is set individually corresponding to each first condition. The first condition is that the operator is performing an action other than an action corresponding to a user. The second condition is that the operator has returned to an action (or posture) or state corresponding to a user.
[0088] The first condition was that the operator operated the keyboard or computer mouse. Corresponding to this first condition, the second condition was that the operator was no longer operating the keyboard or computer mouse and returned to facing forward.
[0089] In the manual mode, the operator terminal 16 determines whether there is a key input or a mouse input, and if there is a key input or a mouse input, it determines that the first condition is met.
[0090] In addition, in the automatic mode, the operator terminal 16 determines whether there is a key input or a mouse input, and if the state of no key input or mouse input continues for a second predetermined time (for example, 30 seconds) and the operator's face returns to facing forward, it is determined that the second condition corresponding to the first first condition is met. Whether the operator's face is facing forward is determined based on the current direction of the operator's face.
[0091] The current face orientation of the operator is calculated or detected based on the current captured image (i.e., face image) with the orientation of the face image of the operator facing the camera 66 as a reference. The face orientation of the operator when facing the camera 66 is the front orientation. In this embodiment, the face orientation when the operator's face is facing forward is set to 0 degrees, and the angle of the current face orientation (three-dimensional direction) of the operator is calculated. However, the face orientation of the operator can be detected by the difference between multiple feature points extracted from the face image of the operator facing the camera 66 and multiple feature points extracted from the current face image. When detecting the face orientation (head orientation) of the operator based on the captured image, an image processing library such as MediaPipe Pose or MediaPipe Face Mesh can be used. It is also possible to extract the face (head) region of the operator from the captured image and detect the face orientation (head orientation) of the operator from changes in the face (head) region. The same applies to the calculation or detection of the face orientation of the operator below.
[0092] The second condition (1) is that the operator's face is not recognized or is facing in a direction other than the front. Corresponding to the first condition, the second condition (2) is that the operator's face returns to facing the front.
[0093] In the operator terminal 16, in manual mode, a known face recognition process is performed on the captured image of the operator to determine whether the operator's face has been recognized. If the operator's face has not been recognized, it is determined that the second first condition is met. If the operator's face has been recognized, it is further determined whether the operator's face is deviated from the front by a predetermined angle (e.g., 60 degrees) or more. Specifically, it is determined whether the angle of the current face orientation is deviated by a predetermined angle in the left / right and / or up / down directions from a reference direction. However, the angle in the direction of face rotation, i.e., the angle in the direction of tilting the head sideways, is excluded, because the operator's face is facing forward.
[0094] For face recognition processing, you can use image processing libraries such as Face API (a face recognition application provided by Microsoft (company name)) or MediaPipe Face Mesh.
[0095] If it is determined that the operator's face is displaced from the front in any direction by a predetermined angle or more, it is determined that the second first condition is met. However, if it is determined that the operator's face is not displaced from the front in either the up / down and / or left / right directions by a predetermined angle or more, it is not determined that the second first condition is met.
[0096] In this embodiment, if it is determined that the operator's face is shifted by more than a predetermined angle, it is determined that the second first condition is met, but it may also be determined that the second first condition is met if the state in which the operator's face is shifted by more than a predetermined angle continues for a certain period of time (for example, 5 seconds), because it is possible that the operator may look away for an instant.
[0097] In addition, in the automatic mode, the operator terminal 16 determines whether the operator's face is facing forward, and if it is determined that the operator's face is facing forward, it determines that the second condition corresponding to the second first condition is met, and if it is not determined that the operator's face is facing forward, it does not determine that the second conditions corresponding to the two first conditions are met.
[0098] In this embodiment, when it is determined that the operator's face is facing forward, it is determined that the second condition corresponding to the second first condition is met, but it may also be determined that the second condition corresponding to the second first condition is met when the state in which the operator's face is facing forward continues for a certain period of time (for example, 5 seconds). This is because it is conceivable that the operator may look straight ahead for an instant while performing other work, etc.
[0099] The third first condition is that the operator has performed a predetermined first movement pattern. For example, the predetermined first movement pattern is the movement of taking notes. Corresponding to this first condition, the third second condition is that the operator has performed a predetermined second movement pattern. For example, the predetermined second movement pattern is the movement of turning one's face from a slightly lowered position to facing forward.
[0100] In the manual mode, the operator terminal 16 detects (estimates) whether or not the operator is taking notes by pattern recognition using a neural network based on a captured image (still image or video) of the operator. If it is detected that the operator is taking notes, it is determined that the third first condition is met.
[0101] Furthermore, in the automatic mode, the operator terminal 16 detects (estimates) whether the operator is turning his / her face from a slightly lowered position to the front by pattern recognition using a neural network based on a captured image (still image or video). However, instead of pattern recognition, as described above, it is also possible to detect whether the operator's face is facing forward based on facial feature points extracted from the facial image. When it is detected that the operator is turning his / her face from a lowered position to the front (or that the operator is facing forward), it is determined that the second condition corresponding to the third first condition is satisfied.
[0102] In this way, the operation mode can be switched from the manual mode to the automatic mode by the operator performing an action other than the action of responding to the user, and can be switched from the automatic mode to the manual mode by performing an action returning to the action of responding to the user. In other words, the operation mode can be switched seamlessly between the manual mode and the automatic mode without the operator being aware of it.
[0103] As described above, in the automatic mode, the avatar's movements are automatically controlled, allowing the operator to perform other tasks while interacting with the user.
[0104] The reason why the operation mode is changed from automatic mode to manual mode when it is determined that the second condition corresponding to the first condition is satisfied is to correctly determine that the operator has finished the task that they were performing when the operation mode was switched from manual mode to automatic mode and has returned to a state in which they can attend to the user. However, even when any of the first conditions is satisfied to switch the operation mode from manual mode to automatic mode, if it is detected that there is no key input or mouse input and that the operator is facing forward, it is possible to determine that the operator has returned to a state in which they can attend to the user, so it is also possible to set only the first second condition described above for each first condition.
[0105] Fig. 6 shows an example of the operator's side response screen 200 displayed on the display device 60 of the operator terminal 16. As shown in Fig. 6, the operator's side response screen 200 has a display area 202 in the center of the screen, and a button image 204 below the display area 202.
[0106] The display area 202 is an area where user video data is output. A video of the user being attended to by the operator is displayed. The button image 204 is an icon for ending the user's attendance. When the button image 204 is turned on, a notification of the end of the attendance is sent to the user terminal 12.
[0107] Fig. 7 shows an example of a memory map 300 of the RAM 22 built into the user terminal 12. As shown in Fig. 7, the RAM 22 includes a program storage area 302 and a data storage area 304. The program storage area 302 stores an interactive processing program, which is an example of an information processing program executed by the user terminal 12 of this embodiment.
[0108] The interactive processing programs executed on the user terminal 12 include a main processing program 302a, an operation detection program 302b, a communication program 302c, an image generation program 302d, an image output program 302e, a photography program 302f, an avatar control program 302g, a sound detection program 302h, and a sound output program 302i.
[0109] The main processing program 302a is a program for executing the main routine of the interactive processing of the user terminal 12 of this embodiment.
[0110] The operation detection program 302b is a program for detecting operation data 304a input from the input device 28 in accordance with a user's operation and storing the data in the data storage area 304.
[0111] The communication program 302c is a program for communicating (sending and receiving data, etc.) with an external device, in this embodiment, the operator terminal 16.
[0112] Image generation program 302d is a program for generating image data corresponding to all or part of a screen (such as 100) or image to be displayed on display device 30, using image generation data 304b.
[0113] The image output program 302e is a program for outputting image data generated according to the image generation program 302d to the display device 30. Therefore, a screen or image corresponding to the image data is displayed on the display device 30.
[0114] The photographing program 302f is a program for causing the camera 36 to perform photographing processing and storing user video data 304c input from the camera 36 in the data storage area 304.
[0115] The avatar control program 302g is a program for moving the avatar using the avatar movement data 304e.
[0116] The sound detection program 302h is a program for detecting the user's voice input from the microphone 32 and storing in the data storage area 304 user voice data 304d corresponding to the detected voice.
[0117] The sound output program 302i is a program for outputting the operator voice data 304f received from the operator terminal 16 to the speaker 34. Therefore, the voice corresponding to the operator voice data 304f or the voice obtained by converting this voice into the voice of an avatar is output from the speaker 34.
[0118] Although not shown in the figure, the program storage area 302 stores the operating system and middleware of the user terminal 12, as well as a browser and other application programs other than the interactive processing program of the present application.
[0119] Data storage area 304 stores operation data 304a, image generation data 304b, user video data 304c, user voice data 304d, avatar movement data 304e, operator voice data 304f, and the like.
[0120] The operation data 304a is data of a user's operation detected in accordance with the operation detection program 302b. The operation data 304a is deleted from the data storage area 304 after being used for processing by the CPU 20.
[0121] Image generation data 304b is data for generating a screen or image to be displayed on display device 30 of user terminal 12, and includes data for generating avatar image 102.
[0122] The user video data 304c is data of video captured by the camera 36. This user video data 304c is basically data of video including the user, but if the user leaves their seat or otherwise moves out of the range of the camera 36, the video may not include the user's video. When the user video data 304c is transmitted to the operator terminal 16 by the CPU 20, it is erased from the data storage area 304.
[0123] The user voice data 304d is data of the user's voice detected by the microphone 32. When the user voice data 304d is transmitted to the operator terminal 16 by the CPU 20, it is deleted from the data storage area 304.
[0124] Avatar movement data 304e is manual movement data or automatic movement data received from operator terminal 16. Avatar movement data 304e is deleted from data storage area 304 after being used for processing by CPU 20.
[0125] The operator voice data 304f is operator voice data received from the operator terminal 16. When the operator voice data 304f is output to the speaker 34 by the CPU 20, it is deleted from the data storage area 304.
[0126] Although not shown, the data storage area 304 stores other data required to execute the interactive processing, and is provided with a timer (counter) and flags required to execute the interactive processing.
[0127] Fig. 8 shows an example of a memory map 500 of the RAM 52 built into the operator terminal 16. As shown in Fig. 8, the RAM 52 includes a program storage area 502 and a data storage area 504. The program storage area 502 stores an interactive processing program, which is an example of an information processing program executed by the operator terminal 16 of this embodiment.
[0128] The interactive processing programs executed on the operator terminal 16 include a main processing program 502a, an operation detection program 502b, a communication program 502c, an image generation program 502d, an image output program 502e, a photography program 502f, an action data generation program 502g, an action data selection program 502h, a face recognition program 502i, an orientation detection program 502j, an action pattern detection program 502k, an action mode switching program 502m, a sound detection program 502n, and a sound output program 502p.
[0129] The main processing program 502a is a program for executing the main routine of the interactive processing of the operator terminal 16 in this embodiment.
[0130] The operation detection program 502b is a program for detecting operation data 504a input from the input device 58 in accordance with an operator's operation, and temporarily storing the data in the data storage area 504.
[0131] The communication program 502c is a program for communicating (sending and receiving data, etc.) with an external device, in this embodiment, the user terminal 12.
[0132] Image generation program 502d is a program for using image generation data 504b to generate image data corresponding to a screen (such as 200) or image to be displayed on display device 60. However, when generating image data corresponding to operator-side response screen 250 as shown in FIG. 6, user video data 504f is also used.
[0133] The image output program 502e is a program for outputting the image data generated according to the image generation program 502d to the display device 60. Therefore, a screen or an image corresponding to the image data is displayed on the display device 60.
[0134] The photographing program 502f is a program for causing the camera 66 to perform photographing processing and temporarily storing the operator video data 504c input from the camera 66 in the data storage area 504.
[0135] The motion data generation program 502g is a program for detecting the motion of the operator from the operator video data 504c in the manual mode and generating manual motion data 504e for controlling the motion of the avatar.
[0136] The action data selection program 502h is a program for selecting automatic action data 504h for actions corresponding to what the operator says in the automatic mode. However, the automatic action data 504h can also select automatic action data 504h for actions corresponding to the emotions of the operator or user.
[0137] The face recognition program 502i is a program for recognizing the face of the operator from the operator video data 504c.
[0138] The direction detection program 502j is a program for detecting the direction of the operator's face from the operator video data 504c.
[0139] The motion pattern detection program 502k is a program for detecting whether the motion of the operator is a predetermined motion pattern, in this embodiment, a predetermined first motion pattern or a predetermined second motion pattern.
[0140] The operation mode switching program 502m is a program for switching the operation mode between the manual mode and the automatic mode.
[0141] The sound detection program 502n is a program for detecting the operator's voice input from the microphone 62 and temporarily storing in the data storage area 504 operator voice data 504d corresponding to the detected voice.
[0142] The sound output program 502p is a program for outputting user voice data 504g received from the user terminal 12 to the speaker 64. Therefore, the speaker 64 outputs a voice corresponding to the user voice data 504g.
[0143] Although not shown, the program storage area 502 stores the operating system and middleware of the operator terminal 16, as well as a browser and other application programs other than the interactive processing program of the present application.
[0144] Fig. 9 shows an example of the specific contents of data storage area 504 shown in Fig. 8. Data storage area 504 stores operation data 504a, image generation data 504b, operator video data 504c, operator voice data 504d, manual operation data 504e, user video data 504f, user voice data 504g, automatic operation data 504h, operation mode data 504i, and orientation data 504j.
[0145] The operation data 504a is data of the operation of the operator detected in accordance with the operation detection program 502b. The operation data 504a is deleted from the data storage area 504 after being used for processing by the CPU 50.
[0146] Image generation data 504b is data used to generate a screen or image to be displayed on display device 60 of operator terminal 16.
[0147] The operator video data 504c is data of video captured by the camera 66. This operator video data 504c is basically data of video including the operator, but if the operator leaves his / her seat or otherwise moves out of the range of the camera 66, the video may not include the operator. Once the operator video data 504c is used for processing by the CPU 50, it is deleted from the data storage area 504.
[0148] The operator voice data 504d is data of the voice of the operator detected by the microphone 62. Once the operator voice data 504d has been used for processing by the CPU 50, it is deleted from the data storage area 504.
[0149] The manual action data 504e is data for controlling the action of the avatar generated based on the operator video data 504c. When the manual action data 504e is transmitted to the user terminal 12 by the CPU 50, it is deleted from the data storage area 504.
[0150] The user video data 504f is user video data received from the user terminal 12. Once the user video data 504f has been used for processing by the CPU 50, it is deleted from the data storage area 504.
[0151] The user voice data 504g is user voice data received from the user terminal 12. When the user voice data 504g is output to the speaker 64 by the CPU 50, it is deleted from the data storage area 504.
[0152] The automatic action data 504h is data for controlling a plurality of actions of a pre-generated avatar, and includes data corresponding to each of a plurality of pre-set actions.
[0153] The operation mode data 504i is data for identifying whether the operation mode is the manual mode or the automatic mode.
[0154] The orientation data 504j is data on the orientation of the operator's face detected according to the orientation detection program 502j. In this embodiment, the orientation data 504j is data on the angle of the current orientation of the face (three-dimensional direction) based on the direction in which the face faces forward.
[0155] The data storage area 504 also stores an input flag 504k, a face non-recognition flag 504m, and a motion recognition flag 504n.
[0156] The input flag 504k is a flag for determining whether or not there is a key input or mouse input by the operator. If there is a key input or mouse input by the operator, the input flag 504k is turned on, and if there is no key input or mouse input by the operator for a second predetermined time (30 seconds in this embodiment), the input flag 504k is turned off.
[0157] The face unrecognition flag 504m is a flag for determining whether the face of the operator is not recognized. If the face of the operator is not recognized, the face unrecognition flag 504m is turned on, and if the face of the operator is recognized, the face unrecognition flag 504m is turned off.
[0158] The action recognition flag 504n is a flag for determining whether a predetermined first action pattern or a predetermined second action pattern of the operator has been recognized. When the predetermined first action pattern of the operator has been recognized, the action recognition flag 504n is turned on, and when the predetermined second action pattern has been recognized, the action recognition flag 504n is turned off.
[0159] Although not shown, the data storage area 504 stores other data required to execute the interactive processing, and is provided with a timer (counter) and flags required to execute the interactive processing.
[0160] 10 is a flow diagram showing the dialogue processing of the CPU 20 of the user terminal 12. Although not shown in the figure, in parallel with the dialogue processing, the CPU 20 performs operations such as detecting operation data and storing the operation data, causing the camera 36 to perform a photographing process and storing user video data, detecting audio and storing user audio data, and receiving and storing data transmitted from the operator terminal 16. The dialogue processing of the CPU 20 will be described below with reference to FIG. 10, but overlapping content will be briefly described.
[0161] 10, when the CPU 20 of the user terminal 12 starts the dialogue process, in step S1, it establishes a connection with the operator terminal 16, and in step S3, it displays the user's side greeting screen 100 as shown in FIG. 4 on the display device 30. At this time, since the user terminal 12 has not received the avatar movement data 304e, the user's side greeting screen 100 including the upright avatar image 102 is displayed, for example. At this time, the avatar may also perform an action (and speak) to greet the user.
[0162] In the next step S5, it is determined whether the dialogue has ended. Here, CPU 20 determines whether there has been an instruction to end the dialogue from the user, and whether there has been a notification of the end of the dialogue from operator terminal 16. However, if the dialogue is to be ended in accordance with the user's instruction, the operator terminal 16 is notified that the dialogue has ended.
[0163] If step S5 returns "YES," that is, if the dialogue has ended, the dialogue processing ends. On the other hand, if step S5 returns "NO," that is, if the dialogue has not ended, step S7 determines whether data has been received from operator terminal 16. Here, CPU 20 determines whether avatar movement data 304e, or avatar movement data 304e and operator voice data 304f, has been received.
[0164] If "NO" in step S7, that is, if data has not been received from the operator terminal 16, the process proceeds to step S17. On the other hand, if "YES" in step S7, that is, if data has been received from the operator terminal 16, the process proceeds to step S9, where it is determined whether or not operator voice data 304f is present.
[0165] If step S9 returns "YES," that is, if operator voice data 304f is present, then in step S11, operator voice data 304f is output to speaker 34. Then, in step S13, avatar image data for performing an action in accordance with avatar movement data 304e is generated in accordance with the output of operator voice data 304f and output to display device 30, and the process proceeds to step S17. Accordingly, avatar 102 is updated, and the avatar displayed on user-side response screen 100 speaks to the user while making gestures and changing the direction of its face. At this time, avatar 102 is lip-synchronized with the voice output from speaker 34.
[0166] On the other hand, if step S9 is “NO,” that is, if operator voice data 304f is not available, in step S15, avatar image data for performing an action in accordance with avatar action data 304e is generated and output to display device 30, and the process proceeds to step S17. Accordingly, avatar image 102 is updated, and the avatar displayed on user-side response screen 100 makes gestures and changes the direction of its face.
[0167] In step S17, it is determined whether or not there is voice input. If "YES" in step S17, that is, if there is voice input, in step S19, the user video data 304c and the user voice data 304d are transmitted to the operator terminal 16, and the process returns to step S5. On the other hand, if "NO" in step S17, that is, if there is no voice input, in step S21, the user video data 304c is transmitted to the operator terminal 16, and the process returns to step S5.
[0168] 11 and 12 are flow diagrams showing the interactive processing of the CPU 50 of the operator terminal 16. Although not shown, the CPU 50 executes, in parallel with the interactive processing, an image capturing process, a process for detecting operation data and voice data, and a process for receiving various data from the user terminal 12. The interactive processing of the CPU 50 will be described below, but the same processing content will be briefly described. At the start of the interactive processing, the operating mode is set to manual mode.
[0169] 11, when the CPU 50 of the operator terminal 16 starts the dialogue processing, in step S101, it establishes a connection with the user terminal 12, and in step S103, it displays the operator's response screen 200 on the display device 60. At this time, the operator terminal 16 has not received the user video data 504f, so the user's video is not displayed in the display area 202.
[0170] In the following step S105, it is determined whether the operation mode is the manual mode by referring to the operation mode data 504i. If the answer is "YES" in step S105, that is, if the operation mode is the manual mode, then in step S107, manual operation data 504e is generated from the operator video data 504c, and the process proceeds to step S111. On the other hand, if the answer is "NO" in step S105, that is, if the operation mode is the automatic mode, then in step S109, automatic operation data 504h is selected according to a predetermined method, and the process proceeds to step S111.
[0171] In step S111, operation data, i.e., manual operation data 504e or automatic operation data 504h, is transmitted to the user terminal 12, and in step S113, it is determined whether voice data has been detected. If "YES" in step S113, that is, if voice data has been detected, then in step S115, operator voice data 504d is transmitted to the user terminal 12, and the process proceeds to step S117 shown in Fig. 12. On the other hand, if "NO" in step S113, that is, if voice data has not been detected, the process proceeds to step S117.
[0172] As shown in Fig. 12, in step S117, it is determined whether data has been received. Here, the CPU 50 determines whether user video data 504f, or user video data 504f and user voice data 504g, have been received. If "NO" in step S117, that is, if data has not been received, the process proceeds to step S125. On the other hand, if "YES" in step S117, that is, if data has been received, it is determined in step S119 whether user voice data 504g is present.
[0173] If the result in step S119 is "YES", that is, if the user voice data 504g is present, then in step S121 the user video data 504f and the user voice data 504g are output, and the process proceeds to step S125. On the other hand, if the result in step S119 is "NO", that is, if the user voice data 504g is not present, then in step S123 the user video data 504f is output, and the process proceeds to step S125.
[0174] In step S125, it is determined whether the dialogue has ended. Here, the CPU 50 determines whether there is an instruction to end the dialogue from the operator, and whether a notification of the dialogue end has been received from the user terminal 12. If the answer is "NO" in step S125, that is, if the dialogue has not ended, the process returns to step S105. On the other hand, if the answer is "YES" in step S125, that is, if the dialogue has ended, the dialogue processing ends. However, if the dialogue is to be ended in accordance with the operator's instruction, the user terminal 12 is notified that the dialogue has ended.
[0175] 13 to 16 are flow charts showing the mode switching process of the CPU 50 of the operator terminal 16. The mode switching process is executed in parallel with the dialogue process shown in FIGS.
[0176] 13, when the CPU 50 starts the mode switching process, it sets the operation mode to the manual mode in step S201, and determines whether the operation mode is the manual mode in step S203. If the determination in step S203 is "NO," that is, if the operation mode is the automatic mode, the process proceeds to step S229 shown in FIG.
[0177] On the other hand, if "YES" in step S203, that is, if the operation mode is the manual mode, it is determined in step S205 whether there is a key input or mouse input. If "YES" in step S205, that is, if there is a key input or mouse input, the input flag 504k is turned on in step S207, and the process proceeds to step S219.
[0178] If step S205 is "NO", that is, if there is no key input or mouse input, face recognition processing is executed in step S207, and it is determined whether face recognition is incorrect in step S211. If step S211 is "YES", that is, if face recognition is incorrect, the process proceeds to step S217.
[0179] On the other hand, if step S211 is "NO," that is, if the face recognition is successful, then in step S213, a direction detection process is executed, and in step S215, it is determined whether the direction of the face has deviated from the front by 60 degrees or more. In this step S215, it is determined whether the direction of the face indicated by the direction data 504j detected in step S213, i.e., the angle, is 60 degrees or more.
[0180] If "NO" in step S215, that is, if the face direction is not deviated from the front by 60 degrees or more, the process proceeds to step S221 shown in Fig. 14. On the other hand, if "YES" in step S215, that is, if the face direction is deviated from the front by 60 degrees or more, the face non-recognition flag 504m is turned on in step S217, the operation mode is set to automatic mode in step S219, and the process proceeds to step S229.
[0181] 14, in step S221, a motion recognition process is executed, and in step S223, it is determined whether a predetermined first motion pattern has been recognized. If "NO" in step S223, that is, if the predetermined first motion pattern has not been recognized, the process returns to step S205 shown in FIG.
[0182] On the other hand, if "YES" in step S223, that is, if the predetermined first movement pattern is recognized, the movement recognition flag 504n is turned on in step S225, the movement mode is set to the automatic mode in step S227, and the process proceeds to step S229.
[0183] As shown in Fig. 15, in step S229, it is determined whether or not the input flag 504k is on. If "NO" in step S229, that is, if the input flag 504k is off, the process proceeds to step S249 shown in Fig. 16. On the other hand, if "YES" in step S229, that is, if the input flag 504k is on, it is determined in step S231 whether or not there is any key input or mouse input.
[0184] If step S231 is "NO", that is, if there is key input and / or mouse input, then in step S233 the no-input time is reset (i.e., set to 0), and the process returns to step S203 shown in Fig. 13. On the other hand, if step S231 is "YES", that is, if there is no key input or mouse input, then in step S235 it is determined whether or not the no-input time is being counted.
[0185] If step S235 is "NO", that is, if the no-input time is not being counted, then in step S237, counting of the no-input time is started, and the process returns to step S203. On the other hand, if step S235 is "YES", that is, if the no-input time is being counted, then in step S239, it is determined whether the no-input time has exceeded a second predetermined time (for example, 30 seconds).
[0186] If step S239 is "NO", that is, if the time without input has not elapsed the second predetermined time, the process returns to step S203. On the other hand, if step S239 is "YES", that is, if the time without input has elapsed the second predetermined time, face recognition processing is executed in step S241, and it is determined in step S243 whether the operator's face has been recognized.
[0187] If "NO" in step S243, that is, if the operator's face has not been recognized, the process returns to step S203. On the other hand, if "YES" in step S243, that is, if the operator's face has been recognized, the input flag 504k is turned off in step S245, the operation mode is set to the manual mode in step S247, and the process returns to step S203.
[0188] As described above, if the result of step S229 is "NO," then in step S249 shown in Fig. 16, it is determined whether the face non-recognition flag 504m is on. If the result of step S249 is "YES," that is, if the face non-recognition flag 504m is on, then in step S251, a direction detection process is executed, and in step S253, it is determined whether the operator's face is facing forward. However, if it is strictly determined whether the face direction is 0 degrees, it is determined in most cases that the face is not facing forward, so the face is determined to be facing forward even if it is off by about 5 to 10 degrees.
[0189] If "NO" in step S253, that is, if the operator's face is not facing forward, the process returns to step S203. On the other hand, if "YES" in step S253, that is, if the operator's face is facing forward, the face unrecognition flag 504m is turned off in step S255, and the operation mode is set to manual mode in step S257, and the process returns to step S203.
[0190] Also, if "NO" in step S249, that is, if the face non-recognition flag 504m is off, a movement recognition process is executed in step S259, and it is determined in step S261 whether or not a predetermined second movement pattern has been recognized.
[0191] If "NO" in step S261, that is, if the predetermined second movement pattern is not recognized, the process returns to step S203. On the other hand, if "YES" in step S261, that is, if the predetermined second movement pattern is recognized, the movement recognition flag 504n is turned off in step S263, and the process proceeds to step S257.
[0192] According to this embodiment, in a manual mode in which the avatar's movements are manually controlled, if the movements of the operator who is responding to the user satisfy a first condition, the mode is set to an automatic mode in which the avatar's movements are automatically controlled, and if the movements of the operator satisfy a second condition corresponding to the first condition, the mode is set to the manual mode, thereby enabling smooth switching between the manual mode and the automatic mode.
[0193] Furthermore, according to this embodiment, the operation mode is switched based on the action of the operator who is responding to the user, so the manual mode and automatic mode can be switched without the operator being aware of it, and therefore the user will hardly feel any discomfort or annoyance.
[0194] In this embodiment, when the operator starts to serve the user, the operation mode is set to the manual mode, but it can also be set to the automatic mode. In such a case, the operation mode is set to the manual mode when any of the second conditions is satisfied.
[0195] In this embodiment, when the operation mode is set to automatic mode, all of the avatar's movements are automatically generated regardless of the operator's movements. However, some of the avatar's movements may also be automatically generated. For example, when automatic mode is set in response to the detection of key input or mouse input, the avatar's hand and arm movements may be automatically generated, and the facial expressions and movements may be controlled according to the operator's facial expressions and movements. Even in this case, the operator can still use their hands to perform other tasks while interacting with the user.
[0196] Furthermore, in this embodiment, when the operation mode is set to the automatic mode, the operation of the avatar is automatically controlled, but the speech of the avatar may also be automatically controlled. In such a case, for example, a chatbot function is further provided in the operator terminal 16, and the speech of the avatar (i.e., the answer) to the speech of the user (i.e., the question) is determined by the chatbot.
[0197] Furthermore, in this embodiment, three first conditions and corresponding second conditions are set, but only one or two first conditions and corresponding second conditions may be set. Furthermore, one or more other first conditions and corresponding second conditions may be added.
[0198] It should be noted that the order of processing steps in the flow chart shown in this embodiment can be changed if the same results are obtained.
[0199] Furthermore, the various screens, images and specific numerical values given in this embodiment are merely examples and can be changed as needed. [Explanation of symbols]
[0200] 10. Information Processing Systems 12...User terminal 14...Network 16...Operator terminal 20, 50...CPU 22, 52...RAM 24, 54...Communication I / F 26, 56... Input / output I / F 28, 58...input device 30, 60…display device 32, 62...Mike 34, 64...speakers 36, 66...camera
Claims
1. An information processing device used by an operator who responds to a user, a camera for photographing the operator; an avatar movement control means for controlling, in a manual mode, the movement of an avatar of the operator displayed on a display device of a user terminal used by the user, based on an image captured by the camera; and An information processing device comprising: a mode switching means for switching from the manual mode to an automatic mode that automatically controls the movement of the avatar when the movement of the operator satisfies a first condition based on an image captured by the camera; and for switching from the automatic mode to the manual mode when the movement of the operator satisfies a second condition corresponding to the first condition in the automatic mode.
2. 2. The information processing device according to claim 1, wherein the first condition is that the operator has performed an action other than an action of responding to the user, and the second condition is that the operator has returned to an action or state of responding to the user.
3. The information processing apparatus according to claim 2 , further comprising a plurality of said first conditions and a plurality of second conditions corresponding to each of said first conditions.
4. the avatar movement control means generates movement data for controlling the movement of the avatar based on the image captured by the camera; The information processing apparatus according to claim 1 , further comprising a transmitting means for transmitting the motion data to the user terminal.
5. the avatar movement control means generates the avatar image reflecting the movement of the operator based on the image captured by the camera; The information processing apparatus according to claim 1 , further comprising a transmitting unit configured to transmit the avatar image to the user terminal.
6. An information processing program executed on an information processing device that is used by an operator who responds to a user and has a camera that photographs the operator, The processor of the information processing device an avatar movement control step of controlling, in a manual mode, a movement of an avatar of the operator displayed on a display device of a user terminal used by the user based on an image captured by the camera; and an information processing program that executes a mode switching step of switching from the manual mode to an automatic mode that automatically controls the movement of the avatar when the movement of the operator satisfies a first condition based on an image captured by the camera, and switching from the automatic mode to the manual mode when the movement of the operator satisfies a second condition corresponding to the first condition in the automatic mode.
7. 1. An information processing method for an information processing device that is used by an operator who responds to a user and has a camera that photographs the operator, The processor of the information processing device In a manual mode, the operation of an avatar of the operator displayed on a display device of a user terminal used by the user is controlled based on an image captured by the camera; An information processing method, comprising: when the operator's movement satisfies a first condition based on an image captured by the camera, switching from the manual mode to an automatic mode that automatically controls the movement of the avatar; and when the operator's movement in the automatic mode satisfies a second condition corresponding to the first condition, switching from the automatic mode to the manual mode.
Citation Information
Patent Citations
Communication system, reception terminal device, and program thereof
JP2021056940A