Information processing system, information processing device, information processing program, and information processing method

The system facilitates natural interactions among multiple AI avatars and a user by managing coordinated dialogue through a central control unit and dialogue control units, addressing limitations in existing single-role avatar systems.

JP2026045783APending Publication Date: 2026-03-13AVITA INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing information processing systems with AI avatars are limited to single-role interactions, preventing multiple avatars and a user from engaging in discussions or meetings.

Method used

A system comprising a central control unit and multiple dialogue control units that manage interactions among multiple AI avatars and a user, using a large language model to generate and select AI utterances based on user input, ensuring only one avatar speaks at a time, and employing methods like random selection, priority selection, and goal-oriented dialogue management.

Benefits of technology

Enables multiple AI avatars and a user to interact naturally, simulating human-like discussions or meetings by allowing coordinated dialogue and facilitating user interaction skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026045783000001_ABST
    Figure 2026045783000001_ABST
Patent Text Reader

Abstract

This invention provides an information processing system, information processing device, method, and program for engaging in discussions or meetings with multiple AI avatars, similar to those between humans. [Solution] The information processing system (10) comprises a central control unit (12), a plurality of dialogue control units (16), and a user terminal (18). In a role-playing discussion or meeting, when a user speaks, the user terminal transmits the user's utterance audio data to the plurality of dialogue control units via the central control unit. Each of the plurality of dialogue control units inputs the user's utterance text data into a large-scale language model (16a), and each obtains AI utterance text data of multiple AI avatars with different personalities or values ​​from the large-scale language model and transmits it to the central control unit. The central control unit converts one AI utterance text data selected in a predetermined manner into audio data and transmits it to the user terminal. The user terminal outputs the AI's utterance as audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an information processing system, an information processing apparatus, an information processing program, and an information processing method, and particularly, for example, to an information processing system, an information processing apparatus, an information processing program, and an information processing method in which an avatar (hereinafter referred to as an "AI avatar") that operates by artificial intelligence interacts with a user.

Background Art

[0002] An example of an information processing system in the background art is disclosed in Patent Document 1. In the conversation control device of Patent Document 1, two avatars with different roles (subordinate, beginner, non-expert, superior, senior, expert) are displayed, and one avatar given a role such as a superior, senior, or expert mainly answers the user's questions, and the other avatar given a role such as a subordinate, beginner, or non-expert supports the user's understanding by supplementing or echoing the explanations by the one avatar.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the above background art, the two avatars only speak according to their respective roles, and it is not possible to construct a system in which a plurality of avatars and a user discuss or hold a meeting.

[0005] Therefore, the main object of this invention is to provide a novel information processing system, information processing apparatus, information processing program, and information processing method.

[0006] Another object of this invention is to provide an information processing system, information processing device, information processing program, and information processing method that enable multiple AI avatars and a user to interact with each other as if they were human beings. [Means for solving the problem]

[0007] The first invention is an information processing system comprising a central control unit that controls the interaction between a user and multiple AI avatars operated by multiple artificial intelligences with different personalities or values, a user terminal used by the user, and multiple dialogue control units that control each of the multiple AI avatars interacting with the user, wherein the central control unit comprises user utterance transmission means that transmits dialogue history data and user utterance data to the multiple dialogue control units when it receives user utterance data from the user terminal, AI utterance receiving means that receives AI utterance data which is the utterance data of each AI avatar from the multiple dialogue control units in response to the user utterance data, selection means that selects one AI utterance data from the multiple AI utterance data according to a predetermined method, and AI utterance transmission means that transmits the one AI utterance data selected by the selection means to the user terminal.

[0008] The second invention is subordinate to the first invention, and the prescribed method is a method for selecting one AI speech data using a random number, using the probability of speech set for each of a plurality of AI avatars.

[0009] The third invention is subordinate to the first invention, and the prescribed method is a method for selecting one AI speech data to be preferentially spoken in response to a specific utterance by a user.

[0010] The fourth invention is subordinate to the first invention, and the prescribed method selects one AI speech data to be spoken using an AI that determines whether the conversation is progressing toward a dialogue goal set for the user.

[0011] The fifth invention is dependent on any of the first to fourth inventions, wherein each of the plurality of dialogue control devices receives user utterance data transmitted by the user utterance transmission means, generates AI utterance data for the user utterance data, and determines whether to utter utterance content corresponding to the AI ​​utterance data, and the selection means selects one AI utterance data from one or more AI utterance data that the determination result indicates to utter utterance content, according to a predetermined method.

[0012] The sixth invention is an information processing device for controlling a conversation between a user and a plurality of AI avatars with different personalities or values, comprising: user utterance transmission means for transmitting conversation history data and user utterance data to a plurality of dialogue control devices that control each of the plurality of AI avatars that converse with the user when user utterance data is received from a user terminal used by the user; AI utterance receiving means for receiving AI utterance data, which is the utterance data of each AI avatar, from the plurality of dialogue control devices in response to the user utterance data; selection means for selecting one AI utterance data from the plurality of AI utterance data according to a predetermined method; and AI utterance transmission means for transmitting the one AI utterance data selected by the selection means to the user terminal.

[0013] The seventh invention is an information processing program executed by an information processing device that controls a conversation between a user and a plurality of AI avatars with different personalities or values, and causes the processor of the information processing device to execute a user utterance transmission step in which, upon receiving user utterance data from a user terminal used by the user, conversation history data and user utterance data are transmitted to a plurality of conversation control devices that control each of the plurality of AI avatars that converse with the user; an AI utterance reception step in which, in relation to the user utterance data, AI utterance data which is the utterance data of each AI avatar is transmitted from the plurality of conversation control devices; a selection step in which one AI utterance data is selected from the plurality of AI utterance data according to a predetermined method; and an AI utterance transmission step in which the one AI utterance data selected in the selection step is transmitted to the user terminal.

[0014] The eighth invention is an information processing method for an information processing device that controls a conversation between a user and a plurality of AI avatars with different personalities or values, wherein the processor of the information processing device, upon receiving user utterance data from a user terminal used by the user, transmits conversation history data and user utterance data to a plurality of dialogue control devices that control each of the plurality of AI avatars that converse with the user, receives AI utterance data, which is the utterance data of each AI avatar, from the plurality of dialogue control devices in relation to the user utterance data, selects one AI utterance data from the plurality of AI utterance data according to a predetermined method, and transmits the selected one AI utterance data to the user terminal. [Effects of the Invention]

[0015] According to this invention, multiple AI avatars and a user can interact with each other in a manner similar to that between humans.

[0016] The aforementioned objectives, other objectives, features, and advantages of this invention will become even clearer from the detailed description of the embodiments described below with reference to the drawings. [Brief explanation of the drawing]

[0017] [Figure 1] Figure 1 shows an information processing system according to one embodiment of this invention. [Figure 2] Figure 2 is a block diagram showing an example of the electrical configuration of the central control unit. [Figure 3] Figure 3 is a block diagram showing an example of the electrical configuration of a dialogue control device. [Figure 4] Figure 4 is a block diagram showing an example of the electrical configuration of a user terminal. [Figure 5] Figure 5 shows an example of an interactive image displayed on the user terminal's display device. [Figure 6] Figure 6 shows an example of the RAM memory map of the central control unit. [Figure 7] Figure 7 shows an example of the RAM memory map of an interactive control device. [Figure 8]FIG. 8 is a diagram showing an example of a memory map of the RAM of the user terminal. [Figure 9] FIG. 9 is a flowchart showing a part of an example of the dialogue management process of the CPU of the central control device. [Figure 10] FIG. 10 is a flowchart that is another part of an example of the dialogue management process of the CPU of the central control device and follows FIG. 9. [Figure 11] FIG. 11 is a flowchart showing an example of the dialogue control process of the CPU of the dialogue control device. [Figure 12] FIG. 12 is a flowchart showing an example of the dialogue control process of the CPU of the user terminal. [Figure 13] FIG. 13 is a flowchart showing an example of the dialogue image display process of the CPU of the user terminal.

Embodiments for Carrying Out the Invention

[0018] Referring to FIG. 1, an information processing system 10 (hereinafter simply referred to as "system 10"), which is an embodiment of this invention, includes a central control device 12. The central control device 12 is communicably connected to a plurality of dialogue control devices 16, user terminals 18, etc. via a network 14. Also, a large language model 16a is communicably connected to each of the plurality of dialogue control devices 16. However, the large language model 16a may be provided in the dialogue control device 16, or may be communicably connected to the dialogue control device 16 via the network 14.

[0019] This system 10 is used to improve the user's dialogue ability in a situation where multiple people discuss or hold a meeting by allowing the user to interact with a plurality of AI avatars, that is, by allowing the user to perform dialogue training. Furthermore, it is used to improve the facilitation ability.

[0020] The central control unit 12 is an information processing device such as a server that comprehensively controls the entire system 10. The dialogue control device 16 is an information processing device such as a personal computer that controls the AI ​​avatar and automatically interacts with the user. The user terminal 18 is an information processing device such as a personal computer that is used by the user to discuss or hold meetings with multiple AI avatars.

[0021] Furthermore, the central control unit 12 and each of the multiple dialogue control units 16, as well as the central control unit 12 and the user terminal 18, can communicate bidirectionally using the P2P (Peer to Peer) method via the network 14.

[0022] Network 14 consists of an IP network (or IP network) including the Internet, and an access network (or access network) for accessing this IP network. The access network can include public telephone networks, mobile phone networks, wired LANs, wireless LANs, CATV (Cable Television), etc. Furthermore, WebRTC (Web Real-Time Communication) is used to realize bidirectional communication using the P2P method. For this purpose, appropriate elements such as a signaling server, STUN server, and TURN server (not shown) are provided, but since these are publicly known, a detailed explanation is omitted.

[0023] Figure 2 is a block diagram showing an example of the electrical configuration of the central control unit 12. As shown in Figure 2, the central control unit 12 includes a CPU 20. The CPU 20 is connected to RAM 22, a communication interface (hereinafter referred to as "communication I / F") 24, and an input / output interface (hereinafter referred to as "input / output I / F") 26 via an internal bus.

[0024] The CPU 20 is a processor that controls the central control unit 12 and, by extension, the entire system 10. However, instead of the CPU 20, a System-on-a-Chip (SoC) that includes multiple functions such as CPU functionality and GPU (Graphics Processing Unit) functionality may be provided.

[0025] RAM22 is the main memory of the central control unit 12 and is used as the work area or buffer area of ​​the CPU 20. Although not shown in the diagram, the central control unit 12 is also provided with an HDD and ROM as auxiliary storage devices. However, non-volatile memory such as an SSD may be used instead of the HDD, or in addition to the HDD. Various software such as the operating system, middleware, and various application programs are stored in the auxiliary storage devices.

[0026] The communication interface 24 is a wired interface that, under the control of the CPU 20, transmits and receives control signals and data between each of the multiple interaction control devices 16 and external computers such as user terminals 18 via the network 14. However, a wireless interface for connecting to a wireless LAN can also be used as the communication interface 24.

[0027] Input devices 28 and display devices 30 are connected to the input / output interface 26. The input devices 28 include keyboards, computer mice, and touch panels. The display device 30 is, for example, a liquid crystal display.

[0028] The input / output interface 26 outputs operation data (or operation information) received from the input device 28 to the CPU 20. The input / output interface 26 also outputs image data generated by the CPU 20 to the display device 30, causing the display device 30 to display a screen or image corresponding to the image data. Note that the configuration of the central control unit 12 shown in Figure 2 is an example and is not limited to this configuration.

[0029] Figure 3 is a block diagram showing an example of the electrical configuration of the dialogue control device 16. As shown in Figure 3, the dialogue control device 16 includes a CPU 40. The CPU 40 is connected to the RAM 42, communication I / F 44, and input / output I / F 46 via an internal bus.

[0030] The CPU 40 is a processor that oversees the overall control of the interaction control device 16. However, instead of the CPU 40, an SoC (System on a Chip) that includes multiple functions such as CPU and GPU functions may be provided.

[0031] RAM 42 is the main memory of the dialogue control device 16 and is used as the work area or buffer area of ​​the CPU 40. Although not shown in the diagram, the dialogue control device 16 is also provided with an HDD and ROM as auxiliary storage devices. However, non-volatile memory such as an SSD may be used instead of the HDD, or in addition to the HDD. Various software such as the operating system, middleware, and various application programs are stored in the auxiliary storage devices.

[0032] The communication interface 44 is a wired interface for sending and receiving control signals and data to and from an external computer, such as the central control unit 12, via the network 14 under the control of the CPU 40. However, a wireless interface for connecting to a wireless LAN can also be used as the communication interface 44.

[0033] Input devices 48 and display devices 50 are connected to the input / output interface 46. The input devices 48 include keyboards, computer mice, and touch panels. The display device 50 is, for example, a liquid crystal display.

[0034] The input / output interface 46 outputs operation data (or operation information) received from the input device 48 to the CPU 40. The input / output interface 46 also outputs image data generated by the CPU 40 to the display device 50, causing the display device 50 to display a screen corresponding to the image data. Note that the configuration of the interactive control device 16 shown in Figure 3 is an example and is not limited to this configuration.

[0035] Figure 4 is a block diagram showing an example of the electrical configuration of the user terminal 18 shown in Figure 1. As shown in Figure 4, the user terminal 18 includes a CPU 60. The CPU 60 is connected to the RAM 62, communication I / F 64, and input / output I / F 66 via an internal bus.

[0036] The CPU 60 is a processor that controls the overall operation of the user terminal 18. However, instead of the CPU 60, an SoC (System on a Chip) that includes multiple functions such as CPU and GPU functions may be provided.

[0037] RAM62 is the main memory of the user terminal 18 and is used as the work area or buffer area of ​​the CPU 60. Although not shown in the diagram, the user terminal 18 is also equipped with an HDD and ROM as auxiliary storage devices. However, non-volatile memory such as an SSD may be used instead of the HDD, or in addition to the HDD. Various software such as the operating system, middleware, and various application programs are stored in this auxiliary storage device.

[0038] The communication interface 64 is a wired interface for sending and receiving control signals and data to and from an external computer, such as the central control unit 12, via the network 14 under the control of the CPU 60. However, a wireless interface for connecting to a wireless LAN can also be used as the communication interface 64.

[0039] The input / output interface 66 is connected to input devices 68, display devices 70, microphones 72, speakers 74, and cameras 76, among others. Input devices 68 include keyboards, computer mice, and touch panels. Display devices 70 are, for example, liquid crystal displays.

[0040] The input / output interface 66 outputs operation data (or operation information) received from the input device 68 to the CPU 60. The input / output interface 66 also outputs image data generated by the CPU 60 to the display device 70, causing the display device 70 to display a screen or image corresponding to the image data. However, image data received from an external computer (for example, the central control unit 12) may also be output by the CPU 60.

[0041] Furthermore, the input / output interface 66 converts the user's voice detected by the microphone 72 into digital audio data and outputs it to the CPU 60, and converts the audio data output by the CPU 60 into an analog audio signal and outputs it from the speaker 74.

[0042] Furthermore, the input / output interface 66 outputs data of an image (i.e., a captured image) including the user captured by the camera 76 to the CPU 60. The audio data output from the CPU 60 includes the speech data (audio data) of the AI ​​avatar received from the central control unit 12. The image of the user captured by the camera 76 may be a still image or a moving image. Note that the configuration of the user terminal 18 shown in Figure 4 is just an example and is not limited to this configuration.

[0043] In such a system 10, a user using a user terminal 18 can interact with multiple AI avatars controlled by each of the multiple dialogue control devices 16. In other words, the user engages in dialogue training in situations involving discussion or meetings with multiple people by discussing or holding a meeting in a group that includes multiple AI avatars. In the discussion or meeting dialogue training, the user may be presented with only the agenda, or with the agenda and its objectives. If only the agenda is presented to the user, the discussion or meeting simply takes place. If the agenda and its objectives are presented, the user acts as a facilitator, guiding the discussion or meeting to achieve the objectives.

[0044] However, multiple AI avatars are realized by inputting prompts containing different personas (personalities) into each large-scale language model 16a. In other words, multiple AI avatars with different personalities or values ​​participate in the discussion or meeting. A persona is information about a virtual person used to determine the personality or values ​​(hereinafter simply referred to as "type") of an AI avatar, and includes name, age, gender, personality, occupation, likes, dislikes, linguistic personality, etc.

[0045] However, the prompt also includes instructions in addition to the persona. In this embodiment, the instructions describe how to generate (or decide on) and output a response (utterance content) by referring to the dialogue history in response to the user's utterance, and how to determine whether the AI ​​avatar itself should speak along with the utterance content, and output the result of that determination.

[0046] The user operates the user terminal 18 to begin using the conversational training (or role-playing) service provided by the central control unit 12. For example, the number and type of AI avatars are selected by the user or the administrator of the central control unit 12. In this embodiment, the avatar images are predetermined according to the type of AI avatar, and when the type of AI avatar is selected, the avatar images are also selected at the same time.

[0047] Once the number of AI avatars is selected, the central control unit 12 establishes a connection with each of the same number of dialogue control units 16, and sends a prompt containing a persona corresponding to the selected type of AI avatar to each of the established dialogue control units 16. Upon receiving the prompt, each dialogue control unit 16 inputs it into the large-scale language model 16a connected to it. Subsequently, the user engages in discussion or conference with multiple AI avatars by communicating between each dialogue control unit 16 and the user terminal 18 via the central control unit 12.

[0048] Figure 5 shows an example of a dialogue image 100 displayed on the display device 70 of the user terminal 18 during dialogue training. The dialogue image 100 shown in Figure 5 is an image that displays images of participants taking part in a discussion or meeting. Here, there are three AI avatars and four participants, including the user. Therefore, the dialogue image 100 includes four display areas 102, 104, 106, and 108.

[0049] Display area 102 is the area for displaying the user's image. Display area 104 is the area for displaying the image of the first AI avatar. Display area 106 is the area for displaying the image of the second AI avatar. Display area 108 is the area for displaying the image of the third AI avatar.

[0050] However, in the dialogue image 100, text information that can identify the user and each AI avatar is displayed at the top of display areas 102, 104, 106, and 108. Display area 102 displays "You" as text information representing the user. Display areas 104, 106, and 108 display "First AI Avatar," "Second AI Avatar," and "Third AI Avatar" as text information representing each AI avatar, but in reality, the name listed in the persona is displayed for each AI avatar.

[0051] The user's image is included in the image captured by camera 76. Image data corresponding to each AI avatar's image is generated by the central control unit 12 and transmitted to the user terminal 18. Therefore, the image data for the dialogue image 100 is generated using the captured image data and the image data corresponding to each AI avatar's image.

[0052] In this embodiment, the central control unit 12 pre-stores action data for the speech and viewing actions of each AI avatar. Image data is generated for AI avatars that speak to perform speech actions, and image data is generated for AI avatars that do not speak to perform viewing actions. The generated image data is then transmitted to the user terminal 18. For example, speech actions include moving the head back and forth and making hand gestures, as well as blinking and changes in facial expression. However, when the user terminal 18 speaks, the image of the speaking AI avatar is lip-synced. Viewing actions include slightly moving the head from side to side, as well as blinking and changes in facial expression.

[0053] In this embodiment, the AI ​​avatar performs a speech action when it speaks, and a viewing action otherwise. Therefore, the AI ​​avatar performs a viewing action when the user speaks and when other AI avatars speak.

[0054] Alternatively, instead of the central control unit 12 generating image data of the AI ​​avatar, the central control unit 12 may send operation data of the AI ​​avatar to the user terminal 18, and the user terminal 18 may generate image data of the AI ​​avatar.

[0055] In discussions or meetings, the AI ​​avatar responds to the user's utterances. However, since it would be unnatural for multiple AI avatars to speak simultaneously, in situations where multiple AI avatars are capable of speaking, the system is designed to allow only one selected AI avatar to speak, thereby achieving a human-like system of deferring speech.

[0056] In this embodiment, when a user speaks, the dialogue history is referenced, and each AI avatar, i.e., each large-scale language model 16a participating in the discussion or meeting, generates utterance content and determines whether or not to utter the generated utterance content.

[0057] Specifically, the user terminal 18 detects the user's voice, i.e., utterances, and transmits the detected voice data (hereinafter referred to as "user utterance voice data") to the central control unit 12. Upon receiving the user utterance voice data, the central control unit 12 performs speech recognition on the received user utterance voice data and converts it into text data (hereinafter referred to as "user utterance text data"), and transmits the converted user utterance text data to each of the dialogue control units 16 (hereinafter referred to as "participating dialogue control units 16") that control the AI ​​avatars participating in the discussion or meeting. Subsequently, the central control unit 12 adds the user utterance text data to the dialogue history data. In other words, the central control unit 12 updates the dialogue history data. After updating the dialogue history data, the central control unit 12 transmits the updated dialogue history data to each of the participating dialogue control units 16.

[0058] Each participating dialogue control device 16, upon receiving user utterance text data, inputs the received user utterance text data and dialogue history data to the large-scale language model 16a connected to it. The large-scale language model 16a refers to the dialogue history data and generates text data of the AI ​​avatar's response (utterance) to the user utterance text data, i.e., the user's utterance. It also determines whether the AI ​​avatar will utter and transmits the generated AI avatar utterance text data (hereinafter referred to as "AI utterance text data") and the determination result to the participating dialogue control device 16 to which the large-scale language model 16a is connected. However, the AI ​​utterance text data generated by the large-scale language model 16a and transmitted to the central control device 12 via the dialogue control device 16 also includes identification information to identify the AI ​​avatar that utters the content corresponding to the AI ​​utterance text data.

[0059] Each of the participating dialogue control devices 16 transmits AI utterance text data and judgment results to the central control device 12 upon receiving them. Therefore, the central control device 12 receives (acquires) the AI ​​utterance text data and judgment results of each AI avatar participating in the discussion or meeting. The central control device 12 selects, according to a predetermined method, one AI utterance text data from an AI avatar whose judgment result indicates it will speak, to be transmitted to the user terminal 18. However, if the judgment results for all AI avatars indicate they will not speak, the central control device 12 decides not to transmit any AI utterance text data from any of the AI ​​avatars to the user terminal 18. In other words, it is determined that none of the AI ​​avatars will speak. Furthermore, in this embodiment, the AI ​​utterance text data of an AI avatar determined to speak is converted into the AI ​​avatar's voice data (hereinafter referred to as "AI utterance voice data") and transmitted to the user terminal 18.

[0060] Another method involves assigning a probability of speaking to each AI avatar and then randomly selecting the AI ​​utterance text data for the AI ​​avatar that will speak. In this case, the probability of speaking can be changed according to the number of times each AI avatar has spoken. For example, the probability of speaking can be increased for AI avatars that have spoken infrequently, decreased for AI avatars that have spoken infrequently, or both.

[0061] Another method involves selecting the AI ​​utterance text data of one AI avatar that will be prioritized (or most likely to be prioritized) in response to a specific utterance by the user. Specific utterances include user utterances that include the name set for the AI ​​avatar, and user utterances that include (quote) content spoken by the AI ​​avatar.

[0062] Furthermore, another method involves using an AI (hereinafter referred to as the "scenario determination AI") that determines whether the dialogue is progressing toward the goal set by the user during dialogue training, and then selecting the speech data of one AI avatar to speak. A large-scale language model can be used as the scenario determination AI. This scenario determination AI preferentially (or with a high probability) selects speech from an AI avatar that is close to or approaching the goal set by the user, or preferentially (or with a high probability) selects speech from an AI avatar that is moving away from (or contrary to) the goal set by the user, or speech from an AI avatar that expresses an irrelevant opinion. In the former case, achieving the goal is relatively easy, so the facilitator's skills can be improved at a low level. In the latter case, achieving the goal is relatively difficult, so the facilitator's skills can be improved at a high level. In other words, it is possible to set not only a goal, but also a level, and conduct dialogue training.

[0063] For example, the scenario determination AI is implemented by inputting a prompt into a large-scale language model that contains instructions to select the utterance from each input, i.e., the utterance of the AI ​​avatar, that is most or least relevant to the goal, along with the goal itself, which is the goal. However, unlike the large-scale language model 16a, the large-scale language model used as the scenario determination AI is communicated to the central control unit 12.

[0064] In this embodiment, the central control unit 12 selects speech data of one AI avatar to send to the user terminal 18, converts the selected AI avatar's speech text data into AI avatar's speech audio data, and sends it to the user terminal 18. The user terminal 18 outputs the received AI speech audio data to the speaker 74. The central control unit 12 also adds the selected AI avatar's speech text data to the dialogue history data. In other words, the central control unit 12 updates the dialogue history data. After updating the dialogue history data, the central control unit 12 sends the updated dialogue history data to each of the participating dialogue control units 16.

[0065] In this way, the user speaks, a selected AI avatar responds, and this is repeated to perform dialogue training, i.e., role-playing of discussions and meetings. However, as mentioned above, there may be cases where neither AI avatar responds to the user's speech. Also, if the AI ​​avatar's speech is relatively long, the user may interrupt the AI ​​avatar while it is speaking.

[0066] Figure 6 shows an example of the memory map 200 of the RAM 22 built into the central control unit 12. As shown in Figure 6, the RAM 22 includes a program storage area 202 and a data storage area 204. The program storage area 202 stores the information processing program executed by the central control unit 12 in this embodiment. This information processing program of the central control unit 12 includes a main processing program 202a, a communication program 202b, a dialogue management program 202c, an AI image generation program 202d, a speech recognition program 202e, and a speech conversion program 202f, among others.

[0067] The main processing program 202a is a program for executing the main information processing routine of the central control unit 12 in this embodiment.

[0068] The communication program 202b is a program for communicating (sending and receiving data, etc.) with external devices such as the dialogue control device 16 and the user terminal 18.

[0069] The dialogue management program 202c is a program for managing the dialogue between the user and the AI ​​avatar. It also selects the AI ​​avatar's utterances to send to the user terminal 18 and records the dialogue history of the user's and the AI ​​avatar's utterances (dialogue).

[0070] The AI ​​image generation program 202d is a program that uses image generation data 204f to generate image data (hereinafter referred to as "AI image data") corresponding to the images of each AI avatar participating in a discussion or meeting.

[0071] The speech recognition program 202e is designed to recognize user speech data and transcribe it into text. The transcribed user speech is then added to the dialogue history.

[0072] The voice conversion program 202f is a program for converting the speech data of an AI avatar selected to speak into the voice data of the speaking AI avatar.

[0073] Although not shown in the diagram, the program storage area 202 also stores other programs necessary for executing the information processing program of the central control unit 12.

[0074] The data storage area 204 stores dialogue history data 204a, user utterance voice data 204b, received AI utterance text data 204c, judgment result data 204d, selected AI utterance text data 204e, image generation data 204f, and AI image data 204g, among others.

[0075] Dialogue history data 204a is dialogue history data that records user utterances and AI avatar utterances in a way that allows for speaker identification and in chronological order.

[0076] User utterance voice data 204b is the user's voice utterance data received from the user terminal 18. This user utterance voice data 204b is transmitted to each of the participating dialogue control devices 16, and the data, which has been speech-recognized and transcribed into text, is added to the dialogue history data 204a.

[0077] The received AI speech text data 204c is the speech data of the AI ​​avatar's text received from each of the participating dialogue control devices 16.

[0078] The judgment result data 204d is data relating to the judgment result for each of the AI ​​utterance text data of the AI ​​avatar received from each of the participating dialogue control devices 16.

[0079] The selected AI speech text data 204e is the speech data of the AI ​​avatar that has been selected to speak. This selected AI speech text data 204e is converted into AI speech voice data, which is the voice of the speaking AI avatar, and sent to the user terminal 18.

[0080] Image generation data 204f consists of data such as polygon data, texture data, and motion data used to generate AI image data 204g.

[0081] AI image data 204g is image data for each AI avatar generated according to the AI ​​image generation program 202d.

[0082] Although not shown in the diagram, the data storage area 204 stores other data necessary for the central control unit 12 to perform information processing, and also contains other timers (counters) and flags necessary for performing information processing.

[0083] Figure 7 shows an example of the memory map 300 of the RAM 42 built into each of the dialogue control devices 16. As shown in Figure 7, the RAM 42 includes a program storage area 302 and a data storage area 304. The program storage area 302 stores the information processing program that is executed in each of the dialogue control devices 16 in this embodiment. This information processing program for the dialogue control device 16 includes a main processing program 302a, a communication program 302b, and a speech generation program 302c, among others.

[0084] The main processing program 302a is a program for executing the main routine of information processing of the dialogue control device 16 in this embodiment.

[0085] The communication program 302b is a program for communicating with external devices such as the central control unit 12.

[0086] The speech generation program 302c is a program that inputs dialogue history data 304a and user utterance text data 304b into the large-scale language model 16a, retrieves AI utterance text data and judgment result data of the AI ​​avatar from the large-scale language model 16a, and transmits the retrieved AI utterance text data and judgment result data of the AI ​​avatar to the central control unit 12.

[0087] Although not shown in the diagram, the program storage area 302 also stores other programs necessary for executing the information processing program of the dialogue control device 16.

[0088] Various types of data are stored in the data storage area 304. These types of data include dialogue history data 304a, user utterance text data 304b, AI utterance text data 304c, and judgment result data 304d.

[0089] Dialogue history data 304a is dialogue history data received from the central control unit 12.

[0090] User utterance text data 304b is user utterance data received from the central control unit 12.

[0091] AI speech text data 304c is speech data of the AI ​​avatar's text received from the large-scale language model 16a, with the AI ​​avatar's identification information added to it.

[0092] The judgment result data 304d is data about the judgment result of whether the AI ​​avatar received from the large-scale language model 16a speaks, and the AI ​​avatar's identification information is added to it.

[0093] Although not shown in the diagram, the data storage area 304 stores other data necessary for the dialogue control device 16 to perform information processing, and also contains timers (counters) and flags necessary for performing information processing.

[0094] Figure 8 shows an example of a memory map 400 of the RAM 62 built into the user terminal 18. As shown in Figure 8, the RAM 62 includes a program storage area 402 and a data storage area 404. The program storage area 402 stores the information processing program to be executed in the user terminal 18 of this embodiment. This information processing program for the user terminal 18 includes a main processing program 402a, an operation detection program 402b, a communication program 402c, a shooting program 402d, an image generation program 402e, a display program 402f, a voice detection program 402g, and a voice output program 402h, etc.

[0095] The main processing program 402a is a program for executing the main routine for information processing of the user terminal 18 in this embodiment.

[0096] The operation detection program 402b is a program for detecting operation data 404a input from the input device 68 in accordance with user operations.

[0097] The communication program 402c is a program for communicating with external devices such as the central control unit 12.

[0098] The shooting program 402d is a program for capturing images (still images or moving images) using the camera 76.

[0099] The image generation program 402e is a program for generating image data corresponding to various screens or images using image generation data 404b and the like. When generating image data for the dialogue image 100, the captured image data 404c and AI image data 404f are also used.

[0100] The display program 402f is a program for outputting image data generated according to the image generation program 402e to the display device 70.

[0101] The voice detection program 402g is a program for detecting voice using microphone 72.

[0102] The audio output program 402h is a program for outputting audio from speaker 74.

[0103] Although not shown in the diagram, the program storage area 402 also stores other programs, such as a browser program, which are necessary to execute the information processing program of the user terminal 18.

[0104] Various types of data are stored in the data storage area 404. These types of data include operation data 404a, image generation data 404b, captured image data 404c, user speech voice data 404d, AI speech voice data 404e, and AI image data 404f.

[0105] Operation data 404a is data representing the user's operation status with respect to the input device 68, as detected by the operation detection program 402b.

[0106] Image generation data 404b consists of data such as polygon data and texture data used to generate image data by the image generation program 402e.

[0107] The captured image data 404c is the data of an image captured by the camera 76 according to the shooting program 402d.

[0108] User speech data 404d is data of the user's voice detected by microphone 72.

[0109] AI speech data 404e is speech data of the AI ​​avatar's voice received from the central control unit 12.

[0110] AI image data 404f is image data of each AI avatar received from the central control unit 12.

[0111] Although not shown in the diagram, the data storage area 404 stores other data necessary for the user terminal 18 to perform information processing, and also contains timers (counters) and flags necessary for performing information processing.

[0112] Figures 9 and 10 are flowcharts showing an example of dialogue management processing, which is an example of information processing by the CPU 20 of the central control unit 12. The CPU 20 of the central control unit 12 starts dialogue management processing when the number and types of AI avatars to participate in the discussion or meeting are selected by a user of a user terminal 18 using the dialogue training service or by an administrator of the central control unit 12.

[0113] As shown in Figure 9, when the CPU 20 starts the dialogue management process, in step S1 it establishes a connection state with both the user terminal 18 and the participating dialogue control device 16.

[0114] In the next step, S3, it is determined whether the interaction has ended. Here, the CPU 20 determines whether it has received notification from the user terminal 18 that the interaction has ended. If the answer in step S3 is "YES," that is, if the interaction has ended, the interaction management process is terminated. On the other hand, if the answer in step S3 is "NO," that is, if the interaction has not ended, the process proceeds to step S5.

[0115] In step S5, it is determined whether user utterance voice data 204b has been received. If the answer in step S5 is "NO," that is, if user utterance voice data 204b has not been received, the process returns to step S5.

[0116] On the other hand, if the answer in step S5 is "YES," that is, if user utterance voice data 204b is received, then in step S7, the user utterance voice data 204b is transmitted to each of the participating dialogue control devices 16. In the next step S9, the user utterance voice data 204b is speech-recognized to generate user utterance text data, and in step S11, the user utterance text data is transmitted to each of the participating dialogue control devices 16.

[0117] In the following step S13, AI image data 204g is generated to cause all AI avatars to perform viewing actions, and in step S15, the generated AI image data 204g is sent to the user terminal 18.

[0118] Furthermore, in step S17, the user utterance text data is added to the dialogue history data 20a, and in step S19, the dialogue history data 204a is sent to each of the participating dialogue control devices 16, and the process proceeds to step S21 shown in Figure 10.

[0119] As shown in Figure 10, in step S21, it is determined whether AI utterance text data and judgment result data have been received from each of the participating dialogue control devices 16. In other words, the CPU 20 determines whether there is received AI utterance text data 204c and judgment result data 204d.

[0120] If the answer in step S21 is "NO," that is, if the received AI utterance text data 204c and the judgment result data 204d have not been received, the process returns to step S21. On the other hand, if the answer in step S21 is "YES," that is, if the received AI utterance text data 204c and the judgment result data 204d have been received, the process proceeds to step S23 to determine whether utterances from all AI avatars are unnecessary. Here, the CPU 20 determines whether the judgment result data 204d indicates that no utterances should be made for any of the AI ​​avatars.

[0121] If the answer in step S23 is "YES," that is, if speech from all AI avatars is not needed, return to step S3. On the other hand, if the answer in step S23 is "NO," that is, if speech from all AI avatars is not needed, proceed to step S25.

[0122] In step S25, the AI ​​speech text data of AI avatar 1 that is to speak is selected from the AI ​​speech text data of AI avatars that have been determined to speak, according to a predetermined method. In other words, the selected AI speech text data 204e is determined from the received AI speech text data 204c that has been determined to speak.

[0123] In the next step S27, the selected AI speech text data 204e is converted into AI speech and the AI ​​speech audio data is sent to the user terminal 18. In the following step S29, AI image data 204g is generated that causes the speaking AI avatar to perform a speech action and the other AI avatars to perform a viewing action, and in step S31, the generated AI image data 204g is sent to the user terminal 18.

[0124] Then, in step S33, the selected AI utterance text data 204e is added to the dialogue history data 204a, and in step S35, the dialogue history data 204a is sent to each of the participating dialogue control devices 16, and the process returns to step S3.

[0125] Figure 11 is a flowchart showing an example of the dialogue control processing included in the information processing of the CPU 40 of the dialogue control device 16 shown in Figure 3. The dialogue control processing is executed in each of the participating dialogue control devices 16. When the CPU 40 of each participating dialogue control device 16 receives a dialogue control start instruction from the central control device 12, it starts the dialogue control processing. However, the dialogue control start instruction includes information about the persona set for the AI ​​avatar. Therefore, prior to starting the dialogue control processing, the CPU 40 of each dialogue control device 16 inputs a prompt containing the above-mentioned instruction (command) and persona information to the large-scale language model 16a, causing the large-scale language model 16a (and the dialogue control device 16) to function as an AI avatar. In addition, although not shown in the diagram, the CPU 40 also performs processing to receive control signals or data from other devices in parallel with the dialogue control processing.

[0126] As shown in Figure 11, when the CPU 40 starts the interactive control process, in step S101 it establishes a connection with the central control unit 12.

[0127] In the next step, S103, it is determined whether or not dialogue history data has been received. If the answer in step S103 is "NO," that is, if dialogue history data has not been received, the process proceeds to step S107. On the other hand, if the answer in step S103 is "YES," that is, if dialogue history data has been received, in step S105, the dialogue history data 304a is stored (or updated) in RAM 42, and the process proceeds to step S107.

[0128] In step S107, it is determined whether or not user utterance text data 304b has been received. If the answer in step S107 is "NO," that is, if user utterance text data 304b has not been received, the process proceeds to step S111. On the other hand, if the answer in step S107 is "YES," that is, if user utterance text data 304b has been received, in step S109, the dialogue history data 304a and user utterance text data 304b are input into the large-scale language model 16a, and the process proceeds to step S111.

[0129] In step S111, it is determined whether there is a response from the large-scale language model 16a. In other words, the CPU 40 determines whether it has received the AI ​​utterance text data 304c and the judgment result data 304d. If the answer in step S111 is "NO," that is, if there is no response from the large-scale language model 16a, the process proceeds to step S115.

[0130] On the other hand, if the answer in step S111 is "YES," that is, if there is a response from the large-scale language model 16a, then in step S113, the AI ​​utterance text data 304c and the judgment result data 304d are sent to the central control unit 12, and the process proceeds to step S115.

[0131] In step S115, it is determined whether the interaction has ended. Here, the CPU 40 determines whether there is a notification from the central control unit 12 that the interaction has ended. If the answer in step S115 is "NO", that is, if the interaction has not ended, the process returns to step S103. On the other hand, if the answer in step S115 is "YES", that is, if the interaction has ended, the interaction control process is terminated.

[0132] Figure 12 is a flowchart showing an example of the dialogue control processing included in the information processing of the CPU 60 of the user terminal 18 shown in Figure 4. The CPU 60 of the user terminal 18, in accordance with the user's operation, accesses the dialogue training website provided by the central control unit 12, selects the number and type of AI avatars to participate in the discussion or meeting, and then starts the dialogue control processing. Although not shown in the figure, the CPU 40 also performs processing to receive control signals or data from other devices in parallel with the dialogue control processing.

[0133] As shown in Figure 12, when the CPU 60 starts the interactive control process, in step S201 it establishes a connection with the central control unit 12.

[0134] In the next step, S203, it is determined whether or not there is user utterance input. Here, the CPU 60 determines whether or not it has detected user voice. If the answer in step S203 is "NO," that is, if there is no user utterance input, the process proceeds to step S213.

[0135] On the other hand, if the answer in step S203 is "YES," that is, if there is user utterance input, then in step S205, the user utterance voice data 404d is sent to the central control unit 12, and the process proceeds to step S207.

[0136] In step S207, it is determined whether or not the AI ​​speech voice data 404e has been received. If the answer in step S207 is "YES," that is, if the AI ​​speech voice data 404e has been received, the output of the AI ​​speech voice data 404e is started in step S211, and the process proceeds to step S213. Therefore, if the user speaks while the AI ​​avatar is speaking, as will be described later, when the process returns from step S213 to step S3, the answer in step S3 will be "YES," and the user can interrupt the AI ​​avatar's speech and speak.

[0137] On the other hand, if the answer in step S207 is "NO," that is, if the AI ​​speech voice data 404e has not been received, then in step S209, it is determined whether the user speech voice data 404d has been sent to the central control unit 12 and whether a predetermined time has been waited. Although not shown in the diagram, when the CPU 60 executes the process in step S205, it resets and starts the timer and counts the waiting time.

[0138] If the answer in step S209 is "NO," that is, if the user speech voice data 404d has not been sent to the central control unit 12 and a predetermined time has not been waited, the process returns to step S207. On the other hand, if the answer in step S209 is "YES," that is, if the user speech voice data 404d has been sent to the central control unit 12 and a predetermined time has been waited, the process determines that there is no speech from the AI ​​avatar and proceeds to step S213.

[0139] In step S213, it is determined whether the interaction has ended. Here, the CPU 60 determines whether the user has given an instruction to end the interaction. If the answer in step S213 is "NO", that is, if the interaction has not ended, the process returns to step S203. On the other hand, if the answer in step S213 is "YES", that is, if the interaction has ended, in step S215, the interaction has ended and the interaction control process is terminated by notifying the central control unit 12 of the end of the interaction.

[0140] Figure 13 is a flowchart showing an example of the dialogue image display processing included in the information processing of the CPU 60 of the user terminal 18 shown in Figure 4. When the CPU 60 of the user terminal 18 starts the dialogue control processing, it also starts (executes) this dialogue image display processing in parallel. Therefore, the CPU 40 executes the process of receiving control signals or data from other devices in parallel with the dialogue image display processing. Although not shown in the diagram, the CPU 40 also executes the camera 76's capture processing in parallel with the dialogue image display processing.

[0141] As shown in Figure 13, when the CPU 60 starts the dialogue image display process, in step S251 it determines whether there is an output of AI speech. If the answer in step S251 is "YES", that is, if there is an output of AI speech, in step S253 it generates dialogue image data using the image generation data 404b, captured image data 404c, and AI image data 404f to make the speaking AI avatar lip-sync, and then proceeds to step S257.

[0142] On the other hand, if the answer in step S251 is "NO," that is, if there is no output from the AI ​​utterance, then in step S255, dialogue image data is generated using the image generation data 404b, the captured image data 404c, and the AI ​​image data 404f, and the process proceeds to step S257. Note that since there is no output from the AI ​​utterance, lip-syncing is not performed in step S255.

[0143] In step S257, the dialogue image data generated in step S251 or step S253 is output to the display device 70. In other words, the dialogue image 100 shown in Figure 5 is displayed.

[0144] Then, in step S259, it is determined whether the dialogue has ended. If the answer in step S259 is "NO", the process returns to step S251. On the other hand, if the answer in step S259 is "YES", the dialogue image display process is terminated.

[0145] According to this embodiment, the system outputs the utterance of one AI avatar selected in a predetermined manner from among the utterances of multiple AI avatars generated in response to the user's utterance, allowing the user to discuss or hold a meeting with multiple AI avatars in a manner similar to human interaction.

[0146] Furthermore, according to this embodiment, setting certain goals for users during discussions or meetings can improve their facilitation skills.

[0147] In the above embodiment, the AI ​​avatar and the user spoke using voice, but they may also speak using text.

[0148] Furthermore, in the above embodiment, the large-scale language model generates speech from the AI ​​avatar and determines whether the AI ​​avatar speaks. However, it is not necessary to determine whether the AI ​​avatar speaks. In such a case, the CPU of the central control unit does not perform the processing in step S23, but in step S25 selects one AI speech text data from the AI ​​speech text data of all AI avatars participating in the discussion or meeting, according to a predetermined method.

[0149] Furthermore, while the above-described embodiment explained the case where a CG (Computer Graphics) avatar is used as the AI ​​avatar, it is also possible to use a robot avatar as part or all of multiple AI avatars. In such cases, a robot that also has the functions of the dialogue control device is provided instead of the dialogue control device, and the robot is placed near the user. The robot's speech is generated by a large-scale language model, and it is determined whether the robot should speak or not. As for the robot's actions, speech action data and viewing action data are stored in advance, and robots that can speak perform speech action, while robots that do not speak perform viewing action. However, the robot is equipped with movable parts such as arms, and speech action and viewing action are expressed by the movement of these movable parts. However, for robots that cannot blink or / or change facial expressions, blinking and / or changes in facial expressions are not included in the speech action and viewing action.

[0150] Furthermore, the order in which each step in the flowcharts shown in the above embodiments is changed if the same result can be obtained. Also, the various screens and specific numerical values ​​shown in this specification and drawings are merely examples and can be changed as needed. [Explanation of Symbols]

[0151] 10. Information Processing Systems 12 ... Central Control Unit 14…Network 16 ... Dialogue control device 18 ... User terminal

Claims

1. An information processing system comprising: a central control unit that controls the interaction between a user and multiple avatars operated by multiple artificial intelligences with different personalities or values ​​(hereinafter referred to as "AI avatars"); a user terminal used by the user; and multiple dialogue control units that control each of the multiple AI avatars that interact with the user, The aforementioned central control unit is User utterance transmission means that, upon receiving user utterance data from the user terminal, transmits dialogue history data and the user utterance data to the plurality of dialogue control devices. AI speech receiving means that receives AI speech data, which is the speech data of each AI avatar, from the plurality of dialogue control devices in relation to the user speech data, A selection means for selecting one AI speech data from a plurality of AI speech data according to a predetermined method, and An information processing system comprising an AI speech transmission means for transmitting the AI ​​speech data selected by the selection means to the user terminal.

2. The information processing system according to claim 1, wherein the predetermined method is a method of selecting one AI speech data using a random number based on the probability of each of the plurality of AI avatars that speaks.

3. The information processing system according to claim 1, wherein the predetermined method is a method for selecting one AI speech data to be preferentially spoken in response to a specific utterance by the user.

4. The information processing system according to claim 1, wherein the predetermined method is to select the AI ​​speech data to be spoken by using an AI that determines whether the conversation is progressing toward the dialogue goal set for the user.

5. Each of the plurality of dialogue control devices receives the user utterance data transmitted by the user utterance transmission means, generates the AI ​​utterance data corresponding to the user utterance data, and determines whether to utter the utterance content corresponding to the AI ​​utterance data. The information processing system according to any one of claims 1 to 4, wherein the selection means selects one AI speech data from one or more AI speech data indicating that the judgment result is to utter speech content, in accordance with the predetermined method.

6. An information processing device that controls interaction between a user and multiple AI avatars with different personalities or values, User utterance transmission means that, upon receiving user utterance data from a user terminal used by the user, transmits dialogue history data and the user utterance data to a plurality of dialogue control devices that control each of the plurality of AI avatars that interact with the user. AI speech receiving means that receives AI speech data, which is the speech data of each AI avatar, from the plurality of dialogue control devices in relation to the user speech data, A selection means for selecting one AI speech data from a plurality of AI speech data according to a predetermined method, and An information processing apparatus comprising an AI speech transmission means for transmitting the 1 AI speech data selected by the selection means to the user terminal.

7. An information processing program executed by an information processing device that controls the interaction between a user and multiple AI avatars with different personalities or values, The processor of the aforementioned information processing device, A user utterance transmission step in which, upon receiving user utterance data from a user terminal used by the user, the user utterance data and the dialogue history data are transmitted to a plurality of dialogue control devices that control each of the plurality of AI avatars that interact with the user. AI speech reception step: Receiving AI speech data, which is the speech data of each AI avatar, from the plurality of dialogue control devices with respect to the user speech data. A selection step of selecting one AI speech data from a plurality of AI speech data according to a predetermined method, and An information processing program that causes the program to execute an AI speech transmission step, which transmits the AI ​​speech data selected in the selection step to the user terminal.

8. An information processing method for an information processing device that controls interaction between a user and multiple AI avatars with different personalities or values, The processor of the aforementioned information processing device is When user utterance data is received from the user terminal used by the user, the dialogue history data and the user utterance data are transmitted to a plurality of dialogue control devices that control each of the plurality of AI avatars that interact with the user. With respect to the user utterance data, the plurality of dialogue control devices receive AI utterance data, which is the utterance data of each AI avatar. Select one AI speech data from a plurality of AI speech data according to a predetermined method, An information processing method for transmitting the selected AI speech data 1 to the user terminal.

Citation Information

Patent Citations

  • Conversation control program, conversation control method, and conversation control device

    JP7334800B2