system

The system addresses the challenge of monotonous streams by generating conversational messages for virtual users, ensuring interactive and engaging broadcasts through natural interactions with virtual listeners, thereby improving viewer engagement and streamer popularity.

JP2026041461APending Publication Date: 2026-03-10SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Streamers often face challenges in engaging with few viewers, leading to monotonous streams and difficulty in attracting more viewers due to the lack of effective tools for interactive conversations.

Method used

A system that generates conversational messages for virtual users, analyzes user responses, and generates appropriate next messages using a generative AI model, allowing broadcasters to interact naturally with virtual listeners even when there are few real viewers.

Benefits of technology

Enables interactive and lively broadcasts from the early stages, enhancing viewer engagement and streamer popularity by facilitating natural and realistic conversations with virtual users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026041461000001_ABST
    Figure 2026041461000001_ABST
Patent Text Reader

Abstract

Provide a system. A method for generating conversational messages for virtual users, means for transmitting a conversation message of the virtual user to a user terminal; means for receiving a user response and generating a next conversational message based on the response; means for transmitting the next conversation message to the user terminal again; A system including:
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Before a streamer gains sufficient popularity, they often receive few chat messages from viewers, making it difficult for them to interact with them. As a result, their streams become monotonous, making it difficult for them to attract more viewers. Another problem is the lack of support tools to help streamers smoothly engage in conversations with viewers and increase the appeal of their streams. [Means for solving the problem]

[0005] The present invention solves the above problem by providing a system including a means for generating a conversational message for a virtual user, a means for transmitting the conversational message for the virtual user to a user terminal, a means for receiving a user's response and generating a next conversational message based on the response, and a means for transmitting the next conversational message to the user terminal again. Specifically, this system automatically generates chat messages from a virtual listener, allowing the broadcaster to continue the conversation even when there are few viewers, thereby improving the appeal of the broadcast. Furthermore, the means for analyzing the user's response and generating the next conversational message realizes a more natural and realistic conversation, facilitating smooth interaction with viewers.

[0006] A "virtual user" refers to a simulated user generated by a computer program rather than a real person.

[0007] "Conversation message" refers to text data that represents the content of statements sent and received by users and virtual users through chat.

[0008] "User terminal" refers to communication devices such as computers and smartphones used by broadcasters and viewers.

[0009] "User response" refers to the messages and reactions entered and sent by the broadcaster through the user's device.

[0010] "Means for generating" refers to a program or device that has the function of creating new data or information based on a specific process or algorithm.

[0011] "Transmission means" refers to the communications technology or functionality used to move data or information from one point to another.

[0012] "Means for receiving" refers to a program or device that receives and processes the transmitted data or information.

[0013] "Analysis" refers to the process of examining received data or information in detail to understand its meaning and content. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0022] [First embodiment]

[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0035] The present invention is a listener AI system that enables broadcasters to have natural chats with virtual users. The system operates mainly with a server, terminals, and users.

[0036] System Program

[0037] Server processing

[0038] First, we will explain the processing performed by the server. The server includes a program that generates conversation messages for virtual users. The server starts this program when a broadcaster starts broadcasting, and generates multiple virtual listener accounts based on the settings. This makes it possible to send messages to the chat box immediately after the broadcast starts.

[0039] The server analyzes the responses received from the streamer's user device in real time. During this analysis phase, the streamer's message content is broken down into keywords and the appropriate conversation message to be sent next is generated. For example, if the streamer replies, "Today I'll talk about game strategies," the server analyzes the keyword "game strategies" and generates the next message: "That's fun! Which game?"

[0040] Processing by the terminal

[0041] Next, we will explain the processing performed by the terminal. When broadcasting starts, the broadcaster's terminal sends a start notification to the server. When a conversation message is sent from the server by the virtual listener, the terminal displays it in the chat box.

[0042] The broadcaster checks the chat message through the terminal and enters a response, which is then sent back to the server for analysis.

[0043] User operation

[0044] The streamer reacts to messages from the virtual listeners that appear in the chat box on their device, making the streamer feel as if they are having a conversation with a real viewer. For example, if the initial message from the virtual listener says, "Nice to meet you! What will we talk about today?", the streamer will respond, "Today I'll talk about game strategies."

[0045] In response, the server generates the next conversation message and sends it to the device, allowing the broadcaster to continue the conversation. As the broadcast progresses, even if the number of real viewers increases, the device will appropriately display messages from both virtual listeners and real viewers, allowing the broadcaster to respond to both.

[0046] Specific examples

[0047] When a broadcaster starts broadcasting, the server first generates a virtual listener message, "What are we talking about today?", and sends it to the terminal. When the broadcaster responds, "Today we'll talk about movies," the server analyzes this response and generates and sends the next message, "Which movie?"

[0048] In this way, the server, terminals, and users work together to allow the broadcaster to continue broadcasting while engaging in natural chat with virtual listeners. This system enables interactive broadcasting even at the initial stage when there are only a few viewers, helping broadcasters to increase their popularity.

[0049] The processing flow will be explained below.

[0050] Step 1:

[0051] The server registers the broadcaster's user ID in the database. At this time, it also sets the number of virtual listener accounts and parameters required for message generation as the initial settings for the listener AI.

[0052] Step 2:

[0053] The distributor (user) presses the distribution start button. The distributor's terminal sends this distribution start information to the server.

[0054] Step 3:

[0055] The server receives a notification from the streamer that the stream has started, then launches the listener AI and generates multiple virtual listener accounts based on the streamer's settings.

[0056] Step 4:

[0057] The server uses an AI engine to generate an initial chat message. For example, it generates a message like "Nice to meet you! What will you talk about today?" and sends it to the broadcaster's device.

[0058] Step 5:

[0059] The streamer's device displays the chat message sent from the server in the chat box. The streamer checks this and enters a response such as "Today I'll be talking about game strategies."

[0060] Step 6:

[0061] The distributor's terminal transmits the answer entered by the distributor to the server.

[0062] Step 7:

[0063] The server receives the streamer's response and uses an AI engine to analyze it and extract the keyword "game strategy."

[0064] Step 8:

[0065] The server generates the next conversation message based on the analysis result. For example, it generates the message "That's fun! Which game is it?" and sends it to the broadcaster terminal again.

[0066] Step 9:

[0067] The streamer's device will again display the message sent from the server in the chat box. The streamer will continue to reply to the AI's messages.

[0068] Step 10:

[0069] The process loops from step 6 to step 9, and the conversation with the virtual listener continues until the streamer ends the chat. As the stream progresses, messages from real viewers are also displayed, allowing the streamer to respond to both.

[0070] This allows streamers to have interactive conversations with virtual listeners, deepening their bond with their viewers as they continue streaming.

[0071] Example 1

[0072] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0073] When live streamers have few viewers in the early stages of a stream, they may experience little interaction with viewers, resulting in a lack of excitement. As a result, streamers often struggle to attract viewers' interest, and their streams can feel monotonous until the number of viewers increases. Furthermore, with few chat messages from viewers, streamers may be unsure of how to move the conversation forward. To solve these problems, a system is needed that allows streamers to have natural conversations with virtual viewers and realize interactive streams.

[0074] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0075] In this invention, the server includes: a means for generating conversational messages for virtual users; a means for transmitting the conversational messages for virtual users to a user terminal; a means for receiving a user's response and generating a next conversational message based on the response; a means for transmitting the next conversational message to the user terminal; a means for generating multiple virtual listener accounts based on a broadcast start notification received from the user terminal; a means for analyzing the generated conversational messages in real time and breaking down the broadcaster's input message into keywords; and a means for executing the above-mentioned means using a generative AI model. This allows virtual listener accounts to be generated from the early stages of broadcasting, enabling natural conversations to take place, enabling the broadcaster to always broadcast lively and interactively. Furthermore, by analyzing the broadcaster's messages in real time and generating appropriate conversational messages, smooth conversation progression is ensured even as the number of viewers increases.

[0076] A "virtual user" is not a real person, but a character generated by the system and designed to interact with the user.

[0077] "Conversation Message" means text-based communication content sent by a virtual user or user.

[0078] "User Terminal" means a device used by a broadcaster, such as a computer, tablet, or smartphone, that includes software for sending, receiving, and broadcasting chat messages.

[0079] A "generative AI model" refers to an artificial intelligence algorithm that performs natural language processing, and is a technology for generating interactive messages based on user input messages.

[0080] A "prompt" is an instruction given to a generative AI model that provides it with the information it needs to generate an appropriate conversational message.

[0081] "Analysis" is the process of breaking down a received message to understand or make sense of its content.

[0082] A "distribution start notification" is a signal sent to the server when a distributor starts distribution, and serves as a trigger to start up each part of the system.

[0083] A "virtual listener account" is an account created by the system to act as a virtual user.

[0084] "Real-time" refers to near-instant processing or response, with minimal delay.

[0085] "Keywords" refer to important words or phrases in a user's message, and are the building blocks for generating the next conversational message.

[0086] The present invention provides a system that supports broadcasters in progressing with broadcasts while engaging in natural chat with virtual users. This system operates with a server, a terminal, and a user as its main components. Specific embodiments of this system are described below.

[0087] Server processing

[0088] The server includes a program that generates conversational messages from virtual users. This program uses a generative AI model that performs natural language processing. Specifically, OpenAI's (registered trademark) GPT-4 (registered trademark) can be used. The server launches this program when the broadcaster starts broadcasting, and generates multiple virtual listener accounts based on the settings. This allows messages from virtual users to be streamed in the chat box immediately after the broadcast begins.

[0089] The server also analyzes responses received from the broadcaster's user terminal in real time. This analysis involves breaking down the broadcaster's message content into keywords and generating the appropriate conversational message to be sent next. For example, if the broadcaster replies, "Today, I'll talk about movies," the server analyzes the keyword "movie" and generates the next message, "Which movie are you going to talk about?"

[0090] Processing by the terminal

[0091] When streaming begins, the device sends a start notification to the server. This notification is sent via an HTTP request or WebSocket. When a conversation message is sent from the server by the virtual listener, the device displays it in the chat box. The streamer checks the chat message through the device and enters a response. This response is sent back to the server and used for analysis. The device uses the streaming software OBS Studio or a dedicated streaming app.

[0092] User operation

[0093] The broadcaster reacts to the virtual listener's messages displayed in the chat box on the device, making the broadcaster feel as if they were having a conversation with a real viewer. For example, if the initial virtual listener's message is "Hello! What will you talk about today?", the broadcaster can respond with "I'll talk about movies today." This response is sent to the server, where it is analyzed in real time and the next conversation message is generated.

[0094] Specific examples

[0095] For example, when a broadcaster starts broadcasting, the server first generates a message for the virtual listener asking, "What are we talking about today?" and sends it to the terminal. If the broadcaster responds, "I'll talk about movies today," the server analyzes this response and generates and sends the next message, "Which movie are you talking about?" This allows for a natural chat with the virtual listener.

[0096] Prompt Sentence Examples

[0097] “If the streamer responds, ‘Today we’re going to talk about movies,’ suggest the next conversation message they should generate from their virtual listeners.”

[0098] This system allows streamers to engage with viewers through interactive broadcasts even at the early stages when the number of viewers is still small. As the number of real viewers increases, the system can also provide an environment where streamers can respond smoothly by displaying messages from both virtual listeners and real viewers appropriately.

[0099] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0100] Step 1:

[0101] The terminal sends a notification of the start of distribution to the server.

[0102] Specific operation: When the streamer presses the "Start Streaming" button in the streaming software (e.g., OBS Studio), the device sends a notification to the server via an HTTP request or WebSocket to start streaming.

[0103] Input: "Start distribution" operation of distribution software.

[0104] Output: A notification is sent to the server, and the server goes into a live broadcast state.

[0105] Step 2:

[0106] The server creates multiple virtual listener accounts.

[0107] Specific operation: The server uses the generative AI model to generate profile information (name, icon, etc.) for the virtual listener based on the settings and creates a virtual listener account.

[0108] Input: Notification from device that distribution has started.

[0109] Output: Multiple virtual listener accounts are created on the server.

[0110] Step 3:

[0111] The server generates an initial message.

[0112] What it does: Uses a generative AI model to generate an initial message from a virtual listener (e.g., "Hello! What will we talk about today?").

[0113] Input: Virtual listener account information.

[0114] Output: Initial message from the virtual listener.

[0115] Step 4:

[0116] The server sends the initial message for the virtual listener to the terminal.

[0117] Specific operation: The server generates an initial message and sends it to the terminal via a WebSocket or HTTP request.

[0118] Input: The initial message generated.

[0119] Output: The initial message of the virtual listener is sent to the terminal.

[0120] Step 5:

[0121] Displays messages received by the terminal in the chat box.

[0122] Specific operation: Displays the virtual listener's message in the chat box of the distribution software.

[0123] Input: The initial message sent by the server.

[0124] Output: The message that appears in the chat box.

[0125] Step 6:

[0126] The user (broadcaster) responds to the virtual listener's messages in the chat box.

[0127] Specific operation: The streamer types messages through the chat box and interacts with the virtual listeners. For example, the streamer types, "Today I'll talk about movies."

[0128] Input: The virtual listener's message and the broadcaster's response.

[0129] Output: The streamer's input message.

[0130] Step 7:

[0131] The device sends the distributor's response to the server.

[0132] Specific operation: Sends the broadcaster's input message to the server in real time.

[0133] Input: The streamer's input message.

[0134] Output: The publisher's response sent to the server.

[0135] Step 8:

[0136] The server parses the publisher's response.

[0137] How it works: The server uses the generative AI model to analyze the broadcaster's message by keyword and generate the next appropriate conversational message. For example, it analyzes the keyword "movie" and generates the message "Which movie are you talking about?"

[0138] Input: The broadcaster's response message.

[0139] Output: Analyzed keywords and the next conversation message.

[0140] Step 9:

[0141] The server generates the following conversation message:

[0142] Specific operation: Based on the analysis results, the generative AI model is used to generate the next conversation message.

[0143] Input: Parsed keywords.

[0144] Output: Next conversation message.

[0145] Step 10:

[0146] The server generates a message and sends it to the terminal.

[0147] Specific operation: The next generated conversation message is sent to the device via WebSocket or HTTP request.

[0148] Input: The next generated conversation message.

[0149] Output: The next conversation message sent to the terminal.

[0150] Step 11:

[0151] Displays messages received by the terminal in the chat box.

[0152] Specific operation: The following conversation message will be displayed in the chat box of the distribution software.

[0153] Input: The next conversation message sent from the server.

[0154] Output: The message that appears in the chat box.

[0155] The above processing steps allow the broadcaster to continue broadcasting while engaging in natural chat with the virtual listeners. This allows for interactive broadcasting even when the number of viewers is small from the beginning, contributing to the broadcaster's popularity.

[0156] (Application example 1)

[0157] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0158] In live streaming and content distribution, streamers can feel isolated in the early stages when there are few viewers. This can make interactive streaming difficult, leading to a decline in streamer motivation and a decline in viewer interest. Furthermore, there is a lack of effective means to properly manage messages for both real viewers and virtual listeners.

[0159] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0160] In this invention, the server includes means for generating conversational messages of a virtual user, means for transmitting the conversational messages of the virtual user to a terminal device, means for receiving responses from the user and generating the next conversational message based on the responses, means for analyzing the response messages and generating natural conversation based on prompt sentences using a generative AI model, and means for appropriately processing and displaying the messages of the virtual user and the messages of the actual viewers. This makes it easier for the broadcaster to maintain interactive conversations even in the early stages when the number of viewers is small, thereby improving the quality of the broadcast.

[0161] "Virtual user conversation messages" refers to messages of users virtually created using a generative AI model.

[0162] "Terminal device" refers to a device used by a broadcaster, including a smartphone or head-mounted display.

[0163] "User response" refers to a reply message entered by a broadcaster in response to a conversation message from a virtual user.

[0164] "Generative AI model" refers to an algorithm or system that uses generative AI technology to generate natural-sounding conversations.

[0165] A "prompt sentence" is the base text input to a generative AI model that determines the content of the next conversational message generated.

[0166] "Natural conversation" refers to messages generated by a generative AI model based on prompts that feel like real human conversations.

[0167] "Messages and actual viewer messages" refers to messages generated by virtual users or AI, as well as messages entered by viewers in real time.

[0168] "Means for appropriate processing and display" refers to the technology and algorithms that distinguish between messages from virtual users and messages from actual viewers and display them in a way that is easy for users to understand.

[0169] The present invention provides a system for generating conversational messages of virtual users and enabling broadcasters to broadcast interactive live content. The system operates using a server, a terminal device, and a broadcaster as its main components.

[0170] Server Processing

[0171] The server includes a program that generates conversation messages for virtual users. When a broadcast starts, the server starts this program and generates multiple accounts for the virtual users. As a result, messages from the virtual users are automatically displayed in the chat box as soon as the broadcast starts.

[0172] The server has a response message receiving function that analyzes the streamer's messages in real time. It breaks down the message entered by the streamer into keywords and generates the next conversation message based on these keywords. The generated message is input into a generative AI model using prompt sentences to create a natural conversation. For example, if the streamer enters "Today we will talk about movies," the server analyzes the keyword "movie" and generates the next message, "Which movie?"

[0173] Terminal processing

[0174] The terminal device is a device used by the broadcaster, such as a smartphone or head-mounted display. When broadcasting begins, the terminal device sends a start notification to the server. When a message from the virtual user is sent from the server, the terminal device has the function of displaying the message directly in the chat box.

[0175] The terminal device also has the role of transmitting the response message entered by the distributor to the server, which allows the server to analyze and generate the response message.

[0176] User operation

[0177] The broadcaster inputs responses to the virtual user's messages via a chat box displayed on the terminal device. These responses are sent to the server, where they are analyzed and used to generate the next conversation message. This system allows the broadcaster to continuously engage in natural interactions with the virtual user and the actual viewers.

[0178] Specific examples

[0179] When a broadcaster starts broadcasting, the server first generates a message from the virtual user saying, "What's the topic today?" and sends it to the terminal device. When the broadcaster responds, "Today I'll introduce a cooking recipe," the server analyzes this response, generates the next message, "That's great! What kind of cooking recipe is it?" and sends it again to the terminal device.

[0180] Prompt Sentence Examples

[0181] Streamer: Today I'll show you a recipe

[0182] Virtual users:

[0183] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0184] Step 1:

[0185] The server receives the notification that distribution has started.

[0186] Input: Notification of start of broadcasting from the broadcaster's terminal device.

[0187] Operation: The server receives the notification of the start of distribution and starts a program that generates conversational messages for the virtual user.

[0188] Output: Launch of the virtual user conversation message generation program.

[0189] Step 2:

[0190] The server generates an initial conversation message for the virtual user.

[0191] Input: Notification of start of distribution.

[0192] How it works: The server uses a generative AI model to generate an initial conversation message based on a configured prompt, such as "What's the topic today?"

[0193] Output: Initial virtual user conversation message.

[0194] Step 3:

[0195] The server sends an initial conversation message to the terminal device of the virtual user.

[0196] Input: The initial conversation message of the generated virtual user.

[0197] Operation: The server sends the generated initial conversation message to the broadcaster's terminal device.

[0198] Output: The initial conversation message of the virtual user displayed on the broadcaster's terminal device.

[0199] Step 4:

[0200] The broadcaster responds to the initial conversation message.

[0201] Input: The initial conversation message of the virtual user.

[0202] How it works: The broadcaster responds to the initial conversation message through the chat box on the terminal device, for example, typing, "Today I'll introduce you to a cooking recipe."

[0203] Output: The distributor's response message.

[0204] Step 5:

[0205] The terminal device sends the distributor's response message to the server.

[0206] Input: The response message entered by the broadcaster.

[0207] Operation: The terminal device sends the message entered by the broadcaster to the server.

[0208] Output: The publisher's response message sent to the server.

[0209] Step 6:

[0210] The server analyzes the distributor's response message.

[0211] Input: The broadcaster's response message.

[0212] How it works: The server breaks down the broadcaster's response message into keywords and creates a prompt using a generative AI model.

[0213] Output: Analysis results and prompt for the next conversation message.

[0214] Step 7:

[0215] The server generates the following conversation message:

[0216] Input: Prompt statement and analysis result.

[0217] How it works: The server uses a generative AI model to generate the next conversational message based on the prompt, for example, "That's great! What's the recipe?"

[0218] Output: The following conversation message is generated:

[0219] Step 8:

[0220] The server sends the following conversation message to the terminal device:

[0221] Input: The next generated conversation message.

[0222] Operation: The server again sends the next conversation message that has been generated to the broadcaster's terminal device.

[0223] Output: The next conversation message displayed on the terminal.

[0224] The above processing steps enable the broadcaster to broadcast live while maintaining natural interaction with the virtual user.

[0225] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0226] This invention is a system that combines a listener AI system that enables broadcasters to chat naturally with virtual users with an emotion engine that recognizes the user's emotions. This system operates with a server, a terminal, and a user as its main components, and the emotion engine can generate appropriate conversational messages according to the user's emotional state.

[0227] System Program

[0228] Server processing

[0229] First, we will explain the processing performed by the server. The server includes a program that generates conversation messages for virtual users. When a broadcaster starts broadcasting, this program is launched and multiple virtual listener accounts are generated based on the settings. This makes it possible to send messages to the chat box immediately after the broadcast begins.

[0230] The server also has an emotion engine built in that recognizes the emotional state of the streamer when analyzing their response. During this analysis phase, the content of the streamer's message is broken down into keywords, and the emotion engine determines their emotional state (e.g., happy, sad, or angry). The next appropriate conversational message to be sent is generated based on this emotional state. For example, if the streamer replies, "I'm not feeling very well today," the server analyzes the keyword "not feeling well" and their emotional state, and generates the message, "That sounds worrying. Are you okay?"

[0231] Processing by the terminal

[0232] Next, we will explain the processing performed by the terminal. When broadcasting starts, the broadcaster's terminal sends a start notification to the server. When a conversation message is sent from the server by the virtual listener, the terminal displays it in the chat box.

[0233] The streamer checks the chat message through their device and enters a response, which is then sent back to the server and used for analysis by the emotion engine.

[0234] User operation

[0235] The streamer reacts to messages from virtual listeners that appear in the chat box on their device, making the streamer feel as if they are having a conversation with a real viewer. For example, if an initial message from a virtual listener says, "Nice to meet you! What will you talk about today?", the streamer might respond, "I'm not feeling too great today, but I'll talk about movies."

[0236] In response, the server generates the next conversation message and sends it to the device, allowing the streamer to continue the conversation. This system allows streamers to deepen their interaction with viewers while having natural conversations with virtual listeners.

[0237] Specific examples

[0238] When a broadcaster starts broadcasting, the server first generates a message for the virtual listener asking, "What's the topic today?" and sends it to the terminal. When the broadcaster responds, "Today we'll talk about movies," the server analyzes this response with its emotion engine and determines the user's emotional state as "excitement" or "joy." It then generates the next message, "That's exciting! Which movie?" and sends it again to the broadcaster's terminal.

[0239] If the streamer responds, "I'm feeling down today, so I watched a quiet movie," the server analyzes the keyword "depressed" and the emotional state and generates a message saying, "That's terrible. What movie did you watch?"

[0240] In this way, the server, device, user, and emotion engine work together to allow the streamer to have natural, emotionally appropriate conversations with virtual listeners as the stream progresses. This system enables interactive streaming even at the initial stage, when there are only a few viewers, and helps streamers increase their popularity.

[0241] The processing flow will be explained below.

[0242] Step 1:

[0243] The server registers the broadcaster's user ID in the database. At this time, the server sets up the initial settings for the listener AI, including the number of virtual listener accounts, parameters required for message generation, and the emotion engine.

[0244] Step 2:

[0245] The distributor (user) presses the distribution start button. The distributor's terminal sends this distribution start information to the server.

[0246] Step 3:

[0247] The server receives a notification from the streamer that the stream has started, then launches the listener AI and generates multiple virtual listener accounts based on the streamer's settings.

[0248] Step 4:

[0249] The server uses an AI engine to generate an initial chat message. For example, it generates a message like "Nice to meet you! What will you talk about today?" and sends it to the broadcaster's device.

[0250] Step 5:

[0251] The streamer's device displays the chat message sent from the server in the chat box. The streamer checks this and enters a response such as "Today I'll be talking about game strategies."

[0252] Step 6:

[0253] The distributor's terminal transmits the answer entered by the distributor to the server.

[0254] Step 7:

[0255] The server receives the publisher's response, which is analyzed by the emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger).

[0256] Step 8:

[0257] The server generates the next conversation message based on the analysis results of the emotion engine. For example, if the broadcaster replies, "I'm a little sad today, but I'll talk about movies," the server recognizes the emotional state as "sad" and generates the message, "That's terrible. What movie did you see?" The generated message is then sent back to the broadcaster's device.

[0258] Step 9:

[0259] The streamer's device will again display the message sent from the server in the chat box. The streamer will continue to reply to the AI's messages.

[0260] Step 10:

[0261] The process loops from step 6 to step 9, and the conversation with the virtual listener continues until the streamer ends the chat. As the stream progresses, messages from real viewers are also displayed, allowing the streamer to respond to both.

[0262] In this way, the server, device, user, and emotion engine work together to allow the streamer to have natural, emotionally appropriate conversations with virtual listeners as the stream progresses. This system enables interactive streaming even at the initial stage, when there are only a few viewers, and helps streamers increase their popularity.

[0263] Example 2

[0264] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0265] In modern live streaming, a lack of interaction in the early stages when viewers are few can cause streamers to lose motivation. Furthermore, in order for streamers to enjoy natural conversations with viewers, they need to respond in real time, which places a heavy burden on them. A system that can solve these problems and enable streamers to enjoy natural conversations even with a small audience is needed.

[0266] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0267] In this invention, the server includes means for generating a conversational message of a virtual user, means for transmitting the conversational message of the virtual user to an information processing device, means for receiving a response from the broadcaster and analyzing the emotional state of the response, means for generating a next conversational message based on the emotional state, and means for transmitting the next conversational message again to the information processing device. This allows the broadcaster to enjoy natural conversation with the virtual listener even when there are only a few viewers.

[0268] A "virtual user" refers to an artificial participant that does not exist in reality but is generated by a computer program and behaves like a real user.

[0269] "Conversational messages" are messages sent and received in a chat or dialogue format, and refer to sentences or text exchanges generated by users or virtual users.

[0270] An "information processing device" is an electronic device used for calculations, data processing, communication, etc., and includes, for example, computers, smartphones, tablets, etc.

[0271] "Distributor" means an individual or entity that distributes live streaming or real-time video content over the Internet.

[0272] "Emotional state" refers to the emotional state of a person, such as joy, sadness, or anger, that is analyzed from text or audio.

[0273] "Emotion engine" refers to a software module that analyzes user input data and determines the emotions contained in the input.

[0274] "Generating means" refers to technical mechanisms or algorithms that automatically generate data or messages for a specific purpose.

[0275] "Transmitting means" refers to the technical mechanism by which data or messages are transferred from one information processing device to another device or server.

[0276] "Means of receiving" refers to the technical mechanisms for acquiring and processing data or messages sent from outside.

[0277] "Means of analysis" refers to the technical mechanisms used to analyze received data or messages and understand their content and meaning.

[0278] This invention is a listener AI system that enables broadcasters to chat naturally with virtual users, and is a system that combines an emotion engine that recognizes the user's emotions. The main components of this system are a server, a terminal, and a user.

[0279] Server Roles

[0280] The server first generates conversational messages for virtual users. When a broadcaster starts broadcasting, the server creates a virtual listener account and sends an initial message based on the broadcaster's settings to the broadcaster's device. The server is equipped with an emotion engine that analyzes the broadcaster's responses in real time and determines their emotional state.

[0281] As a specific example, if the broadcaster responds, "Today we'll talk about movies," the server analyzes this message with an emotion engine, generates the next conversational message, "That's exciting! Which movie?", and sends it to the device. This emotion engine can use, for example, a natural language processing library implemented in Python or a machine learning model (e.g., BERT).

[0282] Device Role

[0283] The streamer's device sends a notification to the server when the stream starts, and the streamer's response is sent to the server. The virtual listener's message sent from the server is displayed in the chat box on the device. The streamer can react through this chat box and enter their own message.

[0284] Specifically, if a streamer types, "I'm feeling down today, so I watched a quiet movie," this response is sent to the server and analyzed by the emotion engine. The server then generates the next message, "That's terrible. What movie did you watch?" and sends it to the device.

[0285] User Roles

[0286] The streamer reacts to messages from the virtual listeners displayed on their device, making it feel as if they are having a conversation with a real audience. This allows streamers to enjoy natural conversations with virtual listeners even when there are only a few viewers.

[0287] Example prompt sentences

[0288] Examples of prompts to input into the generative AI model could include specific instructions such as, "When the streamer says, 'Today I'll talk about movies,' generate a response from a virtual listener. Also, generate an appropriate response when the streamer is feeling down."

[0289] The feedback generated based on this prompt is concise and specific, making the streamer's conversational experience smoother and more natural.

[0290] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0291] Step 1:

[0292] The server starts a conversation message generating program for the virtual user when the broadcaster starts broadcasting.

[0293] Specific operation: A request to start streaming is sent from the streamer's terminal to the server. Upon receiving this request, the server begins creating a virtual listener account.

[0294] Input: Request to start streaming

[0295] Output: Virtual listener account created

[0296] Step 2:

[0297] The server creates multiple virtual listener accounts based on the configuration.

[0298] Specific operation: The server selects a virtual listener from a list of pre-defined profile data (such as name, icon, attributes, etc.) randomly or based on specific conditions, generates an account, and registers it in the database.

[0299] Input: Profile data list

[0300] Output: Virtual listener account

[0301] Step 3:

[0302] The server generates an initial message using the generated virtual listener account and sends it to the terminal.

[0303] Specific behavior: Generates an initial message from the virtual listener (e.g., "What's the topic today?") and sends it to the streamer's chat box using a real-time communication protocol (e.g., WebSocket).

[0304] Input: Virtual listener account, initial message creation prompt

[0305] Output: Initial message

[0306] Step 4:

[0307] The broadcaster responds to the initial message that appears in the chat box on the device.

[0308] Specific operation: The broadcaster enters a message in the chat box (e.g., "Today I'll talk about movies") and presses the send button. This action sends the message to the server.

[0309] Input: Initial message, broadcaster response

[0310] Output: Streamer's response message

[0311] Step 5:

[0312] The server receives the responding message from the sender and analyzes its contents using an emotion engine.

[0313] Specific operation: The publisher's response message is passed to the emotion engine, which uses natural language processing techniques (e.g., keyword extraction, emotion analysis) to determine the emotional state of the message. For example, the message is determined to be "movie" and "excited."

[0314] Input: Streamer's response message

[0315] Output: Emotional state (excitement, joy, etc.)

[0316] Step 6:

[0317] The server generates the next conversational message based on the emotional state.

[0318] Specific behavior: Using the emotional state and the generative AI model's prompt sentence (e.g., "Generate a virtual listener's response when the streamer responds that they're talking about movies"), generate the next conversational message (e.g., "That's fun! Which movie is it?").

[0319] Input: Emotional state, prompt sentence

[0320] Output: Next conversation message

[0321] Step 7:

[0322] The server transmits the next generated conversation message to the broadcaster's terminal.

[0323] Specific operation: The server sends the generated message to the device and displays it in the chat box, allowing the broadcaster to see the next message and respond.

[0324] Input: Next conversation message

[0325] Output: Conversation messages displayed on the terminal

[0326] (Application example 2)

[0327] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0328] In conventional streaming services, when the number of viewers is small at first, there is little interaction with the streamer, making it difficult to keep viewers engaged. It is also difficult for streamers to grasp viewers' emotions in real time and respond appropriately, resulting in a decline in the quality of interaction with viewers. To solve these issues, a system that makes the dialogue between viewers and streamers more natural and emotional is needed.

[0329] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for generating a conversation message of a virtual user, means for transmitting the conversation message of the virtual user to a user terminal, means for receiving a response from the user and generating a next conversation message based on the response, means for again transmitting the next conversation message to the user terminal, and means for recognizing the emotional state of the user and generating a conversation message based on the emotional state. This allows the broadcaster to have a conversation that is responsive to the viewers' emotions, making interactive broadcasting possible even when there are only a few viewers initially.

[0330] A "virtual user" is a non-existent user that is generated within a computer system and behaves in the same way as a real user.

[0331] "Conversational messages" are messages in text or audio format that a user or virtual user uses in a dialogue.

[0332] "User terminal" refers to a device such as a computer, smartphone, or tablet used by a broadcaster or viewer.

[0333] "Emotional state" refers to the user's feelings such as joy, sadness, anger, excitement, etc.

[0334] An "emotion engine" is software or hardware that recognizes a user's emotional state from their messages and actions.

[0335] "Analysis" is the process of interpreting, breaking down, and understanding user responses and input information.

[0336] A "creation means" is a method or device for performing a specific process or creating a specific result.

[0337] A "server" is a computer system that provides services to other computers and devices over a network.

[0338] In this invention, a system is constructed that allows a broadcaster to chat with a virtual user in a natural and emotional way. Each component of the system will be described in detail below.

[0339] Server processing

[0340] The server first generates a conversational message for the virtual user, using pre-prepared message templates and dialogue models. The generated message is then sent to the user's device. The server also receives and analyzes responses from the broadcaster. For analysis, it uses an emotion engine (for example, Hugging Face's transformers library or a specific emotion recognition model, mrm8488 / distilroberta-finetuned-emotion) to recognize the user's emotional state. Based on this emotional state, the server generates the next conversational message and sends it back to the user's device.

[0341] Processing by the terminal

[0342] At the start of a broadcast, the device sends a notification to the server and receives conversational messages from virtual listeners. These messages are displayed in the chat box, and the broadcaster types a response. This response is sent to the server for further analysis and message generation.

[0343] User operation

[0344] The streamer reacts to messages from the virtual listeners that appear in the chat box on their device, making the streamer feel as if they are having a conversation with a real viewer. For example, if a message from an initial virtual listener appears saying, "Hello! How are you doing today?", the streamer can respond with, "I saw a great movie today." Based on this response, the emotion engine recognizes the streamer's emotional state as "joy" and generates a message saying, "Thanks for sharing your wonderful story! Tell me more!"

[0345] Specific examples

[0346] When a broadcaster starts broadcasting, the server generates a message from the initial virtual listener, "What do you think about the recent news?", and sends it to the device. If the broadcaster responds with "The recent news is a little sad," the server analyzes this response with its emotion engine and determines the user's emotional state as "sad." It then generates a message saying, "That's a sad story...that must have been tough," and sends it back to the device.

[0347] Prompt Sentence Examples

[0348] In an emotion-aware live chat application, design an AI model to generate an appropriate response based on the streamer's input message, "I watched a great movie today."

[0349] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0350] Step 1:

[0351] The server receives a notification that distribution has started. When the broadcaster presses the start distribution button on the device, the device sends a start notification to the server. The server receives this notification and prepares to start distribution.

[0352] Step 2:

[0353] The server generates the initial conversation message for the virtual user. For generation, it uses pre-prepared message templates and dialogue models. This generated message is stored in the server. An example of an inserted template message is "Hello! How are you doing today?"

[0354] Step 3:

[0355] The server sends an initial conversation message to the user terminal of the virtual user that has been created. The message is transferred from the server to the user terminal as a data packet and displayed in a chat box on the terminal.

[0356] Step 4:

[0357] The user terminal displays the received message from the virtual user in a chat box on the screen, and the broadcaster visually confirms the displayed message.

[0358] Step 5:

[0359] The broadcaster inputs text to respond to the virtual user's message. The input text is temporarily stored on the terminal and then sent to the server. An example input might be, "I saw a fun movie today."

[0360] Step 6:

[0361] The server receives the response from the broadcaster and analyzes the message using the emotion engine, which uses Hugging Face's transformers library and the emotion analysis model mrm8488 / distilroberta-finetuned-emotion. The response text is analyzed to determine the associated emotional state (e.g., "joy").

[0362] Step 7:

[0363] The server generates a new conversational message based on the emotional state. This new message will match the broadcaster's emotional state. For example, if the analysis result is "joy," the server generates a message like "Thank you for sharing your wonderful story! Tell me more!"

[0364] Step 8:

[0365] The server then sends the newly generated conversation message to the user terminal again as a data packet, which is then displayed in the chat box of the user terminal.

[0366] Step 9:

[0367] The user terminal displays the new message received again in the chat box, allowing the broadcaster to confirm it and continue the interaction to enter the next response.

[0368] This allows the broadcaster to continue a natural and emotionally appropriate conversation with the virtual user.

[0369] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0370] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0371] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0372] [Second embodiment]

[0373] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0374] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0375] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0376] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0377] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0378] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0379] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0380] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0381] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0382] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0383] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0384] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0385] The present invention is a listener AI system that enables broadcasters to have natural chats with virtual users. The system operates with a server, terminals, and users as its main components.

[0386] System Program

[0387] Server processing

[0388] First, we will explain the processing performed by the server. The server includes a program that generates conversation messages for virtual users. The server starts this program when a broadcaster starts broadcasting, and generates multiple virtual listener accounts based on the settings. This makes it possible to send messages to the chat box immediately after the broadcast starts.

[0389] The server analyzes the responses received from the streamer's user device in real time. During this analysis phase, the streamer's message content is broken down into keywords and the appropriate conversation message to be sent next is generated. For example, if the streamer replies, "Today I'll talk about game strategies," the server analyzes the keyword "game strategies" and generates the next message: "That's fun! Which game?"

[0390] Processing by the terminal

[0391] Next, we will explain the processing performed by the terminal. When broadcasting starts, the broadcaster's terminal sends a start notification to the server. When a conversation message is sent from the server by the virtual listener, the terminal displays it in the chat box.

[0392] The broadcaster checks the chat message through the terminal and enters a response, which is then sent back to the server for analysis.

[0393] User operation

[0394] The streamer reacts to messages from the virtual listeners that appear in the chat box on their device, making the streamer feel as if they are having a conversation with a real viewer. For example, if the initial message from the virtual listener says, "Nice to meet you! What will we talk about today?", the streamer will respond, "Today I'll talk about game strategies."

[0395] In response, the server generates the next conversation message and sends it to the device, allowing the broadcaster to continue the conversation. As the broadcast progresses, even if the number of real viewers increases, the device will appropriately display messages from both virtual listeners and real viewers, allowing the broadcaster to respond to both.

[0396] Specific examples

[0397] When a broadcaster starts broadcasting, the server first generates a virtual listener message, "What are we talking about today?", and sends it to the terminal. When the broadcaster responds, "Today we'll talk about movies," the server analyzes this response and generates and sends the next message, "Which movie?"

[0398] In this way, the server, terminals, and users work together to allow the broadcaster to continue broadcasting while engaging in natural chats with virtual listeners. This system enables interactive broadcasting even at the initial stage when there are only a few viewers, helping broadcasters to increase their popularity.

[0399] The processing flow will be explained below.

[0400] Step 1:

[0401] The server registers the broadcaster's user ID in the database. At this time, it also sets the number of virtual listener accounts and parameters required for message generation as the initial settings for the listener AI.

[0402] Step 2:

[0403] The distributor (user) presses the distribution start button. The distributor's terminal sends this distribution start information to the server.

[0404] Step 3:

[0405] The server receives a notification from the streamer that the stream has started, then launches the listener AI and generates multiple virtual listener accounts based on the streamer's settings.

[0406] Step 4:

[0407] The server uses an AI engine to generate an initial chat message. For example, it generates a message like "Nice to meet you! What will you talk about today?" and sends it to the broadcaster's device.

[0408] Step 5:

[0409] The streamer's device displays the chat message sent from the server in the chat box. The streamer checks this and enters a response such as "Today I'll be talking about game strategies."

[0410] Step 6:

[0411] The distributor's terminal transmits the answer entered by the distributor to the server.

[0412] Step 7:

[0413] The server receives the streamer's response and uses an AI engine to analyze it and extract the keyword "game strategy."

[0414] Step 8:

[0415] The server generates the next conversation message based on the analysis result. For example, it generates the message "That's fun! Which game is it?" and sends it to the broadcaster terminal again.

[0416] Step 9:

[0417] The streamer's device will again display the message sent from the server in the chat box. The streamer will continue to reply to the AI's messages.

[0418] Step 10:

[0419] The process loops from step 6 to step 9, and the conversation with the virtual listener continues until the streamer ends the chat. As the stream progresses, messages from real viewers are also displayed, allowing the streamer to respond to both.

[0420] This allows streamers to have interactive broadcasts through conversations with virtual listeners, deepening their bond with their viewers while continuing to stream.

[0421] Example 1

[0422] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0423] When live streamers have few viewers in the early stages of a stream, they may experience little interaction with viewers, resulting in a lack of excitement. As a result, streamers often struggle to attract viewers' interest, and their streams can feel monotonous until the number of viewers increases. Furthermore, with few chat messages from viewers, streamers may be unsure of how to move the conversation forward. To solve these problems, a system is needed that allows streamers to have natural conversations with virtual viewers and realize interactive streams.

[0424] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0425] In this invention, the server includes: a means for generating conversational messages for virtual users; a means for transmitting the conversational messages for virtual users to a user terminal; a means for receiving a user's response and generating a next conversational message based on the response; a means for transmitting the next conversational message to the user terminal; a means for generating multiple virtual listener accounts based on a broadcast start notification received from the user terminal; a means for analyzing the generated conversational messages in real time and breaking down the broadcaster's input message into keywords; and a means for executing the above-mentioned means using a generative AI model. This allows virtual listener accounts to be generated from the early stages of broadcasting, enabling natural conversations to take place, enabling the broadcaster to always broadcast lively and interactively. Furthermore, by analyzing the broadcaster's messages in real time and generating appropriate conversational messages, smooth conversation progression is ensured even as the number of viewers increases.

[0426] A "virtual user" is not a real person, but a character generated by the system and designed to interact with the user.

[0427] "Conversation Message" means text-based communication content sent by a virtual user or user.

[0428] "User Terminal" means a device used by a broadcaster, such as a computer, tablet, or smartphone, that includes software for sending, receiving, and broadcasting chat messages.

[0429] A "generative AI model" refers to an artificial intelligence algorithm that performs natural language processing, and is a technology for generating interactive messages based on user input messages.

[0430] A "prompt" is an instruction given to a generative AI model that provides it with the information it needs to generate an appropriate conversational message.

[0431] "Analysis" is the process of breaking down a received message to understand or make sense of its content.

[0432] A "distribution start notification" is a signal sent to the server when a distributor starts distribution, and serves as a trigger to start up each part of the system.

[0433] A "virtual listener account" is an account created by the system to act as a virtual user.

[0434] "Real-time" refers to near-instant processing or response, with minimal delay.

[0435] "Keywords" refer to important words or phrases in a user's message, and are the building blocks for generating the next conversational message.

[0436] The present invention provides a system that supports broadcasters in progressing with broadcasts while engaging in natural chat with virtual users. This system operates with a server, a terminal, and a user as its main components. Specific embodiments of this system are described below.

[0437] Server processing

[0438] The server includes a program that generates conversational messages from virtual users. This program uses a generative AI model that performs natural language processing. Specifically, OpenAI's GPT-4 can be used. The server launches this program when the streamer begins streaming, and generates multiple virtual listener accounts based on the settings. This allows messages from virtual users to be streamed in the chat box immediately after the stream begins.

[0439] The server also analyzes responses received from the broadcaster's user terminal in real time. This analysis involves breaking down the broadcaster's message content into keywords and generating the appropriate conversational message to be sent next. For example, if the broadcaster replies, "Today, I'll talk about movies," the server analyzes the keyword "movie" and generates the next message, "Which movie are you going to talk about?"

[0440] Processing by the terminal

[0441] When streaming begins, the device sends a start notification to the server. This notification is sent via an HTTP request or WebSocket. When a conversation message is sent from the server by the virtual listener, the device displays it in the chat box. The streamer checks the chat message through the device and enters a response. This response is sent back to the server and used for analysis. The device uses the streaming software OBS Studio or a dedicated streaming app.

[0442] User operation

[0443] The broadcaster reacts to the virtual listener's messages displayed in the chat box on the device, making the broadcaster feel as if they were having a conversation with a real viewer. For example, if the initial virtual listener's message is "Hello! What will you talk about today?", the broadcaster can respond with "I'll talk about movies today." This response is sent to the server, where it is analyzed in real time and the next conversation message is generated.

[0444] Specific examples

[0445] For example, when a broadcaster starts broadcasting, the server first generates a message for the virtual listener asking, "What are we talking about today?" and sends it to the terminal. If the broadcaster responds, "I'll talk about movies today," the server analyzes this response and generates and sends the next message, "Which movie are you talking about?" This allows for a natural chat with the virtual listener.

[0446] Prompt Sentence Examples

[0447] “If the streamer responds, ‘Today we’re going to talk about movies,’ suggest the next conversation message they should generate from their virtual listeners.”

[0448] This system allows streamers to engage with viewers through interactive broadcasts even at the early stages when the number of viewers is still small. As the number of real viewers increases, the system can also provide an environment where streamers can respond smoothly by displaying messages from both virtual listeners and real viewers appropriately.

[0449] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0450] Step 1:

[0451] The terminal sends a notification of the start of distribution to the server.

[0452] Specific operation: When the streamer presses the "Start Streaming" button in the streaming software (e.g., OBS Studio), the device sends a notification to the server via an HTTP request or WebSocket to start streaming.

[0453] Input: "Start distribution" operation of distribution software.

[0454] Output: A notification is sent to the server, and the server goes into a live broadcast state.

[0455] Step 2:

[0456] The server creates multiple virtual listener accounts.

[0457] Specific operation: The server uses the generative AI model to generate profile information (name, icon, etc.) for the virtual listener based on the settings and creates a virtual listener account.

[0458] Input: Notification from device that distribution has started.

[0459] Output: Multiple virtual listener accounts are created on the server.

[0460] Step 3:

[0461] The server generates an initial message.

[0462] What it does: Uses a generative AI model to generate an initial message from a virtual listener (e.g., "Hello! What will we talk about today?").

[0463] Input: Virtual listener account information.

[0464] Output: Initial message from the virtual listener.

[0465] Step 4:

[0466] The server sends the initial message for the virtual listener to the terminal.

[0467] Specific operation: The server generates an initial message and sends it to the terminal via a WebSocket or HTTP request.

[0468] Input: The initial message generated.

[0469] Output: The initial message of the virtual listener is sent to the terminal.

[0470] Step 5:

[0471] Displays messages received by the terminal in the chat box.

[0472] Specific operation: Displays the virtual listener's message in the chat box of the distribution software.

[0473] Input: The initial message sent by the server.

[0474] Output: The message that appears in the chat box.

[0475] Step 6:

[0476] The user (broadcaster) responds to the virtual listener's messages in the chat box.

[0477] Specific operation: The streamer types messages through the chat box and interacts with the virtual listeners. For example, the streamer types, "Today I'll talk about movies."

[0478] Input: The virtual listener's message and the broadcaster's response.

[0479] Output: The streamer's input message.

[0480] Step 7:

[0481] The device sends the distributor's response to the server.

[0482] Specific operation: Sends the broadcaster's input message to the server in real time.

[0483] Input: The streamer's input message.

[0484] Output: The publisher's response sent to the server.

[0485] Step 8:

[0486] The server parses the publisher's response.

[0487] How it works: The server uses the generative AI model to analyze the broadcaster's message by keyword and generate the next appropriate conversational message. For example, it analyzes the keyword "movie" and generates the message "Which movie are you talking about?"

[0488] Input: The broadcaster's response message.

[0489] Output: Analyzed keywords and the next conversation message.

[0490] Step 9:

[0491] The server generates the following conversation message:

[0492] Specific operation: Based on the analysis results, the generative AI model is used to generate the next conversation message.

[0493] Input: Parsed keywords.

[0494] Output: Next conversation message.

[0495] Step 10:

[0496] The server generates a message and sends it to the terminal.

[0497] Specific operation: The next generated conversation message is sent to the device via WebSocket or HTTP request.

[0498] Input: The next generated conversation message.

[0499] Output: The next conversation message sent to the terminal.

[0500] Step 11:

[0501] Displays messages received by the terminal in the chat box.

[0502] Specific operation: The following conversation message will be displayed in the chat box of the distribution software.

[0503] Input: The next conversation message sent from the server.

[0504] Output: The message that appears in the chat box.

[0505] The above processing steps allow the broadcaster to continue broadcasting while engaging in natural chat with the virtual listeners. This allows for interactive broadcasting even when the number of viewers is small from the beginning, contributing to the broadcaster's popularity.

[0506] (Application example 1)

[0507] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0508] In live streaming and content distribution, streamers can feel isolated in the early stages when there are few viewers. This can make interactive streaming difficult, leading to a decline in streamer motivation and a decline in viewer interest. Furthermore, there is a lack of effective means to properly manage messages for both real viewers and virtual listeners.

[0509] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0510] In this invention, the server includes means for generating conversational messages of a virtual user, means for transmitting the conversational messages of the virtual user to a terminal device, means for receiving responses from the user and generating the next conversational message based on the responses, means for analyzing the response messages and generating natural conversation based on prompt sentences using a generative AI model, and means for appropriately processing and displaying the messages of the virtual user and the messages of the actual viewers. This makes it easier for the broadcaster to maintain interactive conversations even in the early stages when the number of viewers is small, thereby improving the quality of the broadcast.

[0511] "Virtual user conversation messages" refers to messages of users virtually created using a generative AI model.

[0512] "Terminal device" refers to a device used by a broadcaster, including a smartphone or head-mounted display.

[0513] "User response" refers to a reply message entered by a broadcaster in response to a conversation message from a virtual user.

[0514] "Generative AI model" refers to an algorithm or system that uses generative AI technology to generate natural-sounding conversations.

[0515] A "prompt sentence" is the base text input to a generative AI model that determines the content of the next conversational message generated.

[0516] "Natural conversation" refers to messages generated by a generative AI model based on prompts that feel like real human conversations.

[0517] "Messages and actual viewer messages" refers to messages generated by virtual users or AI, as well as messages entered by viewers in real time.

[0518] "Means for appropriate processing and display" refers to the technology and algorithms that distinguish between messages from virtual users and messages from actual viewers and display them in a way that is easy for users to understand.

[0519] The present invention provides a system for generating conversational messages of virtual users and enabling broadcasters to broadcast interactive live content. The system operates using a server, a terminal device, and a broadcaster as its main components.

[0520] Server Processing

[0521] The server includes a program that generates conversation messages for virtual users. When a broadcast starts, the server starts this program and generates multiple accounts for the virtual users. As a result, messages from the virtual users are automatically displayed in the chat box as soon as the broadcast starts.

[0522] The server has a response message receiving function that analyzes the streamer's messages in real time. It breaks down the message entered by the streamer into keywords and generates the next conversation message based on these keywords. The generated message is input into a generative AI model using prompt sentences to create a natural conversation. For example, if the streamer enters "Today we will talk about movies," the server analyzes the keyword "movie" and generates the next message, "Which movie?"

[0523] Terminal processing

[0524] The terminal device is a device used by the broadcaster, such as a smartphone or head-mounted display. When broadcasting begins, the terminal device sends a start notification to the server. When a message from the virtual user is sent from the server, the terminal device has the function of displaying the message directly in the chat box.

[0525] The terminal device also has the role of transmitting the response message entered by the distributor to the server, which allows the server to analyze and generate the response message.

[0526] User operation

[0527] The broadcaster inputs responses to the virtual user's messages via a chat box displayed on the terminal device. These responses are sent to the server, where they are analyzed and used to generate the next conversation message. This system allows the broadcaster to continuously engage in natural interactions with the virtual user and the actual viewers.

[0528] Specific examples

[0529] When a broadcaster starts broadcasting, the server first generates a message from the virtual user saying, "What's the topic today?" and sends it to the terminal device. When the broadcaster responds, "Today I'll introduce a cooking recipe," the server analyzes this response, generates the next message, "That's great! What kind of cooking recipe is it?" and sends it again to the terminal device.

[0530] Prompt Sentence Examples

[0531] Streamer: Today I'll show you a recipe

[0532] Virtual users:

[0533] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0534] Step 1:

[0535] The server receives the notification that distribution has started.

[0536] Input: Notification of start of broadcasting from the broadcaster's terminal device.

[0537] Operation: The server receives the notification of the start of distribution and starts a program that generates conversational messages for the virtual user.

[0538] Output: Launch of the virtual user conversation message generation program.

[0539] Step 2:

[0540] The server generates an initial conversation message for the virtual user.

[0541] Input: Notification of start of distribution.

[0542] How it works: The server uses a generative AI model to generate an initial conversation message based on a configured prompt, such as "What's the topic today?"

[0543] Output: Initial virtual user conversation message.

[0544] Step 3:

[0545] The server sends an initial conversation message to the terminal device of the virtual user.

[0546] Input: The initial conversation message of the generated virtual user.

[0547] Operation: The server sends the generated initial conversation message to the broadcaster's terminal device.

[0548] Output: The initial conversation message of the virtual user displayed on the broadcaster's terminal device.

[0549] Step 4:

[0550] The broadcaster responds to the initial conversation message.

[0551] Input: The initial conversation message of the virtual user.

[0552] How it works: The broadcaster responds to the initial conversation message through the chat box on the terminal device, for example, typing, "Today I'll introduce you to a cooking recipe."

[0553] Output: The distributor's response message.

[0554] Step 5:

[0555] The terminal device sends the distributor's response message to the server.

[0556] Input: The response message entered by the broadcaster.

[0557] Operation: The terminal device sends the message entered by the broadcaster to the server.

[0558] Output: The publisher's response message sent to the server.

[0559] Step 6:

[0560] The server analyzes the distributor's response message.

[0561] Input: The broadcaster's response message.

[0562] How it works: The server breaks down the broadcaster's response message into keywords and creates a prompt using a generative AI model.

[0563] Output: Analysis results and prompt for the next conversation message.

[0564] Step 7:

[0565] The server generates the following conversation message:

[0566] Input: Prompt statement and analysis result.

[0567] How it works: The server uses a generative AI model to generate the next conversational message based on the prompt, for example, "That's great! What's the recipe?"

[0568] Output: The following conversation message is generated:

[0569] Step 8:

[0570] The server sends the following conversation message to the terminal device:

[0571] Input: The next generated conversation message.

[0572] Operation: The server again sends the next conversation message that has been generated to the broadcaster's terminal device.

[0573] Output: The next conversation message displayed on the terminal.

[0574] The above processing steps enable the broadcaster to broadcast live while maintaining natural interaction with the virtual user.

[0575] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0576] This invention is a system that combines a listener AI system that enables broadcasters to chat naturally with virtual users with an emotion engine that recognizes the user's emotions. This system operates with a server, a terminal, and a user as its main components, and the emotion engine can generate appropriate conversational messages according to the user's emotional state.

[0577] System Program

[0578] Server processing

[0579] First, we will explain the processing performed by the server. The server includes a program that generates conversation messages for virtual users. When a broadcaster starts broadcasting, this program is launched and multiple virtual listener accounts are generated based on the settings. This makes it possible to send messages to the chat box immediately after the broadcast begins.

[0580] The server also has an emotion engine built in that recognizes the emotional state of the streamer when analyzing their response. During this analysis phase, the content of the streamer's message is broken down into keywords, and the emotion engine determines their emotional state (e.g., happy, sad, or angry). The next appropriate conversational message to be sent is generated based on this emotional state. For example, if the streamer replies, "I'm not feeling very well today," the server analyzes the keyword "not feeling well" and their emotional state, and generates the message, "That sounds worrying. Are you okay?"

[0581] Processing by the terminal

[0582] Next, we will explain the processing performed by the terminal. When broadcasting starts, the broadcaster's terminal sends a start notification to the server. When a conversation message is sent from the server by the virtual listener, the terminal displays it in the chat box.

[0583] The streamer checks the chat message through their device and enters a response, which is then sent back to the server and used for analysis by the emotion engine.

[0584] User operation

[0585] The streamer reacts to messages from virtual listeners that appear in the chat box on their device, making the streamer feel as if they are having a conversation with a real viewer. For example, if an initial message from a virtual listener says, "Nice to meet you! What will you talk about today?", the streamer might respond, "I'm not feeling too great today, but I'll talk about movies."

[0586] In response, the server generates the next conversation message and sends it to the device, allowing the streamer to continue the conversation. This system allows streamers to deepen their interaction with viewers while having natural conversations with virtual listeners.

[0587] Specific examples

[0588] When a broadcaster starts broadcasting, the server first generates a message for the virtual listener asking, "What's the topic today?" and sends it to the terminal. When the broadcaster responds, "Today we'll talk about movies," the server analyzes this response with its emotion engine and determines the user's emotional state as "excitement" or "joy." It then generates the next message, "That's exciting! Which movie?" and sends it again to the broadcaster's terminal.

[0589] If the streamer responds, "I'm feeling down today, so I watched a quiet movie," the server analyzes the keyword "depressed" and the emotional state and generates a message saying, "That's terrible. What movie did you watch?"

[0590] In this way, the server, device, user, and emotion engine work together to allow the streamer to continue streaming while engaging in natural, emotionally appropriate conversations with virtual listeners. This system enables interactive streaming even at the initial stage, when there are only a few viewers, and helps streamers increase their popularity.

[0591] The processing flow will be explained below.

[0592] Step 1:

[0593] The server registers the broadcaster's user ID in the database. At this time, the server sets up the initial settings for the listener AI, including the number of virtual listener accounts, parameters required for message generation, and the emotion engine.

[0594] Step 2:

[0595] The distributor (user) presses the distribution start button. The distributor's terminal sends this distribution start information to the server.

[0596] Step 3:

[0597] The server receives a notification from the streamer that the stream has started, then launches the listener AI and generates multiple virtual listener accounts based on the streamer's settings.

[0598] Step 4:

[0599] The server uses an AI engine to generate an initial chat message. For example, it generates a message like "Nice to meet you! What will you talk about today?" and sends it to the broadcaster's device.

[0600] Step 5:

[0601] The streamer's device displays the chat message sent from the server in the chat box. The streamer checks this and enters a response such as "Today I'll be talking about game strategies."

[0602] Step 6:

[0603] The distributor's terminal transmits the answer entered by the distributor to the server.

[0604] Step 7:

[0605] The server receives the publisher's response, which is analyzed by the emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger).

[0606] Step 8:

[0607] The server generates the next conversation message based on the analysis results of the emotion engine. For example, if the broadcaster replies, "I'm a little sad today, but I'll talk about movies," the server recognizes the emotional state as "sad" and generates the message, "That's terrible. What movie did you see?" The generated message is then sent back to the broadcaster's device.

[0608] Step 9:

[0609] The streamer's device will again display the message sent from the server in the chat box. The streamer will continue to reply to the AI's messages.

[0610] Step 10:

[0611] The process loops from step 6 to step 9, and the conversation with the virtual listener continues until the streamer ends the chat. As the stream progresses, messages from real viewers are also displayed, allowing the streamer to respond to both.

[0612] In this way, the server, device, user, and emotion engine work together to allow the streamer to continue streaming while engaging in natural, emotionally appropriate conversations with virtual listeners. This system enables interactive streaming even at the initial stage, when there are only a few viewers, and helps streamers increase their popularity.

[0613] Example 2

[0614] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0615] In modern live streaming, a lack of interaction in the early stages when viewers are few can cause streamers to lose motivation. Furthermore, in order for streamers to enjoy natural conversations with viewers, they need to respond in real time, which places a heavy burden on them. A system that can solve these problems and enable streamers to enjoy natural conversations even with a small audience is needed.

[0616] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0617] In this invention, the server includes means for generating a conversational message of a virtual user, means for transmitting the conversational message of the virtual user to an information processing device, means for receiving a response from the broadcaster and analyzing the emotional state of the response, means for generating a next conversational message based on the emotional state, and means for transmitting the next conversational message again to the information processing device. This allows the broadcaster to enjoy natural conversation with the virtual listener even when there are only a few viewers.

[0618] A "virtual user" refers to an artificial participant that does not exist in reality but is generated by a computer program and behaves like a real user.

[0619] "Conversational messages" are messages sent and received in a chat or dialogue format, and refer to sentences or text exchanges generated by users or virtual users.

[0620] An "information processing device" is an electronic device used for calculations, data processing, communication, etc., and includes, for example, computers, smartphones, tablets, etc.

[0621] "Distributor" means an individual or entity that distributes live streaming or real-time video content over the Internet.

[0622] "Emotional state" refers to the emotional state of a person, such as joy, sadness, or anger, that is analyzed from text or audio.

[0623] "Emotion engine" refers to a software module that analyzes user input data and determines the emotions contained in the input.

[0624] "Generating means" refers to technical mechanisms or algorithms that automatically generate data or messages for a specific purpose.

[0625] "Transmitting means" refers to the technical mechanism by which data or messages are transferred from one information processing device to another device or server.

[0626] "Means of receiving" refers to the technical mechanisms for acquiring and processing data or messages sent from outside.

[0627] "Means of analysis" refers to the technical mechanisms used to analyze received data or messages and understand their content and meaning.

[0628] This invention is a listener AI system that enables broadcasters to chat naturally with virtual users, and is a system that combines an emotion engine that recognizes the user's emotions. The main components of this system are a server, a terminal, and a user.

[0629] Server Roles

[0630] The server first generates conversational messages for virtual users. When a broadcaster starts broadcasting, the server creates a virtual listener account and sends an initial message based on the broadcaster's settings to the broadcaster's device. The server is equipped with an emotion engine that analyzes the broadcaster's responses in real time and determines their emotional state.

[0631] As a specific example, if the broadcaster responds, "Today we'll talk about movies," the server analyzes this message with an emotion engine, generates the next conversational message, "That's exciting! Which movie?", and sends it to the device. This emotion engine can use, for example, a natural language processing library implemented in Python or a machine learning model (e.g., BERT).

[0632] Device Role

[0633] The streamer's device sends a notification to the server when the stream starts, and the streamer's response is sent to the server. The virtual listener's message sent from the server is displayed in the chat box on the device. The streamer can react through this chat box and enter their own message.

[0634] Specifically, if a streamer types, "I'm feeling down today, so I watched a quiet movie," this response is sent to the server and analyzed by the emotion engine. The server then generates the next message, "That's terrible. What movie did you watch?" and sends it to the device.

[0635] User Roles

[0636] The streamer reacts to messages from the virtual listeners displayed on their device, making it feel as if they are having a conversation with a real audience. This allows streamers to enjoy natural conversations with virtual listeners even when there are only a few viewers.

[0637] Example prompt sentences

[0638] Examples of prompts to input into the generative AI model could include specific instructions such as, "When the streamer says, 'Today I'll talk about movies,' generate a response from a virtual listener. Also, generate an appropriate response when the streamer is feeling down."

[0639] The feedback generated based on this prompt is concise and specific, making the streamer's conversational experience smoother and more natural.

[0640] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0641] Step 1:

[0642] The server starts a conversation message generating program for the virtual user when the broadcaster starts broadcasting.

[0643] Specific operation: A request to start streaming is sent from the streamer's terminal to the server. Upon receiving this request, the server begins creating virtual listener accounts.

[0644] Input: Request to start streaming

[0645] Output: Virtual listener account created

[0646] Step 2:

[0647] The server creates multiple virtual listener accounts based on the configuration.

[0648] Specific operation: The server selects a virtual listener from a list of pre-defined profile data (such as name, icon, attributes, etc.) randomly or based on specific conditions, generates an account, and registers it in the database.

[0649] Input: Profile data list

[0650] Output: Virtual listener account

[0651] Step 3:

[0652] The server generates an initial message using the generated virtual listener account and sends it to the terminal.

[0653] Specific behavior: Generates an initial message from the virtual listener (e.g., "What's the topic today?") and sends it to the streamer's chat box using a real-time communication protocol (e.g., WebSocket).

[0654] Input: Virtual listener account, initial message creation prompt

[0655] Output: Initial message

[0656] Step 4:

[0657] The broadcaster responds to the initial message that appears in the chat box on the device.

[0658] Specific operation: The broadcaster enters a message in the chat box (e.g., "Today I'll talk about movies") and presses the send button. This action sends the message to the server.

[0659] Input: Initial message, broadcaster response

[0660] Output: Streamer's response message

[0661] Step 5:

[0662] The server receives the responding message from the distributor and analyzes its contents using an emotion engine.

[0663] Specific operation: The publisher's response message is passed to the emotion engine, which uses natural language processing techniques (e.g., keyword extraction, emotion analysis) to determine the emotional state of the message. For example, the message is determined to be "movie" and "excited."

[0664] Input: Streamer's response message

[0665] Output: Emotional state (excitement, joy, etc.)

[0666] Step 6:

[0667] The server generates the next conversational message based on the emotional state.

[0668] Specific behavior: Using the emotional state and the generative AI model's prompt sentence (e.g., "Generate a virtual listener's response when the streamer responds that they're talking about movies"), generate the next conversational message (e.g., "That's fun! Which movie is it?").

[0669] Input: Emotional state, prompt sentence

[0670] Output: Next conversation message

[0671] Step 7:

[0672] The server transmits the next generated conversation message to the broadcaster's terminal.

[0673] Specific operation: The server sends the generated message to the device and displays it in the chat box, allowing the broadcaster to see the next message and respond.

[0674] Input: Next conversation message

[0675] Output: Conversation messages displayed on the terminal

[0676] (Application example 2)

[0677] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0678] In conventional streaming services, when the number of viewers is small at first, there is little interaction with the streamer, making it difficult to keep viewers engaged. It is also difficult for streamers to grasp viewers' emotions in real time and respond appropriately, resulting in a decline in the quality of interaction with viewers. To solve these issues, a system that makes the dialogue between viewers and streamers more natural and emotional is needed.

[0679] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for generating a conversation message of a virtual user, means for transmitting the conversation message of the virtual user to a user terminal, means for receiving a response from the user and generating a next conversation message based on the response, means for again transmitting the next conversation message to the user terminal, and means for recognizing the emotional state of the user and generating a conversation message based on the emotional state. This allows the broadcaster to have a conversation that is responsive to the viewers' emotions, making interactive broadcasting possible even when there are only a few viewers initially.

[0680] A "virtual user" is a non-existent user that is generated within a computer system and behaves in the same way as a real user.

[0681] "Conversational messages" are messages in text or audio format that a user or virtual user uses in a dialogue.

[0682] "User terminal" refers to a device such as a computer, smartphone, or tablet used by a broadcaster or viewer.

[0683] "Emotional state" refers to the user's feelings such as joy, sadness, anger, excitement, etc.

[0684] An "emotion engine" is software or hardware that recognizes a user's emotional state from their messages and actions.

[0685] "Analysis" is the process of interpreting, breaking down, and understanding user responses and input information.

[0686] A "creation means" is a method or device for performing a specific process or creating a specific result.

[0687] A "server" is a computer system that provides services to other computers and devices over a network.

[0688] In this invention, a system is constructed that allows a broadcaster to chat with a virtual user in a natural and emotional way. Each component of the system will be described in detail below.

[0689] Server processing

[0690] The server first generates a conversational message for the virtual user, using pre-prepared message templates and dialogue models. The generated message is then sent to the user's device. The server also receives and analyzes responses from the broadcaster. For analysis, it uses an emotion engine (for example, Hugging Face's transformers library or a specific emotion recognition model, mrm8488 / distilroberta-finetuned-emotion) to recognize the user's emotional state. Based on this emotional state, the server generates the next conversational message and sends it back to the user's device.

[0691] Processing by the terminal

[0692] At the start of a broadcast, the device sends a notification to the server and receives conversational messages from virtual listeners. These messages are displayed in the chat box, and the broadcaster types a response. This response is sent to the server for further analysis and message generation.

[0693] User operation

[0694] The streamer reacts to messages from the virtual listeners that appear in the chat box on their device, making the streamer feel as if they are having a conversation with a real viewer. For example, if a message from an initial virtual listener appears saying, "Hello! How are you doing today?", the streamer can respond with, "I saw a great movie today." Based on this response, the emotion engine recognizes the streamer's emotional state as "joy" and generates a message saying, "Thanks for sharing your wonderful story! Tell me more!"

[0695] Specific examples

[0696] When a broadcaster starts broadcasting, the server generates a message from the initial virtual listener, "What do you think about the recent news?", and sends it to the device. If the broadcaster responds with "The recent news is a little sad," the server analyzes this response with its emotion engine and determines the user's emotional state as "sad." It then generates a message saying, "That's a sad story...that must have been tough," and sends it back to the device.

[0697] Prompt Sentence Examples

[0698] In an emotion-aware live chat application, design an AI model to generate an appropriate response based on the streamer's input message, "I watched a great movie today."

[0699] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0700] Step 1:

[0701] The server receives a notification that distribution has started. When the broadcaster presses the start distribution button on the device, the device sends a start notification to the server. The server receives this notification and prepares to start distribution.

[0702] Step 2:

[0703] The server generates the initial conversation message for the virtual user. For generation, it uses pre-prepared message templates and dialogue models. This generated message is stored in the server. An example of an inserted template message is "Hello! How are you doing today?"

[0704] Step 3:

[0705] The server sends an initial conversation message to the user terminal of the virtual user that has been created. The message is transferred from the server to the user terminal as a data packet and displayed in a chat box on the terminal.

[0706] Step 4:

[0707] The user terminal displays the received message from the virtual user in a chat box on the screen, and the broadcaster visually confirms the displayed message.

[0708] Step 5:

[0709] The broadcaster inputs text to respond to the virtual user's message. The input text is temporarily stored on the terminal and then sent to the server. An example input might be, "I saw a fun movie today."

[0710] Step 6:

[0711] The server receives the response from the broadcaster and analyzes the message using the emotion engine, which uses Hugging Face's transformers library and the emotion analysis model mrm8488 / distilroberta-finetuned-emotion. The response text is analyzed to determine the associated emotional state (e.g., "joy").

[0712] Step 7:

[0713] The server generates a new conversational message based on the emotional state. This new message will match the broadcaster's emotional state. For example, if the analysis result is "joy," the server generates a message like "Thank you for sharing your wonderful story! Tell me more!"

[0714] Step 8:

[0715] The server then sends the newly generated conversation message to the user terminal again as a data packet, which is then displayed in the chat box of the user terminal.

[0716] Step 9:

[0717] The user terminal displays the new message received again in the chat box, allowing the broadcaster to confirm it and continue the interaction to enter the next response.

[0718] This allows the broadcaster to continue a natural and emotionally appropriate conversation with the virtual user.

[0719] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0720] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0721] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0722] [Third embodiment]

[0723] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0724] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0725] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0726] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0727] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0728] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0729] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0730] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0731] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0732] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0733] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0734] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0735] The present invention is a listener AI system that enables broadcasters to have natural chats with virtual users. The system operates with a server, terminals, and users as its main components.

[0736] System Program

[0737] Server processing

[0738] First, we will explain the processing performed by the server. The server includes a program that generates conversation messages for virtual users. The server starts this program when a broadcaster starts broadcasting, and generates multiple virtual listener accounts based on the settings. This makes it possible to send messages to the chat box immediately after the broadcast starts.

[0739] The server analyzes the responses received from the streamer's user device in real time. During this analysis phase, the streamer's message content is broken down into keywords and the appropriate conversation message to be sent next is generated. For example, if the streamer replies, "Today I'll talk about game strategies," the server analyzes the keyword "game strategies" and generates the next message: "That's fun! Which game?"

[0740] Processing by the terminal

[0741] Next, we will explain the processing performed by the terminal. When broadcasting starts, the broadcaster's terminal sends a start notification to the server. When a conversation message is sent from the server by the virtual listener, the terminal displays it in the chat box.

[0742] The broadcaster checks the chat message through the terminal and enters a response, which is then sent back to the server for analysis.

[0743] User operation

[0744] The streamer reacts to messages from the virtual listeners that appear in the chat box on their device, making the streamer feel as if they are having a conversation with a real viewer. For example, if the initial message from the virtual listener says, "Nice to meet you! What will we talk about today?", the streamer will respond, "Today I'll talk about game strategies."

[0745] In response, the server generates the next conversation message and sends it to the device, allowing the broadcaster to continue the conversation. As the broadcast progresses, even if the number of real viewers increases, the device will appropriately display messages from both virtual listeners and real viewers, allowing the broadcaster to respond to both.

[0746] Specific examples

[0747] When a broadcaster starts broadcasting, the server first generates a virtual listener message, "What are we talking about today?", and sends it to the terminal. When the broadcaster responds, "Today we'll talk about movies," the server analyzes this response and generates and sends the next message, "Which movie?"

[0748] In this way, the server, terminals, and users work together to allow the broadcaster to continue broadcasting while engaging in natural chats with virtual listeners. This system enables interactive broadcasting even at the initial stage when there are only a few viewers, helping broadcasters to increase their popularity.

[0749] The processing flow will be explained below.

[0750] Step 1:

[0751] The server registers the broadcaster's user ID in the database. At this time, it also sets the number of virtual listener accounts and parameters required for message generation as the initial settings for the listener AI.

[0752] Step 2:

[0753] The distributor (user) presses the distribution start button. The distributor's terminal sends this distribution start information to the server.

[0754] Step 3:

[0755] The server receives a notification from the streamer that the stream has started, then launches the listener AI and generates multiple virtual listener accounts based on the streamer's settings.

[0756] Step 4:

[0757] The server uses an AI engine to generate an initial chat message. For example, it generates a message like "Nice to meet you! What will you talk about today?" and sends it to the broadcaster's device.

[0758] Step 5:

[0759] The streamer's device displays the chat message sent from the server in the chat box. The streamer checks this and enters a response such as "Today I'll be talking about game strategies."

[0760] Step 6:

[0761] The distributor's terminal transmits the answer entered by the distributor to the server.

[0762] Step 7:

[0763] The server receives the streamer's response and uses an AI engine to analyze it and extract the keyword "game strategy."

[0764] Step 8:

[0765] The server generates the next conversation message based on the analysis result. For example, it generates the message "That's fun! Which game is it?" and sends it to the broadcaster terminal again.

[0766] Step 9:

[0767] The streamer's device will again display the message sent from the server in the chat box. The streamer will continue to reply to the AI's messages.

[0768] Step 10:

[0769] The process loops from step 6 to step 9, and the conversation with the virtual listener continues until the streamer ends the chat. As the stream progresses, messages from real viewers are also displayed, allowing the streamer to respond to both.

[0770] This allows streamers to have interactive broadcasts through conversations with virtual listeners, deepening their bond with their viewers while continuing to stream.

[0771] Example 1

[0772] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0773] When live streamers have few viewers in the early stages of a stream, they may experience little interaction with viewers, resulting in a lack of excitement. As a result, streamers often struggle to attract viewers' interest, and their streams can feel monotonous until the number of viewers increases. Furthermore, with few chat messages from viewers, streamers may be unsure of how to move the conversation forward. To solve these problems, a system is needed that allows streamers to have natural conversations with virtual viewers and realize interactive streams.

[0774] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0775] In this invention, the server includes: a means for generating conversational messages for virtual users; a means for transmitting the conversational messages for virtual users to a user terminal; a means for receiving a user's response and generating a next conversational message based on the response; a means for transmitting the next conversational message to the user terminal; a means for generating multiple virtual listener accounts based on a broadcast start notification received from the user terminal; a means for analyzing the generated conversational messages in real time and breaking down the broadcaster's input message into keywords; and a means for executing the above-mentioned means using a generative AI model. This allows virtual listener accounts to be generated from the early stages of broadcasting, enabling natural conversations to take place, enabling the broadcaster to always broadcast lively and interactively. Furthermore, by analyzing the broadcaster's messages in real time and generating appropriate conversational messages, smooth conversation progression is ensured even as the number of viewers increases.

[0776] A "virtual user" is not a real person, but a character generated by the system and designed to interact with the user.

[0777] "Conversation Message" means text-based communication content sent by a virtual user or user.

[0778] "User Terminal" means a device used by a broadcaster, such as a computer, tablet, or smartphone, that includes software for sending, receiving, and broadcasting chat messages.

[0779] A "generative AI model" refers to an artificial intelligence algorithm that performs natural language processing, and is a technology for generating interactive messages based on user input messages.

[0780] A "prompt" is an instruction given to a generative AI model that provides it with the information it needs to generate an appropriate conversational message.

[0781] "Analysis" is the process of breaking down a received message to understand or make sense of its content.

[0782] A "distribution start notification" is a signal sent to the server when a distributor starts distribution, and serves as a trigger to start up each part of the system.

[0783] A "virtual listener account" is an account created by the system to act as a virtual user.

[0784] "Real-time" refers to near-instant processing or response, with minimal delay.

[0785] "Keywords" refer to important words or phrases in a user's message, and are the building blocks for generating the next conversational message.

[0786] The present invention provides a system that supports broadcasters in progressing with broadcasts while engaging in natural chat with virtual users. This system operates with a server, a terminal, and a user as its main components. Specific embodiments of this system are described below.

[0787] Server processing

[0788] The server includes a program that generates conversational messages from virtual users. This program uses a generative AI model that performs natural language processing. Specifically, OpenAI's GPT-4 can be used. The server launches this program when the streamer begins streaming, and generates multiple virtual listener accounts based on the settings. This allows messages from virtual users to be streamed in the chat box immediately after the stream begins.

[0789] The server also analyzes responses received from the broadcaster's user terminal in real time. This analysis involves breaking down the broadcaster's message content into keywords and generating the appropriate conversational message to be sent next. For example, if the broadcaster replies, "Today, I'll talk about movies," the server analyzes the keyword "movie" and generates the next message, "Which movie are you going to talk about?"

[0790] Processing by the terminal

[0791] When streaming begins, the device sends a start notification to the server. This notification is sent via an HTTP request or WebSocket. When a conversation message is sent from the server by the virtual listener, the device displays it in the chat box. The streamer checks the chat message through the device and enters a response. This response is sent back to the server and used for analysis. The device uses the streaming software OBS Studio or a dedicated streaming app.

[0792] User operation

[0793] The broadcaster reacts to the virtual listener's messages displayed in the chat box on the device, making the broadcaster feel as if they were having a conversation with a real viewer. For example, if the initial virtual listener's message is "Hello! What will you talk about today?", the broadcaster can respond with "I'll talk about movies today." This response is sent to the server, where it is analyzed in real time and the next conversation message is generated.

[0794] Specific examples

[0795] For example, when a broadcaster starts broadcasting, the server first generates a message for the virtual listener asking, "What are we talking about today?" and sends it to the terminal. If the broadcaster responds, "I'll talk about movies today," the server analyzes this response and generates and sends the next message, "Which movie are you talking about?" This allows for a natural chat with the virtual listener.

[0796] Prompt Sentence Examples

[0797] “If the streamer responds, ‘Today we’re going to talk about movies,’ suggest the next conversation message they should generate from their virtual listeners.”

[0798] This system allows streamers to engage with viewers through interactive broadcasts even at the early stages when the number of viewers is still small. As the number of real viewers increases, the system can also provide an environment where streamers can respond smoothly by displaying messages from both virtual listeners and real viewers appropriately.

[0799] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0800] Step 1:

[0801] The terminal sends a notification of the start of distribution to the server.

[0802] Specific operation: When the streamer presses the "Start Streaming" button in the streaming software (e.g., OBS Studio), the device sends a notification to the server via an HTTP request or WebSocket to start streaming.

[0803] Input: "Start distribution" operation of distribution software.

[0804] Output: A notification is sent to the server, and the server goes into a live broadcast state.

[0805] Step 2:

[0806] The server creates multiple virtual listener accounts.

[0807] Specific operation: The server uses the generative AI model to generate profile information (name, icon, etc.) for the virtual listener based on the settings and creates a virtual listener account.

[0808] Input: Notification from device that distribution has started.

[0809] Output: Multiple virtual listener accounts are created on the server.

[0810] Step 3:

[0811] The server generates an initial message.

[0812] What it does: Uses a generative AI model to generate an initial message from a virtual listener (e.g., "Hello! What will we talk about today?").

[0813] Input: Virtual listener account information.

[0814] Output: Initial message from the virtual listener.

[0815] Step 4:

[0816] The server sends the initial message for the virtual listener to the terminal.

[0817] Specific operation: The server generates an initial message and sends it to the terminal via a WebSocket or HTTP request.

[0818] Input: The initial message generated.

[0819] Output: The initial message of the virtual listener is sent to the terminal.

[0820] Step 5:

[0821] Displays messages received by the terminal in the chat box.

[0822] Specific operation: Displays the virtual listener's message in the chat box of the distribution software.

[0823] Input: The initial message sent by the server.

[0824] Output: The message that appears in the chat box.

[0825] Step 6:

[0826] The user (broadcaster) responds to the virtual listener's messages in the chat box.

[0827] Specific operation: The streamer types messages through the chat box and interacts with the virtual listeners. For example, the streamer types, "Today I'll talk about movies."

[0828] Input: The virtual listener's message and the broadcaster's response.

[0829] Output: The streamer's input message.

[0830] Step 7:

[0831] The device sends the distributor's response to the server.

[0832] Specific operation: Sends the broadcaster's input message to the server in real time.

[0833] Input: The streamer's input message.

[0834] Output: The publisher's response sent to the server.

[0835] Step 8:

[0836] The server parses the publisher's response.

[0837] How it works: The server uses the generative AI model to analyze the broadcaster's message by keyword and generate the next appropriate conversational message. For example, it analyzes the keyword "movie" and generates the message "Which movie are you talking about?"

[0838] Input: The broadcaster's response message.

[0839] Output: Analyzed keywords and the next conversation message.

[0840] Step 9:

[0841] The server generates the following conversation message:

[0842] Specific operation: Based on the analysis results, the generative AI model is used to generate the next conversation message.

[0843] Input: Parsed keywords.

[0844] Output: Next conversation message.

[0845] Step 10:

[0846] The server generates a message and sends it to the terminal.

[0847] Specific operation: The next generated conversation message is sent to the device via WebSocket or HTTP request.

[0848] Input: The next generated conversation message.

[0849] Output: The next conversation message sent to the terminal.

[0850] Step 11:

[0851] Displays messages received by the terminal in the chat box.

[0852] Specific operation: The following conversation message will be displayed in the chat box of the distribution software.

[0853] Input: The next conversation message sent from the server.

[0854] Output: The message that appears in the chat box.

[0855] The above processing steps allow the broadcaster to continue broadcasting while engaging in natural chat with the virtual listeners. This allows for interactive broadcasting even when the number of viewers is small from the beginning, contributing to the broadcaster's popularity.

[0856] (Application example 1)

[0857] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0858] In live streaming and content distribution, streamers can feel isolated in the early stages when there are few viewers. This can make interactive streaming difficult, leading to a decline in streamer motivation and a decline in viewer interest. Furthermore, there is a lack of effective means to properly manage messages for both real viewers and virtual listeners.

[0859] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0860] In this invention, the server includes means for generating conversational messages of a virtual user, means for transmitting the conversational messages of the virtual user to a terminal device, means for receiving responses from the user and generating the next conversational message based on the responses, means for analyzing the response messages and generating natural conversation based on prompt sentences using a generative AI model, and means for appropriately processing and displaying the messages of the virtual user and the messages of the actual viewers. This makes it easier for the broadcaster to maintain interactive conversations even in the early stages when the number of viewers is small, thereby improving the quality of the broadcast.

[0861] "Virtual user conversation messages" refers to messages of users virtually created using a generative AI model.

[0862] "Terminal device" refers to a device used by a broadcaster, including a smartphone or head-mounted display.

[0863] "User response" refers to a reply message entered by a broadcaster in response to a conversation message from a virtual user.

[0864] "Generative AI model" refers to an algorithm or system that uses generative AI technology to generate natural-sounding conversations.

[0865] A "prompt sentence" is the base text input to a generative AI model that determines the content of the next conversational message generated.

[0866] "Natural conversation" refers to messages generated by a generative AI model based on prompts that feel like real human conversations.

[0867] "Messages and actual viewer messages" refers to messages generated by virtual users or AI, as well as messages entered by viewers in real time.

[0868] "Means for appropriate processing and display" refers to the technology and algorithms that distinguish between messages from virtual users and messages from actual viewers and display them in a way that is easy for users to understand.

[0869] The present invention provides a system for generating conversational messages of virtual users and enabling broadcasters to broadcast interactive live content. The system operates using a server, a terminal device, and a broadcaster as its main components.

[0870] Server Processing

[0871] The server includes a program that generates conversation messages for virtual users. When a broadcast starts, the server starts this program and generates multiple accounts for the virtual users. As a result, messages from the virtual users are automatically displayed in the chat box as soon as the broadcast starts.

[0872] The server has a response message receiving function that analyzes the streamer's messages in real time. It breaks down the message entered by the streamer into keywords and generates the next conversation message based on these keywords. The generated message is input into a generative AI model using prompt sentences to create a natural conversation. For example, if the streamer enters "Today we will talk about movies," the server analyzes the keyword "movie" and generates the next message, "Which movie?"

[0873] Terminal processing

[0874] The terminal device is a device used by the broadcaster, such as a smartphone or head-mounted display. When broadcasting begins, the terminal device sends a start notification to the server. When a message from the virtual user is sent from the server, the terminal device has the function of displaying the message directly in the chat box.

[0875] The terminal device also has the role of transmitting the response message entered by the distributor to the server, which allows the server to analyze and generate the response message.

[0876] User operation

[0877] The broadcaster inputs responses to the virtual user's messages via a chat box displayed on the terminal device. These responses are sent to the server, where they are analyzed and used to generate the next conversation message. This system allows the broadcaster to continuously engage in natural interactions with the virtual user and the actual viewers.

[0878] Specific examples

[0879] When a broadcaster starts broadcasting, the server first generates a message from the virtual user saying, "What's the topic today?" and sends it to the terminal device. When the broadcaster responds, "Today I'll introduce a cooking recipe," the server analyzes this response, generates the next message, "That's great! What kind of cooking recipe is it?" and sends it again to the terminal device.

[0880] Prompt Sentence Examples

[0881] Streamer: Today I'll show you a recipe

[0882] Virtual users:

[0883] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0884] Step 1:

[0885] The server receives the notification that distribution has started.

[0886] Input: Notification of start of broadcasting from the broadcaster's terminal device.

[0887] Operation: The server receives the notification of the start of distribution and starts a program that generates conversational messages for the virtual user.

[0888] Output: Launch of the virtual user conversation message generation program.

[0889] Step 2:

[0890] The server generates an initial conversation message for the virtual user.

[0891] Input: Notification of start of distribution.

[0892] How it works: The server uses a generative AI model to generate an initial conversation message based on a configured prompt, such as "What's the topic today?"

[0893] Output: Initial virtual user conversation message.

[0894] Step 3:

[0895] The server sends an initial conversation message to the terminal device of the virtual user.

[0896] Input: The initial conversation message of the generated virtual user.

[0897] Operation: The server sends the generated initial conversation message to the broadcaster's terminal device.

[0898] Output: The initial conversation message of the virtual user displayed on the broadcaster's terminal device.

[0899] Step 4:

[0900] The broadcaster responds to the initial conversation message.

[0901] Input: The initial conversation message of the virtual user.

[0902] How it works: The broadcaster responds to the initial conversation message through the chat box on the terminal device, for example, typing, "Today I'll introduce you to a cooking recipe."

[0903] Output: The distributor's response message.

[0904] Step 5:

[0905] The terminal device sends the distributor's response message to the server.

[0906] Input: The response message entered by the broadcaster.

[0907] Operation: The terminal device sends the message entered by the broadcaster to the server.

[0908] Output: The publisher's response message sent to the server.

[0909] Step 6:

[0910] The server analyzes the distributor's response message.

[0911] Input: The broadcaster's response message.

[0912] How it works: The server breaks down the broadcaster's response message into keywords and creates a prompt using a generative AI model.

[0913] Output: Analysis results and prompt for the next conversation message.

[0914] Step 7:

[0915] The server generates the following conversation message:

[0916] Input: Prompt statement and analysis result.

[0917] How it works: The server uses a generative AI model to generate the next conversational message based on the prompt, for example, "That's great! What's the recipe?"

[0918] Output: The following conversation message is generated:

[0919] Step 8:

[0920] The server sends the following conversation message to the terminal device:

[0921] Input: The next generated conversation message.

[0922] Operation: The server again sends the next conversation message that has been generated to the broadcaster's terminal device.

[0923] Output: The next conversation message displayed on the terminal.

[0924] The above processing steps enable the broadcaster to broadcast live while maintaining natural interaction with the virtual user.

[0925] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0926] This invention is a system that combines a listener AI system that enables broadcasters to chat naturally with virtual users with an emotion engine that recognizes the user's emotions. This system operates with a server, a terminal, and a user as its main components, and the emotion engine can generate appropriate conversational messages according to the user's emotional state.

[0927] System Program

[0928] Server processing

[0929] First, we will explain the processing performed by the server. The server includes a program that generates conversation messages for virtual users. When a broadcaster starts broadcasting, this program is launched and multiple virtual listener accounts are generated based on the settings. This makes it possible to send messages to the chat box immediately after the broadcast begins.

[0930] The server also has an emotion engine built in that recognizes the emotional state of the streamer when analyzing their response. During this analysis phase, the content of the streamer's message is broken down into keywords, and the emotion engine determines their emotional state (e.g., happy, sad, or angry). The next appropriate conversational message to be sent is generated based on this emotional state. For example, if the streamer replies, "I'm not feeling very well today," the server analyzes the keyword "not feeling well" and their emotional state, and generates the message, "That sounds worrying. Are you okay?"

[0931] Processing by the terminal

[0932] Next, we will explain the processing performed by the terminal. When broadcasting starts, the broadcaster's terminal sends a start notification to the server. When a conversation message is sent from the server by the virtual listener, the terminal displays it in the chat box.

[0933] The streamer checks the chat message through their device and enters a response, which is then sent back to the server and used for analysis by the emotion engine.

[0934] User operation

[0935] The streamer reacts to messages from virtual listeners that appear in the chat box on their device, making the streamer feel as if they are having a conversation with a real viewer. For example, if an initial message from a virtual listener says, "Nice to meet you! What will you talk about today?", the streamer might respond, "I'm not feeling too great today, but I'll talk about movies."

[0936] In response, the server generates the next conversation message and sends it to the device, allowing the streamer to continue the conversation. This system allows streamers to deepen their interaction with viewers while having natural conversations with virtual listeners.

[0937] Specific examples

[0938] When a broadcaster starts broadcasting, the server first generates a message for the virtual listener asking, "What's the topic today?" and sends it to the terminal. When the broadcaster responds, "Today we'll talk about movies," the server analyzes this response with its emotion engine and determines the user's emotional state as "excitement" or "joy." It then generates the next message, "That's exciting! Which movie?" and sends it again to the broadcaster's terminal.

[0939] If the streamer responds, "I'm feeling down today, so I watched a quiet movie," the server analyzes the keyword "depressed" and the emotional state and generates a message saying, "That's terrible. What movie did you watch?"

[0940] In this way, the server, device, user, and emotion engine work together to allow the streamer to continue streaming while engaging in natural, emotionally appropriate conversations with virtual listeners. This system enables interactive streaming even at the initial stage, when there are only a few viewers, and helps streamers increase their popularity.

[0941] The processing flow will be explained below.

[0942] Step 1:

[0943] The server registers the broadcaster's user ID in the database. At this time, the server sets up the initial settings for the listener AI, including the number of virtual listener accounts, parameters required for message generation, and the emotion engine.

[0944] Step 2:

[0945] The distributor (user) presses the distribution start button. The distributor's terminal sends this distribution start information to the server.

[0946] Step 3:

[0947] The server receives a notification from the streamer that the stream has started, then launches the listener AI and generates multiple virtual listener accounts based on the streamer's settings.

[0948] Step 4:

[0949] The server uses an AI engine to generate an initial chat message. For example, it generates a message like "Nice to meet you! What will you talk about today?" and sends it to the broadcaster's device.

[0950] Step 5:

[0951] The streamer's device displays the chat message sent from the server in the chat box. The streamer checks this and enters a response such as "Today I'll be talking about game strategies."

[0952] Step 6:

[0953] The distributor's terminal transmits the answer entered by the distributor to the server.

[0954] Step 7:

[0955] The server receives the publisher's response, which is analyzed by the emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger).

[0956] Step 8:

[0957] The server generates the next conversation message based on the analysis results of the emotion engine. For example, if the broadcaster replies, "I'm a little sad today, but I'll talk about movies," the server recognizes the emotional state as "sad" and generates the message, "That's terrible. What movie did you see?" The generated message is then sent back to the broadcaster's device.

[0958] Step 9:

[0959] The streamer's device will again display the message sent from the server in the chat box. The streamer will continue to reply to the AI's messages.

[0960] Step 10:

[0961] The process loops from step 6 to step 9, and the conversation with the virtual listener continues until the streamer ends the chat. As the stream progresses, messages from real viewers are also displayed, allowing the streamer to respond to both.

[0962] In this way, the server, device, user, and emotion engine work together to allow the streamer to continue streaming while engaging in natural, emotionally appropriate conversations with virtual listeners. This system enables interactive streaming even at the initial stage, when there are only a few viewers, and helps streamers increase their popularity.

[0963] Example 2

[0964] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0965] In modern live streaming, a lack of interaction in the early stages when viewers are few can cause streamers to lose motivation. Furthermore, in order for streamers to enjoy natural conversations with viewers, they need to respond in real time, which places a heavy burden on them. A system that can solve these problems and enable streamers to enjoy natural conversations even with a small audience is needed.

[0966] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0967] In this invention, the server includes means for generating a conversational message of a virtual user, means for transmitting the conversational message of the virtual user to an information processing device, means for receiving a response from the broadcaster and analyzing the emotional state of the response, means for generating a next conversational message based on the emotional state, and means for transmitting the next conversational message again to the information processing device. This allows the broadcaster to enjoy natural conversation with the virtual listener even when there are only a few viewers.

[0968] A "virtual user" refers to an artificial participant that does not exist in reality but is generated by a computer program and behaves like a real user.

[0969] "Conversational messages" are messages sent and received in a chat or dialogue format, and refer to sentences or text exchanges generated by users or virtual users.

[0970] An "information processing device" is an electronic device used for calculations, data processing, communication, etc., and includes, for example, computers, smartphones, tablets, etc.

[0971] "Distributor" means an individual or entity that distributes live streaming or real-time video content over the Internet.

[0972] "Emotional state" refers to the emotional state of a person, such as joy, sadness, or anger, that is analyzed from text or audio.

[0973] "Emotion engine" refers to a software module that analyzes user input data and determines the emotions contained in the input.

[0974] "Generating means" refers to technical mechanisms or algorithms that automatically generate data or messages for a specific purpose.

[0975] "Transmitting means" refers to the technical mechanism by which data or messages are transferred from one information processing device to another device or server.

[0976] "Means of receiving" refers to the technical mechanisms for acquiring and processing data or messages sent from outside.

[0977] "Means of analysis" refers to the technical mechanisms used to analyze received data or messages and understand their content and meaning.

[0978] This invention is a listener AI system that enables broadcasters to chat naturally with virtual users, and is a system that combines an emotion engine that recognizes the user's emotions. The main components of this system are a server, a terminal, and a user.

[0979] Server Roles

[0980] The server first generates conversational messages for virtual users. When a broadcaster starts broadcasting, the server creates a virtual listener account and sends an initial message based on the broadcaster's settings to the broadcaster's device. The server is equipped with an emotion engine that analyzes the broadcaster's responses in real time and determines their emotional state.

[0981] As a specific example, if the broadcaster responds, "Today we'll talk about movies," the server analyzes this message with an emotion engine, generates the next conversational message, "That's exciting! Which movie?", and sends it to the device. This emotion engine can use, for example, a natural language processing library implemented in Python or a machine learning model (e.g., BERT).

[0982] Device Role

[0983] The streamer's device sends a notification to the server when the stream starts, and the streamer's response is sent to the server. The virtual listener's message sent from the server is displayed in the chat box on the device. The streamer can react through this chat box and enter their own message.

[0984] Specifically, if a streamer types, "I'm feeling down today, so I watched a quiet movie," this response is sent to the server and analyzed by the emotion engine. The server then generates the next message, "That's terrible. What movie did you watch?" and sends it to the device.

[0985] User Roles

[0986] The streamer reacts to messages from the virtual listeners displayed on their device, making it feel as if they are having a conversation with a real audience. This allows streamers to enjoy natural conversations with virtual listeners even when there are only a few viewers.

[0987] Example prompt sentences

[0988] Examples of prompts to input into the generative AI model could include specific instructions such as, "When the streamer says, 'Today I'll talk about movies,' generate a response from a virtual listener. Also, generate an appropriate response when the streamer is feeling down."

[0989] The feedback generated based on this prompt is concise and specific, making the streamer's conversational experience smoother and more natural.

[0990] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0991] Step 1:

[0992] The server starts a conversation message generating program for the virtual user when the broadcaster starts broadcasting.

[0993] Specific operation: A request to start streaming is sent from the streamer's terminal to the server. Upon receiving this request, the server begins creating virtual listener accounts.

[0994] Input: Request to start streaming

[0995] Output: Virtual listener account created

[0996] Step 2:

[0997] The server creates multiple virtual listener accounts based on the configuration.

[0998] Specific operation: The server selects a virtual listener from a list of pre-defined profile data (such as name, icon, attributes, etc.) randomly or based on specific conditions, generates an account, and registers it in the database.

[0999] Input: Profile data list

[1000] Output: Virtual listener account

[1001] Step 3:

[1002] The server generates an initial message using the generated virtual listener account and sends it to the terminal.

[1003] Specific behavior: Generates an initial message from the virtual listener (e.g., "What's the topic today?") and sends it to the streamer's chat box using a real-time communication protocol (e.g., WebSocket).

[1004] Input: Virtual listener account, initial message creation prompt

[1005] Output: Initial message

[1006] Step 4:

[1007] The broadcaster responds to the initial message that appears in the chat box on the device.

[1008] Specific operation: The broadcaster enters a message in the chat box (e.g., "Today I'll talk about movies") and presses the send button. This action sends the message to the server.

[1009] Input: Initial message, broadcaster response

[1010] Output: Streamer's response message

[1011] Step 5:

[1012] The server receives the responding message from the distributor and analyzes its contents using an emotion engine.

[1013] Specific operation: The publisher's response message is passed to the emotion engine, which uses natural language processing techniques (e.g., keyword extraction, emotion analysis) to determine the emotional state of the message. For example, the message is determined to be "movie" and "excited."

[1014] Input: Streamer's response message

[1015] Output: Emotional state (excitement, joy, etc.)

[1016] Step 6:

[1017] The server generates the next conversational message based on the emotional state.

[1018] Specific behavior: Using the emotional state and the generative AI model's prompt sentence (e.g., "Generate a virtual listener's response when the streamer responds that they're talking about movies"), generate the next conversational message (e.g., "That's fun! Which movie is it?").

[1019] Input: Emotional state, prompt sentence

[1020] Output: Next conversation message

[1021] Step 7:

[1022] The server transmits the next generated conversation message to the broadcaster's terminal.

[1023] Specific operation: The server sends the generated message to the device and displays it in the chat box, allowing the broadcaster to see the next message and respond.

[1024] Input: Next conversation message

[1025] Output: Conversation messages displayed on the terminal

[1026] (Application example 2)

[1027] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1028] In conventional streaming services, when the number of viewers is small at first, there is little interaction with the streamer, making it difficult to keep viewers engaged. It is also difficult for streamers to grasp viewers' emotions in real time and respond appropriately, resulting in a decline in the quality of interaction with viewers. To solve these issues, a system that makes the dialogue between viewers and streamers more natural and emotional is needed.

[1029] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for generating a conversation message of a virtual user, means for transmitting the conversation message of the virtual user to a user terminal, means for receiving a response from the user and generating a next conversation message based on the response, means for again transmitting the next conversation message to the user terminal, and means for recognizing the emotional state of the user and generating a conversation message based on the emotional state. This allows the broadcaster to have a conversation that is responsive to the viewers' emotions, making interactive broadcasting possible even when there are only a few viewers initially.

[1030] A "virtual user" is a non-existent user that is generated within a computer system and behaves in the same way as a real user.

[1031] "Conversational messages" are messages in text or audio format that a user or virtual user uses in a dialogue.

[1032] "User terminal" refers to a device such as a computer, smartphone, or tablet used by a broadcaster or viewer.

[1033] "Emotional state" refers to the user's feelings such as joy, sadness, anger, excitement, etc.

[1034] An "emotion engine" is software or hardware that recognizes a user's emotional state from their messages and actions.

[1035] "Analysis" is the process of interpreting, breaking down, and understanding user responses and input information.

[1036] A "creation means" is a method or device for performing a specific process or creating a specific result.

[1037] A "server" is a computer system that provides services to other computers and devices over a network.

[1038] In this invention, a system is constructed that allows a broadcaster to chat with a virtual user in a natural and emotional way. Each component of the system will be described in detail below.

[1039] Server processing

[1040] The server first generates a conversational message for the virtual user, using pre-prepared message templates and dialogue models. The generated message is then sent to the user's device. The server also receives and analyzes responses from the broadcaster. For analysis, it uses an emotion engine (for example, Hugging Face's transformers library or a specific emotion recognition model, mrm8488 / distilroberta-finetuned-emotion) to recognize the user's emotional state. Based on this emotional state, the server generates the next conversational message and sends it back to the user's device.

[1041] Processing by the terminal

[1042] At the start of a broadcast, the device sends a notification to the server and receives conversational messages from virtual listeners. These messages are displayed in the chat box, and the broadcaster types a response. This response is sent to the server for further analysis and message generation.

[1043] User operation

[1044] The streamer reacts to messages from the virtual listeners that appear in the chat box on their device, making the streamer feel as if they are having a conversation with a real viewer. For example, if a message from an initial virtual listener appears saying, "Hello! How are you doing today?", the streamer can respond with, "I saw a great movie today." Based on this response, the emotion engine recognizes the streamer's emotional state as "joy" and generates a message saying, "Thanks for sharing your wonderful story! Tell me more!"

[1045] Specific examples

[1046] When a broadcaster starts broadcasting, the server generates a message from the initial virtual listener, "What do you think about the recent news?", and sends it to the device. If the broadcaster responds with "The recent news is a little sad," the server analyzes this response with its emotion engine and determines the user's emotional state as "sad." It then generates a message saying, "That's a sad story...that must have been tough," and sends it back to the device.

[1047] Prompt Sentence Examples

[1048] In an emotion-aware live chat application, design an AI model to generate an appropriate response based on the streamer's input message, "I watched a great movie today."

[1049] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1050] Step 1:

[1051] The server receives a notification that distribution has started. When the broadcaster presses the start distribution button on the device, the device sends a start notification to the server. The server receives this notification and prepares to start distribution.

[1052] Step 2:

[1053] The server generates the initial conversation message for the virtual user. This is done using pre-prepared message templates and dialogue models. These generated messages are stored in the server. An example of an inserted template message is "Hello! How are you doing today?"

[1054] Step 3:

[1055] The server sends an initial conversation message to the user terminal of the virtual user that has been created. The message is transferred from the server to the user terminal as a data packet and displayed in a chat box on the terminal.

[1056] Step 4:

[1057] The user terminal displays the received message from the virtual user in a chat box on the screen, and the broadcaster visually confirms the displayed message.

[1058] Step 5:

[1059] The broadcaster inputs text to respond to the virtual user's message. The input text is temporarily stored on the terminal and then sent to the server. An example input might be, "I saw a fun movie today."

[1060] Step 6:

[1061] The server receives the response from the broadcaster and analyzes the message using the emotion engine, which uses Hugging Face's transformers library and the emotion analysis model mrm8488 / distilroberta-finetuned-emotion. The response text is analyzed to determine the associated emotional state (e.g., "joy").

[1062] Step 7:

[1063] The server generates a new conversational message based on the emotional state. This new message will match the broadcaster's emotional state. For example, if the analysis result is "joy," the server generates a message like "Thank you for sharing your wonderful story! Tell me more!"

[1064] Step 8:

[1065] The server then sends the newly generated conversation message to the user terminal again as a data packet, which is then displayed in the chat box of the user terminal.

[1066] Step 9:

[1067] The user terminal displays the new message received again in the chat box, allowing the broadcaster to confirm it and continue the interaction to enter the next response.

[1068] This allows the broadcaster to continue a natural and emotionally appropriate conversation with the virtual user.

[1069] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1070] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1071] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1072] [Fourth embodiment]

[1073] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1074] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1075] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1076] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1077] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1078] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1079] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1080] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1081] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1082] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1083] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1084] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1085] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1086] The present invention is a listener AI system that enables broadcasters to have natural chats with virtual users. The system operates with a server, terminals, and users as its main components.

[1087] System Program

[1088] Server processing

[1089] First, we will explain the processing performed by the server. The server includes a program that generates conversation messages for virtual users. The server starts this program when a broadcaster starts broadcasting, and generates multiple virtual listener accounts based on the settings. This makes it possible to send messages to the chat box immediately after the broadcast starts.

[1090] The server analyzes the responses received from the streamer's user device in real time. During this analysis phase, the streamer's message content is broken down into keywords and the appropriate conversation message to be sent next is generated. For example, if the streamer replies, "Today I'll talk about game strategies," the server analyzes the keyword "game strategies" and generates the next message: "That's fun! Which game?"

[1091] Processing by the terminal

[1092] Next, we will explain the processing performed by the terminal. When broadcasting starts, the broadcaster's terminal sends a start notification to the server. When a conversation message is sent from the server by the virtual listener, the terminal displays it in the chat box.

[1093] The broadcaster checks the chat message through the terminal and enters a response, which is then sent back to the server for analysis.

[1094] User operation

[1095] The streamer reacts to messages from the virtual listeners that appear in the chat box on their device, making the streamer feel as if they are having a conversation with a real viewer. For example, if the initial message from the virtual listener says, "Nice to meet you! What will we talk about today?", the streamer will respond, "Today I'll talk about game strategies."

[1096] In response, the server generates the next conversation message and sends it to the device, allowing the broadcaster to continue the conversation. As the broadcast progresses, even if the number of real viewers increases, the device will appropriately display messages from both virtual listeners and real viewers, allowing the broadcaster to respond to both.

[1097] Specific examples

[1098] When a broadcaster starts broadcasting, the server first generates a virtual listener message, "What are we talking about today?", and sends it to the terminal. When the broadcaster responds, "Today we'll talk about movies," the server analyzes this response and generates and sends the next message, "Which movie?"

[1099] In this way, the server, terminals, and users work together to allow the broadcaster to continue broadcasting while engaging in natural chats with virtual listeners. This system enables interactive broadcasting even at the initial stage when there are only a few viewers, helping broadcasters to increase their popularity.

[1100] The processing flow will be explained below.

[1101] Step 1:

[1102] The server registers the broadcaster's user ID in the database. At this time, it also sets the number of virtual listener accounts and parameters required for message generation as the initial settings for the listener AI.

[1103] Step 2:

[1104] The distributor (user) presses the distribution start button. The distributor's terminal sends this distribution start information to the server.

[1105] Step 3:

[1106] The server receives a notification from the streamer that the stream has started, then launches the listener AI and generates multiple virtual listener accounts based on the streamer's settings.

[1107] Step 4:

[1108] The server uses an AI engine to generate an initial chat message. For example, it generates a message like "Nice to meet you! What will you talk about today?" and sends it to the broadcaster's device.

[1109] Step 5:

[1110] The streamer's device displays the chat message sent from the server in the chat box. The streamer checks this and enters a response such as "Today I'll be talking about game strategies."

[1111] Step 6:

[1112] The distributor's terminal transmits the answer entered by the distributor to the server.

[1113] Step 7:

[1114] The server receives the streamer's response and uses an AI engine to analyze it and extract the keyword "game strategy."

[1115] Step 8:

[1116] The server generates the next conversation message based on the analysis result. For example, it generates the message "That's fun! Which game is it?" and sends it to the broadcaster terminal again.

[1117] Step 9:

[1118] The streamer's device will again display the message sent from the server in the chat box. The streamer will continue to reply to the AI's messages.

[1119] Step 10:

[1120] The process loops from step 6 to step 9, and the conversation with the virtual listener continues until the streamer ends the chat. As the stream progresses, messages from real viewers are also displayed, allowing the streamer to respond to both.

[1121] This allows streamers to have interactive broadcasts through conversations with virtual listeners, deepening their bond with their viewers while continuing to stream.

[1122] Example 1

[1123] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1124] When live streamers have few viewers in the early stages of a stream, they may experience little interaction with viewers, resulting in a lack of excitement. As a result, streamers often struggle to attract viewers' interest, and their streams can feel monotonous until the number of viewers increases. Furthermore, with few chat messages from viewers, streamers may be unsure of how to move the conversation forward. To solve these problems, a system is needed that allows streamers to have natural conversations with virtual viewers and realize interactive streams.

[1125] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1126] In this invention, the server includes: a means for generating conversational messages for virtual users; a means for transmitting the conversational messages for virtual users to a user terminal; a means for receiving a user's response and generating a next conversational message based on the response; a means for transmitting the next conversational message to the user terminal; a means for generating multiple virtual listener accounts based on a broadcast start notification received from the user terminal; a means for analyzing the generated conversational messages in real time and breaking down the broadcaster's input message into keywords; and a means for executing the above-mentioned means using a generative AI model. This allows virtual listener accounts to be generated from the early stages of broadcasting, enabling natural conversations to take place, enabling the broadcaster to always broadcast lively and interactively. Furthermore, by analyzing the broadcaster's messages in real time and generating appropriate conversational messages, smooth conversation progression is ensured even as the number of viewers increases.

[1127] A "virtual user" is not a real person, but a character generated by the system and designed to interact with the user.

[1128] "Conversation Message" means text-based communication content sent by a virtual user or user.

[1129] "User Terminal" means a device used by a broadcaster, such as a computer, tablet, or smartphone, that includes software for sending, receiving, and broadcasting chat messages.

[1130] A "generative AI model" refers to an artificial intelligence algorithm that performs natural language processing, and is a technology for generating interactive messages based on user input messages.

[1131] A "prompt" is an instruction given to a generative AI model that provides it with the information it needs to generate an appropriate conversational message.

[1132] "Analysis" is the process of breaking down a received message to understand or make sense of its content.

[1133] A "distribution start notification" is a signal sent to the server when a distributor starts distribution, and serves as a trigger to start up each part of the system.

[1134] A "virtual listener account" is an account created by the system to act as a virtual user.

[1135] "Real-time" refers to near-instant processing or response, with minimal delay.

[1136] "Keywords" refer to important words or phrases in a user's message, and are the building blocks for generating the next conversational message.

[1137] The present invention provides a system that supports broadcasters in progressing with broadcasts while engaging in natural chat with virtual users. This system operates with a server, a terminal, and a user as its main components. Specific embodiments of this system are described below.

[1138] Server processing

[1139] The server includes a program that generates conversational messages from virtual users. This program uses a generative AI model that performs natural language processing. Specifically, OpenAI's GPT-4 can be used. The server launches this program when the streamer begins streaming, and generates multiple virtual listener accounts based on the settings. This allows messages from virtual users to be streamed in the chat box immediately after the stream begins.

[1140] The server also analyzes responses received from the broadcaster's user terminal in real time. This analysis involves breaking down the broadcaster's message content into keywords and generating the appropriate conversational message to be sent next. For example, if the broadcaster replies, "Today, I'll talk about movies," the server analyzes the keyword "movie" and generates the next message, "Which movie are you going to talk about?"

[1141] Processing by the terminal

[1142] When streaming begins, the device sends a start notification to the server. This notification is sent via an HTTP request or WebSocket. When a conversation message is sent from the server by the virtual listener, the device displays it in the chat box. The streamer checks the chat message through the device and enters a response. This response is sent back to the server and used for analysis. The device uses the streaming software OBS Studio or a dedicated streaming app.

[1143] User operation

[1144] The broadcaster reacts to the virtual listener's messages displayed in the chat box on the device, making the broadcaster feel as if they were having a conversation with a real viewer. For example, if the initial virtual listener's message is "Hello! What will you talk about today?", the broadcaster can respond with "I'll talk about movies today." This response is sent to the server, where it is analyzed in real time and the next conversation message is generated.

[1145] Specific examples

[1146] For example, when a broadcaster starts broadcasting, the server first generates a message for the virtual listener asking, "What are we talking about today?" and sends it to the terminal. If the broadcaster responds, "I'll talk about movies today," the server analyzes this response and generates and sends the next message, "Which movie are you talking about?" This allows for a natural chat with the virtual listener.

[1147] Prompt Sentence Examples

[1148] “If the streamer responds, ‘Today we’re going to talk about movies,’ suggest the next conversation message they should generate from their virtual listeners.”

[1149] This system allows streamers to engage with viewers through interactive broadcasts even at the early stages when the number of viewers is still small. As the number of real viewers increases, the system can also provide an environment where streamers can respond smoothly by displaying messages from both virtual listeners and real viewers appropriately.

[1150] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1151] Step 1:

[1152] The terminal sends a notification of the start of distribution to the server.

[1153] Specific operation: When the streamer presses the "Start Streaming" button in the streaming software (e.g., OBS Studio), the device sends a notification to the server via an HTTP request or WebSocket to start streaming.

[1154] Input: "Start distribution" operation of distribution software.

[1155] Output: A notification is sent to the server, and the server goes into a live broadcast state.

[1156] Step 2:

[1157] The server creates multiple virtual listener accounts.

[1158] Specific operation: The server uses the generative AI model to generate profile information (name, icon, etc.) for the virtual listener based on the settings and creates a virtual listener account.

[1159] Input: Notification from device that distribution has started.

[1160] Output: Multiple virtual listener accounts are created on the server.

[1161] Step 3:

[1162] The server generates an initial message.

[1163] What it does: Uses a generative AI model to generate an initial message from a virtual listener (e.g., "Hello! What will we talk about today?").

[1164] Input: Virtual listener account information.

[1165] Output: Initial message from the virtual listener.

[1166] Step 4:

[1167] The server sends the initial message for the virtual listener to the terminal.

[1168] Specific operation: The server generates an initial message and sends it to the terminal via a WebSocket or HTTP request.

[1169] Input: The initial message generated.

[1170] Output: The initial message of the virtual listener is sent to the terminal.

[1171] Step 5:

[1172] Displays messages received by the terminal in the chat box.

[1173] Specific operation: Displays the virtual listener's message in the chat box of the distribution software.

[1174] Input: The initial message sent by the server.

[1175] Output: The message that appears in the chat box.

[1176] Step 6:

[1177] The user (broadcaster) responds to the virtual listener's messages in the chat box.

[1178] Specific operation: The streamer types messages through the chat box and interacts with the virtual listeners. For example, the streamer types, "Today I'll talk about movies."

[1179] Input: The virtual listener's message and the broadcaster's response.

[1180] Output: The streamer's input message.

[1181] Step 7:

[1182] The device sends the distributor's response to the server.

[1183] Specific operation: Sends the broadcaster's input message to the server in real time.

[1184] Input: The streamer's input message.

[1185] Output: The publisher's response sent to the server.

[1186] Step 8:

[1187] The server parses the publisher's response.

[1188] How it works: The server uses the generative AI model to analyze the broadcaster's message by keyword and generate the next appropriate conversational message. For example, it analyzes the keyword "movie" and generates the message "Which movie are you talking about?"

[1189] Input: The broadcaster's response message.

[1190] Output: Analyzed keywords and the next conversation message.

[1191] Step 9:

[1192] The server generates the following conversation message:

[1193] Specific operation: Based on the analysis results, the generative AI model is used to generate the next conversation message.

[1194] Input: Parsed keywords.

[1195] Output: Next conversation message.

[1196] Step 10:

[1197] The server generates a message and sends it to the terminal.

[1198] Specific operation: The next generated conversation message is sent to the device via WebSocket or HTTP request.

[1199] Input: The next generated conversation message.

[1200] Output: The next conversation message sent to the terminal.

[1201] Step 11:

[1202] Displays messages received by the terminal in the chat box.

[1203] Specific operation: The following conversation message will be displayed in the chat box of the distribution software.

[1204] Input: The next conversation message sent from the server.

[1205] Output: The message that appears in the chat box.

[1206] The above processing steps allow the broadcaster to continue broadcasting while engaging in natural chat with the virtual listeners. This allows for interactive broadcasting even when the number of viewers is small from the beginning, contributing to the broadcaster's popularity.

[1207] (Application example 1)

[1208] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1209] In live streaming and content distribution, streamers can feel isolated in the early stages when there are few viewers. This can make interactive streaming difficult, leading to a decline in streamer motivation and a decline in viewer interest. Furthermore, there is a lack of effective means to properly manage messages for both real viewers and virtual listeners.

[1210] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1211] In this invention, the server includes means for generating conversational messages of a virtual user, means for transmitting the conversational messages of the virtual user to a terminal device, means for receiving responses from the user and generating the next conversational message based on the responses, means for analyzing the response messages and generating natural conversation based on prompt sentences using a generative AI model, and means for appropriately processing and displaying the messages of the virtual user and the messages of the actual viewers. This makes it easier for the broadcaster to maintain interactive conversations even in the early stages when the number of viewers is small, thereby improving the quality of the broadcast.

[1212] "Virtual user conversation messages" refers to messages of users virtually created using a generative AI model.

[1213] "Terminal device" refers to a device used by a broadcaster, including a smartphone or head-mounted display.

[1214] "User response" refers to a reply message entered by a broadcaster in response to a conversation message from a virtual user.

[1215] "Generative AI model" refers to an algorithm or system that uses generative AI technology to generate natural-sounding conversations.

[1216] A "prompt sentence" is the base text input to a generative AI model that determines the content of the next conversational message generated.

[1217] "Natural conversation" refers to messages generated by a generative AI model based on prompts that feel like real human conversations.

[1218] "Messages and actual viewer messages" refers to messages generated by virtual users or AI, as well as messages entered by viewers in real time.

[1219] "Means for appropriate processing and display" refers to the technology and algorithms that distinguish between messages from virtual users and messages from actual viewers and display them in a way that is easy for users to understand.

[1220] The present invention provides a system for generating conversational messages of virtual users and enabling broadcasters to broadcast interactive live content. The system operates using a server, a terminal device, and a broadcaster as its main components.

[1221] Server Processing

[1222] The server includes a program that generates conversation messages for virtual users. When a broadcast starts, the server starts this program and generates multiple accounts for the virtual users. As a result, messages from the virtual users are automatically displayed in the chat box as soon as the broadcast starts.

[1223] The server has a response message receiving function that analyzes the streamer's messages in real time. It breaks down the message entered by the streamer into keywords and generates the next conversation message based on these keywords. The generated message is input into a generative AI model using prompt sentences to create a natural conversation. For example, if the streamer enters "Today we will talk about movies," the server analyzes the keyword "movie" and generates the next message, "Which movie?"

[1224] Terminal processing

[1225] The terminal device is a device used by the broadcaster, such as a smartphone or head-mounted display. When broadcasting begins, the terminal device sends a start notification to the server. When a message from the virtual user is sent from the server, the terminal device has the function of displaying the message directly in the chat box.

[1226] The terminal device also has the role of transmitting the response message entered by the distributor to the server, which allows the server to analyze and generate the response message.

[1227] User operation

[1228] The broadcaster inputs responses to the virtual user's messages via a chat box displayed on the terminal device. These responses are sent to the server, where they are analyzed and used to generate the next conversation message. This system allows the broadcaster to continuously engage in natural interactions with the virtual user and the actual viewers.

[1229] Specific examples

[1230] When a broadcaster starts broadcasting, the server first generates a message from the virtual user saying, "What's the topic today?" and sends it to the terminal device. When the broadcaster responds, "Today I'll introduce a cooking recipe," the server analyzes this response, generates the next message, "That's great! What kind of cooking recipe is it?" and sends it again to the terminal device.

[1231] Prompt Sentence Examples

[1232] Streamer: Today I'll show you a recipe

[1233] Virtual users:

[1234] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1235] Step 1:

[1236] The server receives the notification that distribution has started.

[1237] Input: Notification of start of broadcasting from the broadcaster's terminal device.

[1238] Operation: The server receives the notification of the start of distribution and starts a program that generates conversational messages for the virtual user.

[1239] Output: Launch of the virtual user conversation message generation program.

[1240] Step 2:

[1241] The server generates an initial conversation message for the virtual user.

[1242] Input: Notification of start of distribution.

[1243] How it works: The server uses a generative AI model to generate an initial conversation message based on a configured prompt, such as "What's the topic today?"

[1244] Output: Initial virtual user conversation message.

[1245] Step 3:

[1246] The server sends an initial conversation message to the terminal device of the virtual user.

[1247] Input: The initial conversation message of the generated virtual user.

[1248] Operation: The server sends the generated initial conversation message to the broadcaster's terminal device.

[1249] Output: The initial conversation message of the virtual user displayed on the broadcaster's terminal device.

[1250] Step 4:

[1251] The broadcaster responds to the initial conversation message.

[1252] Input: The initial conversation message of the virtual user.

[1253] How it works: The broadcaster responds to the initial conversation message through the chat box on the terminal device, for example, typing, "Today I'll introduce you to a cooking recipe."

[1254] Output: The distributor's response message.

[1255] Step 5:

[1256] The terminal device sends the distributor's response message to the server.

[1257] Input: The response message entered by the broadcaster.

[1258] Operation: The terminal device sends the message entered by the broadcaster to the server.

[1259] Output: The publisher's response message sent to the server.

[1260] Step 6:

[1261] The server analyzes the distributor's response message.

[1262] Input: The broadcaster's response message.

[1263] How it works: The server breaks down the broadcaster's response message into keywords and creates a prompt using a generative AI model.

[1264] Output: Analysis results and prompt for the next conversation message.

[1265] Step 7:

[1266] The server generates the following conversation message:

[1267] Input: Prompt statement and analysis result.

[1268] How it works: The server uses a generative AI model to generate the next conversational message based on the prompt, for example, "That's great! What's the recipe?"

[1269] Output: The following conversation message is generated:

[1270] Step 8:

[1271] The server sends the following conversation message to the terminal device:

[1272] Input: The next generated conversation message.

[1273] Operation: The server again sends the next conversation message that has been generated to the broadcaster's terminal device.

[1274] Output: The next conversation message displayed on the terminal.

[1275] The above processing steps enable the broadcaster to broadcast live while maintaining natural interaction with the virtual user.

[1276] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1277] This invention is a system that combines a listener AI system that enables broadcasters to chat naturally with virtual users with an emotion engine that recognizes the user's emotions. This system operates with a server, a terminal, and a user as its main components, and the emotion engine can generate appropriate conversational messages according to the user's emotional state.

[1278] System Program

[1279] Server processing

[1280] First, we will explain the processing performed by the server. The server includes a program that generates conversation messages for virtual users. When a broadcaster starts broadcasting, this program is launched and multiple virtual listener accounts are generated based on the settings. This makes it possible to send messages to the chat box immediately after the broadcast begins.

[1281] The server also has an emotion engine built in that recognizes the emotional state of the streamer when analyzing their response. During this analysis phase, the content of the streamer's message is broken down into keywords, and the emotion engine determines their emotional state (e.g., happy, sad, or angry). The next appropriate conversational message to be sent is generated based on this emotional state. For example, if the streamer replies, "I'm not feeling very well today," the server analyzes the keyword "not feeling well" and their emotional state, and generates the message, "That sounds worrying. Are you okay?"

[1282] Processing by the terminal

[1283] Next, we will explain the processing performed by the terminal. When broadcasting starts, the broadcaster's terminal sends a start notification to the server. When a conversation message is sent from the server by the virtual listener, the terminal displays it in the chat box.

[1284] The streamer checks the chat message through their device and enters a response, which is then sent back to the server and used for analysis by the emotion engine.

[1285] User operation

[1286] The streamer reacts to messages from virtual listeners that appear in the chat box on their device, making the streamer feel as if they are having a conversation with a real viewer. For example, if an initial message from a virtual listener says, "Nice to meet you! What will you talk about today?", the streamer might respond, "I'm not feeling too great today, but I'll talk about movies."

[1287] In response, the server generates the next conversation message and sends it to the device, allowing the streamer to continue the conversation. This system allows streamers to deepen their interaction with viewers while having natural conversations with virtual listeners.

[1288] Specific examples

[1289] When a broadcaster starts broadcasting, the server first generates a message for the virtual listener asking, "What's the topic today?" and sends it to the terminal. When the broadcaster responds, "Today we'll talk about movies," the server analyzes this response with its emotion engine and determines the user's emotional state as "excitement" or "joy." It then generates the next message, "That's exciting! Which movie?" and sends it again to the broadcaster's terminal.

[1290] If the streamer responds, "I'm feeling down today, so I watched a quiet movie," the server analyzes the keyword "depressed" and the emotional state and generates a message saying, "That's terrible. What movie did you watch?"

[1291] In this way, the server, device, user, and emotion engine work together to allow the streamer to continue streaming while engaging in natural, emotionally appropriate conversations with virtual listeners. This system enables interactive streaming even at the initial stage, when there are only a few viewers, and helps streamers increase their popularity.

[1292] The processing flow will be explained below.

[1293] Step 1:

[1294] The server registers the broadcaster's user ID in the database. At this time, the server sets up the initial settings for the listener AI, including the number of virtual listener accounts, parameters required for message generation, and the emotion engine.

[1295] Step 2:

[1296] The distributor (user) presses the distribution start button. The distributor's terminal sends this distribution start information to the server.

[1297] Step 3:

[1298] The server receives a notification from the streamer that the stream has started, then launches the listener AI and generates multiple virtual listener accounts based on the streamer's settings.

[1299] Step 4:

[1300] The server uses an AI engine to generate an initial chat message. For example, it generates a message like "Nice to meet you! What will you talk about today?" and sends it to the broadcaster's device.

[1301] Step 5:

[1302] The streamer's device displays the chat message sent from the server in the chat box. The streamer checks this and enters a response such as "Today I'll be talking about game strategies."

[1303] Step 6:

[1304] The distributor's terminal transmits the answer entered by the distributor to the server.

[1305] Step 7:

[1306] The server receives the publisher's response, which is analyzed by the emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger).

[1307] Step 8:

[1308] The server generates the next conversation message based on the analysis results of the emotion engine. For example, if the broadcaster replies, "I'm a little sad today, but I'll talk about movies," the server recognizes the emotional state as "sad" and generates the message, "That's terrible. What movie did you see?" The generated message is then sent back to the broadcaster's device.

[1309] Step 9:

[1310] The streamer's device will again display the message sent from the server in the chat box. The streamer will continue to reply to the AI's messages.

[1311] Step 10:

[1312] The process loops from step 6 to step 9, and the conversation with the virtual listener continues until the streamer ends the chat. As the stream progresses, messages from real viewers are also displayed, allowing the streamer to respond to both.

[1313] In this way, the server, device, user, and emotion engine work together to allow the streamer to continue streaming while engaging in natural, emotionally appropriate conversations with virtual listeners. This system enables interactive streaming even at the initial stage, when there are only a few viewers, and helps streamers increase their popularity.

[1314] Example 2

[1315] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1316] In modern live streaming, a lack of interaction in the early stages when viewers are few can cause streamers to lose motivation. Furthermore, in order for streamers to enjoy natural conversations with viewers, they need to respond in real time, which places a heavy burden on them. A system that can solve these problems and enable streamers to enjoy natural conversations even with a small audience is needed.

[1317] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1318] In this invention, the server includes means for generating a conversational message of a virtual user, means for transmitting the conversational message of the virtual user to an information processing device, means for receiving a response from the broadcaster and analyzing the emotional state of the response, means for generating a next conversational message based on the emotional state, and means for transmitting the next conversational message again to the information processing device. This allows the broadcaster to enjoy natural conversation with the virtual listener even when there are only a few viewers.

[1319] A "virtual user" refers to an artificial participant that does not exist in reality but is generated by a computer program and behaves like a real user.

[1320] "Conversational messages" are messages sent and received in a chat or dialogue format, and refer to sentences or text exchanges generated by users or virtual users.

[1321] An "information processing device" is an electronic device used for calculations, data processing, communication, etc., and includes, for example, computers, smartphones, tablets, etc.

[1322] "Distributor" means an individual or entity that distributes live streaming or real-time video content over the Internet.

[1323] "Emotional state" refers to the emotional state of a person, such as joy, sadness, or anger, that is analyzed from text or audio.

[1324] "Emotion engine" refers to a software module that analyzes user input data and determines the emotions contained in the input.

[1325] "Generating means" refers to technical mechanisms or algorithms that automatically generate data or messages for a specific purpose.

[1326] "Transmitting means" refers to the technical mechanism by which data or messages are transferred from one information processing device to another device or server.

[1327] "Means of receiving" refers to the technical mechanisms for acquiring and processing data or messages sent from outside.

[1328] "Means of analysis" refers to the technical mechanisms used to analyze received data or messages and understand their content and meaning.

[1329] This invention is a listener AI system that enables broadcasters to chat naturally with virtual users, and is a system that combines an emotion engine that recognizes the user's emotions. The main components of this system are a server, a terminal, and a user.

[1330] Server Roles

[1331] The server first generates conversational messages for virtual users. When a broadcaster starts broadcasting, the server creates a virtual listener account and sends an initial message based on the broadcaster's settings to the broadcaster's device. The server is equipped with an emotion engine that analyzes the broadcaster's responses in real time and determines their emotional state.

[1332] As a specific example, if the broadcaster responds, "Today we'll talk about movies," the server analyzes this message with an emotion engine, generates the next conversational message, "That's exciting! Which movie?", and sends it to the device. This emotion engine can use, for example, a natural language processing library implemented in Python or a machine learning model (e.g., BERT).

[1333] Device Role

[1334] The streamer's device sends a notification to the server when the stream starts, and the streamer's response is sent to the server. The virtual listener's message sent from the server is displayed in the chat box on the device. The streamer can react through this chat box and enter their own message.

[1335] Specifically, if a streamer types, "I'm feeling down today, so I watched a quiet movie," this response is sent to the server and analyzed by the emotion engine. The server then generates the next message, "That's terrible. What movie did you watch?" and sends it to the device.

[1336] User Roles

[1337] The streamer reacts to messages from the virtual listeners displayed on their device, making it feel as if they are having a conversation with a real audience. This allows streamers to enjoy natural conversations with virtual listeners even when there are only a few viewers.

[1338] Example prompt sentences

[1339] Examples of prompts to input into the generative AI model could include specific instructions such as, "When the streamer says, 'Today I'll talk about movies,' generate a response from a virtual listener. Also, generate an appropriate response when the streamer is feeling down."

[1340] The feedback generated based on this prompt is concise and specific, making the streamer's conversational experience smoother and more natural.

[1341] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1342] Step 1:

[1343] The server starts a conversation message generating program for the virtual user when the broadcaster starts broadcasting.

[1344] Specific operation: A request to start streaming is sent from the streamer's terminal to the server. Upon receiving this request, the server begins creating virtual listener accounts.

[1345] Input: Request to start streaming

[1346] Output: Virtual listener account created

[1347] Step 2:

[1348] The server creates multiple virtual listener accounts based on the configuration.

[1349] Specific operation: The server selects a virtual listener from a list of pre-defined profile data (such as name, icon, attributes, etc.) randomly or based on specific conditions, generates an account, and registers it in the database.

[1350] Input: Profile data list

[1351] Output: Virtual listener account

[1352] Step 3:

[1353] The server generates an initial message using the generated virtual listener account and sends it to the terminal.

[1354] Specific behavior: Generates an initial message from the virtual listener (e.g., "What's the topic today?") and sends it to the streamer's chat box using a real-time communication protocol (e.g., WebSocket).

[1355] Input: Virtual listener account, initial message creation prompt

[1356] Output: Initial message

[1357] Step 4:

[1358] The broadcaster responds to the initial message that appears in the chat box on the device.

[1359] Specific operation: The broadcaster enters a message in the chat box (e.g., "Today I'll talk about movies") and presses the send button. This action sends the message to the server.

[1360] Input: Initial message, broadcaster response

[1361] Output: Streamer's response message

[1362] Step 5:

[1363] The server receives the responding message from the distributor and analyzes its contents using an emotion engine.

[1364] Specific operation: The publisher's response message is passed to the emotion engine, which uses natural language processing techniques (e.g., keyword extraction, emotion analysis) to determine the emotional state of the message. For example, the message is determined to be "movie" and "excited."

[1365] Input: Streamer's response message

[1366] Output: Emotional state (excitement, joy, etc.)

[1367] Step 6:

[1368] The server generates the next conversational message based on the emotional state.

[1369] Specific behavior: Using the emotional state and the generative AI model's prompt sentence (e.g., "Generate a virtual listener's response when the streamer responds that they're talking about movies"), generate the next conversational message (e.g., "That's fun! Which movie is it?").

[1370] Input: Emotional state, prompt sentence

[1371] Output: Next conversation message

[1372] Step 7:

[1373] The server transmits the next generated conversation message to the broadcaster's terminal.

[1374] Specific operation: The server sends the generated message to the device and displays it in the chat box, allowing the broadcaster to see the next message and respond.

[1375] Input: Next conversation message

[1376] Output: Conversation messages displayed on the terminal

[1377] (Application example 2)

[1378] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1379] In conventional streaming services, when the number of viewers is small at first, there is little interaction with the streamer, making it difficult to keep viewers engaged. It is also difficult for streamers to grasp viewers' emotions in real time and respond appropriately, resulting in a decline in the quality of interaction with viewers. To solve these issues, a system that makes the dialogue between viewers and streamers more natural and emotional is needed.

[1380] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for generating a conversation message of a virtual user, means for transmitting the conversation message of the virtual user to a user terminal, means for receiving a response from the user and generating a next conversation message based on the response, means for again transmitting the next conversation message to the user terminal, and means for recognizing the emotional state of the user and generating a conversation message based on the emotional state. This allows the broadcaster to have a conversation that is responsive to the viewers' emotions, making interactive broadcasting possible even when there are only a few viewers initially.

[1381] A "virtual user" is a non-existent user that is generated within a computer system and behaves in the same way as a real user.

[1382] "Conversational messages" are messages in text or audio format that a user or virtual user uses in a dialogue.

[1383] "User terminal" refers to a device such as a computer, smartphone, or tablet used by a broadcaster or viewer.

[1384] "Emotional state" refers to the user's feelings such as joy, sadness, anger, excitement, etc.

[1385] An "emotion engine" is software or hardware that recognizes a user's emotional state from their messages and actions.

[1386] "Analysis" is the process of interpreting, breaking down, and understanding user responses and input information.

[1387] A "creation means" is a method or device for performing a specific process or creating a specific result.

[1388] A "server" is a computer system that provides services to other computers and devices over a network.

[1389] In this invention, a system is constructed that allows a broadcaster to chat with a virtual user in a natural and emotional way. Each component of the system will be described in detail below.

[1390] Server processing

[1391] The server first generates a conversational message for the virtual user, using pre-prepared message templates and dialogue models. The generated message is then sent to the user's device. The server also receives and analyzes responses from the broadcaster. For analysis, it uses an emotion engine (for example, Hugging Face's transformers library or a specific emotion recognition model, mrm8488 / distilroberta-finetuned-emotion) to recognize the user's emotional state. Based on this emotional state, the server generates the next conversational message and sends it back to the user's device.

[1392] Processing by the terminal

[1393] At the start of a broadcast, the device sends a notification to the server and receives conversational messages from virtual listeners. These messages are displayed in the chat box, and the broadcaster types a response. This response is sent to the server for further analysis and message generation.

[1394] User operation

[1395] The streamer reacts to messages from the virtual listeners that appear in the chat box on their device, making the streamer feel as if they are having a conversation with a real viewer. For example, if a message from an initial virtual listener appears saying, "Hello! How are you doing today?", the streamer can respond with, "I saw a great movie today." Based on this response, the emotion engine recognizes the streamer's emotional state as "joy" and generates a message saying, "Thanks for sharing your wonderful story! Tell me more!"

[1396] Specific examples

[1397] When a broadcaster starts broadcasting, the server generates a message from the initial virtual listener, "What do you think about the recent news?", and sends it to the device. If the broadcaster responds with "The recent news is a little sad," the server analyzes this response with its emotion engine and determines the user's emotional state as "sad." It then generates a message saying, "That's a sad story...that must have been tough," and sends it back to the device.

[1398] Prompt Sentence Examples

[1399] In an emotion-aware live chat application, design an AI model to generate an appropriate response based on the streamer's input message, "I watched a great movie today."

[1400] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1401] Step 1:

[1402] The server receives a notification that distribution has started. When the broadcaster presses the start distribution button on the device, the device sends a start notification to the server. The server receives this notification and prepares to start distribution.

[1403] Step 2:

[1404] The server generates the initial conversation message for the virtual user. This is done using pre-prepared message templates and dialogue models. These generated messages are stored in the server. An example of an inserted template message is "Hello! How are you doing today?"

[1405] Step 3:

[1406] The server sends an initial conversation message to the user terminal of the virtual user that has been created. The message is transferred from the server to the user terminal as a data packet and displayed in a chat box on the terminal.

[1407] Step 4:

[1408] The user terminal displays the received message from the virtual user in a chat box on the screen, and the broadcaster visually confirms the displayed message.

[1409] Step 5:

[1410] The broadcaster inputs text to respond to the virtual user's message. The input text is temporarily stored on the terminal and then sent to the server. An example input might be, "I saw a fun movie today."

[1411] Step 6:

[1412] The server receives the response from the broadcaster and analyzes the message using the emotion engine, which uses Hugging Face's transformers library and the emotion analysis model mrm8488 / distilroberta-finetuned-emotion. The response text is analyzed to determine the associated emotional state (e.g., "joy").

[1413] Step 7:

[1414] The server generates a new conversational message based on the emotional state. This new message will match the broadcaster's emotional state. For example, if the analysis result is "joy," the server generates a message like "Thank you for sharing your wonderful story! Tell me more!"

[1415] Step 8:

[1416] The server then sends the newly generated conversation message to the user terminal again as a data packet, which is then displayed in the chat box of the user terminal.

[1417] Step 9:

[1418] The user terminal displays the new message received again in the chat box, allowing the broadcaster to confirm it and continue the interaction to enter the next response.

[1419] This allows the broadcaster to continue a natural and emotionally appropriate conversation with the virtual user.

[1420] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1421] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1422] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1423] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1424] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1425] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1426] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1427] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1428] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1429] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1430] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1431] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1432] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1433] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1434] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1435] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1436] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1437] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1438] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1439] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1440] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1441] The following is further disclosed regarding the above embodiment.

[1442] (Claim 1)

[1443] means for generating conversational messages of virtual users;

[1444] means for transmitting a conversation message of the virtual user to a user terminal;

[1445] means for receiving a user response and generating a next conversational message based on the response;

[1446] means for transmitting the next conversation message to the user terminal again;

[1447] A system including:

[1448] (Claim 2)

[1449] 10. The system of claim 1, further comprising means for generating an initial conversational message of the virtual user.

[1450] (Claim 3)

[1451] 10. The system of claim 1, further comprising means for analyzing the user's response and generating a next conversational message based on the analysis result.

[1452] "Example 1"

[1453] (Claim 1)

[1454] means for generating conversational messages of virtual users;

[1455] means for transmitting a conversation message of the virtual user to a user terminal;

[1456] means for receiving a user response and generating a next conversational message based on the response;

[1457] means for transmitting the next conversation message to the user terminal again;

[1458] means for generating a plurality of virtual listener accounts based on a distribution start notification received from a user terminal;

[1459] A means for analyzing the generated conversation messages in real time and breaking down the broadcaster's input message into keywords;

[1460] means for performing the above-mentioned means using a generative AI model;

[1461] A system including:

[1462] (Claim 2)

[1463] 10. The system of claim 1, further comprising means for generating an initial conversational message of the virtual user.

[1464] (Claim 3)

[1465] 10. The system of claim 1, further comprising means for analyzing the user's response and generating a next conversational message based on the analysis result.

[1466] "Application Example 1"

[1467] (Claim 1)

[1468] means for generating conversational messages of the virtual user;

[1469] means for transmitting a conversation message of the virtual user to a terminal device;

[1470] means for receiving a user response and generating a next conversational message based on the response;

[1471] means for transmitting the next conversation message to the terminal device again;

[1472] A means for analyzing the response message and generating a natural conversation based on the prompt sentence using a generative AI model;

[1473] means for appropriately processing and displaying the virtual user's messages and the actual viewer's messages;

[1474] A system including:

[1475] (Claim 2)

[1476] 10. The system of claim 1, further comprising means for generating an initial conversation message of the virtual user.

[1477] (Claim 3)

[1478] 2. The system according to claim 1, further comprising means for analyzing the user's response and generating a next conversation message based on the analysis result.

[1479] "Example 2: Combining Emotion Engines"

[1480] (Claim 1)

[1481] means for generating conversational messages of virtual users;

[1482] means for transmitting a conversation message of the virtual user to an information processing device;

[1483] means for receiving a response from the publisher and analyzing the emotional state of the response;

[1484] means for generating a next conversational message based on said emotional state;

[1485] means for transmitting the next conversation message to the information processing device again;

[1486] A system including:

[1487] (Claim 2)

[1488] 10. The system of claim 1, further comprising means for generating an initial conversational message of the virtual user.

[1489] (Claim 3)

[1490] 2. The system according to claim 1, further comprising means for analyzing the broadcaster's response and generating a next conversation message based on the analysis result.

[1491] "Application example 2 when combining emotion engines"

[1492] (Claim 1)

[1493] means for generating conversational messages of virtual users;

[1494] means for transmitting a conversation message of the virtual user to a user terminal;

[1495] means for receiving a user response and generating a next conversational message based on the response;

[1496] means for transmitting the next conversation message to the user terminal again;

[1497] means for recognizing an emotional state of a user and generating a conversational message based on the emotional state;

[1498] A system including:

[1499] (Claim 2)

[1500] 10. The system of claim 1, further comprising means for generating an initial conversational message of the virtual user.

[1501] (Claim 3)

[1502] 10. The system of claim 1, further comprising means for analyzing the user's response and generating a next conversational message based on the analysis result. [Explanation of symbols]

[1503] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for generating conversational messages of virtual users; means for transmitting a conversation message of the virtual user to a user terminal; means for receiving a user response and generating a next conversational message based on the response; means for transmitting the next conversation message to the user terminal again; A system including:

2. 2. The system of claim 1, further comprising means for generating an initial conversational message for the virtual user.

3. 2. The system of claim 1, further comprising means for analyzing the user's response and generating a next conversational message based on the analysis result.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A