Servers and Computer Programs

The live streaming system addresses the need for realistic training by using machine learning models to create diverse virtual viewers, allowing streamers to practice and improve their interaction skills within a simulated environment that replicates real-time viewer reactions.

JP7673925B1Active Publication Date: 2025-05-0917LIVE JAPAN INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024157030
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2025-05-09
Estimated Expiration
2044-09-10

AI Technical Summary

Technical Problem

Existing live streaming technologies lack the ability to create realistic and interactive environments for streamers to practice and improve their skills, particularly in simulating interactions with diverse viewer reactions.

Method used

A live streaming system that utilizes multiple machine learning models to create virtual viewers with various personalities and interests, allowing streamers to practice their interactions in a simulated environment that mimics real-time reactions from different types of viewers.

Benefits of technology

Enables streamers to hone their skills in responding to different viewer reactions, improving their engagement and content quality, and provides a realistic training environment that enhances their overall streaming performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007673925000001_ABST
    Figure 0007673925000001_ABST
Patent Text Reader

Abstract

Applying machine learning to live streaming to create new features or improve existing ones. [Solution] The server is equipped with: a means for receiving data including the words and actions of a broadcaster via a network from a terminal of the broadcaster of a live broadcast in which the broadcaster and multiple viewers, including virtual viewers realized by a machine learning model, participate; a means for acquiring the reaction output by a machine learning model that takes the words and actions of the broadcaster as input and outputs the reaction that a viewer with set attributes would make; and a means for transmitting data for realizing the reaction to the terminal via the network.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a server and a computer program. [Background technology]

[0002] With the development of IT technology, the way information is exchanged has also changed. In the Showa era, one-way information transmission such as newspapers and television was the norm. In the Heisei era, mobile phones and personal computers became widespread and Internet communication speeds improved significantly, leading to the rise of instant two-way communication services such as chat services, and on-demand video distribution services became more popular as storage costs decreased. Now, in the Reiwa era, with the increasing functionality of smartphones and further improvements in network speeds such as 5G, services that enable real-time communication through video, especially live streaming services, are rapidly gaining recognition. Live streaming services are seeing an increase in users, especially among young people, as they allow everyone to share the same fun time even when they are far apart.

[0003] Patent Document 1 discloses a technology that uses machine learning to realize a virtual distributor. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent No. 7497002 [Non-patent literature]

[0005] [Non-Patent Document 1] "How to give ChatGPT the personality of King Gilgamesh?", Takayuki Fukatsu, URL: https: / / note.com / fladdict / n / neff2e9d52224 Summary of the Invention [Problem to be solved by the invention]

[0006] By applying machine learning to live streaming, it is possible to implement various functions that were previously impossible or impractical. The technology described in Patent Document 1 is just one example of such functions.

[0007] The present disclosure has been made in consideration of these problems, and its purpose is to create new functions or improve existing functions by applying machine learning to live streaming. [Means for solving the problem]

[0008] One aspect of the present invention relates to a server, which includes: means for receiving data including a behavior of a broadcaster via a network from a terminal of the broadcaster of a live broadcast in which the broadcaster and multiple viewers, including a virtual viewer realized by a machine learning model, participate, means for acquiring a reaction output by a machine learning model that receives the behavior of the broadcaster as input and outputs a reaction that a viewer with a set attribute would make; and means for transmitting data for realizing the reaction to the terminal via the network.

[0009] Another aspect of the present invention is a computer program, which realizes a function of transmitting data including the words and actions of a broadcaster to a terminal of the broadcaster of a live broadcast in which the broadcaster and a plurality of different virtual viewers realized by a plurality of different machine learning models participate, via a network, to a server that provides the live broadcast, a function of receiving data for realizing a plurality of reactions output from the plurality of machine learning models by inputting the words and actions to the plurality of machine learning models from the server via the network, and a function of displaying a plurality of objects representing the plurality of reactions on a display based on the data.

[0010] In addition, any combination of the above components, or mutual substitution of the components or expressions of the present invention between an apparatus, a method, a system, a computer program, a recording medium storing a computer program, etc., are also valid aspects of the present invention. Effect of the Invention

[0011] According to the present invention, new functions can be created or existing functions can be improved by applying machine learning to live streaming. [Brief description of the drawings]

[0012] [Figure 1] FIG. 1 is a schematic diagram showing a configuration of a live distribution system according to a first embodiment. [Diagram 2] FIG. 1 is a schematic diagram for explaining AI training delivery. [Diagram 3] 2 is a block diagram showing the functions and configuration of the user terminal of FIG. 1. [Figure 4] 2 is a block diagram showing the functions and configuration of the server in FIG. 1. [Diagram 5] 5 is a data structure diagram showing an example of a stream DB in FIG. 4. [Figure 6] 5 is a data structure diagram showing an example of a user DB in FIG. 4. [Figure 7] 5 is a data structure diagram showing an example of a gift DB in FIG. 4. [Figure 8] FIG. 5 is a data structure diagram showing an example of an ML model DB in FIG. [Figure 9] 1 is a flowchart showing the flow of a series of processes in AI training delivery. [Figure 10] FIG. 13 is a representative screen diagram of a distribution preparation screen displayed on the display of a distributor's user terminal. [Figure 11] This is a representative screen image of a live streaming room screen for AI training streaming displayed on the display of a streamer's user terminal. [Figure 12]This is a representative screen image of a live streaming room screen for AI training streaming displayed on the display of a streamer's user terminal. [Figure 13] This is a representative screen image of a live streaming room screen for AI training streaming displayed on the display of a streamer's user terminal. [Figure 14] FIG. 13 is a representative screen diagram of a distribution end screen displayed on the display of a user terminal of a distributor. [Figure 15] This is a representative screen image of a live streaming room screen displayed on the display of a user terminal of a real viewer participating in an AI training broadcast. [Figure 16] FIG. 13 is a representative screen diagram of a live streaming selection screen displayed on the display of the user terminal of an active user. [Figure 17] FIG. 11 is a block diagram showing the functions and configuration of a server according to a second embodiment. [Figure 18] 18 is a data structure diagram showing an example of a stream DB in FIG. 17. [Figure 19] 18 is a data structure diagram showing an example of an ML model DB of FIG. 17. [Figure 20] 1 is a flowchart showing the flow of a series of processes in live streaming in which AI viewers participate. [Figure 21] FIG. 13 is a representative screen diagram of a live streaming selection screen displayed on the display of the user terminal of an active user. [Figure 22] FIG. 13 is a representative screen diagram of a live streaming room screen displayed on the display of a viewer's user terminal. [Figure 23] 1 is a block diagram illustrating an example of a hardware configuration of an information processing device according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] Hereinafter, the same or equivalent components, members, processes, and signals shown in each drawing will be given the same reference numerals, and duplicated explanations will be omitted as appropriate. Also, some of the members that are not important for the explanation will be omitted in each drawing.

[0014] (First embodiment) The live distribution system according to the first embodiment provides live distribution in which a real distributor and a plurality of different virtual viewers (hereinafter referred to as AI viewers) realized by a plurality of different machine learning models (hereinafter referred to as ML models) participate. The distributor can simulate live distribution by distributing to the AI ​​viewers as if they were real viewers.

[0015] For example, the following usage scenario is assumed for the live distribution system according to this embodiment. Mr. A started streaming with the dream of becoming a popular streamer. However, he was worried about how to interact with viewers and how to liven up his content. At that time, Mr. A learned about the "AI Liver Training System" provided by the live streaming system according to this embodiment. The system has AI viewers with different personalities and interests, such as avid fans, critical viewers, and people watching for the first time. Party A uses this system to simulate a live broadcast. As Party A speaks, the AI ​​viewers send comments and gifts in real time. In addition, the AI ​​viewers also start conversations with each other, creating the atmosphere of the live broadcast. The system will analyze Party A's performance and evaluate the following: - Speed ​​and appropriateness of response to comments - Choose topics that will interest your audience - Timing of excitement in distribution - Fostering community building among viewers In addition, the system will simulate different situations, such as when a controversy is imminent, when viewership is declining, or when it is increasing rapidly, so that the party can practice how to handle these situations. Once trained, the system will provide detailed feedback, such as specific suggestions like "You tend to be slow to respond to comments" or "Your audience is more interested in certain topics." By using this system regularly, A improves his / her skills. A becomes more confident in actual broadcasts and gradually gains more fans. The system's unique feature is that it can simulate the complex environment unique to live streaming, rather than simply providing one-on-one conversation practice. Through interactions between multiple AI viewers, it recreates a more realistic streaming environment, helping streamers improve their overall skills.

[0016] FIG. 1 is a schematic diagram showing the configuration of a live distribution system 1 according to a first embodiment. The live distribution system 1 provides an interactive live distribution service in which a distributor (also called a live streamer) LV and a viewer (also called an audience) AU (AU1, AU2, ...) can communicate with each other in real time. As shown in FIG. 1, the live distribution system 1 includes a server 10, a user terminal 20 on the distributor side, and a user terminal 30 (30a, 30b, ...) on the viewer side. In addition to distributors who distribute live distribution and viewers who watch live distribution, there are also users who log in to the live distribution platform but do not distribute or watch. Such users are called active users. The distributors, viewers, and active users may be collectively called users. The server 10 may be configured by one or more information processing devices connected to a network NW. The user terminals 20 and 30 may be, for example, mobile terminals such as smartphones, tablet terminals, laptop PCs, recorders, portable game consoles, and wearable devices, or may be stationary devices such as desktop PCs. The server 10, the user terminal 20, and the user terminal 30 are connected to each other so as to be able to communicate with each other via various wired or wireless networks NW.

[0017] The live distribution system 1 involves a distributor LV, a viewer AU, an administrator (not shown) who manages the server 10, and an AI viewer. The distributor LV is a person who transmits content in real time by recording and filming his / her own songs, talks, performances, fortune telling, game commentary, etc. on his / her own user terminal 20 and uploading the content to the server 10 as is. The administrator provides a platform for live distribution of content on the server 10, and also mediates or manages real-time exchanges between the distributor LV and the viewer AU. The viewer AU accesses the platform with the user terminal 30, selects desired content, and views it. During the live distribution of this content, the viewer AU performs operations to make comments, cheer, or request a fortune telling through the user terminal 30, and the distributor LV who provides the content responds to such comments, cheers, or requests, and the reactions are transmitted to the viewer AU by video and / or audio, thereby establishing two-way communication.

[0018] In this specification, live streaming may be live streaming in which a distributor and a real viewer AU participate, live streaming in which a distributor and multiple viewers including an AI viewer participate, or live streaming in which only an AI viewer participates as a viewer. Live streaming may be realized using existing live streaming technologies such as HTTP Live Streaming, Common Media Application Format, Web Real-Time Communications, Real-Time Messaging Protocol, and MPEG DASH.

[0019] In this specification, "moving image data" refers to data including image data (also called video data) generated by the imaging function of the user terminals 20 and 30, and audio data (also called audio data) generated by the audio input function of the user terminals 20 and 30. The moving image data is played back on the user terminals 20 and 30, allowing users to view the content. In this embodiment, it is assumed that processing such as compression, decompression, encoding, decoding, transcoding, and other processing is performed to change the format, size, and specifications of the data from when the moving image data is generated on the distributor's user terminal until it is played back on the viewer's user terminal. Since the contents (e.g., moving images and audio) represented by the moving image data before and after such processing are substantially unchanged, this embodiment will be described assuming that the moving image data after such processing is the same as the moving image data before such processing. In other words, when moving image data is generated on the distributor's user terminal and then played back on the viewer's user terminal via the server 10, the moving image data generated on the distributor's user terminal, the moving image data passing through the server 10, and the moving image data received and played on the viewer's user terminal are all the same moving image data.

[0020] In the example of Fig. 1, a broadcaster LV is live-streaming a talk. The user terminal 20 of the broadcaster LV records and records the image and audio of the broadcaster LV while he or she is talking, generating video data, and transmits the video data to the server 10 via the network NW. The user terminal 20 also displays the recorded video image VD of the broadcaster LV on the display of the user terminal 20, thereby enabling the broadcast content by the broadcaster LV to be confirmed.

[0021] The user terminals 30a, 30b of the viewers AU1, AU2 who have requested the platform to watch the live broadcast of the broadcaster LV each receive video data related to the live broadcast via the network NW, and play the received video data to display the videos VD1, VD2 on the display and output audio from the speaker. The videos VD1, VD2 displayed on each user terminal 30a, 30b are substantially identical to the video VD captured by the user terminal 20 of the broadcaster LV, and the audio output from each user terminal 30a, 30b is also substantially identical to the audio recorded by the user terminal 20 of the broadcaster LV.

[0022] The recording and recording in the user terminal 20 of the distributor LV and the playback of the video data in the user terminals 30a and 30b of the viewers AU1 and AU2 are performed substantially simultaneously. When one viewer AU1 inputs a comment on the content of the talk of the distributor LV to the user terminal 30a, the server 10 displays the comment in real time on the user terminal 20 of the distributor LV and also on the user terminals 30a and 30b of each viewer AU1 and AU2. When the distributor LV reads the comment and develops a talk that overlaps the content, the video and audio of the talk are output on the user terminals 30a and 30b of each viewer AU1 and AU2, and it is recognized that a conversation between the distributor LV and the viewer AU1 has been established. In this way, the live distribution system 1 realizes live distribution that enables two-way communication, not one-way.

[0023] In this embodiment, an ML model is provided in the server 10, and this ML model is made to learn the activities of real viewers AU in live streaming. Through learning, when the words and actions of a distributor are input, the ML model outputs a reaction that a real viewer with the set attributes would make. An AI viewer is realized by such a trained ML model. The ML model may be realized using a known supervised machine learning technique such as GPT (Generative Pre-trained Transformer)-3, GPT-3.5, and GPT-4 provided by openAI, LLaMA (Large Language Model Meta AI) provided by Meta, and Bloom (BigScience Language Open-science Open-access Multilingual). A personality can be set for the model by inputting constraints and guidelines for action into the prompts of the ML model (see, for example, Non-Patent Document 1).

[0024] A broadcaster can carry out an AI training broadcast. In addition to the broadcaster, multiple different AI viewers participate in the AI ​​training broadcast. In the AI ​​training broadcast, the broadcaster's movements, facial expressions, comments, and other behavior are input into an ML model, and comments, pseudo-gifting, and other reactions output by the ML model in response to the input are provided to the broadcaster as coming from the AI ​​viewer. The AI ​​training broadcast may be configured so that only the AI ​​viewer can participate as a viewer, or so that real viewer AU can participate in addition to the AI ​​viewer.

[0025] 2 is a schematic diagram for explaining AI training distribution. The distributor LV is performing AI training distribution using the user terminal 20. A live distribution room screen 700 for the distributor is displayed on the display of the user terminal 20. The live distribution room screen 700 displays the video image of the distributor LV, comments, and pseudo gift effects.

[0026] The user terminal 20 of the distributor LV generates video data DA by recording the words and actions of the distributor, and transmits it to the server 10 via the network NW. The server 10 preprocesses the received video data DA and obtains image data IM and text data TX. The image data IM represents the images included in the video data DA, and the text data TX includes the voice included in the video data DA converted into text by STT (Speech to text) and comments.

[0027] Image data IM and text data TX are input to an ML model that realizes multiple (three in the example of FIG. 2) AI viewers MLV1, MLV2, and MLV3. Each ML model has different attributes set. Each ML model outputs a reaction that a viewer with the set attributes would make in response to the input of image data IM and text data TX. The output reaction includes a favorability rating 702 for the broadcaster LV (represented by a bar graph in the example of FIG. 2), input of a comment in the AI ​​training broadcast, and use of a pseudo gift. A pseudo gift is an item used by an AI viewer and is valid only in an individual AI training broadcast. A pseudo gift indicates the degree of excitement of an AI training broadcast and produces an effect in the AI ​​training broadcast, but does not affect the broadcaster's reward. In this embodiment, a case where an AI viewer uses a pseudo gift will be described, but in other embodiments, the AI ​​viewer may use a gift that affects the broadcaster's reward. This gift may be the same as the gift used by a real viewer AU.

[0028] The ML model may output the likeability rating 702, the use of pseudo-gifts, and the input of comments in parallel, or may first output the likeability rating 702, and then determine the amount of pseudo-gifts and the amount of comments according to the value of the likeability rating 702. If the likeability rating 702 is high, the amount of pseudo-gifts may be increased, and the amount of positive comments may be increased.

[0029] Even if the same image data IM and text data TX are input, each ML model may output a different reaction. For example, when the first AI viewer MLV1 and the second AI viewer MLV2 make comments and the distributor responds to the comment of the first AI viewer MLV1, the favorability rating 702 of the first AI viewer MLV1 increases, while the favorability rating 702 of the second AI viewer MLV2 decreases. If the interest of the second AI viewer MLV2 is set to baseball and the interest of the third AI viewer MLV3 is set to beauty, when the distributor talks about baseball, the second AI viewer MLV2 outputs many comments, while the third AI viewer MLV3 is silent.

[0030] The responses output by each ML model are input to other ML models. This allows for interaction between ML models (AI viewers) in AI training broadcasts. For example, if the first AI viewer MLV1 uses a large pseudo-gift, the favorability ratings of the other AI viewers MLV2 and MLV3 will increase. If the category of the first AI viewer MLV1 is set to VIP, as the number of comments output by the first AI viewer MLV1 increases, the number of comments output by the other AI viewers MLV2 and MLV3 will decrease. This shows consideration for VIPs.

[0031] In this way, with AI training broadcasts, broadcasters can check the AI ​​viewers' reactions to their own words and actions in real time, so they can learn what words and actions will make the broadcast more exciting and enjoyable by involving AI viewers with the desired attributes.

[0032] In this embodiment, a case is described in which the ML model outputs or updates the AI ​​viewer's likeability rating for the broadcaster, but in other embodiments, instead of or in addition to the likeability rating for the broadcaster, the ML model may output or update the AI ​​viewer's likeability rating for the live broadcast and / or the AI ​​viewer's degree of interest, emotions, or engagement score for the broadcaster or live broadcast.

[0033] Fig. 3 is a block diagram showing the functions and configuration of the user terminal 20 in Fig. 1. The user terminal 30 has the same functions and configuration as the user terminal 20. Each block shown in Fig. 2 and the following block diagrams can be realized in hardware terms by elements and mechanical devices such as a computer CPU, and in software terms by a computer program, etc., but here, functional blocks realized by the cooperation of these are depicted. Therefore, those skilled in the art who have read this specification will understand that these functional blocks can be realized in various ways by combining hardware and software.

[0034] The distributor LV and the viewer AU download and install a live distribution application program (hereinafter, referred to as a live distribution app) according to this embodiment from a download site to the user terminals 20 and 30 via the network NW. Alternatively, the live distribution app may be pre-installed in the user terminals 20 and 30. When the live distribution app is executed by the user terminals 20 and 30, the user terminals 20 and 30 communicate with the server 10 via the network NW to realize various functions. Hereinafter, functions realized by the user terminals 20 and 30 (processors such as CPUs) executing the live distribution app will be described as functions of the user terminals 20 and 30. These functions are actually functions that the live distribution app causes the user terminals 20 and 30 to realize. Note that in other embodiments, these functions may be realized by a computer program written in a programming language such as HTML (HyperText Markup Language), which is transmitted from the server 10 to a web browser of the user terminals 20 and 30 via the network NW and executed by the web browser.

[0035] The user terminal 20 includes a distribution unit 100 that generates video data recording the user's image and voice and provides it to the server 10, a viewing unit 200 that acquires the video data from the server 10 and plays it, and a non-distribution processing unit 400 that processes requests from active users. The user activates the distribution unit 100 when distributing, the viewing unit 200 when viewing, and the non-distribution processing unit 400 when searching for a live broadcast to watch, viewing a distributor's profile, or viewing an archive. A user terminal on which the distribution unit 100 is active is a user terminal on the distributor side, that is, a user terminal on the generating side of video data, a user terminal on which the viewing unit 200 is active is a user terminal on the viewer side, that is, a user terminal on the playing side of video data, and a user terminal on which the non-distribution processing unit 400 is active is a user terminal of an active user.

[0036] The delivery unit 100 includes an imaging control unit 102, an audio control unit 104, a video transmission unit 106, a delivery-side UI control unit 108, and a delivery-side communication unit 110. The imaging control unit 102 is connected to a camera (not shown in FIG. 2) and controls imaging by the camera. The imaging control unit 102 acquires image data from the camera. The audio control unit 104 is connected to a microphone (not shown in FIG. 2) and controls audio input by the microphone. The audio control unit 104 acquires audio data from the microphone. The video transmission unit 106 transmits video data including the image data acquired by the imaging control unit 102 and the audio data acquired by the audio control unit 104 to the server 10 via the network NW. The video transmission unit 106 transmits the video data in real time. That is, the generation of video data by the imaging control unit 102 and the audio control unit 104 and the transmission of the generated video data by the video transmission unit 106 are substantially performed simultaneously.

[0037] The distribution side UI control unit 108 controls the UI for the distributor. The distribution side UI control unit 108 is connected to a display not shown in FIG. 2, and displays a moving image on the display by playing back the moving image data to be transmitted by the moving image transmission unit 106. The distribution side UI control unit 108 is connected to input means such as a touch panel, a keyboard, and a display not shown in FIG. 2, and acquires input from the distributor via the input means. The distribution side UI control unit 108 superimposes a predetermined frame image on the moving image. The frame image includes various user interface objects (hereinafter simply referred to as objects) for receiving input from the distributor, comments input by the viewer, and information acquired from the server 10. The distribution side UI control unit 108 accepts, for example, a tap input on an object by the distributor.

[0038] The delivery-side communication unit 110 controls communication with the server 10 during live delivery. The delivery-side communication unit 110 transmits the contents of the input by the distributor acquired by the delivery-side UI control unit 108 to the server 10 via the network NW. The delivery-side communication unit 110 receives various types of information associated with the live delivery from the server 10 via the network NW.

[0039] The viewing unit 200 includes a viewing side UI control unit 202 and a viewing side communication unit 204. The viewing side communication unit 204 controls communication with the server 10 during live distribution. The viewing side communication unit 204 receives video data related to the live distribution in which the distributor and the viewers participate from the server 10 via the network NW.

[0040] The viewer-side UI control unit 202 controls the UI for the viewer. The viewer-side UI control unit 202 is connected to a display and a speaker (not shown in FIG. 2), and plays the received video data to display a video image on the display and output audio from the speaker. The output of an image on the display and audio from the speaker can be collectively referred to as "video data being played". The viewer-side UI control unit 202 is connected to input means (not shown in FIG. 2), such as a touch panel, a keyboard, or a display, and acquires input by the viewer through the input means. The viewer-side UI control unit 202 superimposes a predetermined frame image on the image of the video data acquired from the server 10. The frame image includes various objects for receiving input from the viewer, comments input by the viewer, and information acquired from the server 10. The viewer-side communication unit 204 transmits the contents of the input by the viewer acquired by the viewer-side UI control unit 202 to the server 10 via the network NW.

[0041] The non-broadcast processing unit 400 includes a non-broadcast UI control unit 402 and a non-broadcast communication unit 404. The non-broadcast UI control unit 402 controls a UI for active users. For example, the non-broadcast UI control unit 402 generates a live broadcast selection screen that displays a list of live broadcasts that can currently be joined and accepts the selection of a live broadcast by an active user, and displays it on the display. The non-broadcast UI control unit 402 generates a profile screen of an arbitrary user and displays it on the display. The non-broadcast UI control unit 402 plays back an archive generated by recording and filming past live broadcasts.

[0042] The non-broadcast communication unit 404 controls communication with the server 10 outside of live broadcasting. The non-broadcast communication unit 404 receives information for generating a live broadcast selection screen, information for generating a profile screen, and archive data from the server 10 via the network NW. The non-broadcast communication unit 404 transmits the contents of input by active users to the server 10 via the network NW.

[0043] Fig. 4 is a block diagram showing the functions and configuration of the server 10 in Fig. 1. The server 10 includes a distribution information providing unit 302, a relay unit 304, a gift processing unit 308, a payment processing unit 310, a training unit 330, a stream DB 314, a user DB 318, a gift DB 320, and an ML model DB 340.

[0044] 5 is a data structure diagram showing an example of the stream DB 314 in FIG. 4. The stream DB 314 holds information on live streaming including AI training streaming currently being performed. The stream DB 314 holds a stream ID that identifies a live streaming (including AI training streaming) on ​​a live streaming platform provided by the live streaming system 1, a broadcaster ID that is a user ID that identifies a broadcaster of the live streaming, a viewer ID that is a user ID that identifies an actual viewer (not an AI viewer) of the live streaming, a flag indicating whether the live streaming is an AI training streaming or not, the status of the AI ​​viewer participating in the live streaming, the total amount of pseudo gifts used in the live streaming, the total number of comments posted in the live streaming, the score of the live streaming, the number of viewers of the live streaming (including AI viewers), and the streaming time of the live streaming, in association with each other.

[0045] In the live streaming platform provided by the live streaming system 1 according to the present embodiment, when a user performs live streaming, the user becomes a streamer, and when the same user watches a live stream by another user, the user becomes a viewer. Therefore, the distinction between streamer and viewer is not fixed, and a user ID registered as a streamer ID at one time may be registered as a viewer ID at another time.

[0046] The flag "Y" indicates that it is an AI training broadcast, and "N" indicates that it is not.

[0047] The status of an AI viewer includes an AI viewer ID that identifies the AI ​​viewer participating in the AI ​​training broadcast, the AI ​​viewer's emotions, the AI ​​viewer's favorability rating for the broadcaster, and the total amount of pseudo gifts used by the AI ​​viewer. The number of AI viewers in an AI training broadcast is calculated by counting the AI ​​viewer IDs included in the corresponding status.

[0048] In this embodiment, the ML model realizing the AI ​​viewer participating in the AI ​​training distribution is generated exclusively for the AI ​​training distribution by duplicating the model data of the original ML model. Therefore, it is possible to have AI viewers with the same attributes participate in two different AI training distributions that are proceeding simultaneously. In the example of FIG. 5, three AI viewers are participating in the AI ​​training distribution "ST1". The AI ​​viewer with the AI ​​viewer ID "MD1_1" is an AI viewer realized by the ML model with the model ID "MD1" described later, outputs the emotion "fun", has a favorability output of "40", and has used "1000" pseudo gifts so far. Three AI viewers are participating in the AI ​​training distribution "ST2". The AI ​​viewer with the AI ​​viewer ID "MD1_2" is an AI viewer realized by the ML model with the same model ID "MD1" as above, outputs the emotion "joy", has a favorability output of "60", and has used "2000" pseudo gifts so far.

[0049] The score is an indicator of the excitement of a live stream. A live stream with a high score is recognized as "exciting" or "popular." The score varies depending on, for example, the number of viewers, the duration of the stream, the number of comments, the content of the comments, the number of shares, the total amount of gifts, the number of viewers who gave gifts, and the number of cheers. The score is reset when the live stream ends. Cheers are digital items that viewers give to the streamer, and unlike gifts, they do not require payment. Once you have given an cheer, you must wait a certain period of time before you can give another cheer. The score is an example of an indicator that shows the streamer's performance in a live stream. The score for AI training streams will be described later.

[0050] FIG. 6 is a data structure diagram showing an example of the user DB 318 of FIG. 4. The user DB 318 holds information about the user. The user DB 318 holds a user ID for identifying the user, points owned by the user, a reward given to the user, attributes of the user, a viewing history of the user, and a distribution history of the user in association with each other. The attributes of the user include the gender of the user, the age of the user, the category of the user, the personality of the user, and the interest of the user. The attributes of the user, except for the category, may be registered by the user himself through a distribution application, or may be predicted by the system using machine learning or the like from the viewing history or distribution history of the user.

[0051] The category is an indicator of the user's performance as a viewer on the live streaming platform. In another embodiment, the level may be an indicator of the user's performance as a broadcaster on the live streaming platform, or may be an indicator of the user's performance as both a broadcaster and a viewer. The category may be determined based on the viewing history. The category may be increased or decreased depending on the total viewing time as a viewer of the live streaming, the number and / or amount of gifts given, the number of comments, etc. Alternatively, the category may be evaluated and determined by an administrator. Alternatively, the category may be automatically determined by a predetermined rule or a machine learning model. In the example of FIG. 6, the category is selected from three categories: "new", "mid", and "VIP".

[0052] Points are electronic value circulated within the live streaming platform. Users purchase points by credit card or other payment methods. Rewards are electronic value defined within the live streaming platform and are an indicator for determining the amount of money that a broadcaster receives from the administrator of the live streaming platform. On the live streaming platform, when a viewer gives a gift to a broadcaster during or outside of a live streaming, the viewer's points are consumed and the broadcaster's reward increases accordingly.

[0053] The viewing history is data that collects activity history related to viewing live streams on a live streaming platform, such as when and for how long a user watched which live stream, what comments they made, how many gifts they gave, etc. The streaming history is data that collects activity history related to streaming of live streams on a live streaming platform, such as when a user live streamed, how long the stream lasted, what comments they made, how many scores or gifts they received, etc.

[0054] Fig. 7 is a data structure diagram showing an example of the gift DB 320 in Fig. 4. The gift DB 320 holds information about gifts that can be used by viewers in live distribution. A gift is electronic data or a digital item that has the following characteristics. -Can be purchased with points or money, or given free of charge. Something that viewers can give to the streamer. Giving a gift to a streamer is also called using a gift or throwing a gift. In some cases, the purchase and use of the gift occur simultaneously as a set, while in other cases the viewer can use the gift at any time after purchasing it. When a viewer gives a gift to a streamer, the streamer will receive a corresponding reward. When a gift is used, an effect associated with the gift may occur. For example, an effect corresponding to the gift may appear on the live streaming room screen.

[0055] The gift DB320 stores a gift ID that identifies a gift, a given reward that is given to a broadcaster when the gift is given to the broadcaster, and a value point that is a value to be paid when the gift is used, in association with each other. A viewer can give a gift to a broadcaster by paying the value points of a desired gift while watching a live broadcast. The payment of the value points may be made by an appropriate electronic payment means, for example, the viewer may pay the value points to the administrator. Alternatively, payment by bank transfer or credit card may be used. The relationship between the given reward and the value points can be arbitrarily set by the administrator. For example, the given reward may be set to the value points. Alternatively, the value points may be set to points obtained by multiplying the given reward by a predetermined coefficient such as 1.2, or the value points may be set to points obtained by adding a predetermined commission point to the given reward.

[0056] FIG. 8 is a data structure diagram showing an example of the ML model DB340 of FIG. 4. The ML model DB340 holds information on an ML model that realizes an AI viewer who is to participate in an AI training distribution. The ML model DB340 holds a model ID that identifies an ML model, attributes set in the ML model, and model data of the ML model in association with each other. The attributes include gender, age, category, personality, and interests, similar to those in FIG. 6. When the ML model is realized by an API (Application Programming Interface), the model data includes a URL (Uniform Resource Locator) to be included in an API call of the ML model.

[0057] In this embodiment, an administrator determines the attributes of an ML model, and causes the ML model to learn the viewing history of a real viewer having the determined attributes. In the example of FIG. 8, the administrator acquires the viewing history of a real user (multiple users are possible) whose user attributes are male, in his / her 30s, VIP, and impatient from the user DB318, and causes the ML model "MD1" to learn the acquired viewing history. As a result, the ML model "MD1" outputs a reaction to the distributor's words and actions that a real viewer having the attributes of male, in his / her 30s, VIP, and impatient would have. Such learning of the ML model may be realized using a known machine learning technique.

[0058] Returning to FIG. 4, when the distribution information providing unit 302 receives a request to start live distribution from the user terminal 20 of the distributor via the network NW, it registers a stream ID that identifies the live distribution and a distributor ID of the distributor of the live distribution in the stream DB 314. When the distribution information providing unit 302 receives a request to provide information about live distribution from the non-distribution communication unit 404 of the user terminal of an active user via the network NW, it refers to the stream DB 314 and generates a list of live distributions that can be currently viewed. The distribution information providing unit 302 transmits the generated list to the user terminal that made the request via the network NW. The non-distribution UI control unit 402 of the user terminal that made the request generates a live distribution selection screen based on the received list, and displays it on the display of the user terminal.

[0059] When the non-delivery UI control unit 402 of the user terminal accepts the selection of live streaming by the active user on the live streaming selection screen, it generates a delivery request including the stream ID of the selected live streaming and transmits it to the server 10 via the network NW. The delivery information providing unit 302 starts providing the live streaming identified by the stream ID included in the received delivery request to the requesting user terminal. The delivery information providing unit 302 updates the stream DB 314 so that the viewer ID of the stream ID includes the user ID of the active user of the requesting user terminal. As a result, the active user becomes a viewer of the selected live streaming.

[0060] The relay unit 304 relays the transmission of video data from the user terminal 20 of the distributor to the user terminal 30 of the viewer in the live distribution started by the distribution information providing unit 302. The relay unit 304 receives a signal indicating a user input by the viewer during the live distribution, i.e., during the playback of the video data, from the viewing side communication unit 204. The signal indicating the user input may be an object designation signal indicating the designation of an object displayed on the display of the user terminal 30, and the object designation signal includes the viewer ID of the viewer, the distributor ID of the distributor performing the live distribution that the viewer is viewing, and an object ID that identifies the object. If the object is a gift icon, the object ID is a gift ID. In that case, the object designation signal is a gift use signal indicating the use of a gift by the viewer for the distributor. Similarly, the relay unit 304 receives a signal indicating a user input by the distributor during the playback of the video data, for example, an object designation signal, from the distribution side communication unit 110 of the distribution unit 100 of the user terminal 20.

[0061] The gift processing unit 308 updates the user DB 318 so as to increase the distributor's reward according to the reward for the gift specified by the gift ID included in the gift use signal. The gift processing unit 308 refers to the gift DB 320 and specifies the reward corresponding to the gift ID included in the received gift use signal. The gift processing unit 308 updates the user DB 318 so as to add the specified reward to the reward corresponding to the distributor ID included in the gift use signal.

[0062] In response to receiving the gift use signal, the payment processing unit 310 processes the payment of the value of the gift by the viewer. The payment processing unit 310 refers to the gift DB 320 and identifies the value points of the gift identified by the gift ID included in the gift use signal. The payment processing unit 310 updates the user DB 318 so as to subtract the identified value points from the points of the viewer identified by the viewer ID included in the gift use signal.

[0063] The training unit 330 manages and controls the AI ​​training distribution. The training unit 330 includes a setting unit 332, a progress processing unit 334, an evaluation unit 336, and a feedback unit 338.

[0064] When the setting unit 332 receives a request to start an AI training distribution from the user terminal 20 of the distributor via the network NW, it registers in the stream DB 314 a stream ID that identifies the AI ​​training distribution, a distributor ID of the distributor of the AI ​​training distribution, a training flag "Y", and the status of the AI ​​viewer. The setting unit 332 registers the AI ​​viewer specified in the start request (the AI ​​viewer specified by the distributor) in the status. For example, when the distributor specifies an AI viewer specified by the ML model "MD1", the setting unit 332 duplicates the model data of the ML model "MD1", assigns the AI ​​viewer ID "MD1_1" to the AI ​​viewer realized by the duplicate, and registers the AI ​​viewer ID "MD1_1" in the status. When the distribution information providing unit 302 completes the setting of the AI ​​viewer, it starts providing the AI ​​training distribution to the user terminal of the distributor that made the start request.

[0065] The progress processing unit 334 receives video data from the user terminal 20 of the distributor via the network NW in the AI ​​training distribution started by the setting unit 332, and transmits reaction data including reactions by the AI ​​viewer to the user terminal 20 via the network NW. The progress processing unit 334 receives a signal indicating a user input by the distributor during playback of the video data, for example an object designation signal, from the distribution-side communication unit 110 of the distribution unit 100 of the user terminal 20. As described above, the video data includes the words and actions of the distributor in the AI ​​training distribution in which the distributor and multiple viewers including an AI viewer realized by an ML model participate.

[0066] The progress processing unit 334 receives the words and actions of the distributor included in the received video data, and acquires the reaction output by the ML model that outputs the reaction that a viewer with the set attributes would make. The progress processing unit 334 extracts image data and text data from the video data and inputs them to the ML models of all AI viewers participating in the AI ​​training distribution. The progress processing unit 334 acquires the reaction output from each ML model in response to the input. The progress processing unit 334 includes the acquired reaction in the reaction data. The reaction data may include the AI ​​viewer status, total amount of pseudo gifts, total number of comments, score, number of viewers, and distribution time of the AI ​​training distribution held in the stream DB 314. The user terminal of the distributor reflects the reaction of the AI ​​viewer (use of gifts, input of comments, etc.) on the live distribution room screen based on the received reaction data, so it can be said that the reaction data is data for realizing the reaction of the AI ​​viewer on the user terminal of the distributor.

[0067] When a pseudo-gift is used by an AI viewer, the progress processing unit 334 updates the status, the total amount of pseudo-gifts, and the score corresponding to the AI ​​training distribution in which the pseudo-gift was used in the stream DB 314. In the example of FIG. 5, when the AI ​​viewer "MD1_1" uses a pseudo-gift of 30 points in the AI ​​training distribution "ST1", 30 points are added to the pseudo-gift of "MD1_1" in the status of "ST1" to become 70, and 30 points are added to the total amount of pseudo-gifts of "ST1" to become 1530. The score is also updated according to a predetermined formula. However, in this embodiment, the pseudo-gift does not affect the remuneration of the distributor, so the remuneration of the distributor "LR1" in FIG. 6 does not change due to the use of the pseudo-gift.

[0068] The progress processing unit 334 inputs the responses output by the ML model of each AI viewer participating in the AI ​​training distribution into the ML models of the other AI viewers.

[0069] The evaluation unit 336 evaluates the AI ​​training distribution and / or the distributor of the AI ​​training distribution based on the reactions output by the ML models corresponding to each of the multiple AI viewers participating in the AI ​​training distribution. The evaluation unit 336 calculates a score for the AI ​​training distribution in real time and updates the stream DB 314 with the calculated score. The score is an index that indicates the excitement of the AI ​​training distribution. The score varies depending on, for example, the number of viewers, the distribution time, the number of comments, the content of the comments, the total amount of pseudo gifts, the number of AI viewers who sent pseudo gifts, and the like. The score, the total amount of pseudo gifts, and the total number of comments are examples of indexes that indicate the performance of the distributor in the AI ​​training distribution. In addition to information obtained from the AI ​​training distribution, the evaluation unit 336 may also use information obtained from non-AI training distributions performed by the same distributor to perform the evaluation.

[0070] When the AI ​​training distribution is completed, the evaluation unit 336 generates an improvement suggestion comment to the distributor of the AI ​​training distribution by analyzing the archive data of the AI ​​training distribution. The generation of the improvement suggestion comment may be realized on a rule-based basis or may be realized using machine learning technology. In the case of a rule-based basis, it can be realized by associating a combination of parameter value ranges such as a score representing the performance of the AI ​​training distribution and the number of viewers with a pre-input improvement suggestion comment and storing them in the server. The evaluation by the evaluation unit 336 may include an evaluation of the reaction of each AI viewer, an evaluation of the reaction of the AI ​​viewer group as a whole (such as the reaction of the beginner viewer group), and a proposal for distribution content. The reaction of the AI ​​viewer group is, for example, the good or bad reception from the beginner group or the good or bad reception from the VIP group.

[0071] The feedback unit 338 transmits the result of the evaluation by the evaluation unit 336 to the user terminal of the distributor of the target AI training distribution via the network NW. When the AI ​​training distribution is in progress, the result of the evaluation includes the score at that time, the status of the AI ​​viewer, the total amount of pseudo gifts, the total number of comments, the number of viewers, and the distribution time. When the AI ​​training distribution has ended, the result of the evaluation includes the score at the end time, the status of the AI ​​viewer, the total amount of pseudo gifts, the total number of comments, the number of viewers, the distribution time, and the generated improvement suggestion comment. Regarding training of the ML model, the server 10 may be configured to improve the sophistication of the corresponding AI viewer's response by learning from actual viewer responses as training data.

[0072] The operation of the live distribution system 1 configured as above will now be described. FIG. 9 is a flowchart showing a series of processing steps in AI training distribution. The server 10 receives a request to start AI training distribution from the user terminal 20 of the distributor via the network NW (S202). The server 10 starts providing live distribution (i.e., AI training distribution) in which multiple AI viewers participate (S204). The server 10 receives video data recording the words and actions of the distributor from the user terminal 20 of the distributor via the network NW (S206). The server 10 extracts the words and actions of the distributor from the received video data (S208). The server 10 inputs the extracted words and actions into an ML model corresponding to the AI ​​viewers (S210). The server 10 acquires multiple reactions output by multiple ML models (S212). The server 10 updates the evaluation parameters of the AI ​​training distribution based on the acquired multiple reactions (S214). The evaluation parameters include the total amount of pseudo gifts, the total number of comments, the score, and the number of viewers.

[0073] The server 10 transmits the acquired multiple responses and the updated evaluation parameters to the distributor's user terminal 20 via the network NW (S216). The server 10 determines whether or not a distribution end instruction has been received from the distributor's user terminal 20 (S218). If not received (N in S218), the process returns to step S206. If received (Y in S218), the server 10 generates evaluation information for the ended AI training distribution (S220). The server 10 transmits the generated evaluation information to the distributor's user terminal 20 via the network NW (S222).

[0074] FIG. 10 is a representative screen diagram of a distribution preparation screen 710 displayed on the display of the user terminal 20 of the distributor. The distribution preparation screen 710 has a distribution start button 712 for receiving an instruction to start normal (non-AI training distribution) live distribution, and a training start button 714 for receiving an instruction to start AI training distribution. The distribution preparation screen 710 is configured to be able to receive various settings for AI training distribution. The distribution preparation screen 710 has a first selection area 716 that allows the distributor to select whether the system randomly selects AI viewers to participate in the AI ​​training distribution, or whether the distributor designates them in advance. The distribution preparation screen 710 has a second selection area 718 that allows the distributor to select whether to allow the AI ​​viewers to enter and leave the AI ​​training distribution midway. The distribution preparation screen 710 has a third selection area 720 that allows the distributor to select whether to allow real (non-AI viewers) viewers to participate in the AI ​​training distribution. When the distributor selects in advance to designate an AI viewer in the first selection area 716, the distribution preparation screen 710 has an AI viewer setting area 722 that displays details of the AI ​​viewer to be participated and enables the distributor to add, modify, or delete. The AI ​​viewer setting area 722 displays attributes and a change button 724 for each AI viewer to be included in the AI ​​training distribution. When the change button 724 is specified (e.g., tapped), the user terminal 20 superimposes a list of AI viewers to be changed to on the distribution preparation screen 710. This list may be generated by the user terminal 20 referring to the ML model DB 340 of the server 10 via the network NW. The AI ​​viewer setting area 722 has an add button 726. When the add button 726 is specified, a list similar to that when the change button 724 is specified is displayed, and information of the AI ​​viewer selected from the list is added to the AI ​​viewer setting area 722.

[0075] When the distributor makes the desired settings in the first selection area 716, the second selection area 718, the third selection area 720, and the AI ​​viewer setting area 722 and presses the start training button 714, the user terminal 20 generates a request to start AI training distribution including the selection and input results in the first selection area 716, the second selection area 718, the third selection area 720, and the AI ​​viewer setting area 722, and transmits it to the server 10 via the network NW.

[0076] The processing when "Specify" is selected in first selection area 716 is as described above, and the AI ​​viewer specified by the distributor in AI viewer setting area 722 becomes the initial participant in the AI ​​training distribution. When "Random" is selected in first selection area 716, the initial participants are randomly selected from the ML models registered in ML model DB 340.

[0077] When "Yes" is selected in the second selection area 718, new AI viewers may enter or leave the room depending on the score of the AI ​​training delivery. For example, many AI viewers may enter the room when the score suddenly rises, or a specific AI viewer may leave the room when the favorability rating of the specific AI viewer falls below a threshold.

[0078] FIG. 11 is a representative screen diagram of a live distribution room screen 700 of an AI training distribution displayed on the display of the user terminal 20 of the distributor. The live distribution room screen 700 includes a plurality of objects representing a plurality of reactions of a plurality of AI viewers, which are generated based on the reaction data received from the server 10. The live distribution room screen 700 has a video image 728 of the distributor obtained by playing the video data to be transmitted by the video transmission unit 106 of the user terminal 20, a comment display area 730, a distribution end button 732, an AI viewer status display area 734, and an evaluation parameter display area 736. The distribution side UI control unit 108 of the user terminal 20 generates the live distribution room screen 700 by superimposing other objects, that is, the comment display area 730, the distribution end button 732, the AI ​​viewer status display area 734, and the evaluation parameter display area 736, on the video image 728 obtained by playing the video data.

[0079] The comment display area 730 may include comments entered by the AI ​​viewer, comments entered by real viewers if real viewers are available for participation, and a notification from the system. The notification from the system may include information indicating which AI viewer gave which pseudo gift to the broadcaster. The broadcasting side UI control unit 108 generates a comment display area 730 including the AI ​​viewer's comments included in the reactions received from the server 10, and includes the generated comment display area 730 in the live broadcasting room screen 700.

[0080] The distribution end button 732 is an object for receiving an instruction from the distributor to stop providing the AI ​​training distribution. When the distribution end button 732 is tapped, the distribution side communication unit 110 generates a distribution end instruction and transmits it to the server 10 via the network NW.

[0081] The AI ​​viewer status display area 734 displays the status of each AI viewer participating in the AI ​​training broadcast. For each AI viewer, the AI ​​viewer status display area 734 displays attributes 738, emotions 740, favorability rating 742, and an object 744 representing the total amount of pseudo gifts used by the AI ​​viewer so far. The broadcasting side UI control unit 108 generates the AI ​​viewer status display area 734 based on the response received from the server 10, and includes the generated AI viewer status display area 734 in the live broadcasting room screen 700.

[0082] The evaluation parameter display area 736 displays the evaluation parameters of the AI ​​training distribution at the current time. The evaluation parameter display area 736 displays the total amount of pseudo gifts, the total number of comments, the score, the number of viewers, and the distribution time. The distribution side UI control unit 108 generates the evaluation parameter display area 736 based on the reaction received from the server 10, and includes the generated evaluation parameter display area 736 in the live distribution room screen 700.

[0083] By looking at the AI ​​viewer status display area 734 and the evaluation parameter display area 736 on the live streaming room screen 700, the broadcaster can see how the AI ​​viewer reacts to his or her own words and actions during the live streaming, and can experiment to find out what words and actions will make the live streaming more exciting.

[0084] FIG. 12 is a representative screen diagram of a live distribution room screen 700 of an AI training distribution displayed on the display of a distributor's user terminal 20. FIG. 12 corresponds to a case where a distributor is on fire. The server 10 judges whether or not the AI ​​training distribution is on fire based on the reaction of the AI ​​viewers participating in the AI ​​training distribution and the value of the evaluation parameter. For example, if the score is lower than a threshold value and the favorability rating of all the AI ​​viewers is lower than a threshold value, the AI ​​training distribution is judged to be on fire. If it is judged that the AI ​​training distribution is on fire, the distribution side UI control unit 108 superimposes an object 748 indicating that the AI ​​training distribution is on fire on the live distribution room screen 700.

[0085] FIG. 13 is a representative screen diagram of a live distribution room screen 700 of an AI training distribution displayed on the display of a distributor's user terminal 20. FIG. 13 corresponds to a case where the evaluation of the AI ​​training distribution is rising sharply. The server 10 judges whether the AI ​​training distribution is in a rising state based on the reaction of the AI ​​viewers participating in the AI ​​training distribution and the value of the evaluation parameter. For example, if the rate of increase in the score is higher than a threshold value and the favorability ratings of all the AI ​​viewers are higher than a threshold value, the AI ​​training distribution is judged to be in a rising state. If it is judged that the AI ​​training distribution is in a rising state, the distribution side UI control unit 108 superimposes an object 746 indicating the rising state on the live distribution room screen 700.

[0086] In this way, by displaying objects 746, 748 indicating the status of the AI ​​training broadcast on the live broadcast room screen 700, the broadcaster can understand at a glance the status of their broadcast.

[0087] 14 is a representative screen diagram of a distribution end screen 750 displayed on the display of the distributor's user terminal 20. When the distributor taps the distribution end button 732, a distribution end instruction is sent to the server 10. The server 10 sends the evaluation result of the ended AI training distribution to the user terminal 20. The distribution side UI control unit 108 generates a distribution end screen 750 based on the evaluation result received from the server 10, and displays it on the display. The distribution end screen 750 has a parameter display area 752 that displays evaluation parameters at the time of the end of the AI ​​training distribution, and an improvement suggestion comment display area 754 that displays improvement suggestion comments.

[0088] FIG. 16 is a representative screen diagram of a live streaming selection screen 600 displayed on the display of a user terminal of an active user. The live streaming selection screen 600 includes thumbnails 602 showing each live streaming in a list of currently viewable live streaming (including AI training streaming set to allow real viewers to participate) received from the server 10. The non-streaming UI control unit 402 generates the live streaming selection screen 600 based on the list of live streaming acquired from the server 10, and displays it on the display. On the live streaming selection screen 600, thumbnails corresponding to AI training streaming are displayed in a distinguishable manner from thumbnails that do not. In the example of FIG. 16, a mark 604 is added to the thumbnail corresponding to the AI ​​training streaming.

[0089] FIG. 15 is a representative screen diagram of a live distribution room screen 608 displayed on the display of the user terminal 30 of a real viewer participating in the AI ​​training distribution. The live distribution room screen 608 displays a video generated by the user terminal 20 of the distributor in real time. The live distribution room screen 608 has a video 610 of the distributor obtained by playing video data received from the server 10, a comment display area 618, a viewing end button 620, an evaluation parameter display area 636, and a message display area 622. The live distribution room screen 608 does not have an area for accepting input of comments or an area for accepting instructions to use gifts. In the example of FIG. 15, the use of gifts and input of comments by real viewers are prohibited in the AI ​​training distribution. The message display area 622 displays that fact. The real viewers can see the process of the distributor's growth through watching the live distribution room screen 608.

[0090] In the above-mentioned embodiment, examples of the database are a hard disk and a semiconductor memory. Furthermore, based on the description in this specification, it will be understood by those skilled in the art who have read this specification that each unit can be realized by a CPU (not shown), an installed application program module, a system program module, a semiconductor memory that temporarily stores the contents of data read from a hard disk, and the like.

[0091] According to the live distribution system 1 of the present embodiment, it is possible to allow AI viewers to participate in live distribution. The AI ​​viewers provide reactions to the distributor's words and actions. This allows the distributor to proceed with live distribution even if there are no real viewers, and also allows the live distribution to be livened up by having AI viewers participate when there are real viewers. In addition, the distributor can hone their distribution skills by conducting AI training distribution with multiple AI viewers. This can provide a useful training means, especially for distributors who have just started distribution.

[0092] (Second embodiment) In the first embodiment, an AI viewer is used to train a broadcaster. In the second embodiment, a broadcaster creates and uses an AI viewer that is his or her "alter ego."

[0093] 17 is a block diagram showing the functions and configuration of the server 50 according to the second embodiment. The server 10 includes a distribution information providing unit 302, a relay unit 304, a gift processing unit 308, a payment processing unit 310, a model generating unit 350, a model deploying unit 352, a stream DB 354, a user DB 318, a gift DB 320, and an ML model DB 356.

[0094] Fig. 18 is a data structure diagram showing an example of the stream DB 354 in Fig. 17. The stream DB 354 holds a stream ID that identifies a live stream on the live streaming platform provided by the live streaming system 1, a broadcaster ID that is a user ID that identifies the broadcaster of the live stream, a viewer ID that identifies the viewers (including AI viewers) of the live stream, the total number of comments posted in the live stream, the score of the live stream, the number of viewers (including AI viewers) of the live stream, and the streaming time of the live stream, in association with each other.

[0095] Fig. 19 is a data structure diagram showing an example of the ML model DB 356 in Fig. 17. The ML model DB 356 holds information on an ML model that is generated by a user or instructed to be generated by a user and that realizes an AI viewer corresponding to the user or another user designated by the user. The ML model DB 356 holds, in association with each other, a model ID that identifies an ML model, a corresponding user ID that identifies a user corresponding to the ML model, model data of the ML model, and a gift budget allocated to the ML model by the user who generated it.

[0096] Returning to FIG. 17, the model generation unit 350 receives a generation request for an AI viewer corresponding to a user (hereinafter referred to as an avatar generation user) or another user designated by the user (hereinafter collectively referred to as an avatar target user) from the user terminal of the user. The model generation unit 350 acquires the viewing history of the avatar target user designated in the received generation request from the user DB 318. The model generation unit 350 generates a trained ML model by having an untrained ML model learn the acquired viewing history. Through training, the trained ML model comes to output a reaction that would be made by the avatar target user (real user). In this sense, the avatar target user is one of the attributes of the corresponding ML model or AI viewer, and therefore the corresponding user ID that identifies the user corresponding to the ML model is included in the attributes of the ML model or AI viewer. It can be said that the ML model that has learned the viewing history of viewer A has the "viewer A attribute". The model generation unit 350 registers information of the generated trained ML model in the ML model DB 356. The model generation unit 350 inquires of the user for whom avatars are to be generated about the gift budget, and registers the answer obtained in the ML model DB 356 .

[0097] The AI ​​viewer uses gifts within the gift budget in a manner similar to the use of simulated gifts in the first embodiment. Gifts used by an AI viewer and gifts used by a real viewer have the same effect. When an AI viewer uses a gift, gift processing unit 308 and payment processing unit 310 perform the same processing as when a real viewer uses a gift.

[0098] The model deployment unit 352 performs a process for allowing an AI viewer realized by the ML model generated by the model generation unit 350 to participate as a viewer in various live broadcasts. For example, the model deployment unit 352 may accept a reservation for allowing an AI viewer to participate in a live broadcast. In this case, the model deployment unit 352 accepts a broadcaster designated by the avatar generation user, and when the broadcaster starts live broadcasting, allows the AI ​​viewer of the avatar target user to participate in the live broadcasting. Alternatively, when an active user designates a thumbnail on the live broadcast selection screen, the model deployment unit 352 may allow the active user to select whether the active user himself or the AI ​​viewer corresponding to the active user will participate. Alternatively, when a viewer is watching a live broadcast by a certain broadcaster and another broadcaster that the viewer follows starts live broadcasting, the model deployment unit 352 may send a push notification to the user terminal of the viewer, and in the push notification, inquire whether to allow the AI ​​viewer corresponding to the viewer to participate in the live broadcast started by the other broadcaster. Alternatively, the model development section 352 may allow the broadcaster to select an AI viewer before the start of live broadcasting or during live broadcasting, and allow the selected AI viewer to participate in the live broadcasting.

[0099] The model development unit 352 may record interactions between the AI ​​viewer and the distributor during live streaming. The model development unit 352 may provide the avatar generation user with the interactions themselves or summaries of the interactions. By obtaining summaries from the multiple generated AI viewers, the avatar generation user can obtain information about the state of multiple live streaming, the level of excitement, the content of conversations, and the like, without actually participating in those live streaming.

[0100] The operation of the live distribution system equipped with the server 50 having the above configuration will now be described. 20 is a flowchart showing a series of processing steps in a live broadcast in which an AI viewer participates. Server 50 determines whether or not the participation conditions for a specific AI viewer in a specific live broadcast are met (S250). The participation conditions are as described above in the explanation of model deployment unit 352, and for example, the participation conditions are met when an active user taps a thumbnail of a live broadcast and then selects to allow the AI ​​viewer corresponding to the active user to participate.

[0101] If the participation conditions are met (Y in S250), the server 50 allows the specific AI viewer to participate in the specific live broadcast (S252). The server 50 receives video data recording the words and actions of the broadcaster from the broadcaster's user terminal 20 via the network NW (S254). The server 50 extracts the words and actions of the broadcaster from the received video data (S256). The server 50 inputs the extracted words and actions into an ML model corresponding to the specific AI viewer (S258). The server 50 acquires the reaction output by the ML model (S260). The server 50 transmits the acquired reaction to the broadcaster's user terminal 20 and the user terminals of other viewers via the network NW (S262).

[0102] The server 50 determines whether a predetermined exit condition is met for the AI ​​viewer participating in the live broadcast (S264). The exit condition is met, for example, when the favorability rating output by the AI ​​viewer falls below a predetermined threshold. If the exit condition is met (Y of S264), the server 50 causes the specific AI viewer to exit the specific live broadcast (S268). If the exit condition is not met (N of S264), the server 50 determines whether a broadcast end instruction has been received from the broadcaster's user terminal 20 (S266). If not received (N of S266), the process returns to step S254. If received (Y of S266), the process ends.

[0103] FIG. 21 is a representative screen diagram of a live streaming selection screen 800 displayed on the display of the user terminal of an active user. The live streaming selection screen 800 includes thumbnails 802 showing each live streaming in a list of currently viewable live streaming received from the server 50. The non-streaming UI control unit 802 generates the live streaming selection screen 800 based on the list of live streaming acquired from the server 50, and displays it on the display. When the active user specifies or taps a thumbnail 802 on the live streaming selection screen 800, the non-streaming UI control unit 402 superimposes a selection area 804 on the live streaming selection screen 800 to allow the active user to select whether the active user himself or the AI ​​viewer corresponding to the active user will participate in the live streaming corresponding to the specified thumbnail 802. The selection area 804 has an enter button 806 corresponding to the active user's own participation in the live streaming, and a join as a bot button 808 corresponding to the AI ​​viewer corresponding to the active user's participation in the live streaming.

[0104] 22 is a representative screen diagram of a live streaming room screen 808 displayed on the display of a viewer's user terminal 30. The live streaming room screen 808 displays in real time a video image generated by the broadcaster's user terminal 20. The live streaming room screen 808 has a video image 810 of the broadcaster obtained by playing video data received from the server 50, a gift object 812, a comment input area 816, a comment display area 818, and an end viewing button 820.

[0105] FIG. 22 corresponds to a case where an active user (different from the viewer viewing the live streaming room screen 808 in FIG. 22) taps the Join with Bot button 808 on the live streaming selection screen 800 in FIG. 21, and an AI viewer corresponding to the active user joins the live streaming. More specifically, a thumbnail C corresponding to the ongoing live streaming A is displayed on the live streaming selection screen 800 through the live streaming room screen 808 in FIG. 22 that is being viewed by viewer B. An active user D different from viewer B views the live streaming selection screen 800, taps the thumbnail C, and further taps the Join with Bot button 808. Then, a message 820 indicating that an AI viewer of active user D has entered the room is displayed in the comment display area 818 of the live streaming room screen 808 in FIG. 22 that viewer B is viewing.

[0106] According to the live distribution system including the server 50 according to the present embodiment, a viewer can have an AI viewer who performs similar viewing activities on his / her behalf participate in a live distribution. Therefore, even when the viewer is busy, asleep at night, or watching another live distribution, the viewer can maintain his / her presence and connection with the distributor by having an AI viewer participate in the live distribution.

[0107] The hardware configuration of the information processing device according to the first and second embodiments will be described with reference to Fig. 23. Fig. 23 is a block diagram showing an example of the hardware configuration of the information processing device according to the embodiments. The illustrated information processing device 900 can realize, for example, the server 10 and the user terminals 20 and 30 in the first embodiment, and can realize the server 50 in the second embodiment.

[0108] The information processing device 900 includes a CPU 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903. The information processing device 900 may also include a host bus 907, a bridge 909, an external bus 911, an interface 913, an input device 915, an output device 917, a storage device 919, a drive 921, a connection port 925, and a communication device 929. The information processing device 900 further includes an imaging device (not shown) such as a camera. The CPU 901 is an example of a hardware configuration for realizing functions realized by components described in this specification. The functions described in this specification may be realized by a circuit programmed to realize the described functions. The circuit programmed to realize the functions described in this specification includes a CPU (a Central Processing Unit), a DSP (Digital Signal Processor), a general-purpose processor, a specific application processor, an integrated circuit, ASICs (Application Specific Integrated Circuits), and / or a combination thereof. In this specification, a unit that realizes a specific function may be realized as a circuit programmed to realize the function.

[0109] The CPU 901 functions as an arithmetic processing device and a control device, and controls the overall operation or a part of the operation in the information processing device 900 according to various programs recorded in the ROM 902, the RAM 903, the storage device 919, or the removable recording medium 923. For example, the CPU 901 controls the overall operation of each functional unit included in the server 10 and the user terminals 20 and 30 in this embodiment. The ROM 902 stores programs and arithmetic parameters used by the CPU 901. The RAM 903 temporarily stores programs used in the execution of the CPU 901 and parameters that change appropriately during the execution. The CPU 901, the ROM 902, and the RAM 903 are mutually connected by a host bus 907 constituted by an internal bus such as a CPU bus. Furthermore, the host bus 907 is connected to an external bus 911 such as a PCI (Peripheral Component Interconnect / Interface) bus via a bridge 909.

[0110] The input device 915 may be, for example, a device operated by a user, such as a mouse, a keyboard, a touch panel, a button, a switch, or a lever, or may be a device that converts a physical quantity into an electric signal, such as a sound sensor such as a microphone, an acceleration sensor, a tilt sensor, an infrared sensor, a depth sensor, a temperature sensor, or a humidity sensor. The input device 915 may be, for example, a remote control device that uses infrared or other radio waves, or an external connection device 927 such as a mobile phone that supports the operation of the information processing device 900. The input device 915 includes an input control circuit that generates an input signal based on information input by the user or a sensed physical quantity and outputs the input signal to the CPU 901. The user operates the input device 915 to input various data to the information processing device 900 or to instruct a processing operation.

[0111] The output device 917 is configured with a device capable of visually or audibly notifying the user of acquired information. The output device 917 can be, for example, a display such as an LCD, a PDP, or an OELD, an audio output device such as a speaker or a headphone, or a printer device. The output device 917 outputs the results obtained by the processing of the information processing device 900 as a video such as text or an image, or as a sound such as audio.

[0112] The storage device 919 is a data storage device configured as an example of a storage unit of the information processing device 900. The storage device 919 is configured, for example, with a magnetic storage device such as a hard disk drive (HDD), a semiconductor storage device, an optical storage device, or a magneto-optical storage device. The storage device 919 stores programs and various data executed by the CPU 901, and various data acquired from the outside.

[0113] The drive 921 is a reader / writer for a removable recording medium 923 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, and is built into or externally attached to the information processing device 900. The drive 921 reads out information recorded on the attached removable recording medium 923 and outputs the information to the RAM 903. The drive 921 also writes information to the attached removable recording medium 923.

[0114] The connection port 925 is a port for directly connecting a device to the information processing device 900. The connection port 925 can be, for example, a Universal Serial Bus (USB) port, an IEEE 1394 port, a Small Computer System Interface (SCSI) port, or the like. The connection port 925 can also be an RS-232C port, an optical audio terminal, a High-Definition Multimedia Interface (HDMI) (registered trademark) port, or the like. By connecting an external connection device 927 to the connection port 925, various types of data can be exchanged between the information processing device 900 and the external connection device 927.

[0115] The communication device 929 is, for example, a communication interface configured with a communication device for connecting to a network NW. The communication device 929 can be, for example, a communication card for a wired or wireless LAN (Local Area Network), Bluetooth (registered trademark), or WUSB (Wireless USB). The communication device 929 may also be a router for optical communication, a router for ADSL (Asymmetric Digital Subscriber Line), or a modem for various communications. The communication device 929 transmits and receives signals, for example, between the Internet and other communication devices using a predetermined protocol such as TCP / IP. The communication network NW connected to the communication device 929 is a network connected by wire or wirelessly, for example, the Internet, a home LAN, infrared communication, radio wave communication, or satellite communication. The communication device 929 realizes a function as a communication unit.

[0116] An imaging device (not shown) such as a camera is a device that captures real space and generates a captured image using an imaging element such as a CCD (Charge Coupled Device) or a CMOS (Complementary Metal Oxide Semiconductor) and various members such as a lens for controlling the formation of a subject image on the imaging element. The imaging device may capture a still image or a moving image.

[0117] The configuration and operation of the live distribution system according to the embodiment have been described above. This embodiment is merely an example, and it will be understood by those skilled in the art that various modifications are possible in the combination of each component and each process, and that such modifications are also within the scope of the present disclosure.

[0118] In the first embodiment, the ML model may be set to realize situations and scenarios that may occur in live distribution. The server 10 sets the ML model randomly or according to instructions from the distributor so that a specific situation is simulated in the AI ​​training distribution. For example, a scene that is likely to ignite is simulated by forcibly setting the emotions of all participating AI viewers to "anger." Alternatively, a scene that is likely to cause viewers to drop out is simulated by forcibly lowering the average favorability rating of the participating AI viewers. Alternatively, the ML model may be set to enable simulations for different time periods, such as lunch breaks, evenings, and late nights.

[0119] In the first embodiment, a case where multiple AI viewers participate in an AI training distribution has been described, but this is not limited thereto. For example, when a distributor is realized by an ML model, a single live distribution in which multiple AI distributors (distributors realized by an ML model) participate may be provided. In this case, the server inputs the output of an AI distributor to another AI distributor to realize a conversation between the AI ​​distributors. Real users who watch the live distribution can enjoy the conversation between the AI ​​distributors and can also participate by sending gifts and comments. The live distribution may be configured so that the viewer can specify the topic of the live distribution. If an AI viewer participates in the live distribution, it becomes possible to enjoy the interaction between multiple AI distributors and AI viewers.

[0120] In the first embodiment, the broadcaster may be configured to set his / her own level. The server 10 allows AI viewers according to the set level to participate in the AI ​​training broadcast. For example, when the broadcaster's level is set to level 1 (beginner), AI viewers who also have the beginner attribute are allowed to participate in the AI ​​training broadcast. When the broadcaster's level is set to level 100 (professional broadcaster), AI viewers who have the VIP attribute are allowed to participate in the AI ​​training broadcast.

[0121] In the first embodiment, the server may be configured to enable the setting of AI viewers when they are captured as a group. For example, the server may be configured to enable the distributor to set half of the AI ​​viewers to AI viewers with beginner attributes. Alternatively, the server may be configured to enable the distributor to set 70% of the AI ​​viewers to AI viewers (army) who pay a monthly fee. The server may generate patterns of groups of viewers by analyzing the viewing history of actual viewers.

[0122] In the first embodiment, the server may be configured to compare and analyze broadcasters who have practiced using the AI ​​training broadcast with other broadcasters, and output the results.

[0123] In the first embodiment, the evaluation unit may acquire and analyze the sound of the AI ​​training distribution (or a normal live distribution) and perform evaluation. Basically, the sound balance is good in live distribution by a high-quality distributor. For example, there is no silent time or relatively little, and even if the distributor is not talking, the background music is played in a balanced manner, allowing the viewer to spend a comfortable time. The volume balance also affects the quality of the live distribution. For example, if the volume balance is poor (especially if it is too low) compared to other live distributions, it leads to a decrease in the quality of the live distribution. Therefore, the evaluation unit may generate an evaluation of the AI ​​training distribution by acquiring and analyzing the presence or absence and length of silent time in the AI ​​training distribution, and / or the sound balance, and / or the volume balance. Machine learning may be used to adjust the volume balance and voice output in the live distribution so that humans (viewers) find it comfortable.

[0124] In the first embodiment, the server may be configured to allow the broadcaster to specify a specific event and set a goal of entering the rankings in the event. In this case, the server allows the broadcaster to select which event he or she wants to succeed in before starting the AI ​​training broadcast. The server sets AI viewers that are the same as or similar to viewers of live broadcasts by broadcasters who have been successful in the event genre in the past (e.g., in the top 5). The evaluation unit evaluates the AI ​​training broadcast by comparing it with live broadcasts by broadcasters who have been successful in the event genre in the past, and provides feedback. According to this example, the broadcaster can obtain feedback from the perspective of making a dream come true (e.g., wanting to appear on a model's runway). It is possible to meet the broadcaster's desire to be number one in this event, whether as a top or non-top broadcaster, and to broadcast on a live broadcast platform for that purpose. Alternatively, the server may predict which event genre (music, modeling, etc.) is likely to be successful based on the similarity between interactions during live streams of streamers who were ranked in the top 5 or other places in past events of the same genre as the specified event and interactions during the AI ​​training stream, and provide feedback. An example of feedback is, "If you improve here, how high would you have placed in this event in the past?" (determined from the similarity of the content of the stream with the winning streamer and the interaction). Beginners and mid-level streamers can be asked to use AI training streams before participating in an event, and they can be encouraged to participate in similar events, such as event genres where they might win. In this case, motivation is increased, and competition between participating streamers is created, leading to the revitalization of the streaming platform.

[0125] In the processing procedures described in this specification, particularly in the processing procedures described using flow diagrams and flowcharts, it is possible to omit some of the steps that make up the processing procedures, to add steps that are not explicitly stated as steps that make up the processing procedures, and / or to rearrange the order of the steps, and processing procedures in which such omissions, additions, or changes in order have been made are included within the scope of this disclosure as long as they do not deviate from the spirit of this disclosure.

[0126] At least a part of the functions realized by the server 10 may be realized by a device other than the server 10, such as the user terminals 20 and 30. At least a part of the functions realized by the user terminals 20 and 30 may be realized by a device other than the user terminals 20 and 30, such as the server 10.

Claims

1. A means for receiving data including the words and actions of a broadcaster from a terminal of the broadcaster of a live broadcast in which the broadcaster and multiple viewers, including a virtual viewer realized by a machine learning model, participate via a network; A means for acquiring a reaction output by a machine learning model that receives the distributor's words and actions as input and outputs a reaction that a viewer with a set attribute would make; and means for transmitting data to the terminal via the network, for causing an object representing the reaction to be displayed on a display of the terminal of the distributor.

2. The server of claim 1, wherein a plurality of different virtual viewers participate in the live broadcast, each of the virtual viewers is set with corresponding attributes, and each of the virtual viewers is realized by a corresponding machine learning model.

3. The server according to claim 1 , wherein the live distribution is configured so that only virtual viewers can participate as viewers.

4. The server of claim 1 , wherein the response output by the machine learning model includes at least one of the emotions, degree of interest, and likeability of a virtual viewer corresponding to the machine learning model toward the live broadcast or the broadcaster.

5. The server of claim 1 , wherein the reaction output by the machine learning model includes the input of comments and / or the use of gifts in the live broadcast by a virtual viewer corresponding to the machine learning model.

6. The server according to claim 2 , wherein the machine learning model corresponding to each virtual viewer receives as input responses output by machine learning models corresponding to other virtual viewers.

7. A means for evaluating the live stream and / or the broadcaster of the live stream based on responses output by machine learning models corresponding to each of the plurality of virtual viewers; The server according to claim 2 , further comprising: means for transmitting a result of the evaluation to the terminal of the distributor via the network.

8. A means for retaining a viewing history of a real user in the live streaming platform provided by the server; and means for training the machine learning model to learn the viewing history, The server of claim 1 , wherein the machine learning model is trained to output a response that the real user would make.

9. Receiving data including the words and actions of a broadcaster from a terminal of the broadcaster of a live broadcast in which the broadcaster and multiple viewers, including a virtual viewer realized by a machine learning model, participate via a network; Obtaining a reaction output by a machine learning model that inputs the distributor's words and actions and outputs a reaction that a viewer with a set attribute would make; and transmitting data to the terminal via the network for causing an object representing the reaction to be displayed on a display of the terminal of the distributor.

10. A terminal of a broadcaster of a live broadcast in which a plurality of different virtual viewers and broadcasters participate, which are realized by a plurality of different machine learning models, A function of transmitting data including the words and actions of the broadcaster to a server that provides the live broadcast via a network; A function of receiving data for realizing a plurality of reactions output from the plurality of machine learning models by inputting the behavior to the plurality of machine learning models from the server via the network; A computer program for realizing a function of displaying a plurality of objects representing the plurality of reactions on a display based on the data.

Citation Information

Patent Citations

  • Live shopping interestingness prediction method based on eye movement features and DeepFM

    CN114706481A

  • Demand forecasting system

    JP2021103444A

  • Systems and methods for streaming - Patents.com

    JP7497002B1

  • Sales Forecasting using Browsing Ratios and Browsing Durations

    US20170337572A1

  • Server device and electronic commerce method

    WO2021106034A1