Real-time live distribution system

The real-time live streaming system addresses the lack of realism and engagement in AI characters by using a virtual human that interacts dynamically with viewers, enhancing communication and revenue generation through sentiment-based responses and interactive features.

JP2025186961APending Publication Date: 2025-12-24CRYSTAL METHOD CO LTD

Patent Information

Application Number
JP2024095462
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-12
Publication Date
2025-12-24

AI Technical Summary

Technical Problem

Existing real-time live streaming systems with AI characters lack realism and fail to enhance communication between broadcasters and viewers, limiting engagement and revenue generation.

Method used

A real-time live streaming system that utilizes a virtual human AI avatar, capable of extracting viewer comments, performing text analysis, setting emotion parameters, and modifying its behavior based on comment sentiment to enhance interaction and generate revenue through features like singing or dancing in response to tips.

Benefits of technology

Enhances communication and engagement between broadcasters and viewers while generating revenue, all at a low cost, by using a virtual human that reflects the streamer's characteristics and responds dynamically to viewer feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025186961000001_ABST
    Figure 2025186961000001_ABST
Patent Text Reader

Abstract

To provide a real-time distribution system that is capable of enhancing communication between a distributor and a viewer at low costs and enables the distributor to obtain a benefit.SOLUTION: A real-time live distribution system for performing real-time live distribution of an idle talk dialog content toward a plurality of viewer terminals comprises a distributor terminal 1 including: an extraction section 11 having a virtual human for extracting viewer comments input from the plurality of viewer terminals in real time; an analysis section 12 for performing text analysis on the extracted viewer comments; a reading section 13 for reading a text using voice synthesis on the basis of a text analysis result; a feeling parameter setting section 14 for setting a virtual human feeling parameter for categorizing the viewer comments, which are at least a positive comment or negative comment, to determine the positive comment as a plus comment and the negative comment as a minus comment; and a change section 15 for changing a way of speaking according to the set feeling parameter.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a real-time live distribution system, and more particularly to a real-time live distribution system that uses a virtual human. [Background technology]

[0002] Conventionally, there are real-time live streaming systems in which broadcasters and viewers broadcast live content in real time via chat. Real-time live streaming can be divided into two types: one where the content is broadcast live at the broadcast time, and another where the content is stored on a server and streamed at a set date and time.

[0003] In a true live stream, a live person must be present at the live streaming location (live venue) at the time of the live stream. This places restrictions on the streamer, as they must travel to the location. Therefore, live streaming apps have recently become popular as a tool for freely streaming live from anywhere. Live streaming apps are becoming more popular because they can be used by installing them on a smartphone or PC with a camera.

[0004] Live streaming broadcasters can be individuals, talents, entertainers, singers, and many other people. Recently, there has been an increase in live streaming content in which broadcasters are not real people, but virtual humanoid or animal-like characters or avatars who communicate with viewers.

[0005] Because characters and avatars lack realism, some systems extract the voice and movement characteristics of real people, such as individuals, actors, singers, and entertainers, to create virtual humans (also called digital humans) that closely resemble the real people, and then use these virtual humans to perform real-time live streaming. Real-time live streaming genres include chatting, musical instruments, karaoke, games, cosplay, dance, official broadcasts, home drinking, game commentary, and reading. For example, there is a real-time live streaming system that creates an "AI VTuber" that responds to comments in the chat box and broadcasts live in real time (Non-Patent Document 1). Here, an AI VTuber refers to a virtual YouTuber that communicates with viewers primarily through AI (artificial intelligence). Non-Patent Document 1 discloses technology that allows an AI VTuber to live stream on YouTube (registered trademark) and automatically respond to comments in the chat box during the live stream. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Japanese Patent Application Publication No. 2023-124275 [Non-patent literature]

[0007] [Non-Patent Document 1] https: / / zenn.dev / makunugi / articles / a4ed9e142526f2 Summary of the Invention [Problem to be solved by the invention]

[0008] According to the technology disclosed in Non-Patent Document 1, it is possible to automatically obtain comments entered by viewers in the chat box, read the comments aloud, and converse with viewers, but there is a problem in that the image of the AI ​​character can only perform simple movements, making it difficult to liven up communication between the broadcaster and viewers.

[0009] Furthermore, as described in paragraph

[0087] , Patent Document 1 discloses a technology for conducting a chat live broadcast by selecting the "chat" genre on the genre-specific live broadcast selection screen on the viewer's user terminal, but no consideration is given to enhancing communication between the broadcaster and the viewer.

[0010] The present invention has been made in consideration of the above problems, and its purpose is to provide a real-time live distribution system that can enhance communication between distributors and viewers at low cost and also generate revenue for distributors. [Means for solving the problem]

[0011] The real-time live streaming system according to the first invention is a real-time live streaming system in which a distributor terminal for distributors and multiple viewer terminals for viewers are connected via a network, and which performs real-time live streaming of chat conversation content from the distributor terminal to the multiple viewer terminals, wherein the distributor uses a virtual human, which is an AI avatar of the distributor himself / herself, as the distributor, and the distributor terminal has: extraction means for extracting viewer comments input from the multiple viewer terminals in real time using the virtual human; analysis means for performing text analysis of the viewer comments extracted by the extraction means; reading means for reading out text using voice synthesis based on the text analysis results analyzed by the analysis means; emotion parameter setting means for classifying the viewer comments as at least positive comments or negative comments, and setting emotion parameters of the virtual human such that positive comments are plus comments and negative comments are minus comments; and modification means for changing the speaking style based on the emotion parameters set by the emotion parameter setting means.

[0012] A real-time live distribution system according to a second aspect of the present invention is the first aspect of the present invention, characterized in that the readout means preferentially picks up and reads out viewer comments that are charged for, such as tips.

[0013] The real-time live distribution system according to the third aspect of the present invention is the first or second aspect of the present invention, characterized in that the virtual human has an activation means that activates a specified action when control logic is set.

[0014] A real-time live distribution system according to a fourth aspect of the present invention is the system of the first aspect, wherein the activation means causes the virtual human to blow a conch shell when a certain number of negative comments are input.

[0015] A real-time live distribution system according to a fifth aspect of the present invention is the system of the first aspect, wherein the virtual human has a change means for changing facial expressions.

[0016] The real-time live streaming system according to the sixth invention is characterized in that, in the first invention, the virtual human reflects the facial image, voice, movements, hobbies, thoughts, self-awareness, and past real-time information of the streamer himself / herself.

[0017] A seventh aspect of the present invention is a real-time live distribution system according to the first aspect of the present invention, characterized in that the virtual human sings a song or performs a dance movement in response to the viewer comments with a fee. [Effects of the Invention]

[0018] According to the first to seventh inventions, it is possible to realize a real-time distribution system that can enhance communication between distributors and viewers at low cost and also earn revenue for distributors. [Brief explanation of the drawings]

[0019] [Figure 1] FIG. 1 is a diagram showing the configuration of a real-time live distribution system to which the present invention is applied. [Figure 2] FIG. 2 is a functional block diagram of the distributor terminal of FIG. [Figure 3] FIG. 3 is a functional block diagram of the virtual human part in the embodiment of the present invention shown in FIG. [Figure 4] FIG. 4 is a diagram showing an example of a display screen of real-time live streaming video content in an embodiment of the present invention. [Figure 5] FIG. 5 is an explanatory diagram of real-time live streaming video content according to an embodiment of the present invention. [Figure 6] FIG. 6 is an explanatory diagram of a triggering unit for real-time live streaming video content in an embodiment of the present invention. [Figure 7] FIG. 7 is a schematic explanatory diagram of a typical real-time live distribution system. [Figure 8] FIG. 8 is a diagram showing an example of emotion parameters according to the embodiment of the present invention. [Figure 9] FIG. 9 is a processing flowchart of a broadcaster terminal of a real-time live distribution system according to an embodiment of the present invention. [Figure 10] FIG. 10 is a processing flowchart of the virtual human unit in the distributor terminal of FIG. [Figure 11] FIG. 11 is a diagram showing an example of learning data of a distributor himself / herself. [Figure 12] FIG. 12 is a diagram showing an example of learning data of the broadcaster's past real information. [Figure 13] FIG. 13 is a diagram showing the configuration of a real-time live distribution system according to another embodiment of the present invention, in which a live commerce unit is added to a distributor terminal. DETAILED DESCRIPTION OF THE INVENTION

[0020] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, exemplary embodiments of the present invention will be described with reference to the accompanying drawings.

[0021] First, a real-time live distribution system 100 according to an embodiment of the present invention will be described with reference to FIGS. 1, 2, 3, and 4. FIG.

[0022] Fig. 1 is a configuration diagram of a real-time live streaming system 100 to which the present invention is applied, and Fig. 2 is a functional block diagram of the broadcaster terminal 1 of Fig. 1. Fig. 3 is a functional block diagram of the virtual human unit 10 in the broadcaster terminal 1 of Fig. 2. Fig. 4 is a diagram showing an example of a display screen for real-time live streaming video content in an embodiment of the present invention, Fig. 5 is an explanatory diagram of real-time live streaming video content in an embodiment of the present invention, and Fig. 6 is an explanatory diagram of an activation unit for real-time live streaming video content in an embodiment of the present invention.

[0023] As shown in Figure 1, a real-time live streaming system 100 to which the present invention is applied comprises a distributor terminal 1 and a plurality of viewer terminals 2a, 2b, ... 2n connected to a network 3, and a real-time live streaming system for chat dialogue content that is transmitted from the distributor terminal 1 to the plurality of viewer terminals 2a, 2b, ... 2n via the network 3 and a video streaming server 4.

[0024] The real-time live distribution system 100 will be described using an example of chat dialogue content, but the type of content provided by a distributor through live distribution etc. is not limited to chat dialogue content. For example, the distributed content may be news, songs, education, acting, cooking, etc.

[0025] The broadcaster terminal 1 broadcasts chat content and other live content in real time. The broadcaster terminal 1 includes a virtual human unit 10, an extraction unit 11, an analysis unit 12, a reading unit 13, an emotion parameter setting unit 14, and a modification unit 15. Figure 4 shows an example of a display screen for chat content (web page) broadcasting live in real time. In the example shown in Figure 4, the virtual human AI Jun is displayed in the center of the display screen, and input viewer comments are displayed in real time on the web screen. A chat replay of the top videos is displayed outside the video display screen. When viewers pay for tips, the viewer comments are displayed at the top, and thank-you comments are also displayed. As shown in Figure 5, the entire display screen is generated by AI, and the generated AI reads out and responds to comments from viewers in real time. Setting the control logic activates a specified action, as shown in Figure 6. For example, if a certain number of negative comments are received, the AI ​​may blow a conch shell. It may also sing or dance.

[0026] The virtual human unit 10, as the broadcaster, causes the virtual human, an AI avatar of the broadcaster, to perform various actions. For example, it is an AI avatar that looks, moves, and talks like the broadcaster. Such a virtual human can be created at low cost using the AI ​​avatar generation software "Deep AI (registered trademark)." Various information about the broadcaster is collected as AI learning data on behalf of the broadcaster, and the virtual human engages in casual conversations through its own inference learning. Because learning data is used, there will be some discrepancy with the broadcaster's actual information. The virtual human unit 10 collects the broadcaster's facial image, voice, actions, hobbies, thoughts, self-awareness, and past real-life information as basic information in the learning data, so it can reflect actions that are close to the broadcaster's own.

[0027] In the extraction unit 11, the virtual human of the virtual human unit 10 extracts viewer comments input from a plurality of viewer terminals 2a, 2b, . . . , 2n in real time.

[0028] The analysis unit 12 performs text analysis on the viewer comments extracted by the extraction unit 11.

[0029] The reading unit 13 reads out the text by voice synthesis based on the text analysis result analyzed by the analysis unit 12. For example, the reading unit 13 may be configured to preferentially pick up and read out viewer comments with a fee, such as tips.

[0030] The emotion parameter setting unit 14 classifies viewer comments into positive comments, neutral normal comments, and negative comments, and sets emotion parameters for the virtual human such that positive comments are plus comments, neutral normal comments are zero comments, and negative comments are minus comments. Note that neutral comments may be omitted and viewer comments may be classified into two types: positive comments and negative comments.

[0031] The change unit 15 changes the speaking style according to the emotion parameters set by the emotion parameter setting unit 14.

[0032] The virtual human unit 10 includes an actuator 21, a facial expression changer 22, a reflector 23, and an action unit 24.

[0033] When the control logic is set, the initiation unit 21 initiates a specified action. For example, the initiation unit 21 may be configured to make the virtual human blow a conch shell when a certain number of negative comments are input.

[0034] The facial expression change unit 22 changes the facial expression of the virtual human, for example, by changing the facial expression according to the emotion parameter values ​​set in the emotion parameter setting unit based on the analysis results of the viewer comments.

[0035] The reflecting unit 23 reflects the face image, voice, movements, hobbies, thoughts, self-awareness, and past real information of the distributor himself / herself in the virtual human.

[0036] The operation unit 24 controls the virtual human to sing or dance in response to viewer comments for which a fee is charged. The operation performed by the operation unit 24 is not limited to singing or dancing. For example, if the distributor is a rakugo performer, the virtual human may be configured to recite rakugo.

[0037] The viewer terminals 2a, 2b, 2n are terminals such as mobile communication terminals (e.g., smartphones), tablet PCs, laptop PCs, and desktop PCs. Any device capable of viewing a web browser can be used. It is also possible to install a real-time live streaming app and view chat dialogue content. Services such as YouTube (registered trademark) and TikTok (registered trademark) can be used as real-time live streaming systems.

[0038] The network 3 may be a wired network or a wireless network. If the network 3 is connected via the Internet, broadband, a mobile network, or the like, it is possible to view real-time live streaming video content.

[0039] The video distribution server 4 can be any of various video distribution servers, for example, YouTube (registered trademark) or TikTok (registered trademark) servers can be used.

[0040] The cloud 5 can use a general-purpose document generation AI 6 for viewer comments. The distributor terminal 1 also has a server function, and can capture a window of video content on the server and distribute it to the video distribution server 4.

[0041] The general-purpose document generation AI 6 is a document generation AI such as Chat-GPT, and can be used to generate natural language documents that serve as the basis for speech synthesis.

[0042] FIG. 7 is an explanatory diagram outlining a typical real-time live streaming system. As shown in FIG. 7, open the web page (see FIG. 4) on the broadcaster terminal 1 in OBS (Open Broadcaster Software) and capture the browser screen on the main unit screen using window capture. Next, set the audio output. Next, set the stream key obtained in YouTube (registered trademark) Studio in OBS. Once the settings are complete, click "Start Streaming" in the lower right corner of OBS. This will allow you to live stream on YouTube (registered trademark).

[0043] In the present invention, the basic flow of distribution operations of the real-time live distribution system is the same as that shown in Fig. 7. Viewer comments posted on YouTube (registered trademark) can be acquired in real time using the YouTube (registered trademark) Data API, as shown in Fig. 7.

[0044] The characteristic emotion parameters of the present invention will be described below. An example of emotion parameters is shown in FIG. 8. In the example of FIG. 8, emotion parameters are categorized into three types. The emotion parameters store the character, emotion parameter type, and indicator, and three types of characters are displayed during real-time live streaming, with indicators showing the level of each. The emotion parameter types "emotion happy," "emotion neutral," and "emotion bad" are always displayed during live streaming. In the example of FIG. 8, emotion parameters are categorized into three types, but they may be categorized into four or more types. For example, emotion parameters such as anger and sadness may be set.

[0045] Emotion parameters are detected for viewer comments and classified according to the content of the comment. For example, the emotion parameter setting unit 14 classifies viewer comments as at least positive or negative, and sets emotion parameters for the virtual human such that positive comments are plus comments and negative comments are minus comments. The emotion parameter setting unit 14 may further classify neutral or normal comments and display neutral or normal comments as zero comments.

[0046] According to the emotion parameters, learning data for tagged good mood, learning data for tagged neutral mood, and learning data for tagged bad mood are created, and the virtual human's speech style is changed depending on which of the three categories the viewer comments fall into. At this time, utterances made when in a good mood, neutral mood, and bad mood are created as inference data, and inference is made according to the emotion parameters.

[0047] FIG. 9 is a processing flowchart of the distributor terminal 1 of FIG.

[0048] The distributor terminal 1 has an extraction step 101 , an analysis step 102 , a reading-out step 103 , an emotion parameter setting step 104 , and a change step 105 .

[0049] In the extraction step 101, a virtual human extracts viewer comments input from a plurality of viewer terminals 2a, 2b, . . . , 2n in real time.

[0050] In the analysis step 102, the viewer comments extracted in the extraction step 101 are subjected to text analysis.

[0051] In the reading step 103, the text is read aloud by voice synthesis based on the text analysis result obtained in the analysis step 102.

[0052] In the emotion parameter setting step 104, the viewer comments are classified as at least positive or negative, and emotion parameters of the virtual human are set so that positive comments are plus comments and negative comments are minus comments. Furthermore, neutral normal comments may also be set.

[0053] In the change step 105, the speaking style is changed according to the emotion parameter set in the emotion parameter setting step 104. For example, if the number of negative comments exceeds a certain number, the emotion parameter display becomes "Emotion Bad," and the character's face changes to look angry. At this time, the speaking style changes to look slightly annoyed. This speaking style can be changed for each virtual human, as shown in FIG. 11. As shown in FIG. 11, the broadcaster terminal 1 has multiple virtual humans of the broadcaster himself, and information such as the number, virtual human name, facial image, voice, hobbies, thoughts, and self-perception is collected as learning data. For example, "AI Jun" with No. "001" stores information such as the facial image, voice data of the broadcaster himself, hobbies such as visiting castles, positive thinking, not worrying, self-perception such as bad at sports, and calm and collected, as shown in FIG. 4. Similarly, information such as "AI ○○" with No. "002" and "AI △△" with No. "003" is also stored. As shown in FIG. 12, information such as "Era," "Episode a," "Episode b," "Episode c," "Episode d," and "Episode e" is collected as real information about the streamer's past. For example, for elementary school students, episodes about school, sports day, bus trips, club activities, and outside of school are collected. Other episodes may include notable episodes of the streamer. For example, information about going to Koshien Baseball Stadium or winning the Informatics Olympiad may be collected.

[0054] FIG. 10 is a processing flowchart of the virtual human unit 10 in FIG.

[0055] The virtual human unit 10 includes an activation step 201 , an expression change step 202 , a reflection step 203 , and an action step 204 .

[0056] In the activation step 201, a control logic is set for the virtual human to activate a specified action. For example, in the activation step 201, when a certain number of negative comments as described above are input, the virtual human blows a conch shell.

[0057] The facial expression of the virtual human being being displayed is changed in the expression change step 202. The facial expression may be changed randomly or in accordance with the state of an emotion parameter.

[0058] In the reflection step 203, the virtual human is operated to reflect the face image, voice, movements, hobbies, thoughts, and self-awareness of the distributor himself / herself as shown in FIG. 11, and the distributor's past real information as shown in FIG.

[0059] The act step 204 may configure the virtual human to sing or perform dance movements in response to paid viewer comments.

[0060] FIG. 13 is a diagram showing the configuration of a distributor terminal 1a in which a live commerce unit 16 is added to the configuration of the distributor terminal 1 in FIG.

[0061] The distributor terminal 1a includes a virtual human unit 10, an extraction unit 11, an analysis unit 12, a reading unit 13, an emotion parameter setting unit 14, a change unit 15, and a live commerce unit 16.

[0062] The configuration of Fig. 13 is the same as that of Fig. 2 except for the live commerce unit 16, so a description thereof will be omitted. The live commerce unit 16 will be described below.

[0063] A real-time live distribution system according to another embodiment of the present invention is similar to that shown in FIG. 1, but has an added function of displaying an online purchase screen by a live commerce unit 16 during live distribution and selling products.

[0064] This will lead to further profitability for the broadcaster. The live commerce unit 16 may present the broadcaster's hobbies, products recommended by the broadcaster, and products related to the live broadcast content. It may also display an affiliate screen.

[0065] In the above embodiment, a real-time distribution system can be realized that can enhance communication between distributors and viewers at low cost and also generate revenue for distributors. Furthermore, by adding a live commerce unit to the distributor terminal 1a, products that have caught the viewers' interest or products recommended by the distributor can be purchased online during real-time live video distribution, thereby further increasing the distributor's profitability. [Explanation of symbols]

[0066] 1, 1a Streamer terminal 2a, 2b, 2n Viewer terminal 3 Network 4. Video distribution server 5. Cloud 6 General-purpose document generation AI 10 Virtual Human Club 11 Extraction part 12 Analysis Department 13 Reading Section 14 Emotion parameter setting section 15 Changes 16 Live Commerce Department 21. Movement unit 22 Facial Expression Change Unit 23 Reflection section 24 Operating unit 100 Real-time live streaming system 101 Extraction Steps 102 Analysis Steps 103 Reading Steps 104 Emotion parameter setting section 105 Change Steps 201 Activation Step 202 Facial Expression Change Steps 203 Reflection Step 204 Operation Steps

Claims

1. A real-time live distribution system in which a distributor terminal for a distributor and a plurality of viewer terminals for viewers are connected via a network, and chat dialogue content is distributed in real time from the distributor terminal to the plurality of viewer terminals, As the broadcaster, a virtual human, which is an AI avatar of the broadcaster, is used, The distributor terminal is an extraction means for extracting viewer comments input from the plurality of viewer terminals in real time by the virtual human; an analysis means for performing text analysis on the viewer comments extracted by the extraction means; a reading means for reading aloud the text by voice synthesis based on the text analysis result analyzed by the analyzing means; an emotion parameter setting means for classifying the viewer comments as at least positive comments or negative comments, and setting emotion parameters of the virtual human such that positive comments are plus comments and negative comments are minus comments; a changing means for changing the speaking style in accordance with the emotion parameters set by the emotion parameter setting means; A real-time live distribution system comprising:

2. 2. The real-time live distribution system according to claim 1, wherein the reading means gives priority to and reads out comments from viewers that are charged for, such as tips.

3. 3. The real-time live distribution system according to claim 1, wherein the virtual human has an initiation means for initiating a specified action when control logic is set.

4. 2. The real-time live distribution system according to claim 1, wherein the activation means causes the virtual human to blow a conch shell when a certain number of negative comments are input.

5. 2. The real-time live distribution system according to claim 1, wherein the virtual human has a change means for changing facial expressions.

6. 2. The real-time live distribution system according to claim 1, wherein the virtual human reflects the face image, voice, movements, hobbies, thoughts, self-awareness, and past real information of the distributor himself / herself.

7. 2. The real-time live distribution system according to claim 1, wherein the virtual human sings or dances in response to the viewer's comments, which are charged for.

Citation Information

Patent Citations

  • server

    JP2023124275A

Cited By

  • Digital human intelligent interaction control method and system in private domain live broadcast scene

    CN122027840A