Systems and methods for live streaming
The live streaming system uses AI to generate real-time, unpredictable video and audio effects based on broadcaster emotions and viewer interactions, addressing the challenge of maintaining user engagement by offering dynamic content.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- 17LIVE JAPAN INC
- Filing Date
- 2024-11-06
- Publication Date
- 2026-05-19
AI Technical Summary
Existing live streaming platforms struggle to maintain user engagement by consistently providing new or unprecedented content and services, as they are limited to predetermined effects and interactions.
A live streaming system that utilizes machine learning and AI models to generate dynamic and real-time video and audio effects based on broadcaster emotions and viewer interactions, allowing for automatic and unpredictable changes in video and audio data.
Enhances user engagement by providing dynamic and engaging content that encourages viewers to interact more and broadcasters to be more expressive, thereby improving the overall experience.
Smart Images

Figure 2026081972000001_ABST
Abstract
Description
Technical Field
[0007] ,
[0001] The present invention relates to streaming, and particularly to live streaming.
Background Art
[0002] As represented by live streaming services, real-time communication on the Internet has penetrated into daily life. Various platforms and providers offer live streaming services, and the competition is fierce. For a platform, it is important to provide the services desired by users.
[0003] Taiwan Patent Application Publication No. TW202231071A discloses a system for live streaming.
Summary of the Invention
[0004] A method according to an embodiment of the present invention is a method for live streaming executed by one or more computers, including the step of obtaining delivery data of live streaming, and the step of generating an effect not limited to predetermined options based on the delivery data.
[0005] A system according to an embodiment of the present invention is a system for live streaming, including one or more processors, and the one or more processors execute machine-readable instructions to perform the step of obtaining delivery data of live streaming and the step of generating an effect not limited to predetermined options based on the delivery data. [[ID=II]]
[0006] A non-transitory computer-readable medium including a program for live streaming, wherein the program causes one or more computers to perform the step of obtaining delivery data of live streaming and the step of generating an effect not limited to predetermined options based on the delivery data.
Brief Description of the Drawings
[0007] [Figure 1] This is a schematic diagram showing the configuration of a live streaming system 1 based on some embodiments of the present invention. [Figure 2] This block diagram shows the functions and configuration of the user terminal 30 in Figure 1, based on some embodiments of the present invention. [Figure 3] This block diagram shows the functions and configuration of the server in Figure 1, based on some embodiments of the present invention. [Figure 4] Figure 3 is a table showing an example data structure for the stream DB310. [Figure 5] Figure 3 is a table showing an example data structure for user DB312. [Figure 6] Figure 3 is a table showing an example data structure for the gift DB314. [Figure 7] This table shows an example data structure for Stream DB310. [Figure 8] This table shows an example data structure for emotion DB332. [Figure 9] This table shows an example data structure for Effect DB334. [Figure 10] This flowchart shows an exemplary flow based on some embodiments of the present invention. [Figure 11] This is an example of real-time effect generation. [Figure 12] This is an example of generating sentiment descriptions based on specified weights. [Figure 13] This is an example of generating emotion descriptions and modified video data. [Figure 14] This is an example of generating emotion descriptions and modified videos. [Figure 15] This is an example of saving modified effects or correlation data. [Figure 16] This is a block diagram showing an example of the hardware configuration of an information processing device based on some embodiments of the present invention. [Modes for carrying out the invention]
[0008] In the following drawings, identical or similar components, members, procedures, or signals are denoted by the same reference numerals, and redundant explanations are omitted where appropriate. Furthermore, some components that are not essential to the description of each drawing are omitted.
[0009] User engagement, including streaming time and viewing time, is crucial for the operation of a streaming platform. It is desirable for the platform to consistently provide new or unprecedented content and services to maintain user engagement.
[0010] Figure 1 is a schematic diagram showing the configuration of a live streaming system 1 based on some embodiments of the present invention. The live streaming system 1 provides a live streaming service for real-time interaction or communication between streaming broadcasters (also called live streamers, anchors, or live broadcasters) LV and viewers (also called audiences) AU (AU1, AU2...). As shown in Figure 1, the live streaming system 1 includes a server 10, user terminals 20 and 30 (30a, 30b...). In some embodiments, the broadcasters and viewers may be collectively referred to as users. The server 10 may include one or more information processing devices connected to a network NW. The user terminals 20 and 30 may be mobile devices such as smartphones, tablets, notebook PCs, recorders, portable game consoles, or wearable devices, or stationary devices such as desktop PCs. The server 10, user terminals 20 and 30 are connected to each other so as to be able to communicate with each other via various wired or wireless networks NW.
[0011] The live streaming system 1 includes a broadcaster (LV), a viewer (AU), and an administrator (or application provider, not shown) who manages the server 10. The broadcaster (LV) is a person who broadcasts content in real time by recording content on their user terminal 20 and uploading it directly or indirectly to the server 10. Examples of such content may include the broadcaster's own songs, talks, performances, gameplay, or any other content. The administrator provides a platform on the server 10 for live streaming of such content and mediates or manages real-time interaction between the broadcaster (LV) and the viewer (AU). The viewer (AU) accesses the platform on their user terminal 30 and selects and watches the desired content. While the selected content is being live-streamed, the viewer (AU) performs actions such as sending comments, cheers, and gifts via their user terminal 30. The broadcaster (LV) broadcasting the content may respond to these comments, cheers, or gifts. Such responses are sent to the viewer (AU) via video and / or audio, establishing two-way communication.
[0012] The term "live streaming" may refer to a data transmission mode that enables content recorded on the user terminal 20 of the broadcaster LV to be played back and viewed substantially in real time on the user terminal 30 of the viewer AU, or it may refer to a live broadcast realized by such a transmission mode. The live stream may be realized using existing live streaming technologies such as HTTP live streaming, CMAF (Common Media Application Format), WebRTC (Web Real-Time Communications), RTMP (Real-Time Messaging Protocol), etc. The live stream includes a transmission mode that allows the viewer AU to view the content with a predetermined delay, simultaneously with the recording of the content by the broadcaster LV. The length of the delay may be such that communication between the broadcaster LV and the viewer AU can be established. The live stream is distinguished from so-called on-demand streaming, in which all recorded data of the content is stored on a server and then provided to the user from the server upon user request.
[0013] The term "video data" as used herein refers to data that includes image data (also called video data) generated using the imaging function of the user terminal 20 or 30, and audio data generated using the audio input function of the user terminal 20 or 30. The video data is played back on the user terminals 20 and 30 so that the user can view the content. In some embodiments, it is assumed that between the generation of the video data on the broadcaster's user terminal and the playback of the video data on the viewer's user terminal, processing is performed on the video data to change its format, size, or data specifications, such as compression, decompression, encoding, decoding, or transcoding. However, since the content represented by the video data (e.g., video and audio) does not substantially change before and after such processing, in this specification, the video data after such processing is described as being identical to the video data before such processing. In other words, when video data is generated on the broadcaster's user terminal and then played back on the viewer's user terminal via the server 10, the video data generated on the broadcaster's user terminal, the video data passing through the server 10, and the video data received and played back on the viewer's user terminal are all the same video data.
[0014] In the example shown in Figure 1, the broadcaster LV provides live streaming data. The broadcaster LV's user terminal 20 records the broadcaster LV's video and audio to generate the streaming data, and the generated data is transmitted to the server 10 via the network NW. At the same time, the user terminal 20 displays the recorded video VD of the broadcaster LV on its display, allowing the user terminal 20 to check the live streaming content currently being broadcast by the broadcaster LV.
[0015] Each of the user terminals 30a and 30b of the viewers AU1 and AU2 that requested to view the live broadcast of the delivery person LV on the platform receives video data related to the live broadcast (hereinafter may be referred to as "video data of the live broadcast") via the network NW, plays back the received video data, displays the videos VD1 and VD2 on the display, and outputs audio from the speaker. The videos VD1 and VD2 respectively displayed on the user terminals 30a and 30b are substantially the same as the video VD captured by the user terminal 20 of the delivery person LV, and the audio output from the user terminals 30a and 30b is substantially the same as the audio recorded by the user terminal 20 of the delivery person LV.
[0016] The recording of video and audio on the user terminal 20 of the delivery person LV and the playback of video data on the user terminals 30a and 30b of the viewers AU1 and AU2 are performed substantially simultaneously. When the viewer AU1 inputs a comment on the content provided by the delivery person LV into the user terminal 30a, the server 10 displays the comment on the user terminal 20 of the delivery person LV in real time, and also displays the comment on the user terminals 30a and 30b of the viewers AU1 and AU2 respectively. When the delivery person LV reads the comment and develops a talk corresponding to the comment, the video and audio of the talk are displayed on the user terminals 30a and 30b of the viewers AU1 and AU2 respectively. This interactive operation is recognized as a conversation being established between the delivery person LV and the viewer AU1. Thereby, in the live streaming system 1, a live broadcast that enables two-way communication instead of one-way communication is realized.
[0017] FIG. 2 is a block diagram showing the functions and configuration of the user terminal 30 of FIG. 1 based on some embodiments of the present invention. The user terminal 20 has the same or similar functions and configuration as the user terminal 30. Each block in FIG. 2 and subsequent block diagrams may be implemented by hardware with elements such as a computer's CPU or mechanical devices, or may be implemented by software with a computer program or the like. The functional blocks may be realized by the coordinated operation between these elements. Therefore, those skilled in the art will understand that these functional blocks can be realized in various forms by the combination of hardware and software.
[0018] The distributor LV and the viewer AU may download a live distribution application program (hereinafter referred to as the live distribution application) from a download site via the network NW and install it on the user terminals 20 and 30. Alternatively, the live distribution application may be pre-installed on the user terminals 20 and 30. When the live distribution application is executed on the user terminals 20 and 30, the user terminals 20 and 30 communicate with the server 10 via the network NW and implement or execute various functions. Hereinafter, the functions implemented by the user terminals 20 and 30 (processors such as CPUs) on which the live distribution application is executed will be described as the functions of the user terminals 20 and 30. These functions are actually realized by the live distribution application on the user terminals 20 and 30. In some embodiments, these functions may be described in a programming language such as HTML (HyperText Markup Language), transmitted from the server 10 to the web browsers of the user terminals 20 and 30 via the network NW, and realized by a computer program executed by the web browsers.
[0019] The user terminal 30 includes a distribution unit 100 and a viewing unit 200. The distribution unit 100 generates video data in which the video and audio of the user (or the user's side) are recorded, and provides the video data to the server 10. The viewing unit 200 receives the video data from the server 10 and plays the video data. The user activates the distribution unit 100 when performing a live stream, and activates the viewing unit 200 when viewing the video. The user terminal on which the distribution unit 100 is activated is the streamer's terminal, i.e., the user terminal that generates the video data. The user terminal on which the viewing unit 200 is activated is the viewer's terminal, i.e., the user terminal on which the video data is reproduced and played.
[0020] The distribution unit 100 includes an imaging control unit 102, an audio control unit 104, a video transmission unit 106, and a broadcaster-side UI control unit 108. The imaging control unit 102 is connected to a camera (not shown in Figure 2) and controls the imaging performed by the camera. The imaging control unit 102 acquires image data from the camera. The audio control unit 104 is connected to a microphone (not shown in Figure 2) and controls the audio input from the microphone. The audio control unit 104 acquires audio data from the microphone. The video transmission unit 106 transmits video data, including the image data acquired by the imaging control unit 102 and the audio data acquired by the audio control unit 104, to the server 10 via the network NW. The video data is transmitted in real time by the video transmission unit 106. In other words, the generation of the video data by the imaging control unit 102 and the audio control unit 104, and the transmission of the generated video data by the video transmission unit 106, are performed substantially simultaneously. The broadcaster-side UI control unit 108 controls the broadcaster's UI (user interface). The broadcaster-side UI control unit 108 may be connected to a display (not shown in Figure 2), and by playing the video data transmitted by the video transmission unit 106, it displays video on the display. The broadcaster-side UI control unit 108 may display operation objects and instruction permission objects on the display and accept input from the broadcaster by tapping these objects.
[0021] The viewing unit 200 includes a viewer-side UI control unit 202, an overlay information generation unit 204, and an input information transmission unit 206. The viewing unit 200 receives video data related to live streaming from the server 10 via the network NW, in which the broadcaster, the viewer (a user of the user terminal 30), and other viewers participate. The viewer-side UI control unit 202 controls the viewer's UI. The viewer-side UI control unit 202 is connected to a display and a speaker (not shown in Figure 2), and plays the received video data to display the video on the display and output sound from the speaker. The state in which the video is output to the display and the sound is output from the speaker can be called the state in which "video data is being played". The viewer-side UI control unit 202 is also connected to input means such as a touch panel, keyboard, and display (not shown in Figure 2), and acquires user input via these input means. The overlay information generation unit 204 overlays a predetermined frame image onto the image generated from the video data from the server 10. The frame image includes various user interface objects (hereinafter simply referred to as "objects") for receiving input from the user, comments entered by the viewer, and / or information obtained from the server 10. The input information transmission unit 206 transmits the user input obtained by the viewer-side UI control unit 202 to the server 10 via the network NW.
[0022] Figure 3 shows a block diagram illustrating the functions and configuration of the server 10 in Figure 1, based on some embodiments of the present invention. The server 10 includes a distribution information provision unit 302, a relay unit 304, a gift processing unit 306, a payment processing unit 308, a stream DB 310, a user DB 312, a gift DB 314, an acquisition unit 320, a processing unit 322, an emotion DB 332, an effect DB 334, and a machine learning DB 336.
[0023] When the broadcaster receives a notification or request to start a live stream via the network NW from the user terminal 20 on the broadcaster's side, the broadcast information provision unit 302 registers the stream ID for identifying the live stream and the broadcaster ID of the broadcaster performing the live stream in the stream DB 310.
[0024] When the distribution information provision unit 302 receives a request for information regarding live streaming from the viewing unit 200 of the user terminal 30 on the viewer side via the network NW, the distribution information provision unit 302 retrieves or checks currently available live streaming from the stream DB 310 and creates a list of available live streaming. The distribution information provision unit 302 sends the created list to the requesting user terminal 30 via the network NW. The viewer-side UI control unit 202 of the requesting user terminal 30 generates a live streaming selection screen based on the received list and displays it on the user terminal 30's display.
[0025] When the input information transmission unit 206 of the user terminal 30 receives the viewer's selection result on the live streaming selection screen, the input information transmission unit 206 generates a distribution request including the stream ID of the selected live streaming and sends the request to the server 10 via the network NW. The distribution information provision unit 302 starts providing the live streaming specified by the stream ID included in the received distribution request to the requesting user terminal 30. The distribution information provision unit 302 updates the stream DB 310 to include the user ID of the viewer of the requesting user terminal 30 in the viewer ID of the stream ID (or the corresponding viewer ID).
[0026] The relay unit 304 relays the video data from the broadcaster's user terminal 20 to the viewer's user terminal 30 in the live stream initiated by the distribution information provision unit 302. During the live stream or playback of the video data, the relay unit 304 receives a signal from the input information transmission unit 206 representing user input from the viewer. The signal representing user input may be an object designation signal that specifies an object displayed on the user terminal 30's display. The object designation signal may include the viewer's viewer ID, the broadcaster ID of the broadcaster of the live stream being viewed by the viewer, and an object ID that identifies the object. If the object is a gift, the object ID is the gift ID. Similarly, the relay unit 304 receives a signal from the distribution unit 100 of the user terminal 20 representing user input made by the broadcaster during playback of the video data (or during the live stream). This signal may be an object designation signal.
[0027] Furthermore, the signal representing the user input may also be a comment input signal that includes a comment entered by the viewer into the user terminal 30 and the viewer's viewer ID. Upon receiving the comment input signal, the relay unit 304 transmits the comment and the viewer ID included in the signal to the broadcaster's user terminal 20 and to the user terminals 30 of other viewers. In these user terminals 20 and 30, the viewer-side UI control unit 202 and the overlay information generation unit 204 display the received comment on a display associated with the viewer ID that was also received.
[0028] The gift processing unit 306 updates the user database 312 by increasing the points of the sender based on the points of the gift identified by the gift ID included in the object designation signal. Specifically, the gift processing unit 306 refers to the gift database 314 to identify the points to be awarded for the gift ID included in the received object designation signal. Subsequently, the gift processing unit 306 updates the user database 312 and adds the identified points to the points of the sender ID (or the points corresponding to the sender ID) included in the object designation signal.
[0029] The payment processing unit 308 processes payment for a gift from a viewer in response to receiving the object designation signal. Specifically, the payment processing unit 308 refers to the gift database 314 to identify the price points of the gift, which are identified by the gift ID included in the object designation signal. The payment processing unit 308 then updates the user database 312 and subtracts the identified price points from the points of the viewer, which are identified by the viewer ID included in the object designation signal.
[0030] Figure 4 is a table showing an exemplary data structure of the Stream DB 310 in Figure 3. The Stream DB 310 holds information about the live stream currently in progress. The Stream DB 310 stores the Stream ID, Streamer ID, and Viewer ID in relation to each other. The Stream ID is an ID used to identify a live stream on the live streaming platform provided by the Live Streaming System 1. The Streamer ID is a user ID used to identify the streamer providing the live stream. The Viewer ID is a user ID used to identify the viewer of the live stream. In some embodiments of the Live Streaming System 1, when a user starts a live stream, that user becomes a streamer, and when the same user watches a live stream broadcast by another user, that user also becomes a viewer. Therefore, the distinction between streamer and viewer is not fixed, and a user ID that is registered as a streamer ID at one time may be registered as a viewer ID at another time.
[0031] Figure 5 is a table showing an exemplary data structure of the user DB 312 in Figure 3. The user DB 312 holds information about users. The user DB 312 stores the user ID and the points in relation to each other. The user ID identifies the user. The points correspond to the points held by the corresponding user. The points are electronic value that circulates within the live streaming platform. In some embodiments, when a streamer receives a gift from a viewer during a live stream, the streamer's points increase by the value corresponding to the gift. The points are used, for example, to determine the amount of compensation (such as money) that the streamer receives from the administrator of the live streaming platform. In some embodiments, when the streamer receives a gift from a viewer, the streamer may be given an amount of money corresponding to the gift instead of points.
[0032] Figure 6 is a table showing an exemplary data structure of the gift DB314 in Figure 3. The gift DB314 holds information about gifts that viewers can use during the live stream. Gifts are electronic data. Gifts may be purchased with points or money, or they may be provided free of charge. Viewers can send gifts to the streamer. Sending a gift to a streamer is also called using a gift, sending a gift, or throwing a gift. Some gifts can be used immediately after purchase, while others can be used later by the viewer at any time after purchase. When a viewer sends a gift to a streamer, the streamer is awarded a corresponding amount of points. When a gift is used, an effect related to that gift may occur. For example, an effect (such as a visual or auditory effect) corresponding to the gift may be displayed on the live stream screen.
[0033] The gift database 314 stores the gift ID, awarded points, and price points in an associated manner. The gift ID is used to identify the gift. The awarded points are the amount of points awarded to the streamer when a gift is sent to them. The price points are the amount of points paid for the use (purchase) of the gift. Viewers can send a desired gift to a streamer by paying the price points for the desired gift while watching a live stream. Payment of the price points can be made by any appropriate electronic payment method. For example, payment may be made by the viewer paying the price points to the administrator. Alternatively, payment may be made by bank transfer or credit card. The administrator can arbitrarily set the relationship between the awarded points and the price points. For example, it may be set so that awarded points = price points. Alternatively, the price points may be set by multiplying the awarded points by a predetermined coefficient such as 1.2, or by adding a predetermined fee point to the awarded points.
[0034] The machine learning DB 336 may store various machine learning models and / or AI models. For example, one or more large language models (LLMs) or generative AI models may be included in the machine learning DB 336. Various sentiment detection models may be implemented in the machine learning DB 336. The machine learning DB 336 may include image detection models, video detection models, or voice detection models that can be used to analyze the emotions of a broadcaster or the atmosphere of the broadcast room. In some embodiments, the machine learning DB 336 may be implemented outside of the server 10.
[0035] Figure 7 is a table showing an exemplary data structure of Stream DB310. Stream DB310 stores and correlates Stream ID, Broadcaster ID, Timestamp, Video Clip Data, Audio Clip Data, and Interaction Data. The Interaction Data includes interaction records within the broadcast room, such as comment data, gift data, share count, and viewer count.
[0036] The timestamp is configured to specify a period or timing. The video clip data stores a URL where the video data of the live stream clipped at the timestamp is stored. The audio clip data stores a URL where the audio data of the live stream clipped at the timestamp is stored. In some embodiments, the video clip data may include both video and audio data. The URL may lead to an area within or outside the server 10. In some embodiments, the video or audio data may be stored on a streaming server different from the server 10. The clipping or sampling process may be performed by the processing unit 322.
[0037] The comment data stores the content and / or number of comments from viewers in the live stream room. In some embodiments, the viewer ID of the viewer who sent the comment is also stored. The gift data stores the type of gift, the number of gifts, and / or the value of the gifts in the live stream room. In some embodiments, the viewer ID of the viewer who sent the gift is also stored. The share count stores the number of viewers who have shared the live stream room with their friends or on social media. The viewer count indicates how many viewers are in the live stream room.
[0038] Figure 8 is a table illustrating the exemplary data structure of the emotion DB332. The emotion DB332 stores the stream ID, timestamp, video emotion tag, audio emotion tag, comment emotion tag, gift emotion tag, interaction score, emotion description, and correlation keywords in an interrelated manner. Each emotion tag indicates or represents the emotion of the broadcaster or the emotion / atmosphere of the broadcast room.
[0039] The video emotion tag stores the emotion detected or determined from the video clip. The audio emotion tag stores the emotion detected or determined from the audio clip. The emotion may be detected by an emotion detection unit (not shown) by referring to the video clip and / or audio clip. In some embodiments, the processing unit 322 may use the machine learning database 336 to analyze the emotion of the corresponding video clip and / or audio clip. The machine learning database 336 may include an AI model, such as an LLM model, that can take video data and / or audio data as input and output the emotion tag. For example, the AI model may analyze the broadcaster's facial expressions, body gestures, movements, voice pitch, and speech content to determine the emotion in real time.
[0040] The comment sentiment tag stores the sentiment detected or determined from the comment content. The sentiment may be detected by a sentiment detection unit (not shown) by referring to the comment content. In some embodiments, the processing unit 322 may use the machine learning database 336 to analyze the sentiment of the corresponding comment content. The machine learning database 336 may include an AI model, such as an LLM model, that can take the comment data as input and output the sentiment tag in real time.
[0041] The gift emotion tag stores the emotion detected or determined from the gift data. The emotion may be detected by an emotion detection unit (not shown) by referring to the gift data. In some embodiments, the processing unit 322 may use the machine learning database 336 to analyze the emotion of the corresponding gift data. The machine learning database 336 may include an AI model, such as an LLM model, which can take the gift data as input and output the emotion tag in real time. For example, the AI model may analyze the visual and / or audio data of the gift and provide the emotion tag in real time. A beautiful, positive image may correspond to a positive emotion. A negative image that evokes disgust may correspond to a negative emotion. The AI model may analyze the price of the gift and determine the level of happiness based on the price.
[0042] In some embodiments, the gifts do not have to be predetermined. For example, a viewer may take a picture of an object near them and send it as a gift in the live stream room. In some embodiments, predetermined gifts may correspond to predetermined sentiment tags. These predetermined sentiment tags may differ among different streamers based on each streamer's past reactions to those gifts.
[0043] The interaction score or interaction level indicates the degree of user engagement or interaction within the live stream. The interaction score or interaction level may be determined by the interaction data in the stream DB 310. This calculation can be determined in various ways depending on the actual situation. The calculation may be performed by the processing unit 322.
[0044] For example, the interaction score may increase based on the positivity of the comments, the number of comments during the period, the value of the gifts, the number of gifts, the number of shares, the number of viewers, and / or the rate of increase in the number of viewers. For example, the interaction score may decrease based on the negativity of the comments, the decrease in the number of comments, the decrease in the number of viewers, and / or the rate of decrease in the number of viewers. Different interaction levels (high, medium, low, etc.) may be defined based on different ranges of the interaction score.
[0045] In some embodiments, the interaction score / interaction level may be determined by the machine learning DB336. For example, based on a prompt such as, "Please provide interaction levels based on interaction data. These interaction levels indicate user engagement within the delivery room. These interaction levels may increase with positive comments and decrease with negative comments," the AI model may analyze the interaction data and provide an interaction score / interaction level. This prompt may be set by the operator of the delivery platform. The interaction score / interaction level may be determined in real time based on real-time interaction data.
[0046] The interaction score may be used to determine the sentiment explanation. For example, the interaction score may be input as part of a prompt to the AI model to generate the sentiment explanation. In some embodiments, a higher interaction score may result in a more positive effect from the sentiment explanation, while a lower interaction score may result in a more negative effect from the sentiment explanation.
[0047] The sentiment description describes the emotions of the broadcaster and / or the atmosphere / situation of the broadcast room. The correlated keywords indicate the main elements that contribute to the sentiment description. In some embodiments, the processing unit 322 inputs data from the stream DB 310 and / or data from the sentiment DB 332 into the machine learning DB 336 to generate the sentiment description and correlated keywords. In some embodiments, video clip data, audio clip data and / or interaction data are input into the machine learning DB 336 to generate video sentiment tags, audio sentiment tags, comment sentiment tags, gift sentiment tags, interaction levels, sentiment descriptions and / or correlated keywords.
[0048] The correlation keyword may be the main element determined by the machine learning DB336 to be most highly correlated with the sentiment description. In some embodiments, the correlation keyword may be determined based on the timestamp of the extracted word / element (e.g., from video clip data, audio clip data, and / or interaction data, and such extraction may be performed, for example, by the processing unit 322) and the timestamp of the sentiment (or timing of change).
[0049] For example, at timestamp t2, there is a comment that says, "It's beautiful" (Figure 7). The video clip includes a smile, and the audio clip includes a cheerful tone of voice. The machine learning DB336 analyzes the above data and provides the emotion description, "The broadcaster is smiling broadly, narrowing their eyes in delight, and responding warmly to the viewer's positive comment with a cheerful tone of voice," along with the correlated keyword, "beautiful."
[0050] For example, at timestamp t3, the gift GT07 displays a cockroach in the streaming room. The video clip shows the streamer feeling frightened and closing their eyes, and the audio clip includes the streamer's scream. The machine learning database 336 analyzes the above data and provides an emotion description, "The streamer is frightened by the cockroach gift, tightly closing their eyes and screaming with their mouth wide open," and the corresponding keyword, "cockroach." The machine learning database 336 determines that the cockroach gift has the highest correlation with the emotion change (from happy to scared).
[0051] Prompts may be entered into the machine learning DB336 to provide the sentiment description and / or the associated keywords. These prompts may be entered or predetermined by the operator of the distribution platform. For example, the prompts may include content such as: "Please generate sentiment descriptions and associated keywords for the distribution room based on the video clip data, audio clip data, and interaction data I provide. The sentiment description describes the emotions of the broadcaster in the distribution room and / or the atmosphere within the distribution room. The associated keywords are the main elements that resulted from the sentiment description determined based on the above data."
[0052] In some embodiments, different weights may be assigned to different elements / data in the determination of the sentiment description by the machine learning DB336. For example, the prompt may include weights assigned to different elements. For example, the prompt may include content such as, "Please focus on interaction data when determining sentiment descriptions and correlated keywords." As shown in Figures 7 and 8, at timestamp t3, the number of viewers decreases to 26, and the interaction level becomes "low." By weighting the interaction data, the machine learning DB336 may provide a sentiment description such as, "The broadcaster is afraid of gifts and is also frustrated because more users are leaving the broadcast room," or it may provide a correlated keyword such as "decrease in viewers." These weights may be set by the operator of the broadcasting platform or the broadcaster. Details regarding weight setting will be explained with reference to Figure 12.
[0053] Figure 9 is a table showing an exemplary data structure of Effect DB334. Effect DB334 stores stream ID, timestamp, sentiment description, correlation keywords, modified video clip data, and modified audio clip data in an associated manner.
[0054] The modified video clip data stores the URL where the modified video clip is stored. The modified audio clip data stores the URL where the modified audio clip is stored. In some embodiments, the modified video clip data may include both video and audio data. The URL may lead to an area within or outside of the server 10. In some embodiments, the modified video or audio data may be stored on a streaming server different from the server 10.
[0055] In some embodiments, the processing unit 322 inputs the video clip data (in the stream DB 310), the audio clip data, the sentiment description, and / or the correlation keywords into the machine learning DB 336 to generate the modified video and audio data. For example, the machine learning DB 336 may include a generative AI model such as Midjourney or Stable Diffusion (or similar models) that can accept video and / or audio data and modify the data based on prompts. The sentiment description and / or the correlation keywords may act as prompts.
[0056] The acquisition unit 320 is configured to acquire and / or store data in the database. For example, in some embodiments, the interaction data in the stream DB 310 may be acquired by the acquisition unit 320 from a monitoring unit within or outside the server 10.
[0057] Figure 10 shows a flowchart illustrating an exemplary flow based on some embodiments of the present invention.
[0058] In step S1000, the acquisition unit 320 acquires video data, audio data, and interaction data. In step S1002, the processing unit 322 and / or the machine learning database 336 analyze the above data to determine sentiment tags and interaction scores / interaction levels, extract keywords from the video data, audio data, and interaction data, and calculate / determine the correlation between sentiment tags, interaction scores, and keywords. The calculation of correlation is used to determine the main elements / correlated keywords that are attributed to sentiment tags. Other methods may be used.
[0059] In step S1004, the processing unit 322 and / or the machine learning database 336 generate correlated keywords and sentiment descriptions based on video data, audio data, sentiment tags, interaction data, and / or correlation data. For example, the processing unit 322 may first determine correlated keywords based on correlation data and then generate sentiment descriptions based on them. For example, it may determine from the correlation data that a cockroach gift is the main element / correlated keyword for the sentiment tag "fear," and then indicate in the sentiment description that the streamer feels "fear" towards "cockroaches." For example, it may determine from the correlation data that a decrease in the number of viewers is the main element / correlated keyword for the sentiment tag "sad," and then indicate in the sentiment description that the streamer feels "sad" about the "decrease in the number of viewers."
[0060] In step S1006, the processing unit 322 and / or the machine learning database 336 generate modified video / audio data based on the previously modified video / audio data according to the sentiment description and / or the correlation keywords. The sentiment description and the correlation keywords may also act as prompts for modifying the video / audio data.
[0061] In step S1008, the modified video / audio data is provided to users, such as broadcasters and viewers in the distribution room.
[0062] This invention provides automatic and dynamic video / audio changes based on the broadcaster's emotions or the atmosphere / engagement of the broadcast room. These video / audio changes may include facial distortion or background decoration. The emotion tags, interaction scores / levels, emotion descriptions, and / or related keywords are determined in real time and are not limited to predetermined options / settings. The changes (or special effects) to the video / audio data are generated in real time and are not limited to predetermined options / settings. This has the potential to improve engagement in the broadcast room. Viewers may be more inclined to try more interactions (such as commenting or sending gifts) to see what effects appear. Broadcasters may be more inclined to try more expressions and actions to see what effects are generated.
[0063] Figure 11 shows an example of real-time effect generation.
[0064] At timing t1, viewer V1 sends a cockroach gift, and a predetermined gift effect is displayed. At timing t2, the system detects that streamer D1 makes a terrified face and screams. The system generates an emotion description that "the streamer is terrified by the cockroach gift, tightly closing their eyes and screaming with their mouth wide open," and determines the corresponding correlated keyword to be "cockroach." At timing t3, the system generates a modified / emphasized facial expression, an emphasized / added cockroach image, and visual / auditory effects representing the scream on streamer D1 within the streaming room. These are all generated on the fly and are not predetermined effects. The original predetermined gift effect may coexist with the additional effects generated by the AI. In some embodiments, audio data such as a "scream" sound may allow the emotion description to have a description related to that sound. This description then leads to a visual or auditory effect generated by the AI model.
[0065] In some embodiments, the broadcaster D1 may be a virtual broadcaster, or a V-liver / V-streamer / virtual character anchor. The system determines the virtual broadcaster's emotion (or emotion description) based on interaction data (such as comment data, gift data, and viewer count), and generates effects / changes on the virtual broadcaster based on that emotion. The virtual broadcaster may be controlled by an artificial intelligence unit (in server 10 or machine learning DB 336) or by a human from the broadcaster's user terminal.
[0066] Effects or modifications applied to video / audio may be further instructed by multiple prompts specifying the style of such effects. For example, additional prompts may be used to instruct the machine learning DB336 to generate animated, distorted, exaggerated, or enlarged versions of the broadcaster image based on emotion.
[0067] Figure 12 shows an example of generating sentiment descriptions based on specified weights.
[0068] Streamer D1 can specify the elements that most influence their emotions from the displayed user interface (UI). This UI is displayed before or during Streamer D1's live stream. The selected items or rankings (game score, my facial expression, gift value) may be stored in the stream DB 310 along with the interaction data. The processing unit 322 and / or the machine learning DB 336 then generate priority prompts such as, "When generating emotion descriptions, please give more weight to game score, my facial expression (or streamer expression), and gift value."
[0069] Subsequently, the processing unit 322 inputs prompts for generating video data, audio data, interaction data, sentiment descriptions (as described above), and priority prompts into the machine learning DB 336, and generates the sentiment descriptions. In some embodiments, the game score is the score of the game being played by the streamer in the streaming room. The game score may be stored as part of the interaction data. The streamer's expressions may be extracted / detected by the expression unit or the machine learning DB 336. The streamer may set priorities, inform viewers what has the most impact on them (through automatically generated effects / changes / deformations), and encourage viewers to support them. Figure 8 also illustrates examples of different sentiment descriptions obtained as a result of different weighting settings.
[0070] Figure 13 shows an example of generating emotion descriptions and modified videos.
[0071] The server receives interaction data such as comment data and gift data from the user. The server stores correlation data (or historical correlation data) that shows how the broadcaster has responded to various comments and gifts. Based on the interaction data (or current interaction data) and the correlation data, the server (by processing unit 322 and / or machine learning DB 336) generates sentiment descriptions (or predictions of sentiment descriptions). Subsequently, video / audio data is received from the broadcaster, modified based on the sentiment descriptions, and modified video / audio data is generated.
[0072] For example, suppose the gift data contains an image of a cockroach. The correlation data indicates that the streamer is afraid of cockroaches. Therefore, the sentiment description / prediction "The streamer is afraid of the cockroach gift" is generated (or predicted). Subsequently, the video / audio data from the streamer is modified using this sentiment description.
[0073] For example, the sentiment description is generated before (or simultaneously with) receiving the video / audio data from the broadcaster. Sentiment detection based on the current video / audio data from the broadcaster is omitted during sentiment description generation. Therefore, modification efficiency may be improved, and computational resources / time may be saved.
[0074] Figure 14 shows an example of sentiment descriptions and modified video generation based on correlation data.
[0075] The correlation data for streamer D1 indicates that cockroach gifts are a major factor in causing the streamer to feel fear. For example, past data shows that when a viewer sends a cockroach gift, streamer D1's emotion tag becomes fear. Therefore, when a signal for a cockroach gift is received from the viewer, the server generates an emotion description (or emotion description prediction) that "the streamer is afraid of the cockroach gift," and then modifies streamer D1's video data accordingly.
[0076] Correlation data for streamer D2 indicates that cockroach gifts are a key factor in exciting the streamer. For example, past data shows that when a viewer sends a cockroach gift, streamer D2's emotion tag becomes excited. Therefore, when a signal for a cockroach gift is received from the viewer, the server generates an emotion description (or emotion description prediction) that "the streamer is excited about the cockroach gift," and then modifies streamer D2's video data accordingly.
[0077] These examples demonstrate that the same gift or the same interaction data (such as comments or viewer count) can elicit different sentiment descriptions and / or different modifying effects for different streamers, based on their respective correlation data. Streamers D1 and D2 may be in different streaming rooms, or they may be in the same streaming room, such as in PK or competitive mode, in which case both the video data of streamer D1 and the video data of streamer D2 will be displayed in the same streaming room.
[0078] The sentiment description is generated based on interaction data (gifts, comments, viewer count, etc.) and correlation data. Sentiment detection of the streamer video may be omitted to improve the efficiency of sentiment description generation. In some embodiments, the sentiment detection of the streamer video (the actual streamer's sentiment, not a predicted sentiment) may be fed back to adjust / improve / update the correlation data (e.g., by comparing it with a predicted sentiment). This can maintain / improve the accuracy of sentiment prediction generation based on the correlation data.
[0079] Figure 15 shows an example of saving modified effect or correlation data. Timings t1 and t2 may be related or separate situations.
[0080] At timing t1, viewer V1 sends a cockroach gift, and streamer D2 is detected as excited. An exaggerated excited expression is generated and displayed. If streamer D2 likes the modified excited expression, they can choose to save the effect, which will then be correlated with the cockroach gift.
[0081] At timing t2, the system (processing unit 322) detects that the generated effect has led to an improvement in interaction data or interaction score. The processing unit 322 displays "Viewers are happy with the effect! Do you want to save the correlation?", and the streamer D2 can choose to save the effect and correlate it with the cockroach gift.
[0082] Correlation data between effects (or modified video / audio data) and key elements may be stored, for example, in the user DB312. Once effects are stored, the next time the key element is detected in the broadcaster's broadcast room (which may be the current broadcast room or a subsequent broadcast room), the same effect will be provided, eliminating the need for the machine learning DB to generate sentiment descriptions or new modified effects. This may further improve efficiency. A customized effects database corresponding to various interaction elements (gifts, comments, viewer count, etc.) may be established for each broadcaster.
[0083] The hardware configuration of the information processing device will be described with reference to Figure 16. Figure 16 is a block diagram showing an example of the hardware configuration of an information processing device based on some embodiments of the present invention. The information processing device 900 shown in the figure can, for example, realize the server 10 and / or the user terminals 20, 30 in some embodiments.
[0084] The information processing device 900 includes a CPU 901, a ROM (read-only memory) 903, and a RAM (random access memory) 905. The information processing device 900 may also include a host bus 907, a bridge 909, an external bus 911, an interface 913, an input device 915, an output device 917, a storage device 919, a drive 921, a connection port 925, and a communication device 929. Furthermore, the information processing device 900 may include an imaging device such as a camera (not shown). In addition to the CPU 901, or in place of the CPU 901, the information processing device 900 may include a DSP (digital signal processor) or an ASIC (application-specific integrated circuit).
[0085] The CPU 901 functions as an arithmetic processing unit and control unit, and controls the operation of all or part of the information processing unit 900 according to various programs stored in the ROM 903, RAM 905, storage device 919, or removable recording medium 923. For example, the CPU 901 controls the overall operation of each functional unit included in the server 10 and the user terminals 20 and 30 in some embodiments. The ROM 903 stores programs and arithmetic parameters used by the CPU 901. The RAM 905 functions as main memory, storing programs used when the CPU 901 is running and parameters that change as appropriate during execution. The CPU 901, ROM 903, and RAM 905 are interconnected by a host bus 907, which may be an internal bus such as a CPU bus. Furthermore, the host bus 907 is connected to an external bus 911 such as a PCI (Peripheral Component Interconnect / Interface) bus via a bridge 909.
[0086] The input device 915 may be a user-operated device such as a mouse, keyboard, touch panel, button, switch, or lever, or a device that converts physical quantities into electrical signals, such as a sound sensor represented by a microphone, an acceleration sensor, a tilt sensor, an infrared sensor, a depth sensor, a temperature sensor, or a humidity sensor. The input device 915 may also be, for example, a remote control device that uses radio waves such as infrared rays, or an external connection device 927 such as a mobile phone that is compatible with the operation of the information processing device 900. The input device 915 includes an input control circuit that generates an input signal based on the information entered by the user or the detected physical quantity and outputs it to the CPU 901. The user inputs various data to the information processing device 900 and instructs its operation by operating the input device 915.
[0087] The output device 917 is a device that can notify the user of acquired information visually or audibly. The output device 917 may be, for example, a display such as an LCD, PDP, or OLED, an audio output device such as a speaker or headphones, or a printer. The output device 917 outputs the processing results from the information processing device 900 as text, images or other visual content, or audio such as sound.
[0088] The storage device 919 is a data storage device and is configured as an example of a storage unit of the information processing device 900. The storage device 919 is, for example, a magnetic storage device such as a hard disk drive (HDD), a semiconductor storage device, an optical storage device, or a magneto-optical storage device. This storage device 919 stores programs executed by the CPU 901, various data, and various data acquired from external sources.
[0089] The drive 921 is a reader / writer for removable recording media 923 such as magnetic disks, optical disks, magneto-optical disks, and semiconductor memory, and is built into or attached to the information processing device 900. The drive 921 reads information recorded on the installed removable recording media 923 and outputs it to the RAM 905. Furthermore, the drive 921 writes data to the installed removable recording media 923.
[0090] The connection port 925 is a port for directly connecting equipment to the information processing device 900. The connection port 925 may be, for example, a USB (Universal Serial Bus) port, an IEEE 1394 port, or a SCSI (Small Computer System Interface) port. Furthermore, the connection port 925 may be an RS-232C port, an optical audio terminal, or an HDMI (High-Definition Multimedia Interface) port. By connecting the external connection device 927 to the connection port 925, various types of data can be exchanged between the information processing device 900 and the external connection device 927.
[0091] The communication device 929 is, for example, a communication interface formed by a communication device for connecting to the network NW. The communication device 929 may be, for example, a communication card for a wired or wireless local area network (LAN), Bluetooth (trademark), or wireless USB (WUSB). Furthermore, the communication device 929 may be a router for optical communication, a router for ADSL (asymmetric digital subscriber line), or a modem for various types of communication. The communication device 929 transmits and receives signals, etc., to and from the Internet or other communication devices using a predetermined protocol such as TCP / IP. The communication network NW connected to the communication device 929 is a network connected by wire or wireless, and may be, for example, the Internet, a home LAN, infrared communication, radio communication, satellite communication, etc. The communication device 929 realizes the function of a communication device.
[0092] The imaging device (not shown) is a device that captures images of real space and generates captured images using an image sensor such as a CCD (charge-coupled device) or CMOS (complementary metal-oxide-semiconductor) and various components such as lenses for controlling the imaging of a subject on the image sensor. The imaging device may capture still images or moving images.
[0093] The configuration and operation of the live streaming system 1 according to this embodiment have been described above. This embodiment is merely one example, and various changes can be made to the combination of each component and each process, and it will be understood by those skilled in the art that such changes are also within the scope of the present invention.
[0094] The processes and procedures described herein can be implemented by software, hardware, or any combination thereof, in addition to those explicitly described. For example, the processes and procedures described herein can be implemented by implementing the logic corresponding to those processes and procedures on a medium such as an integrated circuit, volatile memory, non-volatile memory, non-temporary computer-readable media, or magnetic disk. Furthermore, the processes and procedures described herein can be implemented as computer programs corresponding to those processes and procedures and can be executed by various types of computers.
[0095] Furthermore, the systems or methods described in the above embodiments may be integrated into a program stored on a non-temporary computer-readable medium such as a solid-state memory device, optical disk memory device, or magnetic disk memory device. Alternatively, the program may be downloaded from a server via the Internet and executed by a processor.
[0096] Although the technical content and features of the present invention have been described above, any person with ordinary skill in the art to which the present invention pertains can make many further modifications and alterations without departing from the teachings and disclosures of the present invention. Therefore, the scope of the present invention is not limited to the embodiments already disclosed, but is the scope of the claims, including other modifications and alterations that do not depart from the present invention. [Explanation of symbols]
[0097] 1. Communication System 10 servers 20 User Terminals 30, 30a, 30b User terminals LV streamer AU1, AU2 viewers VD, VD1, VD2 video NW Network 30 User terminals 100 distribution units 102 Imaging control unit 104 Audio Control Unit 106 Video transmission unit 108 Streamer-side UI control unit 200 viewing units 202 Viewer-side UI control unit 204 Superposition Information Generation Unit 206 Input Information Transmission Unit 302 Distribution Information Provision Unit 304 Relay Unit 306 Gift Processing Unit 308 Payment Processing Unit 310 Stream DB 312 User DB 314 Gift DB 320 units acquired 322 Processing Units 332 Emotion DB 334 Effect DB 336 Machine Learning Database 900 Information Processing Equipment 901 CPU 903 ROM 905 RAM 907 Host Bus 909 Bridge 911 External bus 913 Interface 915 Input device 917 Output device 919 Storage devices 921 Drive 923 Removable recording media 925 connection ports 927 External connection device 929 Communication equipment
Claims
1. A method of live streaming performed by a server, The process of acquiring live streaming data, A step of generating an effect based on the streaming data of the live stream, A method comprising, wherein the effect is not limited to predetermined options.
2. The method according to claim 1, further comprising the step of generating a description of the live stream based on the distribution data, wherein the effect is generated based on the description and the description is not limited to predetermined options.
3. The method according to claim 1, characterized in that the generation of the effect includes inputting the distribution data into a generating AI model and providing the effect.
4. The method according to claim 2, characterized in that the generation of the description includes inputting the distribution data into a first generating AI model to provide the description, and the generation of the effect includes inputting the description as a prompt into a second generating AI model to provide the effect, wherein the first generating AI model is the same as or different from the second generating AI model.
5. The method according to claim 3, characterized in that the distribution data includes the emotions of the participants in the live stream, and the effect includes a deformation of the participant's face based on the emotions.
6. Furthermore, the process includes inputting the distribution data into a first generation AI model to generate an emotional description of the live streamer, wherein the emotional description is not limited to predetermined options. The generation of the effect includes inputting the emotion description and the distribution data into a second generative AI model and generating modified distribution data including the effect for the live distribution, wherein the effect is related to the emotion description. The method according to claim 1, characterized in that the first generative AI model is the same as or different from the second generative AI model.
7. The method according to claim 6, further comprising the step of identifying elements related to the sentiment description, wherein the generation of the effect includes inputting the sentiment description, the elements, and the distribution data into the second generating AI model to generate the modified distribution data including the effect for the live distribution, and the effect is related to the sentiment description and the elements.
8. The method according to claim 5, characterized in that the participant is a virtual character controlled by artificial intelligence or a virtual character controlled by a human broadcaster.
9. The method according to claim 1, characterized in that the effect is related to the audio data of the distributed data.
10. Furthermore, the process includes obtaining priority instruction data with respect to the distribution data from the broadcaster of the live stream, A step of generating the description based on the distribution data and the priority instruction data, The method according to claim 2, characterized by including
11. moreover, The process includes generating correlation data between past interaction data related to the live streamer and the past emotions of the streamer, wherein the past interaction data is from the streamer's viewers. The method according to claim 2, characterized in that the above description is generated based on new interaction data and correlation data during the live stream, rather than on video data of the distribution data from the distributor.
12. Furthermore, the process includes displaying an interface that suggests to the streamer the operation to save the effect, The process of obtaining the selection results from the aforementioned distributor, In response to the selection result, the process involves correlating the effect with the interaction elements in the distribution data and saving it; The method according to claim 1, characterized by including
13. Furthermore, the process includes detecting the interaction element in the broadcaster's subsequent live streams, In subsequent live streams, the process of displaying the effect is included, The method according to claim 12, characterized by including
14. The distribution data includes first video data from a first distributor and second video data from a second distributor, the effect is generated on the first video data, and the method further A step of generating another effect on the second video data based on the aforementioned distribution data, This includes, and the aforementioned other effects are not limited to predetermined options, The aforementioned effect and the aforementioned other effect express different emotions. The method according to claim 1, characterized in that
15. A system for live streaming, comprising one or more processors, wherein one or more processors execute machine-readable instructions. The process of acquiring live streaming data, A step of generating an effect based on the streaming data of the live stream, A system characterized in that it performs the following, and the effect is not limited to predetermined options.
16. A non-temporary computer-readable medium containing a program for live streaming, wherein the program is installed on one or more computers. The process of acquiring live streaming data, A step of generating an effect based on the streaming data of the live stream, A computer-readable medium characterized in that it performs the following, and the effect is not limited to predetermined options.