Distribution device, method, and program

The distribution system adjusts AI broadcaster expressions and gestures to match user preferences, improving satisfaction by personalizing the interactive experience.

JP2026062044APending Publication Date: 2026-04-09KDDI CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Conventional live streaming technologies using AI broadcasters are unable to output expressions and gestures that match the preferences of individual users, leading to reduced user satisfaction.

Method used

A distribution system that controls the manner in which an automatically controlled character speaks to match the preferences of at least one user viewing the content, by utilizing user preference data to adjust facial expressions, gestures, and voice tone in real-time responses.

Benefits of technology

Increases user satisfaction by delivering content that aligns with individual viewer preferences, enhancing the interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062044000001_ABST
    Figure 2026062044000001_ABST
Patent Text Reader

Abstract

This invention provides a distribution device that can enhance user satisfaction in the delivery of interactive video content using automatically controlled characters. [Solution] In a distribution device 20 that accepts comments from users 25 and distributes video content in which an automatically controlled character speaks 26, the manner in which the automatically controlled character speaks is controlled 22, 26 to suit the preferences of at least one user among the users viewing the video content, and then the video content is distributed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a distribution device, method, and program.

Background Art

[0002] For example, there is a so-called "Internet live distribution site" such as YouTube (registered trademark) Live of Non-Patent Document 1 that can distribute video and audio over the Internet, and video content is provided. Viewers (users) can communicate with the distributor in real time by posting comments during the distribution, and can enjoy the content being live-performed by the distributor in a more interactive manner.

[0003] Here, in addition to content in which a live human distributor distributes a video directly taken of himself / herself, or an avatar video of himself / herself generated after adding facial expression analysis and gesture analysis to the video, for example, as disclosed in Non-Patent Document 2, it is also possible to perform a live distribution of a so-called "AI distributor" that uses AI (artificial intelligence) instead of a human as the distributor.

[0004] In Non-Patent Document 2, a live distribution is performed using a remotely operated android (robot) by a human, a conversation corpus is constructed by having a conversation with the android as a user as a viewer, and by using the conversation corpus, a conversation system for live distribution that can autonomously converse with viewer users without requiring remote operation by a human for the android is realized.

[0005] In the conversation system, a response to the user is determined based on a conversation scenario, and the following procedure is taken. (1) Automatically select one "comment close to the statement described in the conversation scenario" from the comments of a large number of users watching the live distribution. (2) Determine the response of the AI distributor based on the conversation scenario and perform voice utterance by voice synthesis.

Prior Art Documents

[0006] [Non-Patent Document 1] Google LLC, "YouTube Live", [online], [searched August 23, 2024], Internet<URL:https: / / www.youtube.com / live> [Non-Patent Document 2] Tomonori Kubota; Kohei Ogawa; Hiroshi Ishiguro. Realization and evaluation of a casual conversation system using an android robot in a live streaming media where an unspecified number of users can speak at any time. Transactions of the Japanese Society for Artificial Intelligence, 2018, 33.1: DSH-G_1-13. [Overview of the project] [Problems that the invention aims to solve]

[0007] However, in the conventional live streaming technology using AI broadcasters described above, the AI ​​broadcaster's facial expressions, gestures, and voice information are output only in accordance with the dialogue scenario. Therefore, the AI ​​broadcaster was unable to output expressions that users watching the broadcast would prefer.

[0008] As schematically illustrated in Figure 1, when a user A comments "I ate a hamburger," and an AI streamer responds with "Oh, that's nice" according to a dialogue scenario, conventional technology, as shown in the upper section, uses a fixed expression for the response, such as "cheerful," corresponding to the content of the response. In this case, if user A prefers a relatively calm "expressionless" expression rather than a relatively lively "cheerful" expression, the response expression that matches that preference cannot be realized. As shown in the lower section, for user A, an AI streamer responding with "Oh, that's nice" in an "expressionless" state increases user A's satisfaction with the AI ​​streamer's content. However, conventional technology could not realize a response that reflected user preferences and increased content satisfaction in this way.

[0009] In view of the above-mentioned prior art, the present invention aims to provide a distribution device, method, and program that can enhance user satisfaction in interactive video content distribution by an AI distributor, i.e., generally an automatically controlled character. [Means for solving the problem]

[0010] To achieve the above objective, the present invention provides a distribution device that distributes video content in which an automatically controlled character speaks while accepting comments from users, characterized in that the manner in which the automatically controlled character speaks is controlled to suit the preferences of at least one user among the users viewing the video content, and then the video content is distributed. The present invention also provides a method and program corresponding to the distribution device. [Effects of the Invention]

[0011] According to the present invention, the manner in which an automatically controlled character speaks is controlled to match the preferences of at least one user among the viewers when delivering video content, thereby increasing satisfaction with video content delivery. [Brief explanation of the drawing]

[0012] [Figure 1] This diagram schematically illustrates the difference between conventional methods and desired methods. [Figure 2] This is a diagram illustrating the configuration of a distribution system according to one embodiment. [Figure 3] This is a functional block diagram of a distribution system according to one embodiment. [Figure 4] This is a functional block diagram of a content generation unit according to one embodiment. [Figure 5] This is a flowchart illustrating the operation of a distribution system according to one embodiment. [Figure 6] This diagram illustrates the types of preferences for each user recorded in the user preference database. [Figure 7] It is a flowchart showing details of step S4 in FIG. 5 according to one embodiment. <00000�5>It is a diagram listing schematic explanatory examples of the present embodiment. [Figure 9] It is a diagram showing a schematic example for explaining an embodiment in which a user preference calculation unit calculates the preference type of each user using the information collected by the viewing information collection device. [Figure 10] It is a diagram showing the hardware configuration in a general computer.

MODE FOR CARRYING OUT THE INVENTION

[0013] FIG. 2 is a configuration diagram of a distribution system according to one embodiment. The distribution system 100 includes terminals 10-k (k = 1, 2,..., N) used by a plurality of N users U-k (k = 1, 2,..., N), a distribution device 20, and a viewing information collection device 30. The terminal 10-k can be configured as a computer device used by each user U-k, such as a smartphone, a tablet, a laptop PC (personal computer), or a desktop PC. The distribution device 20 provides a distribution service, and the viewing information collection device 30 collects the viewing information of each user, and any of these can be configured as a computer device as a general server for performing the corresponding processing.

[0014] In the distribution system 100, each terminal 10-k (k = 1, 2,..., N), the distribution device 20, and the viewing information collection device 30 are communicable with each other via a network NW such as the Internet.

[0015] FIG. 3 is a functional block diagram of a distribution system 100 according to an embodiment. As shown in the figure, the terminal 10 includes an input unit 11 and a content playback unit 12. The distribution device 20 includes a user preference calculation unit 21, a user preference DB (database) 22, a dialogue scenario DB 23, a character DB 24, a comment management unit 25, and a content generation unit 26. Note that the terminal 10 shown in FIG. 3 is any one of the N terminals 10-k (k = 1, 2,..., N) shown in FIG. 2. In other words, the common functional block configuration in all of the N terminals 10-k (k = 1, 2,..., N) in FIG. 2 is the terminal 10 in FIG. 3. FIG. 4 is a functional block diagram of the content generation unit 26, and the content generation unit 26 includes a progress management unit 261, a comment selection unit 262, and a generation unit 263.

[0016] FIG. 5 is a flowchart of the operation of a distribution system according to an embodiment. Hereinafter, while explaining the flowchart, the details of the processing contents of each functional block in FIGS. 3 and 4 will be described.

[0017] In step S1, by receiving the work of an administrator or the like, after constructing the user preference DB 22, the process proceeds to step S2. The constructed user preference DB 22 records information on which of the plurality of preset preference types each of the N users U-k (k = 1, 2,..., N) has a preference pref(k). For example, as illustrated in FIG. 6 in advance, three types A, B, and C are defined as the plurality of preference types, and information on which preference each user U-k has corresponds to can be recorded in the user preference DB 22 as follows, for example. pref(1)=A, pref(2)=C, pref(3)=B, …, pref(N-1)=B, pref(N)=A

[0018] In one embodiment, determining which of three preference categories, A, B, or C, each user falls into can be done by distributing a predetermined questionnaire to each user to determine their preference category, receiving the responses to the questionnaire (corresponding to preference survey information) in the user preference calculation unit 21, calculating each user's preference category from the responses using a rule-based method, and saving the calculation results in the user preference DB 22.

[0019] The prescribed questionnaire could, for example, ask users to select which of the well-known broadcasters A1, A2, A3, B1, B2, B3, C1, C2, and C3 they prefer the most. If any of broadcasters A1, A2, or A3 are selected, it would correspond to preference type A; if any of broadcasters B1, B2, or B3 are selected, it would correspond to preference type B; and if any of broadcasters C1, C2, or C3 are selected, it would correspond to preference type C. The results of the questionnaire could be judged based on rules in this way. In this case, each user would input their questionnaire responses (preference survey information) into an input unit 11, which is an input device consisting of hardware such as a touch panel or keyboard, and send it to the broadcasting device 20, and the user preference calculation unit 21 would receive the questionnaire results from each user.

[0020] Alternatively, in an embodiment described later, the user preference calculation unit 21 may receive information regarding each user's viewing status of internet distributions, which has been collected in advance by the viewing information collection device 30, and analyze this information to determine which of a predetermined preference category (for example, preference categories A, B, and C in Figure 6) each user belongs to, and store this information in the user preference DB 22.

[0021] In step S2, the system accepts tasks from administrators and others to construct the dialogue scenario DB23 and character DB24 before proceeding to step S3. For the character DB24, it is sufficient to record arbitrary content for avatars (however, in this embodiment, there are no real people moving the avatars; they are automatically controlled characters operated by a computer acting as so-called AI broadcasters) and character data (data that renders the appearance of the character using CAD data, etc., and voice data for the character to speak arbitrary dialogue text) in the character DB24.

[0022] The character data can be prepared in any existing data format so that motion videos reflecting corresponding gestures and expressions can be realized through gesture data and facial expression data. For example, to realize gestures, predetermined skeletal points of the entire body can be defined and prepared as data so that character videos that move in conjunction with motion data that moves the skeletal points can be realized. Similarly, predetermined feature points (eyes, eyebrows, mouth, etc.) for the face can be defined and prepared as data so that facial expressions change in conjunction with motion data that moves the feature points. Character data can also be prepared for any part of the body, not the whole body, for example, just the upper body including the face. The character data can also be prepared as basic voice data that can generate audio of the character reading aloud any given dialogue text, by having the character utter the text. This voice data can be prepared in any existing voice data format (for example, data formats such as voices for each vowel and consonant, each word, or each phrase). The basic voice data can also be prepared in a format compatible with any existing voice data format, so that when instructions on the manner of speaking (such as speaking quickly, speaking slowly, speaking excitedly, speaking depressed, etc.) are given in addition to the dialogue text, text-to-speech voice data corresponding to the said manner of speaking can be output.

[0023] Details of the dialogue scenario DB23 will be described later, but in order to realize responses that match user preferences, when responding to a user, you should set up the dialogue content in advance by associating a response according to the preference type (each of the predetermined preference types stored in the user preference DB22) and saving it in the dialogue scenario DB23.

[0024] In step S3, the administrator or other personnel configure the distribution content and notify each terminal 10-k of each of the N users Uk (k=1,2,…,N) that the distribution will be carried out. Then, the distribution device 20 starts the distribution, and the process proceeds to step S4.

[0025] The administrator can configure the content of the broadcast by selecting one of the one or more dialogue scenarios stored in the Dialogue Scenario DB23. At the same time, the administrator can select the AI ​​broadcaster who will deliver the broadcast according to the chosen dialogue scenario from among the multiple characters stored in the Character DB24.

[0026] In this embodiment, the "AI streamer" is not limited to using AI as a means of controlling the character from which video and audio are generated; any automatically controlled character can be used. However, for the sake of simplicity, the following explanation will use the "AI streamer" as an example of an "any automatically controlled character."

[0027] Furthermore, the process of setting the distribution content in step S3, and the process of notifying the user that the configured distribution will be carried out, are not processes specific to the distribution system 100 of this embodiment, but are the same as the processes used when general internet live streaming is performed. Therefore, the functional blocks that explicitly show the process in step S3 are omitted from Figures 3 and 4.

[0028] In step S4, the content generation unit 26 of the distribution device 20 generates and transmits content in real time as video and audio data from the AI ​​distributor to the content playback unit 12 of the user's terminal 10, which has received the distribution notification in step S3 and selected to view the distribution, enabling the user to view it. The content playback unit 12 can be implemented using a display and speakers as hardware, enabling the user to view the content in real time.

[0029] In this embodiment, similar to general internet live streaming, common audio and video content generated in real time by the content generation unit 26 is distributed to each of the terminals 10 of multiple users. On the other hand, in the content distribution of this embodiment, when the AI ​​distributor responds to a comment from a specific user, the content can be generated so as to reflect the preferences of that specific user in the response.

[0030] For example, as illustrated in Figure 6, if a specific user's preference type falls under A among A, B, and C, the AI ​​broadcaster's response to that user's posted comment can be generated as content that is "expressionless with many gestures."

[0031] Figure 7 is a flowchart detailing step S4 according to one embodiment. Figure 8 lists schematic explanatory examples of this embodiment, and the example in Figure 8 will be explained as a premise before describing each step in Figure 7 below.

[0032] Examples EX1 and EX2 in Figure 8 show partial examples of dialogue scenarios stored in the dialogue scenario DB23. Example EX1 is an example of a dialogue scenario structure in which the AI ​​presenter takes the lead in providing topics, and along the way accepts and responds to comments from the user, repeating this process. This AI presenter-led style dialogue scenario can take the form of a repeating structure that sequentially advances the dialogue, for example, as shown below. This will generate buzz with the first monologue by the AI ​​streamer, Mono(1). →Wait for the first comment from the user and select a comment. →A response Res(1) is generated in response to the first selection comment, reflecting the user's preferences. (The same applies to subsequent attempts.) The AI ​​streamer will generate buzz with her second monologue, Mono(2). →Wait for a second comment from the user and select the comment. →A response Res(1) is generated for the second selection comment, reflecting the user's preferences.

[0033] Example EX1 in Figure 8 is an example in which the first monologue Mono(1) and the first response Res(1) by the AI ​​broadcaster are listed as three candidate Res(1) = {Res(1-1), Res(1-2), Res(1-3)}. These monologues and responses are linked to the text information of the content spoken by the AI ​​broadcaster as shown in the figure, and parameter information is pre-prepared to enable the rendering control of the AI ​​broadcaster's motion video when they make the utterance, and the generation control of the text voice generated by the AI ​​broadcaster when they make the utterance.

[0034] The three possible responses for the first response, Res(1), are Res(1)={Res(1-1),Res(1-2),Res(1-3)}, each containing keyword candidates for selecting one comment as the target of the response from multiple comments obtained from multiple users. <t1>Candidates = curry, soba, ..." or " <t2>Candidates are pre-set as "Catfish, Foie Gras, ..." etc.

[0035] For example, in Mono(1), an AI streamer says, "It was hot today, wasn't it? What did you all eat?" to introduce a topic to viewers and prompt them to respond. Within a certain time period after this, the AI ​​streamer selects the user comments from which the above-mentioned comments are selected. <t1> , <t2>By performing a keyword search, a certain user U-1 (username) If the response comment from "Tom" is selected as "I ate curry," the AI ​​broadcaster can respond with the following utterance as response Res(1-1): "Tom, I also ate curry. It's delicious, isn't it?"

[0036] Furthermore, the third response (Res(1-3)) in the first response (Res(1)) is intended to be a waiting state in silence if there are no user comments matching the keyword after the monologue (Mono(1)) by the AI ​​broadcaster, or until a match is found. If this waiting state of responses (Res(1-3)) continues for a certain period of time or longer, it may be assumed that there was no response, or that the user was not listening to the AI ​​broadcaster, and the system may proceed to the second monologue (Mono(2)).

[0037] In this case, the preference type of user U-1 is obtained from the user preference DB22, and if it corresponds to preference type C "surprised expression, speaks quickly" in Figure 6, for example, a standard content corresponding to the utterance is added to the response Res(1-1) according to preference type C, for example as shown below, and the video and audio of the response by the AI ​​broadcaster can be generated. The standard definition "Facial expression: Cheerful" will be changed to "Facial expression: Surprised" to reflect preference type C. The standard definition "Tone: Low" will be changed to "Tone: High" to reflect preference type C. The standard definition "Speed: Normal" will be changed to "Speed: Fast" to reflect preference type C. The standard definition "Gesture type: ja" is changed to "Gesture type: jd" to reflect preference type C.

[0038] In the above example, tone and speed refer to the tone (pitch based on fundamental frequency, etc.) and speed of speech, respectively, and are parameters that control the manner in which the AI ​​streamer generates its voice. Facial expression refers to the facial expression of the AI ​​streamer's character, and gesture type refers to the AI ​​streamer's full-body movements, and are parameters that control the rendering of the AI ​​streamer's character video. The aforementioned character DB24 stores character data so that the content generation unit 26 can generate character video data and speech data according to these parameters.

[0039] In the example of EX1 in Figure 8, the AI ​​presenter takes the lead in the dialogue scenario, providing topics to the user, receiving responses from the user sequentially, and responding according to the user's preferences based on the comments selected. On the other hand, the example of EX2 assumes that the AI ​​presenter does not take the lead in providing topics, but rather appropriately selects one comment (presumably a user question) from multiple comments provided by the user, for example, using a random keyword corresponding to the time, and automatically generates a response comment from the AI ​​presenter based on the selected comment using AI (artificial intelligence) functions.

[0040] In Example EX2, as in Example EX1, the AI ​​broadcaster can be controlled to deliver the response comment in a manner appropriate to the user's preference type of the comment selected by the AI ​​broadcaster as the response. In other words, the AI ​​controls α, β, and γ shown in Example EX2 represent types of controls for the visual appearance and vocalization of the character, similar to those exemplified in Example EX1.

[0041] For example, by using a pre-trained LLM (Large-Scale Language Model) and giving a prompt instruction such as, "Randomly pick a user comment from the comments being streamed, identify the user's preference type, and respond to this user with actions and tone appropriate to their preferences," the content generation unit 26 can automatically generate responses to user comments, as in Example EX2, along with the character's appearance and speaking style at that time. Such LLM models can also be stored in the character DB 24 and dialogue scenario DB 23 as one embodiment of character data corresponding to the character's personality, etc., for generating the character's video and voice. (Note that when using LLM models, the LLM model may combine the roles of both the character DB 24 and the dialogue scenario DB 23, and information on how to give prompt instructions to the LLM model for live streaming may be stored in the dialogue scenario DB 23.)

[0042] Example EX3 in Figure 8 is an example of a streaming screen provided to the user via a web browser or the like when an AI streamer delivers content. The streaming screen has three sections: D1 where the AI ​​streamer's video is displayed, D2 where comments posted by multiple users watching the content are displayed in chronological order, and D3 which is a text input field for users watching the content to post comments. The configuration of this streaming screen itself can use the same configuration as a typical internet live stream that accepts comments in real time.

[0043] Based on the example in Figure 8, we will now explain the flowchart in Figure 7 (details of step S4 in Figure 5). Step S41 is a step that confirms the prerequisites for what is always executed when the flowchart in Figure 7 is executed. The flowchart in Figure 7 is configured to advance the automatic generation of content by the AI ​​distributor by repeatedly executing the case-based judgment in step S42 at each time and the execution of one of the processes in steps S51, S52, or S53 according to that judgment at each time. Step S41 is what needs to be executed at all times in order to be repeatedly executed at each time.

[0044] In other words, in step S41, as a process that is constantly executed to carry out automatic content generation, the progress management unit 261 (Figure 4) manages the progress of the dialogue scenario handled by the content (for example, in the example in Figure 8, it manages where the current position is in the dialogue scenario, such as monologue Mono(1) → waiting for comments → response Res(1) → monologue Mono(2) → waiting for comments → response Res(2) → monologue Mono(3) → waiting for comments → response Res(3) →…). Also in step S41, in accordance with the progress of the content distribution, the comment management unit 25 receives the posted comments entered by the viewer user into the input unit 11 of their terminal 10, and saves and manages all comments from users up to the present time by associating the posting user's ID, username, posting time, and posted comment content. Under the continuous processing by the progress management unit 261 and the comment management unit 25, real-time processing from step S42 onward continues, including comment selection by the comment selection unit 262 and real-time generation of video and audio content by the generation unit 263.

[0045] Step S42 determines which of the following (1), (2), or (3) the current time corresponds to by referring to the progress management status by the progress management unit 261. If it is (1), proceed to step S51; if it is (2), proceed to step S52; and if it is (3), proceed to step S53.

[0046] (1) Timing of generating unnecessary behavior that reflects user preferences The timing of (1) can be determined, for example, as the timing at which the AI ​​broadcaster speaks the monologues Mono(1), Mono(2), etc. in Example EX1. (2) Timing for generating response behavior that reflects user preferences The timing of (2) can be determined, for example, as the timing at which the AI ​​broadcaster speaks the responses Res(1), Res(2), etc. in Example EX1. (3) Timing of making selections from user comments The timing in (3) can be determined as the time from, for example, after the utterance of monologues such as Mono(1), Mono(2), ... in Example EX1, until before the AI ​​broadcaster utters the corresponding responses Res(1), Res(2), ...

[0047] Note that (1) and (2) are selected exclusively from each other, but (1) and (3) may be selected simultaneously. That is, for example, even if the AI ​​broadcaster is in the middle of speaking (not yet finished speaking) the monologue Mono(1), Mono(2), etc. of EX1, user comments on the content of the speech are received in real time, so the selection of comments in (3) may be performed in parallel.

[0048] In step S51, the AI ​​broadcaster's video and audio are generated and distributed as standard behavior without reflecting user preferences, and the process returns to step S42. For example, in monologues Mono(1), Mono(2), etc. in example EX1, the AI ​​broadcaster's video and audio are generated and distributed according to standard behavioral patterns and speech patterns that are independent of user preferences and are defined in the character DB24 and linked to the text content of their speech.

[0049] In step S52, the AI ​​broadcaster generates video and audio that reflects user preferences, broadcasts them, and then returns to step S42. The user whose preferences are reflected in step S52 is the user whose comment was selected in step S53, immediately preceding step S52. In step S53, the next comment to respond to is selected from the user comments, the preference type of the corresponding user is obtained, and then the process returns to step S42.

[0050] Specific examples of steps S52 and S53 are as described above in Example EX1. In step S53, the user comment is selected by the comment selection unit 262 from among the comments stored in the comment management unit 25, selecting those where the difference between the associated timestamp (posting time as a past time) and the current time is within a certain value, and which match keywords corresponding to the utterance. If multiple comments match, one comment may be selected based on predetermined criteria such as the oldest or newest posting time.

[0051] This embodiment has been explained by describing the flows in Figures 5 and 7. If the progress management unit 261 determines that the dialogue scenario by the AI ​​presenter has ended, the flows in Figures 5 and 7 should be terminated. In cases where the AI ​​presenter proactively provides topics, as in Example EX1, an end point can be defined within the dialogue scenario. Also, in cases where the AI ​​presenter does not proactively provide topics, as in Example EX2, and the dialogue scenario progresses randomly based on the AI ​​presenter's responses to comments received sequentially from the user, the end can be determined when the duration of the dialogue scenario reaches a threshold.

[0052] As described above, according to the embodiments of the present invention, the AI ​​broadcaster generates content that is tailored to the preferences of the specific user who posted the comment, thereby increasing the satisfaction of users participating in interactive AI broadcasts. Various supplementary matters will be explained below.

[0053] (1) According to embodiments of the present invention, the value of content generation by AI broadcasters, which is an example of the application of AI such as deep learning that has seen remarkable development in recent years, can be effectively expanded by improving the satisfaction of viewing users, thereby contributing to technological innovation and to Goal 9 of the United Nations Sustainable Development Goals (SDGs), "Build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation."

[0054] (2) Figure 9 is a schematic example illustrating an embodiment in which the user preference calculation unit 21 calculates each user's preference type using the information collected by the viewing information collection device 30, which will be described later.

[0055] As shown in the figure, the viewing information collection device 30 stores comments from one or more broadcasters during live internet streaming, linking them with the user ID, a timestamp corresponding to the progress time of the live stream (as a timestamp of when the comment was posted), and the content of the posted comment, in the same manner as the management method used by the comment management unit 25.

[0056] Let D(j,k) be the set of comments posted by user Uk on streamer Sj(j=1,2,…,M). Text analysis is performed on this set of comments, and the degree of positivity of the comments is evaluated as pos(j,k). The streamer Sj=Sj is selected to maximize the degree of positivity pos(j,k). max Therefore, the preference type corresponding to the behavior of the broadcaster with the highest degree of positivity should be used as the preference type of the user Uk. Here, it is sufficient to analyze the broadcast content of each broadcaster in advance and determine which of the preference types A, B, or C in Figure 6 corresponds to the behavior of that broadcaster.

[0057] In addition to the calculation method shown in Figure 9 above, the viewing information collection device 30 may store the user's viewing history of distributed content, and the user preference calculation unit 21 may obtain the user's preference type from the characteristics of the people appearing in the distributed content that the user frequently views, or, if there are multiple people in the distributed content, from the characteristics of the person the user pays the most attention to.

[0058] (3) In the above explanation, we described the tone (pitch), speed of speech, gestures, and facial expressions as items that reflect user preferences in content generation. However, visual and / or auditory performance elements for expressing the emotions of the AI ​​broadcaster may be used as items that can achieve a similar effect.

[0059] Examples of visual effects used to express emotions in AI streamers include, for instance, overlaying an animated depiction of sweat flowing onto the AI ​​streamer's face to express surprise, or overlaying a drawing resembling converging lines to emphasize a joyful emotion onto the AI ​​streamer's face.

[0060] Examples of visual effects used to express emotions in AI streamers include, for example, adding a predetermined sound effect (such as the sound of a vase breaking) superimposed on the AI ​​streamer's speech to express surprise, or adding an emphasis sound echo effect to the AI ​​streamer's speech to express joy.

[0061] (4) In the embodiments described above, we have explained the case in which a response to a comment from a specific user is generated in accordance with the preferences of that specific user. However, depending on the situation, for example, if multiple users post similar comments to each other, it may be more convenient for the progress of the broadcast if the AI ​​broadcaster responds to all of these users at once. For example, in a popular broadcast, there are many comments posted, and it would take a long time for the AI ​​broadcaster to respond to each comment individually. Also, even if there was enough time to respond to each comment individually, responding to all comments would slow down the pace of the broadcast and potentially reduce the value of the broadcast content. On the other hand, it is thought that many users want the broadcaster to respond to their comments, so summarizing comments from multiple users is effective.

[0062] For example, if similar comments like "What did you eat?" (User1) and "What kind of food do you eat?" (User3) are received at once, the AI ​​can respond to them all together. (In this case, if the preferences of the multiple users who posted similar comments differ, the method of response can be determined by a rule-based method such as majority vote.)

[0063] Regarding the handling of similar comments, the comment selection unit 262 monitors the comments stored in the comment management unit 25, applies content classification using any existing method such as clustering or AI processing to comments obtained within a certain period up to the present time, and retrieves multiple comments that are determined to be identical in content. In this case, the comment selection unit 262 may select comments that are identical in content and also meet predetermined criteria, such as having the most comments. The text of the audio content generated by the content generation unit 26 when generating video and audio responses to multiple comments at once may also be generated according to predetermined rules, such as reading aloud all or part of the names of the multiple users who posted the same comment.

[0064] (5) For example, the system may analyze multiple users who have commented within a predetermined time after the start of distribution (e.g., up to 5 minutes after the start of distribution) and decide what background (room background, outdoor background, etc.) to use for distribution. For example, if there are three types of preference categories, A, B, and C, as shown in Figure 6, and the comment management unit 25 counts the user comments (counts are 2 or more if one user makes two or more comments) made by the AI ​​distributor during the first 5 minutes after the start of distribution, and the corresponding preference category counts are count(A), count(B), and count(C), then the system may use the background Back(A), Back(B), and Back(C) corresponding to the preference category with the largest number among these as the background for content distribution by the AI ​​distributor. In this case, for example, if count(A), count(B), and count(C) are totaled after 5 minutes and it is found that count(A) is the largest, then background Back(A) should be set as the background at that point. On the other hand, for the first five minutes of distribution until the tabulation is complete, any designated background other than backgrounds Back(A), Back(B), or Back(C) (for example, a plain background) should be used.

[0065] (6) Depending on the type of topic covered in the broadcasts the user views (politics, sports, cooking, cars, etc.), the AI ​​broadcaster may save the type of broadcaster the user prefers (reflected in facial expressions, gestures, voice, etc.) and then respond according to the type of topic covered in the broadcasts the user listens to (category of spoken content).

[0066] In other words, in the above embodiment, the preference type pref(k) is defined for each user Uk (k=1,2,...,N), but it may also be defined for each topic type j (j=1,2,...,M) such as politics, sports, cooking, cars, etc., so that the preference type is set as pref(k,j) and utilized according to the type of delivered content. For example, in a delivery, when the same AI deliverer responds to the same user, the AI ​​may respond seriously when the topic is politics and energetically when the topic is sports, thus enabling delivery that is more in line with the user's preferences by using response patterns according to the type of topic.

[0067] (7) In this embodiment, it is assumed that the AI ​​broadcaster is able to respond in a manner that suits the preferences of the viewing user (by selecting from a list of candidate response modes). Therefore, it is not advisable to apply this embodiment to cases where the AI ​​broadcaster's response, regardless of the mode, would not suit the user's preferences (for example, when a user does not like the AI ​​broadcaster in the first place). To avoid such situations, when notifying users of the start of broadcasting and encouraging them to watch, the notification of the start of broadcasting may be sent only to users whose preferences match at least one of the AI ​​broadcaster's response modes.

[0068] (8) In the embodiments described above, user preferences were reflected in responses to comments. However, even when not responding to comments, the AI ​​broadcaster may reflect the preferences of at least one user among those watching the broadcast and perform speeches or other actions accordingly. For example, during a certain period immediately after the start of the broadcast (which may include zero comments), the AI ​​broadcaster may provide topics and not respond to user comments. In such cases, the AI ​​broadcaster may proceed with providing topics in a manner that matches the preference type of the largest number of users among those watching the broadcast during that period. That is, the AI ​​broadcaster may express characteristics (facial expressions, gestures, etc.) that are preferred by the majority of the multiple users watching the broadcast while making speeches or other actions. (In this case, the minority will see the AI ​​broadcaster preferred by the majority.)

[0069] For example, similar to the example above, let's assume there are three types of preference categories, A, B, and C, as shown in Figure 6. At a predetermined time after the start of distribution, let's count the number of users viewing content corresponding to preference categories A, B, and C (so-called online users) to obtain num(A), num(B), and num(C). For example, if num(A) is the largest number (the majority), we may generate video and audio content in a manner appropriate to preference category A (this may or may not be a response to comments).

[0070] This embodiment can be realized as a variation of step S51 in the flow shown in Figure 7 above. That is, step S51 was "a scene in which the broadcaster speaks a monologue or the like, which is not a response to a comment, so the broadcaster's video and audio are generated and distributed as standard behavior without reflecting user preferences," but step S51 can be modified to "a scene in which the broadcaster speaks a monologue or the like, which is not a response to a comment, but the broadcaster's video and audio are generated and distributed as behavior that reflects the largest number of user preferences among the current user preferences."

[0071] (9) In the above embodiments, the speech output in the content generated by the AI ​​distributor in the content generation unit 26 is played back as audio data by the speaker and output audibly. Alternatively, or in addition to this, the speech output may be output visually by displaying the speech content as text such as subtitles on the display, or it may be output visually as video through sign language gestures. Furthermore, within the same AI content by the same AI distributor, the speech output may be switched between being output as audio and being output visually as subtitles or sign language gestures depending on the scene.

[0072] (10) Figure 10 shows an example of a hardware configuration in a typical computer device 70. The terminal 10, distribution device 20, and viewing information collection device 30 in the distribution system 100 can be realized as one or more computer devices 70 having such a configuration. When one of the terminal 10, distribution device 20, and viewing information collection device 30 is realized with two or more computer devices 70, information necessary for processing may be sent and received via a network. The computer device 70 includes a CPU (Central Processing Unit) 71 that executes predetermined instructions, a GPU (Graphics Processing Unit) 72 as a dedicated processor that executes some or all of the execution instructions of the CPU 71 on behalf of or in cooperation with the CPU 71, RAM 73 as main memory that provides a work area to the CPU 71 (and GPU 72), ROM 74 as auxiliary memory, a communication interface 75, a display 76 that outputs a display, an input interface 77 that accepts user input via a mouse, keyboard, touch panel, etc., a speaker 78 that outputs sound, and a bus BS for sending and receiving data between these.

[0073] The terminal 10, the distribution device 20, and the viewing information collection device 30 (and their respective functional blocks) can be implemented by a CPU 71 and / or GPU 72 that read and execute predetermined programs corresponding to the functions of each part from ROM 74. Both the CPU 71 and GPU 72 are types of arithmetic units (processors). When display-related processing is performed, the display 76 also operates in conjunction; when communication-related processing for data transmission and reception is performed, the communication interface 75 also operates in conjunction; and when audio output-related processing is performed, the speaker 78 also operates in conjunction. [Explanation of symbols]

[0074] 100...Distribution system, 10...Terminal, 20...Distribution device, 30...Viewing information collection device 11...Input unit, 12...Content playback unit, 21...User preference calculation unit, 22...User preference DB, 23...Dialogue scenario DB, 24...Character DB, 25...Comment management unit, 26...Content generation unit 261...Progress Management Unit, 262...Comment Selection Unit, 263...Generation Unit < / t1>

Claims

1. In a distribution device that delivers video content in which an automatically controlled character speaks while accepting comments from users, A distribution device characterized by controlling the manner in which the automatically controlled character speaks to suit the preferences of at least one user among the users viewing the video content, and then distributing the video content.

2. Controlling it in a manner that suits the aforementioned preferences is The distribution device according to claim 1, further comprising controlling the manner in which the automatically controlled character speaks in response to a comment received from a specific user, in a manner that is in line with the preferences of the specific user.

3. Controlling it in a manner that suits the aforementioned preferences is The tone of the voice of the aforementioned automatically controlled character, The speed of speech of the aforementioned automatically controlled character, and The distribution device according to claim 1 or 2, which includes controlling at least one of the auditory performance expressions for expressing the emotions of the automatically controlled character.

4. Controlling it in a manner that suits the aforementioned preferences is The facial expression of the aforementioned automatically controlled character, Visual performance expressions for expressing the emotions of the aforementioned automatically controlled character, and The distribution device according to claim 1 or 2, which includes controlling at least one of the gestures of the automatically controlled character.

5. Controlling it in a manner that suits the aforementioned preferences is The distribution device according to claim 1, further comprising controlling the manner in which the automatically controlled character speaks in response to similar comments received from multiple users, so as to be in accordance with the preferences of at least one of the multiple users.

6. Controlling it in a manner that suits the aforementioned preferences is The distribution device according to claim 1, which includes counting the number of users currently viewing the video content for each user preference type, and then controlling the content in a manner appropriate to the preference type that yielded the highest count.

7. The distribution device according to claim 1, characterized in that the manner in which the automatically controlled character speaks is controlled to conform to the preferences of at least one user among the users viewing the video content and the category type of the spoken content, and then the video content is distributed.

8. The distribution device according to claim 1, characterized in that it counts the number of users who are watching the video content for each user preference type, and sets the background corresponding to the preference type that yielded the highest count as the background of the video content.

9. The distribution device according to claim 1, characterized in that it outputs the speech of the automatically controlled character audibly and / or visually.

10. In a computer-based distribution method that delivers video content in which an automatically controlled character speaks while accepting comments from users, A distribution method characterized by controlling the manner in which the automatically controlled character speaks to suit the preferences of at least one user among the users viewing the video content, and then distributing the video content.

11. In a program that uses a computer as a distribution device to deliver video content in which an automatically controlled character speaks while accepting comments from users, A program characterized by controlling the manner in which the automatically controlled character speaks to suit the preferences of at least one user among those viewing the video content, and then distributing the video content.