Device and method
The apparatus and method use a generative AI model to create exclusive content for online events by estimating and incorporating participant-specific highlights, addressing the lack of exclusivity in conventional systems.
Patent Information
- Application Number
- PCT/JP2024/016241
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-25
- Publication Date
- 2025-10-30
AI Technical Summary
Conventional systems for participating in online events via the Internet fail to provide a sense of exclusivity to participants, despite being able to express emotions as digital waves.
An apparatus and method utilizing a generative AI model to generate content based on estimated first and second highlight portions of an event, incorporating participant-specific information, to distribute exclusive content.
Generates content that provides participants with a sense of exclusivity by leveraging participant-specific highlights, enhancing the online event experience.
Smart Images

Figure JP2024016241_30102025_PF_FP_ABST
Abstract
Description
Apparatus and method
[0001] One aspect of the present disclosure relates to an apparatus and a method.
[0002] It is becoming common to participate in events such as concerts, theater performances, and sporting events via the Internet. However, compared to events where participants actually go to the venue (hereinafter also referred to as a "real event"), it is difficult to obtain a sense of excitement and exclusivity (or specialness) from events participated in via the Internet (hereinafter also referred to as an "online event"). A system for solving this problem is disclosed in Patent Literature 1. The system described in Patent Literature 1 can express the emotions of event participants as digital waves and provide these digital waves to the participants.
[0003] Japanese Patent Application Laid-Open No. 2021-149135
[0004] However, while the conventional systems described above can provide participants with the aforementioned sense of excitement, they are unable to provide a sense of exclusivity. Therefore, it is conceivable to distribute content that is only available to participants of an online event. However, generating such content requires time and effort. Therefore, it is conceivable to use a generative artificial intelligence (AI) model that can generate a generation result (content) in response to input of various generation request information, in accordance with the instructions, context, questions, and output format indicated in the generation request information, and return the generation result as response information.
[0005] An object of one aspect of the present disclosure is to provide an apparatus and method that can use a generative AI model to generate content that can give participants of an online event a sense of exclusivity.
[0006] An apparatus according to one aspect of the present disclosure includes a reception unit that receives transmission information transmitted from each participant of an event that can be attended via the Internet; an estimation unit that estimates a first highlight portion, which is an exciting portion of the event, and a second highlight portion, which is an exciting portion of the event for each participant, based on the transmission information received by the reception unit; and an output unit that generates generation request information to instruct a generation AI model to generate content to be distributed to each participant, based on the first highlight portion and the second highlight portion estimated by the estimation unit.
[0007] According to one aspect of the present disclosure, content that can give participants of an online event a sense of exclusivity can be generated using a generative AI model.
[0008] FIG. 1 is a diagram illustrating an overview of a distribution device according to an embodiment. FIG. 2 is a diagram illustrating an example of the functional configuration of the distribution device according to an embodiment. FIG. 3 is a diagram illustrating an example of an event distributed from the distribution device. FIG. 4 is a diagram illustrating an example of data included in distribution information. FIG. 5 is a flowchart illustrating an example of a process for generating and distributing content to be distributed to each participant. FIG. 6 is a diagram illustrating a method for extracting a first highlight portion. FIG. 7 is a diagram illustrating a method for extracting a second highlight portion. FIG. 8 is a diagram illustrating an example of content to be distributed to each participant. FIG. 9 is a diagram illustrating an example of the hardware configuration of a computer used in the distribution device according to an embodiment.
[0009] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the description of the drawings, the same elements are designated by the same reference numerals, and duplicate explanations will be omitted. Furthermore, the embodiments of the present disclosure in the following description are specific examples of the present disclosure, and the present disclosure is not limited to these embodiments unless otherwise specified to limit the present disclosure.
[0010] 1 is a diagram illustrating an example of a system configuration including a distribution device (device) 1 according to an embodiment. As illustrated in FIG. 1, the distribution device 1 is configured to be able to communicate with a plurality of user terminals used by a plurality of users via a network. The network may be any communication network, such as a mobile communication network or the Internet.
[0011] The distribution device 1 distributes an event that can be participated in via the Internet, and provides an environment in which users (participants) can participate in the event to terminals 5 connected via the Internet.
[0012] The terminal 5 is a computer device such as a mobile communication terminal or a laptop computer that performs mobile communication. In this embodiment, the terminal 5 is assumed to be a smartphone, but is not limited thereto. The terminal 5 is a portable terminal owned by the user of the terminal 5. For example, as shown in FIG. 3 , the terminal 5 can receive at least one of video, still images, and sound distributed by the event organizer via the distribution device 1, and event participants can send at least one of video, still images, and sound to the event organizer. Note that the video and still images may include text information such as captions, chat, and moving comments on the screen. Information that can be received by the terminal 5 and information that can be sent from the terminal 5 when participating in an event via the Internet will be described in detail below.
[0013] 2 is a diagram illustrating an example of a functional configuration of the distribution device 1 according to the embodiment. As illustrated in FIG. 2, the distribution device 1 includes an event distribution unit 15, a reception unit 11, an estimation unit 12, an output unit 13, a generation unit 21, a storage unit 23, and a content distribution unit 17.
[0014] Each functional block of the distribution device 1 is assumed to function within the distribution device 1, but is not limited to this. For example, some of the functional blocks of the distribution device 1 may function in a computer device different from the distribution device 1 and connected to the distribution device 1 through a network, while appropriately sending and receiving information with the distribution device 1. Furthermore, the functional blocks in the distribution device 1 do not have to be configured as shown in FIG. 2; some functional blocks of the distribution device 1 may be omitted, multiple functional blocks may be integrated into one functional block, or one functional block may be separated into multiple functional blocks.
[0015] Hereinafter, each function of the distribution device 1 shown in FIG. 2 will be described.
[0016] The event distribution unit 15 distributes events that can be participated in via the Internet. Examples of events here include concerts, theater, and sports. Events that can be participated in via the Internet include events held in real space, as shown in FIG. 1, in which still images or videos captured by a camera 3 or the like are distributed via the Internet, and events held in the metaverse space, which is a virtual space constructed on the Internet. The following mainly describes an example in which an event held in real space is distributed via the Internet.
[0017] For example, as shown in FIG. 3 , the event distribution unit 15 distributes data (content) to event participants, which combines on a single screen a background image 81 of the event, event performers 82, chat messages 83 (83A, 83B) sent by the participants, and the number of participants 84 in the event connected via the Internet. Examples of the data distributed from the event distribution unit 15 include, for example, video data and still image data. As described above, the data distributed by the event distribution unit 15 may not only be video or still images as shown in FIG. 3 , but may also be data (content) formed from sound instead of or in addition to these.
[0018] The distribution information distributed by the event distribution unit 15 to participants in the event includes at least one of information regarding the number of participants in the event connected via the Internet, information regarding the performers of the event, information regarding the background of the event, and information regarding chats sent by participants, as shown in Figure 4, for example.
[0019] 2, the event distribution unit 15 may not only distribute data such as moving images or still images as described above, but also provide participants with a function that enables them to enter chat messages. The event distribution unit 15 may also provide a function that enables participants to participate in the event as an avatar.
[0020] The reception unit 11 receives transmission information transmitted from each participant for an event that can be participated in via the Internet. Examples of the transmission information received by the reception unit 11 include at least one of information about the participant, monitoring information about the participant when participating in the event, and information about chats transmitted by the participant.
[0021] Examples of information about participants include the participant's age, gender, appearance, hobbies and interests, information about events they have participated in in the past, etc. Information about participants can be obtained, for example, from a profile that participants are required to enter in advance to participate in an event, or from the results of responses to a questionnaire after participating in an event.
[0022] The monitoring information of participants when participating in an event includes information about the reactions of participants when participating in the event. Examples of participant reactions include the participants' movements, behavior, vocalizations, etc. The participant reactions also include the reactions of the avatars when the participants participate as avatars. The participant reactions can be acquired, for example, from a camera, microphone, etc. mounted on the terminal 5. The participant's movements and behavior also include changing the display mode of the video being distributed via the Internet. If the event (video or still image) distributed from the distribution device 1 to the terminal 5 can be enlarged or reduced, the magnification ratio is included; if the angle at which the event is viewed can be adjusted, the angle (eye position); if the focus position when viewing the event can be adjusted, the focus position is included.
[0023] The information related to the chat sent by the participants includes text information, icons, emoticons, and the like input into the chat function provided by the distribution device 1.
[0024] The estimation unit 12 estimates a first highlight portion, which is an exciting portion of the event, and a second highlight portion, which is an exciting portion of the event for each participant, based on the transmission information and distribution information received by the reception unit 11. The first highlight portion is an exciting portion of the event as a whole, and for example, the more participants who are excited about a portion, the more likely it is to become the first highlight portion. Furthermore, the second highlight portion is an exciting portion for each participant, and may be different for each participant.
[0025] The estimation unit 12 can extract the first highlight portion and the second highlight portion based on, for example, the number of chats sent by participants when an event distributed via the Internet is divided into unit times (e.g., one minute), the presence or absence of high-energy words included in the chats, or the frequency of occurrence of such words. The estimation unit 12 can also extract the first highlight portion based on, for example, the number of participants connecting to the event via the Internet or the number of participants who leave the Internet-connected state when the event is divided into unit times. The estimation unit 12 can also extract the first highlight portion based on, for example, the amount of reactions of performers that can be extracted from videos, still images, or audio distributed via the Internet when the event is divided into unit times, the presence or absence of reactions of high-energy performers, or the frequency of occurrence of reactions of high-energy performers. The estimation unit 12 can also extract the second highlight portion based on, for example, the amount of reactions of participants that can be extracted when monitoring participants at the time of their participation in the event when the event is divided into unit times, the presence or absence of reactions of high-energy participants, or the frequency of occurrence of reactions of high-energy participants. The unit time here may be a predetermined time (for example, one minute), or may be a time period (for example, 30 seconds) before and after the scene in which the number of times exceeds a predetermined threshold.
[0026] The output unit 13 generates a prompt (generation request information) for instructing the generation AI model to generate content to be distributed to each participant, based on the first highlight portion and the second highlight portion estimated by the estimation unit 12. The output unit 13 outputs the generated prompt to the generation unit 21. A prompt is information indicating an instruction or question entered by a user in an interactive system such as a dialogue with an AI model or a command line interface (CLI).
[0027] The prompt output by the output unit 13 expresses, in text, for example, the command to be executed by the generating AI model, the task to be executed by the generating AI model, the performance content to be considered by the generating AI model, the output format of the response information from the generating AI model, etc. The prompt may also include input information (e.g., video or still image data of an event distributed via the Internet) that is the target of the command or task to be executed by the generating AI model. Examples of such input information include, for example, data files with file names including a predetermined extension, such as text data, image data, application-related data, audio data, video data, and still image data.
[0028] The output unit 13 generates, for example, a prompt to generate content including at least one of a video, a still image, and audio of the first highlight portion and the second highlight portion estimated by the estimation unit 12. The output unit 13 may also generate generation request information to generate avatars of the participants based on information about the participants included in the transmitted information, and generate content in which the avatars are combined with video or still images. Examples of prompts output by the output unit 13 will be described in detail later.
[0029] The generation unit 21 is a part that enables the provision of content using a generative AI model. The generative AI model is a model that can generate content in response to a prompt containing input information, based on any one or a combination of the instructions, context, question, and output format indicated by the prompt, and return the content as response information. The prompt can also include input information (in this embodiment, video or still images of an event distributed via the Internet). In this case, the generative AI model generates response information targeted at the input information. The generative AI model may be, for example, an interactive AI model that includes a user interface (UI) for interacting with a user and enables text or voice chat with the user. Examples of such generative AI models include Stable Diffusion, Midjourney, and DALL-E3.
[0030] In the present embodiment, the generation unit 21, i.e., the generation AI model, has been described as being configured as part of the distribution device 1, but it may also be configured as another computer device connected via a network. Note that while only one distribution device 1 is illustrated in FIG. 1, the distribution device 1 may be configured by multiple computer devices, or may include multiple computer devices if the generation unit 21 is configured as a separate entity.
[0031] The content distribution unit 17 distributes the content generated for each participant by the generation unit 21 to each participant via a network. Note that if the event organizer mails media storing the content generated for each participant by the generation unit 21, the content distribution unit 17 may not be provided. Furthermore, the content distribution unit 17 may tokenize the content generated for each participant by the generation unit 21 as an NFT (Non-Fungible Token). This ensures the originality and rarity of the content.
[0032] The storage unit 23 stores (preserves) the transmission information received by the reception unit 11. The storage unit 23 may store the transmission information received by the reception unit 11 on a regular or irregular basis while maintaining the chronological order. The storage unit 23 may also store videos or still images of the event distributed by the event distribution unit 15, or may store content generated by the generation unit 21 to be distributed to each participant.
[0033] Next, an example of the process executed by the distribution device 1 described above, i.e., an example of a method for generating a prompt for generating content to be distributed to each participant, will be described mainly with reference to FIGS. 5 to 8. FIG. 5 is a flowchart showing an example of a process for generating and distributing content to be distributed to each participant. FIG. 6 is a diagram showing an example of a method for extracting a first highlighted portion. FIG. 7 is a diagram showing an example of a method for extracting a second highlighted portion.
[0034] Here, a process will be described in which the distribution device 1 distributes an event that can be participated in via the Internet and generates content to be distributed to each event participant after the event ends. As shown in Fig. 5, the event distribution unit 15 distributes the event via the Internet (step S1). The event distribution unit 15 distributes, for example, an event such as that shown in Fig. 3.
[0035] Next, the reception unit 11 receives transmission information transmitted from each of the event participants (step S2: reception step). The information that the reception unit 11 can receive includes information that can be acquired before the event distribution (e.g., information related to the participant's profile), information that can be acquired during the event distribution (e.g., participant reactions that can be acquired from a camera or microphone, information related to posted chats), and information that can be acquired after the event distribution (e.g., information related to a questionnaire about the event). This information can be acquired from the event distribution unit 15 or the storage unit 23.
[0036] Next, the estimation unit 12 estimates a first highlight portion, which is an exciting part of the event, based on the transmitted information received by the reception unit 11 (step S3: estimation step). As shown in Fig. 6, the first highlight portion can be estimated by extracting information based on at least one of (1) information about chats throughout the event, (2) information about the reactions of performers throughout the event, and (3) information about the number of simultaneous connections throughout the event.
[0037] (1) From the information about chats for the entire event, for example, scenes with a large amount of chat per unit time (e.g., one minute) can be extracted as first highlight portions, or scenes with chats containing highly passionate words can be extracted as first highlight portions. Highly passionate words here include, for example, positive words, words that show goodwill or enjoyment, etc. Specifically, they include words such as cool, cute, wonderful, exciting, or like. Highly passionate words are stored, for example, in the storage unit 23.
[0038] Furthermore, (2) from information about the reactions of the performers throughout the event, it is possible to extract, for example, scenes with a large number of reactions per unit time (e.g., one minute) as first highlight portions, or scenes with highly passionate reactions as first highlight portions. Highly passionate reactions here include, for example, positive reactions, reactions showing goodwill or enjoyment, etc. Specifically, they include reactions such as joy, surprise, excitement, relief, and envy. Highly passionate reactions are stored, for example, in the storage unit 23 as patterns of movement or facial expression. The relevant reactions can be extracted using known methods such as pattern recognition.
[0039] Furthermore, (3) from information regarding the number of simultaneous connections for the entire event, for example, when divided into unit time periods (e.g., one minute), it is possible to extract as a first highlight portion a scene in which the number of participants connecting to the event via the Internet is greater than a predetermined threshold, or to extract as a first highlight portion a scene in which participants continue to participate without leaving the Internet connection for more than the predetermined threshold. Note that the state in which participants continue to participate without leaving the Internet connection can be determined to be a state in which participants are engrossed in the event.
[0040] Next, the estimation unit 12 estimates a second highlight portion, which is an exciting portion of the event, based on the transmitted information received by the reception unit 11 (step S4: estimation step). The second highlight portion can be estimated by extracting information based on at least one of (4) information about the chats of each participant and (5) information about the reactions of the performers of each participant, as shown in FIG.
[0041] (4) From the information about the chats of each participant, for example, scenes with a large amount of chat for each participant in a unit time (for example, one minute) can be extracted as the second highlight portion, or scenes with chats containing highly passionate words for each participant can be extracted as the second highlight portion. The highly passionate words referred to here can be, for example, the same as the words exemplified in (1) above.
[0042] Furthermore, (5) from the information regarding the reactions of each participant, for example, scenes with a large number of reactions for each participant in a unit time (for example, one minute) can be extracted as second highlight portions, and scenes with a highly passionate reaction for each participant can be extracted as first highlight portions. The highly passionate reactions referred to here can be, for example, the same as the words exemplified in (2) above.
[0043] Furthermore, the estimation unit 12 can (6) extract second highlights based on information about each participant. The estimation unit 12 may extract participant values from participant profiles or questionnaire data, for example, or from the participant's behavior when participating in an event. The estimation unit 12 can estimate the preferences of each participant (performers, music, works, etc.) from such information. By taking into account the preferences of each participant, the estimation unit 12 can extract second highlights that are more in line with the participant's true nature.
[0044] Next, the output unit 13 generates a prompt for instructing the generation of content to be distributed to each participant based on the first highlight portion and the second highlight portion estimated by the estimation unit 12 (estimation step) (step S5: output step). The output unit 13 generates a prompt to generate content including, for example, videos or still images of scenes corresponding to the first highlight portion and the second highlight portion. The output unit 13 may generate a prompt including an instruction to increase the length (number) of videos or still images to be captured as content as the excitement level of the first highlight portion increases.
[0045] The output unit 13 also generates prompts to generate content such as video, still images, or sounds that reflect information about the performers (such as their names, appearances, emotions, states (standing, sitting, etc.), and statements), information about the participants (such as their names, appearances, emotions, states (standing, sitting, etc.), and information about the event space (such as the appearance of the walls, ceiling, and floor where the performers and participants are located, and the appearance of the stage where the performers are located, etc.).
[0046] In addition, the output unit 13 generates prompts to generate content that incorporates production methods to reflect the preferences of each participant (showing the performers or participants large, showing the participants small and the performers large, showing a panoramic view from the avatar's point of view, showing a panoramic view of the entire stage, etc.).
[0047] Note that the scene corresponding to the first highlight portion and the scene corresponding to the second highlight portion estimated by the estimation unit 12 may be the same, or the scene corresponding to the first highlight portion and the scene corresponding to the second highlight portion may be different. Furthermore, when extraction is performed at unit time intervals as described above, a moving image or still image of a scene including a first highlight portion refers to a moving image or still image of a unit time that includes a portion determined to be the first highlight portion.
[0048] Next, the output unit 13 outputs the prompt generated in step S5 to the generation unit 21 (generative AI model). The generation unit 21, which has received the prompt, generates content to be distributed to each participant based on the prompt (step S6). The content generated by the generation unit 21 is content (data) that combines, on a single screen, information 91 related to a background image of the event, information 92 related to performers of the event, information 93 and 94 related to chats sent by participants, and information 95 related to avatars of the participants, as shown in FIG. 8, for example.
[0049] Information 91 relating to the background image of the event includes, for example, images of the stage and background where the performers held the event. Information 92 relating to the performers of the event includes information such as videos, still images, and comments of the performers themselves (for example, comments that excited the audience), and profiles of the performers. Information 93 and 94 relating to chats sent by participants includes chat content of participants attending the event (for example, content supporting the performers), chat content of the participants themselves (for example, content supporting the performers), etc. Information 95 relating to the avatars of the participants includes videos, still images, and sounds of the avatars of the participants.
[0050] Next, the content generated by the generation unit 21 is distributed to the participants of the event via the Internet (step S7).
[0051] Next, the effects of the distribution device 1 according to the embodiment will be described.
[0052] According to the distribution device 1 and distribution method of the above embodiment, generation request information (prompts) can be generated that not only incorporates the first highlight, which is the highlight of the entire event, but also generates content that incorporates the second highlight, which is the highlight of each participant. This allows the generation AI model to generate different content for each participant of the online event. As a result, the generation AI model can be used to generate content that gives participants of the online event a sense of exclusivity.
[0053] The transmission information may include at least one of information about the participants, information about chats sent by the participants while participating in the event, and monitoring information about the participants while participating in the event. This allows the unique information of each participant to be reflected when generating content to be distributed to each participant. As a result, content with a sense of exclusiveness or specialness can be generated.
[0054] The estimation unit 12 may extract the first highlight portion and the second highlight portion based on the number of chats sent by participants when the event is divided into unit times, the presence or absence of passionate words in the chats, or the frequency of appearance of such words, thereby improving the accuracy of extraction of the first highlight portion and the second highlight portion.
[0055] Furthermore, the estimation unit 12 may extract the first highlight portion based on the number of participants who connect to the event via the Internet or the number of participants who leave the state of being connected to the Internet when the event is divided into unit times. In this case as well, the accuracy of extraction of the first highlight portion and the second highlight portion can be improved.
[0056] Furthermore, the estimation unit 12 may extract the first highlight portion based on the amount of reactions of performers that can be extracted from videos, still images, or audio distributed over the Internet when the event is divided into unit times, the presence or absence of reactions of performers with high enthusiasm, or the number of occurrences of reactions of performers with high enthusiasm. In this case, the accuracy of extraction of the first highlight portion can be improved.
[0057] Furthermore, the estimation unit 12 may extract the second highlight portion based on the amount of participant reactions extracted when monitoring participants at the time of their participation in the event, the presence or absence of reactions from highly enthusiastic participants, or the number of occurrences of reactions from highly enthusiastic participants, when the event is divided into unit times. In this case, the accuracy of extraction of the second highlight portion can be improved.
[0058] Furthermore, the estimation unit 12 may estimate the preference tendencies of each participant based on information about the participants (information that can be obtained from the participant's profile, questionnaire data, etc.), and extract the second highlight portion based on the estimated preference tendencies. By taking into account such preference tendencies of each participant, the estimation unit 12 can extract the second highlight portion that is more in line with the actual participant.
[0059] The output unit 13 may generate prompts based on the first and second highlight portions as well as video, still images, or audio of the event distributed via the Internet. This generates content for the generation AI model that includes at least one of the video, images, audio, and chat of the first and second highlight portions. This allows for the distribution of exclusive or special content to each participant.
[0060] The output unit 13 may generate a prompt to generate content including at least one of a video, a still image, and an audio of the first highlight portion and the second highlight portion estimated by the estimation unit 12. This allows the content to be generated for the generation AI model by including an avatar generated based on information about the participant included in the transmitted information. As a result, it is possible to deliver exclusive or special content to each participant.
[0061] The output unit 13 may generate a prompt based on the preference tendency of each participant estimated by the estimation unit 12. By taking into account such preference tendency of each participant, the output unit 13 can generate a prompt that allows the generation unit 21 to generate content that feels more exclusive or special to the participants.
[0062] The output unit 13 may generate avatars of the participants based on information about the participants included in the transmitted information, and may generate prompts to generate content in which the avatars are combined with moving or still images, thereby enabling the delivery of content that feels more exclusive or special to the participants.
[0063] The apparatus and method of the present disclosure may have the following configuration.
[0064] [1] A device comprising: a reception unit that receives transmission information transmitted from each participant for an event that can be participated in via the Internet; an estimation unit that estimates a first highlight portion that is an exciting portion of the event and a second highlight portion that is an exciting portion of the event for each participant based on the transmission information received by the reception unit; and an output unit that generates generation request information to instruct a generative AI model to generate content to be distributed to each participant based on the first highlight portion and the second highlight portion estimated by the estimation unit.
[0065] [2] The device described in [1] above, wherein the outgoing information includes at least one of information about the participant, information about chats sent by the participant while participating in the event, and monitoring information about the participant while participating in the event.
[0066] [3] The device described in [1] or [2] above, wherein the estimation unit extracts the first highlight portion and the second highlight portion based on the number of chats sent by the participants when the event is divided into unit times, the presence or absence of high-energy words included in the chats, or the number of times the high-energy words appear.
[0067] [4] The device according to any one of [1] to [2] above, wherein the estimation unit extracts the first highlight portion based on the number of participants connecting to the event via the Internet or the number of participants who leave the state of being connected to the Internet when the event is divided into unit times.
[0068] [5] The device according to any one of [1] to [4] above, wherein the estimation unit extracts the first highlight part based on the amount of reactions of the performers that can be extracted from the video, still images, or audio distributed via the Internet when the event is divided into unit times, the presence or absence of reactions of the performers with high enthusiasm, or the number of occurrences of reactions of the performers with high enthusiasm.
[0069] [6] The device according to any one of [1] to [5] above, wherein the estimation unit extracts the second highlight portion based on the amount of reactions of the participants extracted when the participants are monitored at the time of their participation in the event when the event is divided into unit time periods, whether or not there are reactions of the participants with high enthusiasm, or the number of times reactions of the participants with high enthusiasm appear, or estimates a preference tendency for each participant based on information about the participants included in the transmitted information, and extracts the second highlight portion based on the estimated preference tendency.
[0070] [7] The device according to any one of [1] to [6] above, wherein the output unit generates the generation request information for generating the content including at least one of a moving image, a still image, and an audio of the first highlight portion and the second highlight portion estimated by the estimation unit.
[0071] [8] The device according to any one of [1] to [7] above, wherein the estimation unit estimates a preference tendency for each participant based on information about the participant included in the transmitted information, and the output unit generates the generation request information based on the preference tendency for each participant estimated by the estimation unit.
[0072] [9] The device described in [7] above, wherein, when the event is an event that is held in real space and distributed via the Internet, the output unit generates avatars of the participants based on information about the participants included in the transmission information, and generates the generation request information to generate the content in which the avatars are combined with the video or the still image; and when the event is an event in which the participants participate as avatars in a virtual space constructed on the Internet, the output unit generates the generation request information to generate the content in which avatars participating in the virtual space are combined with the video or the still image.
[0073]
[10] A method comprising: a reception step of receiving transmission information transmitted from each participant for an event that can be participated in via the Internet; an estimation step of estimating a first highlight portion that is an exciting part of the event and a second highlight portion that is an exciting part of the event for each participant, based on the transmission information received in the reception step; and an output step of generating generation request information for instructing a generative AI model to generate content to be distributed to each participant, based on the first highlight portion and the second highlight portion estimated in the estimation step.
[0074] Although one embodiment has been described above, one aspect of the present disclosure is not limited to the above embodiment, and various modifications are possible without departing from the spirit of the present disclosure.
[0075] In the above embodiment, the distribution device 1 has been described as primarily distributing an event held in real space via the Internet. However, the distribution device 1 can also be applied to events held in a metaverse space, which is a virtual space built on the Internet. In such a configuration, for example, the output unit 13 may generate a prompt that generates content in which avatars participating in the virtual space are combined with video or still images of the event. In this case, content that feels more exclusive or special to participants can be distributed.
[0076] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of at least one of hardware and software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wires, wirelessly, etc.) and these multiple devices. The functional block may also be realized by combining software with the single device or the multiple devices.
[0077] Functions include, but are not limited to, judgment, determination, judgment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, election, establishment, comparison, assumption, expectation, regard, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.
[0078] For example, the distribution device 1 according to an embodiment of the present disclosure may function as a computer that performs processing of a prompt generation method for generating content according to the present disclosure. Fig. 9 is a diagram illustrating an example of a hardware configuration of the distribution device 1 according to an embodiment of the present disclosure. The distribution device 1 described above may be physically configured as a computer including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, etc.
[0079] In the following description, the term "device" may be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the distribution device 1 may be configured to include one or more of the devices shown in the figure, or may be configured to exclude some of the devices.
[0080] Each function in the distribution device 1 is realized by loading specified software (programs) onto hardware such as the processor 1001, memory 1002, etc., so that the processor 1001 performs calculations, controls communication via the communication device 1004, and controls at least one of reading and writing data in the memory 1002 and storage 1003.
[0081] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured by a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, the above-mentioned reception unit 11, estimation unit 12, output unit 13, generation unit 21, event distribution unit 15, content distribution unit 17, etc. may be realized by including the processor 1001.
[0082] The processor 1001 also reads programs (program codes), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes in accordance with the programs. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, the reception unit 11, the estimation unit 12, the output unit 13, the generation unit 21, the event distribution unit 15, and the content distribution unit 17 may be implemented by a control program stored in the memory 1002 and running on the processor 1001, and similar implementations may be used for other functional blocks. While the above-described various processes have been described as being executed by one processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The programs may also be transmitted from a network via a telecommunications line.
[0083] The program executed by processor 1001 may be a program that includes a reception step of receiving transmission information transmitted from each participant for an event that can be joined via the Internet; an estimation step of estimating a first highlight part, which is the exciting part of the event, and a second highlight part, which is the exciting part of the event for each participant, based on the transmission information received in the reception step; and an output step of generating a prompt (generation request information) to instruct a generation AI model to generate content to be distributed to each participant based on the first highlight part and the second highlight part estimated by the estimation step.
[0084] The memory 1002 is a computer-readable recording medium and may be configured, for example, by at least one of a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be referred to as a register, a cache, a main memory (primary storage device), etc. The memory 1002 may store executable programs (program codes), software modules, etc. for implementing the prompt generation method according to an embodiment of the present disclosure.
[0085] Storage 1003 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including at least one of memory 1002 and storage 1003. For example, the above-mentioned storage unit 23 may be realized by including storage 1003.
[0086] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as, for example, a network device, a network controller, a network card, a communication module, etc. The communication device 1004 may be configured to include a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc. to realize at least one of frequency division duplex (FDD) and time division duplex (TDD). For example, the above-mentioned reception unit 11, estimation unit 12, output unit 13, generation unit 21, event distribution unit 15, content distribution unit 17, etc. may be realized by including the communication device 1004.
[0087] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. Note that the input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).
[0088] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.
[0089] The distribution device 1 may also be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.
[0090] Notification of information is not limited to the aspects / embodiments described in this disclosure, and may be performed using other methods.
[0091] Each aspect / embodiment described in the present disclosure may be applied to at least one of systems using LTE (Long Term Evolution), LTE-Advanced (LTE-A), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), FRA (Future Radio Access), NR (New Radio), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, UWB (Ultra-Wide Band), Bluetooth (registered trademark), or other suitable systems, and next-generation systems enhanced based on these. Furthermore, a combination of multiple systems (e.g., a combination of at least one of LTE and LTE-A with 5G, etc.) may also be applied.
[0092] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.
[0093] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be transmitted to another device.
[0094] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).
[0095] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).
[0096] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.
[0097] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0098] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.
[0099] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0100] In addition, terms explained in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings.
[0101] As used in this disclosure, the terms "system" and "network" are used interchangeably.
[0102] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like, all of which are considered to be "determining." "Determining" and "determining" may also include resolving, selecting, choosing, establishing, comparing, and the like, all of which are considered to be "determining." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Also, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.
[0103] The terms "connected," "coupled," or any variation thereof, refer to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using one or more wires, cables, and / or printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.
[0104] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0105] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.
[0106] When the terms "include," "including," and variations thereof are used in this disclosure, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, when the term "or" is used in this disclosure, it is not intended to be an exclusive or.
[0107] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.
[0108] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."
[0109] 1...Distribution device (device), 3...Camera, 5...Terminal, 11...Reception unit, 12...Estimation unit, 13...Output unit, 15...Event distribution unit, 17...Content distribution unit, 21...Generation unit, 23...Storage unit, 1001...Processor, 1002...Memory, 1003...Storage, 1004...Communication device, 1005...Input device, 1006...Output device, 1007...Bus
Claims
1. A device comprising: a reception unit that receives transmission information transmitted from each participant of an event that can be joined via the Internet; an estimation unit that estimates a first highlight portion that is an exciting part of the event and a second highlight portion that is an exciting part of the event for each participant based on the transmission information received by the reception unit; and an output unit that generates generation request information to instruct a generative AI model to generate content to be distributed to each participant based on the first highlight portion and the second highlight portion estimated by the estimation unit.
2. The device of claim 1, wherein the outgoing information includes at least one of information about the participant, information about chats sent by the participant while participating in the event, and monitoring information about the participant while participating in the event.
3. The device described in claim 2, wherein the estimation unit extracts the first highlight portion and the second highlight portion based on the number of chats sent by the participants when the event is divided into unit times, the presence or absence of high-energy words contained in the chats, or the number of times such words appear.
4. The device described in claim 1, wherein the estimation unit extracts the first highlight portion based on the number of participants who connect to the event via the Internet or the number who leave the Internet-connected state when the event is divided into unit time periods.
5. The device described in claim 1, wherein the estimation unit extracts the first highlight portion based on the amount of reactions of performers that can be extracted from video, still images, or audio distributed via the Internet when the event is divided into unit times, the presence or absence of reactions of performers with high enthusiasm, or the number of occurrences of reactions of performers with high enthusiasm.
6. The device described in claim 1, wherein the estimation unit extracts the second highlight portion based on the amount of reactions of the participants extracted when the participants are monitored at the time of their participation in the event when the event is divided into unit time periods, whether or not there are reactions from participants with high enthusiasm, or the number of times reactions from participants with high enthusiasm appear, or estimates the preference trends of each participant based on information about the participants included in the transmitted information, and extracts the second highlight portion based on the estimated preference trends.
7. The device according to claim 1, wherein the output unit generates the generation request information for generating the content including at least one of video, still images, or audio of the first highlight portion and the second highlight portion estimated by the estimation unit.
8. The device of claim 1, wherein the estimation unit estimates the preference trends of each participant based on information about the participant included in the transmitted information, and the output unit generates the generation request information based on the preference trends of each participant estimated by the estimation unit.
9. The device described in claim 7, wherein, when the event is an event that is held in real space and distributed via the Internet, the output unit generates avatars of the participants based on information about the participants included in the transmitted information, and generates the generation request information to generate the content in which the avatars are superimposed on the video or the still image; and when the event is an event in which the participants participate as avatars in a virtual space constructed on the Internet, the output unit generates the generation request information to generate the content in which avatars participating in the virtual space are superimposed on the video or the still image.
10. A method comprising: a reception step of receiving transmission information transmitted from each participant of an event that can be attended via the Internet; an estimation step of estimating a first highlight portion that is an exciting part of the event and a second highlight portion that is an exciting part of the event for each participant based on the transmission information received in the reception step; and an output step of generating generation request information for instructing a generative AI model to generate content to be distributed to each participant based on the first highlight portion and the second highlight portion estimated by the estimation step.
Citation Information
Patent Citations
Video distribution system, computer program used therefor, and control method
JP2021197614A
Servers and Computer Programs
JP7433617B1
Cited By
Content distribution system and computer program
JP7820000B1