System, information processing method, and program
Patent Information
- Application Number
- JP2026068240
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2026-04-16
- Publication Date
- 2026-10-01
- Estimated Expiration
- 2046-04-16
AI Technical Summary
【0030】 本開示によれば、複数の動画データを用いて配信者に有用なコンセプト動画を生成できる。
Smart Images

Figure 0007927363000001_ABST
Abstract
Description
[[Technical Field]]
[0001] The present invention relates to a system, an information processing method and a program. [[Background Art]]
[0002] As an invention relating to a conventional system, for example, the system described in Patent Document 1 is known. In this system, one or more computers acquire video data distributed by a distributor to one or more viewers, and obtain performer actions performed by a performer during distribution of the video data and / or action information indicating viewer actions performed by the one or more viewers on the video data during distribution of the video data, and determine a start point and an end point of a first portion of the video data based on the action information. [[Prior Art Document]] [[Patent Document]]
[0003] [[Patent Document 1]] Japanese Unexamined Patent Application Publication No. 2025-147164 [[Summary of the Invention]] [[Problem to be Solved by the Invention]]
[0004] By the way, there is a demand to generate a concept video useful for a distributor using a plurality of pieces of video data.
[0005] Accordingly, an object of the present invention is to provide a system, an information processing method, and a program that can generate a concept video useful for a distributor using a plurality of pieces of video data. [[Means for Solving the Problem]]
[0006] A first aspect is a system including one or more information processing devices, The control circuit of the one or more information processing devices described above functions as a video data acquisition means, a section information acquisition means, a concept video information acquisition means, and an output means. The aforementioned video data acquisition means acquires multiple video data associated with one or more broadcaster IDs, The section information acquisition means acquires a plurality of section pieces related to a plurality of sections used for generating a concept video in the plurality of video data, The concept video information acquisition means uses the plurality of section information to acquire concept video information used when viewing the concept video which includes the plurality of sections, The output means outputs the concept video information to the terminal. It is a system.
[0007] The second form is, The control circuit of the one or more information processing devices described above functions as a means for acquiring specification information. The specification information acquisition means acquires specification information indicating the generation conditions for the concept video, The section information acquisition means acquires the plurality of section information using the specification information. This is the system described in the first form.
[0008] The third form is, The aforementioned specification information is entered by the user of the terminal. This is the system described in the second form.
[0009] The fourth form is, The memory circuit of the one or more information processing devices stores multiple templates for the concept video. The aforementioned specification information indicates the template selected by the user of the terminal from among the multiple templates. This is the system described in the second form.
[0010] The fifth form is, The video data acquisition means acquires multiple video data associated with one or more broadcaster IDs based on the specification information. This is the system described in either the second, third, or fourth form.
[0011] The sixth form is, The specification information includes at least one of the following: the intended use of the concept video, the theme or keywords of the concept video, the structure of the concept video, the length of the concept video, or information indicating the video data to be referenced. This is a system described in any of the second to fifth forms.
[0012] The seventh form is, The aforementioned multiple sections are: sections in which performers or one or more viewers become excited in the multiple videos shown by the multiple video data; sections in which content related to a specific topic is included in the multiple videos shown by the multiple video data; or sections in which typical content is included in the multiple videos shown by the multiple video data. The system is one of the systems described in the first to sixth forms.
[0013] The eighth form is, The control circuit of the one or more information processing devices functions as a parameter information generation means, The parameter information generation means generates a plurality of parameter information indicating parameters related to the distribution of the plurality of video data, The interval information acquisition means acquires the plurality of interval information using the plurality of parameter information. The system is one of the systems described in the first to seventh forms.
[0014] The ninth form is, The parameter information includes at least one parameter that depends on at least one of: the volume of a video represented by video data; the number of comments or the number of stamps for the video represented by the video data; the number of viewers of the video represented by the video data; or the number of actions that occur on an SNS (Social Network Service) after the video represented by the video data is distributed on said SNS, The system according to the eighth embodiment.
[0015] In a tenth embodiment, the control circuit of the one or more information processing apparatuses functions as text information generating means, said text information generating means generates a plurality of pieces of text information by converting speech included in a plurality of videos represented by said plurality of pieces of video data into text, said section information acquiring means acquires said plurality of pieces of section information using said plurality of pieces of text information, The system according to any one of the first to ninth embodiments.
[0016] In an eleventh embodiment, said section information acquiring means acquires, using said plurality of pieces of text information, a plurality of pieces of section information related to fixed-form sections including fixed-form content in the videos represented by said plurality of pieces of video data, The system according to the tenth embodiment.
[0017] In a twelfth embodiment, the control circuit of the one or more information processing apparatuses functions as specification information acquiring means, said specification information acquiring means acquires specification information including a keyword of said concept video, said section information acquiring means acquires said plurality of pieces of section information using the keyword included in said specification information and text represented by said plurality of pieces of text information, The system according to the tenth embodiment.
[0018] In a thirteenth embodiment, The control circuit of the one or more information processing devices described above functions as a means for acquiring specification information. The specification information acquisition means acquires specification information including the keywords of the concept video, Each of the aforementioned video data contains metadata, The video data acquisition means acquires the plurality of video data using the specification information and the plurality of metadata. The system is one of the first to twelfth forms.
[0019] The 14th form is, The interval information acquisition means acquires the plurality of interval information using a machine learning model. The system is one of the first to thirteenth forms.
[0020] The 15th form is, Each of the aforementioned interval information includes information about the interval that the machine learning model has determined to have a high degree of attention. This is the system described in Form 14.
[0021] The 16th form is, The concept video information acquisition means acquires concept video data representing the concept video formed by combining the multiple sections as the concept video information. The system is one of the first to fifteenth forms.
[0022] The 17th form is, The control circuit of the one or more information processing devices described above functions as a means for acquiring specification information. The specification information acquisition means acquires specification information indicating the generation conditions for the concept video, The concept video information acquisition means acquires concept video data as concept video information, which shows a concept video in which the plurality of sections are combined in an order determined based on the specification information. This is the system described in Form 16.
[0023] The 18th form is, The concept video information acquisition means acquires the concept video information by inputting the multiple section information into the generating AI. This is the system described in Form 16.
[0024] The 19th form is, The aforementioned multiple section information is video data showing the video of the aforementioned multiple sections. This is the system described in Form 18.
[0025] The 20th form is, The concept video information acquisition means acquires the concept video information to which video representations have been added by the generating AI to the videos of the plurality of sections. This is the system described in Form 19.
[0026] The 21st form is, The aforementioned concept video information acquisition means acquires the concept video data having a format for posting to social media as the concept video information. This is a system described in any of the 16th to 20th forms.
[0027] The 22nd form is, The aforementioned concept video information indicates the upload location of the concept video on the distribution platform or social media. The system is one of the first to sixteenth forms.
[0028] The 23rd form is, Information processing method, In the information processing method, the control circuit of one or more information processing devices functions as a video data acquisition means, a section information acquisition means, a concept video information acquisition means, and an output means. The aforementioned video data acquisition means acquires multiple video data associated with one or more broadcaster IDs, The section information acquisition means acquires a plurality of section pieces of information related to the section used to generate the concept video in the plurality of video data, The concept video information acquisition means uses the plurality of section information to acquire concept video information used when viewing the concept video which includes the plurality of sections, The output means outputs the concept video information to the terminal. It is an information processing method.
[0029] The 24th form is, It is a program, The program causes the control circuits of one or more information processing devices to function as video data acquisition means, section information acquisition means, concept video information acquisition means, and output means. The aforementioned video data acquisition means acquires multiple video data associated with one or more broadcaster IDs, The section information acquisition means acquires a plurality of section pieces of information related to the section used to generate the concept video in the plurality of video data, The concept video information acquisition means uses the plurality of section information to acquire concept video information used when viewing the concept video which includes the plurality of sections, The output means outputs the concept video information to the terminal. It is a program. [Effects of the Invention]
[0030] According to this disclosure, it is possible to generate a concept video useful to the broadcaster using multiple video data. [Brief explanation of the drawing]
[0031] [Figure 1] Figure 1 is a block diagram of System 1. [Figure 2] Figure 2 is a block diagram of the broadcaster terminal 10. [Figure 3] Figure 3 is a block diagram of server 110. [Figure 4]Figure 4 is a block diagram of the video server 410. [Figure 5] Figure 5 is a flowchart showing the distribution process. [Figure 6] Figure 6 shows the viewer action table. [Figure 7] Figure 7 shows a point table. [Figure 8] Figure 8 shows an example of a video being streamed. [Figure 9] Figure 9 shows the time evolution of the elevation parameter. [Figure 10] Figure 10 is a flowchart showing the concept video generation process. [Figure 11] Figure 11 is a flowchart showing the concept video generation process related to the first modified example. [Figure 12] Figure 12 shows the change in volume over time. [Figure 13] Figure 13 shows an example of interval extraction in the third modified example. [Modes for carrying out the invention]
[0032] (Embodiment) System 1 according to an embodiment of this disclosure will be described with reference to the drawings. [Overview of System 1] First, I will explain the overall configuration of System 1 with reference to the diagram. Figure 1 is a block diagram of System 1.
[0033] System 1, shown in Figure 1, comprises a broadcaster terminal 10, a server 110, a generation AI server 210, a machine learning model server 310, and a video server 410. The broadcaster terminal 10, server 110, generation AI server 210, machine learning model server 310, and video server 410 can communicate with each other via a communication network. The network can be the internet or an intranet, etc.
[0034] The broadcaster terminal 10 is an information processing device used by the video broadcaster. The broadcaster terminal 10 is, for example, a smartphone, a tablet, a home game console, a portable game console, a personal computer, or a standalone virtual reality (VR) head-mounted display.
[0035] Server 110 is an information processing device used by the operator of a video distribution service. Server 110 is a computer. Server 110 can communicate with the distributor terminal 10, the generation AI server 210, the machine learning model server 310, and the video server 410.
[0036] The Generative AI Server 210 is an information processing device equipped with Generative AI. Generative AI is a large-scale artificial intelligence model used in the field of Natural Language Processing (NLP). By learning from large amounts of text data (web pages, books, articles, etc.), Generative AI can understand patterns in human language and effectively perform Natural Language Generation (NLG) tasks. Generative AI is used in many NLP tasks, such as generating responses to specific questions, automatically generating text, summarizing text, translation, sentiment analysis, image generation, and video generation. Note that the Generative AI Server 210 may be an external server provided by a different operator than the operator of Server 110.
[0037] The machine learning model server 310 is an information processing device equipped with a machine learning model. The machine learning model is a trained model that outputs a judgment result, prediction result, or recognition result based on input data. The machine learning model can, but is not limited to, a neural network, decision tree, support vector machine, etc. The machine learning model server 310 can perform predetermined processing based on the output of the machine learning model. Note that the machine learning model server 310 may be an external server provided by a different business operator than the operator of server 110.
[0038] The video server 410 is an information processing device that stores video data. The video server 410 is used by the operator of the distribution platform. The video server 410 stores an archive of the distributor's past streamed videos. Video data can be retrieved via the distribution platform API or by uploading. Both server 110 and video server 410 are operated by the video distribution platform provider.
[0039] [Structure of the broadcaster's device 10] The structure of the broadcaster terminal 10 will be explained with reference to the diagram. Figure 2 is a block diagram of the broadcaster terminal 10.
[0040] As shown in Figure 2, the broadcaster terminal 10 includes a control circuit 12, a memory circuit 14, a network interface 16, a camera 17, a graphics processing circuit 18, a display 20, an audio processing circuit 22, a speaker 24, and an operation unit 26.
[0041] The memory circuit 14 stores programs and data. The memory circuit 14 is, for example, a combination of ROM (Read Only Memory), RAM (Random Access Memory), and storage (for example, flash memory or hard disk). The program includes, for example, the following: • OS (Operating System) programs • Programs for applications that perform information processing (e.g., web browsers, or applications that request the generation of concept videos)
[0042] The data includes, for example, the following: • Databases referenced in information processing • Data obtained by performing information processing (i.e., the results of performing information processing)
[0043] The control circuit 12 implements the functions of the broadcaster terminal 10 by executing the program stored in the memory circuit 14. The control circuit 12 is, for example, a circuit that includes at least one of the following: ·CPU(Central Processing Unit) ·GPU(Graphic Processing Unit) ·ASIC(Application Specific Integrated Circuit) ·FPGA(Field Programmable Gate Array)
[0044] The network interface 16 controls communication between the broadcaster terminal 10 and an external device. The external device is a server 110.
[0045] Camera 17 captures the surroundings of the broadcaster terminal 10 and generates video data. In this embodiment, camera 17 captures the broadcaster.
[0046] The graphics processing circuit 18 displays an image on the display 20 based on the image data generated by the control circuit 12. The display 20 is either a liquid crystal display or an organic EL (Electro-Luminescence) display.
[0047] The audio processing circuit 22 outputs sound to the speaker 24 based on the audio data generated by the control circuit 12.
[0048] The operation unit 26 generates an operation signal based on the broadcaster's operation and outputs the operation signal to the control circuit 12. The operation unit 26 is, for example, a touch panel, a keyboard, a mouse, etc.
[0049] [Server 110 Structure] Next, the structure of server 110 will be explained with reference to the diagram. Figure 3 is a block diagram of server 110.
[0050] As shown in Figure 3, the server 110 includes a control circuit 112, a memory circuit 114, and a network interface 116.
[0051] The memory circuit 114 stores the program PG and data. The memory circuit 114 is, for example, a combination of ROM (Read Only Memory), RAM (Random Access Memory), and storage (for example, flash memory or hard disk). The program includes, for example, the following program: • OS programs • A program for an application that performs information processing.
[0052] The data includes, for example, the following: • Databases referenced in information processing • Data obtained by performing information processing (i.e., the results of performing information processing)
[0053] The control circuit 112 implements the functions of the server 110 by executing the program PG stored in the memory circuit 114. The control circuit 112 is a circuit that includes, for example, at least one of the following: ·CPU(Central Processing Unit) ·GPU(Graphic Processing Unit) ·ASIC(Application Specific Integrated Circuit) ·FPGA(Field Programmable Gate Array)
[0054] The control circuit 112 includes, as functional blocks, a video data acquisition means 120, a section information acquisition means 122, a concept video information acquisition means 124, an output means 126, a specification information acquisition means 128, a text information generation means 130, and a parameter information generation means 132.
[0055] The network interface 116 controls communication between the server 110 and external devices. The external devices are the broadcaster terminal 10, the generation AI server 210, the machine learning model server 310, and the video server 410.
[0056] [Structure of video server 410] Next, the structure of the video server 410 will be explained with reference to the diagram. Figure 4 is a block diagram of the video server 410.
[0057] As shown in Figure 4, the video server 410 includes a control circuit 412, a memory circuit 414, and a network interface 416.
[0058] The memory circuit 414 stores programs and data. The memory circuit 414 is, for example, a combination of ROM (Read Only Memory), RAM (Random Access Memory), and storage (for example, flash memory or hard disk).
[0059] The memory circuit 414 includes a video database 414a. The video database 414a stores multiple video data previously distributed by the distributor and multiple video data distributed by distributors other than the distributor. The video data includes data of a temporally consecutive sequence of videos and audio data. Each video data is associated with the distributor ID of the distributor who distributed that video data.
[0060] Furthermore, each video data is associated with metadata (date and time, title, category, etc.) and chat logs (viewer comments, stamps, etc.). Metadata is attribute information related to the distribution of video data. Metadata indicates the date and time, title, and category of the video file. Chat logs are records of chat messages posted during the distribution of video data. Chat logs include the message body, posting time, poster ID or username, recipient information, emojis, stamps, etc.
[0061] The control circuit 412 implements the functions of the video server 410 by executing the program stored in the memory circuit 414. The control circuit 412 is a circuit that includes, for example, at least one of the following: ·CPU(Central Processing Unit) ·GPU(Graphic Processing Unit) ·ASIC(Application Specific Integrated Circuit) ·FPGA(Field Programmable Gate Array)
[0062] The network interface 416 controls communication between the video server 410 and an external device. The external device is the server 110.
[0063] [System 1 Operation] Next, the operation of System 1 will be explained with reference to the diagram. System 1 performs distribution processing and concept video generation processing. Distribution processing is the process performed by System 1 when a distributor distributes video data. Through distribution processing, the distributor's video data and viewer actions are stored in the video DB 414a of the video server 410. Concept video generation processing is the process performed by System 1 when generating concept video data from multiple video data. Through concept video generation processing, promotional concept video data according to a predetermined concept is generated from multiple past video data of the distributor.
[0064] The control circuit 112 of server 110 reads the programs stored in the memory circuit 114, causing these programs to execute the operations described below in the control circuit 112 of server 110. The programs then cause the control circuit 112 of server 110 to function as a video data acquisition means 120, a section information acquisition means 122, a concept video information acquisition means 124, an output means 126, a specification information acquisition means 128, a text information generation means 130, and a parameter information generation means 132. In the information processing method, the control circuit 112 of server 110 functions as a video data acquisition means 120, a section information acquisition means 122, a concept video information acquisition means 124, an output means 126, a specification information acquisition means 128, a text information generation means 130, and a parameter information generation means 132.
[0065] [Distribution Processing] First, let's explain the distribution process in System 1. Figure 5 is a flowchart of the distribution process. Figure 6 is a viewer action table. Figure 7 is a point table. Figure 8 is an example of video being distributed. Figure 9 is a time-dependent change in the excitement parameter.
[0066] The memory circuit 114 of server 110 stores the viewer action table shown in Figure 6 and the point table shown in Figure 7. As shown in Figure 6, the viewer action table includes items for viewer, viewer action, and time of action occurrence. For example, the viewer action table records that viewer US1 performed a "like" action at 3:10, viewer US1 performed a "comment" action at 3:11, viewer US2 performed a "gift" action at 3:15, and viewer US3 performed a "like" action at 3:15. The viewer action table corresponds to the chat log.
[0067] As shown in Figure 7, each viewer action is associated with points. Specifically, a "like" is associated with 1 point, a comment with 3 points, and a gift with 10 points.
[0068] The broadcaster initiates video streaming by operating the control unit 26 of the broadcaster terminal 10. The control circuit 12 of the broadcaster terminal 10 starts streaming video data (step S1). Specifically, the control circuit 12 of the broadcaster terminal 10 captures the broadcaster with the camera 17 and generates video data showing the broadcaster's video. The control circuit 12 of the broadcaster terminal 10 transmits the video data to the server 110 via the network interface 16. The network interface 116 of the server 110 receives the video data and outputs it to the control circuit 112. The control circuit 112 of the server 110 streams the video data to the viewer terminals.
[0069] Next, the control circuit 112 of the server 110 determines whether or not a viewer action has been acquired (step S2). A viewer action is a viewer action taken by a viewer on a video during the distribution of video data. As shown in Figure 8, viewer actions include actions such as a viewer pressing the "Like" button, a viewer pressing the "Gift" button, and a viewer sending a comment. The video being distributed shown in Figure 8 includes the broadcaster's avatar video and displays of actions from viewers (entry notification, like, comment, gift, etc.). Specifically, the video shown in Figure 8 includes displays of viewer actions such as an entry notification "US1 has entered the room", an action "US2: Like", a comment "US3: "Hello"", and a gift sent "US4: Gift". If a viewer action is acquired, the process proceeds to step S3. If no viewer action is acquired, the process proceeds to step S6.
[0070] If no viewer actions have been received (N in step S2), the control circuit 112 of the server 110 transmits the video data to the viewer terminal via the network interface 116 (step S6). As a result, the viewer terminal receives the video data. The viewer terminal displays the video data, which does not include viewer action displays indicating viewer actions by multiple viewers. Viewer action displays include heart marks, gift marks, and comments. After this, the process returns to step S2.
[0071] When a viewer action is detected (Y in step S2), the control circuit 112 of the server 110 transmits composite video data corresponding to the viewer action to the viewer terminal via the network interface 116 (step S3). Specifically, if the viewer presses the "Like" button, the control circuit 112 of the server 110 generates composite video data that displays a heart mark on the streamed video. Also, if the viewer presses the "Gift" button, the control circuit 112 of the server 110 generates composite video data that displays a gift mark on the streamed video. Furthermore, if the viewer sends a comment, the control circuit 112 of the server 110 generates composite video data that displays the comment on the streamed video.
[0072] Furthermore, the control circuit 112 of the server 110 records viewer actions (step S4). Specifically, the control circuit 112 of the server 110 records the viewer, the type of viewer action, and the time the action occurred in the viewer action table shown in Figure 6.
[0073] Next, the control circuit 112 of the server 110 determines whether or not video streaming has ended (step S5). If video streaming has not ended, this process returns to step S2. If video streaming has ended, this process proceeds to step S7.
[0074] When video streaming ends (Y in step S5), the control circuit 112 of server 110 sends the video data, metadata, and chat log to video server 410 via the network interface 116. The network interface 416 of video server 410 receives the video data, metadata, and chat log and outputs the video data, metadata, and chat log to the control circuit 412. The control circuit 412 of video server 410 stores the video data, metadata, and chat log in video DB 414a. This completes the streaming process.
[0075] As the above distribution process is performed multiple times at different times, multiple video data associated with the same distributor ID are accumulated in video DB414a. These accumulated video data are used in the concept video generation process described later.
[0076] [Concept video generation process] Next, the concept video generation process in System 1 will be explained with reference to the diagram. Figure 10 is a flowchart of the concept video generation process.
[0077] First, the broadcaster inputs specification information by operating the control panel 26 of the broadcaster terminal 10. The specification information indicates the conditions for generating a concept video. The specification information includes at least one of the following: the purpose of the concept video, the theme or keywords of the concept video, the structure of the concept video, the length of the concept video, or information indicating the video data to be referenced. Specifically, the purpose of the concept video may be for announcements, self-introductions, or a compilation of memorable scenes. The theme or keywords of the concept video may be "horror episode," "singing broadcast," or "competition winning scene." The structure of the concept video may follow a structure pattern such as greeting → build-up → closing. The information indicating the video data to be referenced indicates the scope of the video data to be used when generating the concept video. This information may include using only video data from a specific period (for example, the most recent month), using only video data with the highest number of views, or using only videos from broadcasts selected by the broadcaster. The information indicating the video data to be referenced may be a combination of two or more conditions, such as using only video data from a specific period (for example, the most recent month), using only video data with the highest number of views, or using only videos from broadcasts selected by the broadcaster. The input format for the specification information may be a free-text format entered by the broadcaster, or a format in which the broadcaster selects from several pre-prepared templates. The pre-prepared templates include the specifications for the concept video from which the broadcaster wants to generate short promotional videos, self-introduction videos, and compilation videos of memorable scenes. In this case, the memory circuit 114 of the server 110 stores several templates for the concept video. The specification information indicates the template selected by the broadcaster (user) of the broadcaster terminal 10 from among the several templates. In this embodiment, the broadcaster selects a template for a compilation video of memorable scenes and also selects to use only video data from the most recent month for generating the concept video as information indicating the video data to be referenced. The control circuit 12 of the broadcaster terminal 10 acquires specification information indicating that only the template for the highlight reel video and video data from the most recent month will be used to generate the concept video (step S11).
[0078] Next, the control circuit 12 of the broadcaster terminal 10 transmits the specification information and the broadcaster's broadcaster ID to the server 110 via the network interface 16 (step S12). The network interface 116 of the server 110 receives the specification information and the broadcaster's broadcaster ID and outputs the specification information and the broadcaster's broadcaster ID to the control circuit 112. As a result, the control circuit 112 (specification information acquisition means 128) of the server 110 acquires the specification information indicating the generation conditions for the concept video and the broadcaster's broadcaster ID (step S101).
[0079] Next, the control circuit 112 of the server 110 transmits video data request information to the video server 410 via the network interface 116 (step S102). More specifically, the control circuit 112 of the server 110 generates video data request information based on the specification information and distributor ID obtained in step S101. Specifically, the control circuit 112 of the server 110 generates video data request information by extracting the distributor ID, the conditions for the video data to be referenced, and other search conditions from the specification information. The distributor ID is information that specifies the target distributor. The conditions for the video data to be referenced include a time range (e.g., the last month), playback count rank (top N items), theme, keywords, etc. Other search conditions are filtering conditions using metadata. The video data request information includes information to identify multiple video data to be used in generating the concept video data. That is, the video data request information indicates the search conditions for multiple video data to be used in generating the concept video data. The video data request information includes the distributor ID, the theme or keywords of the concept video, or information indicating at least one of the video data to be referenced. In this embodiment, the specification information indicates that only the template for the highlight reel video and video data from the most recent month will be used to generate the concept video. The video data request information may also be a combination of multiple of these conditions. Therefore, the control circuit 112 of the server 110 generates the video data request information by extracting the distributor ID and the information that only the video data from the most recent month will be used to generate the concept video, based on the fact that only the distributor ID, the template for the highlight reel video, and video data from the most recent month will be used to generate the concept video.
[0080] In this embodiment, we will explain using the example where the broadcaster selects a "memorable moments compilation video" template and specifies "video data to be referenced: most recent month". In this case, the control circuit 112 of the server 110 generates video data request information in the following process. (1) The control circuit 112 of the server 110 extracts search conditions for video data from the specification information. Template type: "Compilation video of memorable scenes" Time range: "The last month" Streamer ID: Streamer ID obtained in Step S101 (2) The control circuit 112 of the server 110 generates video data request information based on the extracted conditions. The document states that only the streamer ID and video data from the most recent month will be included. (3) The control circuit 112 of the server 110 generates video data request information and sends it to the video server 410. This video data request information allows the video server 410 to search and extract video data from the specified distributor for the past month.
[0081] The network interface 416 of the video server 410 receives video data request information and outputs the video data request information to the control circuit 412. As a result, the control circuit 412 of the video server 410 acquires the video data request information (step S201).
[0082] Next, the control circuit 412 of the video server 410 retrieves multiple video data from the video DB 414a based on the video data request information. More specifically, the control circuit 412 of the video server 410 extracts multiple video data from the video DB 414a that satisfy the video data request information. In this embodiment, the control circuit 412 of the video server 410 extracts multiple video data distributed by the distributor with the distributor ID within the past month from the video DB 414a. At this time, the control circuit 412 of the video server 410 also extracts metadata and chat logs associated with the extracted video data from the video DB 414a. Then, the control circuit 412 of the video server 410 transmits the multiple video data and the metadata and chat logs associated with the multiple video data to the server 110 via the network interface 416 (step S202).
[0083] The network interface 116 of server 110 receives multiple video data, metadata associated with the multiple video data, and chat logs, and outputs the multiple video data, metadata associated with the multiple video data, and chat logs to the control circuit 112. As a result, the control circuit 112 (video data acquisition means 120) of server 110 acquires multiple video data related to the broadcaster ID, metadata associated with the multiple video data, and chat logs (step S103). Since the video data request information is generated based on the specification information, in steps S102 and S103, the control circuit 112 (video data acquisition means 120) of server 110 acquires multiple video data related to the broadcaster ID based on the specification information.
[0084] Next, the control circuit 112 (parameter information generation means 132) of the server 110 generates a plurality of parameter information indicating parameters related to the distribution of a plurality of video data (step S104). The parameter information includes at least one parameter that depends on the volume of the video indicated by the video data, the number of comments or stamps on the video indicated by the video data, the number of viewers of the video indicated by the video data, or the number of actions that occurred on the SNS (Social Network Service) after the video indicated by the video data was distributed on the SNS. In this embodiment, the parameter information indicates an excitement parameter that depends on the number of comments or stamps on the video indicated by the video data. That is, the parameter information indicates an excitement parameter that indicates the degree of excitement that depends on the number of viewer actions. The method for calculating the excitement parameter will be described below.
[0085] The control circuit 112 of server 110 refers to the chat log contained in each of the multiple video data. The chat log records the viewer, the type of viewer action, and the time the action occurred, as shown in Figure 6. The control circuit 112 of server 110 refers to the point table shown in Figure 7 and obtains points corresponding to each viewer action. Specifically, 1 point is associated with a "like," 3 points with a "comment," and 10 points with a "gift." The control circuit 112 of server 110 calculates the cumulative value of viewer action points within a predetermined time window (for example, 10 seconds) as the excitement parameter. The control circuit 112 of server 110 performs this process for each of the multiple video data. As a result, the control circuit 112 of server 110 obtains the graph shown in Figure 9. In Figure 9, the horizontal axis shows the time of the video, and the vertical axis shows the value of the excitement parameter. In this embodiment, the parameter information is represented by the graph shown in Figure 9.
[0086] The parameter information includes at least one of the following: the volume of the video indicated by the video data, the number of comments or stamps on the video indicated by the video data, the number of viewers of the video indicated by the video data, or the number of actions that occurred on SNS (Social Network Service) after the video indicated by the video data was distributed on SNS. In this embodiment, the parameter information includes an excitement parameter based on the cumulative value of viewer action points. Note that "parameters related to distribution" are not limited to parameters of the video data during distribution. It may also include data such as the number of actions on SNS after distribution, such as comments, likes, and gifts, as well as the number of retweets and quotes.
[0087] Next, the control circuit 112 (interval information acquisition means 122) of the server 110 acquires multiple interval video data (multiple interval information) related to intervals used to generate a concept video in multiple video data using specification information and multiple parameter information (step S105). Interval information is information related to intervals of video data used to generate a concept video. Interval information may be interval video data that is an interval of video data, the start time and end time of an interval of video data, or information that can specify a time, such as the start time and the length of the interval, or the start frame and end frame. Furthermore, interval information is not limited to the above information as long as it can identify an interval of video data, and may also be, for example, a video identifier, score, tag, or reason. In this embodiment, multiple interval information is interval video data that shows the video of multiple intervals.
[0088] Furthermore, the specifications indicate that only the template for the highlight reel video and video data from the most recent month will be used to generate the concept video data. The highlight reel video is a video created by combining the exciting sections from the videos shown in each video data. Therefore, the memory circuit 114 of the server 110 stores information for the highlight reel video template indicating that the exciting sections from the videos shown in each video data will be combined. The concept video then includes multiple exciting sections from multiple videos shown in multiple video data, where the performers or one or more viewers were excited. The control circuit 112 of the server 110 identifies the section in which the excitement parameter exceeds a predetermined threshold p0 as the excitement section A1. The excitement section A1 is identified by a starting point SP1, which is the time when the excitement parameter is greater than or equal to the threshold p0, and an ending point EP1, which is the time when the excitement parameter is less than or equal to the threshold p0. The control circuit 112 of the server 110 performs this process for each of the multiple video data, extracting multiple exciting sections from the multiple video data. The control circuit 112 of server 110 obtains information that identifies the video data (distribution ID, distribution date and time, archive URL, etc.) and information that identifies the exciting sections (start timestamp and end timestamp of the exciting section). For example, the control circuit 112 of server 110 extracts the section from 3 minutes 10 seconds to 3 minutes 45 seconds of video data A, the section from 15 minutes 20 seconds to 16 minutes 05 seconds of video data B, and the section from 22 minutes 40 seconds to 23 minutes 10 seconds of video data C as exciting sections.
[0089] Next, the control circuit 112 of the server 110 obtains segment video data corresponding to each of the extracted peak sections from the multiple video data. The segment video data is video data that includes video and audio from the start timestamp to the end timestamp of the peak section. In this way, the control circuit 112 of the server 110 obtains multiple segment video data from the multiple video data.
[0090] Next, the control circuit 112 (concept video information acquisition means 124) of the server 110 acquires concept video information used when viewing a concept video that includes multiple sections, using multiple section video data (section information) (step S106). Concept video information is information necessary when playing a concept video. Concept video information is concept video data. Furthermore, concept video information is not limited to the concept video data itself, but may also include intermediate metadata. Examples of intermediate metadata include title, thumbnail image, description, draft post text, draft hashtags, and storage location information (URL indicating the storage location of the concept video data). In this embodiment, the control circuit 112 (concept video information acquisition means 124) of the server 110 acquires concept video data that represents a concept video in which multiple sections are combined as concept video information. Therefore, the control circuit 112 (concept video information acquisition means 124) of the server 110 generates concept video data by combining the section video data of multiple exciting sections extracted from the multiple video data in a predetermined order. The predetermined order is, for example, the chronological order of the video data distribution. In this case, the control circuit 112 of the server 110 may perform editing processes such as cutting out unnecessary parts, automatically adding subtitles, converting the aspect ratio (for example, converting a horizontal streaming video to a vertical short video), and adding logos or captions.
[0091] Furthermore, the control circuit 112 (concept video information acquisition means 124) of the server 110 may acquire concept video data having a format for posting to social media as concept video information. For example, the control circuit 112 of the server 110 converts the video to a format appropriate to the social media platform to which it will be posted, such as a vertical short video format or a horizontal standard format.
[0092] Next, the control circuit 112 of the server 110 transmits the concept video data to the video server 410 via the network interface 116 (step S107). The network interface 416 of the video server 410 receives the concept video data and outputs the concept video data to the control circuit 412. The control circuit 412 of the video server 410 acquires the concept video data (step S203) and stores the concept video data in the video DB 414a (step S204).
[0093] Furthermore, the control circuit 112 (output means) of the server 110 transmits (outputs) concept video data (concept video information) to the distributor terminal 10 via the network interface 116 (step S108). The network interface 16 of the distributor terminal 10 receives the concept video data and outputs the concept video data to the control circuit 12. As a result, the control circuit 12 of the distributor terminal 10 acquires the concept video data (step S13). This allows the distributor to display the concept video indicated by the concept video data on the display 20 by operating the operation unit 26 of the distributor terminal 10.
[0094] [effect] According to System 1, a concept video useful to the broadcaster can be generated using multiple video data. More specifically, the control circuit 112 (video data acquisition means 120) of Server 110 acquires multiple video data related to the broadcaster ID. The control circuit 112 (section information acquisition means 122) of Server 110 acquires multiple section video data (section information) related to the sections used to generate the concept video from the multiple video data. The control circuit 112 (concept video information acquisition means 124) of Server 110 uses the multiple section video data (section information) to acquire concept video data (concept video information) used when viewing a concept video that includes multiple sections. The control circuit 112 (output means 126) of Server 110 outputs the concept video data (concept video information) to the broadcaster terminal 10. As a result, System 1 can not only generate a digest from a single video data, but also generate concept video data that is more useful to the broadcaster using multiple video data.
[0095] In System 1, the control circuit 112 (specification information acquisition means 128) of Server 110 acquires specification information indicating the conditions for generating a concept video. The control circuit 112 (section information acquisition means 122) of Server 110 acquires multiple section video data (section information) using the specification information. As a result, System 1 can generate concept video data based on generation conditions that align with the distributor's intentions.
[0096] In System 1, specification information is entered by the broadcaster (user) on the broadcaster terminal 10. This allows the broadcaster to specify the generation conditions themselves, enabling System 1 to generate concept video data that accurately reflects the broadcaster's intentions.
[0097] In System 1, the memory circuit 114 of the server 110 stores multiple templates for concept videos. The specification information indicates the template selected by the broadcaster (user) on the broadcaster terminal 10 from among the multiple templates. This allows the broadcaster to easily input specification information by selecting the desired template from the pre-prepared templates.
[0098] In System 1, the control circuit 112 (video data acquisition means 120) of the server 110 acquires multiple video data related to the distributor ID based on the specification information. As a result, System 1 can appropriately set the range of multiple video data to be acquired based on the specification information, and thus efficiently acquire multiple video data to be used in generating concept videos.
[0099] In System 1, the specification information includes at least one of the following: the intended use of the concept video, the theme or keywords of the concept video, the structure of the concept video, the length of the concept video, or information indicating the video data to be referenced. This allows the distributor to specify the conditions for generating the concept video from various perspectives. As a result, System 1 can generate optimal concept video data according to the intended use.
[0100] In System 1, multiple segments are sections in a video, represented by multiple video data, where the performers or one or more viewers showed excitement. This allows System 1 to automatically extract segments with positive reactions from the performers or one or more viewers and include them in the concept video.
[0101] In System 1, the control circuit 112 (parameter information generation means 132) of Server 110 generates multiple parameter pieces of information indicating parameters related to the distribution of multiple video data. The control circuit 112 (interval information acquisition means 122) of Server 110 acquires multiple interval pieces of information using the multiple parameter pieces of information. As a result, System 1 can objectively extract multiple intervals based on quantitative parameters related to distribution.
[0102] In System 1, the parameter information includes at least one parameter that depends on the volume of the video indicated by the video data, the number of comments or stamps on the video indicated by the video data, the number of viewers of the video indicated by the video data, or the number of actions that occurred on social media after the video indicated by the video data was distributed on social media. This allows System 1 to identify exciting segments based on various indicators such as volume and the reactions of one or more viewers, and to extract video data for segments that closely match the broadcaster's preferences.
[0103] (First variation) The following describes system 1a related to the first modified example. Figure 11 is a flowchart of the concept video generation process related to the first modified example.
[0104] System 1a differs from System 1 in that it uses the URL of the concept video data instead of the concept video data itself as concept video information. The operation of System 1a will be explained below, focusing on the differences from the operation of System 1.
[0105] The process up to step S106 is the same as in Figure 10. In step S107, the control circuit 112 of the server 110 transmits the concept video data to the video server 410 via the network interface 116. The network interface 416 of the video server 410 receives the concept video data and outputs the concept video data to the control circuit 412. The control circuit 412 of the video server 410 acquires the concept video data (step S203) and stores the concept video data in the video DB 414a (step S204).
[0106] Furthermore, the control circuit 412 of the video server 410 generates concept video information based on the concept video data (step S205). The concept video information indicates the upload destination of the concept video on the distribution platform. The concept video information includes a URL for playing the concept video data. In other words, the concept video information indicates the storage location of the concept video data. Thus, in system 1a, the concept video data is not the concept video information, but the URL of the concept video data is the concept video information. The control circuit 412 of the video server 410 transmits the concept video information to the server 110 via the network interface 416 (step S206). The network interface 116 of the server 110 receives the concept video information and outputs the concept video information to the control circuit 112. In this way, through the processing in steps S106, S107 and S118, the control circuit 112 of the server 110 (concept video information acquisition means 124) uses multiple section video data (section information) to acquire concept video information used when viewing a concept video that includes multiple sections (step S118).
[0107] Next, the control circuit 112 (output means 126) of the server 110 transmits (outputs) concept video information including the URL to the distributor terminal 10 via the network interface 116 (step S119). The network interface 16 of the distributor terminal 10 receives the concept video information and outputs the concept video information to the control circuit 12. As a result, the control circuit 12 of the distributor terminal 10 acquires the concept video information (step S23).
[0108] The control circuit 12 of the broadcaster terminal 10 can stream the concept video based on the URL included in the concept video information. Furthermore, the broadcaster can make the concept video public by sharing the URL on social media, etc.
[0109] According to System 1a, the concept video information indicates the upload destination of the concept video on the distribution platform or social media. This eliminates the need for the distributor terminal 10 to download the concept video data. Therefore, the storage capacity consumption of the distributor terminal 10 is reduced. In addition, since the distributor can publish the concept video simply by sharing the URL with third parties, the distribution of the concept video becomes easier.
[0110] (Second variation) The following describes system 1b, which relates to the second modified example. Figure 12 shows the change in volume over time.
[0111] System 1b differs from System 1 in its method of calculating the excitement parameter. System 1b calculates the excitement parameter based on the broadcaster's actions, which include the broadcaster's volume, the content of their speech, and their movements. In this embodiment, System 1b calculates the excitement parameter based on the broadcaster's volume in step S104 of Figure 10. System 1b will be described below, focusing on these differences.
[0112] The video shown in Figure 12 includes a streamer who is screaming, "Ahhh!" As shown, when a streamer is surprised or excited, the volume of their voice increases. The control circuit 112 (parameter information generation means 132) of the server 110 analyzes the audio data contained in each of the multiple video data and calculates the streamer's voice volume along the time axis.
[0113] In Figure 12, the horizontal axis represents the time in the video, and the vertical axis represents the volume value. The control circuit 112 of the server 110 identifies the section in which the volume (excitement parameter) exceeds a predetermined threshold p1 as the excitement section A1. However, in this embodiment, the control circuit 112 of the server 110 identifies the section between the starting point SPa1, which is before the starting point SP1 where the volume is equal to or greater than the threshold p1, and the ending point EPa1, which is before the ending point EP1 where the volume is less than or equal to the threshold p1, as the excitement section A1.
[0114] According to System 1b, exciting segments can be extracted based on the broadcaster's volume. This allows System 1b to extract segments where the broadcaster is excited, even in scenes with little viewer interaction. Furthermore, System 1b can accurately extract exciting segments by using both an excitement parameter based on viewer interaction and an excitement parameter based on volume.
[0115] (Third variation) The system 1c relating to the third modified example is described below. Figure 13 shows an example of interval extraction in the third modified example.
[0116] System 1c differs from System 1 in that it extracts fixed or topical segments based on text information. The following describes System 1c, focusing on these differences.
[0117] In step S11, the distributor inputs specification information. Then, in step S12, the distributor terminal 10 transmits the specification information to the server 110. As a result, in step S101 in Figure 10, the control circuit 112 (specification information acquisition means 128) of the server 110 acquires specification information including the keyword of the concept video. As shown in Figure 13, for example, the keyword is "DB". Then, in step S102, the control circuit 112 of the server 110 transmits video data request information including the keyword contained in the specification information to the video server 410 via the network interface 116. In step S202, the control circuit 412 of the video server 410 extracts multiple video data associated with metadata containing the keyword "DB", which is included in the video data request information acquired in step S201, from the video DB 414a. Furthermore, the control circuit 412 of the video server 410 transmits multiple video data associated with metadata containing "DB" to the server 110 via the network interface 416. As a result, in step S103 of Figure 10, the control circuit 112 of the server 110 acquires multiple video data related to the keyword "DB". In this way, in system 1c, the control circuit 112 (video data acquisition means 120) of the server 110 acquires multiple video data using specification information and multiple metadata.
[0118] In step S104 of Figure 10, the control circuit 112 of the server 110 generates multiple text information instead of setting parameter information. More specifically, the control circuit 112 of the server 110 (text information generation means 130) generates multiple text information by converting the audio contained in multiple videos, which are represented by multiple video data, into text. The control circuit 112 of the server 110 generates the text information using existing speech recognition technology. The text information is associated with the time axis of each video data. The text information is information that has been transcribed from what the broadcaster said.
[0119] In step S105 of Figure 10, the control circuit 112 (section information acquisition means 122) of the server 110 uses multiple text information to acquire multiple section video data (section information) related to a fixed section containing standard content in a video shown by multiple video data. Specifically, the control circuit 112 of the server 110 extracts sections containing specific text (standard expression patterns such as greetings, self-introductions, closing remarks, and announcements) from the text information of the multiple video data as fixed sections. The extraction of fixed sections is not limited to exact string matches, but is performed by determining synonyms and paraphrases. For example, the control circuit 112 of the server 110 identifies greeting phrases such as "Good morning" and "Good evening" in addition to "Hello," or summarizing phrases such as "In short" and "That is to say" as standard expression patterns. For example, as shown in Figure 13, the control circuit 112 of the server 110 extracts the section in which the broadcaster says "Hello" at the beginning of the broadcast as opening section A11. Opening section A11 is a standard section containing a greeting phrase from the broadcaster. Similarly, the control circuit 112 of server 110 extracts the section in which the broadcaster says "In summary" towards the end of the broadcast as closing section A13. Closing section A13 is a standard section containing a closing phrase from the broadcaster. Thus, in system 1c, multiple sections are standard sections containing standard content in a video represented by multiple video data. As described above, the control circuit 112 of server 110 extracts sections containing standard speech content that the broadcaster repeatedly uses as standard sections.
[0120] Furthermore, in step S105 of Figure 10, the control circuit 112 (section information acquisition means 122) of the server 110 acquires multiple section information using keywords included in the specification information and text indicated by multiple text information. Topic section extraction includes keyword matching, synonym matching, and determination based on semantic similarity as methods for matching keywords and text. As shown in Figure 13, for example, if the keyword is "DB", the control circuit 112 of the server 110 extracts the section in which the broadcaster says "Today's topic is DB" as topic section A12. Topic section A12 is a section that contains content related to the keyword. In this way, in system 1c, multiple sections are topic sections that contain content related to a specific topic in a video indicated by multiple video data. In addition, sections where clumsy or awkward moments appear, or sections with cute reactions, etc., can also be extracted as topic sections by system 1c by setting keywords corresponding to these sections.
[0121] System 1c can extract fixed segments and / or topic segments based on text information. This allows System 1c to extract segments that are difficult to identify using excitement parameters.
[0122] (Fourth variation) The following describes System 1d, which relates to the fourth modified example. System 1d differs from System 1 in that it acquires segment video data using a machine learning model.
[0123] In system 1d, in step S105 of Figure 10, the control circuit 112 (interval information acquisition means 122) of server 110 acquires multiple interval video data (interval information) using a machine learning model. Specifically, the control circuit 112 of server 110 transmits the multiple video data acquired in step S103, the audio data contained in the multiple video data, the chat logs associated with the multiple video data, and the metadata to the machine learning model server 310 via the network interface 116. The machine learning model server 310 stores a machine learning model that has learned past clipping examples and viewer reaction data. The machine learning model server 310 provides the machine learning model with at least one of the video data, audio data, chat logs, and metadata as input and calculates an attention score for each interval in the multiple video data. For example, the machine learning model server 310 may identify intervals with an attention score above a predetermined threshold as high-attention intervals, or it may arrange the intervals in order of attention score and identify the top predetermined number of intervals as high-attention intervals. The machine learning model server 310 extracts segment video data for multiple segments identified by the machine learning model from the multiple video data. Thus, in system 1d, each of the multiple segment video data (segment information) contains information about the segment that the machine learning model has determined to have a high degree of attention.
[0124] System 1d can identify highly relevant segments using a machine learning model. This means that System 1d can automatically extract segments that capture viewers' attention without the broadcaster having to set thresholds or input keywords.
[0125] (Fifth variation) The following describes System 1e, which relates to the fifth modified example. System 1e differs from System 1 in that it determines the order in which segment video data are combined based on specification information rather than the chronological order of distribution.
[0126] In system 1e, in step S106 of Figure 10, the control circuit 112 (concept video information acquisition means 124) of server 110 acquires concept video data as concept video information, which shows a concept video in which multiple sections are combined in an order determined based on the specification information. Specifically, the order determined based on the specification information is the order in which the broadcaster entered the specification information in step S11. For example, if the broadcaster selects a template, the order determined based on the specification information is the order defined in the template. The order defined in the template is, for example, greeting → excitement → closing.
[0127] Furthermore, in step S11, the broadcaster may input as specification information that they will arrange the segment video data in order of the magnitude of the excitement parameter.
[0128] System 1e can determine the order in which segment video data are combined based on the specification information. This allows System 1e to generate concept video data with a configuration that aligns with the distributor's intentions.
[0129] (Sixth variation) The following describes System 1f, which relates to the sixth modification. System 1f differs from System 1 in that it uses a generation AI in generating concept video data. The following describes System 1f, focusing on these differences.
[0130] In System 1, in step S106 of Figure 10, the control circuit 112 of Server 110 combines multiple segment video data to generate concept video data. In contrast, in System 1f, in step S106, the control circuit 112 (concept video information acquisition means 124) of Server 110 acquires concept video data (concept video information) by inputting multiple segment video data (segment information) into the generation AI.
[0131] More specifically, the control circuit 112 of the server 110 transmits the multiple section video data and prompts acquired in step S105 to the generation AI server 210 via the network interface 116. In this case, the multiple section video data (section information) is video data that represents videos of multiple sections. The prompts include instructions for generating concept video data by combining the multiple section video data. The prompts also include instructions for adding the video expression described later to the concept video shown by the concept video data. The information input to the generation AI server 210 may include, in addition to the multiple section video data, specification information, section order information, tags corresponding to each section, extraction reason, information indicating the distributor's worldview, usage information, length information, text for subtitle generation, and specifications regarding BGM or sound effects. The generation AI server 210 generates concept video data based on the multiple section video data. Specifically, the generation AI server 210 generates concept video data in which video expressions are added by the generation AI to the videos of multiple sections. Specific examples of video expressions include the following patterns.
[0132] Pattern 1: Adding subtitles according to specifications Pattern 2: Adding narration according to specifications Pattern 3: Adding transition effects to smoothly connect extracted scenes. Pattern 4: Generating a title based on the content of the concept video or the platform where it will be posted. Pattern 5: Adding visual emphasis to highlight key sections, points, or highlights. Pattern 6: Adding background music that suits the atmosphere or purpose of the concept video. Pattern 7: Generating thumbnail candidate images or thumbnail text according to the posting destination. Pattern 8: Adding background images, visual elements, or decorative elements that match the streamer's worldview. Pattern 9: Adding a digest or condensed summary based on the content of the broadcast.
[0133] The generation AI server 210 transmits concept video data with added video representations to the server 110. The network interface 116 of the server 110 receives the concept video data and outputs it to the control circuit 112. As a result, the control circuit 112 (concept video information acquisition means 124) of the server 110 acquires concept video data (concept video information) with video representations added by the generation AI for multiple video segments.
[0134] According to system 1f, the control circuit 112 (concept video information acquisition means 124) of server 110 acquires concept video data (concept video information) by inputting multiple section video data (section information) into the generating AI. As a result, system 1f can generate concept video data that includes video expressions that cannot be realized by simply combining multiple section video data. Furthermore, multiple section video data (section information) is video data that represents videos of multiple sections. As a result, the generating AI can generate video expressions after understanding the content of the video and audio.
[0135] Furthermore, in system 1f, the control circuit 112 (concept video information acquisition means 124) of server 110 acquires concept video data (concept video information) to which video expressions have been added by the AI for multiple video segments. As a result, system 1f can generate concept video data that includes video expressions such as the generation of background images and production materials, the generation of transition videos, and the automatic generation of summary narration.
[0136] (Other embodiments) Furthermore, the system according to the present invention is not limited to systems 1, 1a to 1f, but can be modified within the scope of its gist. Also, the configurations of systems 1, 1a to 1f may be combined in any way.
[0137] Furthermore, the above embodiments and the first to sixth modified examples can be combined as follows.
[0138] First, we will describe the combination of the embodiment and one modification. In the first combination, the embodiment and the first modification are combined. In the first combination, the excitement intervals are extracted based on the cumulative value of viewer action points, while the URL of the concept video data is used as concept video information. In the second combination, the embodiment and the second modification are combined. In the second combination, the excitement intervals are extracted by using an excitement parameter based on volume in addition to an excitement parameter based on the cumulative value of viewer action points. In the third combination, the embodiment and the third modification are combined. In the third combination, in addition to extracting the excitement intervals based on the cumulative value of viewer action points, a fixed interval or topic interval is extracted based on text information. In the fourth combination, the embodiment and the fourth modification are combined. In the fourth combination, in addition to extracting the excitement intervals based on the cumulative value of viewer action points, a machine learning model is used to estimate intervals with high attention. In the fifth combination, the embodiment and the fifth modification are combined. In the fifth combination, the excitement intervals are extracted based on the cumulative value of viewer action points, and the order in which the intervals are combined is determined based on the specification information. In the sixth combination, the embodiment and the sixth modified example are combined. In the sixth combination, the excitement intervals are extracted based on the cumulative value of viewer action points, and concept video data with added video expression is generated by a generating AI.
[0139] Next, we will describe combinations of the embodiment with two modifications. In the seventh combination, the embodiment is combined with the first and second modifications. In the seventh combination, the URL is used as concept video information while extracting the excitement intervals based on the cumulative value of viewer action points and volume. In the eighth combination, the embodiment is combined with the first and third modifications. In the eighth combination, the URL is used as concept video information while extracting the excitement intervals based on the cumulative value of viewer action points and interval extraction based on text information. In the ninth combination, the embodiment is combined with the first and fourth modifications. In the ninth combination, the URL is used as concept video information while extracting the excitement intervals based on the cumulative value of viewer action points and interval estimation using a machine learning model. In the tenth combination, the embodiment is combined with the first and fifth modifications. In the 10th combination, the excitement intervals are extracted based on the cumulative value of viewer action points, and the intervals are combined in an order based on the specification information, while using a URL as concept video information. In the 11th combination, the embodiment, the first modification, and the sixth modification are combined. In the 11th combination, the excitement intervals are extracted based on the cumulative value of viewer action points, and concept video data with added video expression is generated by a generation AI, while using a URL as concept video information. In the 12th combination, the embodiment, the second modification, and the third modification are combined. In the 12th combination, the extraction of excitement intervals based on the cumulative value of viewer action points and volume is used in combination with interval extraction based on text information. In the 13th combination, the embodiment, the second modification, and the fourth modification are combined. In the 13th combination, the extraction of excitement intervals based on the cumulative value of viewer action points and volume is used in combination with interval estimation by a machine learning model. In the 14th combination, the embodiment, the second modification, and the fifth modification are combined. In the 14th combination, the climax sections are extracted based on the cumulative value of viewer action points and volume, and the sections are joined in an order based on the specification information. In the 15th combination, the embodiment, the second modification, and the sixth modification are combined.In the 15th combination, the excitement intervals are extracted based on the cumulative value of viewer action points and volume, and concept video data is generated with added video expression by a generating AI. In the 16th combination, the embodiment, the 3rd modification, and the 4th modification are combined. In the 16th combination, the excitement intervals are extracted based on the cumulative value of viewer action points, while interval extraction based on text information and interval estimation by a machine learning model are used in combination. In the 17th combination, the embodiment, the 3rd modification, and the 5th modification are combined. In the 17th combination, the excitement intervals are extracted based on the cumulative value of viewer action points, intervals are extracted based on text information, and intervals are combined in an order based on specification information. In the 18th combination, the embodiment, the 3rd modification, and the 6th modification are combined. In the 18th combination, the excitement intervals are extracted based on the cumulative value of viewer action points, intervals are extracted based on text information, and concept video data is generated with added video expression by a generating AI. In the 19th combination, the embodiment, the 4th modification, and the 5th modification are combined. In the 19th combination, the excitement intervals are extracted based on the cumulative value of viewer action points, the intervals are estimated by a machine learning model, and the intervals are joined in an order based on the specification information. In the 20th combination, the embodiment, the 4th modification, and the 6th modification are combined. In the 20th combination, the excitement intervals are extracted based on the cumulative value of viewer action points, the intervals are estimated by a machine learning model, and concept video data with added video expression is generated by a generative AI. In the 21st combination, the embodiment, the 5th modification, and the 6th modification are combined. In the 21st combination, the excitement intervals are extracted based on the cumulative value of viewer action points, and concept video data with added video expression is generated by a generative AI in an order based on the specification information.
[0140] Next, we will describe combinations of the embodiment with three variations. In the 22nd combination, the embodiment is combined with the first, second, and third variations. In the 22nd combination, the URL is used as concept video information while using both the extraction of exciting intervals based on the cumulative value of viewer action points and volume, and interval extraction based on text information. In the 23rd combination, the embodiment is combined with the first, second, and fourth variations. In the 23rd combination, the URL is used as concept video information while using both the extraction of exciting intervals based on the cumulative value of viewer action points and volume, and interval estimation by a machine learning model. In the 24th combination, the embodiment is combined with the first, second, and fifth variations. In the 24th combination, the exciting intervals are extracted based on the cumulative value of viewer action points and volume, and the intervals are combined in an order based on the specification information while using the URL as concept video information. In the 25th combination, the embodiment is combined with the first, second, and sixth variations. In the 25th combination, the excitement intervals are extracted based on the cumulative value of viewer action points and volume, and concept video data with added video expression is generated by a generation AI, while using a URL as concept video information. In the 26th combination, the embodiment is combined with the first, third, and fourth variations. In the 26th combination, the excitement intervals are extracted based on the cumulative value of viewer action points, while using a combination of interval extraction based on text information and interval estimation by a machine learning model, while using a URL as concept video information. In the 27th combination, the embodiment is combined with the first, third, and fifth variations. In the 27th combination, the excitement intervals are extracted based on the cumulative value of viewer action points, intervals are extracted based on text information, and intervals are combined in an order based on specification information, while using a URL as concept video information. In the 28th combination, the embodiment, the first modification, the third modification, and the sixth modification are combined.In the 28th combination, the excitement intervals are extracted based on the cumulative value of viewer action points, intervals are extracted based on text information, concept video data with added video expression is generated by a generative AI, and a URL is used as the concept video information. In the 29th combination, the embodiment is combined with the first, fourth, and fifth variations. In the 29th combination, the excitement intervals are extracted based on the cumulative value of viewer action points, intervals are estimated by a machine learning model, intervals are combined in an order based on specification information, and a URL is used as the concept video information. In the 30th combination, the embodiment is combined with the first, fourth, and sixth variations. In the 30th combination, the excitement intervals are extracted based on the cumulative value of viewer action points, intervals are estimated by a machine learning model, concept video data with added video expression is generated by a generative AI, and a URL is used as the concept video information. In the 31st combination, the embodiment is combined with the first, fifth, and sixth variations. In the 31st combination, while extracting exciting intervals based on the cumulative value of viewer action points, concept video data with added video expression by a generating AI is generated in an order based on the specification information, and a URL is used as the concept video information. In the 32nd combination, the embodiment is combined with the second, third, and fourth variations. In the 32nd combination, the extraction of exciting intervals based on the cumulative value of viewer action points and volume, interval extraction based on text information, and interval estimation by a machine learning model are used in combination. In the 33rd combination, the embodiment is combined with the second, third, and fifth variations. In the 33rd combination, the extraction of exciting intervals based on the cumulative value of viewer action points and volume, and interval extraction based on text information are used in combination, and the intervals are joined in an order based on the specification information. In the 34th combination, the embodiment is combined with the second, third, and sixth variations. In the 34th combination, the cumulative value of viewer action points and the extraction of exciting sections based on volume, along with section extraction based on text information, are used in combination to generate concept video data with added video expression by a generating AI.In the 35th combination, the embodiment is combined with the second, fourth, and fifth variations. In the 35th combination, the extraction of excitement intervals based on the cumulative value of viewer action points and volume, along with interval estimation by a machine learning model, is used in combination to combine the intervals in the order based on the specification information. In the 36th combination, the embodiment is combined with the second, fourth, and sixth variations. In the 36th combination, the extraction of excitement intervals based on the cumulative value of viewer action points and volume, along with interval estimation by a machine learning model, is used in combination to generate concept video data with added video expression by a generating AI. In the 37th combination, the embodiment is combined with the second, fifth, and sixth variations. In the 37th combination, excitement intervals are extracted based on the cumulative value of viewer action points and volume, and concept video data with added video expression by a generating AI is generated in the order based on the specification information. In the 38th combination, the embodiment is combined with the third, fourth, and fifth modified examples. In the 38th combination, while extracting excitement intervals based on the cumulative value of viewer action points, interval extraction based on text information and interval estimation by a machine learning model are used in combination, and the intervals are joined in an order based on the specification information. In the 39th combination, the embodiment is combined with the third, fourth, and sixth modified examples. In the 39th combination, while extracting excitement intervals based on the cumulative value of viewer action points, interval extraction based on text information and interval estimation by a machine learning model are used in combination, and concept video data with added video expression is generated by a generative AI. In the 40th combination, the embodiment is combined with the third, fifth, and sixth modified examples. In the 40th combination, while extracting excitement intervals based on the cumulative value of viewer action points, intervals are extracted based on text information, and concept video data with added video expression is generated by a generative AI in an order based on the specification information. In the 41st combination, the embodiment is combined with the 4th, 5th, and 6th modified examples. In the 41st combination, while extracting the excitement intervals based on the cumulative value of viewer action points, the intervals are estimated by a machine learning model, and concept video data with added video representations is generated by a generation AI in the order based on the specification information.
[0141] Next, we will describe combinations of the embodiment and four variations. In the 42nd combination, the embodiment is combined with the first, second, third, and fourth variations. In the 42nd combination, URLs are used as concept video information while using a combination of extraction of exciting intervals based on the cumulative value of viewer action points and volume, interval extraction based on text information, and interval estimation by a machine learning model. In the 43rd combination, the embodiment is combined with the first, second, third, and fifth variations. In the 43rd combination, URLs are used as concept video information while using a combination of extraction of exciting intervals based on the cumulative value of viewer action points and volume, and interval extraction based on text information, and combining intervals in an order based on specification information. In the 44th combination, the embodiment is combined with the first, second, third, and sixth variations. In the 44th combination, the extraction of exciting intervals based on the cumulative value of viewer action points and volume, along with interval extraction based on text information, is used in combination to generate concept video data with added video expression by a generating AI, while using a URL as the concept video information. In the 45th combination, the embodiment is combined with the first, second, fourth, and fifth variations. In the 45th combination, the extraction of exciting intervals based on the cumulative value of viewer action points and volume, along with interval estimation by a machine learning model, is used in combination to combine intervals in an order based on specification information, while using a URL as the concept video information. In the 46th combination, the embodiment is combined with the first, second, fourth, and sixth variations. In the 46th combination, the extraction of exciting intervals based on the cumulative value of viewer action points and volume, along with interval estimation by a machine learning model, is used in combination to generate concept video data with added video expression by a generating AI, while using a URL as the concept video information. In the 47th combination, the embodiment, the first modification, the second modification, the fifth modification, and the sixth modification are combined. In the 47th combination, the excitement intervals are extracted based on the cumulative value of viewer action points and volume, and concept video data with added video expression is generated by a generating AI in the order based on the specification information, while using a URL as the concept video information.In the 48th combination, the embodiment is combined with the first, third, fourth, and fifth variations. In the 48th combination, while extracting excitement intervals based on the cumulative value of viewer action points, interval extraction based on text information and interval estimation by a machine learning model are used in combination, and intervals are joined in an order based on specification information, while using a URL as concept video information. In the 49th combination, the embodiment is combined with the first, third, fourth, and sixth variations. In the 49th combination, while extracting excitement intervals based on the cumulative value of viewer action points, interval extraction based on text information and interval estimation by a machine learning model are used in combination, and concept video data with added video expression by a generating AI is generated, while using a URL as concept video information. In the 50th combination, the embodiment is combined with the first, third, fifth, and sixth variations. In the 50th combination, the excitement intervals are extracted based on the cumulative value of viewer action points, intervals are extracted based on text information, and concept video data with added video expression is generated by a generating AI in an order based on the specification information, while using a URL as the concept video information. In the 51st combination, the embodiment is combined with the first, fourth, fifth, and sixth variations. In the 51st combination, the excitement intervals are extracted based on the cumulative value of viewer action points, intervals are estimated by a machine learning model, and concept video data with added video expression is generated by a generating AI in an order based on the specification information, while using a URL as the concept video information. In the 52nd combination, the embodiment is combined with the second, third, fourth, and fifth variations. In the 52nd combination, the extraction of excitement intervals based on the cumulative value of viewer action points and volume, interval extraction based on text information, and interval estimation by a machine learning model are used in combination, and the intervals are combined in an order based on the specification information. In the 53rd combination, the embodiment, the second modification, the third modification, the fourth modification, and the sixth modification are combined.In the 53rd combination, conceptual video data with added video expression is generated by a generating AI, using a combination of extraction of exciting intervals based on the cumulative value of viewer action points and volume, interval extraction based on text information, and interval estimation by a machine learning model. In the 54th combination, the embodiment is combined with the second, third, fifth, and sixth variations. In the 54th combination, conceptual video data with added video expression is generated by a generating AI in the order specified by the specification information, using a combination of extraction of exciting intervals based on the cumulative value of viewer action points and volume, and interval extraction based on text information. In the 55th combination, the embodiment is combined with the second, fourth, fifth, and sixth variations. In the 55th combination, conceptual video data with added video expression is generated by a generating AI in the order specified by the specification information, using a combination of extraction of exciting intervals based on the cumulative value of viewer action points and volume, and interval estimation by a machine learning model. In the 56th combination, the embodiment, the third modification, the fourth modification, the fifth modification, and the sixth modification are combined. In the 56th combination, while extracting the excitement intervals based on the cumulative value of viewer action points, interval extraction based on text information and interval estimation by a machine learning model are used in combination, and concept video data with added video representations is generated by a generating AI in the order based on the specification information.
[0142] Next, the combinations of the embodiment and the five modified examples will be described. In the 57th combination, the embodiment is combined with the first, second, third, fourth, and fifth modified examples. In the 57th combination, the extraction of excitement intervals based on the cumulative value of viewer action points and volume, interval extraction based on text information, and interval estimation by a machine learning model are used in combination, and the intervals are joined in an order based on the specification information, while using a URL as the concept video information. In the 58th combination, the embodiment is combined with the first, second, third, fourth, and sixth modified examples. In the 58th combination, the extraction of excitement intervals based on the cumulative value of viewer action points and volume, interval extraction based on text information, and interval estimation by a machine learning model are used in combination, and concept video data with added video expression by a generating AI is generated, while using a URL as the concept video information. In the 59th combination, the embodiment is combined with the first, second, third, fifth, and sixth modified examples. In combination 59, the extraction of exciting intervals based on the cumulative value of viewer action points and volume, along with interval extraction based on text information, is used in combination. Concept video data with added video expression is generated by a generating AI in an order based on specification information, while using a URL as the concept video information. In combination 60, the embodiment is combined with the first, second, fourth, fifth, and sixth variations. In combination 60, the extraction of exciting intervals based on the cumulative value of viewer action points and volume, along with interval estimation by a machine learning model, is used in combination. Concept video data with added video expression is generated by a generating AI in an order based on specification information, while using a URL as the concept video information. In combination 61, the embodiment is combined with the first, third, fourth, fifth, and sixth variations. In the 61st combination, while extracting the peak intervals based on the cumulative value of viewer action points, interval extraction based on text information and interval estimation using a machine learning model are used in combination, and concept video data with added video representations is generated by a generation AI in the order based on the specification information, while using URLs as concept video information.In the 62nd combination, the embodiment, the second modification, the third modification, the fourth modification, the fifth modification, and the sixth modification are combined. In the 62nd combination, the cumulative value of viewer action points and the extraction of excitement intervals based on volume, interval extraction based on text information, and interval estimation by a machine learning model are used in combination to generate concept video data with added video expression by a generating AI in the order based on the specification information.
[0143] Finally, the combination of the embodiment and all of the modifications will be described. In the 63rd combination, the embodiment is combined with all of the first to sixth modifications. In the 63rd combination, the cumulative value of viewer action points and the extraction of excitement intervals based on volume, interval extraction based on text information, and interval estimation by a machine learning model are used in combination to generate concept video data with video representations added by a generating AI in the order based on the specification information, while using URLs as concept video information.
[0144] The first to sixth combinations described above are combinations of embodiments and modified examples, but modified examples can also be combined without going through embodiments. For example, by combining the first and sixth modified examples, it is possible to generate concept video data with added video representation by a generating AI, while using a URL as concept video information. Furthermore, by combining the third, fourth, and sixth modified examples, it is possible to generate concept video data with added video representation by a generating AI by using both interval extraction based on text information and interval estimation by a machine learning model. In this way, the first to sixth modified examples can be combined arbitrarily without going through embodiments, as long as they do not contradict each other.
[0145] In the above embodiment, the specification information acquisition means 128, the text information generation means 130, and the parameter information generation means 132 are not essential components. That is, the control circuit 112 of the server 110 only needs to function as at least the video data acquisition means 120, the interval information acquisition means 122, the concept video information acquisition means 124, and the output means 126. The acquisition of specification information, the generation of text information, the generation of parameter information, the addition of video representation by the generation AI server 210, and the interval estimation by the machine learning model server 310 may be adopted as needed.
[0146] The broadcaster may also input approval or rejection of the concept video data by operating the operation unit 26 of the broadcaster terminal 10. If the broadcaster inputs rejection, the control circuit 12 of the broadcaster terminal 10 sends rejection information to the server 110. The rejection information indicates the reason why the user rejected the data. For example, the rejection information indicates that the keywords to be re-extracted, the range of video data to be referenced, the composition pattern of the concept video, the format according to the posting destination, the caption conditions, or the production conditions did not conform to the broadcaster's intentions. When the control circuit 112 of the server 110 receives rejection information, it regenerates the concept video data based on the rejection information by re-extracting multiple section information, changing the combination order of multiple section information, or re-inputting into the generation AI. On the other hand, if the control circuit 12 of the broadcaster terminal 10 receives approval information, the control circuit 112 of the server 110 or the control circuit 12 of the broadcaster terminal 10 may output the approved concept video data.
[0147] Furthermore, concept video information is not limited to video data representing a concept video. Concept video information may include location information for section A of video 1, location information for section B of video 2, and location information for section C of video 3, and may also include instructions to play sections A, B, and C continuously. In this case, the concept video information includes the URLs for videos 1 to 3, the start and end times corresponding to each URL, and instructions to play sections A, B, and C continuously.
[0148] Furthermore, the entity that instructs the generation of concept video data is not limited to the broadcaster. A third party, such as a fan of the broadcaster, may also instruct the generation of concept video data. For example, a fan may select a target broadcaster, input specification information, and generate concept video data from the broadcaster's previously broadcast video data. In this case, from the standpoint of copyright protection, the control circuit 112 of the server 110 may include a flow in which it sends a notification to the broadcaster terminal 10 requesting the broadcaster's consent, and generates the concept video data after obtaining the broadcaster's consent.
[0149] Furthermore, the streamer ID is not limited to the ID of a single streamer. The number of streamer IDs can be one or more, and may be multiple. In other words, systems 1,1a to 1f may generate concept video data from video data of multiple streamers. For example, if multiple streamers conduct a collaborative stream, systems 1,1a to 1f may extract characteristic sections from the video data of each streamer and generate a digest video data of the collaboration.
[0150] Furthermore, all or part of the processing described as being performed by server 110 may be performed by the broadcaster terminal 10. In this case, the output processing is that the broadcaster terminal 10 displays the concept video on the display 20 using the concept video data.
[0151] The functions of the generation AI server 210 and the machine learning model server 310 may be integrated into server 110. Furthermore, the generation AI server 210 and the machine learning model server 310 may be external servers provided by a different operator than the operator of server 110.
[0152] Furthermore, parameters related to streaming are not limited to data collected during the stream. Parameters related to streaming may also include data such as comments, retweets, and quotes on social media after the stream. Additionally, parameters related to streaming may include the number of concurrent viewers and the number of game events that occurred.
[0153] Furthermore, systems 1,1a to 1f may be inserted between the sections from which transition videos generated by the generation AI have been extracted.
[0154] Furthermore, the editing process of the concept video (cutting out unnecessary parts, automatically adding subtitles, converting aspect ratios, adding logos and text overlays, etc.) may be performed by an algorithm or by a generative AI such as a large-scale language model (LLM).
[0155] The server 110 may include multiple information processing devices.
[0156] Note that the broadcaster's device 10 is not limited to a PC. The broadcaster's device 10 may be a smartphone, tablet, or game console.
[0157] Furthermore, at least a portion of the processing described herein as being performed by a specific device may be performed by any information processing device. In addition, the series of processing performed by each device described herein may be implemented using software, hardware, or a combination of software and hardware.
[0158] Furthermore, the control circuit 112 (video data acquisition means 120) of the server 110 may narrow down the target of acquisition of multiple video data based on the specification information. For example, if the specification information indicates "popular scene compilation," the control circuit 112 (video data acquisition means 120) of the server 110 may prioritize the acquisition of video data with a high number of views. In this case, the memory circuit 114 of the server 110 stores information for the popular scene compilation video template indicating that it will acquire video data with a high number of views among the videos indicated by each video data. Also, if the specification information indicates "niche scenes," the control circuit 112 (video data acquisition means 120) of the server 110 may acquire video data with a low number of views.
[0159] The control circuit 112 (concept video information acquisition means 124) of the server 110 may acquire concept video data having a format for posting to SNS as concept video information. The format for posting to SNS is a format that corresponds to the SNS to which it is posted, such as a vertical short video format or a horizontal standard format. This format includes not only the video format (vertical / horizontal, etc.) but also information that needs to be attached. Specifically, it may include attached information that differs for each SNS, such as the video title, thumbnail image, poster information (distributor name, profile, etc.), and description (caption). In this case, the control circuit 112 of the server 110 transmits the concept video data having a format for posting to SNS to the distributor terminal 10 via the network interface 116. The concept video information may also indicate the upload destination of the concept video on SNS. That is, if the distribution platform or SNS upload function is provided, link information (URL, etc.) that identifies the upload destination may be output as concept video information.
[0160] The server 110 or the broadcaster terminal 10 may output the concept video data approved by the broadcaster as video data in a format suitable for posting on social media.
[0161] Furthermore, the server 110 or the distributor terminal 10 may output link information indicating the upload destination on the distribution platform or SNS, corresponding to the concept video approved by the distributor.
[0162] Furthermore, the control circuit 112 of the server 110 may extract sections that include self-introduction phrases, fan names, or catchphrases that the broadcaster repeatedly uses, as standard sections.
[0163] The metadata may include not only the delivery date and time, title, and category, but also tags, delivery time slots, etc. The control circuit 112 of the server 110 may use this metadata to narrow down the target video data or segment information according to the specification information.
[0164] The parameter information may also include parameters that depend on the frequency of laughter obtained by analyzing the audio contained in the video data. The control circuit 112 of the server 110 may extract intervals in which the parameters exceed a threshold. [Explanation of Symbols]
[0165] 1: System 10: Streamer's device 12: Control circuits 14:Memory circuit 16: Network Interface 17: Camera 18: Graphics Processing Circuit 20: Display 22: Audio Processing Circuit 24: Speaker 26:Operation section 110: Server 112: Control circuits 114:Memory circuit 116: Network Interface 120: Method for acquiring video data 122: Means for obtaining section information 124: Method for obtaining concept video information 126: Output means 128: Means for obtaining specification information 130: Text information generation means 132: Parameter information generation means 210: Generation AI Server 310: Machine Learning Model Server 410: Video Server 412: Control circuit 414:Memory circuit 414a: Video Database 416: Network Interface
Claims
1. A system comprising one or more information processing devices, The control circuit of the one or more information processing devices described above functions as a video data acquisition means, a section information acquisition means, a concept video information acquisition means, and an output means. The aforementioned video data acquisition means acquires multiple video data associated with one or more broadcaster IDs, The section information acquisition means acquires a plurality of section pieces related to a plurality of sections used for generating a concept video in the plurality of video data, The concept video information acquisition means uses the plurality of section information to acquire concept video information used when viewing the concept video which includes the plurality of sections, The output means outputs the concept video information to the terminal. system.
2. The control circuit of the one or more information processing devices described above functions as a means for acquiring specification information. The specification information acquisition means acquires specification information indicating the generation conditions for the concept video, The section information acquisition means acquires the plurality of section information using the specification information. The system according to claim 1.
3. The aforementioned specification information is entered by the user of the terminal. The system according to claim 2.
4. The memory circuit of the one or more information processing devices stores multiple templates for the concept video. The aforementioned specification information indicates the template selected by the user of the terminal from among the multiple templates. The system according to claim 2.
5. The video data acquisition means acquires multiple video data associated with one or more distributor IDs based on the specification information. The system according to any one of claims 2 to 4.
6. The specification information includes at least one of the following: the intended use of the concept video, the theme or keywords of the concept video, the structure of the concept video, the length of the concept video, or information indicating the video data to be referenced. The system according to any one of claims 2 to 4.
7. The aforementioned multiple sections are: sections in the multiple videos shown by the multiple video data in which the performers or one or more viewers become excited; sections in the multiple videos shown by the multiple video data in which the content related to a specific topic is included; or sections in the multiple videos shown by the multiple video data in which the content is repetitive. The system according to any one of claims 1 to 4.
8. The control circuit of the one or more information processing devices functions as a parameter information generation means. The parameter information generation means generates a plurality of parameter information indicating parameters related to the distribution of the plurality of video data, The interval information acquisition means acquires the plurality of interval information using the plurality of parameter information. The system according to any one of claims 1 to 4.
9. The parameter information includes at least one parameter that depends on the volume in the video indicated by the video data, the number of comments or stamps on the video indicated by the video data, the number of viewers of the video indicated by the video data, or the number of actions that occurred on the SNS (Social Network Service) after the video indicated by the video data was distributed on the SNS. The system according to claim 8.
10. The control circuit of the one or more information processing devices functions as a text information generation means. The text information generation means generates multiple text information by converting the audio contained in the multiple videos represented by the multiple video data into text, The interval information acquisition means acquires the plurality of interval information using the plurality of text information. The system according to any one of claims 1 to 4.
11. The section information acquisition means uses the plurality of text information to acquire a plurality of section pieces related to a fixed section containing fixed content in a video shown by the plurality of video data. The system according to claim 10.
12. The control circuit of the one or more information processing devices described above functions as a means for acquiring specification information. The specification information acquisition means acquires specification information including the keywords of the concept video, The interval information acquisition means acquires the plurality of interval information using the keywords included in the specification information and the text indicated by the plurality of text information. The system according to claim 10.
13. The control circuit of the one or more information processing devices described above functions as a means for acquiring specification information. The specification information acquisition means acquires specification information including the keywords of the concept video, Each of the aforementioned video data contains metadata, The video data acquisition means acquires the plurality of video data using the specification information and the plurality of metadata. The system according to any one of claims 1 to 4.
14. The interval information acquisition means acquires the plurality of interval information using a machine learning model. The system according to any one of claims 1 to 4.
15. Each of the aforementioned interval information includes information about the interval that the machine learning model has determined to have a high degree of attention. The system according to claim 14.
16. The concept video information acquisition means acquires concept video data representing the concept video formed by combining the multiple sections as the concept video information. The system according to any one of claims 1 to 4.
17. The control circuit of the one or more information processing devices described above functions as a means for acquiring specification information. The specification information acquisition means acquires specification information indicating the generation conditions for the concept video, The concept video information acquisition means acquires concept video data as concept video information, which shows a concept video in which the plurality of sections are combined in an order determined based on the specification information. The system according to claim 16.
18. The concept video information acquisition means acquires the concept video information by inputting the multiple section information into the generating AI. The system according to claim 16.
19. The aforementioned multiple section information is video data showing the video of the aforementioned multiple sections. The system according to claim 18.
20. The concept video information acquisition means acquires the concept video information in which video representations have been added by the generation AI to the videos of the plurality of sections. The system according to claim 19.
21. The aforementioned concept video information acquisition means acquires the concept video data having a format for posting to social media as the concept video information. The system according to claim 16.
22. The aforementioned concept video information indicates the upload location of the concept video on the distribution platform or social media. The system according to any one of claims 1 to 4.
23. Information processing method, In the information processing method, the control circuit of one or more information processing devices functions as a means for acquiring video data, a means for acquiring section information, a means for acquiring concept video information, and an output means. The aforementioned video data acquisition means acquires multiple video data associated with one or more broadcaster IDs, The section information acquisition means acquires a plurality of section pieces of information related to the section used to generate the concept video in the plurality of video data, The concept video information acquisition means uses the plurality of section information to acquire concept video information used when viewing the concept video which includes the plurality of sections, The output means outputs the concept video information to the terminal. Information processing methods.
24. It is a program, The program causes the control circuits of one or more information processing devices to function as video data acquisition means, section information acquisition means, concept video information acquisition means, and output means. The aforementioned video data acquisition means acquires multiple video data associated with one or more broadcaster IDs, The section information acquisition means acquires a plurality of section pieces of information related to the section used to generate the concept video in the plurality of video data, The concept video information acquisition means uses the plurality of section information to acquire concept video information used when viewing the concept video which includes the plurality of sections, The output means outputs the concept video information to the terminal. program.
Citation Information
Patent Citations
System and data manufacturing method
JP2025147164A
Server and computer program
JP2026064188A