Moving image generation device

WO2026167865A1PCT designated stage Publication Date: 2026-08-13NTT DOCOMO INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2026-08-13

Smart Images

  • Figure JP2025004335_13082026_PF_FP_ABST
    Figure JP2025004335_13082026_PF_FP_ABST
Patent Text Reader

Abstract

This moving image generation device comprises: a division unit that, when the time length of first moving image data exceeds a first time length, divides the first moving image data into moving image data having the first time length or less to output a plurality of pieces of divided moving image data including first divided moving image data; a determination unit that, on the basis of the first divided moving image data, determines a plurality of pieces of speech data corresponding one-to-one to a plurality of speeches included in the first divided moving image data, each of the plurality of pieces of speech data indicating at least a corresponding speech by text; and an information acquisition unit that acquires information relating to second moving image data output from an image language model by inputting, to the image language model, a prompt including the first divided moving image data, a first text data group including speech data corresponding to the first divided moving image data, and a second text data group.
Need to check novelty before this filing date? Find Prior Art

Description

Video generation device

[0001] The present invention relates to a video generation device.

[0002] Conventionally, there has been a technique for extracting a portion suitable for a short video from a video using a multimodal large language model (LLM) capable of analyzing the content of the video.

[0003] For example, the video generation device disclosed in Patent Document 1 includes an acquisition unit, an analysis unit, an editing plan unit, and a video editing unit. The acquisition unit acquires a video file. The analysis unit acquires an analysis result including time-series speech information, which is text information obtained by writing down the speech included in the video file in time series. The editing plan unit transmits the time-series speech information and command information for instructing to output an editing plan for editing the time-series speech information according to a desired editing policy information to a dialogue-type AI using a language model, and receives the editing plan from the dialogue-type AI. The video editing unit edits the video file according to the editing plan to generate an edited video.

[0004] Japanese Patent No. 7538574

[0005] When analyzing the content of a video using a multimodal LLM to extract a short video, since there is a limit to the length of the video that can be input to the multimodal LLM, it is impossible to analyze a long video.

[0006] An object of the present disclosure is to provide a video generation device capable of analyzing a long video as compared with the prior art.

[0007] The video generation device according to this disclosure includes: a splitting unit that outputs a plurality of split video data including a first split video data by splitting the first video data into video data of a length equal to or less than the first split video data when the length of the first video data exceeds a first length; a determination unit that determines a plurality of dialogue data that correspond one-to-one with a plurality of lines of dialogue included in the first split video data based on the first split video data, and each of the plurality of dialogue data indicates at least the corresponding lines of dialogue in text; and an information acquisition unit that acquires information about a second video data output from the image language model by inputting a prompt to an image language model that includes the first split video data, a first group of text data including the first split video data, and a second group of text data, wherein the second group of text data includes one or more lines of dialogue data from among the plurality of lines of dialogue data corresponding to the first video data that do not overlap with the lines of dialogue included in the first group of text data.

[0008] According to this disclosure, it becomes possible to analyze longer video durations compared to conventional technologies.

[0009] A block diagram showing an example of the overall configuration of video generation system 1A. A block diagram showing an example of the configuration of video generation device 10A. A diagram showing an example of a text data group TG. A diagram showing an example of a prompt PP1. A diagram showing an example of video information MI. A diagram showing an example of the flow of various data and information used in video generation device 10A. A flowchart showing an example of the operation of video generation device 10A. A block diagram showing an example of the configuration of video generation device 10B. A diagram showing an example of the flow of various data and information used in video generation device 10B. A flowchart showing an example of the operation of video generation device 10B. A flowchart showing an example of the operation of video generation device 10B. A block diagram showing an example of the overall configuration of video generation system 1C according to modification 1.

[0010] 1: The video generation system 1A according to the first embodiment will be described below with reference to Figures 1 to 7.

[0011] 1-1: Configuration of the First Embodiment 1-1-1: Overall Configuration Diagram 1 is a block diagram showing an example of the overall configuration of the video generation system 1A according to this embodiment. As shown in Figure 1, the video generation system 1A comprises a video generation device 10A, an audio analysis server 20, and a generation server 30. The video generation device 10A, the audio analysis server 20, and the generation server 30 are connected to each other via a communication network NET so that they can communicate with one another.

[0012] Note that the fact that Figure 1 shows only one video generation device 10A, one audio analysis server 20, and one generation server 30 is merely an example.

[0013] The video generation device 10A is a device for user U[1] to obtain information about second video data MD2 based on first video data MD1 and to generate second video data MD2 using said information. Second video data MD2 is, for example, video data extracted from a part of first video data MD1. The video generation device 10A obtains information about second video data MD2 from the generation server 30 by inputting a prompt PP1, which includes first video data MD1 and a plurality of dialogue data LDs corresponding to a plurality of lines of dialogue contained in first video data MD1, to the image language model PLM stored in the generation server 30. As described above, if second video data MD2 is video data extracted from a part of first video data MD1, the prompt PP1 includes an instruction to extract a part of first video data MD1. This instruction is an example of a "first instruction".

[0014] The audio analysis server 20 receives the first video data MD1 from the video generation device 10A and transcribes the multiple lines of dialogue contained in the first video data MD1 to determine multiple first text data TD1 that correspond one-to-one with the lines of dialogue. The determined multiple first text data TD1 are transmitted from the audio analysis server 20 to the video generation device 10A.

[0015] The generation server 30 receives the above prompt PP1 from the video generation device 10A. The generation server 30 also stores an image language model PLM. One example of the image language model PLM is Gemini (registered trademark). The generation server 30 inputs the above prompt PP1 received from the video generation device 10A into the image language model PLM. The image language model PLM outputs information regarding the second video data MD2. One example of this information is information indicating the location from which the second video data MD2 is extracted from the first video data MD1.

[0016] 1-1-2: Diagram 2 of the video generation device configuration is a block diagram showing an example configuration of the video generation device 10A. As shown in Figure 2, the video generation device 10A comprises a processing unit 11A, a storage device 12A, a display device 13, an input device 14, and a communication device 15. Each element of the video generation device 10A is interconnected by one or more buses for communicating information.

[0017] The processing unit 11A is a processor that controls the entire video generation device 10A. The processing unit 11A is configured using, for example, one or more chips. The processing unit 11A is configured using, for example, a central processing unit (CPU) that includes interfaces with peripheral devices, an arithmetic unit, and registers. Some or all of the functions of the processing unit 11A may be implemented by hardware such as a DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), PLD (Programmable Logic Device), and FPGA (Field Programmable Gate Array). The processing unit 11A executes various processes in parallel or sequentially.

[0018] The storage device 12A is a recording medium that can be read from and written to by the processing device 11A. The storage device 12A includes, for example, non-volatile memory and volatile memory. Non-volatile memory is, for example, ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory), and EEPROM (Electrically Erasable Programmable Read Only Memory). Volatile memory is, for example, RAM (Random Access Memory).

[0019] The storage device 12A stores multiple programs, including the control program PR1A, for execution by the processing device 11A. The storage device 12A functions as a work area for the processing device 11A. Furthermore, the storage device 12A may store, for example, the prompt PP1 described later.

[0020] The display device 13 is a device that displays images and text information. The display device 13 displays various images under the control of the processing device 11A. For example, various display panels such as liquid crystal display panels and organic EL display panels are preferably used as the display device 13.

[0021] The input device 14 is a device that accepts operations from user U[1]. For example, the input device 14 is configured to include a keyboard, touchpad, touch panel, or pointing device such as a mouse. Here, if the input device 14 is configured to include a touch panel, it may also serve as the display device 13.

[0022] The communication device 15 is hardware that acts as a transmitting and receiving device for communicating with other devices. The communication device 15 is also called, for example, a network device, a network controller, a network card, or a communication module. The communication device 15 may be equipped with a connector for wired connection and an interface circuit corresponding to the connector. The communication device 15 may also be equipped with a wireless communication interface. Examples of connectors and interface circuits for wired connection include products compliant with wired LAN, IEEE1394, and USB. Examples of wireless communication interfaces include products compliant with wireless LAN and Bluetooth®.

[0023] The processing unit 11A functions as a communication control unit 111, a video acquisition unit 112, a text data acquisition unit 113A, a determination unit 114A, a prompt acquisition unit 115A, an input unit 116, an information acquisition unit 117, a display control unit 118, and a generation unit 119, for example, by reading and executing the control program PR1A from the storage device 12A.

[0024] The communication control unit 111 causes the communication device 15 to send and receive various information, various data, and various signals between it and the voice analysis server 20 and the generation server 30. The communication control unit 111 may also cause the communication device 15 to send and receive various information, various data, and various signals between it and other devices on the communication network NET besides the voice analysis server 20 and the generation server 30.

[0025] The video acquisition unit 112 acquires the first video data MD1. The first video data MD1 may be video data input by user U[1] to the video generation device 10A using the input device 14. Alternatively, the first video data MD1 may be video data acquired from a device (not shown) located on the communication network NET.

[0026] The first video data MD1 acquired by the video acquisition unit 112 is transmitted to the voice analysis server 20 by the communication control unit 111. As described above, the voice analysis server 20 transcribes the multiple lines of dialogue contained in the first video data MD1 and determines multiple first text data TD1 that correspond one-to-one with the multiple lines of dialogue. The determined multiple first text data TD1 are transmitted from the voice analysis server 20 to the video generation device 10A.

[0027] The text data acquisition unit 113A acquires multiple first text data TD1 transmitted from the speech analysis server 20.

[0028] The determination unit 114A determines a plurality of dialogue data LDs that correspond one-to-one to a plurality of dialogues by adding identification data DD, which identifies the dialogue, to each of the plurality of first text data TD1 acquired by the text data acquisition unit 113A. Furthermore, the determination unit 114A determines a text data group TG that includes the plurality of first text data TD1 and the plurality of dialogue data LDs.

[0029] Figure 3 shows an example of a text data set TG.

[0030] The text data group TG has five items: "start", "end", "period", "text", and "lim_text". The "start" item stores data indicating the point in time when a line of dialogue begins to be spoken in the first video data MD1. The "end" item stores data indicating the point in time when the line of dialogue ends in the first video data MD1. The "period" item stores data indicating the duration of the line of dialogue in the first video data MD1. The "text" item stores data showing the transcription result of the line of dialogue. The first text data TD1 refers to the data showing the transcription result of the line of dialogue. The "lim-text" item stores dialogue data LD, which is the first text data TD1 stored in the "text" item with identification data DD added to identify the line of dialogue. In the example shown in Figure 3, the identification data DD is the sequential number in which the dialogue corresponding to the first text data TD1 occurs. The dialogue data LD is data in which the identification data DD that identifies the dialogue is attached to each of the multiple first text data TD1, and there is a one-to-one correspondence between multiple lines of dialogue.

[0031] In Figure 2, the prompt acquisition unit 115A acquires a prompt PP1 which includes the above-mentioned multiple dialogue data LDs and the first video data MD1.

[0032] Figure 4 shows an example of prompt PP1. Prompt PP1 includes the first block BL1 to the fourth block BL4.

[0033] The first block BL1 includes an outline of instructions to the image language model PLM stored in the generation server 30. In the example shown in Figure 4, the first block BL1 includes instructions to extract multiple short videos that are part of the first video data MD1 from the video content corresponding to the first video data MD1.

[0034] Block 2, BL2, contains a transcription data file that includes a text data group TG, and a description of the text data group TG. As described above, the text data group TG includes multiple dialogue data LDs.

[0035] The third block, BL3, includes a video file corresponding to the first video data MD1 and a description of the video content corresponding to that video file. In the example shown in Figure 4, the description includes the title of the video content, "What's the point of taking love so seriously? #1".

[0036] The fourth block, BL4, contains details of instructions for the image language model PLM stored in the generation server 30. In the example shown in Figure 4, the details of these instructions include instructions for analyzing the video content and dialogue content included in the video content. The details of these instructions also include extracting multiple short videos appropriate to the video content evenly from the entire video content. The multiple short video data SDs representing these multiple short videos include the second video data MD2. The parts extracted from the first video data MD1 differ for each of the multiple short video data SDs. The details of these instructions also include the length of each short video being equal to five consecutive lines of dialogue data LD included in the "lim_text" item of the text data group TG. The details of these instructions also include outputting the short videos with the reason for extraction and a title. Finally, the details of these instructions include outputting a summary of the contents of the three short videos. In other words, the details of these instructions include outputting a summary of the second video data MD2.

[0037] In the above prompt PP1, the instruction indicating the extraction of a portion of the first video data MD1 is an example of a "first instruction". In the said "first instruction", the number of dialogue data LDs corresponding to the length of each short video is an example of a "first number". Also, in the above prompt PP1, the instruction indicating the output of the reason for extracting the second video data MD2 from the first video data MD1 is an example of a "second instruction". Also, in the above prompt PP1, the instruction indicating the output of the title of the second video data MD2 is an example of a "third instruction". Also, in the above prompt PP1, the instruction indicating the output of the summary of the second video data MD2 is an example of a "fourth instruction". Also, in the above prompt PP1, the instruction indicating the output of the reason for extracting the second video data MD2 from the first video data MD1, the title of the second video data MD2, and the summary of the second video data MD2 is an example of a "fifth instruction".

[0038] In Figure 2, the prompt acquisition unit 115A may acquire the prompt PP1 entered by user U[1] using the input device 14. Alternatively, the prompt acquisition unit 115A may acquire the prompt PP1 stored in the storage device 12A.

[0039] The input unit 116 inputs the prompt PP1 acquired by the prompt acquisition unit 115A to the image language model PLM stored in the generation server 30.

[0040] The information acquisition unit 117 acquires video information MI, which is information related to the second video data MD2 output from the image language model PLM, when the input unit 116 inputs the prompt PP1 to the image language model PLM.

[0041] Figure 5 shows an example of video information MI. The video information MI includes Block 5 BL5 to Block 8 BL8.

[0042] The fifth block BL5 includes the dialogue data LD included in the first short video, the reason for extracting the first short video from the first video data MD1, and information regarding the title of the first short video. Note that the dialogue data LD included in the information indicates the location where the first short video is extracted from the first video data MD1.

[0043] The sixth block BL6 includes the dialogue data LD included in the second short video, the reason for extracting the second short video from the first video data MD1, and information regarding the title of the second short video. Note that the dialogue data LD included in the information indicates the location where the second short video is extracted from the first video data MD1.

[0044] The seventh block BL7 includes the dialogue data LD included in the third short video, the reason for extracting the third short video from the first video data MD1, and information regarding the title of the third short video. Note that the dialogue data LD included in the information indicates the location where the third short video is extracted from the first video data MD1.

[0045] The eighth block BL8 indicates the title of the video content corresponding to the first video data MD1 and information regarding the overall outline of the first to third short videos.

[0046] Note that each of the first to third short videos corresponds to the second video data MD%.

[0047] In FIG. 2, the display control unit 118 causes the display device 13 to display the video information MI acquired by the information acquisition unit 117. As an example, when the information acquisition unit 117 acquires, as the video information MI, the reason for extracting the second video data MD2 from the first video data MD1 output from the image language model PLM, the title of the second video data MD2, and the outline of the second video data MD2, the display control unit 118 causes the display device 13 to display the reason for extracting the second video data MD2 from the first video data MD1 output from the image language model PLM, the title of the second video data MD2, and the outline of the second video data MD2.

[0048] The generation unit 119 generates short video data SD based on the video information MI acquired by the information acquisition unit 117. As an example, when the user U[1] manually generates a short video based on the video information MI displayed on the display device 13, the generation unit 119 generates short video data SD based on the operation information indicating the operation content of the user U[1] input to the input device 14. As another example, the generation unit 119 itself may generate short video data SD based on the video information MI acquired by the information acquisition unit 117.

[0049] FIG. 6 is a diagram showing an example of the flow of various data and information used in the video generation device 10A.

[0050] As described above while referring to FIG. 2, the input unit 116 inputs the prompt PP1 to the image language model PLM stored in the generation server 30. The prompt PP1 includes a text data group TG and first video data MD1. The text data group TG includes a plurality of line data LD. The plurality of line data LD correspond one-to-one with a plurality of lines. Also, the plurality of line data LD are determined by adding identification data DD for identifying the lines to each of the plurality of first text data TD1 obtained by transcribing the plurality of lines included in the first video data MD1. In FIG. 6, for simplicity of explanation, the first video data MD1, the text data group TG, and the prompt PP1 are depicted individually, but as described above while referring to FIG. 4, the prompt PP1 includes the first video data MD1 and the text data group TG.

[0051] The information acquisition unit 117 acquires video information MI output from the image language model PLM. The generation unit 119 generates first short video data SD[1], second short video data SD[2], and third short video data SD[3] based on the video information MI. The first short video data SD[1] corresponds to the second video data MD2[1]. The second short video data SD[2] corresponds to the second video data MD2[2]. The third short video data SD[3] corresponds to the second video data MD2[3].

[0052] In the example shown in Figure 6, the third short video data SD[3] is a video extracted from the section between 40 minutes 15 seconds and 40 minutes 35 seconds of the first video data MD1. If the first video data MD1 is long and identification data DD is not attached to each of the first text data TD1, the timestamp in the third short video data SD[3] may be inaccurate. Alternatively, in this case, the generation unit 119 may extract a section from the first video data MD1 that is shifted from the section indicated by the video information MI as the third short video data SD[3]. In this embodiment, since identification data DD indicating the order in which the dialogue occurs is attached to each of the first text data TD1 as identification data DD to identify the dialogue, the timestamp in the third short video data SD[3] is accurate. Furthermore, the generation unit 119 can more accurately extract the portion indicated by the video information MI from the first video data MD1 as the third short video data SD[3].

[0053] In this embodiment, since identification data DD is attached to each of the first text data TD1, the instruction statement indicating the extraction of a portion of the first video data MD1 in prompt PP1 is shortened, and the processing unit 11A can efficiently perform the video generation process. Furthermore, by attaching identification data DD to each of the first text data TD1, the processing unit 11A can more easily recognize that each line of dialogue is different from the others and the order of the lines of dialogue. As a result, the processing load on the processing unit 11A is reduced, and short video data SD can be generated with limited resources. Consequently, as described above, the timestamp of the short video data SD becomes more accurate, and the video generation device 10A can extract the short video data SD more accurately.

[0054] 1-2: Diagram 7 of the operation of the first embodiment is a flowchart showing an example of the operation of the video generation device 10A.

[0055] In step S1, the processing unit 11A provided in the video generation device 10A functions as a video acquisition unit 112. The processing unit 11A acquires the first video data MD1. The processing unit 11A also functions as a communication control unit 111. The processing unit 11A transmits the first video data MD1 to the audio analysis server 20.

[0056] In step S2, the processing unit 11A functions as a text data acquisition unit 113A. The processing unit 11A acquires a plurality of first text data TD1 transmitted from the speech analysis server 20.

[0057] In step S3, the processing unit 11A functions as a determination unit 114A. The processing unit 11A determines a plurality of dialogue data LDs that correspond one-to-one to a plurality of dialogues by adding identification data DDs that identify dialogues to each of the plurality of first text data TD1s acquired in step S2. Furthermore, the processing unit 11A determines a text data group TG that includes the plurality of first text data TD1s and the plurality of dialogue data LDs.

[0058] In step S4, the processing unit 11A functions as a prompt acquisition unit 115A. The processing unit 11A acquires a prompt PP1 which includes a plurality of dialogue data LDs and a first video data MD1.

[0059] In step S5, the processing unit 11A functions as an input unit 116. The processing unit 11A inputs the prompt PP1 acquired in step S4 to the image language model PLM stored in the generation server 30.

[0060] In step S6, the processing unit 11A functions as an information acquisition unit 117. The processing unit 11A acquires video information MI, which is information relating to the second video data MD2 output from the image language model PLM.

[0061] In step S7, the processing unit 11A functions as a display control unit 118. The processing unit 11A displays the video information MI acquired in step S6 on the display device 13.

[0062] In step S8, the processing unit 11A functions as a generation unit 119. The processing unit 11A generates short video data SD based on the video information MI acquired in step S6.

[0063] 1-3: Effects of the First Embodiment The video generation device 10A according to this embodiment comprises a text data acquisition unit 113A, a determination unit 114A, and an information acquisition unit 117. The text data acquisition unit 113A acquires a plurality of first text data TD1 that correspond one-to-one to a plurality of lines of dialogue contained in the first video data MD1, based on the first video data MD1. The determination unit 114A determines a plurality of line of dialogue data LD that correspond one-to-one to a plurality of lines of dialogue by adding identification data DD that identifies the lines of dialogue to each of the plurality of first text data TD1. The information acquisition unit 117 acquires information about the second video data MD2 output from the image language model PLM by inputting a prompt PP1 including the plurality of line of dialogue data LD and the first video data MD1 to the image language model PLM. The identification data DD indicates the order in which the lines of dialogue occur.

[0064] By having the above configuration, the video generation device 10A can obtain more accurate timestamps compared to conventional technology.

[0065] Specifically, in the video generation device 10A, each of the first text data TD1 is accompanied by identification data DD indicating the order in which the dialogue occurs, thus making it possible to make the timestamp of the second video data MD2 accurate.

[0066] Furthermore, in the video generation device 10A, the above-mentioned prompt PP1 includes a first instruction indicating that a portion of the first video data MD1 is to be extracted. The second video data MD2 is video data obtained by extracting a portion of the first video data MD1. The information acquisition unit 117 acquires information regarding the second video data MD2, which indicates the portion from which the second video data MD2 is extracted from the first video data MD1.

[0067] By having the above configuration, the video generation device 10A can more accurately extract the portion indicated by the video information MI from the first video data MD1 as the second video data MD2.

[0068] Furthermore, in the video generation device 10A, the above-mentioned first instruction indicates that a portion corresponding to a first number of consecutive lines of dialogue will be extracted from the first video data MD1.

[0069] The video generation device 10A, with the above configuration, can specify the location from which to extract the second video data MD2 from the first video data MD1 using the number of lines of dialogue.

[0070] Furthermore, in the video generation device 10A, the above-mentioned prompt PP1 includes a second instruction indicating that the reason for extracting the second video data MD2 from the first video data MD1 is to be output. The information acquisition unit 117 acquires the reason output from the image language model PLM as information related to the second video data MD2.

[0071] The video generation device 10A, with the above configuration, can add to the second video data MD2 the reason why the second video data MD2 was extracted from the first video data MD1.

[0072] Furthermore, in the video generation device 10A, the above-mentioned prompt PP1 includes a third instruction indicating that the title of the second video data MD2 should be output. The information acquisition unit 117 acquires the title output from the image language model PLM as information related to the second video data MD2.

[0073] The video generation device 10A, with the above configuration, can add a title to the second video data MD2.

[0074] Furthermore, in the video generation device 10A, the above-mentioned prompt PP1 includes a fourth instruction indicating that an outline of the second video data MD2 should be output. The information acquisition unit 117 acquires an outline of the second video data MD2 output from the image language model PLM as information regarding the second video data MD2.

[0075] The video generation device 10A, with the above configuration, can append an overview of the second video data MD2 to the second video data MD2.

[0076] Furthermore, in the video generation device 10A, the above-mentioned prompt PP1 includes a fourth instruction indicating that the reason for extracting the second video data MD2 from the first video data MD1, the title of the second video data MD2, and an overview of the second video data MD2 will be output. The information acquisition unit 117 acquires the reason for extracting the second video data MD2 from the first video data MD1, the title of the second video data MD2, and an overview of the second video data MD2, which are output from the image language model PLM, as information related to the second video data MD2. The video generation device 10A includes a display control unit 118. The display control unit 118 causes the reason for extracting the second video data MD2 from the first video data MD1, the title of the second video data MD2, and an overview of the second video data MD2, which are output from the image language model PLM, to be displayed on the display device 13.

[0077] By having the above configuration, the video generation device 10A allows user U[1] to recognize the reason for extracting the second video data MD2 from the first video data MD1, the title of the second video data MD2, and an overview of the second video data MD2.

[0078] Furthermore, in the video generation device 10A, the above-mentioned first instruction indicates that information regarding multiple short video data SDs, the portions extracted from the first video data MD1, differ from each other, will be output. The information acquisition unit 117 acquires information regarding multiple short video data SDs output from the image language model PLM as information regarding the second video data MD2. The second video data MD2 is included in the multiple short video data SDs.

[0079] By having the above configuration, the video generation device 10A can more accurately extract the portion indicated by the video information MI from the first video data MD1 as multiple short video data SDs.

[0080] 2. Second Embodiment The following description will explain the video generation system 1B according to the second embodiment with reference to Figures 8 to 11. For the sake of simplicity, the following description will primarily focus on the differences between the video generation system 1B according to the second embodiment and the video generation system 1A according to the first embodiment. Furthermore, for components in the video generation system 1B according to the second embodiment that are identical to those in the video generation system 1A according to the first embodiment, the same reference numerals will be used, and their functions may be omitted.

[0081] 2-1: Configuration of the Second Embodiment 2-1-1: Overall Configuration The video generation system 1B according to this embodiment includes a video generation device 10B instead of a video generation device 10A, compared to the video generation system 1A according to the first embodiment. In other respects, the overall configuration of the video generation system 1B is the same as the overall configuration of the video generation system 1A, and therefore its illustration is omitted.

[0082] 2-1-2: The configuration diagram 8 of the video generation device is a block diagram showing an example configuration of the video generation device 10B. As shown in Figure 8, the video generation device 10B, compared to the video generation device 10A, is equipped with a processing unit 11B instead of a processing unit 11A, and a storage device 12B instead of a storage device 12A.

[0083] The storage device 12B stores the control program PR1B in place of the control program PR1A stored in the storage device 12A.

[0084] The processing unit 11B functions as a communication control unit 111, a video acquisition unit 112, a text data acquisition unit 113B, a determination unit 114B, a prompt acquisition unit 115B, an input unit 116, an information acquisition unit 117, a display control unit 118, a generation unit 119, and a splitting unit 120, for example, by reading and executing the control program PR1B from the storage device 12B.

[0085] If the duration of the first video data MD1 exceeds the first duration, the splitting unit 120 splits the first video data MD1 into video data of duration less than or equal to the first duration, thereby outputting a plurality of split video data DVDs, including the first split video data DVD[1] and the second split video data DVD[2].

[0086] As an example, the splitting unit 120 may use a feature change detection model that has learned feature changes between frames to select a plurality of specific frames from a plurality of frames contained in the first video data MD1, and select a frame that satisfies a pre-set splitting condition from among the plurality of specific frames as the splitting frame for splitting the first video data MD1. In this case, the splitting unit 120 may input one frame extracted from the plurality of frames and at least one other frame positioned before and after that one frame to the feature change detection model, and select the plurality of specific frames based on its output. Furthermore, each time the "one frame extracted from the plurality of frames" is shifted and changed, the splitting unit 120 may obtain an output from the feature change detection model, and if the output from the feature change detection model satisfies a pre-set condition, that one frame may be designated as the specific frame. Furthermore, the splitting unit 120 may use a relationship detection model that has learned the relationships between frames to classify the above-mentioned multiple specific frames into a single frame and an unrelated frame, which is a specific frame that is not related to the single frame, and select the unrelated frame as the splitting frame for splitting the first video data MD1. One example of such an unrelated frame is a frame in which a scene has changed.

[0087] Each of the multiple divided video data VDs, divided by the division unit 120, is transmitted to the audio analysis server 20 by the communication control unit 111. As described above, the audio analysis server 20 transcribes the multiple lines of dialogue contained in each of the divided video data VDs, and determines a plurality of first text data TD1s that correspond one-to-one with the multiple lines of dialogue for each divided video data VD. The determined plurality of first text data TD1s are transmitted from the audio analysis server 20 to the video generation device 10B. Note that when the plurality of divided video data VDs, divided by the division unit 120, are transmitted to the audio analysis server 20, the first video data MD1 is not transmitted to the audio analysis server 20.

[0088] The text data acquisition unit 113B has the same functions as the text data acquisition unit 113A in the first embodiment.

[0089] Furthermore, the text data acquisition unit 113B acquires multiple first text data TD1 transmitted from the audio analysis server 20 for each segmented video data DVD.

[0090] Similar to the determination unit 114A in the first embodiment, the determination unit 114B determines a plurality of dialogue data LDs that correspond one-to-one to a plurality of dialogues by adding identification data DD, which identifies the dialogue, to each of the plurality of first text data TD1 acquired by the text data acquisition unit 113B.

[0091] In other words, the determination unit 114B determines a plurality of dialogue data LD[1] that correspond one-to-one with the plurality of lines of dialogue contained in the first divided video data VD[1], based on the first divided video data VD[1]. Each of the plurality of dialogue data LD[1] indicates at least the corresponding line of dialogue in text.

[0092] Similarly, the determination unit 114B determines a plurality of dialogue data LD[2] that correspond one-to-one to the plurality of lines of dialogue contained in the second divided video data DVD[2], based on the second divided video data DVD[2]. Each of the plurality of dialogue data LD[2] indicates at least the corresponding line of dialogue in text.

[0093] The determination unit 114B determines a first text data group TG[1] which includes a first text data TD1[1] corresponding to the first segmented video data VD[1] and a dialogue data LD[1] corresponding to the first segmented video data VD[1].

[0094] Similarly, the determination unit 114B determines a second text data group TG[2] which includes a first text data TD1[2] corresponding to the second split video data VD[2] and dialogue data LD[2] corresponding to the second split video data VD[2].

[0095] The prompt acquisition unit 115B has the same function as the prompt acquisition unit 115A in the first embodiment.

[0096] Furthermore, the prompt acquisition unit 115B acquires a prompt PP2[1] which includes the first segmented video data VD[1], the first text data group TG[1], and the second text data group TG[2]. Here, the second text data group TG[2] includes one or more dialogue data LDs corresponding to the first video data MD1 that do not overlap with the dialogue data LD[1] included in the first text data group TG[1].

[0097] Furthermore, the prompt acquisition unit 115B acquires a prompt PP2[2] which includes the second segmented video data VD[2], the first text data group TG[1], and the second text data group TG[2]. Similarly, the second text data group TG[2] includes one or more dialogue data LDs corresponding to the first video data MD1 that do not overlap with the dialogue data LD[1] included in the first text data group TG[1].

[0098] Prompt PP2 may include a "first instruction" indicating that a portion of the first segmented video data VD[1] is to be extracted, similar to prompt PP1 according to the first embodiment. In this "first instruction," the number of dialogue data LDs corresponding to the length of each short video is an example of the "first number." Prompt PP2 may also include a "second instruction" indicating that the reason for extracting the second video data MD2 from the first segmented video data VD[1] is to be output. Prompt PP2 may also include a "third instruction" indicating that the reason for extracting the second video data MD2 from the first segmented video data VD[1], the title of the second video data MD2, and an overview of the second video data MD2 are to be output.

[0099] The input unit 116, similar to the input unit 116 in the first embodiment, inputs the prompt PP2 acquired by the prompt acquisition unit 115B to the image language model PLM stored in the generation server 30.

[0100] The information acquisition unit 117, similar to the information acquisition unit 117 in the first embodiment, acquires video information MI, which is information related to the second video data MD2 output from the image language model PLM, when the input unit 116 inputs the prompt PP2 to the image language model PLM.

[0101] Figure 9 shows an example of the flow of various data and information used in the video generation device 10B. For the sake of simplicity, unlike Figure 2, the prompt PP and video information MI are not shown in Figure 9.

[0102] The splitting unit 120 divides the first video data MD1 into three parts: the first split video data DVD[1], the second split video data DVD[2], and the third split video data DVD[3], each of which is less than or equal to the first time length.

[0103] The determination unit 114B determines a first text data group TG[1] which includes a first text data TD1[1] corresponding to the first segmented video data VD[1] from among a plurality of first text data TD1, and a dialogue data LD[1] corresponding to the first segmented video data VD[1] from among a plurality of dialogue data LDs. The determination unit 114B also determines a second text data group TG[2] which includes a first text data TD1[2] corresponding to the second segmented video data VD[2] from among a plurality of first text data TD1, and a dialogue data LD[2] corresponding to the second segmented video data VD[2] from among a plurality of dialogue data LDs. The determination unit 114B also determines a third text data group TG[3] which includes a first text data TD1[3] corresponding to the third segmented video data VD[3] from among a plurality of first text data TD1, and a dialogue data LD[3] corresponding to the second segmented video data VD[3] from among a plurality of dialogue data LDs.

[0104] The input unit 116 inputs the prompt PP2[1] to the image language model PLM stored in the generation server 30. The prompt PP2[1] includes a first text data group TG[1], a second text data group TG[2], a third text data group TG[3], and a first segmented video data VD[1]. The information acquisition unit 117 acquires the video information MI[1] output from the image language model PLM.

[0105] Furthermore, the input unit 116 inputs the prompt PP2[2] to the image language model PLM stored in the generation server 30. The prompt PP2[2] includes a first text data group TG[1], a second text data group TG[2], a third text data group TG[3], and a second segmented video data VD[2]. The information acquisition unit 117 acquires the video information MI[2] output from the image language model PLM.

[0106] Furthermore, the input unit 116 inputs the prompt PP2[3] to the image language model PLM stored in the generation server 30. The prompt PP2[3] includes a first text data group TG[1], a second text data group TG[2], a third text data group TG[3], and a third segmented video data VD[3]. The information acquisition unit 117 acquires the video information MI[3] output from the image language model PLM.

[0107] The information acquisition unit 117 determines the overall video information MI by combining video information MI[1], video information MI[2], and video information MI[3].

[0108] The generation unit 119 generates a first short video data SD[1] based on the video information MI[1] contained in the video information MI. The generation unit 119 also generates a second short video data SD[2] based on the video information MI[2] contained in the video information MI. The generation unit 119 also generates a third short video data SD[3] based on the video information MI[3] contained in the video information MI. The first short video data SD[1] corresponds to the second video data MD2[1]. The second short video data SD[2] corresponds to the second video data MD2[2]. The third short video data SD[3] corresponds to the second video data MD2[3].

[0109] In Figure 9, the first segmented video data DVD[1] and the second segmented video data DVD[2] are consecutive, and the second text data group TG[2] may include one or more dialogue data LD[2] corresponding to the second segmented video data DVD[2].

[0110] Furthermore, the second segmented video data VD[2] may be at least one of the video data from the beginning portion of the first video data MD1 and the video data from the end portion of the first video data MD1. In this case, the second text data group TG[2] includes one or more dialogue data LDs corresponding to at least one of the video data from the beginning portion of the first video data MD1 and the video data from the end portion of the first video data MD1.

[0111] 2-2: Figures 10 and 11 of the second embodiment are flowcharts showing an example of the operation of the video generation device 10B.

[0112] In step S11, the processing unit 11B provided in the video generation device 10B functions as a video acquisition unit 112. The processing unit 11B acquires the first video data MD1.

[0113] In step S12, the processing unit 11B determines the time length of the first video data MD1 acquired in step S11. If the time length of the first video data MD1 exceeds the first time length of n minutes (YES in step S12), the processing unit 11B executes the process in step S20. If the time length of the first video data MD1 is less than or equal to the first time length of n minutes (NO in step S12), the processing unit 11B functions as the communication control unit 111. The processing unit 11B transmits the first video data MD1 to the voice analysis server 20. After that, the processing unit 11B executes the process in step S13.

[0114] In step S13, the processing unit 11B functions as a text data acquisition unit 113B. The processing unit 11B acquires a plurality of first text data TD1 transmitted from the speech analysis server 20.

[0115] In step S14, the processing unit 11B functions as a determination unit 114B. The processing unit 11B determines a plurality of dialogue data LDs that correspond one-to-one to a plurality of dialogues by adding identification data DD, which identifies the dialogue, to each of the plurality of first text data TD1s acquired in step S13. Furthermore, the processing unit 11B determines a text data group TG that includes the plurality of first text data TD1s and the plurality of dialogue data LDs.

[0116] In step S15, the processing unit 11B functions as a prompt acquisition unit 115B. The processing unit 11B acquires a prompt PP2 which includes multiple dialogue data LDs and a first video data MD1.

[0117] In step S16, the processing unit 11B functions as an input unit 116. The processing unit 11B inputs the prompt PP2 acquired in step S15 to the image language model PLM stored in the generation server 30.

[0118] In step S17, the processing unit 11B functions as an information acquisition unit 117. The processing unit 11B acquires video information MI, which is information relating to the second video data MD2 output from the image language model PLM.

[0119] In step S18, the processing unit 11B functions as a display control unit 118. The processing unit 11B displays the video information MI acquired in step S17 on the display device 13.

[0120] In step S19, the processing unit 11B functions as a generation unit 119. The processing unit 11A generates short video data SD based on the video information MI acquired in step S17.

[0121] In step S20, the processing unit 11B functions as a splitting unit 120. The processing unit 11B divides the first video data MD1 into m segmented video data VD, each with a first time length of n minutes or less. Herein, m is an integer of 2 or more. The processing unit 11B also functions as a communication control unit 111. The processing unit 11B transmits each of the m segmented video data VD to the audio analysis server 20.

[0122] In step S21, the processing unit 11B adds 1 to i. i is an integer, and its initial value is 0. If, as a result of adding 1 to i, i is less than or equal to m (i ≤ m in step S21), the processing unit 11B executes the process in step S22. If, as a result of adding 1 to i, i exceeds m (i > m in step S21), the processing unit 11B executes the process in step S27.

[0123] In step S22, the processing unit 11B functions as a text data acquisition unit 113B. The processing unit 11B acquires a plurality of first text data TD1 transmitted from the speech analysis server 20.

[0124] In step S23, the processing unit 11B functions as a determination unit 114B. The processing unit 11B determines a plurality of dialogue data LDs that correspond one-to-one to a plurality of dialogues by adding identification data DDs that identify dialogues to each of the plurality of first text data TD1s acquired in step S13. Furthermore, the processing unit 11B determines a text data group TG that includes the plurality of first text data TD1s and the plurality of dialogue data LDs.

[0125] In step S24, the processing unit 11B functions as a prompt acquisition unit 115B. The processing unit 11B acquires a prompt PP2 which includes multiple dialogue data LDs and segmented video data VDs.

[0126] In step S25, the processing unit 11B functions as an input unit 116. The processing unit 11B inputs the prompt PP2 acquired in step S24 to the image language model PLM stored in the generation server 30.

[0127] In step S26, the processing unit 11B functions as an information acquisition unit 117. The processing unit 11B acquires video information MI, which is information relating to the second video data MD2 output from the image language model PLM.

[0128] In step S27, the processing unit 11B functions as an information acquisition unit 117. The processing unit 11B determines the overall video information MI by combining all the video information MI. After that, the processing unit 11B executes the process in step S18.

[0129] 2-3: Effects of the Second Embodiment The video generation device 10B according to this embodiment comprises a splitting unit 120, a determination unit 114B, and an information acquisition unit 117. When the time length of the first video data MD1 exceeds a first time length, the splitting unit 120 splits the first video data MD1 into video data of a first time length or less, thereby outputting a plurality of split video data VDs including the first split video data VD[1]. Based on the first split video data VD[1], the determination unit 114B determines a plurality of dialogue data LD[1] that correspond one-to-one to a plurality of lines of dialogue contained in the first split video data VD[1]. Each of the plurality of dialogue data LDs indicates at least the corresponding lines of dialogue in text. The information acquisition unit 117 acquires information about the second video data MD2 output from the image language model PLM by inputting a prompt PP2 to the image language model PLM, which includes the first segmented video data VD[1], a first text data group TG[1] including dialogue data LD[1] corresponding to the first segmented video data VD[1], and a second text data group TG[2]. The second text data group TG[2] includes one or more dialogue data LDs from among the multiple dialogue data LDs corresponding to the first video data MD1 that do not overlap with the dialogue data LD[1] included in the first text data group TG[1].

[0130] By having the above configuration, the video generation device 10B can analyze longer videos compared to conventional technology.

[0131] Specifically, the video generation device 10B divides the first video data MD1 into multiple segmented video data VD, each segmented with a duration less than or equal to the duration of the video that can be input to the multimodal LLM. This enables the analysis of longer videos compared to conventional technology.

[0132] Furthermore, in the video generation device 10B, the second text data group TG[2] includes one or more dialogue data LDs that correspond to video data that is continuous with the first segmented video data VD[1] from the first video data MD1.

[0133] The video generation device 10B, with the above configuration, can analyze multiple consecutive segmented video data VDs.

[0134] Furthermore, in the video generation device 10B, the second text data group TG[2] includes one or more dialogue data LDs corresponding to at least one of the video data from the beginning portion of the first video data MD1 and the video data from the end portion of the first video data MD1.

[0135] The video generation device 10B, with the above configuration, can analyze at least one of the leading and trailing portions of the first video data MD1, which is a segmented video data VD.

[0136] Furthermore, in the video generation device 10B, the multiple segmented video data VD includes the second segmented video data VD[2]. The second text data group TG[2] includes one or more dialogue data LD[2] from among the multiple dialogue data LDs that correspond to the second segmented video data VD[2].

[0137] The video generation device 10B, by having the above configuration, can limit the second text data group TG[2] to text data group TG in units of the second segmented video data VD[2].

[0138] Furthermore, in the video generation device 10B, each of the multiple dialogue data LDs includes text data TD that shows the corresponding dialogue in text and identification data DD that identifies the corresponding dialogue.

[0139] By having the above configuration, the video generation device 10B can obtain more accurate timestamps compared to conventional technology.

[0140] Specifically, in the video generation device 10B, each of the first text data TD1 is accompanied by identification data DD indicating the order in which the dialogue occurs, thereby making it possible to make the timestamp of the second video data MD2 accurate.

[0141] Furthermore, in the video generation device 10B, prompt PP2 includes a first instruction indicating that a portion of the first segmented video data VD[1] is to be extracted. The second video data MD2[1] is video data obtained by extracting a portion of the first segmented video data VD[1].

[0142] By having the above configuration, the video generation device 10B can more accurately extract the portion indicated by the video information MI from the first segmented video data VD[1] as the second video data MD2[1].

[0143] Furthermore, in the video generation device 10B, the first instruction indicates that a first number of consecutive lines of dialogue will be extracted from the first segmented video data DVD[1].

[0144] The video generation device 10B, with the above configuration, can specify the location from which to extract the second video data MD2[1] from the first segmented video data VD[1] using the number of lines of dialogue.

[0145] Furthermore, in the video generation device 10B, prompt PP2 includes a second instruction indicating that the reason for extracting the second video data MD2[1] from the first segmented video data VD[1] is to be output. The information acquisition unit 117 acquires the reason output from the image language model PLM.

[0146] The video generation device 10B, with the above configuration, can add a note to the second video data MD2[1] explaining why the second video data MD2[1] was extracted from the first segmented video data VD[1].

[0147] Furthermore, in the video generation device 10B, prompt PP2 includes a third instruction indicating that the reason for extracting the second video data MD2[1] from the first segmented video data VD[1], the title of the second video data MD2[1], and an overview of the second video data MD2[1] will be output. The information acquisition unit 117 acquires the reason for extracting the second video data MD2[1] from the first segmented video data VD[1], the title of the second video data MD2[1], and an overview of the second video data MD2[1] output from the image language model PLM. The video generation device 10B includes a display control unit 118. The display control unit 118 causes the reason for extracting the second video data MD2[1] from the first segmented video data VD[1], the title of the second video data MD2[1], and an overview of the second video data MD2[1] output from the image language model PLM to be displayed on the display device 13.

[0148] By having the above configuration, the video generation device 10B allows user U[k] to recognize the reason for extracting the second video data MD2[1] from the first segmented video data VD[1], the title of the second video data MD2[1], and an overview of the second video data MD2[1].

[0149] 3. Modifications The present disclosure is not limited to the embodiments illustrated above. Specific examples of modifications are given below.

[0150] 3-1: Modification 1 In the first embodiment, the video generation system 1A comprises a video generation device 10A, an audio analysis server 20, and a generation server 30. In the second embodiment, the video generation system 1B comprises a video generation device 10B, an audio analysis server 20, and a generation server 30. However, in addition to these components, the video generation system 1A and the video generation system 1B may also include a terminal 40 that is communicably connected to the video generation device 10A or the video generation device 10B via a communication network NET.

[0151] Figure 12 is a block diagram showing an example of the overall configuration of a video generation system 1C according to Modification 1. As shown in Figure 12, the video generation system 1C comprises a video generation device 10A, an audio analysis server 20, a generation server 30, and terminals 40[1] to 40[j]. The video generation device 10A, the audio analysis server 20, the generation server 30, and terminals 40[1] to 40[j] are connected to each other via a communication network NET, where j is an integer of 1 or more.

[0152] In Figure 12, user U uses terminal 40. Also, user U[1] uses terminal 40[1]. User U[2] uses terminal 40[2]. User U[k] uses terminal 40[k]. User U[j] uses terminal 40[j]. k is an integer between 1 and j, inclusive. In the following explanation, user U[k] will be used as a representative example of user U, and terminal 40[k] will be used as a representative example of terminal 40.

[0153] User U[k] may upload the first video data MD1 to the video generation device 10A using terminal 40[k]. Alternatively, user U[k] may upload the prompt PP1 to the video generation device 10A using terminal 40[k].

[0154] In this case, the display control unit 118 of the video generation device 10A causes the video information MI to be displayed on a display device (not shown) provided on the terminal 40[k].

[0155] 3-2: Modification 2 In the video generation device 10B according to the second embodiment, as an example shown in Figure 9, the splitting unit 120 splits the first video data MD1 into three split video data VDs: the first split video data VD[1], the second split video data VD[2], and the third split video data VD[3]. However, the splitting unit 120 may split the first video data MD1 into any number of split video data VDs. As an example, the splitting unit 120 may split the first video data MD1 into two split video data VDs: the first split video data VD[1] and the second split video data VD[2].

[0156] In this case, as an example, the first text data group TG[1] and the second text data group TG[2] may include multiple dialogue data LDs corresponding to the first video data MD1. In other words, the first text data group TG[1] and the second text data group TG[2] may include the entirety of multiple dialogue data LDs corresponding to the first video data MD1.

[0157] 3-3: Modification 3 In the video generation device 10B according to the second embodiment, if the time length of the first video data MD1 exceeds the first time length, the splitting unit 120 splits the first video data MD1 into video data of the first time length or less, thereby outputting a plurality of split video data VDs, and the output plurality of split video data VDs were transmitted to the audio analysis server 20. Subsequently, the plurality of lines of dialogue contained in each of the plurality of split video data VDs were transcribed, and for each split video data VD, a plurality of first text data TD1 corresponding one-to-one to the plurality of lines of dialogue were determined.

[0158] However, in the video generation device 10B according to the second embodiment, as in the first embodiment, as a result of transmitting the first video data MD1 acquired by the video acquisition unit 112 to the audio analysis server 20, a plurality of first text data TD1 that correspond one-to-one to a plurality of lines of dialogue contained in the entire first video data MD1 may be determined.

[0159] In this case, the splitting unit 120 may split the first video data MD1 into a plurality of split video data VD based on the plurality of first text data TD1. For example, the splitting unit 120 may split the first video data MD1 into a plurality of split video data VD by using a point where the correlation between two consecutive first text data TD1 is low as the splitting point.

[0160] 3-4: Modification 4 In the video generation device 10B according to the second embodiment, the input unit 116 individually inputs, for example, the prompt PP2[1] corresponding to the first segmented video data VD[1], the prompt PP2[2] corresponding to the second segmented video data VD[2], and the prompt PP2[3] corresponding to the third segmented video data VD[3] to the image language model PLM. However, the input unit 116 may also input a combined prompt PP2, which is a merge of the contents of prompt PP2[1], the contents of prompt PP2[2], and the contents of prompt PP2[3], to the image language model PLM.

[0161] 3-5: Modification 5 In the above embodiment, prompt PP1 or prompt PP2 is input to the image language model PLM stored in the generation server 30, and the image language model PLM outputs video information MI. The information acquisition unit 117 acquires the video information MI. The generation unit 119 generates short video data SD based on the video information MI.

[0162] However, the image language model PLM may output short video data SD instead of video information MI.

[0163] 4. Other (1) In the embodiments described above, ROM and RAM were given as examples for the storage device 12A and storage device 12B, but other suitable storage media include flexible disks, magneto-optical disks (e.g., compact disks, digital multipurpose disks, Blu-ray® disks), smart cards, flash memory devices (e.g., cards, sticks, key drives), CD-ROMs (Compact Disc-ROMs), registers, removable disks, hard disks, floppy® disks, magnetic strips, databases, servers, and other appropriate storage media.

[0164] (2) In the embodiments described above, the information, signals, etc. may be represented using any of the various different techniques. For example, the data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0165] (3) In the embodiments described above, the input and output information may be stored in a specific location (e.g., memory) or managed using a management table. The input and output information may be overwritten, updated, or appended to. The output information may be deleted. The input information may be transmitted to other devices.

[0166] (4) In the embodiments described above, the determination may be made by a value represented using one bit (0 or 1), by a boolean value (true or false), or by a numerical comparison (for example, a comparison with a predetermined value).

[0167] (5) The processing procedures, sequences, flowcharts, etc., exemplified in the embodiments described above may be rearranged in order, as long as there is no contradiction. For example, the methods described in this disclosure present various step elements using an exemplary order and are not limited to the specific order presented.

[0168] (6) Each function illustrated in Figures 1 to 12 is realized by any combination of at least one of hardware and software. Furthermore, the method of realizing each function block is not particularly limited. That is, each function block may be realized using one device that is physically or logically coupled, or it may be realized using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wired or wireless connections). A function block may also be realized by combining software with the one or more devices described above.

[0169] (7) The programs illustrated in the embodiments described above should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc., whether they are called software, firmware, middleware, microcode, hardware description languages ​​or by other names.

[0170] Furthermore, software, instructions, information, etc., may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using at least one of wired technology (such as coaxial cable, fiber optic cable, twisted pair, or digital subscriber line (DSL)) and wireless technology (such as infrared or microwave), then at least one of these wired and wireless technologies is included in the definition of a transmission medium.

[0171] (8) In each of the above-mentioned forms, the terms “system” and “network” shall be used interchangeably.

[0172] (9) The information, parameters, etc. described in this disclosure may be expressed using absolute values, relative values ​​from a given value, or other corresponding information.

[0173] (10) In the embodiments described above, the terminal 40 may be a mobile station (MS). A mobile station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or several other appropriate terms. In this disclosure, terms such as “mobile station,” “user terminal,” “user equipment (UE),” and “terminal” may be used interchangeably.

[0174] (11) In the embodiments described above, the terms “connected,” “coupled,” or any variation thereof, mean any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are “connected” or “coupled” with each other. The coupling or connection between elements may be a physical coupling or connection, a logical coupling or connection, or a combination thereof. For example, “connection” may be reinterpreted as “access.” As used in this disclosure, two elements may be considered to be “connected” or “coupled” with each other using at least one of one or more wires, cables and printed electrical connections, and, in some non-limiting and non-exclusive examples, electromagnetic energy having wavelengths in the radio frequency domain, microwave domain and optical (both visible and invisible) domain.

[0175] (12) In the embodiments described above, the phrase "based on" does not mean "based solely on" unless otherwise specified. In other words, the phrase "based on" means both "based solely on" and "based at least on".

[0176] (13) The terms “determining” and “determining” as used in this disclosure may encompass a wide variety of actions. “Determining” may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, or inquiring (e.g., searching in a table, database, or other data structure), or ascertaining. “Determining” may also include receiving (e.g., receiving information), transmitting (e.g., sending information), inputting, outputting, or accessing (e.g., accessing data in memory). Furthermore, "judgment" and "decision" can include considering something as having been "judged" or "decided" after resolving, selecting, choosing, establishing, comparing, etc. In other words, "judgment" and "decision" can include considering something as having been "judged" or "decided" after some action. Also, "judgment (decision)" can be reinterpreted as "assuming," "expecting," or "considering."

[0177] (14) In the embodiments described above, where “include,” “including,” and variations thereof are used, these terms are intended to be inclusive, as is the term “comprising.” Furthermore, the term “or” as used in this disclosure is not intended to be exclusive OR.

[0178] (15) In the present disclosure, if articles are added by translation, such as a, an, and the in English, the present disclosure may include the fact that the noun following these articles is plural.

[0179] (16) In this disclosure, the term “A and B are different” may mean “A and B are different from each other.” The term may also mean “A and B are each different from C.” Terms such as “separate” and “combine” may be interpreted in the same way as “different.”

[0180] (17) Each aspect / embodiment described herein may be used individually, in combination, or switched between as needed during implementation. Furthermore, notification of the specified information (e.g., notification that "it is X") is not limited to explicit notification, but may also be implicit (e.g., by not providing such notification).

[0181] Although the present disclosure has been described in detail above, it will be clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the intent and scope of the present disclosure as defined by the claims. Accordingly, the descriptions in the present disclosure are illustrative and not restrictive in any way.

[0182] 1A...Video generation system, 1B...Video generation system, 1C...Video generation system, 10A...Video generation device, 10B...Video generation device, 11A...Processing device, 11B...Processing device, 12A...Storage device, 12B...Storage device, 13...Display device, 14...Input device, 15...Communication device, 20...Voice analysis server, 30...Generation server, 40...Terminal, 111...Communication control unit, 112...Video acquisition unit, 113A...Text data acquisition unit, 113B...Text data acquisition unit, 114A...Decision unit, 114B...Decision unit, 115A...Prompt acquisition unit, 115B ...Prompt acquisition unit, 116...Input unit, 117...Information acquisition unit, 118...Display control unit, 119...Generation unit, 120...Splitting unit, BL...Block, DD...Identification data, LD...Dialogue data, MD1...First video data, MD2...Second video data, MI...Video information, NET...Communication network, PLM...Image language model, PP...Prompt, PR1A...Control program, PR1B...Control program, SD...Short video data, TD...Text data, TD1...First text data, TG...Text data group, U...User, VD...Split video data

Claims

1. A video generation device comprising: a splitting unit that outputs a plurality of split video data including a first split video data by splitting the first video data into video data of a length equal to or less than the first split video data when the length of the first video data exceeds a first length; a determination unit that determines a plurality of dialogue data that correspond one-to-one with a plurality of lines of dialogue contained in the first split video data based on the first split video data, and each of the plurality of dialogue data indicates at least the corresponding lines of dialogue in text; and an information acquisition unit that acquires information about a second video data output from the image language model by inputting a prompt to an image language model that includes the first split video data, a first group of text data including the first split video data, and a second group of text data, wherein the second group of text data includes one or more lines of dialogue from among the plurality of lines of dialogue data corresponding to the first video data that do not overlap with the lines of dialogue included in the first group of text data.

2. The video generation apparatus according to claim 1, wherein the first text data group and the second text data group include a plurality of dialogue data corresponding to the first video data.

3. The video generation apparatus according to claim 1, wherein the second group of text data includes one or more lines of dialogue corresponding to video data that is continuous with the first segmented video data among the first video data.

4. The video generation device according to claim 1, wherein the second group of text data includes one or more lines of dialogue corresponding to at least one of the video data from the beginning portion of the first video data and the video data from the end portion of the first video data.

5. The video generation apparatus according to claim 1, wherein the plurality of segmented video data includes a second segmented image data, and the second group of text data includes one or more lines of dialogue data from the plurality of lines of dialogue data that correspond to the second segmented image data.

6. The video generation device according to claim 1, wherein each of the plurality of dialogue data includes text data that shows the corresponding dialogue in text and identification data that identifies the corresponding dialogue, and the identification data indicates the order in which the dialogue occurs.

7. The video generation apparatus according to claim 1, wherein the prompt includes a first instruction indicating the extraction of a portion of the first segmented video data, and the second video data is video data obtained by extracting a portion of the first segmented video data.

8. The video generation device according to claim 7, wherein the first instruction indicates that a portion corresponding to a first number of consecutive lines of dialogue is extracted from the first segmented video data.

9. The video generation apparatus according to claim 7, wherein the prompt includes a second instruction indicating that the reason for extracting the second video data from the first segmented video data is to be output, and the information acquisition unit acquires the reason output from the image language model.

10. The video generation apparatus according to claim 7, wherein the prompt includes a third instruction indicating that the reason for extracting the second video data from the first segmented video data, the title of the second video data, and a summary of the second video data are to be output, the information acquisition unit acquires the reason for extracting the second video data from the first segmented video data, the title of the second video data, and a summary of the second video data output from the image language model, and the display control unit causes the reason for extracting the second video data from the first segmented video data, the title of the second video data, and a summary of the second video data output from the image language model to be displayed on a display device.