Video generation device, video generation method, video generation program, and video generation system
The video generation system addresses the lack of user control in existing technologies by using interactive AI and voice analysis to create videos that meet user-defined policies, enabling high-quality video creation without specialized knowledge or equipment.
Patent Information
- Application Number
- PCT/JP2024/045665
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-19
- Filing Date
- 2024-12-24
- Publication Date
- 2025-10-23
AI Technical Summary
Existing video editing technologies do not consider the structure and appearance of generated videos, limiting user control over the final product, especially for users without specialized knowledge or equipment.
A video generation system that utilizes an interactive AI and voice analysis to create a video editing plan based on user input, allowing for the generation of videos with specified constraints, visual effects, and additional content, using a user terminal, interactive AI service, and voice analysis service to edit videos according to desired policies.
Enables users to create high-quality videos in their desired format without requiring video editing skills or specialized equipment, by generating videos that adhere to user-defined policies and constraints.
Smart Images

Figure JP2024045665_23102025_PF_FP_ABST
Abstract
Description
Video generation device, video generation method, video generation program, and video generation system
[0001] The present invention relates to a video generation device, a video generation method, a video generation program, and a video generation system. This invention claims priority from Japanese Patent Application No. 2024-068392 filed on April 19, 2024, and the contents of that application are incorporated by reference into this application in designated states where incorporation by reference of documents is permitted.
[0002] Patent document 1 describes a video editing device that includes an edited video generation unit that generates edited video data of the length desired by the user and including the content of the part that the user wants to view, in response to a user request.
[0003] Japanese Patent Application Laid-Open No. 2021-087180
[0004] The above technology involves inputting index information to identify each section of video data divided into predetermined lengths, thereby identifying the part that the user wants to view, and generating edited video data that includes that content. Therefore, the structure and appearance of the generated video are not necessarily taken into consideration.
[0005] An object of the present invention is to generate a moving image in a manner desired by a user.
[0006] The present application includes multiple means for solving at least part of the above problems, examples of which are as follows: A moving image generation device according to one aspect of the present invention is characterized by having: an acquisition unit that acquires a moving image file; an analysis unit that acquires an analysis result including time-series utterance information, which is text information obtained by transcribing utterances included in the moving image file in time series; an editing plan unit that transmits the time-series utterance information and command information that instructs an editing device to output an editing plan for editing the time-series utterance information in accordance with desired editing policy information to an interactive AI using a language model and receives the editing plan from the interactive AI; and a moving image editing unit that edits the moving image file in accordance with the editing plan and generates an edited moving image.
[0007] In the above moving image generating device, the editing planning section may include, in the command information, constraints on the edited moving image obtained by the editing plan.
[0008] In the above moving image generating device, the editing planning section may include, in the command information, configuration information about the edited moving image obtained according to the editing plan.
[0009] In the above moving image generating device, the editing planning section may include in the command information a designation of a moving image, a still image or an audio to be added to the edited moving image obtained by the editing plan.
[0010] In the above moving image generating device, the editing planning section may include, in the command information, a designation of a visual effect to be used in the edited moving image obtained according to the editing plan.
[0011] In addition, in the above-mentioned video generation device, the editing plan may include information for constructing the edited video by connecting together partial videos whose start and end positions on a time axis within the video file are specified, and the video editing unit may generate the edited video by cutting out the partial videos from the video file and connecting them together.
[0012] In addition, in the above-mentioned video generation device, the editing plan may include information for connecting partial videos that specify start and end positions on the elapsed time axis within the video file to create the edited video, and designation of video, still images, or audio to be added before and after the partial videos, and the video editing unit may generate the edited video by cutting out the partial videos from the video file, connecting them together, and adding the video, still images, or audio to be added.
[0013] In addition, in the above-mentioned video generation device, the editing plan may include information for connecting partial videos having specified start and end positions on a time axis within the video file to create the edited video, and a specification of visual effects to be used at the joints of the partial videos, and the video editing unit may generate the edited video by cutting out the partial videos from the video file, connecting them together, and applying the specified visual effects to the joints.
[0014] In addition, in the above-mentioned video generation device, the analysis unit may fast-forward edit the speech audio contained in the video file while maintaining the chronological order, and pass it on to a predetermined speech-to-text conversion unit to obtain the chronological speech information.
[0015] In addition, in the above-mentioned video generation device, the analysis unit may identify the speaker of the speech audio included in the video file, extract the speech audio for each speaker while maintaining a time series, and transfer the text information obtained to a predetermined speech-to-text conversion unit to integrate the text information to obtain the time series speech information.
[0016] In addition, in the above-mentioned video generation device, the editing plan may be described in a predetermined format language, and the editing plan unit may include definition information about the format language used to describe the editing plan in the command information.
[0017] In addition, a video generation method according to another aspect of the present invention is a video generation method using a video generation device, the video generation device including a processor, the processor performing the following steps: an acquisition step of acquiring a video file; an analysis step of acquiring an analysis result including time-series utterance information, which is text information obtained by transcribing utterances included in the video file in chronological order; an editing planning step of transmitting the time-series utterance information and command information instructing an editing plan to edit the time-series utterance information in accordance with desired editing policy information to an interactive AI using a language model and receiving the editing plan from the interactive AI; and a video editing step of editing the video file in accordance with the editing plan and generating an edited video.
[0018] In addition, a video generation program according to another aspect of the present invention is a video generation program that causes an information processing device to generate a video, wherein the information processing device has a processor and causes the processor to perform the following steps: an acquisition step of acquiring a video file; an analysis step of acquiring an analysis result including time-series utterance information, which is text information obtained by transcribing utterances included in the video file in chronological order; an editing planning step of transmitting the time-series utterance information and command information instructing the processor to output an editing plan for editing the time-series utterance information in accordance with desired editing policy information to an interactive AI using a language model and receiving the editing plan from the interactive AI; and a video editing step of editing the video file in accordance with the editing plan and generating an edited video.
[0019] In addition, another aspect of the present invention provides a video generation system comprising a user terminal and a video generation device communicatively connected to the user terminal, wherein the video generation device comprises an acquisition unit that acquires a video file from the user terminal via communication, an analysis unit that acquires an analysis result including time-series utterance information, which is text information obtained by transcribing utterances contained in the video file in chronological order, an editing planning unit that transmits the time-series utterance information and command information that instructs an editing plan to be written and output for editing the time-series utterance information in accordance with desired editing policy information to an interactive AI using a language model and receives the editing plan from the interactive AI, and a video editing unit that edits the video file in accordance with the editing plan and generates an edited video.
[0020] According to the present invention, it is possible to provide a technique for generating a video in a format desired by a user.
[0021] Problems, configurations, and effects other than those described above will become apparent from the following description of the embodiments.
[0022] 1 is a diagram illustrating an overview of a video production system according to an embodiment; FIG. 2 is a diagram illustrating a configuration of a video production system according to an embodiment; FIG. 3 is a diagram illustrating an example of the data structure of material information; FIG. 4 is a diagram illustrating an example of the data structure of time-series speech information; FIG. 5 is a diagram illustrating an example of the data structure of editing policy information; FIG. 6 is a diagram illustrating an example of the data structure of command information; FIG. 7 is a diagram illustrating an example of the data structure of an editing plan; FIG. 8 is a diagram illustrating an example of the hardware configuration of a video production device; FIG. 9 is a diagram illustrating an example of a video production flow (video material registration); FIG. 10 is a diagram illustrating an example of a video production flow (editing policy registration); FIG. 11 is a diagram illustrating an example of a video material registration screen; FIG. 12 is a diagram illustrating an example of a new material registration screen; FIG. 13 is a diagram illustrating an example of an editing policy registration screen; FIG. 14 is a diagram illustrating an example of a new editing policy registration screen.
[0023] A video generation system 1 to which an embodiment according to one aspect of the present invention is applied will be described below with reference to the drawings. In the following embodiment, when necessary for convenience, the description will be divided into multiple sections or embodiments. However, unless otherwise specified, they are not unrelated to each other, and one is related to the other in terms of partial or full modification, details, supplementary explanation, etc.
[0024] Furthermore, in the following embodiments, when referring to the number of elements (including the number, numerical value, amount, range, etc.), unless otherwise specified or when it is clearly limited to a specific number in principle, it is not limited to that specific number and may be more or less than the specific number.
[0025] Furthermore, it goes without saying that in the following embodiments, the components (including element steps, etc.) are not necessarily essential unless specifically stated otherwise or unless they are clearly considered essential in principle.
[0026] Similarly, in the following embodiments, when referring to the shapes, positional relationships, etc. of components, etc., it is intended to include those that are substantially similar or similar to those shapes, etc., unless otherwise specified or when it is considered that this is clearly not the case in principle. This also applies to the above numerical values and ranges.
[0027] In addition, in all the drawings for explaining the embodiments, the same components are generally designated by the same reference numerals, and repeated explanations thereof will be omitted.
[0028] In recent years, with the spread of networks and various electronic devices (personal computers, tablet devices, smartphones, etc.), an environment is being created in which videos can be created and published anytime, anywhere. For example, it is becoming possible for anyone to easily shoot videos using a smartphone or the like and post them anywhere and at any time to social networking services (SNSs) or video sharing sites that anyone can access. However, high-quality videos that attract attention are often created by editors with specialized knowledge who invest a lot of time and effort.
[0029] Therefore, in an embodiment of the present invention, a video creation system 1 is provided that accepts a video editing policy desired by a user and automatically creates a video in accordance with the policy. The video creation system 1 allows a user to create a video in a format desired by the user, even if the user does not have video editing skills or does not have the equipment and environment for video creation.
[0030] 1 is a diagram showing an overview of a video generation system according to this embodiment. In the video generation system 1, a user uses a user terminal 400 that the user uses and a group of devices that are communicatively connected to the user terminal 400 via a communication path. The group of devices includes a video generation device 100, a group of devices that provide an interactive AI service 200, and a group of devices that provide a voice analysis service 300.
[0031] For example, the group of devices providing the interactive AI service 200, the group of devices providing the voice analysis service 300, and the video generation device 100 may be a cloud computer connected via the Internet, or a server device managed by the owner of the video generation device 100, the group of devices providing the interactive AI service 200, and the group of devices providing the voice analysis service 300. Furthermore, without being limited to this, a wearable device such as a user's smartwatch may be used as the user terminal 400.
[0032] When the user terminal 400 communicates with the device group (including the video generation device 100, the device group that provides the interactive AI service 200, and the device group that provides the voice analysis service 300), they are connected via a communication path such as a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, a mobile phone network, or a short-range wireless communication such as Bluetooth (registered trademark), or a communication network that is a combination of these. The communication path 50 may be a VPN (Virtual Private Network) on a wireless communication network such as a mobile phone network.
[0033] By using the video creation system 1, a user can create a video in the format desired. Specifically, the user uses the user terminal 400 to register a video material file containing recorded speech and environmental sounds and a video editing policy (1 and 2) in the video creation system 1, and then applies for video creation. The video creation device 100 extracts only the audio component from the video material and requests the audio analysis service 300 to analyze it as an audio file (3) for the video material. The audio analysis service 300 analyzes the audio file for the video material and returns an analyzed text file (4) to the video creation device 100, which associates the timing of speech with the spoken words and sentences.
[0034] The video production device 100 sends the analyzed text file obtained from the audio analysis service 300 and command information (5) including the video editing policy to the interactive AI service 200, requesting the creation of an editing plan. At this time, the video production device 100 does not send actual video material files or audio files to the interactive AI service 200, but sends the analyzed text file. The interactive AI service 200 creates an editing plan that satisfies the constraints specified in the command information, constraints on the edited video obtained by the editing plan, configuration information, designation of video, still images or audio to be added, designation of visual effects, etc., and returns it to the video production device 100 as an editing plan (6).
[0035] When the video generation device 100 receives the editing plan from the interactive AI service 200, it performs video editing processing in accordance with the editing plan, creates an edited video (7), and provides it to the user terminal 400. This allows the user to utilize the provided edited video.
[0036] 2 is a configuration diagram of a video generation system according to an embodiment. The video generation system 1 includes a video generation device 100, an interactive AI service 200 that can communicate with the video generation device 100 via a communication path 50, a voice analysis service 300, and a user terminal 400.
[0037] The moving image generating device 100 includes a storage unit 110, a processing unit 120, an input / output unit 140, and a communication unit 150, which are communicably connected to one another via a bus or the like.
[0038] The storage unit 110 includes material information 111 , time-series speech information 112 , editing policy information 113 , command information 114 , an editing plan 115 , and an edited video 116 .
[0039] 3 is a diagram showing an example data structure of material information. The material information 111 stores information on multiple source videos to be used in generating videos. The material information 111 includes a user 111A, a video title 111B, a video file path 111C, a description 111D, an analysis completion flag 111E, and an analysis result 111F.
[0040] User 111A is information that distinguishes a user from other users. Video title 111B is the title of the video to be registered as material. Video file path 111C is the storage location on the file system of the video to be registered as material, or a URI (Uniform Resource Identifier). Description 111D is information that explains in natural language the content of the video to be registered as material. Analyzed flag 111E is information that indicates whether analysis by the voice analysis service 300 has been completed. Analysis result 111F is analyzed text that is information on the results of analysis by the voice analysis service 300.
[0041] 4 is a diagram showing an example of the data structure of time-series utterance information. The time-series utterance information 112 is information that stores the text of utterances made in a video in chronological order, with the time elapsed in the video being treated as a time series. The time-series utterance information 112 includes an utterance start time 112A, an utterance end time 112B, and utterance text (words) 112C.
[0042] The speech start time 112A and the speech end time 112B are information specifying the start and end timings of speech made in a video, respectively, based on the elapsed time from the start time of the video (time within the video). The speech text (word) 112C is a word spoken between the speech start time 112A and the speech end time 112B. However, it is not limited to a word, and may be a sentence or a phrase of a certain length.
[0043] 5 is a diagram showing an example of the data structure of editing policy information. The editing policy information 113 is information about the editing policy of the video to be generated. The editing policy information 113 includes a title 113A, a content goal 113B, constraints 113C, a content structure 113D, a resource file 113E, and an editing plan format 113F.
[0044] Title 113A is the title of the editing policy or the title of the video to be generated. Content goal 113B is information such as the image that the video to be generated aims to achieve and the psychological changes that the viewer is expected to experience (such as making the viewer feel happy or relaxed when watching). Constraints 113C are information on constraints for creating the video, such as the length (playback time) of the video to be generated. Content structure 113D is information on the structure of the video to be generated, such as connecting three consecutive videos with a visual effect transition. Resource file 113E is information on the video materials to be used in the video to be generated. Editing plan format 113F is information specifying the format of the editing plan for generating the video. The format of the editing plan may be a known format or may be defined in an extension language conforming to SGML (Standard Generalized Markup Language) or the like.
[0045] 6 is a diagram showing an example of the data structure of command information. The command information 114 is a command (prompt) for causing the interactive AI service 200 to perform processing. The command for generating a video according to this embodiment, for example, specifies the editing policy information 113 and instructs the creation of an editing plan in accordance with the editing policy, and is written in natural language.
[0046] 7 is a diagram showing an example of the data structure of an editing plan. The editing plan 115 describes editing information in a predetermined format, for example by specifying tags for components to be assigned to times within the video to be generated, to serve as planning information for creating the video.
[0047] An outline of the format of the editing plan according to this embodiment will be described below. First, the editing plan can broadly include three types of elements: "shot," "view," and "attach." The "shot" tag is a collection of multiple "view" tags. "view" specifies a source file and related information. Source files include videos (including the start and end times of the portions to be used within the video material) and images (the magnification rate and on-screen layout of the image file), and related information includes color and gradation specifications. "attach" specifies an element to be displayed in addition to the source specified by "view" (for images, this includes the size, layout, and start and end times within the video to be generated; for audio, this includes the audio volume and start and end times within the video to be generated).
[0048] For example, cut editing information such as assigning an excerpt from a source video extracted by specifying a time within the source video to one of the components ("views"), playing multiple such excerpts continuously with transitions in between, and then adding a time to display a QR code (registered trademark) for accessing other videos in the channel ("attaches") is described.
[0049] For example, the editing plan may include, as an editing plan, information for creating an edited video by joining together partial videos whose start and end positions on the elapsed time axis within the video files to be used as raw materials. The editing plan may also include, as an editing plan, information for creating an edited video by joining together partial videos whose start and end positions on the elapsed time axis within the video files to be used as raw materials, and specifications for videos, still images, or audio to be added before and after the partial videos. The editing plan may also include, as an editing plan, information for creating an edited video by joining together partial videos whose start and end positions on the elapsed time axis within the video files to be used as raw materials, and specifications for visual effects to be used at the joins between the partial videos.
[0050] Returning to the explanation of Fig. 2, the processing unit 120 includes an acquisition unit 121, an analysis unit 122, an editing planning unit 123, and a video editing unit 124.
[0051] The acquisition unit 121 acquires a video file. The analysis unit 122 acquires, from a voice analysis service, an analysis result including time-series speech information, which is text information obtained by transcribing speech included in the video file in time series. The analysis unit 122 may also fast-forward edit the speech included in the video file while maintaining the time series, and pass the text information to a predetermined voice-to-text conversion unit (voice analysis service 300) to obtain the time-series speech information. Alternatively, the analysis unit 122 may identify speakers of speech included in the video file, extract the text information for each speaker while maintaining the time series, and transfer the text information to a predetermined voice-to-text conversion unit (voice analysis service 300) to obtain the time-series speech information.
[0052] The editing planning unit 123 transmits the time-series utterance information and command information instructing the interactive AI (the interactive AI service 200) using a language model to output an editing plan for editing the time-series utterance information in accordance with desired editing policy information, and receives an editing plan from the interactive AI. The editing planning unit may also include, in the command information, constraints on the edited video obtained by the editing plan. The editing planning unit 123 may also include, in the command information, configuration information on the edited video obtained by the editing plan. The editing planning unit 123 may also include, in the command information, designation of a video, still image, or audio to be added to the edited video obtained by the editing plan. The editing planning unit 123 may also include, in the command information, designation of a visual effect to be used in the edited video obtained by the editing plan. The editing planning unit 123 may also include, in the command information, definition information on a format language for describing the editing plan.
[0053] The video editing unit 124 edits the video files in accordance with the editing plan included in the editing plan to generate an edited video. Specifically, the video editing unit 124 generates the edited video by cutting out partial videos from video material files and splicing them together. The video editing unit 124 may also generate the edited video by adding additional videos, still images, or audio. The video editing unit 124 may also generate the edited video by cutting out partial videos from video files, splicing them together, and applying specified visual effects to the splices.
[0054] The input / output unit 140 controls input and output to and from the video generating device 100. For example, the input / output unit 140 accepts various types of input, such as various contact inputs such as typing, touching, and flick input, or various types of inputs such as gaze input. The input / output unit 140 also outputs information to the user. The output information includes various types of output information such as screens, presentation information, advertisements, and videos.
[0055] The communication unit 150 communicates via the communication path 50 with a group of devices that provide the interactive AI service 200, a group of devices that provide the voice analysis service 300, a user terminal 400, and other terminals that communicate via the Internet.
[0056] The interactive AI service 200 is a service that provides the functions of so-called generation AI, such as GPT and Gemini, via an API (Application Programming Interface). The interactive AI service 200 provides natural language prompts to the generation AI to generate the desired results. In this embodiment, the generation AI generates an editing plan for generating a video.
[0057] The voice analysis service 300 performs voice analysis using known technology such as the Google TTS API. Upon receiving an audio file, the voice analysis service 300 transcribes the speech in the audio file into text, and outputs, as analyzed text, the text of the speech content for each utterance included in the audio file and information specifying the start and end times of the utterance.
[0058] The user terminal 400 is a terminal used by a user. The user terminal 400 may be a smartphone terminal of the user, a personal computer (PC), or the like. Furthermore, without being limited to this, the user terminal 400 may be a wearable device such as a smart watch of the user.
[0059] 8 is a diagram showing an example of the hardware configuration of a moving image generation device. The moving image generation device 100 has a hardware configuration realized by the housing of a so-called server device, workstation, personal computer, smartphone, or tablet terminal. The moving image generation device 100 includes a processor 101, a memory 102, a storage 103, an input device 104, a display device 105, a communication device 106, and a bus connecting each device.
[0060] The processor 101 is an arithmetic device such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit).
[0061] The memory 102 is a memory device such as a RAM (Random Access Memory).
[0062] The storage 103 is a non-volatile storage device capable of storing digital information, such as a hard disk drive, a solid state drive (SSD), or a flash memory.
[0063] The input device 104 is a device that accepts input from one or more of a keyboard, a mouse, a touch panel, and a microphone. The display device 105 is a device that displays one or more of various output devices such as an organic EL (Electro-Luminescence) display.
[0064] The communication device 106 is a network interface card (NIC) or the like that communicates with other devices via a network.
[0065] In addition, the device providing the interactive AI service 200, the device providing the voice analysis service 300, and the user terminal 400 also have hardware configurations that are approximately the same as those of the video generation device 100.
[0066] The processing unit 120, acquisition unit 121, analysis unit 122, editing planning unit 123, and video editing unit 124 of the video generating device 100 are realized by a program that causes the processor 101 to perform processing. This program is stored in the memory 102, storage 103, or a ROM device (not shown), and is loaded onto the memory 102 for execution and executed by the processor 101.
[0067] The memory unit 110 of the moving image generating device 100 is realized by the memory 102 and the storage 103. The input / output unit 140 is realized by the input device 104 and the display device 105. The communication unit 150 is realized by the communication device 106. The above is an example of the hardware configuration of the moving image generating device 100.
[0068] The configuration of the moving image production device 100 can be further divided into more components depending on the processing content, or can be divided so that one component performs even more processing.
[0069] Furthermore, each processing unit (the processing unit 120, the acquisition unit 121, the analysis unit 122, the editing planning unit 123, and the video editing unit 124) may be constructed using dedicated hardware (such as an ASIC or a GPU) that realizes the respective functions. Furthermore, the processing of each processing unit may be executed by a single piece of hardware or by multiple pieces of hardware.
[0070] Next, the operation of the video generation system 1 in this embodiment will be described.
[0071] 9 is a diagram showing an example of a video creation flow (video material registration). The video creation flow (video material registration) starts when a user requests its start in a web browser or application software (hereinafter sometimes simply referred to as a browser) on the user terminal 400.
[0072] The acquisition unit 121 of the video production device 100 generates a video material registration screen and displays it on the user terminal 400 (step S001). Specifically, the acquisition unit 121 generates a video material registration screen on which the user manages a list of videos that have been previously registered. The acquisition unit 121 then transmits display information for the generated video material registration screen to the user terminal 400.
[0073] Then, the browser of the user terminal 400 displays a video material registration screen and sends a video material registration request to the video generation device 100, along with information including the video material file to be registered, the video title, and explanatory information (step S002).
[0074] The acquisition unit 121 acquires the video material file etc. (step S003). Specifically, the acquisition unit 121 registers the user, video title, video file, and description in the material information 111.
[0075] Then, the analysis unit 122 performs video analysis (audio portion extraction) (step S004). Specifically, the analysis unit 122 separates and acquires audio components from the acquired video file.
[0076] Then, the analysis unit 122 performs video analysis (fast-forward audio generation) (step S005). Specifically, the analysis unit 122 fast-forward edits the audio components included in the acquired video file while maintaining the chronological order. For example, when processing a video file of an utterance from 0 minutes 15 seconds into the video until 0 minutes 27 seconds into the video (the utterance lasts 12 seconds), the analysis unit 122 edits the video at four times the normal speed, reducing the data size of the audio file so that the time from the start of the utterance to the end of the utterance is three seconds.
[0077] Then, the analysis unit 122 performs video analysis (audio analysis request) (step S006). Specifically, the analysis unit 122 requests the audio analysis service 300 to analyze the audio file of the fast-forwarded audio created in step S005 by sending it via an API or the like.
[0078] The audio analysis service 300 performs an audio analysis process on the transmitted audio file of the fast-forwarded audio (step S007). Specifically, the audio analysis service 300 generates a speech timing analyzed text of the raw video in which the speech timing and speech content of the raw video are associated and recorded, and transmits the text to the video production device 100.
[0079] Then, the analysis unit 122 performs video analysis (creation of time-series information) (step S007). Specifically, the analysis unit 122 stores the received analyzed text in the time-series utterance information 112, sets the analysis completion flag 111E to "completed," and stores reference information to the time-series utterance information 112 in the analysis result 111F. At this time, if the data structures of the analyzed text and the time-series utterance information 112 are different, the analysis unit 122 may convert the information of the analyzed text so that the time information is returned from the fast-forward state to the normal speed state and store the converted information as the time-series utterance information 112, or may convert the time information so that the time information is returned from the fast-forward state to the normal speed state and then convert it into the data structure of the time-series utterance information 112 and store the converted information.
[0080] The above is an example of the video generation flow (video material registration). According to the video generation flow (video material registration), for a video registered as video material, it is possible to obtain time-series speech information by analyzing text information of speech and the timing of that speech in the video.
[0081] 10 is a diagram showing an example of a video creation flow (editing policy registration). The video creation flow (editing policy registration) starts when the user requests its start on the browser of the user terminal 400.
[0082] The editing planning unit 123 of the video production device 100 generates an editing policy registration screen and displays it on the user terminal 400 (step S101). Specifically, the editing planning unit 123 generates an editing policy registration screen that manages a list of editing policies that the user has previously registered. Then, the editing planning unit 123 transmits display information for the generated editing policy registration screen to the user terminal 400.
[0083] Then, the browser of the user terminal 400 displays an editing policy registration screen and sends an editing policy registration request to the video production device 100, along with information including the editing policy title to be registered, the video material to be registered, and the order (step S102).
[0084] The editing planning unit 123 receives the editing policy and the like (step S103). Specifically, the editing planning unit 123 registers the editing policy title, the content goal, constraints, content structure, and editing plan format based on the order, and resource files based on the registered video material in the editing policy information 113. The editing planning unit 123 interprets the natural language written in the order to identify the content goal, constraints, content structure, and editing plan format included in the order.
[0085] Then, the editing planning unit 123 performs editing preparation (creating command information) (step S104). Specifically, the editing planning unit 123 creates command information 114. For example, the editing planning unit 123 replaces the designated portion of the editing policy data in the command information 114 with the content of the editing policy information 113, and generates a prompt to be passed to the interactive AI service 200.
[0086] Then, the editing planning unit 123 performs editing preparation (planning request) (step S105). Specifically, the editing planning unit 123 transmits the command information 114 created in step S104 and the speech timing analyzed text of the raw video to the interactive AI service 200 via an API or the like.
[0087] The interactive AI service 200 then performs editing planning processing in accordance with the transmitted command information (step S106). Specifically, the interactive AI service 200 uses the speech timing analyzed text of the raw video to perform cut editing, focusing on important parts and statements that are evaluated as interesting or intriguing, taking into consideration the speech content (meaning) and speech timing, and creates an editing plan to incorporate transitions and attachments according to the order to meet the specified length. The interactive AI service 200 generates the planned editing content as an editing plan in a specified format and transmits it to the video production device 100.
[0088] Then, the video editing unit 124 performs video editing (creates an edited video) in accordance with the editing plan 115 (step S107). Specifically, when the video editing unit 124 receives the transmitted editing plan, it stores it in the editing plan 115 in the storage unit 110. Then, the video editing unit 124 performs video editing (creates an edited video) in accordance with the editing plan 115, stores the edited video obtained as a result of the video editing in the edited video 116 in the storage unit 110, and also transmits it to the user terminal 400. Note that the video editing unit 124 may post the edited video obtained as a result of the video editing on a website so that it can be downloaded and transmit a link to the user terminal 400, or may upload the edited video from the user terminal 400 to a video sharing site designated in advance.
[0089] The above is an example of the video generation flow (editing policy registration). According to the video generation flow (editing policy registration), video materials can be edited to obtain an edited video in accordance with the time-series speech information obtained by analyzing the video registered as video materials and the editing plan created using the editing policy. Therefore, it can be said that a video in the format desired by the user can be generated.
[0090] 11 is a diagram showing an example of a video material registration screen. A video material registration screen example 600 displays information including at least a video title 611 and explanatory information 615 for each registered video material file 610. Additionally, the video material registration screen example 600 includes an editing policy display button 601 for receiving an instruction to transition to an editing policy list screen, a new registration button 602 for receiving an instruction to newly register a video material, and, for each registered video material file 610, a video file name 612, a content analysis status 613, and a delete button 614 for canceling the registration of the video material.
[0091] The content analysis status 613 is information indicating whether or not time-series speech information obtained by analyzing speech text information and the timing of the speech in the video has been obtained for the registered video material. When an input is received from the editing policy display button 601, the screen transitions to an example of an editing policy registration screen, which will be described later. When an input is received from the new registration button 602, the screen transitions to an example of a new material registration screen, which will be described later.
[0092] 12 is a diagram showing an example of a new material registration screen. A screen example 650 of the new material registration screen includes at least a video title 651, a video file name 652, a reference button 653 for referencing and inputting a file path indicating the storage location of the video file specified by the video file name 652, a description input field 654 for the material file, a close button 655 for receiving an instruction to transition to the video material registration screen, and a register button 656 for receiving an instruction to register the video material, for the video material to be registered by the user.
[0093] The material file description input field 654 accepts a description of the material content in free text. For example, in the case of video material, the material file description input field 654 accepts a synopsis or a description of scenes at each time in the video. When a registration button 656 accepts an instruction to register video material, it executes the registration process of step S003 in the video generation flow (video material registration).
[0094] 13 is a diagram showing an example of an editing policy registration screen. An example screen 700 of the editing policy registration screen displays, for each registered editing policy 710, at least an editing policy name 711, a delete button 712 for canceling the registration of the editing policy, an order 713 that is the specific content of the editing policy, a create editing plan button 714 for receiving an instruction to create an editing plan, information 715 explaining the plot of the video to be created according to the editing plan, and a create video button 716 for receiving an instruction to generate an edited video in accordance with the editing plan.
[0095] The order 713 is text information that describes the editing policy (including constraints and composition conditions) in natural language. For example, the order 713 may include restrictions or guidelines on the length of the video to be generated, designations of video, still images, and audio to be added to the edited video, or designations of visual effects to be used in the edited video.
[0096] When input is received by the edit plan creation button 714, it is accepted as an instruction to create an edit plan, and steps S104 to S107 of the video creation flow (editing policy registration) are executed. Synopsis 715 displays the synopsis of the edited video indicated by the edit plan (e.g., chapter structure, video playback time, etc.). When input is received by the video creation button 716, it is accepted as an instruction to create a video in accordance with the created edit plan, and step S107 of the video creation flow (editing policy registration) is executed.
[0097] Furthermore, the example screen 700 of the editing policy registration screen includes a registered video material list display button 701 and a new registration button 702. When input is received from the registered video material list display button 701, the screen transitions to the example screen 600 of the video material registration screen. When input is received from the new registration button 702, the screen transitions to the example screen of a new editing policy registration screen, which will be described later.
[0098] 14 is a diagram showing an example of a new editing policy registration screen. A screen example 750 of the new editing policy registration screen includes at least an editing policy name 751 for the editing policy to be registered by the user, a video file name 752 of the video material to be edited, a reference button 753 for referencing and inputting a file path indicating the storage location of the video file specified by the video file name 752, an order input field 754 for receiving specific details of the editing policy, a close button 755 for receiving an instruction to transition to the screen example 700 of the editing policy registration screen, and a register button 756 for receiving an instruction to register the editing policy.
[0099] The order input field 754 accepts free-text instructions (additional information to the prompt) regarding the content of the editing policy. Specifically, the order input field 754 accepts restrictions or guidelines for the length of the video to be generated, designations of video, still images, and audio to be added to the edited video, or designations of visual effects to be used in the edited video. For example, the order input field 754 accepts free-text instructions as the content of the editing policy, such as "It consists of three scenes, and visual effects are added to the transitions between each scene to avoid sudden changes in subject and brightness. The background music should be upbeat, and a QR code should be displayed for 10 seconds at the end of the video. The entire video should be within five minutes."
[0100] When the registration button 656 receives an instruction to register an editing policy, it executes the registration process of step S103 of the video generation flow (editing policy registration).
[0101] The above is the video creation system 1 as one embodiment of the present invention. As in the above embodiment, the video creation system 1 allows a user to create a video in a format desired by the user even if the user does not have video editing skills or does not have the equipment environment for video creation.
[0102] The present invention is not limited to the above-described embodiment. Various modifications of the above-described embodiment are possible within the scope of the technical concept of the present invention. For example, in the above-described embodiment, the video generation device 100 obtains a video editing plan using the interactive AI service 200. However, the present invention is not limited to this. For example, the video generation device 100 itself may operate a generation AI specialized for video generation to generate an edited video.
[0103] Alternatively, in the above embodiment, the video generation device 100 analyzes the speech of the raw video using the audio analysis service 300, but this is not limited to this. For example, the video generation device 100 itself may operate a generation AI specialized in audio analysis to generate time-series speech information.
[0104] Furthermore, the functions of the video production device 100 may be realized by a cloud service configured with one or more computers.
[0105] Furthermore, the technical elements of the above-described embodiments may be applied independently, or may be divided into multiple parts such as program parts and hardware parts and applied.
[0106] The present invention has been described above mainly with reference to the embodiments.
[0107] 1...Video generation system, 50...Communication path, 100...Video generation device, 110...Memory unit, 111...Material information, 112...Time-series speech information, 113...Editing policy information, 114...Command information, 115...Editing plan, 116...Edited video, 120...Processing unit, 121...Acquisition unit, 122...Analysis unit, 123...Editing planning unit, 124...Video editing unit, 140...Input / output unit, 150...Communication unit, 200...Interactive AI service, 300...Speech analysis service, 400...User terminal.
Claims
1. A video generation device having: an acquisition unit that acquires a video file; an analysis unit that acquires an analysis result including time-series speech information, which is text information obtained by transcribing speech included in the video file in chronological order; an editing plan unit that transmits the time-series speech information and command information that instructs an editing plan to be output for editing the time-series speech information in accordance with desired editing policy information to an interactive AI using a language model, and receives the editing plan from the interactive AI; and a video editing unit that edits the video file in accordance with the editing plan and generates an edited video.
2. A video generation device according to claim 1, wherein the editing planning unit includes constraints on the edited video obtained by the editing plan in the command information.
3. A video generating device according to claim 1, wherein the editing planning unit includes, in the command information, configuration information about the edited video obtained by the editing plan.
4. A video generation device according to claim 1, characterized in that the editing planning unit includes in the command information a specification of a video, still image or audio to be added to the edited video obtained by the editing plan.
5. A video generation device according to claim 1, characterized in that the editing planning unit includes in the command information a specification of visual effects to be used in the edited video obtained by the editing plan.
6. A video generation device according to claim 1, wherein the editing plan includes information for constructing the edited video by connecting together partial videos whose start and end positions on a time axis within the video file are specified, and the video editing unit generates the edited video by cutting out the partial videos from the video file and connecting them together.
7. A video generation device as described in claim 1, wherein the editing plan includes information for connecting partial videos with specified start and end positions on a time axis within the video file to create the edited video, and designation of video, still images or audio to be added before and after the partial videos, and the video editing unit generates the edited video by cutting out the partial videos from the video file, connecting them together, and adding the video, still images or audio to be added.
8. A video generation device as described in claim 1, wherein the editing plan includes information for connecting partial videos with specified start and end positions on a time axis within the video file to create the edited video, and designation of visual effects to be used at the joints of the partial videos, and the video editing unit generates the edited video by cutting out the partial videos from the video file and connecting them together, and applying the visual effects specified at the joints.
9. A video generation device according to claim 1, wherein the analysis unit fast-forward edits the speech included in the video file while maintaining the time series, and passes it to a predetermined speech-to-text conversion unit to obtain the time-series speech information.
10. A video generation device according to claim 1, wherein the analysis unit identifies the speaker of the speech included in the video file, extracts the speech for each speaker while maintaining the time series, and passes the text information obtained to a predetermined speech-to-text conversion unit to integrate it to obtain the time series speech information.
11. A video generation device according to claim 1, wherein the editing plan is described in a predetermined format language, and the editing plan unit includes definition information about the format language in which the editing plan is described in the command information.
12. A video generation method using a video generation device, wherein the video generation device includes a processor, and the processor performs the following steps: an acquisition step of acquiring a video file; an analysis step of acquiring an analysis result including time-series utterance information, which is text information obtained by transcribing utterances included in the video file in chronological order; an editing planning step of transmitting the time-series utterance information and command information instructing an interactive AI using a language model to output an editing plan for editing the time-series utterance information in accordance with desired editing policy information, and receiving the editing plan from the interactive AI; and a video editing step of editing the video file in accordance with the editing plan and generating an edited video.
13. A video generation program that causes an information processing device to generate a video, wherein the information processing device has a processor, and causes the processor to perform the following steps: an acquisition step of acquiring a video file; an analysis step of acquiring an analysis result including time-series utterance information, which is text information obtained by transcribing utterances included in the video file in chronological order; an editing planning step of transmitting the time-series utterance information and command information instructing an editing plan to edit the time-series utterance information in accordance with desired editing policy information to an interactive AI using a language model, and receiving the editing plan from the interactive AI; and a video editing step of editing the video file in accordance with the editing plan and generating an edited video.
14. A video generation system comprising a user terminal and a video generation device communicatively connected to the user terminal, wherein the video generation device has an acquisition unit that acquires video files from the user terminal via communication; an analysis unit that acquires analysis results including time-series utterance information, which is text information obtained by transcribing utterances included in the video file in chronological order; an editing planning unit that transmits the time-series utterance information and command information that instructs an editing plan to be output for editing the time-series utterance information in accordance with desired editing policy information to an interactive AI using a language model, and receives the editing plan from the interactive AI; and a video editing unit that edits the video file in accordance with the editing plan and generates an edited video.
Citation Information
Patent Citations
Content generation system, program, and recording medium
JP2006048465A
Moving picture editing program
JP2008236138A
Video digest apparatus and video editing program
JP2010011409A
Content processing system, terminal device, and program
JP2019110480A
Content generation system and content generation method
JP2020140326A