Moving image generation device, moving image generation method, moving image generation program, and moving image generation system

The video generation system addresses the lack of user customization in existing technologies by using an interactive AI to analyze audio and apply user-defined policies, enabling personalized video creation.

JP2025164419AActive Publication Date: 2025-10-30川口史睦
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024068392
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-19
Publication Date
2025-10-30
Estimated Expiration
2044-04-19

AI Technical Summary

Technical Problem

Existing video editing technologies do not consider the structure and appearance of generated videos, limiting user customization.

Method used

A video generation system that includes an acquisition unit, analysis unit, editing planning unit, and video editing unit, utilizing an interactive AI to generate videos based on user-defined editing policies, incorporating constraints, visual effects, and additional media elements.

Benefits of technology

Enables users to create videos in a format desired by them, even without video editing skills or specialized equipment, by analyzing audio content and generating edited videos based on user preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025164419000001_ABST
    Figure 2025164419000001_ABST
Patent Text Reader

Abstract

To provide a technique for generating moving images in a format desired by a user.SOLUTION: A moving image generation device includes: an acquisition unit that acquires a moving image file; an analysis unit that acquires analysis results including time-series speech information, which is text information obtained by transcribing a speech contained in the moving image file in chronological order; an editing plan unit that transmits the time-series speech information and command information indicating the output of an editing plan to edit the time-series speech information according to desired editing policy information to an interactive AI using a language model, and receives the editing plan from the interactive AI; and a moving image editing unit that edits the moving image file according to the editing plan and generates an edited moving image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a moving image generating device, a moving image generating method, a moving image generating program, and a moving image generating system. [Background technology]

[0002] Patent document 1 describes a video editing device that includes an edited video generation unit that generates edited video data of the length desired by the user and including the content of the portion that the user wants to view, in response to a user request. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-087180 Summary of the Invention [Problem to be solved by the invention]

[0004] The above technology involves inputting index information to identify each section of video data divided into predetermined lengths, thereby identifying the part that the user wants to view, and generating edited video data that includes that content. Therefore, the structure and appearance of the generated video are not necessarily taken into consideration.

[0005] An object of the present invention is to generate a moving image in a manner desired by a user. [Means for solving the problem]

[0006] The present application includes multiple means for solving at least part of the above problems, examples of which are as follows: A moving image generation device according to one aspect of the present invention is characterized by having: an acquisition unit that acquires a moving image file; an analysis unit that acquires an analysis result including time-series utterance information, which is text information obtained by transcribing utterances included in the moving image file in time series; an editing planning unit that transmits the time-series utterance information and command information that instructs an editing device to output an editing plan for editing the time-series utterance information in accordance with desired editing policy information to an interactive AI using a language model and receives the editing plan from the interactive AI; and a moving image editing unit that edits the moving image file in accordance with the editing plan and generates an edited moving image.

[0007] In the above moving image generating device, the editing planning section may include, in the command information, constraints on the edited moving image obtained by the editing plan.

[0008] In the above moving image generating device, the editing planning section may include, in the command information, configuration information about the edited moving image obtained according to the editing plan.

[0009] In the above moving image generating device, the editing planning section may include in the command information a designation of a moving image, a still image or an audio to be added to the edited moving image obtained by the editing plan.

[0010] In the above moving image generating device, the editing planning section may include, in the command information, a designation of a visual effect to be used in the edited moving image obtained according to the editing plan.

[0011] In addition, in the above-mentioned video generation device, the editing plan may include information for constructing the edited video by connecting together partial videos whose start and end positions on a time axis within the video file are specified, and the video editing unit may generate the edited video by cutting out the partial videos from the video file and connecting them together.

[0012] In addition, in the above-mentioned video generation device, the editing plan may include information for connecting partial videos that specify start and end positions on the elapsed time axis within the video file to create the edited video, and designation of video, still images, or audio to be added before and after the partial videos, and the video editing unit may generate the edited video by cutting out the partial videos from the video file, connecting them together, and adding the video, still images, or audio to be added.

[0013] In addition, in the above-mentioned video generation device, the editing plan may include information for connecting partial videos having specified start and end positions on a time axis within the video file to create the edited video, and a specification of visual effects to be used at the joints of the partial videos, and the video editing unit may generate the edited video by cutting out the partial videos from the video file, connecting them together, and applying the specified visual effects to the joints.

[0014] In addition, in the above-mentioned video generation device, the analysis unit may fast-forward edit the speech audio contained in the video file while maintaining the chronological order, and pass it on to a predetermined speech-to-text conversion unit to obtain the chronological speech information.

[0015] In addition, in the above-mentioned video generation device, the analysis unit may identify the speaker of the speech audio included in the video file, extract the speech audio for each speaker while maintaining a time series, and transfer the text information obtained to a predetermined speech-to-text conversion unit to integrate the text information to obtain the time series speech information.

[0016] In addition, in the above-mentioned video generation device, the editing plan may be described in a predetermined format language, and the editing plan unit may include definition information about the format language used to describe the editing plan in the command information.

[0017] In addition, a video generation method according to another aspect of the present invention is a video generation method using a video generation device, the video generation device including a processor, the processor performing the following steps: an acquisition step of acquiring a video file; an analysis step of acquiring an analysis result including time-series utterance information, which is text information obtained by transcribing utterances included in the video file in chronological order; an editing planning step of transmitting the time-series utterance information and command information instructing an editing plan to edit the time-series utterance information in accordance with desired editing policy information to an interactive AI using a language model and receiving the editing plan from the interactive AI; and a video editing step of editing the video file in accordance with the editing plan and generating an edited video.

[0018] In addition, a video generation program according to another aspect of the present invention is a video generation program that causes an information processing device to generate a video, the information processing device having a processor, and causing the processor to perform the following steps: an acquisition step of acquiring a video file; an analysis step of acquiring an analysis result including time-series utterance information, which is text information obtained by transcribing utterances included in the video file in chronological order; an editing planning step of transmitting the time-series utterance information and command information instructing the processor to output an editing plan for editing the time-series utterance information in accordance with desired editing policy information to an interactive AI using a language model and receiving the editing plan from the interactive AI; and a video editing step of editing the video file in accordance with the editing plan and generating an edited video.

[0019] In addition, another aspect of the present invention provides a video generation system comprising a user terminal and a video generation device communicatively connected to the user terminal, wherein the video generation device comprises an acquisition unit that acquires a video file from the user terminal via communication, an analysis unit that acquires an analysis result including time-series utterance information, which is text information obtained by transcribing utterances contained in the video file in chronological order, an editing planning unit that transmits the time-series utterance information and command information that instructs an editing plan to be written and output for editing the time-series utterance information in accordance with desired editing policy information to an interactive AI using a language model and receives the editing plan from the interactive AI, and a video editing unit that edits the video file in accordance with the editing plan and generates an edited video. [Effects of the Invention]

[0020] According to the present invention, it is possible to provide a technique for generating a video in a format desired by a user.

[0021] Problems, configurations, and effects other than those described above will become apparent from the following description of the embodiments. [Brief explanation of the drawings]

[0022] [Figure 1] 1 is a diagram illustrating an overview of a video generation system according to an embodiment. [Figure 2] 1 is a configuration diagram of a video generation system according to an embodiment. [Figure 3] FIG. 10 is a diagram illustrating an example of a data structure of material information. [Figure 4] FIG. 10 is a diagram illustrating an example of a data structure of time-series speech information. [Figure 5] FIG. 10 is a diagram illustrating an example of a data structure of editing policy information. [Figure 6] FIG. 10 is a diagram illustrating an example of a data structure of command information. [Figure 7] FIG. 10 is a diagram illustrating an example of the data structure of an editing plan. [Figure 8] FIG. 2 is a diagram illustrating an example of a hardware configuration of a moving image generating device. [Figure 9]FIG. 10 is a diagram illustrating an example of a video generation flow (video material registration). [Figure 10] FIG. 10 is a diagram illustrating an example of a video generation flow (editing policy registration). [Figure 11] FIG. 10 is a diagram showing an example of a video material registration screen. [Figure 12] FIG. 10 is a diagram showing an example of a new material registration screen. [Figure 13] FIG. 10 is a diagram showing an example of an editing policy registration screen. [Figure 14] FIG. 10 is a diagram showing an example of a new editing policy registration screen. DETAILED DESCRIPTION OF THE INVENTION

[0023] A video generation system 1 to which an embodiment according to one aspect of the present invention is applied will be described below with reference to the drawings. In the following embodiments, when necessary for convenience, the description will be divided into multiple sections or embodiments. However, unless otherwise specified, they are not unrelated to each other, and one is related to the other as a partial or complete modification, detail, supplementary explanation, etc.

[0024] Furthermore, in the following embodiments, when referring to the number of elements (including the number, numerical value, amount, range, etc.), unless otherwise specified or when it is clearly limited to a specific number in principle, it is not limited to that specific number and may be more or less than the specific number.

[0025] Furthermore, it goes without saying that in the following embodiments, the components (including element steps, etc.) are not necessarily essential unless otherwise specified or unless they are clearly considered essential in principle.

[0026] Similarly, in the following embodiments, when referring to the shapes, positional relationships, etc. of components, etc., it is intended to include those that are substantially similar or similar to those shapes, etc., unless otherwise specified or when it is considered that this is clearly not the case in principle. This also applies to the above numerical values ​​and ranges.

[0027] In addition, in all the drawings for explaining the embodiments, the same components are generally designated by the same reference numerals, and repeated explanations thereof will be omitted.

[0028] In recent years, the spread of networks and various electronic devices (personal computers, tablet devices, smartphones, etc.) has created an environment in which videos can be created and published anywhere, anytime. For example, it is becoming possible for anyone to easily shoot videos using a smartphone or other device and post them anywhere, anytime to publicly accessible social networking services (SNS) and video sharing sites. However, high-quality videos that attract attention are often the result of the time and effort of editors with specialized knowledge.

[0029] Therefore, in the embodiment according to the present invention, a video creation system 1 is available that accepts a video editing policy desired by a user and automatically creates a video in accordance with the policy. The video creation system 1 allows a user to create a video in a format desired by the user even if the user does not have video editing skills or does not have the equipment environment for video creation.

[0030] 1 is a diagram showing an overview of a video generation system according to this embodiment. In the video generation system 1, a user uses a user terminal 400 that the user uses and a group of devices that are communicably connected to the user terminal 400 via a communication path. The group of devices includes a video generation device 100, a group of devices that provide an interactive AI service 200, and a group of devices that provide a voice analysis service 300.

[0031] For example, the group of devices providing the interactive AI service 200, the group of devices providing the voice analysis service 300, and the video generation device 100 may be a cloud computer connected via the Internet, or a server device managed by the owner of the video generation device 100, the group of devices providing the interactive AI service 200, and the group of devices providing the voice analysis service 300. Furthermore, without being limited to this, a wearable device such as a user's smartwatch may be used as the user terminal 400.

[0032] When the user terminal 400 communicates with the device group (including the video generation device 100, the device group providing the interactive AI service 200, and the device group providing the voice analysis service 300), they are connected via a communication path which is a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, a mobile phone network, short-range wireless communication such as Bluetooth (registered trademark), or a combination of these. The communication path 50 may be a VPN (Virtual Private Network) on a wireless communication network such as a mobile phone network.

[0033] By using the video creation system 1, a video can be created in the format desired by the user. Specifically, the user uses the user terminal 400 to register a video material file containing recorded speech and environmental sounds and a video editing policy (1 and 2) in the video creation system 1, and then applies for video creation. The video creation device 100 extracts only the audio component from the video material and requests the audio analysis service 300 to analyze it as an audio file (3) for the video material. The audio analysis service 300 analyzes the audio file for the video material and returns an analyzed text file (4) to the video creation device 100, which associates the timing of speech with the spoken words and sentences.

[0034] The video generation device 100 sends the analyzed text file obtained from the audio analysis service 300 and command information (5) including the video editing policy to the interactive AI service 200, requesting the creation of an editing plan. At this time, the video generation device 100 does not send the actual video material files or audio files to the interactive AI service 200, but sends the analyzed text file. The interactive AI service 200 creates an editing plan that satisfies the constraints specified in the command information, the constraints on the edited video obtained by the editing plan, the composition information, the designation of video, still images or audio to be added, the designation of visual effects, etc., and returns it to the video generation device 100 as an editing plan (6).

[0035] When the video generation device 100 receives the editing plan from the interactive AI service 200, it performs video editing processing according to the editing plan, creates an edited video (7), and provides it to the user terminal 400. This allows the user to utilize the provided edited video.

[0036] 2 is a configuration diagram of a video generation system according to an embodiment. The video generation system 1 includes a video generation device 100, an interactive AI service 200 that can communicate with the video generation device 100 via a communication path 50, a voice analysis service 300, and a user terminal 400.

[0037] The moving image generating device 100 includes a storage unit 110, a processing unit 120, an input / output unit 140, and a communication unit 150, which are communicably connected to one another via a bus or the like.

[0038] The storage unit 110 includes material information 111, time-series speech information 112, editing policy information 113, command information 114, an editing plan 115, and an edited video 116.

[0039] 3 is a diagram showing an example data structure of material information. Material information 111 stores information on multiple source videos to be used in generating videos. Material information 111 includes user 111A, video title 111B, video file path 111C, description 111D, analyzed flag 111E, and analysis result 111F.

[0040] User 111A is information that distinguishes a user from other users. Video title 111B is the title of the video to be registered as material. Video file path 111C is the storage location on the file system of the video to be registered as material, or a URI (Uniform Resource Identifier). Description 111D is information that explains in natural language the content of the video to be registered as material. Analyzed flag 111E is information that indicates whether analysis by the voice analysis service 300 has been completed. Analysis result 111F is analyzed text that is information on the result of analysis by the voice analysis service 300.

[0041] 4 is a diagram showing an example of the data structure of time-series utterance information. Time-series utterance information 112 is information that stores the text of utterances made in a video in order, with the time elapsed in the video being in chronological order. Time-series utterance information 112 includes utterance start time 112A, utterance end time 112B, and utterance text (words) 112C.

[0042] Utterance start time 112A and utterance end time 112B are information specifying the start timing and end timing of an utterance made in a video by the time elapsed from the start time of the video (time in the video), respectively. Utterance text (word) 112C is a word uttered between utterance start time 112A and utterance end time 112B. However, it is not limited to a word, and may be a sentence or phrase of a certain length.

[0043] 5 is a diagram showing an example data structure of editing policy information. Editing policy information 113 is information about the editing policy of a video to be generated. Editing policy information 113 includes title 113A, content goal 113B, constraints 113C, content structure 113D, resource files 113E, and editing plan format 113F.

[0044] Title 113A is the title of the editing policy or the title of the video to be generated. Content goal 113B is information such as the image that the video to be generated aims to create and the psychological changes that the viewer is aiming to achieve (to make them feel happy or calm when they watch). Constraints 113C is information on constraints for creating the video, such as the length (playback time) of the video to be generated. Content structure 113D is information on the structure of the video to be generated, such as connecting three consecutive videos with a visual effect transition. Resource file 113E is information on the video materials to be used in the video to be generated. Editing plan format 113F is information specifying the format of the editing plan for generating the video. The format of the editing plan may be a known format or may be defined in an extensible language that complies with SGML (Standard Generalized Markup Language) or the like.

[0045] 6 is a diagram showing an example of the data structure of command information. The command information 114 is a command (prompt) for causing the interactive AI service 200 to perform processing. The command for generating a video according to this embodiment, for example, specifies the editing policy information 113 and instructs the creation of an editing plan in accordance with the editing policy, and is written in natural language.

[0046] 7 is a diagram showing an example of the data structure of an editing plan. The editing plan 115 describes editing information in a predetermined format, for example by specifying tags for components to be assigned to times within the video to be generated, to create planning information for creating a video.

[0047] An outline of the format of the editing plan according to this embodiment will be explained. First, the editing plan can broadly include three types of elements: "shot," "view," and "attach." The "shot" tag is a collection of multiple "views." A "view" specifies a source file and related information. A source file includes video (including the start and end times of the portion to be used in the video material) and images (the magnification rate and on-screen layout of the image file), and related information includes color and gradation specifications. "Attach" specifies an element to be displayed by adding it to the source specified by "view" (for an image, this includes the size, layout, and start and end times within the video to be generated; for audio, this includes the audio volume and start and end times within the video to be generated).

[0048] For example, cut editing information such as assigning an excerpt from a source video extracted by specifying a time within the source video to one of the components ("views"), playing multiple such excerpts continuously with transitions in between, and then attaching a time to display a QR code (registered trademark) to access other videos in the channel ("attaches") is described.

[0049] For example, the editing plan may include, as an editing plan, information for creating an edited video by joining together partial videos whose start and end positions on the elapsed time axis within the video files to be used as raw materials. The editing plan may also include, as an editing plan, information for creating an edited video by joining together partial videos whose start and end positions on the elapsed time axis within the video files to be used as raw materials, and specifications for videos, still images, or audio to be added before and after the partial videos. The editing plan may also include, as an editing plan, information for creating an edited video by joining together partial videos whose start and end positions on the elapsed time axis within the video files to be used as raw materials, and specifications for visual effects to be used at the joins between the partial videos.

[0050] Returning to the explanation of Fig. 2, the processing unit 120 includes an acquisition unit 121, an analysis unit 122, an editing planning unit 123, and a video editing unit .

[0051] The acquisition unit 121 acquires a video file. The analysis unit 122 acquires, from a voice analysis service, an analysis result including time-series speech information, which is text information obtained by transcribing speech included in the video file in time series. The analysis unit 122 may also fast-forward edit the speech included in the video file while maintaining the time series, and transfer the text information to a predetermined voice-to-text conversion unit (voice analysis service 300) to obtain the time-series speech information. Alternatively, the analysis unit 122 may identify speakers of speech included in the video file, extract the text information for each speaker while maintaining the time series, and transfer the text information to a predetermined voice-to-text conversion unit (voice analysis service 300) to obtain the time-series speech information.

[0052] The editing planning unit 123 transmits the time-series utterance information and command information instructing the dialogue AI (the dialogue AI service 200) using a language model to output an editing plan for editing the time-series utterance information in accordance with desired editing policy information, and receives an editing plan from the dialogue AI. The editing planning unit may also include, in the command information, constraints on the edited video obtained according to the editing plan. The editing planning unit 123 may also include, in the command information, configuration information on the edited video obtained according to the editing plan. The editing planning unit 123 may also include, in the command information, designation of a video, a still image, or audio to be added to the edited video obtained according to the editing plan. The editing planning unit 123 may also include, in the command information, designation of a visual effect to be used in the edited video obtained according to the editing plan. The editing planning unit 123 may also include, in the command information, definition information on a format language for describing the editing plan.

[0053] The video editing unit 124 edits video files in accordance with the editing plan included in the editing plan to generate an edited video. Specifically, the video editing unit 124 generates the edited video by cutting out partial videos from video material files and splicing them together. The video editing unit 124 may also generate the edited video by adding additional videos, still images, or audio. The video editing unit 124 may also generate the edited video by cutting out partial videos from video files, splicing them together, and applying designated visual effects to the splices.

[0054] The input / output unit 140 controls input and output to and from the video generating device 100. For example, the input / output unit 140 accepts various types of input, such as various contact inputs such as typing, touching, and flick input, or various types of inputs such as gaze input. The input / output unit 140 also performs output to the user. The output information includes various types of output information such as screens, presentation information, advertisements, and videos.

[0055] The communication unit 150 communicates via a communication path 50 with a group of devices that provide an interactive AI service 200, a group of devices that provide a voice analysis service 300, a user terminal 400, and other terminals that communicate via the Internet.

[0056] The interactive AI service 200 is a service that provides the functions of so-called generation AI, such as GPT and Gemini, via an API (Application Programming Interface). The interactive AI service 200 provides commands (prompts) in natural language to the generation AI, causing it to generate a desired result. In this embodiment, the generation AI generates an editing plan for generating a video.

[0057] The voice analysis service 300 performs voice analysis using known technology such as Google TTS API, etc. When the voice analysis service 300 receives an audio file, it transcribes the speech in the audio file into text and outputs, as analyzed text, the text of the speech content for each utterance included in the audio file and information specifying the start and end times of the utterance.

[0058] The user terminal 400 is a terminal used by a user. The user terminal 400 may be a smartphone terminal, a PC (Personal Computer), or the like. Furthermore, without being limited to this, the user terminal 400 may be a wearable device such as a smart watch.

[0059] 8 is a diagram showing an example of the hardware configuration of a moving image generation device. The moving image generation device 100 has a hardware configuration realized by the housing of a so-called server device, workstation, personal computer, smartphone, or tablet terminal. The moving image generation device 100 includes a processor 101, a memory 102, a storage 103, an input device 104, a display device 105, a communication device 106, and a bus connecting each device.

[0060] The processor 101 is an arithmetic device such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit).

[0061] The memory 102 is a memory device such as a RAM (Random Access Memory).

[0062] The storage 103 is a non-volatile storage device capable of storing digital information, such as a hard disk drive, a solid state drive (SSD), or a flash memory.

[0063] The input device 104 is a device that accepts input from one or more of a keyboard, a mouse, a touch panel, and a microphone. The display device 105 is a device that displays one or more of various output devices such as an organic EL (Electro-Luminescence) display.

[0064] The communication device 106 is a network interface card (NIC) or the like that communicates with other devices via a network.

[0065] The device that provides the interactive AI service 200, the device that provides the voice analysis service 300, and the user terminal 400 also have substantially the same hardware configuration as the video generation device 100.

[0066] The processing unit 120, acquisition unit 121, analysis unit 122, editing planning unit 123, and video editing unit 124 of the above-described video generating device 100 are realized by a program that causes the processor 101 to perform processing. This program is stored in the memory 102, storage 103, or a ROM device (not shown), and is loaded onto the memory 102 for execution and executed by the processor 101.

[0067] The memory unit 110 of the moving image generating device 100 is realized by the memory 102 and the storage 103. The input / output unit 140 is realized by the input device 104 and the display device 105. The communication unit 150 is realized by the communication device 106. The above is an example of the hardware configuration of the moving image generating device 100.

[0068] The configuration of the moving image production device 100 can be further divided into more components depending on the processing content, and can also be divided so that one component performs even more processing.

[0069] Furthermore, each processing unit (processing unit 120, acquisition unit 121, analysis unit 122, editing planning unit 123, and video editing unit 124) may be constructed using dedicated hardware (ASIC, GPU, etc.) that realizes the respective functions. Furthermore, the processing of each processing unit may be executed by one piece of hardware or by multiple pieces of hardware.

[0070] Next, the operation of the video generation system 1 in this embodiment will be described.

[0071] 9 is a diagram showing an example of a video generation flow (video material registration). The video generation flow (video material registration) starts when a user requests its start in the web browser or application software (hereinafter sometimes simply referred to as the browser) of user terminal 400.

[0072] The acquisition unit 121 of the video generating device 100 generates a video material registration screen and displays it on the user terminal 400 (step S001). Specifically, the acquisition unit 121 generates a video material registration screen that allows the user to manage a list of videos that have been registered in the past. Then, the acquisition unit 121 transmits display information for the generated video material registration screen to the user terminal 400.

[0073] Then, the browser of the user terminal 400 displays a video material registration screen and sends a video material registration request to the video generation device 100, along with information including the video material file to be registered, the video title, and explanatory information (step S002).

[0074] The acquisition unit 121 acquires a moving image material file and the like (step S003). Specifically, the acquisition unit 121 registers the user, the moving image title, the moving image file, and a description in the material information 111.

[0075] Then, the analysis unit 122 performs video analysis (audio portion extraction) (step S004). Specifically, the analysis unit 122 separates and acquires audio components from the acquired video file.

[0076] Then, the analysis unit 122 performs video analysis (fast-forward audio generation) (step S005). Specifically, the analysis unit 122 fast-forward edits the audio components included in the acquired video file while maintaining the chronological order. For example, when processing a video file of an utterance from 0 minutes 15 seconds into the video when the utterance starts to 0 minutes 27 seconds into the video when the utterance ends (the utterance lasts 12 seconds), the analysis unit 122 edits the video at four times the normal speed and reduces the data size of the audio file so that the time from the start of the utterance to the end of the utterance is three seconds.

[0077] Then, the analysis unit 122 performs video analysis (audio analysis request) (step S006). Specifically, the analysis unit 122 requests the audio analysis service 300 to analyze the audio file of the fast-forwarded audio created in step S005 by transmitting it via an API or the like.

[0078] The audio analysis service 300 performs audio analysis processing on the transmitted audio file of the fast-forwarded audio (step S007). Specifically, the audio analysis service 300 generates a speech timing analyzed text of the raw video in which the speech timing and speech content of the raw video are associated and recorded, and transmits the generated text to the video production device 100.

[0079] Then, the analysis unit 122 performs video analysis (creation of time-series information) (step S007). Specifically, the analysis unit 122 stores the received analyzed text in the time-series utterance information 112, sets the analysis completion flag 111E to "completed", and stores reference information to the time-series utterance information 112 in the analysis result 111F. At this time, if the data structures of the analyzed text and the time-series utterance information 112 are different, the analysis unit 122 may convert the information of the analyzed text so that the time information is returned from the fast-forward state to the normal speed state and store it as the time-series utterance information 112, or may convert the time information so that the time information is returned from the fast-forward state to the normal speed state and then convert it into the data structure of the time-series utterance information 112 and store it.

[0080] The above is an example of the video generation flow (video material registration). According to the video generation flow (video material registration), for a video registered as video material, it is possible to obtain time-series speech information by analyzing the text information of the speech and the timing of the speech in the video.

[0081] 10 is a diagram showing an example of a video generation flow (editing policy registration). The video generation flow (editing policy registration) starts when the user requests its start on the browser of the user terminal 400.

[0082] The editing planning unit 123 of the video generation device 100 generates an editing policy registration screen and displays it on the user terminal 400 (step S101). Specifically, the editing planning unit 123 generates an editing policy registration screen that manages a list of editing policies that the user has previously registered. Then, the editing planning unit 123 transmits display information for the generated editing policy registration screen to the user terminal 400.

[0083] Then, the browser of the user terminal 400 displays an editing policy registration screen and sends an editing policy registration request to the video generation device 100 together with information including the editing policy title to be registered, the video material to be registered, and the order (step S102).

[0084] The editing planning unit 123 receives the editing policy and the like (step S103). Specifically, the editing planning unit 123 registers the editing policy title, the content goal, constraints, content structure, and editing plan format based on the order, and resource files based on the registered video material in the editing policy information 113. The editing planning unit 123 interprets the natural language written in the order to identify the content goal, constraints, content structure, and editing plan format included in the order.

[0085] Then, the editing planning unit 123 performs editing preparation (creating command information) (step S104). Specifically, the editing planning unit 123 creates command information 114. For example, the editing planning unit 123 replaces the designated portion of the editing policy data in the command information 114 described above with the content of the editing policy information 113, and generates a prompt to be passed to the interactive AI service 200.

[0086] Then, the editing planning unit 123 performs editing preparation (planning request) (step S105). Specifically, the editing planning unit 123 transmits the command information 114 created in step S104 and the speech timing analyzed text of the raw video to the interactive AI service 200 via an API or the like.

[0087] The interactive AI service 200 then performs editing planning processing in accordance with the transmitted command information (step S106). Specifically, the interactive AI service 200 uses the speech timing analyzed text of the raw video to perform cut editing, focusing on important parts and statements that are evaluated as interesting or intriguing, taking into consideration the speech content (meaning) and speech timing, and creates an editing plan to incorporate transitions and attachments according to orders to meet the specified length. The interactive AI service 200 generates the planned editing content as an editing plan in a specified format and transmits it to the video production device 100.

[0088] Then, the video editing unit 124 performs video editing (creates an edited video) in accordance with the editing plan 115 (step S107). Specifically, when the video editing unit 124 receives the transmitted editing plan, it stores it in the editing plan 115 in the storage unit 110. Then, the video editing unit 124 performs video editing (creates an edited video) in accordance with the editing plan 115, stores the edited video obtained as a result of the video editing in edited video 116 in the storage unit 110, and also transmits it to the user terminal 400. Note that the video editing unit 124 may post the edited video obtained as a result of the video editing on a website so that it can be downloaded and transmit a link to the user terminal 400, or may upload the edited video from the user terminal 400 to a video sharing site specified in advance.

[0089] The above is an example of the video generation flow (editing policy registration). According to the video generation flow (editing policy registration), video materials can be edited to obtain an edited video in accordance with the time-series speech information obtained by analyzing the video registered as video materials and the editing plan created using the editing policy. Therefore, it can be said that a video in the format desired by the user can be generated.

[0090] 11 is a diagram showing an example of a video material registration screen. A screen example 600 of the video material registration screen displays information including at least a video title 611 and explanatory information 615 for each registered video material file 610. In addition, the screen example 600 of the video material registration screen includes an editing policy display button 601 that accepts an instruction to transition to an editing policy list screen, a new registration button 602 that accepts an instruction to newly register a video material, and, for each registered video material file 610, a video file name 612, a content analysis status 613, and a delete button 614 that cancels the registration of the video material.

[0091] The content analysis status 613 is information indicating whether or not time-series speech information obtained by analyzing the speech text information and the timing of the speech in the video has been obtained for the registered video material. When an input is received by the editing policy display button 601, the screen transitions to an example of an editing policy registration screen, which will be described later. When an input is received by the new registration button 602, the screen transitions to an example of a new material registration screen, which will be described later.

[0092] 12 is a diagram showing an example of a new material registration screen. A screen example 650 of the new material registration screen includes at least a video title 651, a video file name 652, a reference button 653 for referencing and inputting a file path indicating the storage location of the video file specified by the video file name 652, a material file description input field 654, a close button 655 for receiving an instruction to transition to the video material registration screen, and a register button 656 for receiving an instruction to register the video material, for the video material to be registered by the user.

[0093] The material file description input field 654 accepts a description of the material content in free text. For example, in the case of video material, the material file description input field 654 accepts a synopsis or a description of the scene for each time in the video. When a registration button 656 accepts an instruction to register video material, it performs the registration process of step S003 of the video generation flow (video material registration).

[0094] 13 is a diagram showing an example of an editing policy registration screen. An example screen 700 of the editing policy registration screen displays, for each registered editing policy 710, at least an editing policy name 711, a delete button 712 for canceling the registration of the editing policy, an order 713 that is the specific content of the editing policy, a create editing plan button 714 for receiving instructions to create an editing plan, information 715 explaining the plot of the video to be created according to the editing plan, and a generate video button 716 for receiving instructions to generate an edited video in accordance with the editing plan.

[0095] The order 713 is text information that describes the editing policy (including constraints and composition conditions) in natural language. For example, the order 713 may include restrictions and guidelines for the length of the video to be generated, designations for video, still images, and audio to be added to the edited video, or designations for visual effects to be used in the edited video.

[0096] When input is received by the edit plan creation button 714, it is accepted as an instruction to create an edit plan, and steps S104 to S107 of the video creation flow (editing policy registration) are executed. Synopsis 715 displays the synopsis of the edited video indicated by the edit plan (for example, chapter structure, video playback time, etc.). When input is received by the video creation button 716, it is accepted as an instruction to create a video in accordance with the created edit plan, and step S107 of the video creation flow (editing policy registration) is executed.

[0097] Furthermore, the example screen 700 of the editing policy registration screen includes a registered video material list display button 701 and a new registration button 702. When input is received by the registered video material list display button 701, the screen transitions to the example screen 600 of the video material registration screen. When input is received by the new registration button 702, the screen transitions to the example screen of a new editing policy registration screen, which will be described later.

[0098] 14 is a diagram showing an example of a new editing policy registration screen. A screen example 750 of the new editing policy registration screen includes at least an editing policy name 751 for the editing policy to be registered by the user, a video file name 752 of the video material to be edited, a reference button 753 for referencing and inputting a file path indicating the storage location of the video file specified by the video file name 752, an order input field 754 for receiving specific details of the editing policy, a close button 755 for receiving an instruction to transition to the screen example 700 of the editing policy registration screen, and a register button 756 for receiving an instruction to register the editing policy.

[0099] The order input field 754 accepts free-text instructions (additional information to the prompt) regarding the content of the editing policy. Specifically, the order input field 754 accepts restrictions and guidelines for the length of the video to be generated, designations for video, still images, and audio to be added to the edited video, or designations for visual effects to be used in the edited video. For example, the order input field 754 accepts free-text instructions as the content of the editing policy, such as "It will consist of three scenes, and visual effects will be added to the transitions between each scene to avoid sudden changes in subject and brightness. The background music will be upbeat, and there will be 10 seconds to display a QR code at the end of the video. The entire video should be within five minutes."

[0100] When the registration button 656 receives an instruction to register an editing policy, it executes the registration process of step S103 of the video generation flow (editing policy registration).

[0101] The above is the video creation system 1 as one embodiment of the present invention. As in the above embodiment, the video creation system 1 allows a user to create a video in a desired format even if the user does not have video editing skills or does not have the equipment environment for video creation.

[0102] The present invention is not limited to the above-described embodiments. Various modifications of the above-described embodiments are possible within the scope of the technical concept of the present invention. For example, in the above-described embodiments, the video generation device 100 obtains a video editing plan using the interactive AI service 200. However, the present invention is not limited to this. For example, the video generation device 100 itself may operate a generation AI specialized for video generation to generate an edited video.

[0103] Alternatively, in the above embodiment, the video generation device 100 analyzes the speech of the raw video using the audio analysis service 300, but this is not limited to this. For example, the video generation device 100 itself may operate a generation AI specialized in audio analysis to generate time-series speech information.

[0104] Furthermore, the functions of the video production device 100 may be realized by a cloud service configured with one or more computers.

[0105] Furthermore, the technical elements of the above-described embodiments may be applied independently, or may be divided into multiple parts such as program parts and hardware parts and applied.

[0106] The present invention has been described above mainly with reference to the embodiments. [Explanation of symbols]

[0107] 1···Video generation system, 50···Communication path, 100··Video generation device, 110···Memory unit, 111···Material information, 112···Time-series speech information, 113···Editing policy information, 114···Command information, 115···Editing plan, 116···Edited video, 120···Processing unit, 121···Acquisition unit, 122···Analysis unit, 123···Editing plan unit, 124···Video editing unit, 140···Input / output unit, 150···Communication unit, 200···Interactive AI service, 300···Speech analysis service, 400···User terminal.

Claims

1. an acquisition unit that acquires a video file; an analysis unit that acquires an analysis result including time-series utterance information, which is text information obtained by transcribing utterances included in the video file in time series; an editing planning unit that transmits the time-series utterance information and command information instructing the system to output an editing plan for editing the time-series utterance information in accordance with desired editing policy information to an interactive AI using a language model, and receives the editing plan from the interactive AI; a video editing unit that edits the video file in accordance with the editing plan and generates an edited video; A video generation device having the above.

2. The video generation device according to claim 1, the editing planning unit includes, in the command information, constraints on the edited video obtained by the editing plan; A video generation device characterized by:

3. The video generation device according to claim 1, the editing planning unit includes, in the command information, configuration information about the edited video obtained by the editing plan; A video generation device characterized by:

4. The video generation device according to claim 1, the editing planning unit includes in the command information a designation of a moving image, a still image, or an audio to be added to the edited moving image obtained by the editing plan; A video generation device characterized by:

5. The video generation device according to claim 1, the editing planning unit includes, in the command information, a designation of a visual effect to be used in the edited video obtained according to the editing plan; A video generation device characterized by:

6. The video generation device according to claim 1, the editing plan includes information for constructing the edited video by connecting partial videos whose start and end positions on the elapsed time axis within the video file are specified, the video editing unit generates the edited video by cutting out the partial video from the video file and connecting the partial video. A video generation device characterized by:

7. The video generation device according to claim 1, The editing plan includes information for constructing the edited video by connecting partial videos, each of which has a specified start position and end position on a time axis within the video file, and designation of videos, still images, or audio to be added before and after the partial videos; the video editing unit extracts the partial video from the video file, connects them together, and adds the video, still image, or audio to be added, thereby generating the edited video; A video generation device characterized by:

8. The video generation device according to claim 1, the editing plan includes information for constructing the edited video by joining together partial videos whose start and end positions on a time axis within the video file are specified, and designation of visual effects to be used at the joins between the partial videos; the video editing unit extracts the partial videos from the video file, joins them together, and applies the designated visual effect to the joins to generate the edited video; A video generation device characterized by:

9. The video generation device according to claim 1, the analysis unit fast-forward edits the speech included in the video file while maintaining the time series, and transfers the edited speech to a predetermined speech-to-text conversion unit to obtain the time-series speech information; A video generation device characterized by:

10. The video generation device according to claim 1, the analysis unit identifies speakers of speech sounds included in the video file, extracts the speech sounds while maintaining a time series for each speaker, and transfers the extracted text information to a predetermined speech-to-text converter to integrate the text information to obtain the time-series speech information; A video generation device characterized by:

11. The video generation device according to claim 1, the editing plan is written in a predetermined format language; the editing plan unit includes, in the command information, definition information about the format language that describes the editing plan; A video generation device characterized by:

12. A video generation method using a video generation device, comprising: The video generation device includes a processor, The processor: an acquisition step of acquiring a video file; an analysis step of acquiring an analysis result including time-series utterance information, which is text information obtained by transcribing utterances included in the video file in time series; an editing planning step of transmitting the time-series utterance information and command information instructing the output of an editing plan for editing the time-series utterance information in accordance with desired editing policy information to an interactive AI using a language model, and receiving the editing plan from the interactive AI; a video editing step of editing the video file in accordance with the editing plan to generate an edited video; A video generation method that implements the above.

13. A moving image generation program that causes an information processing device to generate a moving image, The information processing device includes a processor, the processor, an acquisition step of acquiring a video file; an analysis step of acquiring an analysis result including time-series utterance information, which is text information obtained by transcribing utterances included in the video file in time series; an editing planning step of transmitting the time-series utterance information and command information instructing the user to output an editing plan for editing the time-series utterance information in accordance with desired editing policy information to an interactive AI using a language model, and receiving the editing plan from the interactive AI; a video editing step of editing the video file in accordance with the editing plan to generate an edited video; A video generation program that performs the above.

14. A video creation system including a user terminal and a video creation device communicably connected to the user terminal, The video generation device an acquisition unit that acquires a video file from the user terminal via communication; an analysis unit that acquires an analysis result including time-series utterance information, which is text information obtained by transcribing utterances included in the video file in time series; an editing planning unit that transmits the time-series utterance information and command information instructing the system to output an editing plan for editing the time-series utterance information in accordance with desired editing policy information to an interactive AI using a language model, and receives the editing plan from the interactive AI; a video editing unit that edits the video file in accordance with the editing plan and generates an edited video; A video generation system characterized by:

Citation Information

Patent Citations

  • Content generation system, program, and recording medium

    JP2006048465A

  • Video digest apparatus and video editing program

    JP2010011409A

  • Content processing system, terminal device, and program

    JP2019110480A

  • Content generation system and content generation method

    JP2020140326A

  • Method and Apparatus for Generating Video Descriptions

    US20130120654A1