Video editing system and video editing method

The video editing system simplifies video editing by using a learning model to divide and transcribe audio sections and detect scene changes, providing a user-friendly interface for efficient editing without requiring high-spec hardware.

JP7747863B1Active Publication Date: 2025-10-01J STREAM
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024228976
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-10-01
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

Conventional video editing devices require dedicated software with numerous functions, leading to slow performance, complex user interfaces, and the need for high-spec hardware, making video editing cumbersome and time-consuming.

Method used

A video editing system utilizing a learning model to divide videos into voiced and unvoiced sections, transcribe audio, detect scene changes, and provide an editing interface with objects representing these elements, allowing for simple editing operations without high-spec hardware.

Benefits of technology

Enables efficient video editing with a user-friendly interface and simplified operations, reducing the need for high-performance hardware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007747863000001_ABST
    Figure 0007747863000001_ABST
Patent Text Reader

Abstract

Edit videos with simple operations without requiring high-spec hardware. [Solution] A video editing system 1 inputs a target video to be edited into a video editing server 2, and uses a learning model that has machine-learned the division of voiced and unvoiced sections in the video, the division of phrase sections, and the detection of scene change points to divide the audio data for the target video into voiced and unvoiced sections, transcribe the voiced sections and divide them into phrase sections, detect scene change points, and provide the user with an editing screen 100 on which phrase section objects 103, unvoiced section objects 104, and scene change point objects 105 are arranged as a video audio document 102. On the editing screen 100, the video audio document 102 is edited in accordance with operations on the phrase section objects 103, unvoiced section objects 104, and scene change point objects 105, and an edited video is created by editing the target video in accordance with the editing results of the video audio document 102.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a video editing system and a video editing method for editing a video. [Background technology]

[0002] 2. Description of the Related Art Conventionally, information processing devices such as personal computers function as video editing devices by being equipped with software (programs) for editing videos.

[0003] For example, the video editing device disclosed in Patent Document 1 includes an edited video generation unit that generates edited video data in accordance with a user's request, the edited video data having the length desired by the user and including the content of the portion that the user wants to view. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent Publication No. 2021-087180 Summary of the Invention [Problem to be solved by the invention]

[0005] However, conventional video editing devices require dedicated software for editing videos (video editing software) to be installed on an information processing device such as a personal computer. The dedicated video editing software installed on the video editing device is versatile and has numerous functions, which makes the video editing device slow and unresponsive when editing videos using the video editing software. Furthermore, because such video editing software provides many functions, the user interface becomes complex and editing videos requires many steps and time. Furthermore, the video editing device requires high-spec hardware to be installed on the information processing device, such as a personal computer, that runs the video editing software.

[0006] In consideration of the above circumstances, the present invention aims to provide a video editing system and a video editing method that enable video editing with simple operations without requiring high-spec hardware. [Means for solving the problem]

[0007] In order to solve the above problem, the video editing system of the present invention inputs a target video to be edited, and uses a learning model that has machine-learned the division of voiced and unvoiced sections in the video, the division of phrase sections within the voiced sections, and the detection of scene change points in the video to divide the audio data of the target video into voiced and unvoiced sections, transcribe the voiced sections of the audio data and divide them into phrase sections, detect scene change points from the target video, and provide the user with an editing screen in which phrase section objects indicating phrase sections of the data, silent section objects indicating silent sections of the audio data, and scene change point objects indicating scene change points of the target video are arranged in chronological order as a series of video and audio documents, and edits the video and audio documents in accordance with operations on the phrase section objects, silent section objects, and / or scene change point objects on the editing screen, and creates an edited video by editing the target video in accordance with the editing results of the video and audio documents.

[0008] In addition, in order to solve the above-mentioned problems, the video editing method of the present invention inputs a target video to be edited, and uses a learning model that has machine-learned the division of voiced and unvoiced sections in the video, the division of phrase sections within the voiced sections, and the detection of scene change points in the video to divide the audio data of the target video into voiced and unvoiced sections, transcribe the voiced sections of the audio data and divide them into phrase sections, detect scene change points from the target video, and provide the user with an editing screen in which phrase section objects indicating phrase sections of the data, silent section objects indicating silent sections of the audio data, and scene change point objects indicating scene change points of the target video are arranged in chronological order as a series of video and audio documents, and on the editing screen, edit the video and audio documents in accordance with operations on the phrase section objects, the silent section objects, and / or the scene change point objects, and create an edited video by editing the target video in accordance with the editing results of the video and audio documents. [Effects of the Invention]

[0009] According to the present invention, it is possible to provide a video editing system and a video editing method that allow video editing with simple operations without requiring high-spec hardware. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a block diagram illustrating a video editing system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a schematic diagram showing an example of an editing screen provided by a video editing server in the video editing system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0011] First, the overall configuration of a video editing system 1 according to an embodiment of the present invention will be described with reference to Fig. 1. As shown in Fig. 1, the video editing system 1 includes a video editing server 2 that executes video editing and a terminal device 3 for performing video editing operations, and the video editing server 2 and the terminal device 3 are communicably connected via a network 5 such as the Internet. The video editing system 1 is a system in which, with a video including video data and audio data as a target video to be edited, the video editing server 2 edits the target video in accordance with editing operations on the target video on the terminal device 3 and provides the edited target video to the terminal device 3.

[0012] Although FIG. 1 shows an example in which the video editing system 1 has one terminal device 3 for one video editing server 2, a plurality of terminal devices 3 may be provided for one video editing server 2.

[0013] The video editing server 2 includes, for example, a control unit 10, a storage unit 11, a communication unit 12, a display unit 13, and an operation unit 14.

[0014] The control unit 10 controls the various units and functions of the video editing server 2, and is configured by a computer such as a CPU (Central Processing Unit), and is connected to a storage unit 11 and a communication unit 12. The storage unit 11 is configured to have memories such as a ROM (Read Only Memory) and a RAM (Random Access Memory), and a recording medium such as a hard disk, and stores programs and data for controlling the various units and functions of the video editing server 2. The communication unit 12 is an interface for connecting the video editing server 2 to the network 5, and in other words, connects the video editing server 2 to the terminal device 3 via the network 5.

[0015] The control unit 10 controls various components and functions of the video editing server 2 by executing arithmetic processing based on programs and data stored in the storage unit 11. For example, the control unit 10 operates as a video input unit 20, a section detection unit 21, a translation unit 22, an editing screen providing unit 23, and a video editing unit 24 by executing a program stored in the storage unit 11. The video input unit 20, the section detection unit 21, the translation unit 22, the editing screen providing unit 23, and the video editing unit 24 realize the video input step, the section detection step, the translation step, the editing screen providing step, and the video editing step of the video editing method according to the present invention.

[0016] Video input unit 20 inputs a target video to be edited that has been transmitted or transferred (uploaded) from terminal device 3 and stores it in storage unit 11. For example, video input unit 20 provides a website via network 5 where the target video can be uploaded, and inputs the uploaded target video via this website displayed on a browser installed in terminal device 3. Video input unit 20 may accept the input (upload) of the target video via a website that displays an editing screen 100 (see FIG. 2 ), which will be described later, for editing the target video, or may accept the input (upload) of the target video via a website that displays an input screen (not shown) dedicated to inputting the target video.

[0017] The section detection unit 21 divides the audio data of the target video input by the video input unit 20 into voiced sections consisting of successive voiced sounds and unvoiced sections consisting of successive unvoiced sounds, and stores the time information of the voiced and unvoiced sections in the target video (the playback start time and playback end time in the target video) in memory unit 11, associating them with the voiced and unvoiced sections, respectively.

[0018] Furthermore, the section detection unit 21 transcribes the voiced sections of the audio data, extracts words, classifies the minimum set of words that constitutes a translatable sentence as a phrase section, and associates time information of the phrase section in the target video (playback start time and playback end time in the target video) with the phrase section and stores it in the storage unit 11. For example, the section detection unit 21 classifies the target video into phrase sections and unvoiced sections in WebVTT units.

[0019] Furthermore, the section detection unit 21 detects scene change points where scenes change from the target moving image, and stores time information of the scene change points in the target moving image (scene change times in the target moving image) in the storage unit 11 in association with the scene change points.

[0020] The video editing server 2 can use a learning model that has been machine-learned by artificial intelligence to classify voiced and unvoiced sections in a video, classify phrase sections in the voiced sections, and detect scene change points in the video, regardless of whether it is an input target video or a video that can be acquired via the network 5, and can use artificial intelligence such as a large-scale language model (LLM) that uses this learning model. The video editing server 2 may function as an artificial intelligence server that creates, stores, and uses this learning model itself, or an external artificial intelligence server may create, store, and use this learning model.

[0021] Then, the section detection unit 21 uses artificial intelligence that utilizes this learning model to divide the audio data of the target video into voiced and unvoiced sections. Furthermore, the section detection unit 21 also uses artificial intelligence that utilizes this learning model to transcribe the voiced sections of the audio data and divide the voiced sections into phrase sections. Furthermore, the section detection unit 21 also uses artificial intelligence that utilizes this learning model to detect scene change points from the target video.

[0022] The translation unit 22 translates phrase sections in the voiced sections of the audio data of the target moving image detected by the section detection unit 21, and stores the translation results in the storage unit 11 in association with the phrase sections.

[0023] For example, the translation unit 22 may use an external translation server to translate each phrase section of the target video. In this case, regardless of whether the target video is an input video or a video obtainable via the network 5, the translation unit 22 causes the external translation server to translate the phrase sections of the voiced sections in the video, and stores a learning model obtained by machine learning the translation results using artificial intelligence in the external translation server. The translation unit 22 then causes the artificial intelligence of the external translation server to translate each phrase section of the target video using this learning model.

[0024] Alternatively, the translation unit 22 may translate each phrase section of the target video through internal processing of the video editing server 2. In this case, the translation unit 22, as internal processing of the video editing server 2, generates a learning model by machine learning the translation results of phrase sections in voiced sections in the video using artificial intelligence, regardless of whether the target video has been input or the video can be acquired via the network 5, and stores the model in the storage unit 11 or a database (not shown). Then, the translation unit 22 translates each phrase section of the target video using artificial intelligence that uses this learning model.

[0025] The editing screen providing unit 23 provides the terminal device 3 with an editing screen 100 as shown in FIG. 2 for editing the target video input by the video input unit 20. For example, the editing screen providing unit 23 provides, via the network 5, a website showing the editing screen 100 on which editing operations for the target video are possible, and accepts editing operations for the target video via this website displayed on a browser installed in the terminal device 3. The editing screen 100 provided by the editing screen providing unit 23 has an edit confirmation button 101, and in response to operation of the edit confirmation button 101, the editing screen providing unit 23 creates editing information for the target video corresponding to the editing operations for the target video and stores it in the storage unit 11.

[0026] Specifically, on the editing screen 100, the editing screen providing unit 23 arranges in chronological order phrase section objects 103 indicating phrase sections of the audio data detected by the section detection unit 21, silent section objects 104 indicating silent sections of the audio data, and scene change point objects 105 indicating scene change points of the target video, and displays them as a series of video audio documents 102 (text) that are like documents (texts) of the audio data of the target video, so that each object can be selected and operated.

[0027] The editing screen providing unit 23 accepts editing of the video / audio document 102 (transcript) consisting of phrase section objects 103, silent section objects 104, and scene change point objects 105 on the editing screen 100, and the editing of the video / audio document 102 on the editing screen 100 becomes provisional editing of the target video. The editing screen providing unit 23 confirms the editing of the video / audio document 102 in response to operation of the edit confirmation button 101, and creates editing information for the target video corresponding to the editing of the video / audio document 102.

[0028] The phrase section object 103, the silent section object 104, and the scene change point object 105 displayed on the editing screen 100 are each individual GUI buttons, and the editing screen providing unit 23 accepts drag-and-drop operations of the phrase section object 103, the silent section object 104, and the scene change point object 105 on the editing screen 100.

[0029] The characters of the corresponding phrase section are displayed in the phrase section object 103. The silent section object 104 and the scene change point object 105 are displayed with marks or the like that allow them to be distinguished from other objects; for example, the silent section object 104 is displayed with a green circular mark (shown as a hollow circle in FIG. 2), and the scene change point object 105 is displayed with a red circular mark (shown as a black circle in FIG. 2).

[0030] For example, the editing screen providing unit 23 accepts a delete operation that specifies a phrase section object 103, a silent section object 104, and a scene change point object 105 in order to delete a phrase section, a silent section, or a scene change point from the target video. The editing screen providing unit 23 accepts a position change operation that specifies a phrase section object 103, a silent section object 104, and a scene change point object 105 and a change position in order to change the position of a phrase section, a silent section, or a scene change point in the target video. The editing screen providing unit 23 accepts a split operation that specifies one phrase section object 103 and a split position in order to split one phrase section into two or more phrase sections. The editing screen providing unit 23 accepts an combine operation that specifies two or more phrase section objects 103 in order to combine two or more phrase sections into one phrase section.

[0031] Furthermore, the editing screen providing unit 23 accepts an event addition operation for the phrase section object 103, the silent section object 104 and / or the scene change point object 105 in order to add a predetermined event to a phrase section of the audio data, a silent section of the audio data and / or a position of the target video corresponding to a scene change point of the target video on the editing screen 100, and edits the video audio document 102 in accordance with the event addition operation. For example, the editing screen providing unit 23 displays a GUI button for an event object 106 on the editing screen 100 to accept various event addition operations, and accepts a drag-and-drop operation of the event object 106.

[0032] At this time, the editing screen providing unit 23 may accept, on the editing screen 100, a designation operation for an object to which an event is to be added (the phrase section object 103, the silent section object 104, or the scene change point object 105) and a designation operation for a period to which the event is to be added (i.e., a playback period during which the event is to be played). Furthermore, when accepting an event addition operation for the phrase section object 103 and / or the silent section object 104, the editing screen providing unit 23 may accept, on the editing screen 100, a designation operation for an addition time in the phrase section and / or the silent section (the start time and / or end time of the event in the phrase section and / or the silent section).

[0033] For example, the editing screen providing unit 23 accepts an event adding operation for specifying and inserting a title (such as a target video or an eye-catching video showing each scene), subtitles (such as Japanese subtitles or English subtitles), captions (such as text overlaid on video), audio (such as narration), music, images, and scene transitions (such as fade-in and fade-out) as events on the editing screen 100. The editing screen providing unit 23 may also accept an event adding operation for inserting other events, such as double speed or normalization, into the phrase section object 103 or the silent section object 104 on the editing screen 100.

[0034] The video editing server 2 (section detection unit 21) may analyze the video of each scene of the target video segmented by scene change points, and predict candidates for events (title, subtitle, caption, sound, music, image, scene change) to be inserted into each scene using artificial intelligence. The editing screen providing unit 23 may provide the predicted candidate events on the editing screen 100.

[0035] Furthermore, the editing screen providing unit 23 may display on the editing screen 100 a scene screen 110 such as a thumbnail or preview corresponding to the video data of the target image in the phrase section, silent section, or scene change point for the object (phrase section object 103, silent section object 104, scene change point object 105) selected in the video audio document 102.

[0036] When the editing of the video / audio document 102 is confirmed in response to the operation of the edit confirmation button 101 on the editing screen 100 provided by the editing screen providing unit 23, the video editing unit 24 transcodes the target video, creates an edited video by editing the target video in response to the editing information of the target video, and stores the edited video in the storage unit 11. The video editing unit 24 stores the created edited video in the storage unit 11, and also transmits or transfers (downloads) it to the terminal device 3 for output.

[0037] For example, when a deletion operation of a phrase section object 103, a silent section object 104, or a scene change point object 105 is confirmed on the editing screen 100, the video editing unit 24 edits the target video by deleting the corresponding phrase section, silent section, or scene change point.

[0038] When the position change operation of the phrase section object 103, the silent section object 104, or the scene change point object 105 on the editing screen 100 is confirmed, the video editing unit 24 changes the positions of the corresponding phrase section, silent section, or scene change point to the specified change positions and edits the target video.

[0039] When a dividing operation or a merging operation of the phrase section object 103 is confirmed on the editing screen 100, the video editing unit 24 divides or merges the corresponding phrase sections and edits the target video so as to make them WebVTT units.

[0040] When an event addition operation is confirmed for a phrase section object 103, a silent section object 104, or a scene change point object 105 on the editing screen 100, the video editing unit 24 adds an event to the corresponding phrase section, silent section, or scene change point to edit the target video.

[0041] The video editing unit 24 may automatically send or transfer the edited video to the terminal device 3 upon completion of creation of the edited video, or may transfer the edited video to the terminal device 3 in response to a download operation on the terminal device 3.

[0042] For example, the video editing unit 24 may provide a website through the network 5 where the edited video can be downloaded, and the edited video may be downloaded via this website displayed on a browser installed in the terminal device 3. The video editing unit 24 may accept the output (download) of the edited video via a website that displays the editing screen 100, or may accept the output (download) of the edited video via a website that displays an output screen dedicated to outputting the edited video.

[0043] The terminal device 3 is a device for a user to perform video editing operations, and is configured, for example, as a personal computer, a smartphone, a tablet terminal, etc. The terminal device 3 is configured, for example, to include a control unit 30, a storage unit 31, a communication unit 32, a display unit 33, and an operation unit 34.

[0044] The control unit 30 controls the various units and functions of the terminal device 3, and is configured as a computer such as a CPU, and is connected to a storage unit 31 and a communication unit 32. The storage unit 31 is configured to have memories such as ROM and RAM, and recording media such as a hard disk, and stores programs and data for controlling the various units and functions of the terminal device 3. The storage unit 31 stores target videos to be edited and edited videos edited by the video editing server 2. The communication unit 32 is an interface for connecting the terminal device 3 to the network 5, and in other words, connects the terminal device 3 to the video editing server 2 via the network 5.

[0045] Particularly in this embodiment, the terminal device 3 is equipped with a browser, and uses the browser to display on the display unit 33 an editing screen 100 provided by the video editing server 2 connected via the network 5. The terminal device 3 can use the operation unit 34 to operate uploading the target video, editing the video / audio document 102 of the target video, and downloading the edited video via the editing screen 100 displayed on the display unit 33 by the browser.

[0046] In this embodiment, as described above, the video editing system 1 inputs a target video to be edited by the video input unit 20 in the video editing server 2, and the section detection unit 21 uses a learning model that has been machine-learned to classify voiced and unvoiced sections in the video, classify phrase sections in the voiced sections, and detect scene change points in the video to classify the audio data of the target video into voiced and unvoiced sections, transcribe the voiced sections of the audio data and classify them into phrase sections, detect scene change points from the target video, and the editing screen providing unit 23 extracts the sentences of the audio data of the target video using the learning model. An editing screen 100 is provided to the user in which a series of video and audio documents 102 are arranged in chronological order, including clause section objects 103 indicating clause sections, silent section objects 104 indicating silent sections of the audio data of the target video, and scene change point objects 105 indicating scene change points of the target video.The user edits the video and audio document 102 on the editing screen 100 in response to operations on the clause section objects 103, silent section objects 104 and / or scene change point objects 105, and a video editing unit 24 creates an edited video by editing the target video in response to the editing results of the video and audio document 102.

[0047] In other words, the video editing method of the present invention inputs a target video to be edited, and uses a learning model that has machine-learned the division of voiced and unvoiced sections in the video, the division of phrase sections within the voiced sections, and the detection of scene change points in the video to divide the audio data of the target video into voiced and unvoiced sections, transcribe the voiced sections of the audio data of the target video and divide them into phrase sections, detect scene change points from the target video, and provide the user with an editing screen 100 in which phrase section objects 103 indicating phrase sections of the audio data of the target video, silent section objects 104 indicating silent sections of the audio data of the target video, and scene change point objects 105 indicating scene change points of the target video are arranged in chronological order as a series of video and audio documents 102, and edits the video and audio document 102 on the editing screen 100 in accordance with operations on the phrase section objects 103, silent section objects 104 and / or scene change point objects 105, and creates an edited video by editing the target video in accordance with the editing results of the video and audio document 102.

[0048] With this configuration, according to this embodiment, on the terminal device 3 performing the editing operation of the target video, the editing results of the video / audio document 102, which includes phrase section objects 103 indicating phrase sections of the audio data of the target video, silent section objects 104 indicating silent sections of the audio data of the target video, and scene change point objects 105 indicating scene change points of the target video, are reflected directly in the editing of the target video on the editing screen 100, and an edited video is created. Therefore, on the terminal device 3, editing points can be easily identified during video editing, and a user interface (UI) / user experience (UX) that is visually and operationally simple can be used. In this way, the video editing system 1 of the present invention allows video editing with simple operations without requiring high-spec hardware.

[0049] In addition, in the video editing system 1 of this embodiment, the editing screen providing unit 23 in the video editing server 2 accepts an event addition operation on the editing screen 100 for the phrase section object 103, the silent section object 104 and / or the scene change point object 105 in order to add a predetermined event to a position of the target video corresponding to a phrase section of the audio data, a silent section of the audio data and / or a scene change point of the target video, and the video editing unit 24 creates an edited video by editing the target video in accordance with the event addition operation.

[0050] With this configuration, according to this embodiment, the target video can be edited so that events are added to the phrase section object 103, the silent section object 104, and the scene change point object 105 of the video / audio document 102 simply by performing an event addition operation on the editing screen 100.

[0051] In the above embodiment, an example was described in which, in response to operation of the edit confirmation button 101 on the editing screen 100, the editing screen providing unit 23 confirms the editing of the video / audio document 102 and creates editing information for the target video corresponding to the editing of the video / audio document 102, and the video editing unit 24 creates an edited video by editing the target video in accordance with the editing information of the target video, but the present invention is not limited to this example.

[0052] In another example, in response to operation of the edit confirmation button 101 on the editing screen 100, the editing of the video / audio document 102 is confirmed on the editing screen 100 of the terminal device 3, editing information for the target video corresponding to the editing of the video / audio document 102 is created, the editing information for the target video is sent from the terminal device 3 to the video editing server 2, and the video editing unit 24 creates an edited video by editing the target video in accordance with the editing information for the target video received from the terminal device 3.

[0053] The present invention can be modified as appropriate within the scope that does not contradict the gist or idea of ​​the invention that can be read from the claims and the entire specification, and video editing systems and video editing methods that involve such modifications are also included in the technical idea of ​​the present invention. [Explanation of symbols]

[0054] 1. Video editing system 2. Video editing server 3 Terminal Devices 5. Network 10 Control Unit 11 Storage section 12 Communications Department 13 Display section 14 Control section 20 Video input section 21 Section detection unit 22 Translation Department 23 Editing Screen Provider 24 Video Editing Department 30 Control Unit 31 Storage section 32 Communications Department 33 Display section 34 Control section 100 Editing screen 101 Edit confirmation button 102 Video and audio documents 103 Phrase Section Object 104 Silent Interval Objects 105 Scene Change Point Object 106 Event Objects 110 Scene Screen

Claims

1. Enter the target video to be edited, Using a learning model that has been machine-learned to classify voiced and unvoiced sections in a video, classify phrase sections within the voiced sections, and detect scene change points in the video, the audio data of the target video is classified into voiced and unvoiced sections, the voiced sections of the audio data are transcribed and classified into phrase sections, and scene change points are detected from the target video; providing a user with an editing screen in which phrase section objects indicating phrase sections of the audio data, silent section objects indicating silent sections of the audio data, and scene change point objects indicating scene change points of the target video are arranged in chronological order as a series of video and audio documents; A video editing system characterized by editing the video and audio document in accordance with operations on the editing screen of the phrase section object, the silent section object and / or the scene change point object, and creating an edited video by editing the target video in accordance with the editing results of the video and audio document.

2. The video editing system according to claim 1, characterized in that, on the editing screen, an event addition operation is accepted for the phrase section object, the silent section object and / or the scene change point object in order to add a predetermined event to a position of the target video corresponding to a phrase section of the audio data, a silent section of the audio data and / or a scene change point of the target video, and the edited video is created by editing the target video in accordance with the event addition operation.

3. Enter the target video to be edited, Using a learning model that has been machine-learned to classify voiced and unvoiced sections in a video, classify phrase sections within the voiced sections, and detect scene change points in the video, the audio data of the target video is classified into voiced and unvoiced sections, the voiced sections of the audio data are transcribed and classified into phrase sections, and scene change points are detected from the target video; providing a user with an editing screen in which phrase section objects indicating phrase sections of the audio data, silent section objects indicating silent sections of the audio data, and scene change point objects indicating scene change points of the target video are arranged in chronological order as a series of video and audio documents; A video editing method characterized by editing the video and audio document in accordance with operations on the editing screen of the phrase section object, the silent section object and / or the scene change point object, and creating an edited video by editing the target video in accordance with the editing results of the video and audio document.

Citation Information

Patent Citations

  • Moving picture reproducing apparatus, moving picture reproducing method, and its computer program

    JP2003309814A

  • Motion picture reproducing apparatus and method, and computer program therefor

    JP2008118688A

  • Method, device and program for audio search

    JP2011069845A

  • Video analysis device, video analysis method, and video analysis program

    JP2014039137A

  • Moving image editing device, moving image editing method, and computer program

    JP2021087180A