Method for automatically editing video through detection of screen switching time point and recommendation of to-be-connected camera images
The automatic video editing method addresses the time-consuming process of manual screen transition detection and camera selection by using AI to detect transitions and recommend camera images, significantly reducing editorial effort and time in video content production.
Patent Information
- Application Number
- PCT/KR2024/011359
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-11
- Filing Date
- 2024-08-02
- Publication Date
- 2025-06-19
AI Technical Summary
The existing video editing process requires significant effort and time from editors to manually detect screen transition points and select camera footage for editing, especially when dealing with multi-channel footage from various angles.
An automatic video editing method that uses artificial intelligence to detect screen transition points in multi-channel camera images and recommends new camera images to be connected at these points, reducing the need for manual editing.
This method drastically reduces the work effort and time required for editors to produce video content by automating the detection of screen transitions and recommendation of camera images, thereby improving efficiency in video content production.
Smart Images

Figure KR2024011359_19062025_PF_FP_ABST
Abstract
Description
Automatic video editing method using screen transition point detection and connected camera video recommendation
[0001] The present invention relates to a video editing technology, and more particularly, to a video editing method for producing a single video content by selectively connecting camera images acquired by shooting a shooting scene from multiple angles.
[0002] In order to produce video content, the shooting site must first be filmed from multiple angles to acquire multi-channel footage, and then the editor must repeatedly review all of the acquired multi-channel footage, decide on screen transitions when deemed necessary, and then select the camera footage to be connected.
[0003] Because screen transitions require an understanding of the situation, such as changes in content, changes in the person speaking, and the duration of the main camera, the editor must carefully review the footage and make a judgment call.
[0004] However, this editing process in video content production requires a lot of effort and time from the editor, which is a major obstacle to video content production. This difficulty increases in proportion to the shooting time and the number of camera channels.
[0005] The present invention has been devised to solve the above problems, and the purpose of the present invention is to provide a video editing method that automatically detects screen transition points by analyzing multi-channel camera images acquired by shooting a shooting site from various angles based on artificial intelligence, and automatically recommends new camera images to be connected at the detected screen transition points, as a means of reducing the work effort and time required for an editor to produce video content.
[0006] In order to achieve the above object, a video automatic editing method according to one embodiment of the present invention comprises the steps of: receiving a plurality of camera images of the same scene; detecting a screen transition point from the input camera images; and recommending camera images that will compose a video from the detected screen transition point to the next screen transition point.
[0007] The input stage may be receiving multiple camera images taken from different angles of the shooting scene.
[0008] The detection step may be to detect whether a screen transition occurs for each camera image, and if a screen transition is detected in a number of camera images greater than or equal to a predetermined number, the detection may be performed at the time of a screen transition.
[0009] The automatic video editing method according to the present invention further includes a step of extracting features from camera images; and the detection step may be detecting whether a screen is switched from the extracted features.
[0010] The detection step may be to detect whether a screen has been switched from the extracted features using a machine learning model trained to estimate whether a screen has been switched from the extracted features.
[0011] Features may include conversational features, behavioral features, reaction features, and background features.
[0012] The recommendation step may be to recommend camera images using a machine learning model trained to recommend new camera images to be connected at the screen transition point by inputting features extracted from camera images before the screen transition point and features extracted from camera images after the screen transition point.
[0013] The recommendation step may be to recommend multiple camera images based on the accuracy of the machine learning model.
[0014] The automatic video editing method according to the present invention may further include a step of generating automatic video content editing information by synthesizing detected screen transition points in the entire section of the video and synthesizing recommended camera image information at the screen transition points.
[0015] According to another aspect of the present invention, a video automatic editing system is provided, comprising: an input unit for receiving a plurality of camera images of the same scene; a detection unit for detecting a screen transition point from the input camera images; and a recommendation unit for recommending camera images to compose a video from the detected screen transition point to the next screen transition point.
[0016] According to another aspect of the present invention, a video automatic editing method is provided, comprising: a step of detecting a screen transition point from multiple camera images of the same scene; a step of recommending camera images that will compose an image from the detected screen transition point to the next screen transition point; and a step of generating automatic editing information by synthesizing the detected screen transition points and recommended camera image information.
[0017] According to another aspect of the present invention, a video automatic editing system is provided, comprising: a detection unit for detecting a screen transition point from multiple camera images of the same scene; a recommendation unit for recommending camera images to compose an image from the detected screen transition point to the next screen transition point; and a generation unit for generating automatic editing information by synthesizing the detected screen transition points and recommended camera image information.
[0018] As described above, according to embodiments of the present invention, multi-channel camera images acquired by shooting a shooting scene from various angles are analyzed based on artificial intelligence to automatically detect screen transition points, and camera images to be newly connected are automatically recommended at the detected screen transition points, thereby drastically reducing the work effort and time required for an editor to produce video content.
[0019] Figure 1 is a video automatic editing system according to one embodiment of the present invention;
[0020] Figure 2 is a video automatic editing method according to another embodiment of the present invention.
[0021] Hereinafter, the present invention will be described in more detail with reference to the drawings.
[0022] An embodiment of the present invention presents a method for automatically editing a video by detecting screen transition points and recommending connected camera footage. This technology automatically detects screen transition points by analyzing multi-channel camera footage captured from various angles of a shooting scene using artificial intelligence, and automatically recommends new camera footage to be connected at the detected screen transition points.
[0023] Unlike the existing method in which the screen transition point and the connected camera image were determined solely by the editor's judgment, in the embodiment of the present invention, these tasks are performed automatically, thereby drastically reducing the editor's work effort and time required for producing video content.
[0024] FIG. 1 is a diagram illustrating the configuration of an automatic video editing system according to one embodiment of the present invention. The automatic video editing system according to the embodiment of the present invention is configured to include an image input unit (110), an image feature extraction unit (120), a screen transition detection unit (130), a connection camera recommendation unit (140), and an editing information generation unit (150).
[0025] The video input unit (110) receives and inputs multiple camera videos taken from different angles of the shooting site through multiple channels.
[0026] The image feature extraction unit (120) extracts features from multi-channel camera images input from the image input unit (110). Feature extraction is performed on a per-camera image basis. If the camera images are N-channel, features are extracted from the camera images of the first channel, features are extracted from the camera images of the second channel, ..., features are extracted from the camera images of the Nth channel.
[0027] Meanwhile, the extracted features include conversation features, behavioral features, and reaction features, and background features can be included as needed.
[0028] Feature extraction can be performed using a machine learning model trained to extract features from input camera images. These machine learning models can be implemented separately for each feature. For example, a machine learning model could be configured to extract conversational features, a machine learning model to extract behavioral features, a machine learning model to extract reaction features, and a machine learning model to extract background features.
[0029] The screen transition detection unit (130) detects the screen transition point based on the features extracted by the image feature extraction unit (120). Specifically, the screen transition detection unit (130) detects whether a screen transition occurs for each multi-channel camera image, and if a screen transition is detected in a predetermined number or more, for example, more than half of the camera images, it is detected as the screen transition point.
[0030] Whether or not a screen has been switched can be detected using a machine learning model trained to estimate whether or not a screen has been switched by inputting features extracted by the image feature extraction unit (120) in a time series.
[0031] The connection camera recommendation unit (140) recommends camera images that will compose the video from the screen transition point detected by the screen transition detection unit (130) to the next screen transition point. It recommends the main camera image that will continue the video content from the screen transition point.
[0032] Camera video recommendations can utilize a linked camera recommendation model, a machine learning model trained to recommend new camera videos to be connected at the transition point, by inputting features extracted from camera videos before and after the transition point. The linked camera recommendation model can be supervised using a training dataset consisting of labels from the final PGM video and inputting multi-channel camera videos owned by the broadcasting station.
[0033] Meanwhile, since the connected camera recommendation model outputs recommended camera images along with their accuracy, the recommended camera images can be ranked based on their accuracy. Accordingly, the connected camera recommendation unit (140) can recommend a set number of camera images in descending order of accuracy as candidate images after the screen transition point.
[0034] The editing information generation unit (150) synthesizes the screen transition points detected by the screen transition detection unit (130) throughout the entire video section, and synthesizes camera image information recommended by the connection camera recommendation unit (140) at the screen transition points to generate automatic editing information for the video content.
[0035] FIG. 2 is a diagram illustrating the flow of a video automatic editing method according to another embodiment of the present invention.
[0036] For automatic video editing, the video input unit (110) receives multiple camera images taken from different angles of the shooting scene in multi-channel format (S210), and the video feature extraction unit (120) extracts features from the multi-channel camera images input in step S210 (S220).
[0037] The screen transition detection unit (130) detects the screen transition point based on the features extracted in step S220 (S230). Then, the connection camera recommendation unit (140) recommends camera images that will compose the image from the screen transition point detected in step S240 to the next screen transition point (S240).
[0038] The next editing information generation unit (150) synthesizes the screen transition points detected in step S230 across the entire video section and synthesizes the recommended camera image information in step S240 to generate automatic editing information for the video content (S250).
[0039] So far, a preferred embodiment of a method for automatically editing a video through screen transition point detection and recommended connected camera images has been described in detail.
[0040] In the above embodiment, unlike the existing method where the screen transition point and the connected camera images were determined solely by the editor's judgment, the multi-channel camera images acquired by shooting the shooting site from various angles were analyzed based on artificial intelligence to automatically detect the screen transition points, and the newly connected camera images were automatically recommended at the detected screen transition points.
[0041] This will drastically reduce the amount of work and time required for editors to produce video content.
[0042] Meanwhile, it goes without saying that the technical idea of the present invention can also be applied to a computer-readable recording medium containing a computer program that performs the functions of the device and method according to the present embodiment. In addition, the technical idea according to various embodiments of the present invention can be implemented in the form of computer-readable code recorded on a computer-readable recording medium. The computer-readable recording medium can be any data storage device that can be read by a computer and store data. For example, the computer-readable recording medium can be a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, an optical disk, a hard disk drive, etc. In addition, the computer-readable code or program stored on the computer-readable recording medium can be transmitted through a network connected between computers.
[0043] In addition, although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.
Claims
1. A step of receiving multiple camera images of the same scene; A step of detecting a screen transition point from input camera images; A video automatic editing method, characterized by including a step of recommending camera images to compose a video from a detected screen transition point to the next screen transition point.
2. In claim 1, The input stage is, A video automatic editing method characterized by receiving multiple camera images taken from different angles at a shooting site.
3. In claim 1, The detection step is, A video automatic editing method characterized by detecting whether a screen transition occurs for each camera image, and detecting the screen transition point when a screen transition is detected in a number of camera images greater than or equal to a predetermined number.
4. In claim 3, further comprising a step of extracting features from camera images; The detection step is, A video automatic editing method characterized by detecting whether a screen transition occurs from extracted features.
5. In claim 4, The detection step is, A video automatic editing method characterized by detecting whether a screen has changed from extracted features using a machine learning model trained to estimate whether a screen has changed from extracted features.
6. In claim 1, The features are, A method for automatically editing a video, characterized in that it includes dialogue features, action features, reaction features and background features.
7. In claim 4, The recommended steps are: A video automatic editing method characterized by recommending camera images using a machine learning model trained to recommend new camera images to be connected at the screen transition point by inputting features extracted from camera images before the screen transition point and features extracted from camera images after the screen transition point.
8. In claim 7, The recommended steps are: A video automatic editing method characterized by recommending multiple camera images based on the accuracy of a machine learning model.
9. In claim 1, A video automatic editing method, further comprising: a step of generating automatic editing information of video content by synthesizing detected screen transition points in the entire section of the video and synthesizing recommended camera image information at the screen transition points.
10. Input section for receiving multiple camera images of the same scene; A detection unit that detects the screen transition point from input camera images; An automatic video editing system, characterized by including a recommendation unit that recommends camera images to compose a video from a detected screen transition point to the next screen transition point.
11. A step of detecting the screen transition point from multiple camera images of the same scene; A step for recommending camera images to compose the image from the detected screen transition point to the next screen transition point; A video automatic editing method, characterized by including a step of generating automatic editing information by synthesizing detected screen transition points and recommended camera image information.
12. A detection unit that detects the point in time when the screen changes from multiple camera images of the same scene; A recommendation unit that recommends camera images to compose the video from the detected screen transition point to the next screen transition point; A video automatic editing system, characterized by including a generation unit that generates automatic editing information by synthesizing detected screen transition points and recommended camera image information.
Citation Information
Patent Citations
Moving image editing apparatus, information terminal, moving image editing method, and moving image editing program
JP2013232813A
Video editing method with music source
KR1020180080642A
Bird Collision Preventing Laminate
KR1020250004963A
Video Cut-editing online server according to category based on reflection of the purpose of video production and operating method of the online server
KR102499873B1
KR20190060027A