Scene division method for managing and analyzing video content

The scene-by-scene video separation method addresses the limitations of conventional video segmentation by using feature analysis and machine learning to group shots into meaningful scenes, enhancing video content management and analysis for improved service quality in various applications.

WO2025127771A1PCT designated stage expired Publication Date: 2025-06-19KOREA ELECTRONICS TECH INST
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/095982
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-11
Filing Date
2024-08-02
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Conventional video segmentation technologies that detect shot boundaries based on screen transitions, camera angle changes, and color information fluctuations are inadequate for effective management and analysis of video content, as they result in short units of content that are not meaningful for analysis.

Method used

A scene-by-scene separation method that uses feature analysis to group shots into meaningful scenes, employing machine learning models for shot boundary detection and feature extraction, including conversational, behavioral, reaction, and background features, to calculate similarities and group shots accordingly.

Benefits of technology

This method enables effective management and analysis of video content by dividing it into meaningful scene units, improving service quality in video summarization, editing, and search applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024095982_19062025_PF_FP_ABST
    Figure KR2024095982_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a scene division method for managing and analyzing video content. A method for dividing a video according to an embodiment of the present invention receives a video, divides the video into shot units, and groups the divided shots, thereby dividing the video into scene units. Accordingly, the shots divided from the video are grouped into scene units through feature analysis to divide the video into scene units, thus making it possible to effectively manage and analyze video content and improve service quality in the fields of video summarization, editing, and searching.
Need to check novelty before this filing date? Find Prior Art

Description

Scene-by-scene separation method for video content management and analysis

[0001] The present invention relates to video processing technology, and more particularly, to a method and system for dividing a video into scenes for effective management and analysis of video content.

[0002] The continued growth of video content, driven by the advancement of the Internet, is further increasing the need for effective management and analysis of video content. Unlike images, video content is comprised of numerous frames. Therefore, for efficient management and analysis, it is necessary to segment it into meaningful units.

[0003] Most of the conventional video segmentation techniques detect shot boundaries based on screen transitions, changes in camera angle, and changes in color information, regardless of the content.

[0004] However, since these separate shot-unit videos can only contain short units of content, there are limitations in effectively managing and analyzing video content based on them.

[0005] The present invention has been devised to solve the above problems, and the purpose of the present invention is to provide a video separation method that groups shots separated from a video into scenes through feature analysis, as a means to enable effective management and analysis of video content.

[0006] A video separation method according to one embodiment of the present invention for achieving the above object includes: a step of receiving a video; a first separation step of separating the video into shot units; and a second separation step of grouping the separated shots and separating the video into scene units.

[0007] The first separation step may be to separate the video into shots using a machine learning model trained to input a video and separate and output it into shots.

[0008] The video segmentation method according to the present invention further includes a step of extracting features from each of the separated shots; and the second segmentation step may be to group the separated shots by comparing the extracted features.

[0009] The second separation step may be to calculate similarities between features of shots belonging to a specific section and group the shots based on the calculated similarities.

[0010] The second separation step may be to calculate the final similarity by averaging the feature-wise similarities.

[0011] Features may include conversational features, behavioral features, and reaction features.

[0012] Features may further include background features.

[0013] The video separation method according to the present invention may further include a step of synthesizing frame number information by scene unit; and a step of generating the synthesized frame number information as scene information.

[0014] The generated scene information can be used as a reference to separate videos for video management and analysis.

[0015] According to another aspect of the present invention, a video separation system is provided, characterized by including a first separation unit that receives a video and separates it into shot units; and a second separation unit that groups the separated shots and separates the video into scene units.

[0016] According to another aspect of the present invention, a video separation method is provided, comprising: a first separation step of separating a video into shot units; a second separation step of separating a video into scene units by grouping the separated shots; and a step of generating scene information of the separated video.

[0017] According to another aspect of the present invention, a video separation system is provided, comprising: a first separation unit for separating a video into shot units; a second separation unit for separating a video into scene units by grouping the separated shots; and a generation unit for generating scene information of the separated video.

[0018] As described above, according to embodiments of the present invention, by grouping shots separated from a video into scenes through feature analysis and separating the video into scenes, effective management and analysis of video content is possible, thereby improving service quality in the fields of video summarization, editing, and search.

[0019] Figure 1 is a video content separation system according to one embodiment of the present invention;

[0020] Figure 2 shows the calculation of similarity between shots.

[0021] FIG. 3 is a diagram illustrating a flow of a video content separation method according to another embodiment of the present invention.

[0022] Hereinafter, the present invention will be described in more detail with reference to the drawings.

[0023] An embodiment of the present invention proposes a scene-by-scene separation method for video content management and analysis. This technique performs scene separation of video content, a prerequisite for efficient management and analysis of the ever-increasing volume of video content.

[0024] Unlike the existing method of separating videos according to screen transitions, changes in camera angles, and changes in color information of the video, the embodiment of the present invention separates videos by scene units, which are meaningful units of the video.

[0025] For reference, a video is composed of multiple scenes, a scene is composed of multiple shots, and a shot is composed of multiple frames.

[0026] FIG. 1 is a diagram illustrating the configuration of a video content separation system according to one embodiment of the present invention. The video content separation system according to the embodiment of the present invention is configured to include a shot-by-shot separation unit (110), a scene-by-scene separation unit (120), a scene information generation unit (130), and a video management / analysis unit (140).

[0027] The shot-by-shot separation unit (110) receives video content from a video source and separates the input video into shot-by-shot segments. To this end, the shot-by-shot separation unit (110) may utilize a machine learning model trained to receive a video, separate it into shot-by-shot segments through shot boundary detection, and output the resulting video.

[0028] The scene unit separation unit (120) is configured to ultimately separate video content into scenes by grouping shots separated by the shot unit separation unit (110) into scenes.

[0029] To this end, the scene unit separation unit (120) first extracts features from each shot separated by the shot unit separation unit (110). The extracted features include dialogue features, action features, reaction features, etc., and background features can also be included as needed.

[0030] Feature extraction can be performed using machine learning models trained to extract features from input shot-by-shot videos. These machine learning models can be implemented separately for each feature. For example, a machine learning model could be configured to extract dialogue features, a machine learning model to extract behavioral features, a machine learning model to extract reaction features, and a machine learning model to extract background features.

[0031] The next scene-by-scene separation unit (120) compares the extracted features and groups adjacent shots with similar features. Specifically, the scene-by-scene separation unit (120) calculates the similarities between features for shots within a specific section.

[0032] Similarity calculation is performed by calculating the similarity between all shots and the remaining shots, as shown in Figure 2, and similarity calculation is performed separately for each feature.

[0033] That is, similarities between shots are calculated for conversation features, similarities between shots are calculated for action features, similarities between shots are calculated for reaction features, and similarities between shots are calculated for background features.

[0034] The next scene unit separation unit (120) averages the calculated similarities to produce a final similarity between adjacent shots. Then, the scene unit separation unit (120) groups shots with high final similarity.

[0035] Through this, the video content is divided into scenes by the following scene unit separation unit (120).

[0036] The scene information generation unit (130) synthesizes frame number information for each scene in a video segmented into scenes by the scene unit separation unit (120), and generates the synthesized frame number information as scene information.

[0037] The video management / analysis unit (140) separates video content into scenes by referring to the scene information generated by the scene information generation unit (130), and then performs video management and analysis tasks such as video summarization, video editing, and video search.

[0038] FIG. 3 is a diagram illustrating a flow of a video content separation method according to another embodiment of the present invention.

[0039] To separate video content, first, a shot-by-shot separation unit (110) receives video content from a video source (S210), and separates the input video into shot-by-shot units in step S210 (S220).

[0040] The next scene unit separation unit (120) extracts features from each shot separated in step S220 (S230). The extracted features include dialogue features, action features, reaction features, etc., and background features may also be included as needed.

[0041] Thereafter, the scene unit separation unit (120) calculates similarities between shots based on the features of the shots extracted in step S230 (S240) and groups similar adjacent shots based on the calculated similarities (S250). As a result, the video content is separated into scenes.

[0042] The next scene information generation unit (130) synthesizes frame number information by scene for the video divided into scene units by step S250 (S260) and generates the synthesized frame number information as scene information (S270).

[0043] The scene information generated in step S270 is used to separate the video into scenes prior to video management and analysis tasks such as video summarization, video editing, and video search for video content.

[0044] So far, a preferred embodiment of a scene-by-scene separation method and system for video content management and analysis has been described in detail.

[0045] In the above embodiment, unlike the existing method of separating a video according to a change in screen transition, a change in camera angle, or a change in color information of the video, in the embodiment of the present invention, a video is separated by scene unit, which is a meaningful unit of the video.

[0046] By using an artificial intelligence model, video content is divided into meaningful units called scenes, so that it can be applied to many fields such as video summarization, video editing, and video search.

[0047] Meanwhile, it goes without saying that the technical idea of ​​the present invention can also be applied to a computer-readable recording medium containing a computer program that performs the functions of the device and method according to the present embodiment. In addition, the technical idea according to various embodiments of the present invention can be implemented in the form of computer-readable code recorded on a computer-readable recording medium. The computer-readable recording medium can be any data storage device that can be read by a computer and store data. For example, the computer-readable recording medium can be a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, an optical disk, a hard disk drive, etc. In addition, the computer-readable code or program stored on the computer-readable recording medium can be transmitted through a network connected between computers.

[0048] In addition, although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.

Claims

1. Step of receiving video; The first separation step is to separate the video into shot units; A video separation method characterized by including a second separation step of separating a video into scenes by grouping separated shots.

2. In claim 1, The first separation step is, A video segmentation method characterized by segmenting a video into shots using a machine learning model trained to input a video and segment and output the video into shots.

3. In claim 1, further comprising a step of extracting features from each of the separated shots; The second separation stage is, A video segmentation method characterized by grouping separated shots by comparing extracted features.

4. In claim 3, The second separation stage is, A video segmentation method characterized by calculating similarities between features of shots belonging to a specific section and grouping shots based on the calculated similarities.

5. In claim 4, The second separation stage is, A video segmentation method characterized by calculating final similarity by averaging feature-specific similarities.

6. In claim 3, The features are, A video segmentation method characterized by including conversation features, action features, and reaction features.

7. In claim 6, The features are, A video segmentation method characterized by including additional background features.

8. In claim 1, A step of synthesizing frame number information on a scene-by-scene basis; and A video separation method further comprising: a step of generating synthesized frame number information into scene information.

9. In claim 8, The generated scene information is, A video separation method characterized by being referenced for separating a video for video management and analysis.

10. A first separation unit that receives a video and separates it into shot units; A video separation system characterized by including a second separation unit that separates a video into scenes by grouping separated shots.

11. The first separation step of separating the video into shot units; A second separation step that groups the separated shots and separates the video into scenes; A video separation method, characterized by comprising a step of generating scene information of a separated video.

12. A first separation unit that separates the video into shot units; A second separation unit that groups the separated shots and separates the video into scenes; A video separation system, characterized by including a generation unit that generates scene information of a separated video.

Citation Information

Patent Citations

  • Method for video shot classification

    KR1020070107628A

  • Method, system and recording medium storing a computer program for building moving picture search database and method for searching moving picture using the sa ...

    KR1020080112975A

  • A method for indexing video frames with slide titles through synchronization of video lectures with slide notes

    KR1020120126953A

  • Integrated logistics management system

    KR1020210155128A

  • Glued fiberboard floor

    KR1020240168136A