Method for extracting important video scene time points on basis of scene changes and event recognition, and device therefor
The proposed device and method for video scene change detection improve accuracy and efficiency by dividing frames into local images, extracting HSV histograms, and analyzing similarities, effectively addressing the limitations of existing technologies.
Patent Information
- Application Number
- PCT/KR2024/019451
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-12
AI Technical Summary
Existing methods for detecting scene changes in videos are prone to errors due to sensitivity to small movements and lighting changes, and they incur high computational costs and slow processing speeds, especially in complex scenes and low-light environments.
A device and method that divide each frame of a video into n*m local images, extract HSV histograms for each local image, and detect similarity between corresponding local images in adjacent frames using a similarity detection unit, thereby minimizing errors and improving processing efficiency.
The method effectively extracts the points in time when scene changes occur and when events start or end within a video, while reducing errors and enhancing processing speed, even in challenging conditions like low-light environments.
Smart Images

Figure KR2024019451_12062025_PF_FP_ABST
Abstract
Description
Method and device for extracting important scene points in a video based on scene transition and event recognition
[0001] The present invention relates to a device and method for automatically extracting the point in time at which a scene changes and the point in time at which an event occurs or ends within an input video.
[0002] The content described in this section merely provides background information for the present embodiment and does not constitute prior art.
[0003] Scene change detection is considered a crucial technical challenge in the fields of video editing and analysis. Scene changes are crucial elements that break the continuity of a video and play a key role in editing, summarizing, retrieval, and search optimization processes.
[0004] Previously, there were pixel-based methods for detecting scene transitions. These methods analyze changes in pixel values between consecutive frames of a video. If a significant change is detected, a scene transition is considered to have occurred. However, these methods are sensitive to even subtle background movements or lighting changes, which can lead to errors.
[0005] Due to these critical issues, feature-based scene transition detection methods have emerged. These methods extract image features, such as edges or textures, and analyze changes in these features to detect scene transitions. While effective in complex scenes, they suffer from high computational costs and slow processing speeds.
[0006] While these traditional methods can be effective for simple scene transitions, they are prone to errors due to various factors, such as dynamic video content, complex backgrounds, and rapid camera movements. In particular, scene transition detection in low-light environments is extremely difficult. Furthermore, these methods face limitations in processing speed and efficiency when processing large amounts of video data in real time.
[0007] Therefore, there is a need for a new method to detect scene changes within a video.
[0008] One embodiment of the present invention aims to provide a device and method for automatically extracting the point in time at which a scene changes and the point in time at which an event occurs or ends within an input image while minimizing errors.
[0009] According to one aspect of the present invention, a scene change recognition device is provided, characterized by including a local image separation unit that separates each frame of an input image into n*m local images, a histogram extraction unit that extracts an HSV (hue, saturation, brightness) histogram for each local image within each frame, and a similarity detection unit that detects similarity between each local image and the corresponding frames based on histograms extracted from local images at corresponding locations within two adjacent frames.
[0010] According to one aspect of the present invention, the histogram extraction unit is characterized by converting RGB values within each frame of the input image into HSV space.
[0011] According to one aspect of the present invention, the histogram extraction unit is characterized in that it converts each frame and each regional image into HSV (hue, saturation, brightness) space and then extracts a histogram for each of H·S·V.
[0012] According to one aspect of the present invention, the similarity detection unit is characterized in that it determines whether the similarity of local images at corresponding locations within adjacent two frames is equal to or greater than a preset first reference value.
[0013] According to one aspect of the present invention, the similarity detection unit is characterized in that, when the similarity is greater than or equal to a preset first criterion, it determines that the local images of corresponding locations within adjacent two frames are similar to each other.
[0014] According to one aspect of the present invention, the similarity detection unit is characterized in that, when the similarity is less than a preset first reference value, it determines that the local images of corresponding locations within adjacent two frames are not similar to each other.
[0015] According to one aspect of the present invention, a scene change point extraction device is provided, which comprises a scene change recognition unit that receives an image and recognizes whether a scene changes within the image, and a point extraction unit that outputs a point in time at which a scene changes within the image based on an analysis result of the scene change recognition unit, wherein the scene change recognition unit comprises a local image separation unit that separates each frame of the input image into n*m local images, a histogram extraction unit that extracts an HSV (hue, saturation, brightness) histogram for each local image within each frame, and a similarity detection unit that detects similarity between each local image and the corresponding frames based on histograms extracted from local images at corresponding positions within two adjacent frames.
[0016] According to one aspect of the present invention, a scene change recognition device is provided, comprising: a local image separation unit that separates each frame of an input image into n*m local images; a histogram extraction unit that extracts an HSV (hue, saturation, brightness) histogram for each local image within each frame; a similarity detection unit that detects similarity between each local image and the corresponding frame based on histograms extracted from local images at corresponding positions within adjacent two frames; and a memory unit that stores a time point of a frame determined by the similarity detection unit to have a different similarity from an adjacent frame.
[0017] According to one aspect of the present invention, the similarity detection unit is characterized in that, when the similarity is less than a preset first reference value, it determines that the local images of corresponding locations within adjacent two frames are not similar to each other.
[0018] According to one aspect of the present invention, the similarity detection unit is characterized in that it determines whether the number of regional images determined to be similar to each other within both frames is less than or equal to a preset second reference value compared to the total number of regional images within each frame.
[0019] According to one aspect of the present invention, the similarity detection unit is characterized in that, when the number of regional images determined to be similar to each other within both frames is less than or equal to a preset second reference value compared to the total number of regional images within each frame, it detects that a scene has changed between the two frames.
[0020] According to one aspect of the present invention, a scene change point extraction device is provided, which includes a scene change recognition unit that receives an image and recognizes whether a scene changes within the image, and a point in time extraction unit that outputs the point in time at which a scene changes within the image based on an analysis result of the scene change recognition unit, wherein the scene change recognition unit includes a region image separation unit that separates each frame of the input image into n*m region images, a histogram extraction unit that extracts an HSV (hue, saturation, brightness) histogram for each region image within each frame, a similarity detection unit that detects similarity between each region image and the corresponding frame based on histograms extracted from region images at corresponding positions within adjacent two frames, and a memory unit that stores the point in time of a frame determined by the similarity detection unit to have a different similarity from an adjacent frame.
[0021] According to one aspect of the present invention, a method for extracting a point in time when a scene changes in an image by a scene change point extraction device is provided, the method comprising: a separation process for dividing each frame in the image into a preset number of regional images; an extraction process for extracting an HSV histogram of each regional image; a calculation process for calculating a similarity of corresponding regional images between adjacent frames based on the extracted HSV histogram; an identification process for determining the number of regional images having a similarity greater than or equal to a preset first criterion; a judgment process for determining whether the number of regional images having a similarity greater than or equal to the preset first criterion is less than or equal to a preset second criterion; and a detection process for detecting whether a scene change has occurred based on a result of the judgment process.
[0022] A scene change point extraction device is provided, characterized by including a scene change recognition unit that receives a video input and recognizes whether a scene changes within the video, an event recognition unit that recognizes whether an event occurs or ends within each identical scene section recognized as the same scene by the scene change recognition unit, a section merging unit that merges the scene change point recognized by the scene change recognition unit and the event occurrence and end points recognized by the event recognition unit, and a point extraction unit that outputs the point in time at which a scene changes within the video to the outside based on the analysis results of each recognition unit or the merging results of the section merging unit.
[0023] As described above, according to one aspect of the present invention, there is an advantage in that the point in time at which a scene changes and the point in time at which an event occurs or ends can be automatically extracted within an input image while minimizing errors.
[0024] FIG. 1 is a diagram illustrating the configuration of a scene transition point extraction device according to one embodiment of the present invention.
[0025] FIG. 2 is a diagram illustrating the configuration of a scene transition recognition unit according to one embodiment of the present invention.
[0026] Figures 3 and 4 are drawings illustrating a process in which a scene change recognition unit recognizes a scene change point in time according to one embodiment of the present invention.
[0027] FIGS. 5 and 6 are diagrams illustrating a process of storing a scene in which a memory unit has been switched according to one embodiment of the present invention.
[0028] FIGS. 7 and 8 are diagrams illustrating a process in which an event recognition unit recognizes an event within a scene section according to one embodiment of the present invention.
[0029] FIG. 9 is a diagram illustrating a process in which a section merging unit according to one embodiment of the present invention merges a time point recognized by a scene transition recognition unit and a time point recognized by an event recognition unit.
[0030] FIG. 10 is a flowchart illustrating a method for a scene transition recognition unit to recognize a scene transition point according to one embodiment of the present invention.
[0031] FIG. 11 is a flowchart illustrating a method for extracting a main scene by a scene transition point extraction device according to one embodiment of the present invention.
[0032] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention. Throughout the description of each drawing, similar reference numerals have been used to designate similar components.
[0033] Terms such as first, second, A, and B may be used to describe various components, but these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component. The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.
[0034] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0035] The terminology used in this application is solely for the purpose of describing specific embodiments and is not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly dictates otherwise. It should be understood that terms such as "comprise" or "have" in this application do not preclude the presence or possibility of addition of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification.
[0036] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.
[0037] Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless expressly defined in this application.
[0038] In addition, each configuration, process, procedure or method included in each embodiment of the present invention may be shared within a scope that is not technically inconsistent with each other.
[0039] FIG. 1 is a diagram illustrating the configuration of a scene transition point extraction device according to one embodiment of the present invention.
[0040] Referring to FIG. 1, a scene change point extraction device (100, hereinafter abbreviated as “device”) according to one embodiment of the present invention includes a scene change recognition unit (110), an event recognition unit (120), a section merging unit (130), and a point extraction unit (140).
[0041] The device (100) receives an arbitrary video input, analyzes and extracts the points in time when a scene changes within the video and the points in time when an event occurs within the same scene section. The device (100) recognizes the points in time when a scene changes within the video and the points in time when an event occurs, extracts the points in time, and notifies the external device, thereby allowing the device user to confirm only the points in time when a scene changes within the video without making any additional effort.
[0042] The scene change recognition unit (110) receives an image and recognizes whether a scene changes within the image. The scene change recognition unit (110) divides each frame within the image into multiple regional images and extracts a histogram of each regional image to detect the similarity between adjacent frames. If the similarity between adjacent frames is below a certain level, the scene change recognition unit (110) determines that a scene has changed between adjacent frames. The scene change recognition unit (110) stores the point in time (frame) at which the scene changed.
[0043] The event recognition unit (120) recognizes whether an event has occurred or ended within each identical scene section recognized as the same scene by the scene transition recognition unit (110). The event recognition unit (120) recognizes whether an object appears within each identical scene section, whether the movement or action of an existing object has changed, or whether voice data has changed, thereby recognizing whether an event has occurred or ended within the corresponding scene section. Similarly, the event recognition unit (120) also stores the time (frame) at which the event occurred and the time at which the event ended.
[0044] The section merging unit (130) merges the scene change time recognized by the scene change recognition unit (110) with the event occurrence time and end time recognized by the event recognition unit (120). The device user may want to see only the scene change time or only the event occurrence / end time, but there may also be cases where they want to see both. Accordingly, the section merging unit (130) merges each time point recognized by each recognition unit (110, 120) so that the device user can check both. As described above, the event recognition unit (120) recognizes whether an event occurs or ends within the same scene section. Therefore, the occurrence or end of an event typically exists within a specific scene section. The section merging unit (130) utilizes this characteristic and merges the occurrence or end time of the event within the scene section when the occurrence or end of an event occurs within a scene section. The device user can collectively check which events have occurred or ended within each scene section as well as each scene transition.
[0045] The point-in-time extraction unit (140) outputs the point-in-time of scene transition or occurrence / end of an event within the video to the outside based on the analysis results of each recognition unit (110, 120) or the merging results of the section merging unit (130). The point-in-time extraction unit (140) extracts the point-in-time (frame) at which a scene transition or an event occurs or ends based on the analysis results of each recognition unit (110, 120). The point-in-time extraction unit (140) extracts the corresponding point-in-time within the video and notifies the outside (user) so that the user can more conveniently recognize the point-in-time at which a scene transition occurs or an event occurs / ends within a scene section. Alternatively, the point-in-time extraction unit (140) allows the device user to collectively recognize all points-in-time at which a scene transition occurs or an event occurs / ends within a scene section based on the external notification of the result of merging by the section merging unit (130).
[0046] The device (100) extracts only the scene transition point and / or the point in time when an event occurs or ends in the input video, thereby enabling the user of the device to easily recognize the corresponding points in the video without having to check the video individually.
[0047] FIG. 2 is a diagram illustrating a configuration of a scene change recognition unit according to one embodiment of the present invention, FIGS. 3 and 4 are diagrams illustrating a process in which a scene change recognition unit according to one embodiment of the present invention recognizes a scene change point in time, and FIGS. 5 and 6 are diagrams illustrating a process in which a memory unit according to one embodiment of the present invention stores a changed scene.
[0048] Referring to FIG. 2, a scene transition recognition unit (110) according to one embodiment of the present invention includes a local image separation unit (210), a histogram extraction unit (220), a similarity detection unit (230), and a memory unit (240).
[0049] The local image separation unit (210) separates each frame of the input image into n*m local images. The local image separation unit (210) separates each frame into multiple local images as illustrated in Fig. 3. The local image separation unit (210) separates the i-th frame into multiple local images, and the separated local images can be defined as follows.
[0050]
[0051] The defined local image refers to the local image at the (x, y) coordinates within the i-th frame. In this way, the local image separation unit (210) separates each frame into an equal number of local images. Accordingly, local images of corresponding locations between adjacent frames can be generated.
[0052] The histogram extraction unit (220) extracts an HSV (hue, saturation, brightness) histogram for each regional image within each frame. The input image has an RGB value for each pixel. The histogram extraction unit (220) converts the RGB values within each frame into the HSV (hue, saturation, brightness) space. The histogram extraction unit (220) converts each frame and each regional image into the HSV (hue, saturation, brightness) space, and then extracts a histogram for each of H, S, and V. The histogram is defined as follows.
[0053]
[0054] The defined histogram refers to a histogram for a local image at the (x, y) coordinates within the i-th frame. As illustrated in Fig. 4, the histogram can be implemented in a form that accumulates H·S·V values of each local image converted into HSV spatial values. The histogram extraction unit (220) extracts histograms for H·S·V of each local image in this form.
[0055] The similarity detection unit (230) detects the similarity between each regional image and the corresponding frames based on the histogram extracted from the regional images at corresponding locations within adjacent two frames (i-th frame and i+1-th frame).
[0056] The similarity detection unit (230) detects the similarity of local images of corresponding locations using the following formula.
[0057]
[0058] Here, H1 and H2 represent histograms of local images at corresponding locations within adjacent two frames, and Sl(H1, H2) represents the similarity of the local images, respectively.
[0059] The similarity detection unit (230) determines whether the similarity of the corresponding regional images calculated through the aforementioned formula is greater than or equal to a preset first criterion. If the similarity of the corresponding regional images is greater than or equal to the preset first criterion, the similarity detection unit (230) determines that the two regional images are similar to each other. Otherwise, the similarity detection unit (230) determines that the two regional images are not similar to each other.
[0060] The similarity detection unit (230) determines the aforementioned similarity for all local images of corresponding locations within adjacent two frames (i-th frame and i+1-th frame).
[0061] When the number of regional images (in corresponding locations) that are judged to be similar to each other in both frames is less than or equal to a preset second criterion compared to the total number of regional images in each frame, the similarity detection unit (230) detects that the scene has changed between the two frames. On the other hand, when the number of regional images (in corresponding locations) that are judged to be similar to each other in both frames exceeds the preset second criterion, the similarity detection unit (230) detects that the scene has continued.
[0062] The memory unit (240) stores the point of time of a frame determined by the similarity detection unit (230) to have a different similarity (that the scene has changed) from an adjacent frame. As illustrated in FIG. 5, the memory unit (240) may store the point of time of a frame determined to have a different similarity (that the scene has changed) from an adjacent frame. Alternatively, as illustrated in FIG. 6, the memory unit (240) may store metadata for a frame determined to have a different similarity (that the scene has changed) from an adjacent frame. By storing the point of time or metadata of the corresponding frame in this way, the memory unit (240) enables the point of time extraction unit (140) to recognize at which point (frame) in the video a scene has changed based on the information stored in the memory unit (240) and to transmit or output this to an external device.
[0063] FIGS. 7 and 8 are diagrams illustrating a process in which an event recognition unit recognizes an event within a scene section according to one embodiment of the present invention.
[0064] Referring to FIG. 7, the event recognition unit (120) recognizes whether an event has occurred or ended within each identical scene section recognized by the scene transition recognition unit (110).
[0065] The event recognition unit (120) receives an image as an input value and includes an artificial intelligence learning model that has learned to recognize whether an object appears, whether movement or action has occurred or changed in the object, or whether audio data has occurred or changed within the image. Accordingly, the event recognition unit (120) receives an image, more specifically, an image in which each scene section is divided by the scene transition recognition unit (110), and outputs the aforementioned output value (appearance of an object, etc.).
[0066] As illustrated in FIG. 7, the event recognition unit (120) recognizes whether an object appears or disappears within a specific scene section, or whether movement or motion occurs in an existing object. Here, the object may be an inanimate object such as an object or device, as illustrated in FIG. 8, a living object such as a human or an animal, or a specific situation such as a fire outbreak. The event recognition unit (120) recognizes an object or object movement within an input image using the aforementioned artificial intelligence learning model.
[0067] Meanwhile, the event recognition unit (120) recognizes the occurrence or change of voice data within the input image using the artificial intelligence learning model described above. The event recognition unit (120) analyzes the frequency, vibration frequency, amplitude, size, etc. of the voice data within the image to recognize whether voice data has occurred or changed. The occurrence of an event (emergency situation) can be determined based on the occurrence of voice data, such as when an emergency bell, control sound, or a person's scream is generated, and the end of a situation or event can be determined based on the cessation of an ongoing conversation.
[0068] The event recognition unit (120) recognizes various situations as described above, stores the time point of the recognized frame, or stores metadata for the frame. Accordingly, the time point extraction unit (140) recognizes the time point (frame) at which an event occurred or ended within the video based on the information stored in the event recognition unit (120), and transmits or outputs this information to an external source.
[0069] FIG. 9 is a diagram illustrating a process in which a section merging unit according to one embodiment of the present invention merges a time point recognized by a scene transition recognition unit and a time point recognized by an event recognition unit.
[0070] The section merging unit (130) recognizes the scene transition time recognized by the scene transition recognition unit (110) and the event occurrence / end time recognized by the event recognition unit (120). The section merging unit (130) recognizes each time point and merges each event occurrence / end time within each scene section. The section merging unit (130) allows the device user to collectively check the scene transition time and the event occurrence / end time.
[0071] FIG. 10 is a flowchart illustrating a method for a scene transition recognition unit to recognize a scene transition point according to one embodiment of the present invention.
[0072] The local image separation unit (210) separates each frame in the image into a preset number of local images (S1010).
[0073] The histogram extraction unit (220) extracts the HSV histogram of each regional image (S1020).
[0074] The similarity detection unit (230) calculates the similarity of corresponding regional images between adjacent frames based on the extracted HSV histogram (S1030). The similarity detection unit (230) determines whether the similarity of corresponding regional images between adjacent frames is greater than or equal to a preset first criterion.
[0075] The similarity detection unit (230) determines the number of regional images whose similarity is greater than a preset first standard value (S1040).
[0076] The similarity detection unit (230) determines whether the number of images in the corresponding region is less than or equal to a preset second standard (S1050).
[0077] If the number of images in the corresponding region is less than or equal to the preset second criterion, the similarity detection unit (230) determines that a scene change has occurred (S1060).
[0078] If the number of images in the corresponding region exceeds the preset second criterion, the similarity detection unit (230) determines that there is no scene change (S1070).
[0079] FIG. 11 is a flowchart illustrating a method for extracting a main scene by a scene transition point extraction device according to one embodiment of the present invention.
[0080] The scene change recognition unit (110) recognizes the scene change point in the received image (S1110).
[0081] The event recognition unit (120) recognizes whether an event has occurred or ended within each identical scene section (S1120).
[0082] The section merging unit (130) merges the scene change time recognized by the scene change recognition unit (110) and the event occurrence / end time recognized by the event recognition unit (120) (S1130).
[0083] The point extraction unit (140) outputs externally the scene change point recognized by the scene change recognition unit (110), the event occurrence / end point recognized by the event recognition unit (120), or the point merged by the section merge unit (130) (S1140).
[0084] Although FIGS. 10 and 11 describe each process as being executed sequentially, this is merely an illustrative description of the technical idea of one embodiment of the present invention. In other words, a person skilled in the art to which one embodiment of the present invention pertains may modify and apply various modifications and variations, such as changing the order described in each drawing and executing it without departing from the essential characteristics of one embodiment of the present invention, or executing one or more of each process in parallel. Therefore, FIGS. 10 and 11 are not limited to a chronological order.
[0085] Meanwhile, the processes illustrated in FIGS. 10 and 11 can be implemented as computer-readable codes on a computer-readable recording medium. A computer-readable recording medium includes all types of recording devices that store data that can be read by a computer system. That is, a computer-readable recording medium includes storage media such as magnetic storage media (e.g., ROM, floppy disks, hard disks, etc.) and optical reading media (e.g., CD-ROMs, DVDs, etc.). In addition, a computer-readable recording medium can be distributed across network-connected computer systems, so that the computer-readable codes can be stored and executed in a distributed manner.
[0086] The scene change point extraction device (100) may be any type of mobile device, such as a wearable device. The scene change point extraction device (100) may include a central processing unit implemented as a controller, integrated circuit, microchip, computer, or other computing device.
[0087] The scene transition point extraction device (100) may include a memory module. The memory module may include RAM, ROM, flash memory, a hard drive, or a device capable of storing machine-readable and executable instructions accessible by a central processing unit.
[0088] The memory module can store instructions from the central processing unit to instruct each component in the scene change point extraction device (100) to perform the aforementioned operations when the central processing unit operates.
[0089] The instructions may include one or more logic or algorithms written in any programming language, such as machine language, which can be executed directly by the processor, or assembly language, object-oriented programming (OOP), scripting language, microcode, etc., which can be compiled or assembled into machine-readable and executable instructions and stored in a memory module. Alternatively, the machine-readable and executable instructions may be written in a hardware description language (HDL), such as logic implemented by a field-programmable gate array (FPGA) configuration or an application-specific integrated circuit (ASIC).
[0090] The above description is merely an example of the technical idea of the present embodiment, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential characteristics of the present embodiment. Therefore, the present embodiments are not intended to limit the technical idea of the present embodiment, but rather to explain it, and the scope of the technical idea of the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of rights of the present embodiment.
[0091]
[0092] **This patent is the result of research conducted with the support of the National IT Industry Promotion Agency (NIPA) funded by the Ministry of Science and ICT (Ministry of Science and ICT) of the Republic of Korea (Project ID: 2710006746, Subproject ID: II220419, Project Name: (Detail 6-4) Development of Unconstrained Robust Precision Motion High-Speed Sensing Technology Specialized for Manufacturing Sites).
[0093]
[0094] CROSS-REFERENCE TO RELATED APPLICATION
[0095] **This patent application claims priority under 35 USC § 119(a) to Korean Patent Application No. 10-2023-0177626, filed in Korea on December 8, 2023, the entire contents of which are incorporated by reference herein. Furthermore, if this patent application claims priority in countries other than the United States for the same reasons, the entire contents of which are incorporated by reference herein.
Claims
1. A regional image separation unit that separates each frame of the input image into n*m regional images; A histogram extraction unit that extracts HSV (hue, saturation, brightness) histograms for each local image in each frame; and A similarity detection unit that detects the similarity between each regional image and its corresponding frames based on the histogram extracted from the regional images at corresponding locations in adjacent two frames. A scene transition recognition device characterized by including a .
2. In paragraph 1, The above histogram extraction unit, A scene change recognition device characterized by converting RGB values within each frame of an input image into HSV space.
3. In paragraph 2, The above histogram extraction unit, A scene change recognition device characterized by extracting histograms for each of H, S, and V after converting each frame and each regional image into HSV (hue, saturation, brightness) space.
4. In paragraph 1, The above similarity detection unit, A scene change recognition device characterized by determining whether the similarity of local images at corresponding locations within adjacent two frames is greater than a preset first criterion.
5. In paragraph 4, The above similarity detection unit, A scene change recognition device characterized in that when the similarity is greater than a preset first criterion, the local images of corresponding positions within adjacent two frames are judged to be similar to each other.
6. In paragraph 4, The above similarity detection unit, A scene change recognition device characterized in that when the similarity is less than a preset first criterion, it is determined that the local images of corresponding positions in adjacent two frames are not similar to each other.
7. A scene change recognition unit that receives a video input and recognizes whether a scene changes within the video; and Based on the analysis results of the scene change recognition unit above, it includes a point in time extraction unit that outputs the point in time when a scene is changed within the video to the outside. The above scene transition recognition unit, A regional image separation unit that separates each frame of the input image into n*m regional images; A histogram extraction unit that extracts HSV (hue, saturation, brightness) histograms for each local image in each frame; and A scene change point extraction device characterized by including a similarity detection unit that detects similarity between each regional image and the corresponding frames based on histograms extracted from regional images at corresponding locations in adjacent two frames.
8. In paragraph 7, The above similarity detection unit, A scene change recognition device characterized in that it determines whether the number of regional images judged to be similar to each other within both frames is less than or equal to a preset second criterion compared to the total number of regional images within each frame.
9. In paragraph 8, The above similarity detection unit, A scene change recognition device characterized in that it detects that a scene has changed between the two frames when the number of regional images judged to be similar to each other within both frames is less than or equal to a preset second criterion compared to the total number of regional images within each frame.
10. In a method for extracting the point in time at which a scene change occurs in a video, A separation process that divides each frame in an image into a preset number of regional images; Extraction process to extract HSV histograms of each regional image; A calculation process that calculates the similarity of corresponding local images between adjacent frames based on the extracted HSV histogram; A process of identifying the number of regional images whose similarity exceeds a preset first criterion; A judgment process for determining whether the number of regional images with a similarity greater than or equal to a preset first criterion is less than or equal to a preset second criterion; and A detection process that detects whether a scene change has occurred based on the judgment result of the above judgment process. A method for extracting scene transition points, characterized by including:
Citation Information
Patent Citations
Method of measuring similarity degree of digital animation content, method of managing animation content using the same, and management system for animation content using the method of managing animation content
JP2010074832A
Device for detecting scene change and method for detecting scene change
KR1020130061865A
Scene cut frame detecting apparatus and method
KR1020170090868A
Method of playing image into music and server supplying program for performing the same
KR102089207B1
Using change of scene to trigger automatic image capture
US20200213509A1