Method and device for generating edited video
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2026-08-12
Smart Images

Figure PAT00002_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a method for generating an edited image and an apparatus for generating an edited image centered on an object, which can compensate for the incomplete detection of an object. Background Technology
[0002] Recently, techniques for detecting people in videos and tracking them using deep learning technology have been introduced. Additionally, there is a growing trend of users consuming vertical video content such as YouTube Shorts, Reels, and TikTok. To cater to this, video editors are editing horizontal video content to create additional vertical video content that includes people from the horizontal content.
[0003] An important aspect of this technology is accurately recognizing and detecting targets from multiple frames using deep learning technology. For example, when producing an edited video that tracks a key figure for 10 seconds, the key figure must be accurately recognized and detected from the multiple frames that make up the 10 seconds to generate a natural-looking edited video.
[0004] However, due to the limitations of current deep learning technology, it is impossible to accurately detect objects (e.g., people). For instance, cases may occur where the background is incorrectly detected as a person, people are not detected at all, or other people are detected along with the main character. Although a natural-looking edited video is generated only when the main character is continuously and accurately detected for 10 seconds, problems have arisen where tracking is interrupted or other objects (other people, animals, backgrounds, etc.) are tracked instead of the main character due to the imperfections of deep learning technology. The problem to be solved
[0005] The present invention aims to solve the aforementioned problems, and the objective of the present invention is to provide a method for generating an edited image and an apparatus for generating an edited image centered on a said object, which can compensate for the incomplete detection of the object. means of solving the problem
[0006] A method for generating an edited video according to the present invention comprises the steps of: obtaining a detection result in which an object is detected in a plurality of frames constituting a unit of video; obtaining a reference point of an important object in the unit of video using the detection result; and generating a supplementary result that supplements the detection result using the reference point.
[0007] In this case, the step of generating an edited video including the important object using the above supplementary result may be further included.
[0008] Meanwhile, the step of obtaining a reference point of the important object may include a step of selecting an important object among a plurality of objects detected across the plurality of frames.
[0009] In this case, each of the plurality of frames is divided into a plurality of zones, and the step of selecting the important object may include the step of calculating the number of detection areas detected across the plurality of frames for each zone, the step of selecting an important zone using the number of detection areas for each zone, and the step of selecting an object within the important zone as the important object.
[0010] In this case, the step of selecting important zones using the number of detection areas per zone may include the step of selecting important zones by further using the size of the detection area within the zone along with the number of detection areas per zone.
[0011] Meanwhile, the step of calculating the number of detection areas detected across the plurality of frames for each of the above zones may include, when a detection area exists in a first zone of a first frame and a detection area exists in a second zone adjacent to the first zone of a second frame, a step of integrating the first zone and the second zone to create an integrated zone, and a step of calculating the number of detection areas detected across the plurality of frames in the integrated zone.
[0012] Meanwhile, the step of obtaining a reference point of the important object may further include the step of selecting the average position of the important object across the plurality of frames as the reference point.
[0013] Meanwhile, the step of generating the supplementary result may include generating a reference position using the reference point, and generating the supplementary result in which a part of the detection area within the detection result is deleted and a new detection area is added to the detection result using the reference position.
[0014] In this case, the above reference position may be the position of a detection area having a distance less than or equal to a threshold value from the reference position before the update.
[0015] Meanwhile, the step of generating the supplementary result in which a portion of the detection area within the detection result is deleted and a new detection area is added to the detection result using the above reference position may include the steps of setting the reference point as the above reference position, determining the first frame having a detection area having a distance of less than or equal to the threshold value from the above reference position, and setting the location of the detection area having a distance of less than or equal to the threshold value from the above reference position as the new reference position.
[0016] Meanwhile, the step of generating the supplementary result in which a portion of the detection area within the detection result is deleted and a new detection area is added to the detection result using the reference position may include the step of setting the location of the detection area existing in the specific frame as the new reference position when a detection area having a distance less than or equal to a threshold value exists in the specific frame.
[0017] Meanwhile, the step of generating the supplementary result in which a portion of the detection area within the detection result is deleted and a new detection area is added to the detection result using the reference position may include, for the plurality of frames, a step of selecting a detection area for the important object among the detection areas for a plurality of objects detected across the plurality of frames using the distance between the reference position and the detection area.
[0018] In this case, the step of generating the supplementary result, in which a portion of the detection area within the detection result is deleted and a new detection area is added to the detection result using the reference position, may further include the step of generating the supplementary result by deleting a detection area from the detection result where the distance from the reference position is greater than a threshold value.
[0019] Meanwhile, the step of generating the supplementary result in which a portion of the detection area within the detection result is deleted and a new detection area is added to the detection result using the reference position may include the step of generating a reference position using the reference point, determining whether a detection area for the important object exists for the plurality of frames using the reference position, and adding a new detection area using the detection area of the previous frame or the detection area of the subsequent frame for the frame in which a detection area for the important object does not exist.
[0020] In this case, the step of adding a new detection area using the detection area of the previous frame or the detection area of the subsequent frame may include, for a frame in which the detection area does not exist, the step of adding a new detection area having the same location as the detection area of the previous frame.
[0021] Meanwhile, the editing video generation device according to the present invention includes a communication unit for acquiring content, and a control unit for acquiring a detection result of detecting an object in a plurality of frames constituting a unit of video within the content, acquiring a reference point of an important object in the unit of video using the detection result, and generating a supplementary result that supplements the detection result using the reference point.
[0022] In this case, the control unit can generate an edited image including the important object using the supplementary result.
[0023] Meanwhile, the control unit can select important objects among a plurality of objects detected across the plurality of frames.
[0024] In this case, each of the above-mentioned plurality of frames is divided into a plurality of zones, and the control unit calculates the number of detection areas detected across the plurality of frames for each zone, selects an important zone using the number of detection areas for each zone, and can select an object within the important zone as the important object.
[0025] Meanwhile, a computer program according to the present invention is stored in a non-transient readable storage medium to execute a method for generating an edited video, comprising the steps of: obtaining a detection result for detecting an object in a plurality of frames constituting a unit of video; obtaining a reference point of an important object in the unit of video using the detection result; and generating a supplementary result that supplements the detection result using the reference point. Brief explanation of the drawing
[0026] FIG. 1 is a block diagram illustrating the components of an edited image generation device according to the present invention. FIG. 2 is a flowchart for explaining a method for generating an edited image according to the present invention. FIG. 3 is a diagram illustrating a detection area in which an object is detected in a frame according to the present invention. FIGS. 4 and FIGS. 5 are drawings for explaining a method of selecting important objects according to the present invention. FIG. 6 is a drawing illustrating a reference point of an important object according to the present invention. FIG. 7 is a diagram illustrating the detection results generated by an object recognition model according to the present invention. FIG. 8 is a diagram illustrating detection results and supplementary results generated by an object recognition model according to the present invention. FIGS. 9 to 12 are drawings for explaining a method of sequentially performing complementary operations on multiple frames that constitute a single unit of video. Specific details for implementing the invention
[0027] Hereinafter, embodiments of the present invention will be described in more detail with reference to the attached drawings. Embodiments of the present invention may be modified in various forms, and the scope of the present invention should not be interpreted as being limited to the embodiments below. These embodiments are provided to more completely explain the present invention to those with average knowledge in the art. Furthermore, although specific terms have been used in the drawings and specification of the present invention, they are used only for the purpose of explaining the present invention and are not used to limit the meaning or the scope of the present invention as described in the claims. Therefore, those with ordinary knowledge in the art will understand that various modifications and equivalent alternative embodiments are possible therefrom. Accordingly, the true technical scope of protection of the present invention should be determined by the technical spirit of the appended claims.
[0028] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components regardless of drawing symbols will be assigned the same reference number, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably solely for the ease of drafting the specification and do not inherently possess distinct meanings or roles. Furthermore, in describing embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description will be omitted. Additionally, the attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification; the technical concept disclosed in this specification is not limited by the attached drawings, and it should be understood that they include all modifications, equivalents, and substitutions that fall within the spirit and technical scope of the present invention.
[0029] Terms including ordinal numbers such as first, second, etc., and a, b, c, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.
[0030] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0031] The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "having" are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0032] For convenience of explanation in implementing the present invention, the components may be described in detail; however, these components may be implemented within a single device or module, or a single component may be divided and implemented across multiple devices or modules.
[0033] FIG. 1 is a block diagram illustrating the components of an edited image generation device according to the present invention.
[0034] An edited image generating device (100) according to the present invention (hereinafter referred to as “device (100)”) may include a communication unit (110), a control unit (120), and a memory (130).
[0035] The communication unit (110) may include a communication interface for communicating with an external device. Additionally, the communication unit (110) may receive content from an external device. Here, the content may be a video composed of a plurality of video frames (hereinafter referred to as “frames”). Additionally, the communication unit (110) may transmit a supplementary result or an edited video generated using the supplementary result to an external device.
[0036] The control unit (120) is composed of one or more processors and can control the overall operation of the device (100).
[0037] The memory (130) can store a program or other instructions for the operation of the device (100).
[0038] Additionally, the memory (130) may store an artificial intelligence model, and the artificial intelligence model may include a shot detection model (131) and an object recognition model (132). The artificial intelligence model may be implemented by various known artificial intelligence algorithms. One or more commands constituting the artificial intelligence model may be stored in the memory (130). The artificial intelligence model is read out and executed by the control unit (120), and the operation of the artificial intelligence model described below can also be viewed as the operation of the control unit (120).
[0039] A shot detection model (131) can extract a shot from the content. Here, a shot is the smallest unit of video data obtained from continuous camera movement without editing by the video producer, and can be distinguished from a scene, which is formed by combining several consecutive shots that have the same meaning.
[0040] Furthermore, a shot is a section of the screen where the camera composition is maintained, and if the composition changes, it can be distinguished as a different shot. For example, if the video composition changes (e.g., when the camera capturing the same object switches from the left camera to the right camera), the video captured by the left camera and the video captured by the right camera are extracted as separate shots.
[0041] However, the term "shot" in this specification is not limited thereto, and if the position of an object in the image changes abruptly, the image before the change and the image after the change may be extracted as separate shots.
[0042] A shot can be detected by various known algorithms. For example, the shot detection model (131) can be configured as a shot boundary detection model. In this case, the shot detection model (131) can extract the boundaries of a shot from the content and extract one or more shots using the boundaries of the shot. Here, a shot can be composed of multiple consecutive frames. For example, a shot with a time length of 10 seconds can be composed of 200 consecutive frames.
[0043] In this specification, an object may be a human face. However, it is not limited thereto, and an object may be a person, a part of a person's body (e.g., upper body), an animal, a plant, a part of an animal (e.g., an animal's face), etc.
[0044] The object recognition model (132) can recognize and detect objects (e.g., human faces) from a frame. The object recognition model (132) can be composed of various known algorithms, for example, the object recognition model (132) can be an MTCNN-based face recognition model.
[0045] The object recognition model (132) can recognize an object within a frame and output a detection area indicating the location or region of the object. For example, the object recognition model (132) can output a bounding box surrounding the object as a detection area. However, it is not limited to this, and it is also possible to output the area of the object, the boundary of the object, etc.
[0046] FIG. 2 is a flowchart for explaining a method for generating an edited image according to the present invention.
[0047] The method for generating an edited video according to the present invention may include the steps of: obtaining a detection result in which an object is detected in a plurality of frames constituting a video unit (S210); obtaining a reference point of an important object in a video unit using the detection result (S220); generating a supplementary result by supplementing the detection result using the reference point (S230); and generating an edited video including an important object using the supplementary result (S240).
[0048] FIG. 3 is a diagram illustrating a detection area in which an object is detected in a frame according to the present invention.
[0049] The control unit (120) can obtain detection results of detecting objects in multiple frames constituting a single unit of video (S210).
[0050] Here, a single video unit may refer to the previously described shot, but is not limited thereto; videos classified according to scenes, content, or other criteria may also be used as a single video unit. However, it may be preferable to use a shot with minimal change in a person's position as a single video unit.
[0051] The control unit (120) can detect objects from multiple frames constituting a unit of video using an object recognition model (132).
[0052] Specifically, the object recognition model (132) can detect a first object (311) from a first frame (310) among a plurality of frames constituting a unit of video, and output a first detection area (321) indicating the location or area of the first object (311). Additionally, the object recognition model (132) can detect a second object (312) from a first frame (310) and output a second detection area (322) indicating the location or area of the second object (312).
[0053] By repeating the same operation for multiple frames constituting a unit of video, the control unit (120) can obtain a detection result in which an object is detected in multiple frames constituting a unit of video.
[0054] Here, the detection result may include detection regions representing one or more objects existing in multiple frames, such as one or more detection regions in a first frame and one or more detection regions in a second frame. For example, if a first object and a second object exist in the first frame, two detection regions may be acquired; if the first object exists in the second frame, one detection region may be acquired; and if the second object exists in the third frame, one detection region may be acquired. The detection regions thus acquired may constitute a single detection result.
[0055] It should be kept in mind that, assuming one detection area is acquired in the first frame and one detection area is acquired in the second frame, the device (100) does not know whether the detection area acquired in the first frame and the detection area acquired in the second frame are pointing to the same object. In one case, the detection area in the first frame may be the result of detecting the main character A, and the detection area in the second frame may be the result of detecting the supporting character B. In another case, the detection area in the first frame may be the result of detecting the main character A, and the detection area in the second frame may also be the result of detecting the main character A. However, the device (100) cannot know whether the two detection areas are pointing to the same object or to different objects.
[0056] It should also be kept in mind that the detection area output by the object recognition model (132) is not necessarily accurate. For example, the object recognition model (132) may detect protagonist A in the first frame and output a detection area, but may not output a detection area in the second frame because it fails to detect protagonist A. As another example, the object recognition model (132) may output a detection area by misidentifying the background as an object, even though there is no target object (e.g., a human face) in the frame.
[0057] Next, the control unit (120) can obtain a reference point for an important object in a unit of video using the detection result of detecting an object in a plurality of frames (S220).
[0058] With respect to S220, as previously explained, multiple objects may appear in multiple frames that constitute a single unit of video. In order to generate an edited video centered on an important object, the control unit (120) may select an important object among multiple objects detected across multiple frames. Here, an important object is an object that must be included when generating an edited video later, and generally, it may be the protagonist or main character of the content, or a character that needs to be highlighted in the corresponding shot. A method for selecting an important object will be explained with reference to FIGS. 4 and 5.
[0059] FIGS. 4 and FIGS. 5 are drawings for explaining a method of selecting important objects according to the present invention.
[0060] Referring to FIG. 4, the control unit (120) can divide a plurality of frames (410, 420, 430, 440) constituting a unit of video into a plurality of zones (a, b, c, d, e). For convenience of explanation, the number of frames constituting a unit of video is shown as four, but is not limited thereto.
[0061] Here, multiple zones (a, b, c, d, e) can be distinguished by zone boundaries extending in the vertical direction of the frame. This is the result of considering that multiple objects (e.g., people) are located much more frequently in the left-right direction than in the up-down direction on the frame.
[0062] Although the number of zones in FIG. 4 is depicted as five, this is for convenience of explanation, and the frame may be divided into fewer or more zones. In addition, the criteria for distinguishing multiple zones (a, b, c, d, e) can be applied equally to multiple frames (410, 420, 430, 440) that constitute a single unit of video.
[0063] Multiple sections may have the same width. For example, if the width of the frame is 1900, there may be 19 sections with a width of 100.
[0064] Referring to FIG. 5, the control unit (120) can determine which area the detection area is located in. For example, the control unit (120) can determine the center point of the detection area and the area where the center point is located. For example, referring to the first frame (410) of FIG. 5a, the detection area (411) and the center point (511) of the detection area are shown. In this case, the control unit (120) can determine that the detection area is located in the fourth area (d) based on the fact that the center point (511) of the detection area is located in the fourth area (d).
[0065] The control unit (120) can calculate the number of detection areas detected across multiple frames for each zone.
[0066] Specifically, the control unit (120) can calculate the number of detection areas detected across multiple frames for the first area (a). More specifically, no detection area is located in the first area (a) of the first frame (410), no detection area is located in the first area (a) of the second frame (420), a detection area (431) (more specifically, the center point (531) of the detection area) is located in the first area (a) of the third frame (430), and no detection area is located in the first area (a) of the fourth frame (440). In this case, the control unit (120) can calculate the number of detection areas detected in the first area (a) as 1.
[0067] As another example, the control unit (120) can calculate the number of detection areas detected across multiple frames for the second area (b). More specifically, the detection area (421) (more specifically, the center point (521) of the detection area) is located in the second area (b) of the second frame (420), the detection area (441) (more specifically, the center point (541) of the detection area) is located in the second area (b) of the fourth frame (440), the detection area does not exist in the second area (b) of the first frame (410), and the detection area does not exist in the second area (b) of the third frame (430). In this case, the control unit (120) can calculate the number of detection areas detected in the second area (b) as 2.
[0068] In another example, the control unit (120) can calculate the number of detection areas detected across multiple frames for the fourth area (d). More specifically, the detection area (411) (more specifically, the center point (511) of the detection area) is located in the fourth area (d) of the first frame (410), the detection area (422) (more specifically, the center point (522) of the detection area) is located in the fourth area (d) of the second frame (420), the detection area does not exist in the fourth area (d) of the third frame (430), and the detection area does not exist in the fourth area (d) of the fourth frame (440). In this case, the control unit (120) can calculate the number of detection areas detected in the fourth area (d) as 2.
[0069] Meanwhile, when calculating the number of objects per zone, the “zone” may include at least one of a single zone or an integrated zone. Specifically, multiple zones that initially divide the frame may be referred to as a “single zone,” and a zone formed by integrating adjacent single zones may be referred to as an “integrated zone.”
[0070] The control unit (120) can create an integrated area by integrating the second area (b) and the first area (a) when there is a detection area in the second area (b) of the second frame and a detection area in the first area (a) adjacent to the second area (b) of the third frame.
[0071] Specifically, referring to the second frame (420) of FIG. 5b, a detection area (421) is located in the second zone (b). Other zones adjacent to the second zone (b) are the first zone (a) and the third zone (c).
[0072] And with reference to the third frame (430) of FIG. 5c, a detection area (431) is located in the first area (a) adjacent to the second area (b). In this case, the control unit (120) can combine the first area (a) and the second area (b) to create a first-second integrated area (a, b).
[0073] The control unit (120) can calculate the number of detection areas detected across multiple frames for the first-second integrated zone (a, b). For example, as previously described, the first number of detection areas detected in the first single zone (a) is 1, and the second number of detection areas detected in the second single zone (b) is 2. In this case, the control unit (120) can calculate the number of detection areas detected in the first-second integrated zone (a, b) as 3 by adding the first number and the second number.
[0074] As previously explained, the device (100) does not know whether the object detected in the first single zone (a) and the object detected in the second single zone (b) are the same object. However, if the objects are detected in adjacent single zones, the device (100) may treat the objects detected in the first-second combined zone (a, b) as the same object after integrating the first single zone (a) and the second single zone (b).
[0075] In the following description, important zones are selected using the number of detection areas in a single zone and the number of detection areas in an integrated zone, but this is not limited thereto, and the method may also be implemented by selecting important zones based solely on the number of detection areas in a single zone without creating an integrated zone.
[0076] The control unit (120) can select an important area using the number of detection areas per area and select an object within the important area as an important object.
[0077] In one embodiment, the control unit (120) may select the area with the largest number of detection areas among the plurality of areas as the important area. For example, it is assumed that in a plurality of frames within a video unit, the detection area is detected 3 times in the first-2 integrated area (a, b), the detection area is detected 0 times in the third single area (c), the detection area is detected 2 times in the fourth single area (d), and the detection area is detected 0 times in the fifth single area (e). In this case, the control unit (120) may select an object within the first-2 integrated area, which has the largest number of detection areas among the plurality of areas, as the important object.
[0078] For example, referring to FIG. 3, it is assumed that actor A (311) appeared across the first single zone (a) and the second single zone (b) in one shot and appeared for a longer period than actor B (312). In this case, the detection area within the first-second integrated zone (a, b) is counted the most, so the first-second integrated zone (a, b) is selected as the important zone, and actor A (311) within the important zone can be selected as the important object in that shot. That is, based on the pattern in content where the important character appears for a longer period in that shot, the object within the zone with the most counted detection area can be selected as the important object.
[0079] In another embodiment, the control unit (120) can select important zones by further utilizing the size of the detection area within the zone along with the number of detection areas per zone.
[0080] Specifically, the control unit (120) may select a region having the maximum number of detection regions among a plurality of regions as a candidate region. Additionally, the control unit (120) may select a region having a number of detection regions that has a difference of less than or equal to a threshold value from the maximum number as a candidate region. For example, a first-second integrated region (a,b) having the maximum number of detection regions (3) may be selected as a candidate region. Also, assuming the threshold value is 2, a fourth single region (d) having a number of detection regions (2) that has a difference of less than or equal to the threshold value from the maximum number of detection regions (3) may be selected as a candidate region.
[0081] In this case, the control unit (120) may select a candidate area with a larger detection area size among a plurality of candidate areas as an important area. More specifically, the control unit (120) may calculate the average size of one or more detection areas (e.g., bounding boxes) detected in the first candidate area, calculate the average size of one or more detection areas (e.g., bounding boxes) detected in the second candidate area, and select the first candidate area with the larger average size as an important area. Then, the control unit (120) may select an object within the important area (first candidate area) as an important object.
[0082] For example, assume that in a single shot, Actor A appears in the 1st-2nd integrated zone (a, b) and Actor B appears in the 4th zone (o), appearing for similar durations. However, assume that in that shot, Actor A's face is captured closer to the screen on average than Actor B's face. In this case, objects within the 1st-2nd integrated zone (a, b) and objects within the 4th zone (d) are counted the most, and since objects within the 1st-2nd integrated zone (a, b) are detected as larger, Actor A within the 1st-2nd integrated zone (a, b) can be selected as the important object in that shot. In other words, based on the pattern in content where important characters appear more frequently and larger on the screen in a given shot, the object with a larger detection area can be selected as the important object.
[0083] In the present invention, the “important object” can be selected by the fact that it is located within an important area without the device (100) identifying who the object is (for example, without identifying the characteristics of the object to distinguish between person A and person B, and then assigning ID 1 to person A and ID 2 to person B).
[0084] Various characters may appear in a single unit of video. And according to the present invention, by selecting important objects (important characters) based on the number of detected objects and the size of the detection area, there is an advantage in being able to appropriately set important characters that should be the target of tracking in a single unit of video.
[0085] Next, the control unit (120) can obtain a reference point of an important object in a unit of video. Specifically, the control unit (120) can select the average position of an important object across multiple frames as the reference point. Here, the reference point may be referred to as Pseudo Ground Truth.
[0086] For example, it is assumed that a unit of video consists of ten frames, and that the first detection area within the first-second integration zone (a, b) of the first frame is the first position on the screen, the second detection area of the first-second integration zone (a, b) of the second frame is the second position on the screen, the third detection area of the first-second integration zone (a, b) of the fourth frame is the third position on the screen, and the fourth detection area of the first-second integration zone (a, b) of the fifth frame is the fourth position on the screen.
[0087] In this case, the control unit (120) may select the average position of the first to fourth positions as the reference point of the important object. For example, the control unit (120) may calculate the average position of the first to fourth positions by averaging the center point of the first detection area, the center point of the second detection area, the center point of the third detection area, and the center point of the fourth detection area.
[0088] Due to the imperfections of deep learning technology, objects may not be detected in some frames. For example, even if person A actually appears in the first through sixth frames, there may be cases where the object recognition model detects person A only in the first, second, fourth, and fifth frames.
[0089] However, the present invention has the advantage of being able to appropriately set the average position of an object to be tracked despite the imperfections of deep learning technology by selecting the average position of an object detected across multiple frames in an important area as a reference point.
[0090] Next, the control unit (120) can generate a supplementary result by supplementing the detection result using a reference point of an important object (S230).
[0091] As explained above, due to the limitations of current deep learning technology, cases may occur where the background is incorrectly detected as a person, where people are not detected at all, or where people other than important figures are detected together. Therefore, the control unit (120) performs a correction on the detection result generated by the object recognition model in S210. That is, the control unit (120) removes the detection area of the rest while leaving only the detection area of the important object, and performs the operation of generating a detection area in cases where the object recognition model failed to extract a detection area even though an important object exists. Then, the control unit (120) generates a correction result for the detection result using the reference point of the important object obtained earlier.
[0092] In this specification, terms such as “position” and “distance” may refer to a position, distance, etc., on a screen. For example, a plurality of frames constituting a unit to be displayed on a screen of size 1920*1200 may also have a size of 1920*1200. Calculating the distance between a first point of the first frame and a second point of the second frame may mean calculating the distance between the position of the first point on the screen of the first frame (e.g., coordinates (600, 400)) and the position of the second point on the screen of the second frame (e.g., coordinates (500, 300)).
[0093] The control unit (120) can generate a reference position using a reference point and generate a supplementary result in which a part of the detection area within the detection result is deleted and a new detection area is added to the detection result using the reference position.
[0094] When generating a supplementary result, there is a method for fixing the reference position as a reference point and a method for updating the reference position. The method for updating the reference position is explained in FIGS. 9 to 12, and the method for fixing the reference position as a reference point is explained first in FIGS. 7 and 8.
[0095] With the reference position fixed as the reference point, a supplementary result can be generated by deleting some of the detection areas and adding new detection areas. FIG. 7 explains the operation of deleting some of the detection areas, and FIG. 8 explains the operation of adding new detection areas.
[0096] FIG. 6 is a drawing illustrating a reference point of an important object according to the present invention.
[0097] FIG. 7 is a diagram illustrating the detection results generated by an object recognition model according to the present invention.
[0098] The control unit (120) can generate a supplementary result in which some of the detection areas within the detection result are deleted.
[0099] First, the control unit (120) can generate a reference position using the reference point (611 in FIG. 6) calculated in S220. Specifically, the control unit (120) can set the reference point (611 in FIG. 6) as the reference position (611 in FIG. 6).
[0100] In addition, the control unit (120) can select the detection area (721, 741 in FIG. 7) for important objects among the detection areas (711, 721, 722, 731, 741 in FIG. 7) for multiple objects detected across multiple frames, using the distance between the reference position (611) and the detection area for multiple frames that constitute a unit of video.
[0101] Specifically, the control unit (120) can select a detection area located at a distance less than or equal to a threshold value from a reference position (611) as a detection area for an important object.
[0102] In one embodiment, the control unit (120) can set a margin range (621) centered on a reference position (611). Although the margin range (621) is illustrated in FIG. 6 as being in the shape of a square, it is not limited thereto and can be set in various ways, for example, as a circle having a specific radius from the reference position (611). Then, the control unit (120) can determine whether a detection area (more specifically, the center point of the detection area) for a plurality of objects detected across a plurality of frames belongs to the margin range (621). Then, the control unit (120) can select the detection area belonging to the margin range (621) as the detection area for important objects.
[0103] For example, referring to FIG. 7a, it can be seen that in the first frame (710), a first detection area (711) for an object other than the important object is created. The distance of the first detection area (711) of the first frame (710) from the reference position (611) is greater than the threshold value. In this case, the control unit (120) can remove the first detection area (711) from the detection result, where the distance from the reference position (611) is greater than the threshold value.
[0104] Referring to FIG. 7b for another example, it can be seen that in the second frame (720), a first detection area (721) for an important object and a second detection area (722) for another object are created. Since the distance of the second detection area (722) of the second frame from the reference position (611) is greater than the threshold value, the control unit (120) can remove the second detection area (722), where the distance from the reference position (611) is greater than the threshold value, from the detection results. On the other hand, since the distance of the first detection area (721) of the second frame from the reference position is less than or equal to the threshold value, the control unit (120) can select the first detection area (721), where the distance from the reference position is less than or equal to the threshold value. In this case, the second detection area (722) is removed from the detection results, and the first detection area (721) can be maintained as is.
[0105] Referring to FIG. 7c as another example, it can be seen that a first detection area (731) for the background is created as the object recognition model incorrectly determines the background as a person. Since the distance of the first detection area (731) of the third frame from the reference position (611) is greater than the threshold value, the control unit (120) can remove the first detection area (731) from the detection result where the distance from the reference point position (611) is greater than the threshold value.
[0106] Referring to FIG. 7d as another example, it can be seen that in the fourth frame (740), a first detection area (741) for an important object is created. Since the distance from the reference position (611) of the first detection area (741) in the fourth frame is less than or equal to a threshold value, the control unit (120) can select the first detection area (741) for which the distance from the reference position is less than or equal to the threshold value. In this case, the first detection area (721) in the detection result can be maintained as is.
[0107] The supplementary results generated in the process of FIGS. 7a to 7d are summarized as follows. The detection results detected by the object recognition model include the first detection area (711) of the first frame (710), the first detection area (721) and the second detection area (722) of the second frame (720), the first detection area (731) of the third frame (730), and the first detection area (741) of the fourth frame (740). However, the control unit (120) can select the first detection area (721) of the second frame (720) and the first detection area (741) of the fourth frame (740) through the operation described above, and generate a supplementary result including the first detection area (721) of the second frame (720) and the first detection area (741) of the fourth frame (740).
[0108] Next, we will explain the operation of adding a new detection area when a detection area for an important object does not exist.
[0109] FIG. 8 is a diagram illustrating detection results and supplementary results generated by an object recognition model according to the present invention.
[0110] The control unit (120) can generate a supplementary result in which a new detection area is added to the detection result.
[0111] First, the control unit (120) can generate a reference position using the reference point (611 in FIG. 6) calculated in S220. Specifically, the control unit (120) can set the reference point (611 in FIG. 6) as the reference position (611 in FIG. 6).
[0112] Additionally, the control unit (120) can determine whether a detection area for an important object exists for a plurality of frames constituting a unit of video by using a reference position. For a frame in which a detection area for an important object does not exist, the control unit (120) can add a new detection area by using the detection area of the previous frame or the detection area of the subsequent frame. Here, the previous frame may mean a frame one frame prior, and the subsequent frame may mean a frame one frame subsequent, but is not limited thereto.
[0113] For example, FIG. 8a illustrates a first frame (810) within a unit of video, and FIG. 8b illustrates a second frame (820) within a unit of video. It can be seen that the object recognition model detected an object in the first frame (810) and created a detection area (811), whereas in the second frame (820), it failed to detect an object and thus failed to create a detection area.
[0114] Referring to FIG. 8a, the detection area (811) within the first frame (810) has a distance from the reference position (611) that is less than or equal to a threshold value. That is, since the detection area for the important object already exists, the control unit (120) can maintain the detection area (811) of the first frame (810) as is.
[0115] Referring to FIG. 8b, in the second frame (820), there is no detection area where the distance from the reference position (611) is less than or equal to the threshold value, which means that there is no detection area for important objects in the second frame (820). In this case, the control unit (120) can create a new detection area in the second frame (820) using the detection area (811) of the previous frame. Specifically, the control unit (120) can add a new detection area to the second frame (820) that has the same position as the detection area (811) of the previous frame.
[0116] FIG. 8c illustrates a second frame (820) with a new detection area (821) added. That is, the control unit (120) can generate a supplementary result with a new detection area (821) added to the second frame (820).
[0117] The supplementary results generated in the process of FIGS. 8a to 8c are summarized as follows. The detection result detected by the object recognition model is the first detection area (811) of the first frame (810), and there is no detection result in the second frame (820). Then, the control unit (120) can select the first detection area (811) of the first frame (810) through the operation described above and newly generate the second detection area (821) of the second frame (820). Accordingly, the control unit (120) can generate a supplementary result including the first detection area (811) of the first frame (810) and the second detection area (821) of the second frame (820).
[0118] It was previously explained that for a frame in which there is no detection area for important objects, a new detection area is added using the detection area of the previous frame or the detection area of the subsequent frame. The detection area of the previous frame and the detection area of the subsequent frame may be detection areas for important objects. That is, the detection area of the current frame can be generated based on the detection area of the previous frame which is at a distance of less than or equal to the threshold value from the reference position (611), and the detection area of the current frame can be generated based on the detection area of the subsequent frame which is at a distance of less than or equal to the threshold value from the reference position (611 in FIG. 6).
[0119] Meanwhile, the control unit (120) can generate a supplementary result by performing both the operation of deleting some of the detection areas described in FIG. 7 and the operation of adding a new detection area described in FIG. 8. Accordingly, other detection areas excluding the detection area for important objects are removed, and a supplementary result can be generated in which a detection area for important objects is added to a frame in which important objects were not detected.
[0120] Next, a method for sequentially performing a supplementary operation on multiple frames constituting a unit of video while updating a reference position is described in FIGS. 9 to 12.
[0121] FIGS. 9 to 12 are drawings for explaining a method of sequentially performing complementary operations on multiple frames that constitute a single unit of video.
[0122] Performing complementary operations sequentially on multiple frames that constitute a unit of video may mean starting the search from the first frame constituting the unit of video and proceeding toward the last frame to perform the complementary operations.
[0123] The control unit (120) can generate a reference position using the reference point generated in S220. Initially, the reference point generated in S220 can be set as the reference position. The control unit (120) uses the reference position to generate a supplementary result in which a part of the detection area within the detection result is deleted and a new detection area is added to the detection result, and unlike the embodiments in FIGS. 7 and 8, the reference position can be updated in the embodiments of FIGS. 9 to 12.
[0124] The control unit (120) can search for the first frame in which a detection area for an important object exists (i.e., a detection area having a distance less than or equal to a threshold from the current reference position (911)).
[0125] Specifically, referring to FIG. 9, the control unit (120) can first search for the first frame (930) (a frame with an index of 0) among a plurality of frames. More specifically, the control unit (120) can determine whether there is a detection area for an important object in the first frame (930). If there is no detection area for an important object in the first frame (930) (i.e., there is no detection area having a distance less than or equal to a threshold from the current reference position (911)), the control unit (120) can search for the next frame, the second frame (940) (a frame with an index of 1).
[0126] The control unit (120) can determine whether there is a detection area for an important object in the second frame (940). Referring to FIG. 9, since there is a detection area (941) for an important object in the second frame (940) (i.e., there is a detection area (941) having a distance less than or equal to a threshold from the current reference position (911)), the control unit (120) can determine that the first frame in which there is a detection area for an important object is the second frame (940). In this case, referring to FIG. 10, the control unit (120) can select and not delete the detection area (941) for an important object (a detection area (941) having a distance less than or equal to a threshold from the current reference position (911)) within the second frame (940). Additionally, the control unit (120) can delete other detection areas (942 in FIG. 9) (i.e., detection areas where the distance from the current reference position (911) is greater than the threshold value) excluding the detection area for important objects in the second frame (940).
[0127] Meanwhile, the detection area for important objects initially exists in the second frame (940), and there is no detection area for important objects in the first frame (930). For the first frame in which there is no detection area for important objects, the control unit (120) can add a new detection area using the detection area of the subsequent frame. Specifically, referring to FIG. 11, the control unit (120) can add a new detection area (1111) to the first frame (930) that has the same location as the detection area (941) of the subsequent frame (second frame (940)). Additionally, the control unit (120) can delete other detection areas (931 in FIG. 9) (i.e., detection areas where the distance from the current reference position (911) is greater than the threshold value) excluding the detection area for important objects in the first frame (930).
[0128] Here, it is explained that the first frame in which the detection area for the important object exists is the second frame (940), and therefore there is only one frame prior to it (the first frame (930)), but this is not limited thereto, and there may be multiple frames prior to the first frame. In this case, the control unit (120) may add a detection area having the same position as the detection area in the first frame in which the detection area for the important object exists for the multiple frames prior to the first frame in which the detection area for the important object exists.
[0129] The reference position may be the location of a detection area having a distance less than or equal to the threshold from the reference position prior to the update. Specifically, the control unit (120) initially set the reference point calculated in S220 as the reference position (911) and determined the first frame (second frame (940)) having a detection area having a distance less than or equal to the threshold from the reference position (911). In this case, the control unit (120) may set the location of the detection area (941) having a distance less than or equal to the threshold from the reference position (911) (more specifically, the center point of the detection area (941)) as the new reference position.
[0130] Next, the control unit (120) can search for a third frame (950) (a frame with an index of 2).
[0131] Specifically, referring to FIG. 12a, the control unit (120) can determine whether there is a detection area for an important object using the current reference position (position of the detection area (941)) in the third frame (950). More specifically, the control unit (120) can determine whether there is a detection area having a distance less than or equal to a threshold value from the current reference position (position of the detection area (941)) for the third frame (950).
[0132] For the third frame (950) in which there is no detection area for important objects, the control unit (120) can add a new detection area using the detection area (941) of the previous frame (940).
[0133] Specifically, referring to FIG. 12b, the control unit (120) can add a new detection area (1121) having the same location as the detection area (941) of the previous frame (940) for the third frame (950) in which no detection area exists.
[0134] Additionally, the control unit (120) can delete a detection area (951 in FIG. 12a) from the third frame (950) where the distance from the current reference position (position of the detection area (941)) is greater than the threshold value.
[0135] Meanwhile, in the third frame (950), no detection area for the important object was found. Therefore, the control unit (120) may not update the current reference position (position of the detection area (941)).
[0136] Next, the control unit (120) can search for a fourth frame (960) (a frame with an index of 3).
[0137] Specifically, with reference to FIG. 12a, the control unit (120) can determine whether there is a detection area for an important object by using the distance between the current reference position (location of the detection area (941)) and the detection area (961) for the fourth frame (960). More specifically, the control unit (120) can determine whether there is a detection area having a distance less than or equal to a threshold value from the current reference position (location of the detection area (941)) for the fourth frame (960).
[0138] In the fourth frame (960), there exists a detection area (961) having a distance less than or equal to a threshold value from the current reference position (location of the detection area (941)). Therefore, the control unit (120) may select and not delete the detection area (961) having a distance less than or equal to a threshold value from the current reference position (location of the detection area (941)).
[0139] Meanwhile, if a detection area having a distance less than or equal to the threshold value exists in a specific frame, the control unit (120) can set the detection area existing in the specific frame as a new reference position. Specifically, in the fourth frame (960), there exists a detection area (961) having a distance less than or equal to the threshold value from the current reference position (the position of the detection area (941)). Therefore, the control unit (120) can set the detection area (961) existing in the fourth frame (960) (more specifically, the center point of the detection area existing in the fourth frame and having a distance less than or equal to the threshold value from the current reference position) as a new reference position. In this way, the reference position can be updated whenever a detection area having a distance less than or equal to the threshold value from the current reference position is found. That is, the reference position may be the location of the detection area having a distance less than or equal to the threshold value from the reference position prior to the update.
[0140] The control unit (120) can sequentially search from the first frame to the last frame of a unit of video to perform a supplementary operation.
[0141] Assuming that there are only four frames (930, 940, 950, 960) in one unit of video, the detection result of the object recognition model illustrated in FIG. 9 includes the first detection result (931) of the first frame (930), the first detection result (941) and the second detection result (942) of the second frame (940), the first detection result (951) of the third frame (950), and the first detection result (961) of the fourth frame (960).
[0142] On the other hand, referring to the supplementary results illustrated in FIG. 12b, in the first frame (930), the existing first detection result (931) was deleted and a new detection result (1111 in FIG. 11) was added. Also, in the second frame (940), the existing second detection result (942) was deleted and the first detection result (941) was maintained. Also, in the third frame (950), the reference first detection result (951) was removed and a new detection result (1121) was added. Also, in the fourth frame (960), the existing first detection result (961) was maintained.
[0143] As a result, when examining the supplementary results of Fig. 12b, it can be seen that all detection areas pointing to other people or backgrounds have been removed, and detection areas appearing in multiple frames point only to the main person. Additionally, while the main person was not detected in some frames within the detection results, the supplementary results of Fig. 12b show that detection areas pointing to the main person exist in all frames.
[0144] In the embodiment described in FIGS. 7 and 8 (an embodiment in which the reference position is fixed as a reference point), under the premise that the reference point is the average position of the important object, detection areas far from the reference point are removed frame by frame, and if there is no detection area close to the reference point in the frame, a detection area is created by referring to the detection area of the previous or subsequent frame. Accordingly, detection areas pointing only to the important object are created without omission in all frames, thereby enabling the creation of an edited video containing the important object without error.
[0145] In the embodiment described in FIGS. 9 to 12 (an embodiment for updating the reference position), even though there is no detection area in which an important object is detected in some frames due to the incompleteness of the object recognition model, a detection area near the reference point (the average position of the important object in multiple frames) is found by sequentially searching from the preceding frame, under the premise that a frame in which an important object is detected must exist. Additionally, under the premise that the important object cannot move far in adjacent frames, detection areas far from the current reference position are removed, and if there is no detection area close to the current reference position, a detection area is created by referencing the detection area of the previous or subsequent frame. Accordingly, detection areas pointing to the important object can be selected by reflecting even the movement of the important object, and detection areas pointing only to the important object are created in every frame without omission, thereby enabling the creation of an edited video containing the important object without error.
[0146] Next, the control unit (120) can generate an edited video containing an important object using the supplementation result (S240). Here, the edited video containing an important object may be a video in which the important object is placed at the center (or near the center) of the video and the important object is highlighted (for example, the important object appears as the main character).
[0147] Specifically, the control unit (120) can generate an edited image in which the detection area is placed at the center (or near the center) by using a plurality of detection areas within the frame of the supplementary result.
[0148] For example, the control unit (120) can generate an edited video by cropping a unit of video. At this time, the control unit (120) can generate the first frame of the edited video by cropping the first frame so that the detection area of the first frame in the supplementary result is placed in the center, and generate the second frame of the edited video by cropping the second frame so that the detection area of the second frame in the supplementary result is placed in the center.
[0149] As an example of an edited video, while a single unit of video is a horizontal video (i.e., a video where the horizontal length of the video is longer than the vertical length), the edited video may be a vertical video (i.e., a video where the vertical length of the video is longer than the horizontal length). Additionally, the control unit (120) can crop a single unit of video to fit a size set by the user, and set the crop area so that the detection area is centered. Additionally, the control unit (120) can crop a single unit of video without screen shaking according to the Focus View algorithm. Accordingly, an edited video can be generated that tracks a key person within a single unit of video and prevents screen shaking.
[0150] When generating an edited video according to the prior art, the quality of the edited video may be lower due to the imperfections of the deep learning technology. For example, according to the prior art, an edited video would be generated using the detection results of the object recognition model illustrated in FIG. 9. In this case, the first frame of the edited video would focus on another person (the person detected by 931 in FIG. 9) instead of the main person, the second frame of the edited video would focus on one of the two people, and the third frame of the edited video would focus on the background. In contrast, according to the present invention, there is an advantage in that an edited video can be generated by accurately tracking only the main person within a single video unit. Furthermore, while the prior art may have frames where tracking is interrupted because there is no detection area at all, the present invention can provide uninterrupted tracking because a detection area exists in every frame.
[0151] Existing deep learning-based methods for tracking people track the entire body rather than the face, which presented a problem where tracking was impossible when the entire body was not visible. This disadvantage was particularly pronounced in content where the entire body is often not visible, such as in dramas, and where video is edited primarily around the actor's face. Although a centroid-tracking algorithm exists as a means to solve this, the algorithm had the problem that video could not be generated automatically because the operator had to set an ID for every first frame of a video shot. However, according to the present invention, the object is designated as a human face, an object recognition model (132) is trained to recognize and detect the human face, and the actions described above can be performed on the human face, thus providing the advantage of improving the quality of the edited video. Furthermore, the present invention solves the problem of the operator having to select an ID for each shot every time, and has the advantage of being able to perform video editing for various content and a large number of shots within the content very quickly and accurately in a fully automated manner without human manual intervention.
[0152] According to the present invention, by introducing a reference point (Pseudo Ground Truth (Pseudo GT)), important objects can be robustly detected from noisy detection results, and tracking based on important objects can be enabled by supplementing noisy detection results.
[0153] In addition, according to the present invention, content produced in horizontal video format is edited into vertical video format and implemented using automated technology. Therefore, it is possible to enable automated editing of not only dramas but also entertainment content, and to provide an automated system that reduces the labor previously performed manually by video editors.
[0154] The foregoing invention may be implemented as computer-readable code on a medium on which a program is recorded. A computer-readable medium includes all types of recording devices in which data that can be read by a computer system is stored. Examples of computer-readable media include HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc. Additionally, the computer may include a control unit. Accordingly, the above detailed description should not be interpreted restrictively in all respects and should be considered exemplary. The scope of the invention should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the invention are included within the scope of the invention.
Claims
Claim 1 A method for generating an edited video, comprising: a step of obtaining a detection result in which an object is detected in a plurality of frames constituting a unit of video; a step of obtaining a reference point of an important object in the unit of video using the detection result; and a step of generating a supplementary result that supplements the detection result using the reference point. Claim 2 A method for generating an edited image, further comprising the step of generating an edited image including the important object using the supplementary result in claim 1. Claim 3 In claim 1, the step of obtaining a reference point of the important object comprises the step of selecting an important object among a plurality of objects detected across the plurality of frames; a method for generating an edited image. Claim 4 In claim 3, each of the plurality of frames is divided into a plurality of zones, and the step of selecting the important object comprises: a step of calculating the number of detection areas detected across the plurality of frames for each zone; a step of selecting an important zone using the number of detection areas for each zone; and a step of selecting an object within the important zone as the important object. Claim 5 In claim 4, the step of selecting an important area using the number of detection areas per zone comprises the step of selecting the important area using the size of the detection area within the zone in addition to the number of detection areas per zone; a method for generating an edited image. Claim 6 In claim 4, the step of calculating the number of detection areas detected across the plurality of frames for each of the above zones comprises: a step of integrating the first zone and the second zone to create an integrated zone when a detection area exists in the first zone of the first frame and a detection area exists in the second zone adjacent to the first zone of the second frame; and a step of calculating the number of detection areas detected across the plurality of frames in the integrated zone; a method for generating an edited image. Claim 7 In claim 3, the step of obtaining a reference point of the important object further comprises the step of selecting the average position of the important object across the plurality of frames as the reference point; a method for generating an edited video. Claim 8 A method for generating an edited image according to claim 1, wherein the step of generating the supplementary result comprises: generating a reference position using the reference point, and generating the supplementary result in which a part of the detection area within the detection result is deleted and a new detection area is added to the detection result using the reference position. Claim 9 In claim 8, the above reference position is the position of a detection area having a distance less than or equal to a threshold value from the reference position before update, a method for generating an edited image. Claim 10 In claim 9, the step of generating the supplementary result in which a part of the detection area within the detection result is deleted and a new detection area is added to the detection result using the reference position comprises: setting the reference point as the reference position, determining the first frame having a detection area having a distance less than or equal to the reference position and a threshold value, and setting the position of the detection area having a distance less than or equal to the reference position and a threshold value as the new reference position; a method for generating an edited image. Claim 11 In claim 9, the step of generating the supplementary result in which a part of the detection area within the detection result is deleted and a new detection area is added to the detection result using the reference position comprises: a step of setting the location of the detection area existing in the specific frame as the new reference position when a detection area having a distance less than or equal to a threshold value exists in the specific frame. Claim 12 In claim 8, the step of generating the supplementary result in which a part of the detection area within the detection result is deleted and a new detection area is added to the detection result using the reference position comprises: a step of selecting a detection area for the important object among the detection areas for a plurality of objects detected across the plurality of frames using the distance between the reference position and the detection area for the plurality of frames; a method for generating an edited image. Claim 13 In claim 12, the step of generating the supplementary result, in which a part of the detection area within the detection result is deleted and a new detection area is added to the detection result using the reference position, further comprises the step of generating the supplementary result by deleting a detection area from the detection result where the distance from the reference position is greater than a threshold value. Claim 14 In claim 8, the step of generating the supplementary result in which a part of the detection area within the detection result is deleted and a new detection area is added to the detection result using the reference position comprises: generating a reference position using the reference point; determining whether a detection area for the important object exists for the plurality of frames using the reference position; and adding a new detection area using the detection area of the previous frame or the detection area of the subsequent frame for the frame in which a detection area for the important object does not exist; comprising a method for generating an edited image. Claim 15 In claim 14, the step of adding a new detection area using the detection area of the previous frame or the detection area of the subsequent frame comprises the step of adding a new detection area having the same location as the detection area of the previous frame for a frame in which the detection area does not exist; a method for generating an edited image. Claim 16 An editing video generating device comprising: a communication unit for acquiring content; and a control unit for acquiring a detection result of detecting an object in a plurality of frames constituting a unit video within the content, acquiring a reference point of an important object in the unit video using the detection result, and generating a supplementary result that supplements the detection result using the reference point. Claim 17 In claim 16, the control unit is an editing image generating device that generates an editing image including the important object using the supplementary result. Claim 18 In claim 16, the control unit is an editing image generating device that selects important objects among a plurality of objects detected across the plurality of frames. Claim 19 An editing image generating device according to claim 18, wherein each of the plurality of frames is divided into a plurality of zones, and the control unit calculates the number of detection areas detected across the plurality of frames for each zone, selects an important zone using the number of detection areas for each zone, and selects an object within the important zone as the important object. Claim 20 A computer program stored in a non-transient readable storage medium for executing a method for generating an edited video, comprising: a step of obtaining a detection result in which an object is detected in a plurality of frames constituting a unit of video; a step of obtaining a reference point of an important object in the unit of video using the detection result; and a step of generating a supplementary result that supplements the detection result using the reference point.