Video content production apparatus and video content creation method

KR103015694B1Active Publication Date: 2026-09-04SK BROADBAND
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
KR1020240006910
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2026-09-04
Estimated Expiration
2044-01-16

Smart Images

  • Figure 112024005994658-PAT00004_ABST
    Figure 112024005994658-PAT00004_ABST
Patent Text Reader

Abstract

The present invention relates to a method for automatically generating new video content by utilizing a plurality of original video content.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a method for automatically generating new video content by utilizing a plurality of original video content. Background Technology

[0002] With the recent emergence of diverse media such as one-person media, interest in video content production and the demand for various video content production are gradually increasing.

[0003] However, general video content production requires a large workforce, including not only writing a script that fits the concept after establishing the concept, but also cinematographers and editors.

[0004] Therefore, despite the high interest in video content production and the demand for diverse content, the reality is that actual video content production faces significant difficulties, as the tasks required are difficult for an individual to perform alone, as described above.

[0005] Accordingly, the present invention proposes a method to utilize the many original video contents produced and held by broadcasting stations and content production companies for the production of new video content. The problem to be solved

[0006] The present invention was created in consideration of the above-mentioned circumstances, and the objective of the present invention is to automatically generate new video content by utilizing a plurality of original video contents. means of solving the problem

[0007] A video content production device according to one embodiment of the present invention for achieving the above objective comprises: a memory including instructions; and a processor that, by executing the instructions, obtains two or more object segmentation images that match a keyword combination of a content script among object segmentation images identified from a plurality of original video content data sets, and generates new video content by combining the two or more object segmentation images.

[0008] More specifically, the processor can classify the object image based on the class of the object image within the original video content and generate the object segmentation image based on pixel information extracted for each classification of the object image.

[0009] Specifically, the processor can classify the object image into mask images of different colors between classes of the object image in the I frame of the Group Of Picture (GOP), and generate the object segmentation image based on fixed pixel information extracted according to the color of the mask image.

[0010] Specifically, the above data set may include the fixed pixel information and the predicted pixel information which predicts the movement of the object image by classification in at least one of the P frame and the B frame based on the fixed pixel information.

[0011] Specifically, when the original image content between the two or more object segmented images is different, the processor may perform at least one of frame synchronization between the two or more object segmented images and resolution conversion by referring to a dataset of each original image content in which the two or more object segmented images are identified.

[0012] A method for generating video content performed in a video content production device according to an embodiment of the present invention for achieving the above objective comprises: an image acquisition step of acquiring two or more object segmentation images that match a keyword combination of a content script among object segmentation images identified from a plurality of original video content datasets; and a video generation step of generating new video content by combining the two or more object segmentation images.

[0013] Specifically, the above method may further include a data analysis step of classifying the object images based on the class of the object images within the original video content and generating the object segmentation images based on pixel information extracted for each classification of the object images.

[0014] Specifically, the data analysis step can classify the object image into mask images of different colors between classes of the object image in the I frame of the Group Of Picture (GOP), and generate the object segmentation image based on fixed pixel information extracted according to the color of the mask image.

[0015] Specifically, the above data set may further include the above fixed pixel information and predicted pixel information that predicts the movement of the object image by classification in at least one of the P frame and the B frame with reference to the above fixed pixel information.

[0016] Specifically, the image generation step may perform at least one of frame synchronization between the two or more object segmented images and resolution conversion by referring to a dataset of each original image content identified in the two or more object segmented images when the original image content between the two or more object segmented images is different. Effects of the invention

[0017] According to the video content production device and video content generation method of the present invention, object segmentation images by object class are generated from a plurality of original video contents and managed as a data set, and new video content is generated by combining object segmentation images that match keyword combinations of content scripts among the object segmentation images identified from the data set, thereby improving the efficiency of video content production. Brief explanation of the drawing

[0018] FIG. 1 is an exemplary diagram illustrating a video content production environment according to an embodiment of the present invention. FIG. 2 is a configuration diagram for explaining an image content production device according to an embodiment of the present invention. FIG. 3 is an example diagram of original image content for explaining an object image according to an embodiment of the present invention. FIG. 4 is a flowchart for explaining a method for generating video content according to an embodiment of the present invention. Specific details for implementing the invention

[0019] Hereinafter, various embodiments of the present invention will be described with reference to the attached drawings.

[0020] In one embodiment of the present invention, a technique for producing video content is described.

[0021] Amidst the flood of modern information, the number of video contents continues to increase, and in particular, with the recent emergence of diverse media such as one-person media, interest in video content production and the demand for the production of various video contents are gradually rising.

[0022] However, general video content production requires a large workforce, including not only writing a script that fits the concept after establishing the concept, but also cinematographers and editors.

[0023] Therefore, despite the high interest in video content production and the demand for diverse content, the reality is that actual video content production faces significant difficulties, as the tasks required are difficult for an individual to perform alone, as described above.

[0024] Accordingly, in one embodiment of the present invention, new video content is to be automatically generated by utilizing many original video contents produced and held by broadcasting stations and content production companies.

[0025] In this regard, FIG. 1 illustrates an exemplary video content production environment according to one embodiment of the present invention.

[0026] As illustrated in FIG. 1, in an image content production environment according to one embodiment of the present invention, the configuration may include an image content production device (200) that generates new image content by utilizing a plurality of original image contents.

[0027] The video content production device (200) stores each data set of multiple original video content in the video archive (100) in the learning result DB (300), and automatically generates new video content by combining object segmentation images identified from the data sets when necessary.

[0028] Such a video content production device (200) can regenerate new video content using a deep learning-based learning model, and for this purpose, it can be implemented in the form of a computing device or server equipped with software (e.g., an application).

[0029] If the video content production device (200) is implemented in the form of a server, for example, it may be implemented in the form of a web server, a database server, a proxy server, etc., and one or more of various software that enables a network load balancing mechanism or service device to operate on the internet or another network may be installed, and through this, it may also be implemented as a computerized system.

[0030] In the video content production environment according to one embodiment of the present invention, new video content can be automatically generated by utilizing a plurality of original video contents based on the above-described configuration. Below, the configuration of a video content production device (200) for realizing this will be described in more detail.

[0031] FIG. 2 shows the configuration of a video content production device (200) according to one embodiment of the present invention.

[0032] As illustrated in FIG. 2, an image content production device (200) according to one embodiment of the present invention may be configured to include a memory containing instructions and a processor that executes instructions within the memory.

[0033] In particular, in the case of a processor according to one embodiment of the present invention, it may have a functional configuration including a data analysis unit (210), an image acquisition unit (220), and an image generation unit (230) according to the implementation function according to the execution of instructions.

[0034] Above, the video content production device (200) according to the embodiment of the present invention can automatically generate new video content through the functional configuration of the aforementioned processor, and below, a more detailed explanation of each functional configuration for realizing this will be provided.

[0035] The data analysis department (210) is responsible for the function of analyzing original video content.

[0036] More specifically, the data analysis unit (210) analyzes multiple original video contents within the video archive (100) and stores the data set, which is the analysis result for each original video content, in the learning result DB (300).

[0037] At this time, the data analysis unit (210) can generate an object segmentation image from the original video content through analysis of the original video content.

[0038] In this regard, object images are classified based on the class of the object images within the original video content, and object segmentation images are generated based on pixel information extracted for each classification of the object images.

[0039] For example, if the original video content is a dog running in a field, as illustrated in Fig. 3, the object image class can be classified into 'field' and 'dog', and a segmented image of 'field' and a segmented image of 'dog' can be generated using pixel information extracted for each of 'field' and 'dog'.

[0040] To examine this in more detail, the data analysis unit (210) extracts an I-frame from the Group Of Picture (GOP) of the original video content and converts it into a black and white image, and classifies the object images into mask images of different colors between the classes of object images in the I-frame converted into a black and white image.

[0041] Additionally, when an object image is classified using a mask image, the data analysis unit (210) generates an object segmentation image for each object image based on fixed pixel information (e.g., RGB information, resolution, location information) extracted according to the color of the mask image.

[0042] In this way, the data analysis unit (210) derives a data set containing the results of generating object segmentation images as an analysis result for the original video content. In particular, to support the creation of new video content using object segmentation images, this data set further includes predicted pixel information that predicts the movement of object images by classification in at least one of the P frame and the B frame by referencing fixed pixel information extracted from the I frame.

[0043] Meanwhile, the data set according to one embodiment of the present invention may include not only the information described above but also various attribute information, such as, for example, Duration information and identification information (name) for supporting keyword-based object segmentation image search (identification).

[0044] The image acquisition unit (220) is responsible for the function of acquiring an object segmentation image.

[0045] More specifically, the image acquisition unit (200) acquires object segmentation images related to new video content from the learning result DB (300).

[0046] At this time, the image acquisition unit (200) determines a keyword combination from the content script and can acquire two or more object segmentation images that match the keyword combination of the content script among the object segmentation images identified from the data set of the learning result DB (300).

[0047] Here, the fact that keyword combinations and object segmentation images are matched can be interpreted to mean that at least some of the words among the identification information (names) assigned to the dataset to support object segmentation image search (identification) match at least some of the keyword combinations.

[0048] The video generation unit (230) is responsible for the function of generating new video content.

[0049] More specifically, when the video generation unit (230) obtains two or more object segmented images that match the keyword combination of the content script, it generates new video content in the form of combining the two or more object segmented images.

[0050] At this time, the image generation unit (230) determines whether the original image content is the same between two or more object segmented images that match the keyword combination of the content script, and can determine the method of combining the two or more object segmented images according to the result of the determination.

[0051] In this regard, when the original video content between two or more object segmented images that match the keyword combination of the content script is different, the video generation unit (230) refers to the data set of each original video content in which the two or more object segmented images are identified, and performs resolution conversion to overcome the frame synchronization and resolution difference between the two or more object segmented images.

[0052] For reference, frame synchronization can be performed, for example, by generating a new PCR (Program Clock Reference) and synchronizing between object segmentation images based on it, and resolution transformation can be performed, for example, by using interpolation and linear interpolation to add data or utilize neighbor data.

[0053] Meanwhile, the video generation unit (230) can combine the two or more object segmented images as they are without the aforementioned frame synchronization and resolution conversion if the original video content between two or more object segmented images that match the keyword combination of the content script is identical to each other.

[0054] As seen above, according to the configuration of the video content production device (200) according to one embodiment of the present invention, it is possible to generate object segmentation images by object class from a plurality of original video contents and manage them as a data set, and to generate new video content by combining object segmentation images that match the keyword combination of the content script among the object segmentation images identified from the data set, thereby improving the efficiency of video content production.

[0055] Hereinafter, a method for generating video content according to an embodiment of the present invention will be described with reference to FIG. 4.

[0056] For convenience of explanation, in the following description, the video content production device (200) described with reference to FIG. 2 will be referred to as the entity performing the method of creating video content.

[0057] First, the video content production device (200) collects and analyzes multiple original video contents within the video archive (100), and stores the data set, which is the analysis result for each original video content, in the learning result DB (300) (S120-S130).

[0058] At this time, the video content production device (200) can generate an object segmentation image from the original video content through analysis of the original video content.

[0059] In this regard, object images are classified based on the class of the object images within the original video content, and object segmentation images are generated based on pixel information extracted for each classification of the object images.

[0060] For example, if the original video content is a dog running in a field as illustrated in Fig. 3, which was previously exemplified, the object image class can be classified into 'field' and 'dog', and a segmented image of 'field' and a segmented image of 'dog' can be generated using pixel information extracted for each of 'field' and 'dog'.

[0061] To examine this in more detail, the video content production device (200) extracts an I-frame from the Group Of Picture (GOP) of the original video content and converts it into a black and white image, and classifies the object images into mask images of different colors between the classes of object images in the I-frame converted into a black and white image.

[0062] Additionally, when an object image is classified using a mask image, the video content production device (200) generates an object segmentation image for each object image based on fixed pixel information (e.g., RGB information, resolution, location information) extracted according to the color of the mask image.

[0063] In this way, the video content production device (200) derives a data set containing the result of generating an object segmentation image as an analysis result of the original video content, and in particular, to support the creation of new video content using the object segmentation image, this data set further includes predicted pixel information that predicts the movement of the object image by classification in at least one of the P frame and the B frame by referencing fixed pixel information extracted from the I frame.

[0064] Meanwhile, the data set according to one embodiment of the present invention may include not only the information described above but also various attribute information, such as, for example, Duration information and identification information (name) for supporting keyword-based object segmentation image search (identification).

[0065] Next, the video content production device (200) obtains an object segmentation image related to the new video content from the learning result DB (300) (S140).

[0066] At this time, the video content production device (200) determines a keyword combination from the content script and can obtain two or more object segmented images that match the keyword combination of the content script among the object segmented images identified from the data set of the learning result DB (300).

[0067] Here, the fact that keyword combinations and object segmentation images are matched can be interpreted to mean that at least some of the words among the identification information (names) assigned to the dataset to support object segmentation image search (identification) match at least some of the keyword combinations.

[0068] Afterwards, when the video content production device (200) obtains two or more object segmented images that match the keyword combination of the content script, it generates new video content in the form of combining the two or more object segmented images (S150-S170).

[0069] At this time, the video content production device (200) can determine whether the original video content is identical between two or more object segmented images that match the keyword combination of the content script, and determine the method of combining the two or more object segmented images according to the result of the determination.

[0070] In this regard, when the original video content between two or more object segmented images that match the keyword combination of the content script is different, the video content production device (200) refers to the data set of each original video content in which the two or more object segmented images are identified, and performs resolution conversion to overcome the frame synchronization and resolution difference between the two or more object segmented images.

[0071] For reference, frame synchronization can be performed, for example, by generating a new PCR (Program Clock Reference) and synchronizing between object segmentation images based on it, and resolution transformation can be performed, for example, by using interpolation and linear interpolation to add data or utilize neighbor data.

[0072] Meanwhile, the video content production device (200) can combine the two or more object segmented images as they are without the aforementioned frame synchronization and resolution conversion if the original video content between two or more object segmented images that match the keyword combination of the content script is identical to each other.

[0073] As seen above, according to the method for generating video content according to one embodiment of the present invention, it is possible to generate object segmentation images by object class from a plurality of original video contents and manage them as a dataset, and to generate new video content by combining object segmentation images that match the keyword combination of the content script among the object segmentation images identified from the dataset, thereby improving the efficiency of video content production.

[0074] Meanwhile, a method for generating image content according to one embodiment of the present invention may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either individually or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the present invention, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the present invention, and vice versa.

[0075] Although the present invention has been described in detail with reference to preferred embodiments, the present invention is not limited to the above-described embodiments, and the technical concept of the present invention extends to the scope in which various modifications or alterations are possible by anyone with ordinary knowledge in the technical field to which the present invention belongs, without departing from the gist of the present invention as claimed in the following claims. Industrial applicability

[0076] According to the video content production device and video content generation method of the present invention, since new video content can be automatically generated by utilizing multiple original video contents, it overcomes the limitations of existing technology. As such, it is an invention with industrial applicability, as it not only offers sufficient potential for the commercialization or business of the applied device rather than just the use of related technology, but is also practically and clearly implementable. Explanation of the symbols

[0077] 100: Video Archive 200: Video content production device 210: Data Analysis Department 220: Image Acquisition Department 230: Image generation unit 300: Training Result DB

Claims

Claim 1 Memory containing instructions; The present invention includes a processor that, by executing the above command, obtains two or more object segmentation images that match a keyword combination of a content script among object segmentation images identified from a plurality of original video content datasets, and generates new video content by combining the two or more object segmentation images; the processor extracts an I-frame from a Group Of Picture (GOP) based on the class of an object image within the original video content and converts it into a grayscale image; classifies the object image from the I-frame converted into a grayscale image using mask images of different colors between the classes of the object image; generates the object segmentation image based on fixed pixel information extracted according to the color of the mask image; includes the fixed pixel information and predicted pixel information that predicts the movement of the object image by classification in at least one of a P-frame and a B-frame based on the fixed pixel information in the dataset; and if the original video content between the two or more object segmentation images is different, performs at least one of frame synchronization between the two or more object segmentation images and resolution conversion by referring to the dataset of each original video content in which the two or more object segmentation images are identified. Video content production device featuring Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 delete Claim 6 A method for generating video content performed in a video content production device, comprising: an image acquisition step of acquiring two or more object segmentation images that match a keyword combination of a content script among object segmentation images identified from a plurality of original video content datasets; The method comprises a video generation step of generating new video content by combining two or more object segmented images, wherein the method comprises a data analysis step of extracting an I frame from a Group Of Picture (GOP) based on the class of an object image within the original video content and converting it into a black and white image, classifying the object image into a mask image of different colors between the classes of the object image in the I frame converted into a black and white image, and generating the object segmented image based on fixed pixel information extracted according to the color of the mask image, wherein the data set includes the fixed pixel information and predicted pixel information generated by predicting the movement of the object image by classification in at least one of a P frame and a B frame included in the same GOP based on the fixed pixel information, wherein the video generation step is characterized by performing at least one of frame synchronization and resolution conversion between the two or more object segmented images based on the fixed pixel information and predicted pixel information included in the data set of each original video content in which the two or more object segmented images are identified, when the original video content between the two or more object segmented images is different. Claim 7 delete Claim 8 delete Claim 9 delete Claim 10 delete Claim 11 A computer program stored on a medium to execute the method of claim 6, combined with hardware.

Citation Information

Patent Citations

  • Appratus and method for processing image

    KR1020150056381A

  • Encoding a privacy masked image

    KR1020180071957A

  • Video generation method and server performing thereof

    KR102560609B1