Method and apparatus for providing short-form video content through automatic editing of broadcast content
The method and device address the inefficiencies of manual editing by using multi-modal analysis to automatically generate short-form video content from broadcast data, ensuring rapid supply and compatibility across platforms.
Patent Information
- Application Number
- PCT/KR2025/003713
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-28
- Filing Date
- 2025-03-24
- Publication Date
- 2025-12-04
AI Technical Summary
Existing methods for creating short-form video content from broadcast content are cumbersome and inefficient, requiring extensive human review and manual marking, which hinders rapid supply and often misses important editing points, and are not adaptable to different playback platforms.
A method and device that utilize multi-modal analysis of video, audio, and text data to automatically select highlight sections, assign weights based on chat data, and convert frames to a preset ratio for generating short-form video content, compatible with various playback platforms.
Enables rapid supply of meaningful short-form video content that stimulates viewer interest and is playable across different display types and platforms, overcoming the inefficiencies of manual editing.
Smart Images

Figure KR2025003713_04122025_PF_FP_ABST
Abstract
Description
Method and device for providing short-form video content through automatic editing of broadcast content
[0001] The present invention relates to a method and device for providing short-form video content through automatic editing of broadcast content.
[0002] Recently, the demand for concise and short-form video content (e.g., shorts, binge-watching, etc.) based on various social media (SNS) platforms is increasing, and the supply is also rapidly increasing accordingly.
[0003] For consumers who need to multitask and focus on a variety of content, this type of video content is serving as an efficient means, and the viewership is also expanding beyond the digitally-savvy younger generation to include middle-aged and older generations.
[0004] This video content is produced and provided for a variety of purposes, and has the advantage of being easily viewable anytime, anywhere and easy to produce at a low cost, so it is actively utilized as a marketing tool by various companies.
[0005] Therefore, to keep pace with the rapidly increasing demand for video content, technologies need to be developed to enable rapid creation and distribution.
[0006] The background technology of the invention has been prepared to facilitate a better understanding of the present invention. It should not be construed as an admission that the matters described in the background technology of the invention constitute prior art.
[0007] Previously, editors would edit long video data corresponding to the original (hereinafter referred to as "broadcast content") to create video content with a shorter playback time, or in the case of live broadcasts, the broadcaster (streamer) would mark the broadcast time at specific points during the broadcast (such as parts they found interesting, impressive sections, or sections that viewers responded well to) and transmit it to the editors for use in the editing process.
[0008] However, the existing method is not only cumbersome and difficult, requiring extensive review by human (e.g., editors) of all broadcast content, but also hinders the rapid supply of video content to meet growing demand. Furthermore, it's difficult for broadcasters to manually mark each video during a live broadcast, making it easy to miss important editing points.
[0009] Additionally, video content can be individually generated by viewers who select preferred segments while watching a broadcast. In other words, video content containing different segments can be created by various users based on a single broadcast content.
[0010] Accordingly, the inventor of the present invention has invented a method and device capable of performing multi-modal analysis based on various types of data included in broadcast content to extract highlight sections, and automatically editing the highlight sections to create and provide short-form video content.
[0011] Accordingly, the problem to be solved by the present invention is to provide a method and device for providing short-form video content through automatic editing of broadcast content, which can provide short-form video content composed of more meaningful highlight sections by performing analysis based on at least two or more of video data, audio data, and text data included in broadcast content, and selecting highlight sections by assigning weights based on chat data as needed, and stimulate the interest of viewers to induce more access.
[0012] Furthermore, the problem to be solved by the present invention is to provide a method and device for providing short-form video content through automatic editing of broadcast content, which not only enables rapid supply in line with the demand of users (consumers, viewers, etc.) by automatically converting each frame corresponding to highlight sections into at least one screen ratio to generate short-form video content, but also enables the short-form video content to be played regardless of the type of display and / or playback platform of the playback device.
[0013] The tasks of the present invention are not limited to the tasks mentioned above, and other tasks not mentioned will be clearly understood by those skilled in the art from the description below.
[0014] In order to solve the above-described problem, a method for providing short-form video content through automatic editing of broadcast content according to an embodiment of the present invention is provided. The method may include the steps of: obtaining video data, audio data, and text data from broadcast content; outputting analysis information using a multi-modal model having at least two or more of the video data, the audio data, and the text data as inputs; outputting at least one highlight section selected based on a score for each frame of the broadcast content using a learned artificial intelligence model having the analysis information as inputs; converting frames corresponding to each of the at least one highlight section at a preset ratio; and generating short-form video content composed of the at least one highlight section based on the converted frame.
[0015] According to a feature of the present invention, the artificial intelligence model may include a large-scale language model, and the step of inputting the analysis information into the artificial intelligence model may include the step of converting a generation command for selecting at least one highlight section into a prompt that can be input into the large-scale language model; and the step of feeding the broadcast content, the prompt, the analysis information, and the broadcast information into the large-scale language model.
[0016] According to a feature of the present invention, the multi-modal model can generate the analysis information by extracting a feature vector for each frame of the broadcast content based on at least two of the video data, the audio data, and the text data.
[0017] According to a feature of the present invention, the step of converting the frame corresponding to each of the at least one highlight section to a preset ratio may include the step of identifying a target object among at least one object included in each of the frames based on the analysis information; the step of identifying an area of the target object in each of the frames; and the step of cropping each of the frames according to the preset ratio so as to include an area of the target object.
[0018] According to a feature of the present invention, the step of converting the frame corresponding to each of the at least one highlight section to a preset ratio may edit each frame using at least one of an adaptive method, a letterbox method, a stretch method, a pillarbox method, a montage method, an edge blending method, and a generative background extension method.
[0019] According to a feature of the present invention, the artificial intelligence model can calculate a score for each frame based on at least one of the video feature, the audio feature, and the keyword feature of each frame, and select at least one highlight section based on the score for each frame.
[0020] According to a feature of the present invention, when chat data for the broadcast content is further input, the artificial intelligence model extracts chat features for each frame of the broadcast content from the chat data, calculates weights, and assigns the calculated weights to scores for each frame of the broadcast content to select at least one highlight section.
[0021] According to a feature of the present invention, the step of generating the short-form video content may include the steps of: confirming the timestamp of each of the at least one highlight section; generating a video by combining frames corresponding to the confirmed timestamp among all frames converted at the preset ratio in chronological order; and performing rendering on the generated video so that it can be played on at least one preset playback platform.
[0022] According to a feature of the present invention, the text data is provided as an output by inputting the audio data into a voice recognition model, and is generated by extracting only the voice from the audio data and converting it into text, and the preset ratio can be set to at least one of 3:2, 4:3, 5:4, 16:10, 9:16 16:9, 1.85:1, 2.35:1 and 1:1 depending on the screen ratio of the display of the playback device or the playback platform.
[0023] In order to solve the above-described problem, an apparatus for providing short-form video content through automatic editing of broadcast content is provided according to an embodiment of the present invention. The apparatus includes a communication interface; a memory; and a processor operably connected to the communication interface and the memory, wherein the processor is configured to obtain video data, audio data, and text data from broadcast content, and output analysis information using a multi-modal model having at least two or more of the video data, the audio data, and the text data as inputs, and output at least one highlight section selected based on a score for each frame of the broadcast content using a learned artificial intelligence model having the analysis information as input, and then output the at least one highlight section when the at least one highlight section is output, and convert frames corresponding to each of the at least one highlight section at a preset ratio, and generate short-form video content composed of the at least one highlight section based on the converted frame.
[0024] The present invention performs analysis based on at least two of video data, audio data, and text data included in broadcast content, and selects highlight sections by assigning weights based on chat data as needed, thereby providing short-form video content comprised of more meaningful highlight sections, and stimulates the interest of viewers to induce more access.
[0025] In addition, the present invention automatically converts each frame corresponding to highlight sections into at least one aspect ratio to generate short-form video content, thereby enabling rapid supply in line with the demand of users (consumers, viewers, etc.), and has the effect of enabling the short-form video content to be played back regardless of the type of display and / or playback platform of the playback device.
[0026] The effects according to the present invention are not limited to those exemplified above, and more diverse effects are included within the present invention.
[0027] FIG. 1 is a schematic diagram illustrating a system for providing short-form video content through automatic editing of broadcast content according to one embodiment of the present invention.
[0028] FIG. 2 is a block diagram showing the configuration of a service server for providing short-form video content through automatic editing of broadcast content according to one embodiment of the present invention.
[0029] FIG. 3 is a block diagram showing the configuration of a user terminal that uses a broadcast service provided from a service server according to one embodiment of the present invention.
[0030] FIG. 4 is a schematic conceptual diagram illustrating a series of data processing operations for performing automatic editing of broadcast content according to one embodiment of the present invention.
[0031] FIG. 5 is a flowchart schematically illustrating a method for providing short-form video content through automatic editing of broadcast content according to one embodiment of the present invention.
[0032] FIG. 6 is a drawing for explaining a specific operation of selecting a highlight section based on a score for each frame calculated according to one embodiment of the present invention.
[0033] FIG. 7 is a drawing for explaining a specific operation of converting a frame corresponding to each of at least one highlight section at a preset ratio according to one embodiment of the present invention.
[0034] FIG. 8 is a drawing showing an example of a screen ratio that can be converted through automatic editing of broadcast content according to one embodiment of the present invention.
[0035] FIG. 9 is a drawing showing an example of converting the screen ratio of broadcast content to at least one preset screen ratio according to one embodiment of the present invention.
[0036] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below, but may be implemented in various different forms. These embodiments are provided solely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims. In connection with the description of the drawings, similar reference numerals may be used for similar components.
[0037] In this document, the expressions "has," "may have," "includes," or "may include" indicate the presence of a feature (e.g., a number, function, operation, or component such as a part), but do not exclude the presence of additional features.
[0038] In this document, the expressions "A or B," "at least one of A and / or B," or "one or more of A or / and B" can include all possible combinations of the listed items. For example, "A or B," "at least one of A and B," or "at least one of A or B" can all refer to cases where (1) at least one A is included, (2) at least one B is included, or (3) at least one A and at least one B are included.
[0039] The terms "first," "second," "first," or "second," as used herein, may describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, without limiting the components. For example, a first user device and a second user device may represent different user devices, regardless of order or importance. For example, without departing from the scope of the rights set forth in this document, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component.
[0040] When it is said that a component (e.g., a first component) is "(operatively or communicatively) coupled with / to" or "connected to" another component (e.g., a second component), it should be understood that the component is directly coupled to the other component, or can be connected via another component (e.g., a third component). Conversely, when it is said that a component (e.g., a first component) is "directly coupled to" or "directly connected to" another component (e.g., a second component), it should be understood that no other component (e.g., a third component) exists between the first component and the other component.
[0041] The expression "configured to" as used herein can be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" does not necessarily mean something that is "specifically designed to" in hardware. Instead, in some contexts, the expression "a device configured to" can mean that the device, together with other devices or components, is "capable of." For example, the phrase "a processor configured (or set) to perform A, B, and C" may mean a dedicated processor (e.g., an embedded processor) for performing those operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform those operations by executing one or more software programs stored in a memory device.
[0042] The terms used in this document are used only to describe specific embodiments and may not be intended to limit the scope of other embodiments. The singular expression may include the plural expression unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by those of ordinary skill in the art described in this document. Terms defined in general dictionaries among the terms used in this document may be interpreted as having the same or similar meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined in this document. In some cases, even if a term is defined in this document, it cannot be interpreted to exclude the embodiments of this document.
[0043] The individual features of the various embodiments of the present invention can be partially or wholly combined or combined with each other, and as can be fully understood by those skilled in the art, various technical connections and operations are possible, and each embodiment can be implemented independently of each other or can be implemented together in a related relationship.
[0044] For clarity in the interpretation of this specification, the terms used in this specification are defined below.
[0045] For clarity in the interpretation of this specification, the terms used in this specification are defined below.
[0046] The device referred to as “service server” hereinafter may mean, but is not limited to, a single physically independent server according to the present invention, and may also be a single virtual machine, and is intended to encompass a single module or program or docker operating on a single virtual or physical machine.
[0047] The term "broadcast service" as used herein refers to services provided by a service server and may encompass a variety of services based on broadcast content. For example, it may include services that broadcast broadcast content in real time or broadcast it after it has been uploaded or produced, and services that create and / or broadcast short-form video content based on the broadcast content. Furthermore, it may include services that collect and / or manage response information regarding the broadcast content and provide it. Furthermore, it may include services that support sponsorship and / or payments based on the broadcast content. In other words, the broadcast service may encompass not only services that simply broadcast broadcast content, but also derivative services that may be provided based on the broadcast content, and is not limited thereto.
[0048] As used herein, the term "Large Language Model (LLM)" may refer to a language model capable of performing Natural Language Processing (NLP) tasks. The large language model may be fine-tuned based on at least one pre-stored data set or trained based on analysis information for each pre-stored broadcast content, and may be trained to output at least one highlight section for a specific broadcast content in response to a request based on the specific broadcast content.
[0049] Here, the analysis information for each pre-stored broadcast content may be information / data generated by performing multi-modal analysis using at least two data among video data, audio data, and text data extracted from each broadcast content. However, during learning, the video data, audio data, and text data extracted from each broadcast content may all be converted to text and used. In addition, at least one highlighted section may be provided as is to at least one user, or a short-form video content composed of these sections may be generated and provided to at least one user.
[0050] Hereinafter, the present invention will be described in detail by describing a preferred embodiment of the present invention with reference to the attached drawings.
[0051] FIG. 1 is a schematic diagram illustrating a system for providing short-form video content through automatic editing of broadcast content according to one embodiment of the present invention.
[0052] Referring to FIG. 1, a short-form video content provision system (hereinafter referred to as a "service provision system") (1000) based on broadcast content according to one embodiment of the present invention may include a service server (100), an administrator terminal (200), and a viewer terminal (300). In this case, the administrator terminal (200) and the viewer terminal (300) are terminals of users who use the broadcast service provided by the service server (100), and may be collectively referred to as user terminals.
[0053] Meanwhile, although the artificial intelligence model (150) is illustrated separately from the service server (100) in FIG. 1, this is only for convenience of explanation, and the artificial intelligence model (150) may be stored in the service server (100) and operated as a module. However, as another embodiment, the artificial intelligence model (150) may be stored in a separate learning model service server (not shown), and the service server (100) may be linked with the learning model service server. The artificial intelligence model (150) according to an embodiment of the present invention may be composed of a convolutional neural network (CNN), a recurrent neural network (RNN), and / or a vision transformer (ViT), and this is only an example and does not limit the scope of the present invention.
[0054] First, the service server (100) may correspond to a web server and refers to a device for automatically generating short-form video content based on broadcast content in the service providing system of the present invention.
[0055] This service server (100) can be connected to an administrator terminal (200) and a viewer terminal (300) to collect or receive all information / data related to broadcast services. For this purpose, it can be linked with a separate API (Application Programming Interface). The API refers to a collection of screen configurations, various functions, etc. required for application developers to easily develop programs that run on an operating system. By utilizing the API, various information about ongoing broadcast content or multiple video contents (including edited contents, clip contents, etc.) related to the broadcast content can be collected in real time.
[0056] Meanwhile, the service server (100) is connected to an administrator terminal (200) and can provide a broadcasting service by transmitting broadcasting content currently in progress or already produced by the administrator to the viewer in response to a broadcasting service provision request. In other words, the broadcasting content can be provided in real time to the viewer terminal (300) by the service server (100) or can be provided to the viewer terminal (300) after uploading is complete.
[0057] That is, the service server (100) can transmit broadcast content that is being produced or has already been produced by the administrator terminal (200) to the viewer terminal (300) of the viewer who requested provision of broadcast content in real time through streaming.
[0058] To this end, the service server (100) may be configured to include a multi-modal model configured to perform multi-modal analysis based on various types of data included in broadcast content, and at least one artificial intelligence model (150) configured to output at least one highlight section based on analysis information output as a result of the multi-modal analysis. Here, at least one artificial intelligence model (150) is pre-trained and may include a large-scale language model.
[0059] The service server (100) can automatically edit broadcast content using a multi-modal model and / or an artificial intelligence model (150) at the request of the administrator terminal (200), thereby creating short-form video content, and provide the same to the administrator terminal (200) and / or the viewer terminal (300).
[0060] That is, the service server (100) inputs analysis information output from a multi-modal model into an artificial intelligence model (150), receives output information including information on at least one highlight section based on the analysis information from the artificial intelligence model (150), and then converts frames corresponding to each of at least one highlight section at a preset ratio based on the output information to generate short-form video content.
[0061] Meanwhile, the service server (100) is connected to a viewer terminal (300) and can collect or receive response information for at least one viewer who watches (or receives) broadcast content or video content, and manage the same. Here, the response information includes data collected in preset time units during the time the broadcast content or video content is transmitted, and can include data on at least one of the viewer, viewing time, chat, response (for example, an act of expressing positive or negative emotions according to a preset method), and support.
[0062] When the service server (100) inputs various data extracted from broadcast content in a multi-modal model, it inputs this response information together so that weights are given to the analysis result values, thereby enabling analysis information including more accurate analysis results to be output.
[0063] Meanwhile, the administrator terminal (200) may collectively refer to a terminal possessed by an administrator who produces broadcast content among users registered on the service server (100) to utilize the broadcast service. Although only one administrator terminal (200) is illustrated in FIG. 1, it may be configured in multiple units, similar to the viewer terminal (300), and the number and type thereof are not limited.
[0064] Each manager requests provision of a broadcast service to the service server (100), and in response, receives a broadcast service from the service server (100) that transmits broadcast content that is in progress or has been completed by the manager to at least one viewer terminal (300).
[0065] In addition, each manager requests the service server (100) to provide short-form video content for the broadcast content, and in response, receives a broadcast service that provides short-form video content generated based on the broadcast content from the service server (100) and transmits it to at least one viewer terminal (300).
[0066] Meanwhile, the viewer terminal (300) may collectively refer to a terminal possessed by each viewer who is registered on the service server (100) to use the broadcast service and receives and views broadcast content and / or short-form video content.
[0067] Each viewer can be provided with broadcast content or short-form video content through the display of a separate playback device connected or linked to the viewer terminal (300) or a separate platform installed on the viewer terminal (300), and can input reaction information through at least one input method while the broadcast content or short-form video content is displayed and played. At this time, the reaction information may be data generated by input or operation by the viewer (e.g., chat, response, support, etc.) as described above, or may be data recorded by the viewer's viewing or access (e.g., viewing time, number of views, etc.).
[0068] Meanwhile, the administrator terminal (200) and the viewer terminal (300) may be a computer, UMPC (Ultra Mobile PC), workstation, netbook, PDA (Personal Digital Assistants), portable computer, web tablet, wireless phone, mobile phone, smart phone, pad, smart watch, wearable terminal, e-book, PMP (portable multimedia player), portable game console, navigation device, black box or digital camera, other mobile communication terminal, etc., on which a plurality of application programs (i.e., applications) desired by the administrator and / or viewer can be installed and executed, and are not limited to the aforementioned terminals.
[0069] Based on this service provision system (1000), an analysis is performed based on at least two of video data, audio data, and text data included in broadcast content, and when scoring each frame based on the analysis results, highlight sections are selected by assigning weights based on chat data as needed, thereby providing short-form video content composed of more meaningful highlight sections, and stimulating the interest of viewers to induce more access.
[0070] FIG. 2 is a block diagram showing the configuration of a service server for providing short-form video content through automatic editing of broadcast content according to one embodiment of the present invention.
[0071] Referring to FIG. 2, the service server (10) may include a communication interface (110), a memory (120), an I / O interface (130), and a processor (140), and each component may communicate with one another through one or more communication buses or signal lines. Meanwhile, although not shown in FIG. 2, an artificial intelligence model may be stored as a module or stored in a separate learning model service server (not shown) linked to the service server (100).
[0072] The communication interface (110) can be connected to the administrator terminal (200) and the viewer terminal (300) as well as other devices through a wired / wireless communication network to exchange data.
[0073] Meanwhile, the communication interface (110) that enables transmission and reception of such data includes a wired communication port (111) and a wireless circuit (112), wherein the wired communication port (111) may include one or more wired interfaces, for example, Ethernet, Universal Serial Bus (USB), FireWire, etc. In addition, the wireless circuit (112) may transmit and receive data with an external device via an RF signal or an optical signal. In addition, the wireless communication may use at least one of a plurality of communication standards, protocols, and technologies, for example, GSM, EDGE, CDMA, TDMA, Bluetooth, Wi-Fi, VoIP, Wi-MAX, or any other suitable communication protocol.
[0074] The memory (120) can store data for at least one process (algorithm) for providing a broadcasting service or a program reproducing the process. Furthermore, the memory (120) can further store processes for performing other operations, but this is not limited thereto.
[0075] Meanwhile, the memory (120) can store various data used in the service server (100) as well as at least one learning model as needed.
[0076] In various embodiments, the memory (120) may include a volatile or non-volatile storage medium capable of storing various data, commands, and information. For example, the memory (120) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., an SD or XD memory, etc.), RAM, SRAM, ROM, EEPROM, PROM, network storage, cloud, and blockchain database.
[0077] In various embodiments, the memory (120) may store configurations of at least one of an operating system (121), a communication module (122), a user interface module (123), and one or more applications (124).
[0078] An operating system (121) (e.g., embedded operating systems such as LINUX, UNIX, MAC OS, WINDOWS, VxWorks, etc.) may include various software components and drivers to control and manage general system operations (e.g., memory management, storage device control, power management, etc.) and may support communication between various hardware, firmware, and software components.
[0079] The communication module (122) can support communication with other devices through the communication interface (110). The communication module (122) can include various software components for processing data received by the wired communication port (111) or wireless circuit (112) of the communication interface (110).
[0080] The user interface module (123) can receive a viewer's request or input from a keyboard, touch screen, microphone, etc. through an I / O interface (130) and provide a user interface on the display.
[0081] An application (124) may include a program or module configured to be executed by one or more processors (140).
[0082] The I / O interface (130) can connect at least one of input / output devices (not shown) of the service server (100), such as a display, a keyboard, a touch screen, and a microphone, to the user interface module (123). The I / O interface (130) can receive viewer input (e.g., voice input, keyboard input, touch input, etc.) together with the user interface module (123) and process commands according to the received input.
[0083] The processor (140) is connected to a communication interface (110), a memory (120), and an I / O interface (130) to control the overall operation of the service server (100), and can perform various commands through applications or programs stored in the memory (120).
[0084] The processor (140) may correspond to a computing device such as a CPU (Central Processing Unit) or an AP (Application Processor). Furthermore, the processor (140) may be implemented in the form of an integrated chip (IC), such as a SoC (System on Chip) in which various computing devices are integrated. Alternatively, the processor (140) may include a module for calculating an artificial neural network model, such as an NPU (Neural Processing Unit).
[0085] Specifically, the processor (140) obtains video data, audio data, and text data from broadcast content, and outputs analysis information using a multi-modal model that inputs at least two of the obtained video data, audio data, and text data. Thereafter, the processor (140) outputs at least one highlight section selected based on a score for each frame of the broadcast content using a learned artificial intelligence model that inputs the analysis information, and converts a frame corresponding to each of the at least one highlight sections at a preset ratio, and then generates short-form video content composed of at least one highlight section based on the converted frame.
[0086] Here, the multi-modal model can generate analysis information by extracting a feature vector for each frame of broadcast content based on at least two of the input video data, audio data, and text data. At this time, the text data can be output through a speech recognition model that inputs audio data, and may be generated by extracting only the voice from the audio data and converting it into text. In other words, since the text data is based on voice utterances, text data may not exist for frames in the broadcast content where voice utterances do not occur, or in other words, frames that do not contain voice.
[0087] Therefore, when generating analysis information, the multi-modal model can generate analysis information based only on video data and audio data, excluding frames where text data does not exist.
[0088] Meanwhile, when inputting analysis information into the artificial intelligence model (150), the processor (150) may input at least one of broadcast content, broadcast information corresponding to the broadcast content, and a prompt. Here, the broadcast information includes information on at least one of the creation date, registration date, file type, broadcast title, broadcast number, broadcast time, broadcast time, producer, and performer for the broadcast content, and the prompt is required when the artificial intelligence model (150) is a large-scale language model, and includes a sentence requesting the generation of analysis information.
[0089] If the artificial intelligence model (150) is a large-scale language model, the processor (140) may convert a generation command to select at least one highlight section when inputting analysis information into the artificial intelligence model into a prompt that can be input into the large-scale language model, and then input broadcast content, broadcast information corresponding to the broadcast content, and a prompt together into the artificial intelligence model (150) to request generation of analysis information. Here, the generation command may be provided in response to a request from the administrator terminal (200) and may include extraction conditions for generating short-form video content, i.e., for extracting at least one highlight section.
[0090] For example, when an administrator requests the service server (100) to create short-form video content for broadcast content, if a sentence such as “Edit the broadcast content to create a 1-minute video to introduce the functions of a dishwasher” is input as a creation command, the processor (140) converts the creation conditions for creating short-form video content into a prompt based on this sentence.
[0091] Meanwhile, the artificial intelligence model (150) can calculate a score for each frame based on the analysis information and select at least one highlight section based on the score for each frame.
[0092] At this time, if there is response information for broadcast content and at least one data that must be considered when selecting a highlight section according to a request by the administrator terminal (200) is set, the processor (140) can calculate a weight by considering the at least one data and assign a score to each frame.
[0093] For example, if the administrator has set chat data to be considered among the data included in the response information, the processor (140) can additionally input chat data included in the response information when inputting analysis information to the artificial intelligence model (150). Accordingly, the artificial intelligence model (150) can extract chat features for each frame of the broadcast content from the chat data, calculate weights, and perform scoring for the corresponding frame based on the feature vector for each frame while assigning the calculated weights.
[0094] Meanwhile, when converting the frames corresponding to each of at least one highlight section at a preset ratio, the processor (140) may identify a target object and an area of the target object among at least one object included in each frame, and then crop each frame according to the preset ratio so as to include the area of the target object. At this time, the processor (140) may not perform an operation of identifying the target object and the area of the target object, and may determine an area for cropping in each frame based on analysis information output from the multi-modal model.
[0095] At this time, various methods can be applied to convert each frame to a preset ratio. For example, each frame can be edited using at least one of the following methods: adaptive method, letterbox method, stretch method, pillarbox method, montage method, edge blending method, and generative background expansion method. However, this is only one example and is not limited to the methods listed above.
[0096] Meanwhile, when generating short-form video content, the processor (140) verifies the timestamp of each of at least one highlight section, and generates a video by combining frames corresponding to the verified timestamps among all frames converted at a preset ratio in chronological order. Thereafter, the processor (140) may perform rendering on the generated video so that it can be played back on at least one preset display device and / or playback platform.
[0097] FIG. 3 is a block diagram illustrating the configuration of a user terminal utilizing a broadcast service provided from a service server according to one embodiment of the present invention. For convenience of explanation, FIG. 3 will be described below based on an administrator terminal (200). However, as previously explained, the user terminal can collectively refer to both the administrator terminal (200) and the viewer terminal (300), and thus the description thereof is also applicable to the viewer terminal (300).
[0098] Referring to FIG. 3, the administrator terminal (200) may include a memory interface (210), one or more processors (220), and a peripheral interface (230). Various components within the administrator terminal (200) may be connected by one or more communication buses or signal lines.
[0099] The memory interface (210) is connected to the memory (250) and can transmit various data to the processor (220). Here, the memory (250) can include at least one type of storage medium among flash memory type, hard disk type, multimedia card micro type, card type memory (e.g., SD or XD memory, etc.), RAM, SRAM, ROM, EEPROM, PROM, network storage, cloud, and blockchain database.
[0100] In various embodiments, the memory (250) may store web / app applications or programs for receiving broadcast services. Furthermore, the memory (250) may store various information about each administrator and / or viewer, as well as broadcast content, and may store various information obtained through the application or program.
[0101] In various embodiments, the memory (250) may store at least one of an operating system (251), a communication module (252), a graphical user interface module (GUI) (253), a sensor processing module (254), a telephone module (255), and an application module (256). Specifically, the operating system (251) may include instructions for processing basic system services and instructions for performing hardware operations. The communication module (252) may communicate with at least one of one or more other devices, computers, and servers. The graphical user interface module (GUI) (253) may process a graphical user interface. The sensor processing module (254) may process sensor-related functions (e.g., processing voice input received through one or more microphones (292). The telephone module (255) may process telephone-related functions. The application module (256) may perform various functions of a user application, such as electronic messaging, web browsing, media processing, navigation, imaging, and other processing functions. In addition, the administrator terminal (200) can store one or more software applications (256-1, 256-2) (e.g., applications for broadcasting services, etc.) associated with a certain type of service in the memory (250).
[0102] In various embodiments, the memory (250) may store a digital assistant client module (257) (hereinafter, DA client module), and accordingly, may store commands for performing client-side functions of the digital assistant and various user data (258) (e.g., user-customized vocabulary data, preference data, other data such as the user's electronic address book, etc.).
[0103] Meanwhile, the DA client module (257) can obtain voice input, text input, touch input, and / or gesture input of an administrator (user) through various user interfaces (e.g., I / O subsystem (240)) provided in the administrator terminal (200).
[0104] Additionally, the DA client module (257) can output data in audiovisual and tactile forms. For example, the DA client module (257) can output data consisting of a combination of at least two or more of voice, sound, notification, text message, menu, graphic, video, animation, and vibration. In addition, the DA client module (257) can communicate with a digital assistant server (not shown) using a communication subsystem (280).
[0105] In various embodiments, the DA client module (257) may collect additional information about the surroundings of the administrator terminal (200) from various sensors, subsystems, and peripheral devices to construct a context associated with the user input. For example, the DA client module (257) may provide context information along with the user input to a digital assistant server to infer the user's intent. Here, context information that may accompany the user input may include sensor information, such as lighting, ambient noise, ambient temperature, images of the surrounding environment, videos, etc. As another example, the context information may include the physical state of the administrator terminal (200) (e.g., device orientation, device position, device temperature, power level, speed, acceleration, motion pattern, cellular signal strength, etc.). As yet another example, the context information may include information related to the software state of the administrator terminal (200) (e.g., processes running on the administrator terminal (200), installed programs, past and present network activity, background services, error logs, resource usage, etc.).
[0106] In various embodiments, the memory (250) may include added or deleted instructions. Furthermore, the administrator terminal (200) may also include additional configurations other than those illustrated in FIG. 3, or may exclude some configurations.
[0107] The processor (220) can control the overall operation of the administrator terminal (200) and can execute various commands to use the broadcasting service provided by the service server (100) by running an application or program stored in the memory (250).
[0108] The processor (220) may correspond to a computing device such as a CPU (Central Processing Unit) or an AP (Application Processor). In addition, the processor (220) may be implemented in the form of an integrated chip (IC), such as a SoC (System on Chip) that integrates various computing devices that perform machine learning, such as an NPU (Neural Processing Unit).
[0109] In various embodiments, the processor (220) may receive or request provision of various notifications, data, information, etc. through a user interface screen.
[0110] The peripheral interface (230) can be connected to various sensors, subsystems, and peripheral devices to provide data so that the administrator terminal (200) can perform various functions. Here, the function performed by the administrator terminal (200) can be understood as being performed by the processor (220).
[0111] The peripheral interface (230) can receive data from a motion sensor (260), a light sensor (light sensor) (261), and a proximity sensor (262), through which the manager terminal (200) can perform orientation, light, and proximity detection functions, etc. For another example, the peripheral interface (230) can receive data from other sensors (263) (positioning system - GPS receiver, temperature sensor, biometric sensor), through which the manager terminal (200) can perform functions related to the other sensors (263).
[0112] In various embodiments, the administrator terminal (200) may include a camera subsystem (270) connected to a peripheral interface (230) and an optical sensor (271) connected thereto, through which the administrator terminal (200) may perform various photographing functions such as taking pictures and recording video clips.
[0113] In various embodiments, the administrator terminal (200) may include a communication subsystem (280) connected to a peripheral interface (230). The communication subsystem (280) may be comprised of one or more wired / wireless networks and may include various communication ports, radio frequency transceivers, and optical transceivers.
[0114] In various embodiments, the administrator terminal (200) includes an audio subsystem (290) connected to the peripheral interface (230), and the audio subsystem (290) includes one or more speakers (291) and one or more microphones (292), thereby enabling the administrator terminal (200) to perform voice-activated functions, such as voice recognition, voice replication, digital recording, and telephone functions.
[0115] In various embodiments, the administrator terminal (200) may include an I / O subsystem (240) connected to a peripheral interface (230). For example, the I / O subsystem (240) may control a touch screen (243) included in the administrator terminal (200) via a touch screen controller (241).
[0116] For example, the touch screen controller (241) may detect a user's contact and movement or cessation of contact and movement using any one of a plurality of touch sensing technologies, such as capacitive, resistive, infrared, surface acoustic wave technology, proximity sensor array, etc. As another example, the I / O subsystem (240) may control other input / control devices (244) included in the administrator terminal (200) via other input controller(s) (242). As an example, the other input controller(s) (242) may control one or more buttons, rocker switches, thumb wheels, infrared ports, USB ports, and pointer devices, such as a stylus.
[0117] FIG. 4 is a schematic conceptual diagram showing a series of data processing operations for performing automatic editing of broadcast content according to one embodiment of the present invention, and automatic editing can be performed by a service server (100).
[0118] Referring to FIG. 4, when broadcast content is input at the request of the administrator terminal (200), the processor (140) extracts video data, audio data, and text data from the broadcast content, thereby obtaining the data, and then inputs the data into the multi-modal model along with at least one of broadcast information and a prompt. Here, the broadcast information may be provided by the administrator terminal (200) upon request, or may be pre-stored in the service server (100).
[0119] Meanwhile, the multimodal model performs multimodal analysis on various input data types based on broadcast information and prompts. This multimodal analysis can utilize at least two of the following: video data, audio data, and text data contained within the broadcast content. Through this multimodal analysis, data on content, subject matter, background elements, aspect ratio, objects, and people can be included in the analysis.
[0120] For example, while data on screen ratio, objects, and people can be identified from video data, further utilization of audio and / or text data allows for the identification of a target object among multiple objects or analysis of its characteristics. In other words, rather than fragmentary recognition of objects or backgrounds, this allows for a more three-dimensional analysis by identifying dynamic changes in the target object.
[0121] Thereafter, the processor (140) stores the analysis information generated by the multi-modal analysis in a database, inputs the analysis information into the artificial intelligence model (150), and performs a cropping operation by extracting the area of the target object from each frame of the broadcast content based on the analysis information.
[0122] That is, the processor (140) converts each frame of the broadcast content to a preset ratio while the artificial intelligence model (150) extracts at least one highlight section based on the analysis information.
[0123] Through this, the processor (140) leaves only the necessary portion of the entire screen of each frame and cuts out the remaining area, judging it as unnecessary. At this time, the processor (140) performs a cropping operation on each frame so as to satisfy a preset ratio.
[0124] For example, the processor (140) can identify the required area of the entire screen of each frame through a request and broadcast information. If it is confirmed through a request by the administrator terminal (200) that the object to be included in the short-form video content is a specific object, and if it is confirmed through the broadcast information that the broadcast content is for selling the specific object, the specific object can be detected in each frame and the area including the specific object can be determined to satisfy a preset ratio. Meanwhile, if it is confirmed through a request by the administrator terminal (200) that the object to be included in the short-form video content is a specific person, and if it is confirmed through the broadcast information that the broadcast content is for selling the specific object, the specific person and the specific object can be detected in each frame and the area including the specific person can be determined to satisfy a preset ratio, but the area can be determined so that at least part of the area is included together depending on whether the specific object is adjacent, or the area for the specific person and the specific object can be determined separately and the specific person and the specific object can be merged so that they are displayed together on a single screen according to the preset ratio. For this purpose, at least one of the adaptive, letterbox, stretch, pillarbox, montage, edge blending, and generative background expansion methods described above may be used. A different method may be selected and used for each frame, and the method may not be applied equally to all frames.
[0125] Meanwhile, the artificial intelligence model (150) extracts at least one highlight section of the broadcast content based on the input analysis information. At this time, if at least one piece of data is set to be additionally considered among the reaction information, the processor (150) may extract at least one piece of data from the reaction information and input it together when inputting the analysis information into the artificial intelligence model (150).
[0126] Accordingly, the processor (140) can calculate a weight for each frame based on at least one piece of data, and select at least one highlight section by assigning a corresponding weight when performing scoring for each frame.
[0127] Thereafter, the processor (140) performs rendering to combine frames corresponding to at least one highlight section selected by the artificial intelligence model (150) among the entire cropped frames, thereby generating short-form video content.
[0128] The above-described FIG. 4 illustrates a case in which a crop operation is performed on each frame included in broadcast content while at least one highlight section is selected from an artificial intelligence model (150), and then short-form video content is generated using only frames (cropped frames) corresponding to at least one highlight section selected by the artificial intelligence model (150). This is an example. That is, after at least one highlight section is extracted from the artificial intelligence model (150), the crop operation may be performed only on the extracted highlight section, but this is not limited thereto.
[0129] Figure 5 is a flowchart schematically illustrating a method for providing short-form video content through automatic editing of broadcast content according to one embodiment of the present invention. Reference will now be made to Figures 6 through 9 for explanation of each step included in Figure 5.
[0130] Referring to FIG. 5, when a request is received from an administrator terminal (200), the processor (140) extracts and acquires video data, audio data, and text data from broadcast content according to the request (S110). At this time, the processor (140) extracts at least one of video data, audio data, and text data for all frames of the broadcast content, and the data extracted for each frame may be stored by being mapped to a timestamp corresponding to each frame.
[0131] Next, the processor (140) outputs analysis information using a multi-modal model that inputs at least two of the image data acquired by step S110, the audio data, and the text data (S120).
[0132] Specifically, the multi-modal model generates analysis information by extracting a feature vector for each frame of broadcast content through correlations between at least two of the input video data, audio data, and text data.
[0133] For example, a multimodal model extracts video features for each frame of broadcast content based on video data, and audio features for each frame of broadcast content based on audio data. Furthermore, if text data exists, text features (keyword features) for each frame of broadcast content are extracted from the text data. Subsequently, a feature vector can be extracted based on the correlation between at least two of the video features, audio features, and text features.
[0134] Next, the processor (140) uses a learned artificial intelligence model that inputs the analysis information acquired by step S120 to output at least one highlight section selected based on the score for each frame of the broadcast content (S130).
[0135] Referring to Fig. 6, as illustrated in (a), a score may be calculated for each frame of broadcast content. Thereafter, based on the calculated score, some sections may be selected as at least one highlight section according to a preset priority or preset threshold. For example, as illustrated in (b), sections with a high score among the scores of the entire section may be selected according to a preset number, or sections with a large change in score over time may be selected according to a preset number based on the change in the score of the entire section. In addition, sections with a score higher than a preset threshold among the entire section may be selected, and in this case, the number of selected sections may also be set. In this case, the number of selected sections may be automatically determined based on the playback time of the short-form video content to be created, or may be preset by an administrator. If a preset threshold is used to select a segment, but the number of segments selected is deemed insufficient to meet the desired playback time of the short-form video content, additional segments can be added before and after the selected segment, or the preset threshold can be lowered to ensure the required number. However, these are merely examples and are not limited to the aforementioned cases, and various conditions may apply.
[0136] Next, the processor (140) converts the frames corresponding to each of at least one highlight section selected by step S130 at a preset ratio (S140). At this time, the processor (140) extracts an area for cropping from the entire screen of the frame corresponding to each highlight section.
[0137] Referring to FIG. 6, the processor (140) determines the size of the crop area (11) according to a preset ratio in the frame as shown in (a). Thereafter, as shown in (b), the processor confirms the area (22) of the target object to be included in the short-form video content among the plurality of objects detected in the frame, and moves the position of the crop area (11) so that the target object area (21), which is the minimum area including the target object, is included, thereby finally determining the area for cropping and performing cropping.
[0138] At this time, the preset ratio can be set to at least one of 3:2, 4:3, 5:4, 16:10, 9:16 16:9, 1.85:1, 2.35:1, and 1:1, depending on the screen ratio of the display of the playback device or the playback platform, as shown in (a) to (i) of FIG. 7. However, the ratios listed above are only examples, and other ratios may be added and are not limited thereto.
[0139] Meanwhile, as explained above, at least one method may be used to edit each frame, and at least one method for editing each frame may be selected in consideration of the screen ratio of the broadcast content, i.e., the original screen ratio and the screen ratio to be converted.
[0140] Specifically, the processor (140) may convert the original screen ratio of each frame to any one of the preset ratios, but in order to enable rapid distribution regardless of the device, the original screen ratio of each frame may be converted to the preset ratios as illustrated in FIG. 8. Accordingly, the service server (100) may generate short-form video content that can be played on various playback devices' displays and / or platforms for one broadcast content and store the generated content in a database.
[0141] Next, the processor (140) generates short-form video content consisting of at least one highlight section based on the frame converted by step S130 (S150).
[0142] At this time, the processor (140) can generate short-form video content by combining all frames converted at a preset ratio by step S140 and arranging them in chronological order based on their respective timestamps.
[0143] As described above, FIG. 5 illustrates a case in which frames corresponding to each highlight section are converted to a preset ratio after at least one highlight section is selected, and corresponds to another embodiment mentioned while explaining FIG. 4. Accordingly, step S140 of converting frames to a preset ratio can be performed simultaneously with step S120, but in this case, since at least one highlight section has not been selected at that point in time, all frames of the broadcast content are converted to a preset ratio.
[0144] As described above, according to the present invention, when video content created by various users such as broadcast streamers, production companies (producers), and viewers exists, by further utilizing this to automatically create short-form video content, productivity and economic efficiency can be improved through rapid supply according to viewer demand.
[0145] Although the embodiments of the present invention have been described in more detail with reference to the attached drawings, the present invention is not necessarily limited to these embodiments, and various modifications may be implemented without departing from the technical spirit of the present invention. Therefore, the embodiments disclosed in the present invention are not intended to limit the technical spirit of the present invention, but to explain it, and the scope of the technical spirit of the present invention is not limited by these embodiments. Therefore, it should be understood that the embodiments described above are illustrative in all aspects and not restrictive. The protection scope of the present invention should be interpreted by the following claims, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of the rights of the present invention.
[0146] [National Research and Development Project Supporting This Invention]
[0147] [Task ID] Not assigned
[0148] [Assignment Number] A1807-24-1011
[0149] [Ministry Name] Ministry of Science and ICT
[0150] [Name of Project Management (Specialist) Agency] National IT Industry Promotion Agency
[0151] [Research Project Name] 2024 Global SaaS Promotion Project (GSIP)
[0152] [Research Project Name] Development and Commercialization of a Global Generative Short-Form Commerce SaaS
[0153] [Name of the project performing organization] Mobidu Co., Ltd.
[0154] [Research Period] May 1, 2024 - December 31, 2024
Claims
1. A method for providing short-form video content through automatic editing of broadcast content performed by a device, A step of acquiring video data, audio data, and text data from broadcast content; A step of outputting analysis information using a multi-modal model that inputs at least two of the image data, the audio data, and the text data; A step of outputting at least one highlight section selected based on the score for each frame of the broadcast content using a learned artificial intelligence model that uses the above analysis information as input; A step of converting a frame corresponding to each of the at least one highlight section at a preset ratio; and A step of generating short-form video content consisting of at least one highlight section based on the converted frame, A method for providing short-form video content through automatic editing of broadcast content.
2. In paragraph 1, The above artificial intelligence model is, Includes large-scale language models, The step of inputting the above analysis information into the above artificial intelligence model is: A step of converting a generation command that selects at least one highlighted section into a prompt that can be input to the large-scale language model; and Comprising a step of feeding the above broadcast content, the above prompt, the above analysis information and the above broadcast information to the large-scale language model, A method for providing short-form video content through automatic editing of broadcast content.
3. In paragraph 1, The above multi-modal model is, Generating the analysis information by extracting a feature vector for each frame of the broadcast content based on at least two of the video data, the audio data, and the text data. A method for providing short-form video content through automatic editing of broadcast content.
4. In paragraph 3, The step of converting the frames corresponding to each of the above at least one highlight section to a preset ratio is: A step of identifying a target object among at least one object included in each of the above frames; A step of confirming the area of the target object in each of the above frames; and Comprising a step of cropping each frame according to the preset ratio to include an area of the target object, A method for providing short-form video content through automatic editing of broadcast content.
5. In paragraph 1, The step of converting the frames corresponding to each of the above at least one highlight section to a preset ratio is: Editing each frame using at least one of an adaptive method, a letterbox method, a stretch method, a pillarbox method, a montage method, an edge blending method, and a generative background extension method. A method for providing short-form video content through automatic editing of broadcast content.
6. In paragraph 3, The above artificial intelligence model is, Based on the above analysis information, a score is calculated for each frame, and at least one highlight section is selected based on the score for each frame. A method for providing short-form video content through automatic editing of broadcast content.
7. In paragraph 6, The above artificial intelligence model is, If additional chat data for the above broadcast content is input, the chat features for each frame of the broadcast content are extracted from the chat data to calculate weights, and the calculated weights are applied to the scores for each frame of the broadcast content to select at least one highlight section. A method for providing short-form video content through automatic editing of broadcast content.
8. In paragraph 1, The steps for creating the above short-form video content are: A step of checking the timestamp of each of the at least one highlight section; A step of generating an image by combining frames corresponding to the confirmed timestamp among the entire frames converted to the preset ratio in chronological order; and Comprising a step of performing rendering on each of the generated images so that they can be played back on at least one preset playback platform, A method for providing short-form video content through automatic editing of broadcast content.
9. In paragraph 1, The above text data is, The above audio data is input into a voice recognition model and provided as output, and is generated by extracting only the voice from the audio data and converting it into text. The above preset ratio is, Depending on the display ratio of the playback device or the screen ratio of the playback platform, it can be set to at least one of 3:2, 4:3, 5:4, 16:10, 9:16 16:9, 1.85:1, 2.35:1 and 1:
1. A method for providing short-form video content through automatic editing of broadcast content.
10. Communication interface; memory; and A processor operably connected to the communication interface and the memory, The above processor, Obtain video data, audio data, and text data from broadcast content; Using a multi-modal model that inputs at least two of the above image data, the above audio data, and the above text data, analysis information is output, Using a learned artificial intelligence model that inputs the above analysis information, if at least one highlight section selected based on the score for each frame of the broadcast content is output, the at least one highlight section is output, Converting the frames corresponding to each of the above at least one highlight section to a preset ratio, Based on the above-mentioned converted frame, a short-form video content composed of at least one highlight section is generated, A device that provides short-form video content through automatic editing of broadcast content.
Citation Information
Patent Citations
PLC system and method for transmitting control data thereof
KR1020240145273A
Device for protecting terminal part of connector
KR1020250117057A
A Quakeproof Device Capable of Dispersing Vibration, Motor Control Board Having the Same and Operating System of Motor Control Board
KR102836903B1
Social media asset portal
US20180367826A1