Platform for Evaluating Advertising Storyboard
Patent Information
- Application Number
- KR1020250015943
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2026-08-14
Smart Images

Figure PAT00013_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to an advertising storyboard evaluation platform, and more specifically, to an advertising storyboard evaluation platform that automatically generates an advertising storyboard and draft video of an advertising client, which is an order request delivered to creators producing advertising videos, and analyzes the draft video to evaluate whether the content is in line with the latest trends. Background Technology
[0002] Advertisers (e.g., specialized advertising agencies, etc.) can create storyboards and request ad production from ad production companies, creators, etc.
[0003] A storyboard is a video production tool that displays a series of illustrations or images in scene order to visualize each scene according to a script for the purpose of producing videos such as movies, advertisements, and animations.
[0004] Companies commissioning the production of advertising videos need drawing skills or graphic design capabilities to create storyboards in-house, and must possess professional skills and experience in various areas such as camera angles, object composition and placement, and filming techniques.
[0005] However, most advertising video production companies are often unable to create storyboards themselves. Consequently, storyboards are typically drafted by professionals with extensive video production experience, or commissioned from professional storyboard artists or experts in the field.
[0006] In the modern digital advertising market, video content is rapidly spreading through platforms such as YouTube, TikTok, and Social Networking Services (SNS). However, there is a problem in that companies commissioning the production of advertising videos are unable to determine whether the produced video content aligns with the latest trends.
[0007] Furthermore, it is difficult for advertising planners to predict in advance how consumers will actually react to an ad video, and there is a complete lack of tools to determine whether the produced ad video aligns with the latest trends. Prior art literature
[0008] Korean Patent Publication No. 10-2024-0157812 The problem to be solved
[0009] To solve these problems, the present invention aims to provide an advertising storyboard evaluation platform that generates an advertising storyboard and draft video of an advertising client, which is an order request delivered to creators producing advertising videos, and analyzes the draft video to evaluate whether it is content that matches the latest trends. means of solving the problem
[0010] An advertising storyboard evaluation platform according to the features of the present invention for achieving the above objective is,
[0011] An automatic advertising storyboard generation device that receives basic information data including video length and video production purpose, and storyboard configuration data which consists of synopsis text for each cut and reference images for each cut, generates dialogue and narration data for each cut of the advertising storyboard reflecting the synopsis text for each cut, generates scene images for each cut reflecting the reference images for each cut, converts the dialogue and narration data for each cut into speech, and synthesizes them by synchronizing with the scene images for each cut to generate a video file that is an advertising draft; and
[0012] It may include an advertising draft video evaluation device that analyzes a video file received from the advertising storyboard automatic generation device to extract a first core component as a comparison item, collects advertising videos from social media and advertising platforms to extract a second core component as a comparison item, quantitatively evaluates the similarity between the first core component and the second core component, and generates evaluation result information based on whether the advertising video of the video file matches recent trend videos and feeds it back to the advertising storyboard automatic generation device. Effects of the invention
[0013] With the above-described configuration, the present invention automatically generates an advertising storyboard and draft video for an advertising client, thereby simplifying the communication process between the advertising client and the advertising producer and significantly reducing work time.
[0014] The present invention has the effect of efficiently reducing the repetition, modification, and requirements analysis required during the video production request process.
[0015] The present invention evaluates a draft video of an advertising client to determine if it is content that fits the latest trends, provides feedback on the evaluation results to check if trend suitability has improved, and can automatically generate an optimal advertising video. Brief explanation of the drawing
[0016] FIG. 1 is a diagram showing the configuration of an advertising storyboard evaluation platform according to an embodiment of the present invention. FIG. 2 is a diagram showing the configuration of an automatic advertising storyboard generation device according to an embodiment of the present invention. FIG. 3 is a diagram showing the configuration of a storyboard production AI model unit and a video production AI model unit according to an embodiment of the present invention. FIG. 4 is a diagram showing an example of input values of basic information data and storyboard configuration data input to an input data processing unit according to an embodiment of the present invention. FIG. 5 is a drawing showing an example of a user interface screen for inputting the input value of FIG. 4 according to an embodiment of the present invention. FIG. 6 is a diagram showing the configuration of an advertising draft video evaluation device according to an embodiment of the present invention. FIG. 7 is a diagram showing an example of evaluating a draft video according to an embodiment of the present invention and transmitting evaluation result feedback. Specific details for implementing the invention
[0017] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components regardless of drawing symbols are assigned the same reference number, and redundant descriptions thereof will be omitted. Furthermore, in describing the embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description will be omitted.
[0018] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.
[0019] A singular expression includes a plural expression unless the context clearly indicates otherwise.
[0020] In this application, each step described may be performed regardless of the order listed, except where it must be performed in the order listed by a particular causal relationship.
[0021] In this application, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0022] Hereinafter, the advertising storyboard automatic generation device of the present invention will be described with reference to the attached drawings.
[0023] FIG. 1 is a diagram showing the configuration of an advertising storyboard evaluation platform according to an embodiment of the present invention.
[0024] An advertising storyboard evaluation platform (10) according to an embodiment of the present invention may include a creator terminal (20), an advertising storyboard automatic generation device (100), and an advertising draft video evaluation device (200).
[0025] The creator terminal (20) is composed of one or more devices and can receive an advertisement video production request from an advertisement storyboard automatic generation device (100) and transmit quotation information for the advertisement video production request to the advertisement storyboard automatic generation device (100).
[0026] An automatic advertising storyboard generation device (100) receives basic information data including video length and video production purpose, and storyboard composition data which is data for configuring synopsis text for each cut and reference images for each cut, and generates dialogue and narration data for each cut of the advertising storyboard reflecting the synopsis text for each cut, generates scene images for each cut reflecting the reference images for each cut, and converts the dialogue and narration data for each cut into voice and synthesizes them by synchronizing with the scene images for each cut to generate a video file that is an advertising draft.
[0027] The ad draft video evaluation device (200) may include an evaluation AI model that analyzes and evaluates a draft video file received from the ad storyboard automatic generation device (100).
[0028] The ad draft video evaluation device (200) can analyze a video file received from an ad storyboard automatic generation device to extract a first core component that is a comparison item, collect ad videos from social media and ad platforms to extract a second core component that is a comparison item, and quantitatively evaluate the similarity between the first core component and the second core component to generate evaluation result information based on whether the video file matches recent trends and feed it back to the ad storyboard automatic generation device.
[0029] FIG. 2 is a diagram showing the configuration of an automatic advertising storyboard generation device according to an embodiment of the present invention, FIG. 3 is a diagram showing the configuration of a storyboard production AI model unit and a video production AI model unit according to an embodiment of the present invention, FIG. 4 is a diagram showing an example of input values of basic information data and storyboard configuration data input to an input data processing unit according to an embodiment of the present invention, and FIG. 5 is a diagram showing an example of a user interface screen for inputting the input values of FIG. 4 according to an embodiment of the present invention.
[0030] An automatic advertising storyboard generation device (100) according to an embodiment of the present invention may include an input data processing unit (110), a control unit (120), a generative AI processing unit (130), a storage unit (160), a display unit (170), and a communication unit (180). The automatic advertising storyboard generation device (100) may be a device owned by a company that commissioned the production of an advertising video.
[0031] The generative AI (Artificial Intelligence) processing unit (130) may include a storyboard production AI model unit (140) and a video production AI model unit (150).
[0032] As illustrated in FIGS. 4 and 5, the input data processing unit (110) receives basic information data and storyboard configuration data. The basic information data may include video length (6 seconds, 15 seconds, 30 seconds, 60 seconds, etc.), video production purpose information (advertisement, YouTube, SNS, movie, animation, etc.), desired budget (1 million won, 5 million won, 10 million won, etc.), and work schedule (2 weeks, 4 weeks, 8 weeks, etc.). The storyboard configuration data may include synopsis text for each cut and reference images for each cut.
[0033] The storyboard production AI model unit (140) can determine the appropriate number of cuts and advertising style by analyzing basic information data (video length, purpose of video production, desired budget, work schedule), and can extract key elements (background, people, emotions, actions, etc.) for each advertising scene by analyzing synopsis text for each cut and reference images for each cut.
[0034] The storyboard production AI model unit (140) according to an embodiment of the present invention may include a story analysis and cut division unit (141), a cut-by-cut image generation unit (142), and a cut-by-cut editing graphic user interface unit (143).
[0035] The AI model unit (140) for creating storyboards can determine the appropriate number of cuts and advertising style by analyzing basic information data, extract key elements for each advertising scene necessary for image generation by analyzing synopsis text for each cut and reference images for each cut, and generate image data for each cut that represents an advertising storyboard reflecting the extracted key elements.
[0036] The story analysis and cut division unit (141) can divide the basic information data and storyboard configuration data received from the input data processing unit (110) into scenes by applying a Natural Language Processing (NLP) model and / or other analysis methods. For example, the natural language processing model may use natural language processing artificial intelligence language models such as Open API, GPT (Generative Pre-trained Transformer), and BERT (Bidirectional Encoder Representations from Transformers).
[0037] Basic information data can play a role in setting the structure and constraints of the entire storyboard.
[0038] The storyboard production AI model unit (140) can be used as basic data to determine the number of cuts, duration per scene, and screen transition method per cut based on the video length of the basic data, as basic data to determine the direction style, camera angle, and color tone based on the video production purpose information, as basic data to determine the AI automatic generation speed and detail per cut based on the work schedule, and as basic data to determine visual details and additional graphic elements based on the desired budget.
[0039] The number of cuts and directing style may vary depending on whether the video is 15 or 60 seconds long. The production purpose information indicates that the directing style may differ, such as social media advertising (fast tempo) or brand videos (emotional direction). For short schedules, the work schedule focuses on AI automatic generation, while for longer schedules, detailed adjustments for each cut may be possible. Regarding the desired budget, the higher the budget, the more capable the AI can be of generating high-quality images and providing detailed direction.
[0040] Storyboard composition data includes data for configuring images by synopsis (synopsis text per cut and reference data per cut) and can provide key information for generating detailed content for each individual scene.
[0041] When the storyboard production AI model unit (140) explains the elements reflected in the final storyboard in the storyboard composition data, it can use the cut-by-cut synopsis text included in the storyboard composition data as basic data for determining cut descriptions, scene backgrounds, emotions and atmospheres, and the cut-by-cut reference images included in the storyboard composition data as basic data for determining AI image generation, scene styles, camera angles, etc.
[0042] The story analysis and cut division unit (141) recognizes the main entities of the storyboard composition data and analyzes the semantic flow of the sentences to automatically identify the points where the cuts should be divided. For example, "background": "window, cafe, morning sunlight", "character": "female", "emotion": "comfort, relaxation", "action": "drinking coffee, smiling".
[0043] For example, points where cuts need to be divided can be cited as "cut split points": ["a scene of sunlight entering through a window," "a scene of a woman taking a sip of coffee"].
[0044] The story analysis and cut division unit (141) analyzes the synopsis text for each cut of the storyboard composition data using a machine learning-based technology that divides the cuts for each advertising scene, automatically extracts key elements for each cut (e.g., background, character, emotion, action elements, etc.), and generates cut division data in JSON format. Here, the key elements may further include various elements such as cut description, mood, and scene style.
[0045] For example, the AI processing process in the story analysis and cut division section (141) is expressed as an example of cut division data in JSON format as follows.
[0046]
[0047] The story analysis and cut segmentation unit (141) generates cut segmentation data that can describe specific scene descriptions (text) for each cut more naturally and emotionally. For example, "a scene of a woman drinking coffee by the window" can be expanded to "a window where the morning sunlight shines softly, a woman is holding a coffee cup with both hands. Warm sunlight comes in through the window, and a soft coffee aroma spreads." The cut segmentation data in the story analysis and cut segmentation unit (141) can be generated using generative AI models such as GPT-4, Claude, and LLaMA.
[0048] Here, cut segmentation data may represent information that can describe a scene with at least one piece of information, such as image description, background, emotion, cut transition, dialogue, or scene type.
[0049] The cut-by-cut image generation unit (142) can determine the emotions and actions of the main object using cut division data received from the story analysis and cut division unit (141), and detect the location of the main object included in the image, and set the camera composition according to the location, emotions, and actions of the object.
[0050] The AI processing process in the cut-by-cut image generation unit (142) is described as an example of detecting the location of a main subject in JSON format as follows.
[0051]
[0052] The AI processing process in the cut-by-cut image generation unit (142) is expressed as an example of selecting the optimal camera composition in JSON format as follows.
[0053]
[0054] The image generation unit (142) can analyze the cut division data received from the story analysis and cut division unit (141) and set a recommended camera angle according to the scene type.
[0055] In other words, the cut-by-cut image generation unit (142) sets the camera composition to a Close-up and front angle to emphasize the character's emotions in the case of an emotional dialogue scene, sets the camera composition to an Extreme Close-up to emphasize the brand message in the case of product emphasis, sets the camera composition to a Low Angle to make the protagonist look stronger in the case of a scene that heightens tension, and sets the camera composition to a Wide Shot to emphasize harmony with the surrounding environment in the case of a scene that expresses a sense of space.
[0056] The image generation unit (142) for each cut can set a lighting style by analyzing the emotion type from the cut segment data. In other words, the image generation unit (142) sets the lighting style to Soft Natural Light to express emotional advertisements (coffee, healing products) in the case of a warm and emotional atmosphere, sets the lighting style to Hard Key Light (strong contrast lighting) to express action or thriller scenes in the case of a tense atmosphere, sets the lighting style to Low Key Light (dark lighting) to express mystery or fashion advertisements in the case of a dreamy feeling, and sets the lighting style to Spotlight to express high-end product advertisements in the case of advertisements that emphasize products.
[0057] The image generation unit (142) can analyze the emotional type from the cut-divided data and set the color tone to reflect the brand color (e.g., Starbucks uses a green tone, Coca-Cola uses red, etc.) or the emotional effect of the advertisement. In other words, the image generation unit (142) sets a warm orange tone color tone to emphasize the emotional and relaxation effects in the case of an emotional advertisement. In the case of a premium product advertisement, it sets a deep black and gold tone color tone to emphasize the luxurious atmosphere. In the case of a tech and IT advertisement, it sets a cool blue tone color tone to convey a futuristic feeling. In the case of a health and nature advertisement, it sets a green and brown tone color tone to convey a nature-friendly feeling.
[0058] The image generation unit (142) for each cut generates image data for each cut that reflects the set camera angle, lighting style, and color tone.
[0059] To explain in more detail, the cut-by-cut image generation unit (142) analyzes the cut-by-cut synopsis text and the cut-by-cut reference image to automatically extract key elements for each advertising scene (image description, background, emotion, cut transition, dialogue, scene type, etc.) required for image generation, and can generate cut-by-cut image data representing an advertising storyboard that reflects the key elements, camera angle, lighting style, and color tone. The cut-by-cut image data in the cut-by-cut image generation unit (142) can be generated using generative AI models such as Stable Diffusion, DALL-E, Midjourney, and Vision AI.
[0060] The cut-by-cut image generation unit (142) can generate the first final storyboard data by including the cut-by-cut synopsis text, cut-by-cut reference data, and the basic information data in the cut-by-cut image data.
[0061] The AI processing process in the cut-by-cut image generation unit (142) is described as an example of extracting key elements in JSON format as follows.
[0062]
[0063] The AI processing process in the cut-by-cut image generation unit (142) is expressed as an example of cut-by-cut image data in JSON format as follows.
[0064]
[0065] The cut-by-cut image data generated by the cut-by-cut image generation unit (142) may be the final storyboard data.
[0066] The cut-by-cut editing graphic user interface section (143) provides a user interface that visually presents the AI-generated storyboard and supports the user in directly modifying it.
[0067] The cut-by-cut editing graphic user interface unit (143) receives cut-by-cut image data from the cut-by-cut image generation unit (142) and outputs it, and provides a user interface in the form of a drag / drop, slider, or text box for each element so that at least one of the components, such as the cut-by-cut image file, description, camera angle, lighting, color tone, dialogue, and directing technique, can be edited by user input.
[0068] The cut-by-cut editing graphic user interface section (143) can generate a second final storyboard data by editing basic information data and storyboard configuration data with user input. The first and second final storyboard data include basic information data and storyboard configuration data, and in particular, may include at least one piece of information among cut-by-cut image files, descriptions, camera angles, lighting, color tones, dialogue, directing techniques, and brand elements.
[0069] The aforementioned first final storyboard data may be data including cut-by-cut synopsis text, cut-by-cut reference data, and basic information data in the cut-by-cut image data that has not been edited by the user and generated by the cut-by-cut image generation unit (142).
[0070] The second final storyboard data may be data edited by user input, such as basic information data and storyboard configuration data in the cut-by-cut editing graphic user interface section (143).
[0071] The control unit (120) can store the first final storyboard data and the second final storyboard data in the storage unit (160).
[0072] The control unit (120) can output the first final storyboard data and the second final storyboard data in the form of PDF files.
[0073] Examples of the first and second final storyboard data are as follows.
[0074]
[0075] The control unit (120) can transmit the first final storyboard data or the second final storyboard data generated by the cut-by-cut image generation unit (142) or the cut-by-cut editing graphic user interface unit (143) to the video production AI model unit (150). Here, the first and second final storyboard data may be video production work requests.
[0076] The control unit (120) can transmit the first final storyboard data or the second final storyboard data to a plurality of creator terminals (20) through the communication unit (180). The creator terminal (20) may be a terminal device possessed by a creator who produces an advertisement video.
[0077] The communication unit (180) can receive quotation information for advertising video production requests from each creator terminal (20).
[0078] The control unit (120) can output a screen displaying quotation information for each creator terminal (20) to the display unit (170).
[0079] The control unit (120) can select one of the multiple quotation information and generate a contract matching signal and transmit it to the creator terminal (20) corresponding to the selected quotation information.
[0080] The image production AI model unit (150) according to an embodiment of the present invention may include a text generation unit (151), an image generation unit for image production (152), an image generation unit (153), a voice generation unit (154), and a combination unit (155).
[0081] The video production AI model unit (150) can generate cut-by-cut dialogue and narration data of an advertisement storyboard that reflects cut-by-cut synopsis text by analyzing the first and second final storyboard data, generate cut-by-cut scene images that reflect cut-by-cut reference images, and generate cut-by-cut dialogue and narration data by converting them into speech and synthesizing them with cut-by-cut scene images to create a final video file.
[0082] The text generation unit (151) can automatically generate scene-specific dialogue, narration, subtitles, and advertising messages by reflecting the basic information data and storyboard composition data included in the first and second final storyboard data.
[0083] When the text generation unit (151) receives the first and second final storyboard data, it analyzes the basic information data of the received first and second final storyboard data to set the length and tempo of the dialogue according to the video length, analyzes the cut-by-cut synopsis text of the first and second final storyboard composition data to generate dialogue and narration, and analyzes the cut-by-cut reference images of the first and second final storyboard composition data to generate the scene atmosphere and directing style.
[0084] To explain in more detail, the text generation unit (151) generates dialogue and narration based on the description of the synopsis text for each cut, sets a dialogue style by reflecting the camera angle and key elements (background, emotion, action, color tone, etc.) for each advertisement scene, and can generate dialogue and narration data for each cut based on the set dialogue style, scene atmosphere, and directing style.
[0085] Cut-by-cut dialogue and narration data in the text generation unit (151) can be generated using generative AI models such as GPT-4, Claude, and LLaMA.
[0086] The AI processing process in the text generation unit (151) is described as follows, with an example of cut-by-cut dialogue and narration data in JSON format.
[0087]
[0088] The image generation unit (152) for video production can analyze the synopsis text for each cut of the storyboard composition data included in the first and second final storyboard data to extract visual elements (background, character, emotion, lighting, color tone, etc.), and generate image data for each cut based on the reference image for each cut of the storyboard composition data, the extracted visual elements, the video length of the basic information data, and the purpose of video production.
[0089] To explain in more detail, the image generation unit (152) for video production can adjust the scene transition method between images according to the video length of the basic information data, determine the style and composition of the scene based on the scene description for each cut, optimize image generation using camera angle and color information, and smoothly produce the flow between images according to the cut transition method.
[0090] Cut-by-cut image data in the image generation unit (152) for video production can be generated using generative AI models such as Stable Diffusion, DALL-E, and Midjourney.
[0091] The video generation unit (153) generates cut-by-cut video at regular time intervals by applying cut-by-cut image data generated by the video production image generation unit (152), cut-by-cut scene description, camera composition (e.g., "Medium Close-up, 30-degree angle") and cut transition method (e.g., "Fade in").
[0092] The image to video in the image generation unit (153) can be generated using generative AI models such as Runway Gen-2, Make-A-Video, and Pika Labs.
[0093] The AI processing process in the video generation unit (153) is expressed as an example of a cut-by-cut video in JSON format as follows.
[0094]
[0095] The voice generation unit (154) adjusts the length of the voice and background music to match the video length of the basic information data, synthesizes the voice using a TTS (Text to Speech) model for the dialogue and narration data generated by the text generation unit (151) to generate voice data for each cut, analyzes the emotion among the main elements of the image data for each cut to determine the background music style for each cut, and generates background music data that matches the background music style for each cut.
[0096] The voice generation unit (154) combines voice data and background music data for each cut to generate audio files for each cut. For voice generation in the voice generation unit (154), TTS models such as ElevenLabs, OpenAI TTS, and Google WaveNet can be used, and AI models such as AIVA, Mubert, and Meta MusicGen can be used for background music generation.
[0097] The voice generation unit (154) can convert audio files into various formats such as MP3, WAV, and AAC, and can adjust the file length of the audio file so as to synchronize with the video.
[0098] The AI processing process in the voice generation unit (154) is described as an example of generating voice and background music in JSON format as follows.
[0099]
[0100] The combination unit (155) can generate a final video file by synchronizing and synthesizing the cut-by-cut video generated by the video generation unit (153) and the cut-by-cut audio file generated by the voice generation unit (154).
[0101] The control unit (120) can generate advertising video production request information including the first final storyboard data or the second final storyboard data and the final video file, and transmit it to a plurality of creator terminals (20) through the communication unit (180).
[0102] The communication unit (180) can receive quotation information for advertising video production requests from each creator terminal (20).
[0103] The control unit (120) can output a screen displaying quotation information for each creator terminal (20) to the display unit (170).
[0104] The control unit (120) can select one of the multiple quotation information and generate a contract matching signal and transmit it to the creator terminal (20) corresponding to the selected quotation information.
[0105] The advertising video production request information, including the first final storyboard data or the second final storyboard data and the final video file (draft video), may be an order request form (work request form) for requesting the production of an advertising video to a creator terminal (20). The advertising storyboard automatic generation device (100) of the present invention may be an AI model that automatically generates such first final storyboard data or the second final storyboard data and the final video file (draft video).
[0106] FIG. 6 is a diagram showing the configuration of an advertising draft video evaluation device according to an embodiment of the present invention, and FIG. 7 is a diagram showing an example of evaluating a draft video according to an embodiment of the present invention and transmitting evaluation result feedback.
[0107] An advertising draft video evaluation device (200) according to an embodiment of the present invention may include a video upload unit (210), a video analysis AI model unit (220), a trend learning database unit (230), an SNS trend analysis unit (240), a video evaluation unit (250), and a final evaluation feedback provision unit (260).
[0108] The video upload unit (210) is connected to the video production AI model unit (150), receives the final video file, which is an advertisement draft video, from the video production AI model unit (150), and uploads it to the video analysis AI model unit (220).
[0109] The video analysis AI model unit (220) may include a video structure analysis unit (221), a narration analysis unit (222), a color analysis unit (223), and a music sound analysis unit (224).
[0110] The video structure analysis unit (221) analyzes the composition of the video file, which is an advertisement video, by scene and shot, and evaluates the editing style and video development method of the video.
[0111] The video structure analysis unit (221) can analyze the visual flow of an advertisement by recognizing the transition method, duration, and camera angle of each scene. In particular, the video structure analysis unit (221) can extract the video length and cut editing style of the video file as key items for video evaluation.
[0112] In the video structure analysis unit (221), for the analysis of key comparison items (video length, cut editing style, etc.), a CNN (Convolutional Neural Network) based model can be used for Scene Boundary Detection of cut transformations per scene, and an AI model that uses YOLO, Faster R-CNN, etc. to detect key objects within the advertisement video and evaluates the direction and speed of movement of the video to analyze the cut editing style. For example, the total number of scenes may be 6, cut editing style fade in / out 30%, hard cut 70%, camera composition close-up 40%, medium shot 60%, average duration per scene 3 seconds, etc.
[0113] The video structure analysis unit (221) can measure the video length of an uploaded video file by using FFmpeg metadata analysis, which extracts metadata from an advertisement video of an uploaded video to determine the total length, and an OpenCV frame technique, which calculates the total number of frames by dividing it by FPS (Frame Per Second) based on FPS information.
[0114] The video structure analysis unit (221) can analyze the number of cuts and cut editing style, including the scene duration, of an advertisement video uploaded by detecting scene transitions and analyzing cut durations through optical flow analysis, scene boundary detection that detects cut changes using a CNN-based deep learning model, histogram difference analysis that detects cut transitions by analyzing color differences between consecutive frames, and K-means clustering that classifies editing styles by learning similar cut duration patterns.
[0115] The narration analysis unit (222) can extract text (subtitles, advertising phrases, dialogue, narration, etc.) of a video using an OCR (Optical Character Recognition) technique and automatically convert the narration and dialogue using a Speech-to-Text (STT) technique to analyze keywords. For example, it may be an advertising message (“Enjoy this moment, with Starbucks”), a text style (“emotional, warm tone”), and keyword analysis (“Starbucks”, “morning”, “coffee”, etc.).
[0116] The narration analysis unit (222) can analyze the SNS viral elements of challenges and hashtags by using an OCR (Optical Character Recognition) technique that detects subtitles and hashtags within an advertisement video of an uploaded video, a natural language processing-based hashtag analysis technique that calculates the frequency of popular hashtags, and a technique that detects challenge elements using a deep learning model.
[0117] The color analysis unit (223) can evaluate whether the color and production style of the video match the brand and trends.
[0118] The color analysis unit (223) can extract the main colors of the image using a color histogram.
[0119] The color analysis unit (223) can determine the emotion and mood by analyzing the visual tone of the advertisement (bright and lively, calm, etc.) using a CNN-based model. For example, it may be a main color (warm orange, red tone, etc.).
[0120] The color analysis unit (223) can determine the main colors of an advertisement video by using a K-Means Clustering technique to extract 3 to 5 main colors from an advertisement video uploaded by the uploaded video and Hsv history matching to compare color distributions, and can determine whether a filter is applied by detecting brightness and saturation adjustments.
[0121] The music sound analysis unit (224) can extract background music and narration sound effects from the video using the MFCC (Mel-Frequency Cepstral Coefficient) technique.
[0122] The music sound analysis unit (224) can analyze the sound characteristics of background music in the uploaded video advertisement video to determine the genre using the MFCC (Mel Frequency Cepstral Coefficients) technique, analyze the sound source spectrum to detect the rhythm and mood of the background music using the spectrogram analysis technique, measure the BPM (Beats per minute) of the background music using Librosa and Essentia libraries, and analyze the style of the background music using a CNN-based emotion classification model.
[0123] The trend learning database unit (230) can crawl data by linking with various social media and advertising platforms to analyze advertising trends and learn trends through machine learning. Here, social media and advertising platforms may include all platforms capable of collecting various trend data, such as YouTube, TikTok, Instagram, Facebook, Naver, and Google Trends.
[0124] The trend learning database unit (230) needs to analyze the latest trends and preferences, so it collects popular advertising videos from social media and advertising platforms that meet or exceed a certain standard. For example, it can collect advertising videos that have over 1 million views in the last 6 months on YouTube, or advertising videos that have over 100,000 likes and 10,000 shares on TikTok, or advertising videos that have over a certain standard of search volume for advertising-related keywords (e.g., new product advertisement, TikTok ad popularity, etc.) on Naver and Google Trends.
[0125] The trend learning database unit (230) collects advertising videos from social media and advertising platforms and transmits them to the SNS trend analysis unit (240).
[0126] The SNS trend analysis unit (240) can measure the video length of advertising videos on platforms such as YouTube, Shorts, TikTok, and Instagram by using FFmpeg metadata analysis to extract metadata from advertising videos and check the total length, OpenCV frame method to calculate the video length by dividing the total number of frames by FPS based on FPS (Frame Per Second) information, and YouTube API and TikTok API crawling method to collect length data of popular advertising videos and calculate the average value.
[0127] The SNS trend analysis unit (240) can analyze the number of cuts and cut editing style, including the speed of scene transitions, of an advertisement video collected from social media and advertising platforms through optical flow analysis that detects scene transitions and analyzes the duration of cuts, scene boundary detection that detects cut changes using a CNN-based deep learning model, histogram difference analysis that detects cut transitions by analyzing color differences between consecutive frames, and K-means clustering that classifies editing styles by learning similar cut duration patterns.
[0128] The SNS trend analysis unit (240) can determine the main colors of an advertisement video collected from social media and advertising platforms by using a K-Means Clustering technique to extract 3 to 5 main colors from the video and Hsv history matching to compare color distributions, and can determine whether a filter is applied by detecting brightness and saturation adjustments.
[0129] The SNS trend analysis unit (240) can analyze the sound characteristics of the background music in the advertisement video to determine the genre using the MFCC (Mel Frequency Cepstral Coefficients) technique, analyze the sound source spectrum to detect the rhythm and mood of the background music using the spectrogram analysis technique, measure the BPM (Beats per minute) of the background music using the Librosa and Essentia libraries, and analyze the sentiment of the background music using a CNN-based sentiment classification model.
[0130] The SNS trend analysis unit (240) can analyze SNS viral elements of TikTok challenges and popular hashtags by using an OCR (Optical Character Recognition) technique that detects subtitles and hashtags in an advertisement video, a natural language processing-based hashtag analysis technique that calculates the frequency of popular hashtags, and a technique that detects challenge elements using a deep learning model.
[0131] The SNS trend analysis unit (240) can learn the latest trends and preferred advertisements by extracting key components of the video length, cut editing style, color and filter, background music, and SNS viral elements of the most popular advertisement videos on SNS over the past 3 to 6 months. In other words, the SNS trend analysis unit (240) can learn patterns from various advertisements and generate average and representative values for each component.
[0132] The video evaluation unit (250) receives the second core components of the video length, cut editing style, color tone and filter, background music, and SNS viral elements of the advertisement video from the SNS trend analysis unit (240), and receives the first core components of the cut editing style, color tone and filter, background music style, and SNS viral elements of the advertisement video of the video file uploaded from the video upload unit (210) from the video analysis AI model unit (220).
[0133] The video evaluation unit (250) can calculate whether the first core component generated by the video analysis AI model unit (220) and the second core component learned by the SNS trend analysis unit (240) match the trend one-to-one for each component.
[0134] The video evaluation unit (250) calculates a similarity score by comparing the video length of the first core component with the video length of the second core component. The similarity score is a value that quantitatively evaluates the similarity between an advertisement video uploaded by the video analysis AI model unit (220) and a trend video learned by the SNS trend analysis unit (240), and represents the similarity between the two videos as a percentage (%). This similarity score can be calculated in various ways, and a method of measuring similarity between vectors is used. The calculated similarity score is expressed as a value between 0% and 100%, indicating how similar the two videos are.
[0135] Representative similarity calculation methods include similarity calculation algorithms such as Cosine Similarity, Mean Squared Difference (MSD), and Pearson Correlation Coefficient.
[0136] As such, the calculation of similarity scores is a known technique, so a detailed explanation is omitted below.
[0137] In the case where the video evaluation unit (250) has a video length of an advertisement video uploaded from the video analysis AI model unit (220) (e.g., 20 seconds) and a video length of a trend video learned from the SNS trend analysis unit (240) (e.g., 10 to 15 seconds), the video evaluation unit (250) calculates the video length similarity score (first similarity score) as 70% (slightly longer than the trend but no significant difference) using a similarity calculation algorithm.
[0138] In the case where the video evaluation unit (250) has a cut editing style of an advertisement video uploaded from the video analysis AI model unit (220) (e.g., long take of 5 seconds or more) and a cut editing style of a trend video learned from the SNS trend analysis unit (240) (e.g., fast cut of 2 to 3 seconds), the similarity score of the cut editing style (second similarity score) is calculated as 50% (slightly slower than the trend) using a similarity calculation algorithm.
[0139] In the case where the color tone and filter (e.g., strong contrast (neon)) of the advertisement video uploaded from the video analysis AI model unit (220) and the color tone and filter (e.g., warm beige and orange cut 2 to 3 seconds) of the trend video learned from the SNS trend analysis unit (240) are similarity scores (third similarity scores) of the color tone and filter (different from the current trend) are calculated as 40% (different from the current trend) using a similarity calculation algorithm.
[0140] In the case where the video evaluation unit (250) has a background music style of an advertisement video uploaded from the video analysis AI model unit (220) (e.g., slow piano performance, 80 BPM) and a background music style of a trend video learned from the SNS trend analysis unit (240) (e.g., exciting song, 120 BPM or more), the similarity score of the background music style (4th similarity score) is calculated as 60% (slow music but suitable for some trends) using a similarity calculation algorithm.
[0141] In the case of the video evaluation unit (250) having an SNS viral element (e.g., no challenge) of the uploaded advertisement video from the video analysis AI model unit (220) and an SNS viral element (e.g., including challenge) of the trend video learned from the SNS trend analysis unit (240), the similarity score of the SNS viral element (5th similarity score) is calculated as 30% (low likelihood of spreading on TikTok and Reels) using a similarity calculation algorithm.
[0142] The video evaluation unit (250) sets weights by considering the importance of each component and calculates a trend suitability score indicating whether the trend is matched using the following mathematical formula 1.
[0143] [Mathematical Formula 1]
[0144]
[0145] Here, the trend fit score is a quantitatively evaluated value indicating the degree of alignment with recent trends and preferences, the first similarity score is the similarity score of video length, W1 represents the weight corresponding to the first similarity score, the second similarity score is the similarity score of cut editing style, W2 represents the weight corresponding to the second similarity score, the third similarity score is the similarity score of color tone and filters, W3 represents the weight corresponding to the third similarity score, the fourth similarity score is the similarity score of background music style, W4 represents the weight corresponding to the fourth similarity score, the fifth similarity score is the similarity score of social media viral elements, and W5 represents the weight corresponding to the fifth similarity score.
[0146] For example, the video evaluation unit (250) assigns weights based on the importance of each component, setting the video length to 20%, cut editing style to 25%, color tone and filter to 15%, background music style to 20%, and social media viral elements to 20%. The weights may be changed as needed.
[0147] The video evaluation section (250) calculates the trend suitability score as (70 × 0.2) + (50 × 0.25) + (40 × 0.15) + (60 × 0.2) + (30 × 0.2) = 14 + 12.5 + 6 + 12 + 6 = 50.5%.
[0148] The video evaluation unit (250) can convert the similarity score for each component of video length, cut editing style, color tone and filter, background music style, and SNS viral elements into an evaluation score according to a predetermined standard (100 points maximum). In this case, a similarity score of 70% for video length may correspond to an evaluation score of 70 points. As another embodiment, the video evaluation unit (250) can calculate the evaluation score by multiplying the similarity score for each component by a predetermined weight, and can also set an evaluation score corresponding to a specific range of similarity scores by defining evaluation criteria for similarity scores in advance.
[0149] The video evaluation unit (250) can generate evaluation result information including the direction of modification for each component by comparing it with a predetermined modification standard based on the evaluation score.
[0150] An example of evaluation result information is shown in Table 1 below.
[0151] [Table 1]
[0152]
[0153] The video evaluation unit (250) generates evaluation result information including trend suitability scores and transmits it to the final evaluation feedback providing unit (260).
[0154] The existing ad video is 20 seconds long, features a long take style and strong contrast colors, and has a trend suitability score of 50.5%, whereas the revised evaluation results have a video length of 12 seconds, a cut editing style of fast cut editing (2 to 3 seconds), a warm orange tone color scheme, a fast tempo background music (120 BPM), and the addition of TikTok challenge elements.
[0155] The final evaluation feedback providing unit (260) can provide feedback of the evaluation result information to the input data processing unit (110) of the advertising storyboard automatic generation device (100).
[0156] The final evaluation feedback providing unit (260) can feed the evaluation result information to the generative AI processing unit (130) of the advertising storyboard automatic generation device (100).
[0157] The advertising storyboard automatic generation device (100) can receive feedback on evaluation result information, re-evaluate the advertising video to check if trend suitability has improved, and automatically generate the optimal advertising video.
[0158] The technical features disclosed in each embodiment of the present invention are not limited to that embodiment only, and as long as they are not mutually incompatible, the technical features disclosed in each embodiment may be combined and applied to different embodiments.
[0159] Therefore, in each embodiment, the technical features are described primarily, but as long as the technical features are not mutually incompatible, they may be combined and applied together.
[0160] The present invention is not limited to the embodiments described above and the attached drawings, and various modifications and variations may be possible from the perspective of those skilled in the art to which the present invention belongs. Accordingly, the scope of the present invention should be defined not only by the claims of this specification but also by equivalents thereof. Explanation of the symbols
[0161] 10: Ad Storyboard Evaluation Platform 100: Ad Storyboard Automatic Generation Device 110: Input data processing unit 120: Control unit 130: Generative AI Processing Unit 140: Storyboard Generation AI Model Unit 141: Story Analysis and Cut Division Section 142: Image Generation per Cut Section 143: Cut-by-cut editing graphic user interface section 150: Video Production AI Model Unit 151: Text Generation Unit 152: Image generation unit for video production 153: Video generation unit 154: Speech generation unit 155: Combination unit 160: Storage unit 170: Display unit 180: Communications Department 200: Advertising Draft Video Evaluation Device 210: Video Upload Unit 220: Video Analysis AI Model Unit 230: Trend Learning Database Department 240: SNS Trend Analysis Department 250: Video Evaluation Department 260: Final Evaluation Feedback Provision Department
Claims
Claim 1 An advertising storyboard evaluation platform comprising: an automatic advertising storyboard generation device that receives basic information data including video length and video production purpose, and storyboard composition data which is data for configuring synopsis-by-cut images, such as synopsis-by-cut reference images, generates dialogue and narration data for each cut of the advertising storyboard reflecting the synopsis-by-cut text, generates scene images for each cut reflecting the reference images for each cut, converts the dialogue and narration data for each cut into voice, and synthesizes them by synchronizing them with the scene images for each cut to generate a video file that is an advertising draft; and an advertising draft video evaluation device that analyzes the video file received from the automatic advertising storyboard generation device to extract a first core component as a comparison item, collects advertising videos from social media and advertising platforms to extract a second core component as a comparison item, quantitatively evaluates the similarity between the first core component and the second core component, generates an evaluation result based on whether the advertising video of the video file matches recent trend videos, and feeds back to the automatic advertising storyboard generation device. Claim 2 In claim 1, the advertising draft video evaluation device further comprises an advertising storyboard evaluation platform including a video analysis AI (Artificial Intelligence) model unit that analyzes the video file to extract first core components of the video file, such as video length, cut editing style, color tone and filter, background music style, and SNS viral elements. Claim 3 In claim 2, the advertising draft video evaluation device further comprises an advertising storyboard evaluation platform including an SNS trend analysis unit that analyzes advertising videos collected from social media and advertising platforms to extract second core components of the advertising video, such as video length, cut editing style, color tone and filter, background music style, and SNS viral elements. Claim 4 In claim 3, the advertising draft video evaluation device further comprises an advertising storyboard evaluation platform that calculates whether a trend aligns by comparing each component of the video analysis AI model unit and the SNS trend analysis unit one-to-one. Claim 5 In claim 3, the advertising draft video evaluation device further comprises a video evaluation unit that calculates a first similarity score of the video length using a similarity calculation algorithm by comparing the video length of the first core component and the video length of the second core component, calculates a second similarity score of the cut editing style using a similarity calculation algorithm by comparing the cut editing style of the first core component and the cut editing style of the second core component, calculates a third similarity score of the color tone and filter using a similarity calculation algorithm by comparing the color tone and filter of the first core component and the color tone and filter of the second core component, calculates a fourth similarity score of the background music style using a similarity calculation algorithm by comparing the background music style of the first core component and the background music style of the second core component, and calculates a fifth similarity score of the SNS viral element using a similarity calculation algorithm by comparing the SNS viral element of the first core component and the SNS viral element of the second core component. Claim 6 In claim 5, the video evaluation unit sets weights by considering the importance of the video length, the cut editing style, the color tone and filter, the background music style, and the SNS viral element, and calculates a trend suitability score indicating whether it conforms to a trend using the following mathematical formula 1, an advertising storyboard evaluation platform.[Mathematical Formula 1] Here, the trend fit score is a quantitatively evaluated value indicating the alignment between recent trends and preferences, the first similarity score is the video length similarity score, W1 represents the weight corresponding to the first similarity score, the second similarity score is the cut editing style similarity score, W2 represents the weight corresponding to the second similarity score, the third similarity score is the color tone and filter similarity score, W3 represents the weight corresponding to the third similarity score, the fourth similarity score is the background music style similarity score, W4 represents the weight corresponding to the fourth similarity score, the fifth similarity score is the SNS viral element similarity score, and W5 is the weight corresponding to the fifth similarity score. Claim 7 In claim 6, the video evaluation unit converts the similarity score for each component of the video length, the cut editing style, the color tone and filter, the background music style, and the SNS viral element into an evaluation score according to a predetermined standard, and generates evaluation result information including the modification direction and the trend suitability score for each component by comparing it with a predetermined modification standard based on the converted evaluation score, an advertising storyboard evaluation platform. Claim 8 In claim 7, the advertising draft video evaluation device further comprises a final evaluation feedback providing unit that feeds the evaluation result information to the advertising storyboard automatic generation device to modify the input value of the advertising storyboard automatic generation device, thereby forming an advertising storyboard evaluation platform.