Video content instant derivation method and device, electronic equipment and storage medium
Through computer vision and natural language processing technology, combined with brand elements, the picture scenes, objects and emotions are automatically recognized, and the plot templates are matched to generate personalized lines and arranged in a storyboard, which solves the problem of inefficient creation of traditional UGC tools, realizing the instant and dynamic generation of video content and the intelligent implantation of brand information.
Patent Information
- Application Number
- CN202510665299.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-02
AI Technical Summary
Traditional UGC tools lack storyline generation capabilities and cannot automatically generate plot logic that meets the needs of commercial communication. The functions of image recognition, script generation and storyboard arrangement are separated, resulting in inexperienced creation.
Through computer vision and natural language processing technology, image scenes, objects and face emotions are identified, plot templates are matched, personalized lines are generated and automatically stratified, and video content is generated by combining brand elements.
It realizes the instant and dynamic generation of video content, improves creative efficiency and user experience, and supports the intelligent implantation of brand information and scene adaptation.
Smart Images

Figure CN120583282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital content generation, and in particular to a method, device, electronic device and storage medium for instant derivation of video content. Background Art
[0002] Traditional user-generated content (UGC) solutions have the following technical flaws:
[0003] 1) Static Template Limitations: Current mainstream tools (such as Canva and Jianying) only support filling in fixed templates with images and text, and lack the ability to generate storylines. Users must manually arrange storyboards, which takes an average of over two hours to create.
[0004] 2) Single content dimension: Traditional tools cannot automatically generate plot logic that meets the needs of commercial communication based on the content of the materials. Corporate brand elements (LOGO, contact number) need to be added manually later, resulting in low communication efficiency.
[0005] 3) Technical fragmentation: In existing solutions, functional modules such as image recognition, script generation, and storyboard editing are independent of each other and lack end-to-end automated processes.
[0006] The above problems need to be solved urgently. Summary of the Invention
[0007] The purpose of the present invention is to solve one of the technical problems existing in the prior art to at least a certain extent.
[0008] To this end, an object of an embodiment of the present invention is to provide a method for instant derivation of video content, which realizes instant and dynamic generation of video content, improves the efficiency of video content generation and user experience.
[0009] Another object of an embodiment of the present invention is to provide a device for real-time derivation of video content.
[0010] In order to achieve the above technical objectives, the technical solutions adopted by the embodiments of the present invention include:
[0011] In one aspect, an embodiment of the present invention provides a method for instant derivation of video content, comprising the following steps:
[0012] Obtain multiple target images and target brand elements uploaded by the user, and identify target scenes, target objects, and target facial emotions based on the target images;
[0013] According to the target scene, the target object and the target facial emotion, a target plot template is obtained by matching in a preset plot template library;
[0014] Performing semantic analysis on the target image, generating personalized lines based on the semantic analysis results, and then generating a target script based on the target plot template and the personalized lines;
[0015] Determining the storyboard type of each target picture, and assigning a plot structure to the target picture according to the target script and the storyboard type to obtain a storyboard segment sequence;
[0016] The target brand element is integrated into the end of the storyboard sequence to obtain target video content.
[0017] Furthermore, in one embodiment of the present invention, the identifying of the target scene, target object, and target facial emotion according to the target image is specifically as follows:
[0018] The target image is input into a preset scene recognition model, an object detection model, and a facial emotion recognition model respectively to obtain the target scene, the target object, and the target facial emotion.
[0019] Furthermore, in one embodiment of the present invention, the target scenario template is obtained by matching the target scene, the target object, and the target facial emotion in a preset scenario template library, which specifically includes:
[0020] Acquire multiple plot template samples and corresponding scene labels, object labels, and facial emotion labels from the plot template library;
[0021] Determining a scene matching degree between the target image and the plot template sample according to the target scene and the scene label;
[0022] Determining the object relevance between the target image and the plot template sample according to the target object and the object label;
[0023] Determining the emotional consistency between the target image and the plot template sample according to the target facial emotion and the facial emotion label;
[0024] Performing a weighted summation of the scene matching degree, the object relevance, and the emotion consistency according to a preset weight coefficient to obtain a similarity between the target image and the plot template sample;
[0025] The plot template sample with the highest similarity is selected as the target plot template.
[0026] Furthermore, in one embodiment of the present invention, the semantic analysis of the target image, generating personalized lines according to the semantic analysis results, and then generating a target script according to the target plot template and the personalized lines specifically includes:
[0027] Performing semantic analysis on the target image according to the target scene, the target object, and the target facial emotion to obtain a semantic analysis result;
[0028] Inputting the semantic analysis result and the preset prompt template into a large language model to obtain the personalized lines;
[0029] Fill the target plot template with the personalized lines to obtain the target script.
[0030] Furthermore, in one embodiment of the present invention, determining the storyboard type of each target image specifically includes:
[0031] Determine the proportion of faces, the proportion of interactions between people, and the proportion of environmental elements in each of the target images;
[0032] When the face ratio is greater than a preset first threshold, determining that the shot type of the target image is a close-up;
[0033] When the face ratio is less than or equal to the first threshold, and the character interaction ratio is greater than or equal to a preset second threshold and less than a preset third threshold, determining that the shot type of the target image is a medium shot;
[0034] When the proportion of the environmental elements is greater than a preset fourth threshold, it is determined that the frame type of the target image is panoramic.
[0035] Furthermore, in one embodiment of the present invention, allocating a plot structure to the target picture according to the target script and the storyboard type to obtain a storyboard segment sequence specifically includes:
[0036] Determine the introduction, development, climax, and ending of the target script;
[0037] Dividing the target image into a first panoramic shot, a second panoramic shot, a first medium shot, a second medium shot, a first close-up shot, and a second close-up shot according to the shot type;
[0038] Arranging the opening portion by storyboarding according to the first panoramic shot and the first medium shot, and determining the corresponding storyboard durations to obtain an opening storyboard segment;
[0039] Arranging the development portion of the shot according to the second medium shot and the first close-up shot, and determining the corresponding shot duration to obtain a development shot segment;
[0040] Arranging the climax part by storyboarding according to the second close-up shot, and determining the corresponding storyboard duration to obtain a climax storyboard segment;
[0041] Arranging the ending part by storyboarding according to the second panoramic shot, and determining the corresponding storyboard duration to obtain an ending storyboard segment;
[0042] The storyboard segment sequence is generated according to the opening storyboard segment, the development storyboard segment, the climax storyboard segment and the ending storyboard segment.
[0043] Furthermore, in one embodiment of the present invention, the step of fusing the target brand element at the end of the storyboard sequence to obtain the target video content specifically includes:
[0044] Obtaining blank areas of multiple target images corresponding to the end portion of the storyboard sequence by an edge detection algorithm;
[0045] The target brand element is resized according to the blank area, and the target brand element is rendered into a plurality of target images corresponding to the end portion of the storyboard sequence to generate the target video content.
[0046] On the other hand, an embodiment of the present invention provides a device for instantly deriving video content, comprising:
[0047] An image recognition module is used to obtain multiple target images and target brand elements uploaded by users, and identify target scenes, target objects, and target facial emotions based on the target images;
[0048] A plot template matching module is used to match the target plot template in a preset plot template library according to the target scene, the target object and the target facial emotion to obtain a target plot template;
[0049] A script generation module is used to perform semantic analysis on the target image, generate personalized lines according to the semantic analysis results, and then generate a target script according to the target plot template and the personalized lines;
[0050] A storyboard arrangement module is used to determine the storyboard type of each target picture, and assign a plot structure to the target picture according to the target script and the storyboard type to obtain a storyboard segment sequence;
[0051] The brand element fusion module is used to fuse the target brand element at the end of the storyboard sequence to obtain the target video content.
[0052] On the other hand, an embodiment of the present invention provides an electronic device, which includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the method for instant derivation of video content as described above is realized.
[0053] On the other hand, an embodiment of the present invention further provides a storage medium, which is a computer-readable storage medium for computer-readable storage, and stores one or more programs, which can be executed by one or more processors to implement the instant video content derivation method as described above.
[0054] The advantages and benefits of the present invention will be described in part in the following description and will become apparent from the following description or learned through practice of the present invention:
[0055] The embodiment of the present invention obtains multiple target images and target brand elements uploaded by users, identifies target scenes, target objects, and target facial emotions based on the target images, matches the target scenes, target objects, and target facial emotions in a preset plot template library to obtain a target plot template, performs semantic analysis on the target images, generates personalized lines based on the semantic analysis results, and then generates a target script based on the target plot template and personalized lines. The frame type of each target image is determined, and the target image is assigned a plot structure based on the target script and the frame type to obtain a frame segment sequence. The target brand element is integrated into the end of the frame segment sequence to obtain the target video content. The embodiment of the present invention performs scene recognition, object detection, and facial emotion recognition on the target images uploaded by users to match the corresponding target script template. The target script that meets the commercial communication goals is generated based on the semantic analysis results of the target images. The frame type is automatically split based on the composition characteristics of the target images and the plot structure is assigned. The brand element integration realizes the intelligent implantation of brand information and scene adaptation, realizes the instant and dynamic generation of video content, and improves the efficiency of video content generation and the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following introduction is made to the drawings required for use in the embodiments of the present invention. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative work.
[0057] Figure 1 A flowchart of a method for instant derivation of video content provided by an embodiment of the present invention;
[0058] Figure 2 A flowchart of step S101 provided in an embodiment of the present invention;
[0059] Figure 3 A flowchart of step S102 provided in an embodiment of the present invention;
[0060] Figure 4 A step flow chart of step S103 provided in an embodiment of the present invention
[0061] Figure 5 A flowchart of step S104 provided in an embodiment of the present invention;
[0062] Figure 6 Another step flow chart of step S104 provided in an embodiment of the present invention;
[0063] Figure 7 A flowchart of step S105 provided in an embodiment of the present invention;
[0064] Figure 8 A schematic diagram of the structure of a device for instant derivation of video content provided by an embodiment of the present invention;
[0065] Figure 9 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention;
[0066] Figure 10 A schematic diagram of the structure of a storage medium provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0067] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limitations on the present application. It should be noted that, although the functional modules are divided in the system schematic and the logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than the module division in the system schematic or the order in the flow chart. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and no limitation is placed on the order between the steps. The execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0068] In the description of the present invention, the meaning of "a plurality" is two or more. If there is a description of "first" or "second", it is only used to distinguish technical features and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used in this document have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used in this document are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0069] The instant derivation method for video content provided in the embodiments of the present application can be applied to a terminal, can be applied to a server, and can also be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, tablet computer, laptop computer, desktop computer, set-top box, etc.; the server can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the instant derivation method for video content, etc., but is not limited to the above forms.
[0070] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0071] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with the relevant laws, regulations, and standards of the relevant countries and regions. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0072] This invention is suitable for the automated production of user-generated content (UGC) in corporate brand communication scenarios. By integrating computer vision (CV), natural language processing (NLP), dynamic storyboard technology, and brand element intelligent implantation technology, it aims to provide users with an efficient, low-threshold video creation solution, realizing the full-link automated production of "picture → script → storyboard → video", and is particularly suitable for the content production of short videos and video ringtones in corporate brand communication scenarios.
[0073] like Figure 1 FIG2 is a flowchart of a method for instant derivation of video content provided by an embodiment of the present invention, referring to FIG2. Figure 1 The embodiment of the present invention provides a method for instant derivation of video content, which specifically includes the following steps:
[0074] S101, obtaining multiple target images and target brand elements uploaded by a user, and identifying target scenes, target objects, and target facial emotions based on the target images;
[0075] S102, obtaining a target plot template by matching the target scene, target object, and target facial emotion in a preset plot template library;
[0076] S103, performing semantic analysis on the target image, generating personalized lines based on the semantic analysis results, and then generating a target script based on the target plot template and the personalized lines;
[0077] S104, determining the storyboard type of each target picture, and assigning a plot structure to the target picture according to the target script and the storyboard type to obtain a storyboard segment sequence;
[0078] S105. Integrate the target brand elements into the end of the storyboard sequence to obtain target video content.
[0079] The present invention can be used for corporate brand communication and personal content creation. Merchants can quickly generate plot video ringback rings with brand elements by uploading pictures. Ordinary users can also generate high-quality plot video ringback rings through simple operations, reducing production costs.
[0080] The present invention is implemented based on the following technologies:
[0081] Computer Vision (CV): used for image semantic analysis, including scene recognition, object detection, and facial emotion analysis.
[0082] Natural Language Processing (NLP): Used for script generation, matching preset plot templates based on image semantics, and dynamically filling in personalized lines.
[0083] Video storyboard technology: used for intelligent storyboard arrangement, automatically splitting storyboard types (such as close-up, medium shot, and panorama) according to the composition characteristics of the picture, and dynamically adapting to the plot logic.
[0084] Brand communication technology: used for intelligent implantation of corporate logos and contact numbers, supporting dynamic position adjustment and interactive function binding.
[0085] The embodiment of the present invention performs scene recognition, object detection, and facial emotion recognition on the target image uploaded by the user, thereby matching the corresponding target script template, and generates a target script that meets the commercial communication goals in combination with the semantic analysis results of the target image. It automatically splits the storyboard types and allocates the plot structure based on the composition features of the target image, and realizes the intelligent implantation and scene adaptation of brand information through the fusion of brand elements, thereby realizing the instant and dynamic generation of video content, improving the efficiency of video content generation and the user experience.
[0086] like Figure 2 FIG. 1 is a flowchart of step S101 provided in an embodiment of the present invention, referring to FIG. Figure 2 As an optional implementation, the target scene, target object, and target facial emotion are identified based on the target image, specifically:
[0087] S1011. Input the target image into a preset scene recognition model, an object detection model, and a facial emotion recognition model respectively to obtain a target scene, a target object, and a target facial emotion.
[0088] Specifically, the multimodal script generation engine in this embodiment of the present invention integrates computer vision (CV) and natural language processing (NLP) technologies to automate the process from image semantic parsing to plot generation. It uses the YOLOv5 model and OpenFace algorithm to identify scenes (e.g., indoors, outdoors), objects (e.g., coffee cups, cakes), and facial emotions (e.g., joy, surprise) in images.
[0089] like Figure 3 FIG. 1 is a flowchart of step S102 provided in an embodiment of the present invention, referring to FIG. Figure 3 As an optional implementation, a target plot template is obtained by matching the target scene, target object, and target facial emotion in a preset plot template library, which specifically includes:
[0090] S1021, obtaining multiple plot template samples and corresponding scene labels, object labels, and facial emotion labels from a plot template library;
[0091] S1022, determining the scene matching degree between the target image and the plot template sample according to the target scene and the scene label;
[0092] S1023, determining the object relevance between the target image and the plot template sample according to the target object and the object label;
[0093] S1024, determining the emotional consistency between the target image and the plot template sample based on the target facial emotion and the facial emotion label;
[0094] S1025, performing weighted summation of the scene matching degree, object relevance, and emotion consistency according to a preset weight coefficient to obtain the similarity between the target image and the plot template sample;
[0095] S1026: Select the plot template sample with the highest similarity as the target plot template.
[0096] Specifically, the embodiment of the present invention matches the preset plot template library (such as "counterattack", "warmth", and "suspense") based on the identified target scene, target object, and target facial emotion, and calculates the similarity between each plot template sample and the target image through the similarity calculation formula (S = 0.4×scene matching + 0.3×object relevance + 0.3×emotion consistency), and generates the target script based on the template with the highest similarity.
[0097] like Figure 4 FIG. 1 is a flowchart of step S103 provided in an embodiment of the present invention, referring to FIG. Figure 4 As an optional implementation, semantic analysis is performed on the target image, personalized lines are generated based on the semantic analysis results, and then a target script is generated based on the target plot template and the personalized lines, which specifically includes:
[0098] S1031, performing semantic analysis on the target image according to the target scene, target object, and target facial emotion to obtain a semantic analysis result;
[0099] S1032: Input the semantic analysis results and the preset prompt template into the large language model to obtain personalized lines;
[0100] S1033. Fill the target plot template with personalized lines to obtain the target script.
[0101] Specifically, the NLP template engine dynamically populates personalized lines (e.g., "Every steak is meticulously cooked through five steps") based on image semantic analysis results, ultimately generating a script that meets the needs of commercial communication. The engine supports 12 types of plot templates, covering scenarios such as corporate growth histories, customer service cases, and new product launches, ensuring high-quality and diverse content generation.
[0102] like Figure 5 FIG. 1 is a flowchart of step S104 provided in an embodiment of the present invention, referring to FIG. Figure 5As an optional implementation, determining the storyboard type of each target image may include:
[0103] S1040: Determine the proportion of faces, the proportion of interactions between people, and the proportion of environmental elements in each target image;
[0104] S1041: When the face ratio is greater than a preset first threshold, determine that the shot type of the target image is a close-up;
[0105] S1042: When the face ratio is less than or equal to a first threshold, and the person interaction ratio is greater than or equal to a preset second threshold and less than a preset third threshold, determine that the shot type of the target image is medium shot;
[0106] S1043: When the proportion of environmental factors is greater than a preset fourth threshold, determine that the frame type of the target image is panoramic.
[0107] Specifically, the image feature vector is extracted through ResNet50, and the clustering algorithm is combined to determine the shot type (close-up shot: the face accounts for more than 40% and the background is blurred; medium shot: the interactive subject of the characters occupies 30-70% of the screen height; panoramic shot: the environmental elements account for more than 60%).
[0108] like Figure 6 FIG. 1 is another flow chart of step S104 provided in an embodiment of the present invention, referring to FIG. Figure 6 As an optional implementation, the target picture is assigned a plot structure according to the target script and the storyboard type to obtain a storyboard segment sequence, which specifically includes:
[0109] S1044: Determine the opening, development, climax, and ending of the target script according to the target script; and divide the target image into a first panoramic shot, a second panoramic shot, a first medium shot, a second medium shot, a first close-up shot, and a second close-up shot according to the shot type.
[0110] S1045. Arrange the opening portion of the storyboard according to the first panoramic shot and the first medium shot, and determine the corresponding storyboard duration to obtain the opening storyboard segment;
[0111] S1046. Arrange the development portion of the shot according to the second medium shot and the first close-up shot, and determine the corresponding shot duration to obtain a development shot segment.
[0112] S1047, arranging the climax part according to the second close-up shot, and determining the corresponding shot duration to obtain a climax shot segment;
[0113] S1048, arranging the ending part of the storyboard according to the second panoramic shot, and determining the corresponding storyboard duration to obtain the ending storyboard segment;
[0114] S1049: Generate a storyboard segment sequence based on the opening storyboard segment, the development storyboard segment, the climax storyboard segment, and the ending storyboard segment.
[0115] Specifically, the length of storyboards is allocated according to the plot structure of "introduction, development, turn and conclusion" (beginning: 20%, panoramic → medium shot; development: 30%, alternating medium shot + close-up; climax: 40%, fast close-up switching; ending: 10%, panoramic freeze + LOGO display) to ensure that the video rhythm is smooth and in line with the logic of the plot; through dynamic timeline mapping technology, the storyboard type is closely combined with the plot stage. For example, in the climax part, a fast close-up switching of 1 frame per second is used to enhance the visual impact, and in the ending part, the company logo is displayed through a panoramic freeze to enhance brand exposure.
[0116] like Figure 7 FIG. 1 is a flowchart of step S105 provided in an embodiment of the present invention, referring to FIG. Figure 7 As an optional implementation, the target brand element is integrated into the end of the storyboard sequence to obtain the target video content, which specifically includes:
[0117] S1051, obtaining blank areas of multiple target images corresponding to the end of the storyboard sequence using an edge detection algorithm;
[0118] S1052: Adjust the size of the target brand element according to the blank area, and render the target brand element into multiple target images corresponding to the end portion of the storyboard sequence to generate target video content.
[0119] Specifically, an edge detection algorithm automatically selects blank areas in the image to dynamically adjust the company logo's position. The logo begins displaying in the second half of the climax and continues until the end of the video, ensuring a natural transition between the logo and the video content. Furthermore, the contact number is presented verbatim using a typewriter effect, with click-to-dial (on HTML5 pages) and copy-number functionality supported, enhancing the user interaction experience. Furthermore, the system supports multi-platform adaptive rendering technology, automatically adjusting the logo's size and position to ensure consistent display across devices with varying resolutions.
[0120] The overall process of the present invention is described below with reference to specific embodiments.
[0121] 1. Catering enterprise promotional video generation
[0122] 1) Input data:
[0123] Picture: Food preparation in the kitchen (scene: indoors), close-up of dish (object: steak), customer’s smiling face (emotion: joy).
[0124] Brand information: LOGO (position preference: lower right corner), contact number (400-123-4567).
[0125] 2) Processing flow:
[0126] Image semantic analysis: Identifying “dining scene-food-pleasant emotions”.
[0127] Script generation: Match the "warmth" template and generate the lines: "Every steak is carefully cooked through 5 steps."
[0128] 3) Storyboard arrangement:
[0129] Start: Panoramic view of the kitchen (2 seconds);
[0130] Development: Alternating between medium shot of the chef and close-up of the steak (4 seconds);
[0131] Climax: Close-up of the customer’s smiling face (3 seconds) + gradual appearance of the logo;
[0132] Ending: The contact number appears word by word (1 second).
[0133] 4) Output:
[0134] Resolution: 1080×1920;
[0135] Duration: 10 seconds.
[0136] File format: MP4 (H.264 encoding).
[0137] 2. Video generation for new product launches by technology companies
[0138] 1) Input data:
[0139] Images: R&D team at work (scene: office), product prototype (object: smart device), press conference (emotion: excitement).
[0140] Brand information: LOGO (position preference: upper left corner), contact number (400-XXXX-XXXX).
[0141] 2) Processing flow:
[0142] Image semantic parsing: Recognizing “office scene-tech product-excited emotion”.
[0143] Script generation: Match the "Exciting" template and generate the lines: "This device will completely change your lifestyle."
[0144] 3) Storyboard arrangement:
[0145] Beginning: Panorama of the office (2 seconds);
[0146] Development: Alternating between a medium shot of the R&D team and a close-up of the product (4 seconds);
[0147] Climax: Quick close-up switching of the press conference site (3 seconds) + gradual appearance of the logo;
[0148] Ending: The contact number appears word by word (1 second).
[0149] 4) Output:
[0150] Resolution: 1080×1920;
[0151] Duration: 10 seconds
[0152] File format: MP4 (H.264 encoding).
[0153] The above describes the method flow of the embodiment of the present invention. It is understandable that the embodiment of the present invention performs scene recognition, object detection, and facial emotion recognition on the target image uploaded by the user, thereby matching the corresponding target script template, and generating a target script that meets the commercial communication goals based on the semantic analysis results of the target image. Based on the composition characteristics of the target image, the storyboard type is automatically split and the plot structure is allocated. Through the integration of brand elements, the intelligent implantation of brand information and scene adaptation are realized, and the instant and dynamic generation of video content is realized, which improves the efficiency of video content generation and the user experience.
[0154] Compared with the prior art, the embodiments of the present invention have the following advantages:
[0155] 1) Lowering the threshold for creation: Integrating computer vision (CV) and natural language processing (NLP) technologies, the YOLOv5 model is used to identify scenes, objects, and facial emotions in images, match them with a library of preset plot templates, dynamically fill in personalized lines, and generate scripts that meet commercial communication needs. Users do not need professional skills, just upload a photo to generate plot-based ringtones, which stimulates users' creative enthusiasm and enables ordinary users to create with zero basic knowledge.
[0156] 2) Improve content quality: Based on a multimodal script generation engine, the system dynamically generates plots and dialogues that meet commercial communication goals. ResNet50 is used to extract image feature vectors, and clustering algorithms are used to determine the shot type. Shot durations are allocated according to the "introduction, development, turn, and conclusion" plot structure, ensuring a smooth video rhythm that conforms to the plot logic, significantly outperforming manual editing.
[0157] 3) Enhanced Commercial Value: Supports intelligent placement of company logos and contact numbers. An edge detection algorithm automatically selects blank areas in the image to dynamically adjust the company logo's position. The logo begins to display in the second half of the climax and continues until the end of the video, enhancing brand exposure and user interaction.
[0158] 4) Promote technology accessibility: Bring AI video generation technology to ordinary users, support multiple scenarios such as personal creation and corporate communication, stimulate the vitality of the UGC ecosystem, and help small and medium-sized enterprises achieve brand communication at low cost.
[0159] like Figure 8 FIG2 is a schematic diagram showing the structure of a video content instant derivation device provided by an embodiment of the present invention, referring to FIG2. Figure 8 The embodiment of the present invention provides a device for instantly deriving video content, including:
[0160] The image recognition module is used to obtain multiple target images and target brand elements uploaded by users, and identify target scenes, target objects, and target facial emotions based on the target images;
[0161] The plot template matching module is used to match the target plot template in the preset plot template library according to the target scene, target object and target facial emotion;
[0162] The script generation module is used to perform semantic analysis on the target image, generate personalized lines based on the semantic analysis results, and then generate the target script based on the target plot template and personalized lines;
[0163] The storyboard arrangement module is used to determine the storyboard type of each target picture and assign the plot structure to the target picture according to the target script and storyboard type to obtain a storyboard segment sequence;
[0164] The brand element fusion module is used to fuse the target brand element at the end of the storyboard sequence to obtain the target video content.
[0165] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0166] An embodiment of the present invention further provides an electronic device comprising: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for enabling communication between the processor and the memory. When the program is executed by the processor, the aforementioned method for instant video content derivation is implemented. The electronic device can be any intelligent terminal, including a tablet computer and an in-vehicle computer.
[0167] like Figure 9 FIG2 is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention, referring to FIG2 Figure 9 , an embodiment of the present invention provides an electronic device, including:
[0168] The processor 901 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention.
[0169] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called by the processor 901 to execute the instant video content derivation method of the embodiments of the present invention.
[0170] Input / output interface 903, used to implement information input and output;
[0171] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0172] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );
[0173] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0174] like Figure 10 FIG2 is a schematic diagram of the structure of the storage medium provided by the embodiment of the present invention, referring to FIG2 Figure 10 An embodiment of the present invention further provides a storage medium, which is a computer-readable storage medium used for computer-readable storage. The storage medium stores one or more programs 1001, and the one or more programs 1001 can be executed by one or more processors to implement the above-mentioned video content instant derivation method.
[0175] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0176] The embodiment of the present invention also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs Figure 1 The method shown.
[0177] In some optional embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the above-mentioned boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided in an exemplary manner for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operation and logic flow presented herein. Optional embodiments are contemplated in which the order of the various operations is changed and the sub-operations described as a part of a larger operation are performed independently.
[0178] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the above-mentioned functions and / or features can be integrated into a single physical device and / or software module, or one or more functions and / or features can be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More specifically, given the properties, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the routine skills of an engineer. Therefore, a person skilled in the art can implement the present invention set forth in the claims using ordinary skills without undue experimentation. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0179] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the above methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0180] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0181] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable media on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0182] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0183] In the above description of this specification, reference to the terms "one embodiment / example," "another embodiment / example," or "certain embodiments / examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0184] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
[0185] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A method for instant derivation of video content, characterized in that: The following steps are involved: Obtain multiple target images and target brand elements uploaded by the user, and identify target scenes, target objects, and target facial emotions based on the target images; According to the target scene, the target object and the target facial emotion, a target plot template is obtained by matching in a preset plot template library; Performing semantic analysis on the target image, generating personalized lines based on the semantic analysis results, and then generating a target script based on the target plot template and the personalized lines; Determining the storyboard type of each target picture, and assigning a plot structure to the target picture according to the target script and the storyboard type to obtain a storyboard segment sequence; The target brand element is integrated into the end of the storyboard sequence to obtain target video content.
2. A method for instant derivation of video content according to claim 1, characterized in that: The target scene, target object and target facial emotion are identified according to the target image, which is specifically: The target image is input into a preset scene recognition model, an object detection model, and a facial emotion recognition model respectively to obtain the target scene, the target object, and the target facial emotion.
3. A method for instant derivation of video content according to claim 1, characterized in that: The target scenario template is obtained by matching the target scene, the target object, and the target facial emotion in a preset scenario template library, which specifically includes: Acquire multiple plot template samples and corresponding scene labels, object labels, and facial emotion labels from the plot template library; Determining a scene matching degree between the target image and the plot template sample according to the target scene and the scene label; Determining the object relevance between the target image and the plot template sample according to the target object and the object label; Determining the emotional consistency between the target image and the plot template sample according to the target facial emotion and the facial emotion label; Performing a weighted summation of the scene matching degree, the object relevance, and the emotion consistency according to a preset weight coefficient to obtain a similarity between the target image and the plot template sample; The plot template sample with the highest similarity is selected as the target plot template.
4. A method for instant derivation of video content according to claim 1, characterized in that: The step of performing semantic analysis on the target image, generating personalized lines according to the semantic analysis result, and then generating a target script according to the target plot template and the personalized lines specifically includes: Performing semantic analysis on the target image according to the target scene, the target object, and the target facial emotion to obtain a semantic analysis result; Inputting the semantic analysis result and the preset prompt template into a large language model to obtain the personalized lines; Fill the target plot template with the personalized lines to obtain the target script.
5. The method for instant derivation of video content according to claim 1, characterized in that: The determining of the storyboard type of each target image specifically includes: Determine the proportion of faces, the proportion of interactions between people, and the proportion of environmental elements in each of the target images; When the face ratio is greater than a preset first threshold, determining that the shot type of the target image is a close-up; When the face ratio is less than or equal to the first threshold, and the character interaction ratio is greater than or equal to a preset second threshold and less than a preset third threshold, determining that the shot type of the target image is a medium shot; When the proportion of the environmental elements is greater than a preset fourth threshold, it is determined that the frame type of the target image is panoramic.
6. A method for instant derivation of video content according to claim 5, characterized in that: The step of allocating a plot structure to the target picture according to the target script and the storyboard type to obtain a storyboard segment sequence specifically includes: Determine the introduction, development, climax, and ending of the target script; Dividing the target image into a first panoramic shot, a second panoramic shot, a first medium shot, a second medium shot, a first close-up shot, and a second close-up shot according to the shot type; Arranging the opening portion by storyboarding according to the first panoramic shot and the first medium shot, and determining the corresponding storyboard durations to obtain an opening storyboard segment; Arranging the development portion of the shot according to the second medium shot and the first close-up shot, and determining the corresponding shot duration to obtain a development shot segment; Arranging the climax part by storyboarding according to the second close-up shot, and determining the corresponding storyboard duration to obtain a climax storyboard segment; Arranging the ending part by storyboarding according to the second panoramic shot, and determining the corresponding storyboard duration to obtain an ending storyboard segment; The storyboard segment sequence is generated according to the opening storyboard segment, the development storyboard segment, the climax storyboard segment and the ending storyboard segment.
7. A method for instant derivation of video content according to any one of claims 1 to 6, characterized in that: The step of fusing the target brand element at the end of the storyboard sequence to obtain target video content specifically includes: Obtaining blank areas of multiple target images corresponding to the end portion of the storyboard sequence by an edge detection algorithm; The target brand element is resized according to the blank area, and the target brand element is rendered into a plurality of target images corresponding to the end portion of the storyboard sequence to generate the target video content.
8. A device for instant derivation of video content, characterized in that: include: An image recognition module is used to obtain multiple target images and target brand elements uploaded by users, and identify target scenes, target objects, and target facial emotions based on the target images; A plot template matching module is used to match the target plot template in a preset plot template library according to the target scene, the target object and the target facial emotion to obtain a target plot template; A script generation module is used to perform semantic analysis on the target image, generate personalized lines according to the semantic analysis results, and then generate a target script according to the target plot template and the personalized lines; A storyboard arrangement module is used to determine the storyboard type of each target picture, and assign a plot structure to the target picture according to the target script and the storyboard type to obtain a storyboard segment sequence; The brand element fusion module is used to fuse the target brand element at the end of the storyboard sequence to obtain the target video content.
9. An electronic device, characterized in that: The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the steps of the method for instant derivation of video content as described in any one of claims 1 to 7 are realized.
10. A storage medium, which is a computer-readable storage medium and is used for computer-readable storage, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the video content instant derivation method according to any one of claims 1 to 7.