Short drama making method, system, equipment and medium
Through the full-process management method driven by artificial intelligence, multiple AI models are used to realize plot generation, storyboard design and other functions, solving the problem that traditional short drama production relies on man-intensive workflows, and achieving efficient, innovative and low-cost short drama production results.
Patent Information
- Application Number
- CN202510197600.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional short drama production relies on man-intensive workflows and lacks intelligent support, resulting in limited content diversity and innovation, high costs and difficult to control progress. Existing automation tools can only solve problems at specific stages and cannot complete core tasks independently, and still require a lot of manual intervention.
Adopt artificial intelligence-driven full-process management methods, and use large language models, diffusion models, time series models and multimodal large models to realize plot generation, storyboard design, highlight moment recognition and content summary, poster processing and other functions to form an integrated short drama production system.
The short drama production cycle has been greatly shortened, the overall efficiency has been improved, the production cost has been reduced, the content has been enhanced, the innovation and diversity has been enhanced, the system efficiency and robustness have also been improved.
Smart Images

Figure CN120050449A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence technology and video processing technology, and specifically relates to a short drama production method, storage medium, device, and system. Background Art
[0002] Traditional short drama production relies on a labor-intensive workflow. From creative concept to final release, it involves multiple links, and each link requires the participation of professionals. It is easily affected by human factors, resulting in difficult control of the project schedule. There is a lack of intelligent support in aspects such as creative generation and scriptwriting, mainly relying on the experience and inspiration of creators, which limits the diversity and innovation of content. It also involves a large amount of upfront investment, including equipment procurement, venue leasing, actor salaries, etc., which is a heavy burden for small production teams. And it requires complex coordination work. Especially when communicating across departments, poor information transmission may lead to misunderstandings or delays.
[0003] Currently, although there are some automated tools that can help with video editing, special effect adding, etc., these tools usually only focus on a specific stage and do not provide an integrated solution to cover the entire process from creativity to release. They cannot independently complete core tasks such as high-quality scriptwriting, character design, and scene layout, and still require a large amount of manual intervention. Although certain automated tools can reduce the cost of some operations, due to their limited application scope, they have not significantly reduced the total cost of short drama production as a whole. Users need to master the operation skills of multiple different software, increasing the learning curve, and compatibility issues between various tools may also affect work efficiency. Summary of the Invention
[0004] The main purpose of this application is to provide a short drama production method, storage medium, device, and system to overcome or at least partially solve the above problems in the background art. To this end, this application provides the following technical solutions:
[0005] In the first aspect, this application provides a short drama production method, including:
[0006] Plot generation: Generate a plot script that fits the input based on the keywords or outlines input by the creator;
[0007] Shot design: Convert the plot script into shot diagrams of the short drama;
[0008] Highlight moment identification and content summary: Segment the content of the short drama to identify highlight moments; extract video and audio content to generate a content summary that fits the plot;
[0009] Poster processing: Extract key elements based on the highlight moments and generate a poster using the key elements of the highlight moments.
[0010] In at least one embodiment, the process of generating the plot script includes;
[0011] Using a large language model to analyze the input keywords or outline, generating a complete plot script that fits the user input; during the plot generation process, by adjusting the model parameters and introducing style tags, the large language model adjusts the language expression according to the style required by the user.
[0012] In at least one embodiment, the plot script is converted into storyboard images through a diffusion model, and the process includes:
[0013] Information encoding stage: Deeply encode the information in the plot script, where the information includes scenes, characters, and actions, and convert it into encoded information that can be understood and processed by the diffusion model;
[0014] Latent space image representation generation stage: The diffusion model uses a neural network structure to perform non-linear transformation and feature extraction on the encoded information, gradually generating latent feature vectors with specific semantics and structures; the latent feature vectors are mapped to the image space through a decoder network, thereby generating a preliminary image representation. In this process, a diffusion algorithm is used to continuously optimize the details and authenticity of the image representation;
[0015] Optimize shot transitions and frame layouts: The diffusion model introduces an attention mechanism and spatio-temporal constraints; through the attention mechanism, calculate the correlation weights between different information elements, focus on key information, and thus determine the key content of each shot; through spatio-temporal constraints, arrange the shot timing and frame space layout according to the plot rhythm and time sequence.
[0016] In at least one embodiment, the process of identifying the highlight moments is as follows: Use a time series model to process the time series data of the short drama content, and combine a large video understanding model to perform semantic analysis on the video content, so as to achieve accurate segmentation of the short drama content and identification of the highlight moments;
[0017] The process of generating the content summary is as follows: Use a multi-modal large model for processing, respectively analyze the visual information in the video frame and the auditory information in the audio, and then integrate and analyze the two outputs of the model through a set workflow, so as to accurately identify the short drama content and summarize the key points.
[0018] In at least one embodiment, the process of generating the poster includes:
[0019] Use a convolutional neural network to extract key elements from the highlight moments, and the key elements include the main character image and the iconic scene;
[0020] Input the key elements into the diffusion model to generate a poster design template;
[0021] The multi - resolution image generation and cropping algorithm is adopted to process the poster design template to generate posters suitable for different screens.
[0022] Furthermore, the processing process of the multi - resolution image generation and cropping algorithm includes:
[0023] Image generation: The multi - resolution analysis technology is adopted to generate image versions with different resolution levels according to the target size of the poster and the display characteristics of different terminal devices; through the feature fusion and detail enhancement of images with different resolutions, it is ensured that the poster has rich visual details while retaining key information.
[0024] Image cropping: The attention mechanism is introduced in the cropping process. By analyzing the importance and visual attractiveness of each element in the poster, the best cropping area is automatically determined to highlight key elements and core information, and to avoid important content being cropped off.
[0025] In at least one embodiment, the above - mentioned short - play production method further includes user portrait processing, which is used to extract features from the multi - modal data of short - play users and perform dynamic portrait modeling on users, and it includes:
[0026] Collect the behavior data of short - play users and the content preference data for short - plays. The behavior data includes the user's clicks, viewing duration, pauses, jumps, likes, comments, and shares.
[0027] Use the time - series model to analyze the behavior data, construct the user's behavior trajectory, and mine the user's behavior patterns and habits to obtain the required behavior features.
[0028] Use the multi - modal large model to analyze the content preference data, extract the user's preference information, which includes short - play types, keywords, and picture styles, to obtain the required text features and image features.
[0029] The modal model maps the behavior features, text features, and image features into a high - dimensional space through a shared latent space to achieve consistent expression of features and complete feature fusion.
[0030] Based on the fused features, construct a dynamic user portrait to provide accurate user needs and preference information for short - play creation.
[0031] In a second aspect, the present application provides a short - play production system, which includes:
[0032] A plot generation module, which generates a plot script that fits the input according to the keywords or outline input by the creator.
[0033] A storyboard design module, which is used to convert the plot script into storyboards of the short - play.
[0034] A highlight moment recognition and content summary module, which is used to segment the short drama content to identify highlight moments; and extract video and audio content to generate a content summary that fits the plot.
[0035] A poster processing module, which is used to extract key elements based on the highlight moments and generate posters using the key elements of the highlight moments.
[0036] A user portrait processing module, which is used to extract features from the multi-modal data of short drama users and build a dynamic portrait model for the users.
[0037] Thirdly, the present application provides a computer device, including a memory and a processor. Computer program instructions are stored in the memory, and when the computer program instructions are read and run by the processor, the above-mentioned method is executed.
[0038] Fourthly, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a computer, the above-mentioned method is executed.
[0039] The present application can at least achieve the following beneficial effects:
[0040] The present application is a method for conceiving, generating, editing, and publishing short dramas based on AIGC. It not only uses AI-driven full-process management to quickly generate creative concepts, automatically write scripts, and intelligently generate materials through AI algorithms, thereby greatly shortening the production cycle and improving overall efficiency; it uses multi-modal models to analyze massive data to provide creators with novel storylines, character settings, and visual style suggestions, promoting content innovation; it also combines functions such as short drama highlights and content recognition to increase the efficiency and robustness of the system and has a wide adaptation range.
[0041] The present application realizes collaborative innovation through multi-modal technologies. The multi-modal model is applied to short drama creation, promoting collaborative innovation among various related technologies such as image recognition, natural language processing, and audio processing. Different technologies can learn from and complement each other to jointly solve the problems in short drama creation. For example, image recognition technology can provide richer visual references for material generation, natural language processing technology can optimize the quality of script writing and dialogue generation, and audio processing technology can enhance the effects of sound effects and background music. Through this collaborative innovation, multi-modal technologies will continue to improve and develop, providing users with a more immersive short drama experience.
[0042] This application has lowered the technical threshold, enabling more independent producers to participate in short drama creation. It can develop an integrated short drama platform that integrates functions such as creative conceptions and post-editing on an easy-to-use interface, reducing the learning cost and technical difficulty for users. Through a simple and intuitive interface design and operation process, users can create short dramas more conveniently and quickly. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Now, one or more embodiments of this application will be described only by way of example with reference to the accompanying drawings, in which:
[0044] Figure 1 is the overall flowchart of the short drama production method provided by the embodiment of this application;
[0045] Figure 2 is Figure 1 the flowchart of short drama AI production of the method;
[0046] Figure 3 is Figure 1 the flowchart of user portrait processing of the method;
[0047] Figure 4 is the structural block diagram of the short drama production system provided by the embodiment of this application;
[0048] Figure 5 is the structural block diagram of the computer device provided by the embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The following will describe this application in detail with reference to the exemplary embodiments in the accompanying drawings. It should be understood, however, that this application can be implemented in many different forms and should not be construed as limited to the embodiments set forth herein. These embodiments are provided herein to make the disclosure of this application more complete and to fully convey the concept of this application to those skilled in the art.
[0050] First, the noun terms related to the embodiments of this application will be explained. Obviously, the models, algorithms, etc. involved in the following terms are all prior arts in this field, and the technical solutions of the embodiments of this application are developed based on these models, algorithms, and technologies, etc.
[0051] Large Language Model (LLM): The large language model is a commonly used model in the field of natural language processing, used to process various natural language tasks, such as text classification, question answering, dialogue, etc., to generate natural language text or understand the meaning of language text. The large language model is also a general term for deep learning models trained using a large amount of text data, which can better train, analyze, reason, and output key information of the text.
[0052] Diffusion Model: The diffusion model originally stemmed from the field of statistical physics and has been widely applied in computer vision and generative model research in recent years, especially for image generation tasks. The working principle of common diffusion models is to gradually transform high-dimensional data (such as images) into random noise through a series of noise injection and removal steps, and then reverse this process to generate new data.
[0053] Time Series Model: A time series model refers to a statistical model used for analyzing and predicting data that changes over time. It is mainly used to process data arranged in chronological order. They attempt to reveal the underlying patterns behind the data in order to better understand and predict future behavior.
[0054] Large Model for Video Understanding: A video understanding model refers to a model that analyzes and understands video content through artificial intelligence technology. These models can identify objects, actions, and scenes in videos and perform temporal analysis and reasoning.
[0055] Multimodal Large Language Model (MLLM): The multimodal large language model is based on the large language model and is capable of processing and understanding various types of modal information, such as text, images, videos, and audio.
[0056] Deep Learning Model: Deep learning is a branch of machine learning that simulates the information processing method of the human brain by using multi-layer neural networks. Common deep learning models include Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), etc.
[0057] Natural Language Processing Algorithm: Natural Language Processing (NLP) is an interdisciplinary field of computer science, artificial intelligence, and linguistics, mainly studying how to enable computers to understand and process human language. It involves the computerized analysis of text and speech, aiming to develop tools and technologies that can understand and manipulate natural language to perform various tasks.
[0058] User Persona: A user persona refers to a virtual representation of real users and is a target user model built on a series of attribute data. Usually, a user persona is a labeled user model abstracted based on information such as user demographics, web browsing content, online social activities, and consumption behavior.
[0059] Attention Mechanism: The attention mechanism is a method that mimics the human visual and cognitive systems. It allows the neural network to focus on relevant parts when processing input data. By introducing the attention mechanism, the neural network can automatically learn and selectively focus on important information in the input, improving the performance and generalization ability of the model.
[0060] Spatio-temporal Constraint: Spatio-temporal constraint refers to the situation where for a certain event, behavior, or phenomenon, both the time and space of its occurrence or progress are subject to certain restrictions or constraints.
[0061] Multiresolution Analysis (MRA): Multiresolution analysis is a mathematical tool and technique used to decompose complex data or signals into components at different levels of detail. This analysis method is particularly suitable for signal processing and image analysis, capable of capturing local features of the data and providing richer information at different scales.
[0062] Reference Figure 1 , the short drama production method provided by the embodiments of the present application mainly includes two aspects. One is the short drama AI production process, and the other is user portrait processing. The short drama AI production process includes the following steps: plot generation; storyboard design; highlight moment identification and content summary; poster processing. The user portrait processing process includes the following steps: data collection; analyzing the user's behavioral data and content preference data; data mapping, feature consistency expression; constructing a dynamic user portrait. The short drama AI production process is mainly used to provide a full-process creation support tool based on AIGC for short drama creation, from creative generation to finished product optimization, improving content production efficiency and reducing production costs. The user portrait processing process is mainly used to construct a dynamic user portrait, providing accurate user needs and preference information for short drama creation to assist creators in short drama production and delivery.
[0063] Reference Figure 2 , the AI production process of the embodiments of the present application specifically includes the following steps:
[0064] S11. Plot Generation: Generate a plot script that fits the input based on the keywords or outline input by the creator; this step mainly utilizes large language models;
[0065] S12. Storyboard Design: Convert the plot script into storyboard diagrams of the short drama; this step mainly utilizes diffusion models;
[0066] S13. Highlight Moment Identification and Content Summary: Segment the short drama content to identify highlight moments; extract video and audio content to generate a content summary that fits the plot; this step mainly utilizes time series models, video understanding large models, and multimodal large models;
[0067] S14, poster processing: extract key elements based on highlight moments, and generate posters using the key elements of highlight moments; this step mainly utilizes deep learning models and diffusion models.
[0068] In step S11, the generation process of the plot script includes: using the large language model to analyze the input keywords or outline to generate a complete plot script that matches the user input; in the plot generation process, by adjusting model parameters and introducing style tags, the large language model adjusts the language expression according to the style required by the user.
[0069] Specifically, this application uses the excellent language understanding and generation capabilities of the large language model to provide creators with efficient plot creation support. When the creator enters keywords or outlines, the model first performs an in-depth semantic analysis to accurately extract key elements and core themes. Subsequently, based on the rich knowledge and language patterns accumulated during the large-scale pre-training process, the model will add finely tuned prompt words and response logic to generate a complete plot script that is highly consistent with the user input. At this time, if the keywords for the same topic are entered, the large language model generates both the outline and the script; if the outline is entered, the complete plot script is generated directly.
[0070] In the process of plot generation, this application realizes customized output of various styles and tones by flexibly adjusting model parameters and introducing unique style markers. For example, for comedy-style plots, the model will deliberately add humorous elements and light-hearted and humorous language expressions to make the plot interesting and entertaining; for suspense-style plots, it will carefully create a tense atmosphere, cleverly set up suspenseful plots, and be step-by-step and fascinating. In addition, the model can also flexibly switch to other styles such as romance, science fiction, history, etc. according to user needs to ensure that the generated plot script is both in line with the creative intent and highly artistic and ornamental. In this way, creators can not only greatly improve their creative efficiency, but also explore more creative possibilities with the assistance of the model to create unique plot works.
[0071] In step S12, the plot script is converted into storyboard images through a diffusion model. The process includes: Information encoding stage: Deep encoding operations are performed on the information in the plot script. The information includes scenes, characters, and actions, and is transformed into encoded information that can be understood and processed by the diffusion model. Generation stage of latent space image representation: The diffusion model uses a neural network structure to perform non-linear transformations and feature extractions on the encoded information, gradually generating latent feature vectors with specific semantics and structures. The latent feature vectors are mapped to the image space through a decoder network, thereby generating a preliminary image representation. In this process, a diffusion algorithm is used to continuously optimize the details and authenticity of the image representation. Optimization of shot transitions and frame layouts: The diffusion model introduces an attention mechanism and spatio-temporal constraints. Through the attention mechanism, the correlation weights between different information elements are calculated to focus on key information, thereby determining the key content of each shot. Through spatio-temporal constraints, according to the plot rhythm and time sequence, the timing of shots and the spatial layout of the frames are arranged.
[0072] Specifically, in the information encoding stage, the model performs deep encoding operations on information such as scenes, characters, and actions in the plot script. For scene information, the model quantifies and encodes detailed features such as the specific type of the scene (e.g., whether it is an indoor scene or an outdoor scene, and if it is an indoor scene, further subdivided into a living room, bedroom, or office, etc.), the spatial layout of the scene (including the relative position relationships of various objects in the scene, lighting conditions, such as natural light or artificial light, light intensity, and light direction, etc.), and the color atmosphere of the scene (e.g., warm tone or cold tone). For character information, not only the basic appearance features of the character (such as gender, age, hairstyle, skin color, etc.) are encoded, but also information such as the character's expression, posture, and clothing details are deeply analyzed to more comprehensively depict the character image. For action information, the model precisely analyzes key elements such as the type of action (e.g., walking, running, jumping, grabbing, etc.), the amplitude of the action, the starting and ending states of the action, and the time sequence of the action, and transforms these rich action details into an encoded form that can be understood and processed by the model.
[0073] In the generation stage of the latent space image representation, the model constructs the corresponding image representation in the latent space based on the encoded information. This process involves complex mathematical operations and neural network structures. Specifically, the model uses neural network structures such as the Multilayer Perceptron (MLP) to perform non-linear transformations and feature extractions on the encoded information, gradually generating latent feature vectors with specific semantics and structures. These latent feature vectors have specific distribution rules in the latent space. Through the decoder network, they are mapped to the image space to generate a preliminary image representation. In this process, to ensure that the generated image representation has high quality and rationality, the model adopts the Diffusion Model to continuously optimize the details and authenticity of the image representation.
[0074] This application introduces two key technologies, the attention mechanism and spatio-temporal constraints, into the model to further optimize shot transitions and frame layouts.
[0075] In terms of the attention mechanism, the model determines the key content of each shot by calculating the correlation weights between different information elements and focusing on the key information. Specifically, the model adopts the Multi-Head Attention mechanism to simultaneously pay attention to and analyze the input information from multiple different semantic dimensions. For example, when processing scene information, the attention mechanism focuses on the core objects or key regions in the scene; for character information, it pays more attention to the main actions and expression changes of the characters; for action information, it focuses on the key frames and key poses of the actions. In this way, the attention mechanism can determine the most expressive and informative key content for each shot, making the generated storyboard more vivid and attractive.
[0076] In terms of spatio-temporal constraints, the model reasonably arranges the shot timing and frame space layout according to the plot rhythm and time sequence. In the time dimension, the model analyzes the time clues and rhythm changes in the plot script, such as the climax, trough of the plot, and the development rhythm of the plot. Based on this information, the model determines the duration and transition timing of each shot to ensure the coherence and fluency of the entire storyboard sequence in time. For example, in the climax part of the plot, the shot duration may be appropriately shortened to speed up the rhythm and enhance the sense of tension; while in the slow-paced part of the plot, the shot duration can be appropriately extended to create a quiet and steady atmosphere. In the space dimension, the model considers the composition principles and visual balance of the frame, and reasonably arranges the positions and proportional relationships of elements such as characters, scenes, and props in the frame. For example, classic composition methods such as the golden ratio composition method and the symmetric composition method are used to make the frame more beautiful and coordinated. At the same time, the model also flexibly adjusts the perspective and focal length of the frame according to the plot needs to highlight the key content or create a specific visual effect.
[0077] Through the above carefully designed technical solutions, it is possible to generate high-quality, creative and plot-logic-compliant storyboard pictures, providing strong technical support for fields such as film and television production and animation design. Based on the above plot script and storyboard pictures, creators can quickly create short plays, and the creation can be carried out in traditional shooting methods, or in a combination of traditional shooting and video large model processing, or directly generated through the video large model, etc.
[0078] In step S13, the recognition process of the highlight moment is as follows: using a time series model to process the time series data of the short play content, and combining with a large video understanding model to perform semantic analysis on the video content, so as to achieve accurate segmentation of the short play content and recognition of the highlight moment. The generation process of the content summary is as follows: using a multi-modal large model for processing, respectively analyzing the visual information in the video picture and the auditory information in the audio, and then integrating and analyzing the two outputs of the model through a set workflow, so as to accurately identify the short play content and summarize the key points; the content summary is mainly in the form of an introduction and explanation of the short play content.
[0079] Specifically, this application uses a time series model and a large video understanding model to segment the content fragments of the short play, so as to accurately identify the key plots and highlight segments therein. The time series model can adopt a Transformer or LSTM network; the Transformer network, with its powerful parallel computing ability and long sequence modeling advantages, can capture the global features of the short play content in the time dimension; the LSTM network is good at dealing with long-term dependencies in the sequence and effectively modeling the development and changes of the short play plot. The large video understanding model is responsible for performing semantic analysis on the video content, and the two cooperate with each other to achieve accurate segmentation of the short play content fragments and accurate recognition of the highlight moment.
[0080] At the same time, this application also uses a multi-modal large model (including a large video understanding model and a large audio understanding model) to extract video and audio content. With the collaborative operation between models and a set workflow, it realizes the intelligent recognition of the short play content, summarizes the key information, and then generates an introduction and explanation that fits the plot, so as to improve the operation efficiency. Specifically, the large video understanding model is responsible for analyzing visual information such as scenes and character actions in the video picture, and the large audio understanding model focuses on processing auditory information such as lines and sound effects in the audio. Through the workflow designed and arranged by the present invention, the outputs of these two types of models are integrated and analyzed, so as to accurately identify the short play content and summarize the key points.
[0081] In step S14, the process of generating the poster includes: extracting key elements from the highlight moments, where the key elements include the main character image and the iconic scene; inputting the key elements into a diffusion model to generate a poster design template; and using a multi-resolution image generation and cropping algorithm to process the poster design template to generate posters suitable for different screens. The processing process of the multi-resolution image generation and cropping algorithm includes: Image generation: Using multi-resolution analysis technology, according to the target size of the poster and the display characteristics of different terminal devices, generate image versions with different resolution levels; through feature fusion and detail enhancement of images with different resolutions, ensure that the poster retains key information while having rich visual details; Image cropping: The cropping process introduces an attention mechanism, and by analyzing the importance and visual attractiveness of each element in the poster, automatically determine the best cropping area to highlight key elements and core information and avoid important content being cropped out.
[0082] Specifically, first extract key elements from the short play. In this stage, image recognition technology and natural language processing algorithms are used to accurately extract key elements from the short play. For the main character image, through a deep learning model, extract the features of the character images in the short play, identify iconic elements such as the character's facial features, clothing characteristics, and body movements, and use a feature extractor based on a convolutional neural network (CNN) to convert the visual information of the character into a feature vector that can be understood and processed by a computer, so as to accurately locate and capture the unique image of each main character. For the iconic scene, analyze it from the spatio-temporal dimension of the video, identify representative scene images, including key information such as the architectural style, color scheme, and environmental layout of the scene, and assign accurate semantic labels to each iconic scene through semantic analysis technology for subsequent rapid retrieval and utilization of these elements.
[0083] Then, the diffusion model is used to automatically generate a poster design template that fits the content characteristics. Using an advanced diffusion model, the model is based on the principles of probability distribution and random walk, and can accurately control image generation in high-dimensional space. Taking the extracted key elements as input, the diffusion model learns and analyzes a large amount of existing high-quality poster data to explore the style patterns and visual laws. In the process of generating poster design templates, the diffusion model gradually optimizes the image in an iterative manner, starting from the initial random noise, and gradually generates a poster design template that fits the content characteristics according to the characteristics of the key elements and predefined style parameters. For example, if the short play is a costume theme, the model will automatically generate color matching, composition methods, and decorative patterns with ancient elements to ensure that the poster can accurately convey the theme and style of the short play. Finally, a multi-resolution image generation and cropping algorithm is used. This algorithm is based on multi-scale analysis technology and can process posters at different resolutions. In the image generation process, the algorithm generates image versions with different resolution levels according to the target size of the poster and the display characteristics of different terminal devices. By fusion and detail enhancement of images of different resolutions, it is ensured that the poster retains key information while having rich visual details. In the cropping process, the algorithm uses an image cropping method based on the attention mechanism to automatically determine the best cropping area by analyzing the importance and visual appeal of each element in the poster, highlighting key elements and core information, and avoiding the cropping of important content. At the same time, the algorithm will automatically adjust the poster's color balance, contrast, brightness and other visual parameters to adapt to the screen characteristics of different terminal devices, ensuring that the poster can present the best visual effects on different terminal devices (such as mobile phones, tablets, computers, outdoor large screens, etc.). Whether on a high-resolution retina screen or on an ordinary screen with different color display capabilities, the audience can see a clear, beautiful and attractive poster.
[0084] refer to Figure 3 The user portrait processing flow of the embodiment of the present application is used to extract features from the multimodal data of the short drama user and to perform dynamic portrait modeling on the user, and specifically includes the following steps:
[0085] S21. Collecting the behavior data of the users of the skit and the content preference data of the skit, wherein the behavior data includes the users' clicks, viewing time, pauses, jumps, likes, comments and shares;
[0086] S22, using a time series model to analyze the behavior data, construct the user's behavior trajectory, and mine the user's behavior patterns and habits, so as to obtain the required behavior characteristics; using a multimodal large model to analyze the content preference data, extract the user's preference information, the preference information includes the type of skit, keywords and picture style, so as to obtain the required text features and image features;
[0087] S23. The modal model maps the behavioral features, text features, and image features into a high-dimensional space through a shared latent space to achieve consistent expression of features and complete feature fusion.
[0088] S24. Based on the fused features, construct a dynamic user profile to provide accurate user needs and preference information for short drama creation.
[0089] Reference Figure 4 , an embodiment of the present application provides a short drama production system 100, which is used to implement the above short drama production method and includes the following modules:
[0090] A plot generation module 110, which generates a plot script that fits the input according to the keywords or outline input by the creator.
[0091] A storyboard design module 120, which is used to convert the plot script into a storyboard of the short drama.
[0092] A highlight moment recognition and content summary module 130, which is used to segment the short drama content to identify highlight moments; and extract video and audio content to generate a content summary that fits the plot.
[0093] A poster processing module 140, which is used to extract key elements according to the highlight moments and generate a poster using the key elements of the highlight moments.
[0094] A user profile processing module 150, which is used to extract features from the multi-modal data of short drama users and perform dynamic portrait modeling on users.
[0095] Each module of the short drama production system 100 is generally a software module with relatively independent functions, or can be a module combining software and hardware. Each module interacts with each other to facilitate the production of short dramas. The interaction interfaces of each module can be integrated on an easy-to-use interface, reducing the learning cost and technical difficulty of users. Through a simple and intuitive interface design and operation process, users can create short dramas more conveniently and quickly. At the same time, the system framework formed by the cooperation of each module has good scalability, and more applications for actual needs of short dramas can be realized by adding more modules and model training, which is conducive to subsequent iterative updates.
[0096] Finally, the present application also provides a computer device and a computer-readable storage medium. As Figure 4 shown, the computer device 200 includes a memory 210 and a processor 220. A computer program is stored in the memory 210, and when the computer program is read and run by the processor 220, it executes the above method. The computer-readable storage medium has a computer program stored thereon, and when the computer program is run by a computer, it executes the above method.
[0097] The computer device described in the embodiments of the present application is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, servers, mainframe computers, and other suitable computers. The computer device may also represent various forms of mobile devices, such as, personal digital processing devices, smart phones, wearable devices, and other similar computing devices. However, based on the requirement of being able to stably and securely process large-scale data, the computer device should preferably be in the form of a desktop computer, a workstation, a server, a mainframe computer, etc. Further, the components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein. In addition to including a memory and a processor, the computer device may further include a display, etc. These elements are all common components or devices in the art, and the types and models included therein are all conventional selections, which are not elaborated herein in the present application.
[0098] The computer-readable medium described in the embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. A computer-readable storage medium may be, for example (but not limited to), an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.
[0099] It should be understood that all of the above embodiments are exemplary rather than restrictive. Under the concept of this application, those skilled in the art should make various modifications or deformations to the specific embodiments described above, and all of them should fall within the protection scope of this application.
Claims
1. A method for producing a short play, characterized in that: include: Plot generation: Based on the keywords or outlines input by the creator, a plot script that matches the input is generated; Storyboard design: converting the plot script into storyboards for the skit; Highlight moment identification and content summary: segment the content of the short play and identify the highlight moments; extract the video and audio content to generate a content summary that fits the plot; Poster processing: extract key elements based on highlight moments and generate posters using the key elements of highlight moments.
2. The skit production method according to claim 1, characterized in that: The generation process of the plot script includes: Analyze the input keywords or outline using a large language model to generate a complete plot script that matches the user input; During the plot generation process, the large language model adjusts the language expression according to the style required by the user by adjusting model parameters and introducing style markers.
3. The skit production method according to claim 1, characterized in that: The plot script is converted into a storyboard through a diffusion model, and the process includes: Information encoding stage: Deep encoding operations are performed on the information in the plot script, including scenes, characters, and actions, and converted into coded information that can be understood and processed by the diffusion model; Latent space image representation generation stage: The diffusion model uses a neural network structure to perform nonlinear transformation and feature extraction on the encoded information, and gradually generates a latent feature vector with specific semantics and structure; the latent feature vector is mapped to the image space through a decoder network to generate a preliminary image representation. In this process, the diffusion algorithm is used to continuously optimize the details and authenticity of the image representation; Optimize shot switching and screen layout: The diffusion model introduces attention mechanism and time-space constraints. Through the attention mechanism, the correlation weights between different information elements are calculated to focus on key information, thereby determining the key content of each shot. Through time-space constraints, the shot timing and screen space layout are arranged according to the plot rhythm and time sequence.
4. The skit production method according to claim 1, characterized in that: The highlight moment recognition process is as follows: using a time series model to process the time series data of the skit content, and combining the video understanding model to perform semantic analysis on the video content, thereby achieving accurate segmentation of the skit content and recognition of the highlight moments; The generation process of the content summary is as follows: a multimodal large model is used to analyze the visual information in the video screen and the auditory information in the audio respectively, and then the two outputs of the model are integrated and analyzed through a set workflow, so as to accurately identify the content of the short play and summarize the key points.
5. The skit production method according to claim 1, characterized in that: The poster generation process includes: Use convolutional neural networks to extract key elements from highlight moments, including images of main characters and iconic scenes; Inputting the key elements into a diffusion model to generate a poster design template; The poster design template is processed using a multi-resolution image generation and cropping algorithm to generate posters suitable for different screens.
6. The skit production method according to claim 5, characterized in that: The processing of the multi-resolution image generation and cropping algorithm includes: Image generation: Multi-resolution analysis technology is used to generate image versions with different resolution levels according to the target size of the poster and the display characteristics of different terminal devices. By fusing features and enhancing details of images with different resolutions, the poster is ensured to have rich visual details while retaining key information. Image cropping: The attention mechanism is introduced into the cropping process. By analyzing the importance and visual appeal of each element in the poster, the optimal cropping area is automatically determined to highlight key elements and core information and avoid cropping of important content.
7. The skit production method according to claim 1, characterized in that: It also includes user portrait processing, which is used to extract features from the multimodal data of skit users and model dynamic user portraits, including: Collecting the behavior data of the users of the skits and their content preference data, including the users’ clicks, viewing time, pauses, jumps, likes, comments and shares; Use time series models to analyze the behavior data, construct user behavior trajectories, and mine user behavior patterns and habits to obtain the desired behavior characteristics; A multimodal large model is used to analyze the content preference data and extract the user's preference information, which includes the type of skit, keywords and picture style, so as to obtain the required text features and image features; The modal model uses a shared latent space to uniformly map behavioral features, text features, and image features into a high-dimensional space, achieving consistent feature expression and completing feature fusion. Based on the fused features, a dynamic user portrait is constructed to provide accurate user demand and preference information for short drama creation.
8. A short play production system, characterized in that: include: The plot generation module generates a plot script that matches the input based on the keywords or outlines entered by the creator; A storyboard design module, which is used to convert the plot script into a storyboard for a short play; The highlight moment recognition and content summary module is used to segment the content of the short play and identify the highlight moments; it also extracts the video and audio content to generate a content summary that fits the plot; A poster processing module is used to extract key elements according to the highlight moments and generate posters using the key elements of the highlight moments; The user portrait processing module is used to extract features from the multimodal data of short drama users and dynamically model the user portraits.
9. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the computer program is read and executed by the processor, the method according to any one of claims 1 to 7 is executed.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a computer, the method according to any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Pre-notice automatic generation method and product based on movie and television drama plot content
CN117221670A
Industrial micro-drama digital asset production system and method
CN118396263A
Multi-agent-based automatic generation method, system and terminal for micro-drama
CN119383413A
Automated video generation
US12176007B1
Method and apparatus for generating multimedia data, and readable medium and electronic device
WO2023016391A1