High-quality video content automatic generation method and related equipment
By performing multimodal feature extraction, content semantic analysis and deep reinforcement learning optimization on video materials, high-quality video content is generated, and the problem of lack of depth and attractiveness of video generation in the existing technology is solved, and the automatic generation of high-quality video content is achieved.
Patent Information
- Application Number
- CN202510527860.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-25
AI Technical Summary
Existing automated video generation technology is difficult to deeply explore potential story clues in video materials, resulting in the lack of depth and appeal of the generated content, and the lack of effective multimodal feature extraction and analysis mechanisms, limiting video quality and expressiveness.
By extracting the input original video material multimodal feature to obtain the video feature vector set; performing content semantic analysis on the video feature vector set to obtain the video semantic map; using deep reinforcement learning technology to optimize the video semantic map to obtain the optimized video story line; multimodal content is generated based on the optimized video story line to obtain the candidate video clip set; intelligent editing and synthesis of the candidate video clip set to obtain the target high-quality video content.
It has achieved an in-depth understanding of potential story clues in video materials, and the generated video content has higher depth and appeal, which has improved the video quality and expressiveness, and met the market's demand for high-quality video content.
Smart Images

Figure CN120050487A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic video content generation, and particularly to a method for automatically generating high-quality video content and related devices. Background Art
[0002] In today's era of rapid digital development, video content, as one of the main forms of information dissemination, has seen an explosive growth in demand. Whether it is social media, online education, or corporate promotion, the demand for high-quality and attractive video content has reached an unprecedented level. However, traditional video production methods are not only time-consuming and laborious but also require professional skills and knowledge, which discourages a large number of individuals and small enterprises with video content creation needs but lacking corresponding resources. Therefore, exploring a method that can automatically generate high-quality video content has become the key to meeting market demand.
[0003] One of the main existing problems is that although existing technologies can achieve automated video generation to a certain extent, they often neglect the narrative structure and semantic coherence of video content. Many automated attempts are limited to simple image or clip splicing and fail to deeply explore the potential storylines in video materials, resulting in the generated content lacking depth and attraction. In addition, due to the lack of an effective multi-modal feature extraction and analysis mechanism, existing methods are difficult to comprehensively understand the complex information contained in video materials, such as visual, auditory, and emotional elements, thus limiting the quality and expressiveness of the final video.
[0004] To address the above problems, researchers have begun to focus on how to optimize the video content generation process through advanced technical means, such as deep reinforcement learning. This method not only requires the ability to accurately extract and analyze the multi-modal features of video materials but also the ability to transform these features into a logical and appealing storyline. At the same time, with the growing demand for personalized content from users, how to use technical means to intelligently edit and synthesize videos to meet application requirements in different scenarios has also become an important research direction. The existence of these problems has promoted the research and development of more efficient and intelligent methods for automatically generating video content. Summary of the Invention
[0005] The main object of the present invention is to provide a method for automatically generating high-quality video content and related devices, which solves the technical problem that many automated attempts are limited to simple image or clip splicing and fail to deeply explore the potential storylines in video materials, resulting in the generated content lacking depth and attraction.
[0006] To achieve the above object, the present invention provides a method for automatically generating high-quality video content, including the following steps: Perform multi-modal feature extraction on the input original video material to obtain a video feature vector set; Perform content semantic analysis on the video feature vector set to obtain a video semantic graph; Optimize the narrative structure of the video semantic graph through deep reinforcement learning technology to obtain an optimized video story line; Perform multi-modal content generation based on the optimized video story line to obtain a candidate video clip set; Perform intelligent editing and synthesis on the candidate video clip set to obtain the target high-quality video content.
[0007] Furthermore, the performing multi-modal feature extraction on the input original video material to obtain a video feature vector set includes: Perform spatio-temporal dimension decomposition on the input original video material to obtain a video frame sequence set and an audio signal sequence, and perform deep optical flow field calculation on the video frame sequence set to obtain a scene motion feature tensor; Perform hierarchical semantic segmentation on the video frame sequence set based on the scene motion feature tensor to obtain a visual semantic feature map, and perform cross-modal attention mapping on the visual semantic feature map to obtain a visual-semantic association feature matrix; Perform multi-modal temporal fusion on the visual-semantic association feature matrix and the audio signal sequence to obtain a multi-modal feature vector set, and perform context dependency analysis based on the multi-modal feature vector set to obtain a video feature vector set, where the video feature vector set includes a scene semantic relationship graph, temporal dynamic features, and multi-modal interaction patterns.
[0008] Furthermore, the performing content semantic analysis on the video feature vector set to obtain a video semantic graph includes: Perform emotional intensity analysis on the video feature vector set to obtain a multi-dimensional emotional feature vector, and perform temporal correlation calculation on the multi-dimensional emotional feature vector to obtain an emotional evolution trajectory graph, where the emotional evolution trajectory graph includes an emotional change gradient, an emotional density distribution, and an emotional inflection point marker; Perform hierarchical topic mining on the video feature vector set based on the emotional evolution trajectory graph to obtain a topic structure tree, and perform semantic relevance reasoning on the topic structure tree to obtain a topic association network, where the topic association network includes a topic importance score, key event nodes, and a topic migration path; Perform multi-angle scene semantic analysis on the topic association network to obtain a scene semantic feature set, and perform cross-scene semantic mapping based on the scene semantic feature set to obtain a scene transformation matrix, where the scene transformation matrix includes a scene continuity vector, a scene switching rule, and a scene semantic similarity; Perform deep semantic fusion on the scene transformation matrix and the theme association network to obtain a semantic association topology graph, and perform multi-dimensional feature organization based on the semantic association topology graph to obtain a video semantic map, where the video semantic map includes narrative logic links, semantic hierarchical structures, and content main line frameworks.
[0009] Further, optimizing the narrative structure of the video semantic map through deep reinforcement learning technology to obtain an optimized video story line, including: Segment the narrative units of the video semantic map to obtain a sequence of basic narrative segments, and calculate the plot tension of the sequence of basic narrative segments to obtain a narrative tension curve; Construct a narrative state space for the sequence of basic narrative segments based on the narrative tension curve to obtain a narrative state transition network, and perform a reward function mapping on the narrative state transition network to obtain a narrative optimization strategy space, where the narrative optimization strategy space includes state value evaluation, action selection probability, and transfer reward distribution; Perform multiple rounds of exploration sampling on the narrative optimization strategy space to obtain a set of candidate narrative paths, and perform a narrative quality evaluation based on the set of candidate narrative paths to obtain an optimal narrative decision sequence, where the optimal narrative decision sequence includes plot integrity data, narrative coherence metrics, and audience attention prediction; Reconstruct the narrative structure of the optimal narrative decision sequence through deep reinforcement learning technology to obtain an optimized narrative framework, and reconstruct the temporal relationship based on the optimized narrative framework to obtain an optimized video story line, where the optimized video story line includes the main story line context, key plot nodes, and narrative rhythm control parameters.
[0010] Further, constructing a narrative state space for the sequence of basic narrative segments based on the narrative tension curve to obtain a narrative state transition network, including: Extract the peak features of the narrative tension curve to obtain a sequence of key tension nodes, and perform a temporal dependence analysis on the sequence of key tension nodes to obtain a node association graph; Encode the state features of the sequence of basic narrative segments based on the node association graph to obtain a set of narrative state vectors, and calculate the semantic distance of the set of narrative state vectors to obtain a state transition matrix; Divide the action space of the state transition matrix to obtain a narrative decision space, and construct state transition rules based on the narrative decision space to obtain a set of state transition rules; Perform a network topology mapping on the state transition rule set to obtain an initial transition network, and perform structural optimization and organization based on the initial transition network to obtain a narrative state transition network, where the narrative state transition network includes node connection metrics, path reachability indicators, and network stability parameters.
[0011] Furthermore, perform multi-modal content generation based on the optimized video story line to obtain a candidate video clip set, including: Decompose the optimized video story line into scene element configurations to obtain a scene element configuration table, and perform visual style matching on the scene element configuration table to obtain a scene rendering parameter set; Based on the scene rendering parameter set, reconstruct the visual content of the scene element configuration table to obtain a visual scene sequence, and perform dynamic special effect synthesis on the visual scene sequence to obtain a visual effect enhancement sequence; Perform audio emotion mapping on the visual effect enhancement sequence to obtain an emotional audio feature set, and perform audio content generation based on the emotional audio feature set to obtain an audio effect sequence; Based on the audio effect sequence, optimize the audio-visual synchronization of the visual effect enhancement sequence to obtain an initial synthesized clip set, and perform spatio-temporal consistency calibration on the initial synthesized clip set to obtain a calibrated synthesis sequence; Perform quality assessment and screening on the calibrated synthesis sequence to obtain a quality scoring matrix, and perform diversified reordering based on the quality scoring matrix to obtain a candidate video clip set.
[0012] Furthermore, perform intelligent editing and synthesis on the candidate video clip set to obtain target high-quality video content, including: Perform shot boundary detection on the candidate video clip set to obtain a shot segmentation sequence, and perform temporal combination optimization on the shot segmentation sequence to obtain a shot arrangement plan; Based on the shot arrangement plan, design a transition effect for the shot segmentation sequence to obtain a transition special effect sequence, and enhance the visual coherence of the transition special effect sequence to obtain a visually smooth sequence; Perform audio rhythm matching on the visually smooth sequence to obtain an audio-visual fusion sequence, and perform multi-track mixing arrangement based on the audio-visual fusion sequence to obtain a stereo effect plan, where the stereo effect plan includes channel balance parameters, audio transition curves, and spatial sound field layouts; Perform spatio-temporal synchronization integration on the stereo effect plan and the visually smooth sequence to obtain target high-quality video content.
[0013] The present invention also provides a device for automatically generating high-quality video content, including: An extraction module for performing multi-modal feature extraction on the input original video material to obtain a video feature vector set; An analysis module for performing content semantic analysis on the video feature vector set to obtain a video semantic map; An optimization module for optimizing the narrative structure of the video semantic map through deep reinforcement learning technology to obtain an optimized video story line; A generation module for performing multi-modal content generation based on the optimized video story line to obtain a candidate video clip set; A synthesis module for performing intelligent editing and synthesis on the candidate video clip set to obtain the target high-quality video content.
[0014] The present invention also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0015] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0016] A method for automatically generating high-quality video content provided by the present invention includes the following steps: performing multi-modal feature extraction on the input original video material to obtain a video feature vector set; performing content semantic analysis on the video feature vector set to obtain a video semantic map; optimizing the narrative structure of the video semantic map through deep reinforcement learning technology to obtain an optimized video story line; performing multi-modal content generation based on the optimized video story line to obtain a candidate video clip set; performing intelligent editing and synthesis on the candidate video clip set to obtain the target high-quality video content, solving the technical problem that many automated attempts are limited to simple image or segment splicing and fail to deeply explore the potential story clues in the video material, resulting in the generated content lacking depth and attraction, and achieving the technical effect that content semantic analysis can be performed based on the obtained video feature vector set, deeply understanding the themes, emotions and other implicit information contained in the video material, and converting it into a structured video semantic map. This step not only improves the depth of understanding of the video material, but also provides a clear logical framework for the optimization of the narrative structure. Description of the Drawings
[0017] Figure 1 is a schematic diagram of the steps of the method for automatically generating high-quality video content in an embodiment of the present invention; Figure 2 is a structural block diagram of the device for automatically generating high-quality video content in an embodiment of the present invention; Figure 3It is a schematic block diagram of a computer device according to an embodiment of the present invention.
[0018] The implementation, functional features and advantages of the present invention will be further described in conjunction with embodiments with reference to the accompanying drawings. Detailed implementation manners
[0019] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0020] As Figure 1 shown, Figure 1 It is a schematic diagram of the steps of a method for automatically generating high-quality video content according to an embodiment of the present invention; An embodiment of the present invention provides a method for automatically generating high-quality video content, including the following steps: Step S1, perform multi-modal feature extraction on the input original video material to obtain a video feature vector set.
[0021] Specifically, performing multi-modal feature extraction on the input original video material to obtain a video feature vector set is the basic link of the entire method for automatically generating high-quality video content. The core lies in extracting multi-dimensional information from the video material that can comprehensively reflect its content characteristics and converting it into a structured video feature vector set. Specifically, the original video material usually contains multi-modal data such as visual, auditory, and possibly text information. Therefore, corresponding feature extraction algorithms need to be designed for these modalities respectively during the extraction process. For example, in the visual modality, features such as color distribution, object categories, and action postures in the picture can be extracted through a convolutional neural network (CNN); in the auditory modality, the intonation, rhythm, background music, etc. in the audio can be extracted using the short-time Fourier transform or Mel-frequency cepstral coefficients (MFCC); and if the video contains subtitles or speech-to-text content, keywords, sentiment tendencies, and semantic relationships in the text can also be extracted through natural language processing techniques. These multi-modal features are integrated into a unified video feature vector set after being standardized, thus laying the foundation for subsequent content semantic analysis. For example, in the production scenario of a corporate promotional video, the original video material may include the work scenes of company employees, product display segments, and background commentary audio, etc. By performing multi-modal feature extraction on these materials, the system can identify the expressions and actions of employees, the appearance details of products, the key information in the commentary, etc., and finally form a vector set containing various modal features. This not only retains the core information of the original material but also provides rich data support for subsequent generation of high-quality video content.
[0022] Step S2: Perform content semantic analysis on the video feature vector set to obtain a video semantic map.
[0023] Specifically, performing content semantic analysis on the video feature vector set to obtain a video semantic map aims to deeply understand the inherent meaning and structure of video materials, thus providing a logical framework for subsequent steps. First, based on the previously obtained video feature vector set, the system needs to use technologies such as natural language processing and computer vision to analyze the semantic information behind these data. For example, for the video feature vector set extracted from a corporate promotional video, the system will identify and analyze multimodal information such as the actions of people, product details, and the emotional color of background music contained therein, and further explore the potential connections between them. Next, by constructing an association model, the system can organize these scattered information points to form a network that comprehensively reflects the essence of the video content and its interrelationships - namely, the video semantic map. In this map, each node represents a specific content element, such as a key person in a certain scene or the core selling point of a product, while the edges represent the relationships between these elements, such as causal relationships, chronological order, or emotional coherence. Taking a corporate promotional video as an example, if the video shows the cooperation process of company employees on an innovative project, the video semantic map will not only mark the personnel involved in the project, the tools and technologies used, but also reveal how these elements jointly affect the progress of the project, enabling the production team to clearly see the development context of the whole story and providing strong support for optimizing the narrative structure. In this way, the video semantic map formed through the content semantic analysis of the video feature vector set enables the creator to more intuitively grasp the core value and potential story line of the video materials, laying a solid foundation for creating more attractive video content.
[0024] Step S3: Optimize the narrative structure of the video semantic map through deep reinforcement learning technology to obtain an optimized video story line.
[0025] Specifically, the narrative structure of the video semantic graph is optimized through deep reinforcement learning technology to obtain an optimized video story line. This process utilizes advanced machine learning algorithms to enhance the narrative logic of video content and audience attraction. First, based on the video semantic graph generated in the previous steps, the system can identify each content element and its interrelationships, providing the basic data for narrative structure optimization. Next, a deep reinforcement learning model is adopted. This model automatically adjusts the story line to achieve the best performance by simulating the effects of different narrative methods. Specifically, the model will continuously explore the optimal narrative path during the training process according to preset goals (such as maximizing audience engagement or emotional resonance). For example, in the application scenario of a corporate promotional video, if the video aims to show the process of an innovative project from concept to implementation, the deep reinforcement learning technology can analyze which key nodes (such as teamwork, technological innovation points, etc.) should be emphasized and how these nodes should be connected to better tell the whole story. By repeatedly experimenting with different combinations and sequences and based on the feedback mechanism (such as audience reaction simulation or historical data), the system gradually optimizes a story line that is both logical and engaging. This optimized video story line not only ensures the effectiveness of information transmission but also enhances the emotional coherence and attraction of the viewing experience, making the final generated video not only accurately convey the core values of the enterprise but also touch the hearts of the audience and stimulate their emotional resonance. Therefore, optimizing the narrative structure of the video semantic graph with the help of deep reinforcement learning technology has become an indispensable part of creating high-quality video content.
[0026] Step S4: Based on the optimized video story line, multi-modal content generation is performed to obtain a set of candidate video segments.
[0027] Specifically, based on the optimized video story line, multi-modal content generation is performed to obtain a set of candidate video clips. This process transforms the narrative logic formed in the previous step into a specific content presentation form, thereby providing diverse material choices for the final video production. First, the system will clarify the content and emotions to be expressed at each key node according to the optimized video story line, and combine with the previously extracted multi-modal feature vector set to screen out the matching pictures, sounds, and text information from the original video materials. For example, in the scenario of a corporate promotional video, if the story line emphasizes the importance of teamwork, the system will preferentially select those scene segments that show employees collaborating, accompanied by positive background music or commentary, while ensuring that the colors and compositions in the pictures can convey a harmonious atmosphere. Next, through multi-modal content generation technology, the system will not only recombine the existing materials, but also use advanced technologies such as generative adversarial networks (GANs) to supplement the missing content, such as generating transitional shots from specific perspectives or enhancing the expressiveness of certain details. In this way, the system can generate multiple sets of candidate video clips that meet the narrative requirements. These clips not only retain the core information of the original materials, but also make innovative expressions in multiple dimensions such as vision and hearing. Taking a corporate promotional video as an example, the set of candidate video clips may include different styles of edited versions, some focusing on demonstrating the professionalism of the team, while others pay more attention to emotional resonance. This provides a rich selection space for the next intelligent editing and synthesis, and also ensures that the final generated video content can meet diverse communication needs.
[0028] Step S5: Perform intelligent editing and synthesis on the set of candidate video clips to obtain the target high-quality video content.
[0029] Specifically, the candidate video clip set is intelligently edited and synthesized to obtain the target high-quality video content. This process is to integrate the previously generated diverse clips into a coherent and high-quality final video work. First, based on the optimized video storyline, the system analyzes the role of each candidate video clip in the narrative structure, and selects the most suitable clips for combination according to its visual, auditory and emotional expressive characteristics. For example, in the application scenario of corporate promotional videos, if the storyline emphasizes the combination of teamwork and innovative achievements, the system will select those clips that can both show the details of employee collaboration and highlight the highlights of the product, and adjust their order and duration through intelligent algorithms to ensure that the overall rhythm is smooth and the logic is clear. Next, the system will use advanced video editing technology to seamlessly connect the selected clips, such as through timeline alignment, color correction and audio mixing, to eliminate the sense of separation between different clips, while enhancing the consistency of the picture and sound. In addition, intelligent editing will dynamically adjust the content according to the potential needs of the audience, such as adding appropriate transition effects or subtitles to better convey information and attract attention. Taking a corporate promotional video as an example, the system may insert a slow-motion close-up after showing a clip of teamwork to emphasize the details of a key product. This design not only improves the viewing experience of the video, but also strengthens the communication effect of the core information. Ultimately, the target high-quality video content that has been intelligently edited and synthesized not only retains the core value of the original material, but also achieves a high degree of professionalism and attractiveness through multi-modal optimization processing, thereby meeting the needs of corporate promotion and achieving the expected communication effect.
[0030] In a specific embodiment, the multimodal feature extraction is performed on the input original video material to obtain a video feature vector set, including: Decomposing the input original video material in terms of time and space dimensions to obtain a video frame sequence set and an audio signal sequence, and performing deep optical flow field calculation on the video frame sequence set to obtain a scene motion feature tensor; Based on the scene motion feature tensor, the video frame sequence set is subjected to hierarchical semantic segmentation to obtain a visual semantic feature map, and the visual semantic feature map is subjected to cross-modal attention mapping to obtain a visual-semantic association feature matrix; The visual-semantic association feature matrix and the audio signal sequence are subjected to multimodal temporal fusion to obtain a multimodal feature vector set, and context dependency analysis is performed based on the multimodal feature vector set to obtain a video feature vector set, wherein the video feature vector set includes a scene semantic relationship graph, temporal dynamic features and a multimodal interaction pattern.
[0031] Specifically, multi-modal feature extraction is performed on the input original video material to obtain a video feature vector set, which is one of the key steps in realizing the automatic generation of high-quality video content. First, the system needs to decompose the input original video material in the spatio-temporal dimension, splitting the video into a series of consecutive video frame sequences and audio signal sequences. In this process, the video frame sequence contains the temporal evolution information of the scene, while the audio signal sequence provides auditory cues such as background music, dialogue, and environmental sounds. Taking a corporate promotional video as an example, assume that the video shows the cooperation process of the company's team during the development of a new product. Through spatio-temporal dimension decomposition, the system can identify the corresponding images and accompanying sounds for each key moment, laying the foundation for subsequent processing. Next, for the information extracted from the video frame sequence set, the system will perform deep optical flow field calculation to obtain the scene motion feature tensor. Deep optical flow field calculation is a technique used to capture the pixel movement between adjacent frames, which can effectively represent the dynamic changes in the scene, such as the actions of people or the movement of objects. In the application scenario of a corporate promotional video, if the video shows the scene of team members working busily in the laboratory, deep optical flow field calculation can help identify the details of these actions, such as the gestures of employees and the operation of experimental equipment, and then form a tensor representation that describes the dynamic changes of the entire scene. Based on the obtained scene motion feature tensor, the system will perform hierarchical semantic segmentation on the video frame sequence set to obtain the visual semantic feature map. Hierarchical semantic segmentation means dividing the video frames into different semantic categories (such as people, objects, background) and constructing a map depicting the relationships between various elements on this basis. Then, through cross-modal attention mapping on the visual semantic feature map, a visual-semantic association feature matrix can be obtained. This step aims to establish the connection between visual elements and semantic information, enabling the system to not only recognize the content on the screen but also understand its underlying meaning. For example, in a corporate promotional video, the system can distinguish the team members in the foreground and the laboratory equipment in the background through hierarchical semantic segmentation, and reveal the interaction relationship between the two through cross-modal attention mapping, such as an employee operating a specific piece of equipment. Subsequently, the system will perform multi-modal temporal fusion on the obtained visual-semantic association feature matrix and the audio signal sequence to generate a multi-modal feature vector set. This step combines the visual information and audio information of the video, realizing the effective integration of data in two modes, ensuring that the final generated content contains both rich visual details and vivid sound effects. For example, in a corporate promotional video, the system may combine the images showing team cooperation with the positive discussion sounds in the background to enhance the emotional resonance of the audience. Finally, context-dependent relationship analysis is performed based on the multi-modal feature vector set to obtain the video feature vector set. Context-dependent relationship analysis helps the system understand the roles of each segment in the entire storyline and their interactions, including the scene semantic relationship graph, temporal dynamic features, and multi-modal interaction patterns.This means that the system not only has to recognize the content of each individual segment but also understand how they together form a coherent story. For example, in the context of a corporate promotional video, the system can identify the logical connection between the teamwork scene and the product demonstration scene and adjust the presentation of both accordingly, making the overall narrative more smooth and natural. To sum up, through a series of complex processes such as spatio-temporal dimension decomposition, deep optical flow field calculation, hierarchical semantic segmentation, cross-modal attention mapping, multi-modal temporal fusion, and context-dependent relationship analysis of the input original video material, the system can extract a comprehensive and structured set of video feature vectors from the original material. These feature vectors not only contain the basic information of the video content but also reflect the deep semantic relationships and emotional colors contained therein, providing solid data support and technical guarantee for subsequent content semantic analysis, narrative structure optimization, and ultimately the generation of high-quality video content with the desired goals. In the specific application scenario of a corporate promotional video, this refined data processing method helps to accurately capture and present the core values and cultural characteristics of the company, thus effectively improving the quality and influence of the video content.
[0032] In a specific embodiment, the content semantic analysis of the set of video feature vectors to obtain a video semantic map includes: Performing emotional intensity analysis on the set of video feature vectors to obtain a multi-dimensional emotional feature vector, and performing temporal correlation calculation on the multi-dimensional emotional feature vector to obtain an emotional evolution trajectory map, where the emotional evolution trajectory map includes an emotional change gradient, an emotional density distribution, and an emotional inflection point marker; Performing hierarchical theme mining on the set of video feature vectors based on the emotional evolution trajectory map to obtain a theme structure tree, and performing semantic correlation reasoning on the theme structure tree to obtain a theme association network, where the theme association network includes a theme importance score, key event nodes, and a theme migration path; Performing multi-angle scene semantic analysis on the theme association network to obtain a set of scene semantic features, and performing cross-scene semantic mapping based on the set of scene semantic features to obtain a scene conversion matrix, where the scene conversion matrix includes a scene continuity vector, a scene switching rule, and a scene semantic similarity; Performing deep semantic fusion on the scene conversion matrix and the theme association network to obtain a semantic association topology map, and performing multi-dimensional feature organization based on the semantic association topology map to obtain a video semantic map, where the video semantic map includes a narrative logic link, a semantic hierarchical structure, and a content main line framework.
[0033] Specifically, the process of performing content semantic analysis on the video feature vector set to obtain a video semantic map is a key step in transforming the pre-extracted multimodal information into a structured and easily understandable narrative framework. First, the system performs an emotional intensity analysis on the video feature vector set to identify various emotional dimensions contained therein and generate a multi-dimensional emotional feature vector. For example, in a corporate promotional video, by analyzing the expressions of employees when presenting new products, the intonation of their voices, and the choice of background music, different emotional intensities such as positive, anticipatory, or tense in the video can be quantified. Then, by calculating the temporal correlation of these multi-dimensional emotional feature vectors, the system can depict the emotional evolution trajectory map of the entire video. This trajectory map not only includes the emotional change gradient, that is, the rate of change of emotional intensity over time, but also the emotional density distribution, showing the frequency and concentration of specific emotions throughout the video, as well as emotional inflection point markers, indicating the critical moments of significant emotional changes. This step helps to reveal the laws of emotional fluctuations in the video content, thus providing a basis for subsequent theme mining. Based on the obtained emotional evolution trajectory map, the system further conducts hierarchical theme mining on the video feature vector set, aiming to construct a theme structure tree that reflects the core idea of the video content. In this process, the system will identify the main topics and their sub-topics in the video, forming a hierarchical theme network. For example, in a corporate promotional video, it may include "teamwork" as the top-level theme, and "R & D process", "market promotion", etc. as secondary themes. Subsequently, through semantic relevance reasoning on the theme structure tree, the system can establish a theme association network that not only includes the importance scores of each theme but also identifies key event nodes and theme migration paths. This enables the system to understand the logical relationships between various themes and their development context in the video. For example, in a corporate promotional video, the narrative way of transitioning from "teamwork" to "product release" can be clearly shown through the theme association network, enhancing the coherence and attractiveness of the story. Next, the system conducts multi-angle scene semantic analysis on the theme association network to extract a set of scene semantic features. This process involves in-depth interpretation of the meanings of various scenes in the video, considering multiple aspects such as visual elements, audio cues, and text descriptions. For example, in a corporate promotional video, for the scene of showing the team working in the laboratory, the system not only has to analyze the actions of the people and the use of equipment in the picture but also consider whether the background sound effects convey a tense or excited mood. Based on these scene semantic features, the system then performs cross-scene semantic mapping to generate a scene transition matrix. This matrix contains information such as scene continuity vectors, scene switching rules, and scene semantic similarity, helping to ensure the natural and smooth transition between different scenes while maintaining the overall narrative consistency. Finally, through deep semantic fusion of the scene transition matrix and the theme association network, the system constructs a semantic association topology map that comprehensively reflects the structure and meaning of the video content.This topological graph not only shows the interconnections between various elements but also provides rich context information. Based on this, the system organizes multi-dimensional features and finally obtains a video semantic map. As a complex network structure, the video semantic map integrates information from multiple aspects such as narrative logic links, semantic hierarchical structures, and content main line frameworks, providing a solid theoretical foundation for the intelligent editing and synthesis of videos. For example, in the application scenario of corporate promotional videos, the video semantic map can not only guide the production team on how to effectively combine different footage segments but also prompt them on how to adjust the narrative rhythm and emotional expression to better convey the company's brand image and value proposition. In this way, the video semantic map becomes a bridge connecting the original footage and high-quality finished videos, greatly improving the efficiency and quality of video content creation.
[0034] In a specific embodiment, optimizing the narrative structure of the video semantic map through deep reinforcement learning technology to obtain an optimized video story line includes: Segmenting the narrative units of the video semantic map to obtain a sequence of basic narrative segments, and calculating the plot tension of the sequence of basic narrative segments to obtain a narrative tension curve; Constructing a narrative state space for the sequence of basic narrative segments based on the narrative tension curve to obtain a narrative state transition network, and performing a reward function mapping on the narrative state transition network to obtain a narrative optimization strategy space, where the narrative optimization strategy space includes state value evaluation, action selection probability, and transfer reward distribution; Performing multiple rounds of exploration sampling on the narrative optimization strategy space to obtain a set of candidate narrative paths, and performing a narrative quality assessment based on the set of candidate narrative paths to obtain an optimal narrative decision sequence, where the optimal narrative decision sequence includes plot integrity data, narrative coherence metrics, and audience attention prediction; Reorganizing the narrative structure of the optimal narrative decision sequence through deep reinforcement learning technology to obtain an optimized narrative framework, and reconstructing the temporal relationship based on the optimized narrative framework to obtain an optimized video story line, where the optimized video story line includes the main story thread, key plot nodes, and narrative rhythm control parameters.
[0035] Specifically, the process of optimizing the narrative structure of the video semantic graph to obtain the optimized video story line is a crucial step in ensuring that the finally generated video content is attractive and logically coherent. First, the system segments the video semantic graph into narrative units, decomposing complex video content into a series of basic narrative fragment sequences, where each fragment represents an independent but interrelated story element. For example, in a corporate promotional video, different stages such as teamwork, R & D process, and product release can be used as separate basic narrative fragments. Next, the system calculates the plot tension for these basic narrative fragment sequences, constructing a narrative tension curve by analyzing factors such as the emotional intensity and information density of each fragment. This curve shows the change in plot tension of the entire video content from start to finish, helping to identify which parts need to be strengthened to attract the audience's attention and which parts may need to be simplified or deleted. Based on the obtained narrative tension curve, the system further constructs a narrative state space for the basic narrative fragment sequence, forming a narrative state transition network. This network not only depicts the potential connection methods between individual narrative fragments but also provides information on how to smoothly transition from one fragment to the next. To evaluate the effectiveness of different state transitions, the system maps a reward function to the narrative state transition network, thereby obtaining a narrative optimization strategy space. Within this space, the system can evaluate the value of each state, select the optimal action (i.e., the fragment order), and predict the transition reward distribution to ensure that the final story line is both logical and interesting to the audience. For example, in a corporate promotional video, if it is desired to highlight the importance of teamwork, the system may preferentially select those fragments that show employees working together to overcome difficulties and adjust the order and duration of these fragments according to the possible reactions of the audience. Subsequently, the system conducts multiple rounds of exploration sampling on the narrative optimization strategy space to generate a set of candidate narrative paths. Each path is a potential way of storytelling, and by simulating different narrative orders and combinations, the system can explore various possibilities. Based on this, the system evaluates the narrative quality of the set of candidate narrative paths, comprehensively considering factors such as plot integrity data, narrative coherence metrics, and audience attention prediction, and selects the optimal narrative decision sequence from them. For example, in the application scenario of a corporate promotional video, the system may find that a narrative method that particularly emphasizes the spirit of innovation and teamwork can arouse more emotional resonance among the audience than other methods, so this path is selected as the final solution. Finally, through deep reinforcement learning techniques, the system reorganizes the narrative structure of the optimal narrative decision sequence, constructs an optimized narrative framework, and on this basis, reconstructs the temporal relationship to finally obtain the optimized video story line. This process not only includes rearranging individual narrative fragments to form a clear main story line but also includes determining the positions of key plot nodes and setting narrative rhythm control parameters.For example, in a corporate promotional video, the optimized video storyline may start with an inspiring opening, followed by scenes of the team working busily in the laboratory, and then gradually transition to the successful release of the product. Throughout the process, employee interviews and customer feedback are interspersed to enhance the emotional level and narrative depth. By such an approach, the system can not only ensure the logical coherence of the video content and the richness of emotional expression, but also dynamically adjust the narrative rhythm according to the characteristics of the target audience, making the finally generated video more in line with the psychological expectations of the audience, thereby achieving a more effective information transmission and brand promotion purpose. This series of complex processing procedures, from narrative unit segmentation to the final narrative structure reorganization, together constitute an intelligent and efficient video content creation solution, greatly improving the quality and efficiency of video production.
[0036] In a specific embodiment, constructing a narrative state space for the basic narrative segment sequence based on the narrative tension curve to obtain a narrative state transition network includes: Performing peak feature extraction on the narrative tension curve to obtain a sequence of key tension nodes, and performing temporal dependence analysis on the sequence of key tension nodes to obtain a node association map; Performing state feature encoding on the basic narrative segment sequence based on the node association map to obtain a set of narrative state vectors, and calculating the semantic distance of the set of narrative state vectors to obtain a state transition matrix; Performing action space division on the state transition matrix to obtain a narrative decision space, and constructing state transition rules based on the narrative decision space to obtain a set of state transition rules; Performing network topology mapping on the set of state transition rules to obtain an initial transition network, and performing structural optimization organization based on the initial transition network to obtain a narrative state transition network, where the narrative state transition network includes node connection metrics, path reachability indicators, and network stability parameters.
[0037] Specifically, the process of constructing a narrative state space for the basic narrative segment sequence based on the narrative tension curve to obtain a narrative state transition network is an important step to ensure the logical coherence and attractiveness of video content narration. First, the system extracts peak features from the narrative tension curve to identify key tension nodes that can significantly affect the emotional changes of the audience. These key tension nodes represent the moments when the emotional or information density in the video reaches a peak, such as the product release moment in a corporate promotional video or the moment when the team overcomes a major challenge. By performing a temporal dependence analysis on these key tension nodes, the system can understand the mutual relationships and order among different nodes, thereby constructing a node association map. This map not only shows the positions of each key node but also reveals the temporal and logical connections between them, providing a basis for subsequent state feature encoding. Next, based on the obtained node association map, the system encodes the state features of the basic narrative segment sequence to generate a set of narrative state vectors. This process aims to transform each basic narrative segment into a quantifiable form for easy computer processing and analysis. For example, in the application scenario of a corporate promotional video, for a segment showing the team working together to solve a problem, the system encodes it based on various information such as its visual elements, audio cues, and text descriptions to form a unique state vector. Then, the system calculates the semantic distances of these narrative state vector sets to evaluate the similarities and differences between different states, and further constructs a state transition matrix. The state transition matrix not only reflects the transition probabilities between various narrative states but also contains information about the costs or benefits required to transfer from one state to another. This step helps the system understand how to optimize the narrative structure through reasonable state transitions and enhance the viewing experience of the audience. Subsequently, the system divides the state transition matrix into an action space to obtain a narrative decision space. This space defines all possible action (i.e., state transition) options, providing the system with multiple potential narrative path choices. Based on this, the system further constructs a set of state transition rules, which not only include which states can be directly connected but also consider the quality of the connection (such as fluency, logical consistency, etc.). For example, in a corporate promotional video, the system may set a rule that the "team cooperation" state can only transition to the "project progress" or "result display" state, rather than directly jumping to an irrelevant part. To achieve this goal, the system needs to perform a network topology mapping on the set of state transition rules to generate an initial transition network. Although this initial network already preliminarily shows the possible connection ways between states, it still needs to be further optimized to improve its efficiency and stability. Finally, the system optimizes and organizes the structure of the initial transition network to obtain the final narrative state transition network. In this process, the system not only adjusts the connection ways between nodes but also ensures that the entire network has good path reachability and network stability.For example, in a corporate promotional video, the system may enhance the coherence of the story line by adding some transitional segments or rearrange certain segments to reduce the cognitive burden on the audience. The narrative state transition network includes important components such as node connection metrics, path reachability indicators, and network stability parameters. These elements together determine the overall structure and quality of the video content. Through this complex processing flow, the system can effectively optimize the narrative structure of the video, making the final generated content logical and full of emotional resonance, greatly enhancing the attractiveness and influence of the video. In specific applications such as the production of corporate promotional videos, this method not only helps creators better organize materials but also dynamically adjusts the narrative rhythm and focus to meet diverse communication needs.
[0038] In a specific embodiment, encoding the state features of the basic narrative segment sequence based on the node association graph to obtain a set of narrative state vectors, including: Performing temporal feature decomposition on the node association graph to obtain a tension node feature matrix, and performing hierarchical encoding on the tension node feature matrix to obtain a multi-level feature representation; Based on the multi-level feature representation, performing scene element analysis on the basic narrative segment sequence to obtain a set of scene feature combinations, and performing semantic embedding mapping on the set of scene feature combinations to obtain a semantic encoding matrix; Performing narrative unit alignment on the semantic encoding matrix to obtain a narrative structure feature map, and performing multi-dimensional feature fusion based on the narrative structure feature map to obtain a feature fusion tensor; Based on the feature fusion tensor, performing state space projection on the narrative structure feature map to obtain a set of state projection vectors, and performing dimensionality reduction processing on the set of state projection vectors to obtain a state compression representation; Performing vector normalization processing on the state compression representation to obtain an initial state vector, and performing state feature enhancement based on the initial state vector to obtain a set of narrative state vectors.
[0039] Specifically, the process of encoding the state features of the basic narrative segment sequence based on the node association graph to obtain the narrative state vector set is an important link in constructing an efficient narrative structure. First, the system decomposes the temporal features of the node association graph to extract the features of each key tension node, forming a tension node feature matrix. This process aims to transform complex node association information into a processable data form, enabling the system to identify the importance of each node on the time axis and its relationship with other nodes. For example, in a corporate promotional video, the critical moments when the team members overcome technical difficulties and finally successfully launch a product are analyzed and recorded as nodes, forming a detailed feature matrix. Next, the system hierarchically encodes this tension node feature matrix to generate multi-level feature representations. Hierarchical encoding is a method of capturing the deep structure of data through multi-level abstraction, which allows the system to not only focus on the information of individual nodes but also understand the development context of the entire story line. For example, in the application scenario of a corporate promotional video, the system may encode each key node step by step from the micro level (such as a specific technical detail) to the macro level (such as the development process of the entire project) to obtain a comprehensive multi-level feature representation. Based on this multi-level feature representation, the system further analyzes the scene elements of the basic narrative segment sequence, identifies the core elements in each segment, and organizes them into a set of scene feature combinations. For example, in the process of describing team cooperation to solve problems, the system will notice various factors such as the expressions of the characters, the tools used, and the background music, and these elements together form a complete set of scene feature combinations. Subsequently, the system performs semantic embedding mapping on the set of scene feature combinations to generate a semantic encoding matrix. This step converts specific scene features into vector representations in a high-dimensional space, enabling the similarities and differences between different scenes to be quantified mathematically. For example, in a corporate promotional video, different team cooperation scenes may differ due to the specific tasks involved, but through semantic embedding mapping, the system can find the potential common points and differences between them, and thus better understand the content framework of the entire video. Based on the obtained semantic encoding matrix, the system then performs narrative unit alignment to ensure the logical consistency and coherence of each narrative segment, thereby generating a narrative structure feature map. Then, the system uses multi-dimensional feature fusion technology to integrate information from various sources into a unified feature fusion tensor, which contains all the necessary information for subsequent state space projection. After obtaining the feature fusion tensor, the system performs state space projection on it, compressing the high-dimensional information into a more easily processable space to form a set of state projection vectors. To reduce redundant information and improve computational efficiency, the system performs dimensionality reduction processing on the set of state projection vectors to generate a state compression representation. This step is crucial for maintaining the core value of the information while reducing complexity.For example, during the production of a corporate promotional video, by performing dimensionality reduction on a large amount of material, the system can focus on the key parts that best reflect the theme, rather than getting lost in the details. Finally, the system performs vector normalization on the state compression representation to ensure that all vectors are within the same scale range for easy comparison and operation. On this basis, the system further enhances the state features, combines all the information accumulated in the previous steps, and finally generates a set of narrative state vectors. This series of complex processing flows not only improves the depth of understanding of the video content but also provides solid data support for subsequent narrative optimization. Through this method, the corporate promotional video can not only accurately convey the company's core values but also attract the audience's attention in a more engaging way.
[0040] In a specific embodiment, performing multi-modal content generation based on the optimized video storyline to obtain a set of candidate video clips, including: Decompose the scene elements of the optimized video storyline to obtain a scene element configuration table, and perform visual style matching on the scene element configuration table to obtain a set of scene rendering parameters; Reconstruct the visual content of the scene element configuration table based on the set of scene rendering parameters to obtain a visual scene sequence, and perform dynamic special effect synthesis on the visual scene sequence to obtain a visually enhanced sequence; Perform audio emotion mapping on the visually enhanced sequence to obtain a set of emotional audio features, and generate audio content based on the set of emotional audio features to obtain an audio effect sequence; Optimize the audio-visual synchronization of the visually enhanced sequence based on the audio effect sequence to obtain an initial set of synthesized clips, and perform spatio-temporal consistency calibration on the initial set of synthesized clips to obtain a calibrated synthesis sequence; Perform quality assessment and screening on the calibrated synthesis sequence to obtain a quality score matrix, and perform diversified reordering based on the quality score matrix to obtain a set of candidate video clips.
[0041] Specifically, the process of generating multimodal content based on the optimized video storyline to obtain a set of candidate video clips is a key step in transforming a carefully designed narrative structure into a specific, vivid and attractive video content. First, the system decomposes the optimized video storyline into scene elements, extracts the core elements that constitute each scene, and forms a scene element configuration table. For example, in a corporate promotional video, if the storyline emphasizes the importance of teamwork, the scene element configuration table may contain information such as the pictures showing the collaboration of team members, the tools used, and the background environment. Next, the system matches the visual style according to the characteristics of these scene elements, determines the most suitable form of expression and visual effects, and generates a scene rendering parameter set. This step not only determines the visual style of the final video, such as color, light, and composition, but also provides guidance for subsequent visual content reconstruction. For example, for a scene showing teamwork, the system may choose a bright and vibrant visual style to convey a positive atmosphere. Based on the scene rendering parameter set obtained above, the system reconstructs the visual content of the scene element configuration table to generate a series of continuous visual scene sequences. This process uses advanced computer graphics technology to convert abstract element configurations into specific images or animations. For example, in a corporate promotional video, the system can create a realistic laboratory work scene based on the description in the configuration table, including accurate reproduction of experimental equipment and employee operation details. In order to enhance the visual effect, the system will also perform dynamic special effects synthesis on the generated visual scene sequence, adding special effects such as light and shadow transformation and particle effects to enhance the viewing and appeal of the picture. In this way, the originally simple scene is transformed into a more fascinating visual effect enhancement sequence through dynamic special effects synthesis. For example, when showing the release of a new product, you can highlight the highlights of the product and attract the audience's attention by adding flash effects and slow-motion close-ups. Subsequently, the system will perform audio emotion mapping on the visual effect enhancement sequence, analyze the emotional color that each scene needs to convey, and generate an emotional audio feature set based on this. This process takes into account how sound complements visual elements to enhance the overall emotional expression. For example, in a corporate promotional video, when the presentation team successfully overcomes difficulties and launches a new product, the system will select exciting background music and inspiring narration to match the picture, thereby enhancing the audience's emotional resonance. Based on the obtained emotional audio feature set, the system further generates audio content to create an audio effect sequence that meets specific emotional needs. This includes not only the selection of background music, but also multi-level sound design such as dialogue and sound effects, making the entire video equally rich in hearing. Next, the system optimizes the audio and video synchronization of the visual effect enhancement sequence based on the audio effect sequence to ensure the perfect match between the two and generate an initial set of synthetic clips. Audio and video synchronization optimization is a complex process that requires taking into account precise matching on the timeline and consistency of emotional expression.For example, in a corporate promotional video, when the screen switches to a scene where the team celebrates success, the system ensures that the climax of the background music is completely consistent with the cheers and smiling faces in the screen, creating a strong sense of presence. Then, the system calibrates the initial set of synthetic segments for spatiotemporal consistency to ensure the coherence of each segment in the time and space dimensions, avoiding frame skipping or incoordination, and thus generates a calibrated synthetic sequence. This step is critical to ensuring the overall smoothness of the video. Finally, the system performs quality assessment and screening on the calibrated synthetic sequence, and generates a quality scoring matrix through a series of standards (such as clarity, color balance, audio quality, etc.). Based on this matrix, the system can perform diversified reordering of the synthetic segments and select the best quality segments to form a candidate video segment set. For example, in the application scenario of corporate promotional videos, the system may give priority to those segments that best reflect the company culture, product features, and team spirit, while ensuring the logic and emotional coherence of the overall narrative. Through this series of complex processing procedures, the system can not only extract the most expressive content from the original material, but also meet different communication needs in a highly customized way, greatly improving the efficiency of video production and the quality of the finished product.
[0042] In a specific embodiment, the intelligent editing and synthesis of the candidate video clip set to obtain target high-quality video content includes: Performing shot boundary detection on the candidate video clip set to obtain a shot segmentation sequence, and performing time sequence combination optimization on the shot segmentation sequence to obtain a shot arrangement scheme; Based on the lens arrangement scheme, a transition effect is designed for the lens segmentation sequence to obtain a transition special effect sequence, and visual coherence is enhanced for the transition special effect sequence to obtain a visually smooth sequence; Performing audio rhythm matching on the visual smooth sequence to obtain a sound-image fusion sequence, and performing multi-track mixing arrangement based on the sound-image fusion sequence to obtain a stereo sound effect scheme, wherein the stereo sound effect scheme includes channel balance parameters, audio transition curves, and spatial sound field layout; The stereo sound effect scheme and the visual smooth sequence are synchronously integrated in time and space to obtain target high-quality video content.
[0043] Specifically, the process of intelligently editing and synthesizing the candidate video clip set to obtain the target high-quality video content is a crucial step in integrating the previously generated diverse materials into a coherent, high-quality, and attractive final video. First, the system performs shot boundary detection on the candidate video clip set, identifying the starting and ending points of each shot by analyzing the changes and continuity of the video frames, thereby obtaining a shot segmentation sequence. This process ensures that each independent shot can be precisely segmented, providing a basis for subsequent optimization. For example, in a corporate promotional video, the system can identify different shots in the scene of teamwork, such as the transition from a panoramic view of the entire team working to a close-up of a certain member operating the equipment. Next, based on the obtained shot segmentation sequence, the system conducts temporal combination optimization to determine the best shot arrangement plan. This step not only considers the logical order of the story line but also combines factors such as emotional intensity and audience attention to ensure the smooth and natural narration of the entire video. For example, in the application scenario of a corporate promotional video, the system may, according to the narrative needs, place the shots describing the teamwork to solve problems before the product launch to better guide the audience's emotions. Subsequently, based on this shot arrangement plan, the system designs transition effects for the shot segmentation sequence, adding appropriate transition special effects, such as fade-in / fade-out or rotation switching, to form a transition effect sequence. These special effects not only enhance the visual coherence but also improve the pleasure of the viewing experience. To further improve the visual fluency, the system enhances the visual coherence of the transition effect sequence, adjusting the consistency of color, light, and composition to generate a visually smooth sequence. For example, when showing the process of a new product from R & D to market launch, the system can make the transition of the images in each stage more natural and harmonious through fine color correction and light adjustment. Based on the visually smooth sequence, the system performs audio rhythm matching to ensure that the background music, dialogue, and sound effects are synchronized with the video frame changes, generating an audio-visual integration sequence. This process requires precise timeline alignment and meticulous processing of audio elements to achieve the best emotional resonance effect. For example, in a corporate promotional video, when showing the moment when the team successfully launches a new product, the system will select an exciting background music and ensure that the climax part of the music coincides perfectly with the celebration scene in the video. Then, based on the obtained audio-visual integration sequence, the system conducts multi-track mixing arrangement to generate a stereo sound effect plan. In this process, the system not only adjusts the channel balance parameters to make the sound distribution on the left and right channels reasonable but also designs the audio transition curve and a smooth spatial sound field layout to provide an immersive auditory experience. For example, in the scene of describing teamwork, the system can use spatial sound field layout technology to make the audience feel as if they are in a real laboratory environment, hearing the discussions of colleagues around and the operation sounds of equipment. Finally, the system synchronizes and integrates the stereo sound effect plan and the visually smooth sequence in time and space to ensure their perfect cooperation and generate the target high-quality video content.This step involves complex calculations and precise timeline management, with the goal of eliminating any inconsistencies that may affect the viewing experience. For example, in the production of a corporate promotional video, the system carefully checks every shot switching point to ensure that the audio and video elements are seamlessly connected here to avoid any frame skipping or audio and video asynchrony. Through this meticulous processing method, the system can not only ensure the logic of the video content and the richness of emotional expression, but also significantly improve the overall viewing quality and communication effect. Therefore, through the intelligent editing and synthesis of the candidate video clip set, the target high-quality video content finally generated can not only accurately convey the core values and cultural characteristics of the company, but also attract and impress the audience in an attractive way, thereby achieving efficient communication purposes.
[0044] The above describes the method for automatically generating high-quality video content in the embodiment of the present invention. The following describes the device for automatically generating high-quality video content in the embodiment of the present invention. Figure 2 , an embodiment of the high-quality video content automatic generation device in the embodiment of the present invention includes: The extraction module 21 is used to extract multimodal features from the input original video material to obtain a video feature vector set; An analysis module 22, configured to perform content semantic analysis on the video feature vector set to obtain a video semantic graph; An optimization module 23, configured to optimize the narrative structure of the video semantic graph by using deep reinforcement learning technology to obtain an optimized video story line; A generating module 24, configured to generate multimodal content based on the optimized video storyline to obtain a set of candidate video clips; The synthesis module 25 is used to intelligently edit and synthesize the candidate video clip set to obtain target high-quality video content.
[0045] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to the above method embodiment, which will not be repeated here.
[0046] Reference Figure 3 The present invention also provides a computer device in an embodiment, wherein the internal structure of the computer device can be as follows: Figure 3As shown in the figure. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected via a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the above method is implemented.
[0047] Those skilled in the art can understand that Figure 3 the structure shown in the figure is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0048] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0049] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to memory, storage, database, or other media provided by the present invention and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0050] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article or method comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements that are inherent to such process, apparatus, article or method. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, apparatus, article or method comprising such element.
[0051] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for automatically generating high-quality video content, characterized in that: The following steps are involved: Perform multimodal feature extraction on the input original video material to obtain a video feature vector set; Performing content semantic analysis on the video feature vector set to obtain a video semantic graph; Optimizing the narrative structure of the video semantic graph by deep reinforcement learning technology to obtain an optimized video story line; Performing multimodal content generation based on the optimized video storyline to obtain a set of candidate video clips; The candidate video clip set is intelligently edited and synthesized to obtain target high-quality video content.
2. The method for automatically generating high-quality video content according to claim 1, characterized in that: The multimodal feature extraction is performed on the input original video material to obtain a video feature vector set, including: Decomposing the input original video material in terms of time and space dimensions to obtain a video frame sequence set and an audio signal sequence, and performing deep optical flow field calculation on the video frame sequence set to obtain a scene motion feature tensor; Based on the scene motion feature tensor, the video frame sequence set is subjected to hierarchical semantic segmentation to obtain a visual semantic feature map, and the visual semantic feature map is subjected to cross-modal attention mapping to obtain a visual-semantic association feature matrix; The visual-semantic association feature matrix and the audio signal sequence are subjected to multimodal temporal fusion to obtain a multimodal feature vector set, and context dependency analysis is performed based on the multimodal feature vector set to obtain a video feature vector set, wherein the video feature vector set includes a scene semantic relationship graph, temporal dynamic features and a multimodal interaction pattern.
3. The method for automatically generating high-quality video content according to claim 1, characterized in that: The performing content semantic analysis on the video feature vector set to obtain a video semantic graph includes: Performing emotion intensity analysis on the video feature vector set to obtain a multidimensional emotion feature vector, and performing temporal correlation calculation on the multidimensional emotion feature vector to obtain an emotion evolution trajectory diagram, wherein the emotion evolution trajectory diagram includes an emotion change gradient, an emotion density distribution, and an emotion inflection point mark; Based on the emotion evolution trajectory diagram, hierarchical topic mining is performed on the video feature vector set to obtain a topic structure tree, and semantic relevance reasoning is performed on the topic structure tree to obtain a topic association network, wherein the topic association network includes topic importance scores, key event nodes and topic migration paths; Performing multi-angle scene semantic analysis on the subject association network to obtain a scene semantic feature set, and performing cross-scene semantic mapping based on the scene semantic feature set to obtain a scene transition matrix, wherein the scene transition matrix includes a scene continuity vector, a scene switching rule, and a scene semantic similarity; The scene transition matrix and the topic association network are deeply semantically fused to obtain a semantic association topology map, and multi-dimensional feature organization is performed based on the semantic association topology map to obtain a video semantic map, wherein the video semantic map includes narrative logic links, semantic hierarchy and content main line framework.
4. The method for automatically generating high-quality video content according to claim 1, characterized in that: The method of optimizing the narrative structure of the video semantic graph by deep reinforcement learning technology to obtain an optimized video story line includes: Segmenting the video semantic map into narrative units to obtain a basic narrative segment sequence, and calculating the plot tension of the basic narrative segment sequence to obtain a narrative tension curve; Based on the narrative tension curve, a narrative state space is constructed for the basic narrative fragment sequence to obtain a narrative state transfer network, and a reward function is mapped to the narrative state transfer network to obtain a narrative optimization strategy space, wherein the narrative optimization strategy space includes state value evaluation, action selection probability and transfer reward distribution; Performing multiple rounds of exploration sampling on the narrative optimization strategy space to obtain a set of candidate narrative paths, and performing narrative quality assessment based on the set of candidate narrative paths to obtain an optimal narrative decision sequence, wherein the optimal narrative decision sequence includes plot completeness data, narrative coherence measurement, and audience attention prediction; The narrative structure of the optimal narrative decision sequence is reorganized through deep reinforcement learning technology to obtain an optimized narrative framework, and the temporal relationship is reconstructed based on the optimized narrative framework to obtain an optimized video story line, wherein the optimized video story line includes the main story line, key plot nodes and narrative rhythm control parameters.
5. The method for automatically generating high-quality video content according to claim 4, characterized in that: The step of constructing a narrative state space for the basic narrative segment sequence based on the narrative tension curve to obtain a narrative state transition network includes: Extracting peak features of the narrative tension curve to obtain a key tension node sequence, and performing temporal dependency analysis on the key tension node sequence to obtain a node association graph; Based on the node association graph, the basic narrative fragment sequence is encoded with state features to obtain a narrative state vector set, and the narrative state vector set is calculated with semantic distance to obtain a state transition matrix; Performing action space division on the state transition matrix to obtain a narrative decision space, and constructing state transition rules based on the narrative decision space to obtain a state transition rule set; The state transition rule set is mapped to a network topology to obtain an initial transition network, and a structural optimization organization is performed based on the initial transition network to obtain a narrative state transition network, wherein the narrative state transition network includes a node connection metric, a path reachability index, and a network stability parameter.
6. The method for automatically generating high-quality video content according to claim 1, characterized in that: The multimodal content generation based on the optimized video story line to obtain a candidate video segment set includes: Decomposing the optimized video story line into scene elements to obtain a scene element configuration table, and performing visual style matching on the scene element configuration table to obtain a scene rendering parameter set; Reconstructing visual content of the scene element configuration table based on the scene rendering parameter set to obtain a visual scene sequence, and performing dynamic special effects synthesis on the visual scene sequence to obtain a visual effect enhanced sequence; Performing audio emotion mapping on the visual effect enhancement sequence to obtain an emotional audio feature set, and generating audio content based on the emotional audio feature set to obtain an audio effect sequence; Based on the audio effect sequence, the visual effect enhancement sequence is optimized for audio and video synchronization to obtain an initial synthesis segment set, and the initial synthesis segment set is calibrated for time and space consistency to obtain a calibrated synthesis sequence; The calibrated synthetic sequence is subjected to quality assessment screening to obtain a quality score matrix, and diversified reordering is performed based on the quality score matrix to obtain a set of candidate video segments.
7. The method for automatically generating high-quality video content according to claim 1, characterized in that: The intelligently editing and synthesizing the candidate video clip set to obtain target high-quality video content includes: Performing shot boundary detection on the candidate video clip set to obtain a shot segmentation sequence, and performing time sequence combination optimization on the shot segmentation sequence to obtain a shot arrangement scheme; Based on the lens arrangement scheme, a transition effect is designed for the lens segmentation sequence to obtain a transition special effect sequence, and visual coherence is enhanced for the transition special effect sequence to obtain a visually smooth sequence; Performing audio rhythm matching on the visual smooth sequence to obtain a sound-image fusion sequence, and performing multi-track mixing arrangement based on the sound-image fusion sequence to obtain a stereo sound effect scheme, wherein the stereo sound effect scheme includes channel balance parameters, audio transition curves, and spatial sound field layout; The stereo sound effect scheme and the visual smooth sequence are synchronously integrated in time and space to obtain target high-quality video content.
8. A device for automatically generating high-quality video content, characterized in that: include: An extraction module is used to extract multimodal features from the input original video material to obtain a video feature vector set; An analysis module, used for performing content semantic analysis on the video feature vector set to obtain a video semantic graph; An optimization module, used to optimize the narrative structure of the video semantic graph by deep reinforcement learning technology to obtain an optimized video story line; A generation module, configured to generate multimodal content based on the optimized video storyline to obtain a set of candidate video clips; The synthesis module is used to intelligently edit and synthesize the candidate video clip set to obtain target high-quality video content.
9. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Video semantic visualization method
CN102523536A
Intelligent multi-mode animation creation system and creation method
CN116342763A
Video post-editing and video synthesis optimization method
CN116847123A
Method and system for automatically generating story video in meta universe
CN117177003A
Video production method and system based on deep learning
CN117692716A
Cited By
Video material intelligent processing method and device based on dynamic semantic driving
CN120544106A
Video material intelligent processing method and device based on dynamic semantic driving
CN120544106B
Automatic video editing method based on semantic analysis
CN120583283A
Highlight bright spot extraction method and system based on multi-modal model
CN120708124A
Commodity video intelligent generation method based on multi-modal analysis and dynamic narrative architecture
CN120711259A