Method and related equipment for automatically generating high-quality video content

By extracting multimodal feature of video materials and optimizing video storylines, the problem of lack of depth and attractiveness in video generation in the existing technology is solved, and high-quality video content with clear logic and coherence is generated.

CN120050487BActive Publication Date: 2025-08-12SHENZHEN CHUANGYUAN INTERACTIVE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510527860.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-12
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing video generation technology fails to deeply explore potential story clues in video materials, resulting in the lack of depth and appeal of the generated content, and the lack of effective multimodal feature extraction and analysis mechanisms, making it difficult to fully understand the complex information in video materials.

Method used

By extracting the input original video material multimodal feature, generating a video feature vector set, and performing content semantic analysis, optimizing the video story line using deep reinforcement learning technology, and finally intelligent editing and synthesis are performed to generate high-quality video content.

Benefits of technology

It realizes a deep understanding of video materials and the transformation of multi-dimensional information, and generates video content with clear logic, emotional coherence and attractiveness to meet personalized and diverse application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050487B_ABST
    Figure CN120050487B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for automatically generating high-quality video content and related equipment, comprising the following steps: performing multimodal feature extraction on input original video material to obtain a video feature vector set; performing content semantic analysis on the video feature vector set to obtain a video semantic map; optimizing the narrative structure of the video semantic map through deep reinforcement learning technology to obtain an optimized video story line; performing multimodal content generation based on the optimized video story line to obtain a candidate video clip set; and performing intelligent editing and synthesis on the candidate video clip set to obtain target high-quality video content. The method solves the technical problem that many automation attempts are limited to simple image or clip splicing, fail to deeply explore potential story clues in video materials, and result in the generated content lacking depth and appeal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automatic generation of video content, and in particular to a method for automatically generating high-quality video content and related equipment. Background Art

[0002] In today's rapidly developing digital age, demand for video content, as one of the primary forms of information dissemination, is exploding. Whether for social media, online education, or corporate communications, the demand for high-quality, engaging video content has reached unprecedented levels. However, traditional video production methods are not only time-consuming and labor-intensive, but also require specialized skills and knowledge. This prohibits individuals and small businesses that lack the resources to create video content. Therefore, developing a method to automatically generate high-quality video content is crucial to meeting this market demand.

[0003] A major issue is that while existing technologies can achieve a certain degree of automated video generation, they often overlook the narrative structure and semantic coherence of the video content. Many automated attempts are limited to simple image or clip splicing, failing to deeply explore the underlying storylines in the video footage, resulting in a lack of depth and appeal in the generated content. Furthermore, due to the lack of effective multimodal feature extraction and analysis mechanisms, existing methods struggle to fully understand the complex information contained in the video footage, such as visual, auditory, and emotional elements, thus limiting the quality and expressiveness of the final video.

[0004] To address these issues, researchers have begun focusing on optimizing the video content generation process through advanced technologies, such as deep reinforcement learning. This approach requires not only the ability to accurately extract and analyze the multimodal features of video material, but also the ability to transform these features into logical and compelling storylines. At the same time, as user demand for personalized content grows, the use of technology to intelligently edit and synthesize videos to meet the needs of different scenarios has become an important research direction. These challenges are driving the research and development of more efficient and intelligent methods for automatic video content generation. Summary of the Invention

[0005] The main purpose of the present invention is to provide a method and related equipment for automatically generating high-quality video content, which solves the technical problem that many automation attempts are limited to simple image or fragment splicing, fail to deeply explore the potential story clues in the video material, and result in the generated content lacking depth and appeal.

[0006] To achieve the above object, the present invention provides a method for automatically generating high-quality video content, comprising the following steps:

[0007] Perform multimodal feature extraction on the input original video material to obtain a video feature vector set;

[0008] Performing content semantic analysis on the video feature vector set to obtain a video semantic graph;

[0009] Optimizing the narrative structure of the video semantic graph using deep reinforcement learning technology to obtain an optimized video storyline;

[0010] Performing multimodal content generation based on the optimized video storyline to obtain a set of candidate video clips;

[0011] Intelligently edit and synthesize the candidate video clip set to obtain target high-quality video content.

[0012] Furthermore, the multimodal feature extraction is performed on the input original video material to obtain a video feature vector set, including:

[0013] Decomposing the input original video material in terms of time and space to obtain a video frame sequence set and an audio signal sequence, and performing deep optical flow field calculation on the video frame sequence set to obtain a scene motion feature tensor;

[0014] Performing hierarchical semantic segmentation on the video frame sequence set based on the scene motion feature tensor to obtain a visual semantic feature map, and performing cross-modal attention mapping on the visual semantic feature map to obtain a visual-semantic association feature matrix;

[0015] Multimodal temporal fusion is performed on the visual-semantic association feature matrix and the audio signal sequence to obtain a multimodal feature vector set, and context dependency analysis is performed based on the multimodal feature vector set to obtain a video feature vector set, wherein the video feature vector set includes a scene semantic relationship graph, temporal dynamic features and a multimodal interaction pattern.

[0016] Furthermore, performing content semantic analysis on the video feature vector set to obtain a video semantic graph includes:

[0017] Performing emotion intensity analysis on the video feature vector set to obtain a multidimensional emotion feature vector, and performing temporal correlation calculation on the multidimensional emotion feature vector to obtain an emotion evolution trajectory diagram, wherein the emotion evolution trajectory diagram includes an emotion change gradient, an emotion density distribution, and an emotion inflection point mark;

[0018] Based on the emotion evolution trajectory graph, hierarchical topic mining is performed on the video feature vector set to obtain a topic structure tree, and semantic relevance reasoning is performed on the topic structure tree to obtain a topic association network, wherein the topic association network includes topic importance scores, key event nodes, and topic migration paths;

[0019] Performing multi-angle scene semantic analysis on the topic association network to obtain a scene semantic feature set, and performing cross-scene semantic mapping based on the scene semantic feature set to obtain a scene transition matrix, wherein the scene transition matrix includes a scene continuity vector, a scene switching rule, and a scene semantic similarity;

[0020] The scene conversion matrix and the topic association network are deeply semantically fused to obtain a semantic association topology map, and multi-dimensional feature organization is performed based on the semantic association topology map to obtain a video semantic map, wherein the video semantic map includes narrative logical links, semantic hierarchy and content main line framework.

[0021] Furthermore, the narrative structure of the video semantic graph is optimized by deep reinforcement learning technology to obtain an optimized video story line, including:

[0022] Segmenting the video semantic map into narrative units to obtain a basic narrative segment sequence, and calculating plot tension on the basic narrative segment sequence to obtain a narrative tension curve;

[0023] Based on the narrative tension curve, a narrative state space is constructed for the basic narrative segment sequence to obtain a narrative state transition network, and a reward function is mapped to the narrative state transition network to obtain a narrative optimization strategy space, wherein the narrative optimization strategy space includes state value evaluation, action selection probability, and transfer reward distribution;

[0024] Performing multiple rounds of exploration sampling on the narrative optimization strategy space to obtain a set of candidate narrative paths, and performing narrative quality assessment based on the set of candidate narrative paths to obtain an optimal narrative decision sequence, wherein the optimal narrative decision sequence includes plot completeness data, narrative coherence measurement, and audience attention prediction;

[0025] The optimal narrative decision sequence is reorganized into a narrative structure through deep reinforcement learning technology to obtain an optimized narrative framework, and the temporal relationship is reconstructed based on the optimized narrative framework to obtain an optimized video story line, wherein the optimized video story line includes the main story line, key plot nodes and narrative rhythm control parameters.

[0026] Furthermore, constructing a narrative state space for the basic narrative segment sequence based on the narrative tension curve to obtain a narrative state transition network includes:

[0027] Extracting peak features from the narrative tension curve to obtain a key tension node sequence, and performing temporal dependency analysis on the key tension node sequence to obtain a node association graph;

[0028] Performing state feature encoding on the basic narrative segment sequence based on the node association graph to obtain a narrative state vector set, and performing semantic distance calculation on the narrative state vector set to obtain a state transition matrix;

[0029] Performing action space division on the state transition matrix to obtain a narrative decision space, and constructing state transition rules based on the narrative decision space to obtain a state transition rule set;

[0030] The state transition rule set is mapped to a network topology to obtain an initial transition network, and a structural optimization organization is performed based on the initial transition network to obtain a narrative state transition network, wherein the narrative state transition network includes a node connection metric, a path reachability index, and a network stability parameter.

[0031] Furthermore, the multimodal content generation is performed based on the optimized video storyline to obtain a set of candidate video clips, including:

[0032] Decomposing the optimized video storyline into scene elements to obtain a scene element configuration table, and performing visual style matching on the scene element configuration table to obtain a scene rendering parameter set;

[0033] Reconstructing visual content of the scene element configuration table based on the scene rendering parameter set to obtain a visual scene sequence, and performing dynamic special effects synthesis on the visual scene sequence to obtain a visual effect enhanced sequence;

[0034] Performing audio emotion mapping on the visual effect enhancement sequence to obtain an emotional audio feature set, and generating audio content based on the emotional audio feature set to obtain an audio effect sequence;

[0035] Performing audio and video synchronization optimization on the visual effect enhancement sequence based on the audio effect sequence to obtain an initial synthesized segment set, and performing spatiotemporal consistency calibration on the initial synthesized segment set to obtain a calibrated synthesized sequence;

[0036] The calibrated synthesized sequence is subjected to quality assessment screening to obtain a quality score matrix, and diversified reordering is performed based on the quality score matrix to obtain a set of candidate video segments.

[0037] Furthermore, the intelligent editing and synthesis of the candidate video clip set to obtain target high-quality video content includes:

[0038] Performing shot boundary detection on the candidate video clip set to obtain a shot segmentation sequence, and performing time sequence combination optimization on the shot segmentation sequence to obtain a shot arrangement scheme;

[0039] Designing a transition effect for the shot segmentation sequence based on the shot arrangement scheme to obtain a transition special effect sequence, and enhancing the visual coherence of the transition special effect sequence to obtain a visually smooth sequence;

[0040] Performing audio rhythm matching on the visually smooth sequence to obtain a sound-image fusion sequence, and performing multi-track mixing arrangement based on the sound-image fusion sequence to obtain a stereo sound effect scheme, wherein the stereo sound effect scheme includes channel balance parameters, audio transition curves, and spatial sound field layout;

[0041] The stereo sound effect solution and the visual smooth sequence are temporally and spatially synchronized to obtain target high-quality video content.

[0042] The present invention also provides a device for automatically generating high-quality video content, comprising:

[0043] The extraction module is used to extract multimodal features from the input original video material to obtain a video feature vector set;

[0044] An analysis module, configured to perform content semantic analysis on the video feature vector set to obtain a video semantic graph;

[0045] An optimization module, configured to optimize the narrative structure of the video semantic graph using deep reinforcement learning technology to obtain an optimized video storyline;

[0046] A generation module, configured to generate multimodal content based on the optimized video storyline to obtain a set of candidate video clips;

[0047] The synthesis module is used to intelligently edit and synthesize the candidate video clip set to obtain target high-quality video content.

[0048] The present invention also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.

[0049] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.

[0050] The present invention provides a method for automatically generating high-quality video content, comprising the following steps: performing multimodal feature extraction on the input original video material to obtain a video feature vector set; performing content semantic analysis on the video feature vector set to obtain a video semantic map; optimizing the narrative structure of the video semantic map through deep reinforcement learning technology to obtain an optimized video story line; performing multimodal content generation based on the optimized video story line to obtain a set of candidate video clips; and performing intelligent editing and synthesis on the set of candidate video clips to obtain target high-quality video content. This method solves the technical problem that many automation attempts are limited to simple image or clip splicing and fail to deeply explore the potential story clues in the video material, resulting in the generated content lacking depth and appeal. It implements content semantic analysis based on the obtained video feature vector set, which can deeply understand the themes, emotions and other implicit information contained in the video material and convert them into a structured video semantic map. This step not only improves the depth of understanding of the video material, but also provides a clear logical framework for the optimization of the narrative structure. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is a schematic diagram of the steps of a method for automatically generating high-quality video content according to an embodiment of the present invention;

[0052] Figure 2 This is a structural block diagram of an apparatus for automatically generating high-quality video content according to an embodiment of the present invention;

[0053] Figure 3 It is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.

[0054] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0056] like Figure 1 As shown, Figure 1 This is a schematic diagram of the steps of a method for automatically generating high-quality video content in one embodiment of the present invention;

[0057] An embodiment of the present invention provides a method for automatically generating high-quality video content, comprising the following steps:

[0058] Step S1: Perform multimodal feature extraction on the input original video material to obtain a video feature vector set.

[0059] Specifically, multimodal feature extraction is performed on the input raw video material to generate a set of video feature vectors. This process is the foundation of the entire high-quality video content automatic generation method. Its core is to extract multidimensional information from the video material that comprehensively reflects its content characteristics and convert it into a structured set of video feature vectors. Specifically, raw video material typically contains multimodal data such as visual, auditory, and possibly textual information. Therefore, the extraction process requires the design of corresponding feature extraction algorithms for each of these modalities. For example, in the visual modality, convolutional neural networks (CNNs) can be used to extract features such as color distribution, object categories, and gestures in the image. In the auditory modality, short-time Fourier transforms or Mel-frequency cepstral coefficients (MFCCs) can be used to extract information such as intonation, rhythm, and background music from the audio. If the video includes subtitles or speech-to-text content, natural language processing techniques can also be used to extract keywords, sentiment, and semantic relationships from the text. After normalization, these multimodal features are integrated into a unified set of video feature vectors, laying the foundation for subsequent content semantic analysis. For example, in the production of a corporate promotional video, the original video footage might include scenes of company employees at work, product demonstration clips, and background audio commentary. By performing multimodal feature extraction on these clips, the system can identify employee expressions and movements, product details, and key information from the commentary. Ultimately, it forms a vector collection containing multimodal features. This not only preserves the core information of the original footage but also provides rich data support for the subsequent generation of high-quality video content.

[0060] Step S2: performing content semantic analysis on the video feature vector set to obtain a video semantic graph.

[0061] Specifically, the video feature vector set undergoes content semantic analysis to generate a video semantic graph. This process aims to gain a deep understanding of the inherent meaning and structure of the video material, thereby providing a logical framework for subsequent steps. First, based on the previously obtained video feature vector set, the system uses technologies such as natural language processing and computer vision to parse the semantic information underlying this data. For example, for a video feature vector set extracted from a corporate promotional video, the system identifies and analyzes multimodal information such as character actions, product details, and the emotional tone of the background music, further exploring potential connections between them. Next, by constructing an association model, the system organizes these scattered information points into a network that comprehensively reflects the essence of the video content and its interrelationships—the video semantic graph. In this graph, each node represents a specific content element, such as a key figure in a scene or a product's core selling point, while edges represent the relationships between these elements, such as causal relationships, temporal order, or emotional coherence. For example, if a corporate promotional video showcases the collaboration of company employees on an innovative project, the video semantic graph will not only identify the individuals involved, the tools and technologies used, but also reveal how these elements contribute to the project's progress. This allows the production team to clearly understand the story's development and provides strong support for optimizing the narrative structure. This video semantic graph, formed through content semantic analysis of the video's feature vector set, allows creators to more intuitively grasp the core value and underlying storylines of the video material, laying a solid foundation for creating more engaging video content.

[0062] Step S3: Optimize the narrative structure of the video semantic graph through deep reinforcement learning technology to obtain an optimized video story line.

[0063] Specifically, deep reinforcement learning technology is used to optimize the narrative structure of the video semantic graph, resulting in an optimized video storyline. This process utilizes advanced machine learning algorithms to enhance the narrative logic and audience appeal of the video content. First, based on the video semantic graph generated in the previous step, the system identifies individual content elements and their relationships, providing foundational data for narrative structure optimization. Next, a deep reinforcement learning model is employed to automatically adjust the storyline for optimal performance by simulating the effects of different narrative approaches. Specifically, the model continuously explores the optimal narrative path during training, based on pre-set objectives (such as maximizing audience engagement or emotional resonance). For example, in the application scenario of a corporate promotional video, if the video aims to showcase the process from conception to implementation of an innovative project, deep reinforcement learning technology can analyze which key nodes (such as teamwork and technological innovation) should be emphasized, and how these nodes should be connected to better tell the overall story. By repeatedly experimenting with different combinations and sequences, and relying on feedback mechanisms (such as audience response simulations or historical data), the system gradually optimizes a logical and engaging storyline. This optimized video storyline not only ensures effective information delivery but also enhances the emotional coherence and appeal of the viewing experience. The resulting video not only accurately conveys the company's core values but also strikes a chord with viewers, stimulating their emotional resonance. Therefore, optimizing the narrative structure of video semantic graphs using deep reinforcement learning technology has become an essential step in creating high-quality video content.

[0064] Step S4: generating multimodal content based on the optimized video storyline to obtain a set of candidate video clips.

[0065] Specifically, multimodal content generation is performed based on the optimized video storyline to generate a set of candidate video clips. This process transforms the narrative logic formed in the previous step into specific content representations, thereby providing a diverse selection of material for the final video production. First, the system identifies the content and emotion to be expressed at each key node based on the optimized video storyline. Combining this with the previously extracted multimodal feature vectors, the system then selects matching visuals, audio, and text from the original video footage. For example, in a corporate promotional video, if the storyline emphasizes the importance of teamwork, the system will prioritize clips showcasing employee collaboration, complementing them with uplifting background music or commentary, while also ensuring that the colors and composition of the footage convey a harmonious and nuanced atmosphere. Next, using multimodal content generation technology, the system not only recombines the existing footage but also utilizes advanced techniques such as generative adversarial networks (GANs) to supplement missing content, such as generating transition shots from specific perspectives or enhancing the expressiveness of specific details. This allows the system to generate multiple candidate video clip sets that meet narrative requirements. These clips retain the core message of the original footage while offering innovative visual and auditory expression. For example, for a corporate promotional video, the candidate clip set might include editing styles of varying styles, some focusing on showcasing the team's professionalism, others on emotional resonance. This provides a rich selection of options for intelligent editing and synthesis, while also ensuring that the resulting video content meets diverse communication needs.

[0066] Step S5: Intelligently edit and synthesize the candidate video clip set to obtain target high-quality video content.

[0067] Specifically, the candidate video clips are intelligently edited and synthesized to produce the target high-quality video content. This process integrates the previously generated diverse clips into a coherent, high-quality final video. First, based on the optimized video storyline, the system analyzes the role of each candidate video clip in the narrative structure and selects the most appropriate clips for combination based on their visual, auditory, and emotional expressiveness. For example, in the application scenario of a corporate promotional video, if the storyline emphasizes the combination of teamwork and innovative achievements, the system will select clips that both showcase employee collaboration details and highlight product highlights. Using intelligent algorithms, the system adjusts their order and duration to ensure a smooth and logical overall flow. Next, the system uses advanced video editing techniques to seamlessly connect the selected clips, such as through timeline alignment, color correction, and audio mixing, to eliminate the sense of disconnection between different clips and enhance the consistency of visuals and sound. Furthermore, intelligent editing dynamically adjusts the content based on potential audience needs, such as adding appropriate transition effects or subtitles to better convey information and capture attention. For example, in a corporate promotional video, the system might insert a slow-motion close-up after a segment showcasing teamwork to highlight a key product detail. This design not only enhances the video's viewing experience but also strengthens the communication of the core message. Ultimately, the resulting high-quality video content, intelligently edited and synthesized, retains the core value of the original footage while achieving a high degree of professionalism and appeal through multimodal optimization, thus meeting corporate promotional needs and achieving the desired communication effect.

[0068] In a specific embodiment, the multimodal feature extraction is performed on the input original video material to obtain a video feature vector set, including:

[0069] Decomposing the input original video material in terms of time and space to obtain a video frame sequence set and an audio signal sequence, and performing deep optical flow field calculation on the video frame sequence set to obtain a scene motion feature tensor;

[0070] Performing hierarchical semantic segmentation on the video frame sequence set based on the scene motion feature tensor to obtain a visual semantic feature map, and performing cross-modal attention mapping on the visual semantic feature map to obtain a visual-semantic association feature matrix;

[0071] Multimodal temporal fusion is performed on the visual-semantic association feature matrix and the audio signal sequence to obtain a multimodal feature vector set, and context dependency analysis is performed based on the multimodal feature vector set to obtain a video feature vector set, wherein the video feature vector set includes a scene semantic relationship graph, temporal dynamic features and a multimodal interaction pattern.

[0072] Specifically, multimodal feature extraction is performed on the input raw video material to obtain a set of video feature vectors. This process is one of the key steps in automatically generating high-quality video content. First, the system performs spatiotemporal decomposition on the input raw video material, splitting the video into a series of continuous video frame sequences and audio signal sequences. The video frame sequence contains information about the temporal evolution of the scene, while the audio signal sequence provides auditory cues such as background music, dialogue, and ambient sounds. For example, let's say a corporate promotional video depicts the collaborative process of a company team developing a new product. Through spatiotemporal decomposition, the system can identify the images and accompanying sounds corresponding to each key moment, laying the foundation for subsequent processing. Next, the system performs deep optical flow calculation on the information extracted from the video frame sequence to obtain a scene motion feature tensor. Deep optical flow calculation is a technique that captures the movement of pixels between adjacent frames. It can effectively represent dynamic changes in the scene, such as the movement of people or objects. In the application scenario of corporate promotional videos, if the video shows team members busy working in a laboratory, deep optical flow computation can help identify the details of these actions, such as employee gestures and the operation of laboratory equipment, thereby forming a tensor representation that describes the dynamic changes of the entire scene. Based on the scene motion feature tensor obtained above, the system performs hierarchical semantic segmentation on the video frame sequence to obtain a visual semantic feature map. Hierarchical semantic segmentation involves dividing video frames into different semantic categories (such as people, objects, and background) and constructing a map depicting the relationships between these elements. Next, cross-modal attention mapping is performed on the visual semantic feature map to obtain a visual-semantic association feature matrix. This step aims to establish the connection between visual elements and semantic information, enabling the system to not only recognize the content in the picture but also understand the underlying meaning. For example, in a corporate promotional video, the system can use hierarchical semantic segmentation to distinguish between team members in the foreground and laboratory equipment in the background. Cross-modal attention mapping can also reveal the interactive relationship between the two, such as which employee is operating a specific piece of equipment. Subsequently, the system performs multimodal temporal fusion on the obtained visual-semantic association feature matrix and audio signal sequence to generate a multimodal feature vector set. This step combines the visual information and audio information of the video to achieve effective integration of the two modes of data, ensuring that the final generated content contains both rich visual details and vivid sound effects. For example, in a corporate promotional video, the system may combine images showing teamwork with active discussion sounds in the background to enhance the audience's emotional resonance. Finally, context dependency analysis is performed based on the multimodal feature vector set to obtain a video feature vector set. Context dependency analysis helps the system understand the role of each segment in the entire storyline and their interactions, including scene semantic relationship graphs, temporal dynamic features, and multimodal interaction patterns.This means the system must not only identify the content of each individual segment but also understand how they collectively form a coherent story. For example, in the context of a corporate promotional video, the system can identify the logical connection between teamwork scenes and product demonstrations and adjust their presentation accordingly, making the overall narrative more fluid and natural. In summary, by subjecting the raw video input to a series of complex processing methods, including spatiotemporal decomposition, deep optical flow calculation, hierarchical semantic segmentation, cross-modal attention mapping, multimodal temporal fusion, and contextual dependency analysis, the system is able to extract a comprehensive and structured set of video feature vectors from the raw footage. These feature vectors not only contain essential information about the video content but also reflect the deep semantic relationships and emotional undertones inherent therein, providing solid data support and technical assurance for subsequent content semantic analysis, narrative structure optimization, and ultimately, the generation of high-quality video content. In the specific application scenario of corporate promotional videos, this sophisticated data processing approach helps accurately capture and showcase a company's core values and cultural characteristics, effectively improving the quality and impact of the video content.

[0073] In a specific embodiment, performing content semantic analysis on the video feature vector set to obtain a video semantic graph includes:

[0074] Performing emotion intensity analysis on the video feature vector set to obtain a multidimensional emotion feature vector, and performing temporal correlation calculation on the multidimensional emotion feature vector to obtain an emotion evolution trajectory diagram, wherein the emotion evolution trajectory diagram includes an emotion change gradient, an emotion density distribution, and an emotion inflection point mark;

[0075] Based on the emotion evolution trajectory graph, hierarchical topic mining is performed on the video feature vector set to obtain a topic structure tree, and semantic relevance reasoning is performed on the topic structure tree to obtain a topic association network, wherein the topic association network includes topic importance scores, key event nodes, and topic migration paths;

[0076] Performing multi-angle scene semantic analysis on the topic association network to obtain a scene semantic feature set, and performing cross-scene semantic mapping based on the scene semantic feature set to obtain a scene transition matrix, wherein the scene transition matrix includes a scene continuity vector, a scene switching rule, and a scene semantic similarity;

[0077] The scene conversion matrix and the topic association network are deeply semantically fused to obtain a semantic association topology map, and multi-dimensional feature organization is performed based on the semantic association topology map to obtain a video semantic map, wherein the video semantic map includes narrative logical links, semantic hierarchy and content main line framework.

[0078] Specifically, performing content semantic analysis on the video feature vector set to generate a video semantic graph is a key step in transforming the multimodal information extracted earlier into a structured and easily understandable narrative framework. First, the system performs sentiment intensity analysis on the video feature vector set to identify the multiple emotional dimensions contained therein and generate a multidimensional sentiment feature vector. For example, in a corporate promotional video, by analyzing the facial expressions, voice intonation, and background music selection of employees presenting new products, the intensity of different emotions in the video, such as positivity, anticipation, or tension, can be quantified. Next, by calculating the temporal correlation of these multidimensional sentiment feature vectors, the system can depict a trajectory of sentiment evolution throughout the entire video. This trajectory graph not only includes the sentiment gradient—the rate at which sentiment intensity changes over time—but also includes the sentiment density distribution, which shows the frequency and concentration of specific emotions throughout the video, as well as sentiment inflection points, which indicate key moments of significant sentiment change. This step helps reveal patterns in the emotional fluctuations within the video content, thus providing a basis for subsequent topic mining. Based on this resulting sentiment evolution trajectory graph, the system further performs hierarchical topic mining on the video feature vector set, aiming to construct a topic structure that reflects the core ideas of the video content. During this process, the system identifies the main topics and their subtopics within the video, forming a clearly structured topic network. For example, in a corporate promotional video, "Teamwork" might be the top-level topic, with "R&D Process" and "Marketing" as sub-topics. Subsequently, by performing semantic correlation reasoning on the topic structure, the system constructs a topic association network. This network not only includes the importance score of each topic but also identifies key event nodes and topic transition paths. This enables the system to understand the logical relationships between topics and their development within the video. For example, in a corporate promotional video, the narrative transition from "Teamwork" to "Product Launch" can be clearly presented through the topic association network, enhancing the story's coherence and appeal. Next, the system performs multi-angle scene semantic analysis on the topic association network to extract a set of scene semantic features. This process involves in-depth interpretation of the meaning of each scene in the video, taking into account multiple aspects, including visual elements, audio cues, and textual descriptions. For example, in a corporate promotional video showing a team working in a laboratory, the system not only analyzes the characters' movements and equipment usage, but also considers whether the background sound effects convey emotions of tension or excitement. Based on these scene semantic features, the system then performs cross-scene semantic mapping to generate a scene transition matrix. This matrix contains information such as scene continuity vectors, scene switching rules, and scene semantic similarity, helping to ensure natural and smooth transitions between different scenes while maintaining overall narrative consistency. Ultimately, through deep semantic fusion of the scene transition matrix and the topic association network, the system constructs a semantic association topology that comprehensively reflects the structure and meaning of the video content.This topological map not only shows the interconnections between various elements, but also provides rich contextual information. Based on this, the system organizes multi-dimensional features and finally obtains a video semantic map. As a complex network structure, the video semantic map integrates multiple aspects of information such as narrative logical links, semantic hierarchy, and content main line framework, providing a solid theoretical basis for intelligent video editing and synthesis. For example, in the application scenario of corporate promotional videos, the video semantic map can not only guide the production team on how to effectively combine different material clips, but also prompt them on how to adjust the narrative rhythm and emotional expression to better convey the company's brand image and value proposition. In this way, the video semantic map becomes a bridge between original materials and high-quality finished videos, greatly improving the efficiency and quality of video content creation.

[0079] In a specific embodiment, the narrative structure optimization of the video semantic graph using deep reinforcement learning technology to obtain an optimized video story line includes:

[0080] Segmenting the video semantic map into narrative units to obtain a basic narrative segment sequence, and calculating plot tension on the basic narrative segment sequence to obtain a narrative tension curve;

[0081] Based on the narrative tension curve, a narrative state space is constructed for the basic narrative segment sequence to obtain a narrative state transition network, and a reward function is mapped to the narrative state transition network to obtain a narrative optimization strategy space, wherein the narrative optimization strategy space includes state value evaluation, action selection probability, and transfer reward distribution;

[0082] Performing multiple rounds of exploration sampling on the narrative optimization strategy space to obtain a set of candidate narrative paths, and performing narrative quality assessment based on the set of candidate narrative paths to obtain an optimal narrative decision sequence, wherein the optimal narrative decision sequence includes plot completeness data, narrative coherence measurement, and audience attention prediction;

[0083] The optimal narrative decision sequence is reorganized into a narrative structure through deep reinforcement learning technology to obtain an optimized narrative framework, and the temporal relationship is reconstructed based on the optimized narrative framework to obtain an optimized video story line, wherein the optimized video story line includes the main story line, key plot nodes and narrative rhythm control parameters.

[0084] Specifically, optimizing the video semantic graph for narrative structure and obtaining an optimized video storyline is a key step in ensuring the final video content is engaging and logically coherent. First, the system segments the video semantic graph into narrative units, breaking down complex video content into a series of basic narrative segments, each representing an independent yet interconnected story element. For example, in a corporate promotional video, different stages such as teamwork, R&D, and product launch could be represented as separate basic narrative segments. Next, the system calculates plot tension for these basic narrative segment sequences, constructing a narrative tension curve by analyzing factors such as the emotional intensity and information density of each segment. This curve depicts the change in plot tension from the beginning to the end of the video content, helping to identify which parts need to be strengthened to attract the viewer's attention and which parts may need to be simplified or deleted. Based on the resulting narrative tension curve, the system further constructs a narrative state space for the basic narrative segment sequence, forming a narrative state transition network. This network not only depicts the potential connections between narrative segments but also provides information on how to smoothly transition from one segment to the next. To evaluate the effectiveness of different state transitions, the system maps a reward function to the narrative state transition network, resulting in a narrative optimization strategy space. Within this space, the system evaluates the value of each state, selects the optimal action (i.e., the sequence of segments), and predicts the distribution of transition rewards, ensuring that the final storyline is both logical and engaging for the audience. For example, in a corporate promotional video, if the goal is to highlight the importance of teamwork, the system might prioritize segments that show employees working together to overcome difficulties, adjusting the order and duration of these segments based on the audience's likely reactions. The system then performs multiple rounds of exploration sampling in the narrative optimization strategy space to generate a set of candidate narrative paths. Each path represents a potential storytelling approach, and by simulating different narrative sequences and combinations, the system explores multiple possibilities. Based on this, the system evaluates the narrative quality of the candidate narrative paths, comprehensively considering factors such as plot completeness data, narrative coherence measures, and audience attention predictions, to select the optimal narrative decision sequence. For example, in a corporate promotional video application, the system might discover that a narrative approach that emphasizes innovation and teamwork resonates more emotionally than other approaches and select this path as the final solution. Finally, deep reinforcement learning technology is used to restructure the optimal narrative decision sequence, systematically constructing an optimized narrative framework. Based on this framework, the temporal relationships are reconstructed, ultimately resulting in an optimized video storyline. This process not only involves rearranging the narrative segments to form a clear storyline, but also involves determining the location of key plot points and setting parameters for narrative rhythm control.For example, in a corporate promotional video, the optimized video storyline may begin with an inspiring opening, followed by a scene showing the team busy working in the laboratory, and then gradually transition to the successful launch of the product. The whole process is interspersed with employee interviews and customer feedback to enhance the emotional level and narrative depth. Through this method, the system can not only ensure the logical coherence and richness of emotional expression of the video content, but also dynamically adjust the narrative rhythm according to the characteristics of the target audience, so that the final generated video is more in line with the audience's psychological expectations, thereby achieving more effective information transmission and brand promotion purposes. This series of complex processing flows, from the segmentation of narrative units to the final reorganization of the narrative structure, together constitute an intelligent and efficient video content creation solution, which greatly improves the quality and efficiency of video production.

[0085] In a specific embodiment, constructing a narrative state space for the basic narrative segment sequence based on the narrative tension curve to obtain a narrative state transition network includes:

[0086] Extracting peak features from the narrative tension curve to obtain a key tension node sequence, and performing temporal dependency analysis on the key tension node sequence to obtain a node association graph;

[0087] Performing state feature encoding on the basic narrative segment sequence based on the node association graph to obtain a narrative state vector set, and performing semantic distance calculation on the narrative state vector set to obtain a state transition matrix;

[0088] Performing action space division on the state transition matrix to obtain a narrative decision space, and constructing state transition rules based on the narrative decision space to obtain a state transition rule set;

[0089] The state transition rule set is mapped to a network topology to obtain an initial transition network, and a structural optimization organization is performed based on the initial transition network to obtain a narrative state transition network, wherein the narrative state transition network includes a node connection metric, a path reachability index, and a network stability parameter.

[0090] Specifically, constructing a narrative state space for the basic narrative segment sequence based on the narrative tension curve to generate a narrative state transition network is a crucial step in ensuring the video content's narrative logic is coherent and engaging. First, the system extracts peak features from the narrative tension curve to identify key tension nodes that significantly impact the viewer's emotions. These key tension nodes represent moments in the video where emotional or information density reaches peaks, such as a product launch or a moment when a team overcomes a major challenge in a corporate promotional video. By performing temporal dependency analysis on these key tension nodes, the system understands the interrelationships and order between different nodes, thereby constructing a node association graph. This graph not only displays the location of each key node but also reveals the temporal and logical connections between them, providing a foundation for subsequent state feature encoding. Next, based on the node association graph, the system performs state feature encoding on the basic narrative segment sequence, generating a set of narrative state vectors. This process aims to transform each basic narrative segment into a quantifiable form for computer processing and analysis. For example, in a corporate promotional video, for a segment depicting teamwork problem-solving, the system encodes it based on multiple aspects of information, including visual elements, audio cues, and textual descriptions, to form a unique state vector. The system then calculates semantic distances on these narrative state vectors to assess the similarities and differences between different states, ultimately constructing a state transition matrix. The state transition matrix not only reflects the transition possibilities between narrative states but also includes information about the costs or benefits of moving from one state to another. This step helps the system understand how to optimize the narrative structure and enhance the audience's viewing experience through reasonable state transitions. The system then partitions the state transition matrix into an action space, resulting in a narrative decision space. This space defines all possible actions (i.e., state transitions), providing the system with a variety of potential narrative paths to choose from. Based on this, the system further constructs a set of state transition rules. These rules not only define which states can be directly connected but also consider the quality of the connection (such as fluency and logical consistency). For example, in a corporate promotional video, the system might set a rule that states transitioning from "Teamwork" can only transition to "Project Progress" or "Achievement Presentation" states, rather than directly jumping to unrelated sections. To achieve this goal, the system maps the state transition rule set to a network topology, generating an initial transition network. While this initial network initially demonstrates the possible connections between states, it requires further optimization to improve its efficiency and stability. Finally, the system optimizes and organizes the initial transition network to produce the final narrative state transition network. During this process, the system not only adjusts the connections between nodes but also ensures good path accessibility and network stability across the entire network.For example, in a corporate promotional video, the system may enhance the coherence of the storyline by adding some transitional clips, or rearrange certain clips to reduce the cognitive burden on the audience. The narrative state transition network includes important components such as node connection metrics, path accessibility indicators, and network stability parameters. These elements together determine the overall structure and quality of the video content. Through this complex processing flow, the system can effectively optimize the narrative structure of the video, so that the final generated content is both logical and emotionally resonant, greatly enhancing the appeal and influence of the video. In specific applications such as the production of corporate promotional videos, this method not only helps creators better organize materials, but also dynamically adjusts the narrative rhythm and focus to meet diverse communication needs.

[0091] In a specific embodiment, the state feature encoding of the basic narrative segment sequence based on the node association graph to obtain a narrative state vector set includes:

[0092] Performing temporal feature decomposition on the node association graph to obtain a tension node feature matrix, and performing hierarchical encoding on the tension node feature matrix to obtain a multi-level feature representation;

[0093] Performing scene element analysis on the basic narrative fragment sequence based on the multi-level feature representation to obtain a scene feature combination set, and performing semantic embedding mapping on the scene feature combination set to obtain a semantic encoding matrix;

[0094] Performing narrative unit alignment on the semantic coding matrix to obtain a narrative structure feature map, and performing multi-dimensional feature fusion based on the narrative structure feature map to obtain a feature fusion tensor;

[0095] Performing state space projection on the narrative structure feature map based on the feature fusion tensor to obtain a state projection vector group, and performing dimensionality reduction processing on the state projection vector group to obtain a state compression representation;

[0096] Vector normalization is performed on the state compression representation to obtain an initial state vector, and state feature enhancement is performed based on the initial state vector to obtain a narrative state vector set.

[0097] Specifically, encoding the state features of the basic narrative fragment sequence based on the node association graph to obtain a set of narrative state vectors is a crucial step in constructing an efficient narrative structure. First, the system performs temporal feature decomposition on the node association graph to extract the features of each key tension node, forming a tension node feature matrix. This process aims to transform complex node association information into a processable data format, enabling the system to identify the importance of each node on the timeline and its relationship to other nodes. For example, in a corporate promotional video, key moments depicting team members overcoming technical difficulties and ultimately successfully launching a product are analyzed and recorded as nodes, forming a detailed feature matrix. Next, the system performs hierarchical encoding on this tension node feature matrix to generate a multi-level feature representation. Hierarchical encoding is a method that captures the deep structure of data through multiple layers of abstraction. It allows the system to not only focus on the information of individual nodes but also understand the development of the entire storyline. For example, in the application scenario of a corporate promotional video, the system may progressively encode each key node from the micro level (such as a specific technical detail) to the macro level (such as the entire project development process), thereby obtaining a comprehensive multi-level feature representation. Based on this multi-level feature representation, the system further analyzes the underlying narrative segment sequences for scene elements, identifying the core elements within each segment and organizing them into scene feature sets. For example, in a description of a team problem-solving process, the system considers factors such as the characters' expressions, the tools used, and the background music. These elements together constitute a complete scene feature set. The system then performs semantic embedding mapping on this scene feature set to generate a semantic encoding matrix. This step converts specific scene features into vector representations in a high-dimensional space, allowing the similarities and differences between different scenes to be mathematically quantified. For example, in a corporate promotional video, different teamwork scenes may differ due to the specific tasks involved. However, through semantic embedding mapping, the system can identify potential commonalities and differences between them, thereby better understanding the overall content framework of the video. Based on the resulting semantic encoding matrix, the system then performs narrative unit alignment to ensure logical consistency and coherence between the narrative segments, thereby generating a narrative structure feature map. The system then utilizes multidimensional feature fusion techniques to integrate information from various sources into a unified feature fusion tensor, which contains all the necessary information for subsequent state space projection. After obtaining the feature fusion tensor, the system performs state-space projection on it, compressing the high-dimensional information into a more manageable space, forming a state projection vector group. To reduce redundant information and improve computational efficiency, the system performs dimensionality reduction on the state projection vector group to generate a compressed state representation. This step is crucial for preserving the core value of the information while reducing complexity.For example, in the production process of corporate promotional videos, by reducing the dimensions of a large amount of material, the system can focus on the key parts that best reflect the theme without getting lost in the details. Finally, the system performs vector normalization on the compressed state representation to ensure that all vectors are in the same scale range for easy comparison and operation. On this basis, the system further enhances the state features, combines all the information accumulated in the previous steps, and finally generates a narrative state vector set. This series of complex processing flows not only enhances the depth of understanding of the video content, but also provides solid data support for subsequent narrative optimization. In this way, corporate promotional videos can not only accurately convey the company's core values, but also attract the audience's attention in a more engaging way.

[0098] In a specific embodiment, the multimodal content generation based on the optimized video storyline to obtain a set of candidate video segments includes:

[0099] Decomposing the optimized video storyline into scene elements to obtain a scene element configuration table, and performing visual style matching on the scene element configuration table to obtain a scene rendering parameter set;

[0100] Reconstructing visual content of the scene element configuration table based on the scene rendering parameter set to obtain a visual scene sequence, and performing dynamic special effects synthesis on the visual scene sequence to obtain a visual effect enhanced sequence;

[0101] Performing audio emotion mapping on the visual effect enhancement sequence to obtain an emotional audio feature set, and generating audio content based on the emotional audio feature set to obtain an audio effect sequence;

[0102] Performing audio and video synchronization optimization on the visual effect enhancement sequence based on the audio effect sequence to obtain an initial synthesized segment set, and performing spatiotemporal consistency calibration on the initial synthesized segment set to obtain a calibrated synthesized sequence;

[0103] The calibrated synthesized sequence is subjected to quality assessment screening to obtain a quality score matrix, and diversified reordering is performed based on the quality score matrix to obtain a set of candidate video segments.

[0104] Specifically, the process of generating multimodal content based on the optimized video storyline to obtain a set of candidate video clips is a key step in transforming a carefully designed narrative structure into concrete, vivid, and engaging video content. First, the system decomposes the optimized video storyline into scene elements, extracting the core elements that comprise each scene and generating a scene element configuration table. For example, in a corporate promotional video, if the storyline emphasizes the importance of teamwork, the scene element configuration table might include information such as footage depicting team collaboration, the tools used, and the background environment. Next, the system performs visual style matching based on the characteristics of these scene elements, determining the most appropriate presentation and visual effects, and thus generating a set of scene rendering parameters. This step not only determines the visual style of the final video, such as color tone, lighting, and composition, but also provides guidance for subsequent visual content reconstruction. For example, for a scene depicting teamwork, the system might select a bright and vibrant visual style to convey a positive and uplifting atmosphere. Based on the resulting scene rendering parameter set, the system performs visual content reconstruction on the scene element configuration table, generating a series of continuous visual scene sequences. This process utilizes advanced computer graphics technology to transform abstract element configurations into concrete images or animations. For example, in a corporate promotional video, the system can create a realistic laboratory work scene based on the description in the configuration sheet, including accurate reproduction of the laboratory equipment and employee operations. To enhance the visual impact, the system also performs dynamic special effects synthesis on the generated visual scene sequence, adding special effects such as light and shadow transformations and particle effects to enhance the visual appeal and appeal of the image. In this way, the dynamic special effects synthesis transforms the originally simple scene into a more engaging visually enhanced sequence. For example, when showcasing a new product launch, flash effects and slow-motion close-ups can be added to highlight the product's highlights and attract the audience's attention. The system then performs audio emotion mapping on the visually enhanced sequence, analyzing the desired emotional tone of each scene and generating a set of emotional audio features based on this. This process considers how sound complements visual elements to enhance the overall emotional expression. For example, in a corporate promotional video, when the team successfully overcomes difficulties and launches a new product, the system selects inspiring background music and uplifting narration to match the scene, thereby enhancing the audience's emotional resonance. Based on the resulting emotional audio feature set, the system further generates audio content to create an audio effect sequence that meets specific emotional requirements. This includes not only the selection of background music but also multi-layered sound design, including dialogue and sound effects, ensuring the overall video is aurally rich and rich. The system then optimizes the audio and video synchronization of the visual enhancement sequence based on the audio effect sequence, ensuring perfect coordination between the two and generating an initial set of synthesized clips. Optimizing audio and video synchronization is a complex process that requires precise timeline matching and consistent emotional expression.For example, in a corporate promotional video, when the scene switches to a team celebrating a success, the system ensures that the climax of the background music perfectly matches the cheers and smiling faces in the video, creating a strong sense of presence. The system then performs spatiotemporal alignment on the initial set of synthesized clips, ensuring temporal and spatial coherence between the clips to avoid frame skipping or inconsistencies, thereby generating a calibrated synthetic sequence. This step is crucial for ensuring the overall smoothness of the video. Finally, the system performs a quality assessment on the calibrated synthetic sequence, generating a quality score matrix based on a series of criteria such as clarity, color balance, and audio quality. Based on this matrix, the system can perform a diversified re-ranking of the synthesized clips, selecting the highest-quality clips to form the candidate video clip set. For example, in the case of a corporate promotional video, the system might prioritize clips that best embody the company culture, product features, and team spirit, while also ensuring the logical and emotional coherence of the overall narrative. Through this complex processing pipeline, the system not only extracts the most expressive content from the raw footage but also provides highly customized solutions to meet diverse communication needs, significantly improving video production efficiency and the quality of the final product.

[0105] In a specific embodiment, the intelligent editing and synthesis of the candidate video clip set to obtain target high-quality video content includes:

[0106] Performing shot boundary detection on the candidate video clip set to obtain a shot segmentation sequence, and performing time sequence combination optimization on the shot segmentation sequence to obtain a shot arrangement scheme;

[0107] Designing a transition effect for the shot segmentation sequence based on the shot arrangement scheme to obtain a transition special effect sequence, and enhancing the visual coherence of the transition special effect sequence to obtain a visually smooth sequence;

[0108] Performing audio rhythm matching on the visually smooth sequence to obtain a sound-image fusion sequence, and performing multi-track mixing arrangement based on the sound-image fusion sequence to obtain a stereo sound effect scheme, wherein the stereo sound effect scheme includes channel balance parameters, audio transition curves, and spatial sound field layout;

[0109] The stereo sound effect solution and the visual smooth sequence are temporally and spatially synchronized to obtain target high-quality video content.

[0110] Specifically, intelligently editing and synthesizing the candidate video clips to produce the target high-quality video content is a key step in integrating the previously generated diverse material into a coherent, high-quality, and engaging final video. First, the system performs shot boundary detection on the candidate video clips. By analyzing image changes and continuity, it identifies the start and end points of each shot, thereby generating a shot segmentation sequence. This process ensures that each individual shot is accurately segmented, providing a foundation for subsequent optimization. For example, in a corporate promotional video, the system can identify different shots that showcase teamwork, such as switching from a panoramic view of the entire team working to a close-up of a team member operating equipment. Next, based on the resulting shot segmentation sequence, the system performs timing optimization to determine the optimal shot arrangement. This step not only considers the logical order of the storyline, but also incorporates factors such as emotional intensity and viewer attention to ensure a smooth and natural narrative throughout the video. For example, in a corporate promotional video, the system might, based on narrative needs, place shots depicting teamwork and problem-solving before product launches to better guide the audience's emotions. Based on this shot arrangement, the system then designs transition effects for the shot sequence, adding appropriate transition effects such as fades or rotations to create a transition sequence. These effects not only enhance visual coherence but also enhance the viewing experience. To further improve visual fluidity, the system enhances the visual coherence of the transition sequence by adjusting the consistency of color, lighting, and composition to create a visually smooth sequence. For example, when showcasing a new product from R&D to launch, the system can use detailed color correction and lighting adjustments to ensure a more natural and harmonious transition between each stage. Building on this visually smooth sequence, the system then performs audio tempo matching to ensure that background music, dialogue, and sound effects are synchronized with the visual changes, creating a fusion of audio and video. This process requires precise timeline alignment and meticulous processing of audio elements to achieve optimal emotional resonance. For example, in a corporate promotional video, when showcasing the team's successful new product launch, the system selects a stirring background music track and ensures that the climax of the music perfectly matches the celebratory scene. Then, based on the obtained sound and picture fusion sequence, the system will perform multi-track mixing and arrangement to generate a stereo sound effect scheme. In this process, the system not only adjusts the channel balance parameters to ensure a reasonable distribution of sound between the left and right channels, but also designs audio transition curves and smooth spatial sound field layouts to provide an immersive auditory experience. For example, in a scene describing team collaboration, the system can use spatial sound field layout technology to make the audience feel as if they are in a real laboratory environment, hearing the discussions of colleagues around them and the operating sounds of equipment. Finally, the system synchronizes the stereo sound effect scheme and the visual smooth sequence in time and space to ensure perfect coordination between the two and generate the target high-quality video content.This step involves complex calculations and precise timeline management, with the aim of eliminating any inconsistencies that may affect the viewing experience. For example, during the production of a corporate promotional video, the system carefully checks every shot switching point to ensure that the audio and video elements are seamlessly connected here, avoiding any frame skipping or audio and video asynchrony. Through this meticulous processing method, the system can not only ensure the logic of the video content and the richness of emotional expression, but also significantly improve the overall viewing quality and communication effect. Therefore, through the intelligent editing and synthesis of the candidate video clip set, the target high-quality video content finally generated can not only accurately convey the core values ​​and cultural characteristics of the company, but also attract and impress the audience in an extremely attractive way, thereby achieving efficient communication purposes.

[0111] The above describes the method for automatically generating high-quality video content in the embodiment of the present invention. The following describes the device for automatically generating high-quality video content in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, an apparatus for automatically generating high-quality video content includes:

[0112] Extraction module 21, used to extract multimodal features from the input original video material to obtain a video feature vector set;

[0113] An analysis module 22 is configured to perform content semantic analysis on the video feature vector set to obtain a video semantic graph;

[0114] An optimization module 23 is configured to optimize the narrative structure of the video semantic graph using deep reinforcement learning technology to obtain an optimized video storyline;

[0115] A generating module 24 is configured to generate multimodal content based on the optimized video storyline to obtain a set of candidate video segments;

[0116] The synthesis module 25 is used to intelligently edit and synthesize the candidate video clip set to obtain target high-quality video content.

[0117] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to the above method embodiment, which will not be repeated here.

[0118] Reference Figure 3 The embodiment of the present invention further provides a computer device, the internal structure of which can be as follows Figure 3As shown. The computer device includes a processor, memory, display screen, input device, network interface and database connected via a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the above method is implemented.

[0119] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied.

[0120] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the above-described method when executed by a processor. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0121] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media provided herein and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM.

[0122] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.

[0123] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A method for automatically generating high-quality video content, characterized in that: The following steps are involved: Perform multimodal feature extraction on the input original video material to obtain a video feature vector set; Performing content semantic analysis on the video feature vector set to obtain a video semantic graph; Optimizing the narrative structure of the video semantic graph using deep reinforcement learning technology to obtain an optimized video storyline; Performing multimodal content generation based on the optimized video storyline to obtain a set of candidate video clips; Intelligently editing and synthesizing the candidate video clip set to obtain target high-quality video content; The method of optimizing the narrative structure of the video semantic graph by using deep reinforcement learning technology to obtain an optimized video storyline includes: Segmenting the video semantic map into narrative units to obtain a basic narrative segment sequence, and calculating plot tension on the basic narrative segment sequence to obtain a narrative tension curve; Based on the narrative tension curve, a narrative state space is constructed for the basic narrative segment sequence to obtain a narrative state transition network, and a reward function is mapped to the narrative state transition network to obtain a narrative optimization strategy space, wherein the narrative optimization strategy space includes state value evaluation, action selection probability, and transfer reward distribution; Performing multiple rounds of exploration sampling on the narrative optimization strategy space to obtain a set of candidate narrative paths, and performing narrative quality assessment based on the set of candidate narrative paths to obtain an optimal narrative decision sequence, wherein the optimal narrative decision sequence includes plot completeness data, narrative coherence measurement, and audience attention prediction; Restructuring the narrative structure of the optimal narrative decision sequence using deep reinforcement learning technology to obtain an optimized narrative framework, and reconstructing the temporal relationship based on the optimized narrative framework to obtain an optimized video storyline, wherein the optimized video storyline includes the main storyline, key plot nodes, and narrative rhythm control parameters; The step of constructing a narrative state space for the basic narrative segment sequence based on the narrative tension curve to obtain a narrative state transition network includes: Extracting peak features from the narrative tension curve to obtain a key tension node sequence, and performing temporal dependency analysis on the key tension node sequence to obtain a node association graph; Performing state feature encoding on the basic narrative segment sequence based on the node association graph to obtain a narrative state vector set, and performing semantic distance calculation on the narrative state vector set to obtain a state transition matrix; Performing action space division on the state transition matrix to obtain a narrative decision space, and constructing state transition rules based on the narrative decision space to obtain a state transition rule set; Performing network topology mapping on the state transition rule set to obtain an initial transition network, and performing structural optimization organization based on the initial transition network to obtain a narrative state transition network, wherein the narrative state transition network includes a node connection metric, a path reachability index, and a network stability parameter; The step of encoding the state features of the basic narrative fragment sequence based on the node association graph to obtain a narrative state vector set includes: Performing temporal feature decomposition on the node association graph to obtain a tension node feature matrix, and performing hierarchical encoding on the tension node feature matrix to obtain a multi-level feature representation; Performing scene element analysis on the basic narrative fragment sequence based on the multi-level feature representation to obtain a scene feature combination set, and performing semantic embedding mapping on the scene feature combination set to obtain a semantic encoding matrix; Performing narrative unit alignment on the semantic coding matrix to obtain a narrative structure feature map, and performing multi-dimensional feature fusion based on the narrative structure feature map to obtain a feature fusion tensor; Performing state space projection on the narrative structure feature map based on the feature fusion tensor to obtain a state projection vector group, and performing dimensionality reduction processing on the state projection vector group to obtain a state compression representation; Vector normalization is performed on the state compression representation to obtain an initial state vector, and state feature enhancement is performed based on the initial state vector to obtain a narrative state vector set.

2. The method for automatically generating high-quality video content according to claim 1, wherein: The multimodal feature extraction is performed on the input original video material to obtain a video feature vector set, including: Decomposing the input original video material in terms of time and space to obtain a video frame sequence set and an audio signal sequence, and performing deep optical flow field calculation on the video frame sequence set to obtain a scene motion feature tensor; Performing hierarchical semantic segmentation on the video frame sequence set based on the scene motion feature tensor to obtain a visual semantic feature map, and performing cross-modal attention mapping on the visual semantic feature map to obtain a visual-semantic association feature matrix; Multimodal temporal fusion is performed on the visual-semantic association feature matrix and the audio signal sequence to obtain a multimodal feature vector set, and context dependency analysis is performed based on the multimodal feature vector set to obtain a video feature vector set, wherein the video feature vector set includes a scene semantic relationship graph, temporal dynamic features and a multimodal interaction pattern.

3. The method for automatically generating high-quality video content according to claim 1, wherein: The performing content semantic analysis on the video feature vector set to obtain a video semantic graph includes: Performing emotion intensity analysis on the video feature vector set to obtain a multidimensional emotion feature vector, and performing temporal correlation calculation on the multidimensional emotion feature vector to obtain an emotion evolution trajectory diagram, wherein the emotion evolution trajectory diagram includes an emotion change gradient, an emotion density distribution, and an emotion inflection point mark; Based on the emotion evolution trajectory graph, hierarchical topic mining is performed on the video feature vector set to obtain a topic structure tree, and semantic relevance reasoning is performed on the topic structure tree to obtain a topic association network, wherein the topic association network includes topic importance scores, key event nodes, and topic migration paths; Performing multi-angle scene semantic analysis on the topic association network to obtain a scene semantic feature set, and performing cross-scene semantic mapping based on the scene semantic feature set to obtain a scene transition matrix, wherein the scene transition matrix includes a scene continuity vector, a scene switching rule, and a scene semantic similarity; The scene conversion matrix and the topic association network are deeply semantically fused to obtain a semantic association topology map, and multi-dimensional feature organization is performed based on the semantic association topology map to obtain a video semantic map, wherein the video semantic map includes narrative logical links, semantic hierarchy and content main line framework.

4. The method for automatically generating high-quality video content according to claim 1, wherein: The multimodal content generation based on the optimized video storyline to obtain a set of candidate video clips includes: Decomposing the optimized video storyline into scene elements to obtain a scene element configuration table, and performing visual style matching on the scene element configuration table to obtain a scene rendering parameter set; Reconstructing visual content of the scene element configuration table based on the scene rendering parameter set to obtain a visual scene sequence, and performing dynamic special effects synthesis on the visual scene sequence to obtain a visual effect enhanced sequence; Performing audio emotion mapping on the visual effect enhancement sequence to obtain an emotional audio feature set, and generating audio content based on the emotional audio feature set to obtain an audio effect sequence; Performing audio and video synchronization optimization on the visual effect enhancement sequence based on the audio effect sequence to obtain an initial synthesized segment set, and performing spatiotemporal consistency calibration on the initial synthesized segment set to obtain a calibrated synthesized sequence; The calibrated synthesized sequence is subjected to quality assessment screening to obtain a quality score matrix, and diversified reordering is performed based on the quality score matrix to obtain a set of candidate video segments.

5. The method for automatically generating high-quality video content according to claim 1, wherein: The intelligent editing and synthesis of the candidate video clip set to obtain target high-quality video content includes: Performing shot boundary detection on the candidate video clip set to obtain a shot segmentation sequence, and performing time sequence combination optimization on the shot segmentation sequence to obtain a shot arrangement scheme; Designing a transition effect for the shot segmentation sequence based on the shot arrangement scheme to obtain a transition special effect sequence, and enhancing the visual coherence of the transition special effect sequence to obtain a visually smooth sequence; Performing audio rhythm matching on the visually smooth sequence to obtain a sound-image fusion sequence, and performing multi-track mixing arrangement based on the sound-image fusion sequence to obtain a stereo sound effect scheme, wherein the stereo sound effect scheme includes channel balance parameters, audio transition curves, and spatial sound field layout; The stereo sound effect solution and the visual smooth sequence are temporally and spatially synchronized to obtain target high-quality video content.

6. A device for automatically generating high-quality video content, characterized in that: include: The extraction module is used to extract multimodal features from the input original video material to obtain a video feature vector set; An analysis module, configured to perform content semantic analysis on the video feature vector set to obtain a video semantic graph; An optimization module, configured to optimize the narrative structure of the video semantic graph using deep reinforcement learning technology to obtain an optimized video storyline; A generation module, configured to generate multimodal content based on the optimized video storyline to obtain a set of candidate video clips; A synthesis module, configured to intelligently edit and synthesize the candidate video clip set to obtain target high-quality video content; The method of optimizing the narrative structure of the video semantic graph by using deep reinforcement learning technology to obtain an optimized video storyline includes: Segmenting the video semantic map into narrative units to obtain a basic narrative segment sequence, and calculating plot tension on the basic narrative segment sequence to obtain a narrative tension curve; Based on the narrative tension curve, a narrative state space is constructed for the basic narrative segment sequence to obtain a narrative state transition network, and a reward function is mapped to the narrative state transition network to obtain a narrative optimization strategy space, wherein the narrative optimization strategy space includes state value evaluation, action selection probability, and transfer reward distribution; Performing multiple rounds of exploration sampling on the narrative optimization strategy space to obtain a set of candidate narrative paths, and performing narrative quality assessment based on the set of candidate narrative paths to obtain an optimal narrative decision sequence, wherein the optimal narrative decision sequence includes plot completeness data, narrative coherence measurement, and audience attention prediction; Restructuring the narrative structure of the optimal narrative decision sequence using deep reinforcement learning technology to obtain an optimized narrative framework, and reconstructing the temporal relationship based on the optimized narrative framework to obtain an optimized video storyline, wherein the optimized video storyline includes the main storyline, key plot nodes, and narrative rhythm control parameters; The step of constructing a narrative state space for the basic narrative segment sequence based on the narrative tension curve to obtain a narrative state transition network includes: Extracting peak features from the narrative tension curve to obtain a key tension node sequence, and performing temporal dependency analysis on the key tension node sequence to obtain a node association graph; Performing state feature encoding on the basic narrative segment sequence based on the node association graph to obtain a narrative state vector set, and performing semantic distance calculation on the narrative state vector set to obtain a state transition matrix; Performing action space division on the state transition matrix to obtain a narrative decision space, and constructing state transition rules based on the narrative decision space to obtain a state transition rule set; Performing network topology mapping on the state transition rule set to obtain an initial transition network, and performing structural optimization organization based on the initial transition network to obtain a narrative state transition network, wherein the narrative state transition network includes a node connection metric, a path reachability index, and a network stability parameter; The step of encoding the state features of the basic narrative fragment sequence based on the node association graph to obtain a narrative state vector set includes: Performing temporal feature decomposition on the node association graph to obtain a tension node feature matrix, and performing hierarchical encoding on the tension node feature matrix to obtain a multi-level feature representation; Performing scene element analysis on the basic narrative fragment sequence based on the multi-level feature representation to obtain a scene feature combination set, and performing semantic embedding mapping on the scene feature combination set to obtain a semantic encoding matrix; Performing narrative unit alignment on the semantic coding matrix to obtain a narrative structure feature map, and performing multi-dimensional feature fusion based on the narrative structure feature map to obtain a feature fusion tensor; Performing state space projection on the narrative structure feature map based on the feature fusion tensor to obtain a state projection vector group, and performing dimensionality reduction processing on the state projection vector group to obtain a state compression representation; Vector normalization is performed on the state compression representation to obtain an initial state vector, and state feature enhancement is performed based on the initial state vector to obtain a narrative state vector set.

7. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Video post-editing and video synthesis optimization method

    CN116847123A

  • Knowledge base construction method based on video content reading analysis

    CN118966329A

  • System for automatically generating video script creation

    CN119383419A