Intelligent video timeline generation system and method

Through the intelligent video timeline generation system, multimodal deep learning and dynamic weight calculation model, intelligent adaptation of media materials and automatic generation of video timelines are realized, solving the problems of low efficiency and insufficient intelligence in traditional content production processes, and significantly improving the automation and intelligence level of content generation.

CN120223975APending Publication Date: 2025-06-27龙港市融媒体中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510528023.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional content production processes are low efficiency, high cost and slow response speed, which is difficult to meet users' needs for real-time and diversification, and the existing technology is difficult to achieve accurate semantic correlations and scene adaptation between materials, limiting the intelligent level of content generation.

Method used

An intelligent video timeline generation system is proposed, which combines the multi-modal deep learning framework to extract and understand media materials in a joint feature and semantic understanding, and combines dynamic weight calculation model and reinforcement learning optimization model to form a closed-loop iteration process to realize the adaptation of intelligent templates and media materials and the automatic generation of timelines.

Benefits of technology

It significantly improves the automation and intelligence level of content generation, solves the problems of low efficiency in traditional video production, rigid template adaptation and frequent manual intervention, and improves the efficiency and quality of video production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223975A_ABST
    Figure CN120223975A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent video timeline generation system and method, and the method comprises the steps: carrying out the joint feature extraction and semantic understanding of videos, audios and texts through a multi-modal deep learning frame, generating structured metadata, triggering the cross-domain adaptation and material migration based on a fragment type selected by a user, and carrying out the recognition of a video timeline. A dynamic weight calculation model is combined with a multi-head attention mechanism to quantify a matching weight of a template and media assets, a structured timeline is generated by taking maximization of a content adaptation degree as a target, conflict nodes or complementary frames are automatically adjusted through a rule engine, a user can adjust and change the timeline in real time through an interactive interface and trigger reinforcement learning to optimize model parameters, and the user experience is improved. The problems that traditional video production is low in efficiency, template adaptation is rigid and manual intervention is frequent are solved, and the automation and intelligence level of content generation is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of content production, and more specifically, to an intelligent video timeline generation system and method. Background Art

[0002] With the rapid development of emerging media forms such as short videos and live broadcasts, the media industry's demand for content production efficiency, personalization, and intelligence has increased significantly. The traditional content production process mainly relies on manual editing and arrangement, which has inherent defects such as low efficiency, high cost, and slow response speed, and it is difficult to meet users' requirements for real-time and diversification. In addition, the complexity and heterogeneity of multi-modal materials pose challenges to content understanding and structured processing. Existing technologies are difficult to achieve precise semantic association and scene adaptation between materials, further restricting the intelligence level of content generation.

[0003] In the generation of templatized content, user requirements show highly diverse characteristics. However, when existing tools intelligently match user-selected templates with media assets, due to the lack of a dynamic weight calculation mechanism and cross-domain migration ability, the resulting video effect deviates greatly from the user's expectations, and the problem of rigid adaptation is prominent. At the same time, timeline arrangement highly depends on manual experience, and the automated generation logic is insufficient in flexibility and adjustability, making it difficult to support real-time iterative optimization, further exacerbating the frequency and complexity of manual intervention.

[0004] In current technologies, problems such as low material utilization rate, rigid template adaptation, and lack of an automated process closed-loop have become the core bottlenecks restricting the efficient production of the media industry. How to build an end-to-end automated production solution through multi-modal fusion, dynamic decision-making, and closed-loop optimization technologies to improve the efficiency and intelligence level of content generation is the key direction that urgently needs to be broken through in this field. Summary of the Invention

[0005] The purpose of this application is to provide an intelligent video timeline generation system and method to overcome the existing technical defects. By adjusting the timeline in real time through an interactive interface and triggering a reinforcement learning optimization model parameter, a closed-loop iterative process is formed, solving the problems of low efficiency, rigid template adaptation, and frequent manual intervention in traditional video production, and significantly improving the automation and intelligence level of content generation.

[0006] The purpose of this application is achieved through the following technical solutions:

[0007] In a first aspect, this application proposes an intelligent video timeline generation system, which includes:

[0008] A multi-modal scene analysis module for jointly extracting multi-modal features from the input media assets through a multi-modal deep learning framework and performing semantic understanding. The media assets include video, audio, and text;

[0009] The intelligent video compilation setting module is used to trigger cross - domain adaptation according to the selected video compilation type, determine whether the material needs to be migrated, call the pre - processing tool chain to complete the migration operation, and perform verification by calculating the MD5 value;

[0010] The dynamic weight calculation module is used to initialize the feature adaptation matrix according to the selected video compilation type, quantify the matching weight between the template attributes and multi - modal features through the multi - head attention mechanism, and perform intelligent adaptation between the template and the media assets using the constructed dynamic weight calculation model;

[0011] The structured timeline module is used to generate the initial shot sequence, special effect nodes, and voice - over insertion points based on the dynamic weight calculation model with the goal of maximizing the user preference and content adaptation degree, and verify the rationality of the timeline through a rule engine, automatically adjusting the node order or inserting frame - filling materials;

[0012] The interactive feedback module is used to provide functions such as drag - and - drop editing of timeline nodes, weight parameter adjustment, and multi - version comparison through a visual interface, record the user's modification behavior to form a feedback data set, and dynamically update the parameters of the dynamic weight model through a reinforcement learning framework.

[0013] In a possible implementation manner, the multi - modal scene analysis module includes:

[0014] The video processing unit is used to extract the spatio - temporal features of the video, identify the shot transition points and segment independent shots through an object detection algorithm, and extract the shot duration, motion type, and key - frame features;

[0015] The face recognition unit is used to detect the face information in the video and label the face position, expression label, and time stamp;

[0016] The audio processing unit is used to perform speech recognition and speaker separation on the audio, and extract the emotion label and music rhythm features based on a pre - trained emotion classification network;

[0017] The text processing unit is used to perform semantic parsing and keyword extraction on the text, and perform cross - modal semantic alignment;

[0018] The feature fusion unit is used to perform time - axis alignment and weighted fusion of the multi - modal features of video, audio, and text in a unified vector space, and generate structured metadata including shot breakdown, face trajectory, scene semantics, and speech summary.

[0019] In a possible implementation manner, the dynamic weight calculation module includes:

[0020] A time density index calculation unit for dynamically adjusting the logical timeline weight based on shot in and out point data, object segmentation data, speech emotion recognition data, and semantic vectorization data;

[0021] An intelligent error correction feedback unit for detecting logical conflicts or material shortages in the timeline through a rule engine and triggering an automatic frame filling operation.

[0022] In a possible implementation, both the structured timeline module and the interactive feedback module include:

[0023] A real-time preview unit for synchronously displaying the adjusted finished video effect in a visualization interface;

[0024] A sensitive content review unit for supporting users to review the finished video from multiple dimensions of content, effect, theme, and sensitive information, and recording user modification behaviors to construct a feedback data set.

[0025] In a possible implementation, the video processing unit extracts scene classification and object detection features through a convolutional neural network.

[0026] In a possible implementation, the feature fusion unit performs weighted fusion on multi-modal features through a cross-modal attention mechanism to generate unified vectorized data.

[0027] In a possible implementation, the cross-domain adaptation function of the intelligent finished video setting module includes:

[0028] Judging the storage location of the materials required by the logical timeline, and initiating the migration of the source code video if the materials are stored in the edge node;

[0029] Calculating the MD5 value for data verification during the migration process.

[0030] In a second aspect, an intelligent video timeline generation method is also proposed in the embodiments of the present application. The method is applied to the intelligent video timeline generation system according to any item in the first aspect, and includes:

[0031] Performing joint feature extraction and semantic understanding on the input media materials through a multi-modal deep learning framework to generate structured metadata including shot breakdowns, face trajectories, scene semantics, and speech summaries;

[0032] Triggering cross-domain adaptation according to the selected type of the finished video, migrating the materials stored in the edge node, and verifying the data integrity;

[0033] Based on a dynamic weight calculation model, quantifying the matching weights between template attributes and multi-modal features through a multi-head attention mechanism to generate an initial shot sequence and special effect nodes;

[0034] Verify the rationality of the timeline and automatically adjust the node order or insert supplementary frame materials through the rule engine to generate a structured timeline;

[0035] Record the user's adjustment behavior through the interactive feedback module and update the parameters of the dynamic weight model to perform iterative optimization of the timeline.

[0036] In a possible implementation manner, the dynamic weight calculation model is used to: dynamically adjust the weight of the logical timeline based on the time density index, and the time density index is obtained by calculating the in and out points of the shot, object segmentation, speech emotion, and semantic vectorization data;

[0037] Detect timeline conflicts through the rule engine and trigger automatic supplementary frame operations.

[0038] In a possible implementation manner, after generating the structured timeline, the method further includes:

[0039] Real-time preview the effect of the adjusted finished film and conduct multi-dimensional review of sensitive content;

[0040] Record the user's review behavior as a feedback data set and use it for parameter update of the dynamic weight model.

[0041] The main solution of the present application and its various further selection solutions can be freely combined to form multiple solutions, all of which are solutions that can be adopted and claimed by the present application; and in the present application, (each non-conflicting selection) selections can be freely combined with each other and with other selections. Those skilled in the art can understand that there are various combinations according to the prior art and common general knowledge after understanding the solution of the present application, all of which are the technical solutions to be protected by the present application, and will not be enumerated here.

[0042] The present application discloses an intelligent video timeline generation system and method, which jointly extracts features and semantic understanding of video, audio, and text through a multi-modal deep learning framework to generate structured metadata; triggers cross-domain adaptation and material migration based on the type of finished film selected by the user, and uses the dynamic weight calculation model combined with the multi-head attention mechanism to quantify the matching weight between the template and the media assets; generates a structured timeline with the goal of maximizing the content adaptation degree, and automatically adjusts conflict nodes or supplements frames through the rule engine; the user can modify the timeline in real time through the interactive interface and trigger the reinforcement learning to optimize the model parameters, forming a closed-loop iterative process, solving the problems of low efficiency of traditional video production, rigid template adaptation, and frequent manual intervention, and significantly improving the automation and intelligence level of content generation. Description of the Drawings

[0043] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0044] Figure 1 It shows a schematic diagram of the intelligent video timeline generation system proposed in the embodiments of the present application.

[0045] Figure 2 It shows a schematic diagram of the framework of the intelligent video timeline generation system proposed in the embodiments of the present application

[0046] Figure 3 It shows the business flow chart of the intelligent video timeline generation system proposed in the embodiments of the present application. Detailed implementation manners

[0047] The following uses specific specific examples to illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0048] Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.

[0049] In the prior art, the traditional manual editing and production method can no longer meet the rapidly growing demand of users for high-quality and personalized video content. However, the user-generated content on social media platforms has increased sharply, but ordinary users lack professional video editing skills. Therefore, to solve the inefficiency problem in the traditional video editing process and support the automated video production requirements in various application scenarios to meet the rapidly growing multiple.

[0050] Therefore, to solve the above technical problems, the embodiments of the present application propose an intelligent video timeline generation system and method. By applying advanced technologies such as artificial intelligence, computer vision, natural language processing, and speech recognition, through AI structured processing and cross-modal processing, it automatically analyzes multi-modal data such as video, audio, and text, and extracts data such as shot segmentation, face information, scene description, and speech-to-text. It realizes a fully automated intelligent video editing process from material upload to finished product output, supports users to customize video types, modes, and audio settings according to personal needs, significantly improves the efficiency and quality of video production, and reduces the professional skill threshold. Next, a detailed description will be given.

[0051] Please refer to Figure 1 , Figure 1 which shows a schematic diagram of the intelligent video timeline generation system proposed by the embodiments of the present application. The system includes:

[0052] A multi-modal scene analysis module, which is used to extract joint features from the input media materials through a multi-modal deep learning framework to obtain multi-modal features and perform semantic understanding. The media materials include video, audio, and text;

[0053] An intelligent finished product setting module, which is used to trigger cross-domain adaptation according to the selected finished product type, judge whether the materials need to be migrated, and call the preprocessing tool chain to complete the migration operation and perform verification by calculating the MD5 value;

[0054] A dynamic weight calculation module, which is used to initialize the feature adaptation matrix according to the selected finished product type, quantify the matching weight between the template attributes and the multi-modal features through the multi-head attention mechanism, and perform intelligent adaptation of the template and the media materials using the constructed dynamic weight calculation model;

[0055] A structured timeline module, which is used to generate an initial shot sequence, special effect nodes, and voice-over insertion points based on the dynamic weight calculation model with the goal of maximizing the content adaptation degree of the user's preference, and verify the rationality of the timeline through a rule engine, and automatically adjust the node order or insert frame-filling materials;

[0056] An interactive feedback module, which is used to provide functions such as drag-and-drop editing of timeline nodes, weight parameter adjustment, and multi-version comparison through a visual interface, record the user's modification behavior to form a feedback data set, and dynamically update the parameters of the dynamic weight model through a reinforcement learning framework.

[0057] The intelligent video timeline generation system is an automated video editing solution based on multi-modal content analysis. It improves the efficiency and quality of video production through the integration of intelligent, automated, and highly adaptable technologies, reduces the professional skill threshold, and realizes full-process automated processing from material upload to finished product output through the collaborative work of five core modules.

[0058] The multi-modal scene analysis module is responsible for jointly extracting features and understanding semantics from the input media assets (including videos, audios, and texts). Specifically, this module extracts spatio-temporal features from videos through a multi-modal deep learning framework, identifies shot transition points and key frame features, and annotates face positions, identity labels, expressions, and timestamp information through face recognition and tracking technologies. In addition, the module also supports scene classification and motion analysis, identifying the background types and object motion trajectories in the videos. For audio data, the module transcribes the audio into text through speech recognition technology, annotates the timestamps and speaker separation, and extracts emotional labels and music rhythm features. The text data achieves cross-modal semantic alignment through semantic parsing and keyword extraction. Finally, through the timeline alignment module, the video, audio, and text features are fused in a unified vector space to generate structured metadata including shot breakdowns, face trajectories, scene semantics, and speech summaries.

[0059] The intelligent video composition setting module is responsible for triggering cross-domain adaptation according to the selected video composition type and determining whether the materials need to be migrated for processing. Users can choose various video composition types, including face-based composition, speech-based composition, shot-based composition, scene-based composition, or manual composition. The system will automatically collect the corresponding video segments according to the user's selection. In addition, users can also choose video composition modes, such as script montage, intelligent graphics, sports event highlights, high-energy montage highlights, or highlight extraction highlights, to meet the video content and production requirements of different types. The audio settings support voiceovers, original sounds, and dubbings. Users can choose to upload professional dubbing audio files, retain the original video sound, or generate a dubbing that matches the video content through AI. The scene settings allow users to select the target scene (such as "festival celebration" or "tech press conference"), and the system will adjust the video tone, transition effects, and content priorities according to the scene labels to ensure a high degree of match between the video theme and the emotional atmosphere.

[0060] The dynamic weight calculation module quantifies the matching weights between the template attributes and multi-modal features through the multi-head attention mechanism, and constructs a dynamic weight calculation model to achieve the intelligent adaptation of the template to the media assets. Specifically, the module initializes the feature adaptation matrix according to the selected video composition type, and dynamically adjusts the matching weights between the template attributes and multi-modal features through the multi-head attention mechanism. This process ensures a high degree of match between the generated video content and the user's needs, while improving the flexibility and intelligence level of template adaptation.

[0061] The structured timeline module aims to maximize the compatibility between user preferences and content, generating an initial shot sequence, special effect nodes, and voice-over insertion points. The module first generates an initial timeline, including the shot sequence, special effect nodes, and voice-over insertion points. Subsequently, the rule engine validates the rationality of the timeline and automatically adjusts the node order or inserts supplementary frame materials to ensure that the generated timeline has clear logic and coherent content. This module is a key link in the video editing process, directly determining the quality of the final video and the user experience.

[0062] The interactive feedback module provides a more flexible editing experience for users through a visual interface. Users can adjust the video content in real time by dragging and editing timeline nodes, adjusting the weight parameter sliders, and comparing multiple versions. All user modification actions are recorded as a feedback dataset, and the system dynamically updates the parameters of the dynamic weight model through a reinforcement learning framework, thus realizing a closed-loop process of "generation - feedback - regeneration". This module not only improves the user experience but also ensures that the final video result precisely meets the user's needs by iteratively optimizing the generation logic.

[0063] The multi-modal scene analysis module includes:

[0064] A video processing unit for extracting the spatio-temporal features of the video, identifying shot transition points and segmenting independent shots through object detection algorithms, and extracting shot duration, motion type, and key frame features;

[0065] A face recognition unit for detecting face information in the video and annotating face positions, expression labels, and timestamps;

[0066] An audio processing unit for performing speech recognition and speaker separation on the audio, and extracting emotion labels and music rhythm features based on a pre-trained emotion classification network;

[0067] A text processing unit for performing semantic parsing and keyword extraction on the text and achieving cross-modal semantic alignment;

[0068] A feature fusion unit for performing time-axis alignment and weighted fusion of multi-modal features of video, audio, and text in a unified vector space, generating structured metadata including shot breakdowns, face trajectories, scene semantics, and speech summaries.

[0069] The multi-modal scenario analysis module performs joint feature extraction and semantic understanding on media materials such as videos, audios, and texts through a multi-modal deep learning framework. The video processing unit uses object detection algorithms to identify shot transition points and segment independent shots, extracting spatio-temporal features, motion types, and key frame features; the face recognition unit detects and labels the positions, expressions, and timestamps of faces; the audio processing unit extracts emotional tags and music rhythm features through speech recognition and speaker separation technologies; and the text processing unit performs semantic parsing and keyword extraction. The features of each modality are time-axis aligned and weighted fused by the feature fusion unit in a unified vector space, finally generating structured metadata including shot breakdowns, face trajectories, scene semantics, and speech summaries.

[0070] This module adopts a technical path of "modalities processing - cross-modal alignment - feature fusion". By the video processing unit, visual spatio-temporal features are extracted; by the face recognition unit, the character trajectories are constructed; by the audio processing unit, the emotional rhythm is analyzed; and by the text processing unit, the semantic content is parsed. Finally, the feature fusion unit realizes the spatio-temporal alignment and semantic association of multi-modal data, outputting structured metadata with both temporality and semantics.

[0071] Figure 2 The schematic diagram of the intelligent video timeline generation system framework proposed in the embodiment of this application is shown. This framework includes a front-end interaction layer, an ability service layer, and a basic support layer;

[0072] The front-end interaction layer supports users to initiate intelligent video editing with one click. That is, users upload materials, including videos, audios, images, etc., and set video editing parameters, including video type, editing mode, audio configuration, scene selection, etc. The system automatically generates a visual timeline. Users can adjust and submit the synthesis and rendering. After generating the complete video, it supports users to review and submit feedback. The system modifies the timeline content according to the user feedback and re-synthesizes, and so on, supporting the generation of the final version of the video that satisfies users, providing users with an operation experience with low threshold and high flexibility, and meeting the diverse needs from ordinary users to professional creators;

[0073] The ability service layer includes capabilities such as visual models, LLMs, audio cloning, multi-modal AI, etc., as well as services such as multi-modal scenario analysis, intelligent video editing generation, intelligent video editing, synthesis and rendering. Based on capabilities such as video image content analysis, text semantic understanding, replicating speech content with specific tones, and cross-modal data fusion analysis, the system performs multi-modal scenario analysis on the materials uploaded by users, forming structured and vectorized data label data, used to intelligently generate a timeline that meets user needs, supporting automated video editing and enhancement such as transition effects, picture quality enhancement, and intelligent frame filling. Finally, the multi-track data is synthesized into the final video;

[0074] The basic support layer includes cloud host services, network services, storage services, and security services, which are used to provide local or edge node computing resources, dedicated line access, data migration and transmission, data storage (such as OSS, MongoDB, MySQL, etc.), and security services such as waf and data verification.

[0075] After the user uploads the materials and sets the finished video parameters through the front-end interaction layer, the system starts the multi-modal scene analysis module to deeply analyze the materials and generate structured metadata. Subsequently, the intelligent video editing setting module generates editing instructions according to the selected finished video type and mode, the dynamic weight calculation module quantifies the matching weight between the template and the materials, and the structured timeline module generates and optimizes the timeline. Finally, the system outputs the finished video through the synthesis and rendering module for the user to review. The user can adjust the timeline and submit feedback through the interactive feedback module, and the system optimizes the generation logic according to the feedback data set until the final version of the video that satisfies the user is output. Through this intelligent and automated process, the intelligent video timeline generation system provides an efficient and flexible solution for video content production, significantly reducing the production threshold and cost, and meeting the diverse needs of ordinary users to professional creators.

[0076] The dynamic weight calculation module includes:

[0077] The time density index calculation unit is used to dynamically adjust the weight of the logical timeline based on shot in and out point data, object segmentation data, speech emotion recognition data, and semantic vectorization data;

[0078] The intelligent error correction feedback unit is used to detect logical conflicts or missing materials in the timeline through the rule engine and trigger the automatic frame filling operation.

[0079] The dynamic weight calculation module is a key component in the intelligent video timeline generation system, which optimizes the video editing process through two main units: the time density index calculation unit and the intelligent error correction feedback unit.

[0080] The time density index calculation unit is responsible for dynamically adjusting the weight of the logical timeline. This unit analyzes the shot in and out point data to determine the start and end times of each shot; the object segmentation data to identify the positions and movements of different objects in the video; the speech emotion recognition data to evaluate the emotional intensity and changes in the audio; and the semantic vectorization data to understand the semantic information of the video content. By integrating these data, the time density index calculation unit can dynamically adjust the weights of each segment in the timeline, ensuring the fluency and coherence of the video content, while highlighting key plots and emotional climaxes.

[0081] The intelligent error correction and feedback unit detects conflicts in the timeline through a rule engine. This unit can identify unreasonable editing points or content inconsistencies in the timeline, such as jump cuts between shots, audio-video out-of-sync, etc. Once these conflicts are detected, the intelligent error correction and feedback unit triggers an automatic frame filling operation to solve these problems by inserting additional frames or adjusting existing frames, thus ensuring the final output quality of the video. These two units work together, enabling the dynamic weight calculation module to provide intelligent adjustments and optimizations during the video editing process, thereby improving the efficiency and quality of video production.

[0082] Both the structured timeline module and the interactive feedback module include:

[0083] A real-time preview unit for synchronously displaying the adjusted finished video effect in the visualization interface;

[0084] A sensitive content review unit for supporting users to review the finished video from multiple dimensions including content, effect, theme, and sensitive information, and recording user adjustment behaviors to build a feedback dataset.

[0085] The structured timeline module and the interactive feedback module jointly ensure the efficiency and accuracy of video editing in the intelligent video timeline generation system. The real-time preview unit in the structured timeline module can synchronously display the adjusted finished video effect. When editing the timeline, users can see the adjusted video effect in real time through the visualization interface, including changes in the shot sequence, insertion of special effect nodes, synchronization of dubbing, etc. This instant feedback mechanism not only improves the user experience but also significantly enhances the editing efficiency, enabling users to complete high-quality video production faster.

[0086] The sensitive content review unit in the interactive feedback module supports users to review the finished video from multiple dimensions such as content, effect, theme, and sensitive information. Content review ensures that the video content meets user requirements and checks for any missing or irrelevant content; effect review evaluates the visual and audio effects of the video to ensure the coordination of elements such as transitions, special effects, and dubbing; theme review verifies whether the video theme is clear and whether the content is highly relevant to the user-set theme; sensitive information review detects whether there is sensitive or inappropriate content in the video, such as inappropriate remarks, privacy information, etc. The sensitive content review unit builds a feedback dataset by recording user adjustment behaviors, and these feedback data are used in the reinforcement learning framework to dynamically update the parameters of the dynamic weight model, thereby optimizing the generation logic of the system.

[0087] The real-time preview unit and the sensitive content review unit work together to provide users with an efficient and flexible video editing environment. When users adjust the timeline, they can immediately see the effects and ensure the quality and compliance of the content. This dual guarantee mechanism not only improves the efficiency of video production but also ensures the quality and security of the final video, providing users with a comprehensive and intelligent video editing solution that meets the full process needs from creative concept to final output.

[0088] The video processing unit extracts scene classification and object detection features through a convolutional neural network.

[0089] The video processing unit is responsible for in-depth AI structured processing of the materials uploaded by users in the intelligent video timeline generation system. Specifically, this module analyzes the video footage through advanced visual models, accurately detects the scenes, characters, and shot information therein, and stores these key data in a structured manner for subsequent editing and processing.

[0090] The feature fusion unit performs weighted fusion on multi-modal features through a cross-modal attention mechanism to generate unified vectorized data.

[0091] When processing video materials, the feature fusion unit not only focuses on visual elements but also analyzes audio and text content to extract multi-modal features. These features include visual features in the video, acoustic features in the audio, and semantic features in the text. By performing weighted fusion on these features, the module can generate unified vectorized data, which is convenient for computer processing and analysis.

[0092] Finally, the data output by the multi-modal scene analysis module includes structured and vectorized information about shots, characters, scenes, and voices. These data provide a solid foundation for intelligent video editing, enabling the system to automatically generate video content that meets the requirements according to users' needs and preferences.

[0093] The cross-domain adaptation function of the intelligent video generation setting module includes:

[0094] Judging the storage location of the materials required for the logical timeline. If the materials are stored at the edge node, initiate the migration of the source video.

[0095] Perform data verification by calculating the MD5 value during the migration process.

[0096] The cross - domain adaptation function of the intelligent video clip setting module is an important part of the intelligent video timeline generation system, which ensures that materials can be efficiently and securely migrated from different storage locations to the required processing locations. Specifically, this module first determines the storage location of the materials required in the logical timeline. If the materials are stored on the edge nodes, the system will automatically initiate the migration operation of the source video, migrating the materials from the edge nodes to the local storage so that subsequent editing and processing can proceed smoothly.

[0097] During the migration process, to ensure data consistency and integrity, the system calculates the MD5 values of the data before and after migration. The MD5 value is a widely used hash function that can generate a unique fingerprint of the data. By comparing the MD5 values before and after migration, the system can verify whether any changes or damages have occurred to the data during migration. If the MD5 values match, it indicates that the migrated data is complete and consistent and can be used for subsequent video editing processes.

[0098] This cross - domain adaptation function not only improves the efficiency of material processing but also ensures data reliability and security, providing a solid foundation for intelligent video editing. In this way, the intelligent video clip setting module can effectively manage the storage and migration of materials and support the smooth progress of the entire video generation process.

[0099] Figure 3 The business flow chart of the intelligent video timeline generation system proposed in the embodiments of the present application is shown. In multi - modal scenario analysis, materials are uploaded locally or from the system content library. Through AI structured processing and AI cross - modal processing, media asset material AI structured data and vectorized data are output. Then, the user selects parameters such as intelligent video clip type, clip mode, original voiceover for voiceover, and scene based on the multi - modal scenario analysis data to complete the intelligent video clip setting. A structured timeline is generated according to the intelligent video clip parameter information and file address information. Finally, the synthesized and rendered video clip is output for the user to review, and the review data interaction feedback guides the timeline reconstruction, and so on until a video clip satisfactory to the user is output.

[0100] In addition, the present application proposes an intelligent video timeline generation method, which is applied to the above - mentioned intelligent video timeline generation system. The method includes the following steps:

[0101] Joint feature extraction and semantic understanding of the input media asset materials are performed through a multi - modal deep learning framework to generate structured metadata including shot breakdowns, face trajectories, scene semantics, and speech summaries;

[0102] Trigger cross - domain adaptation according to the selected video clip type by the user, perform migration processing on the materials stored on the edge nodes, and verify data integrity;

[0103] Based on the dynamic weight calculation model, the matching weights of template attributes and multi-modal features are quantified through the multi-head attention mechanism to generate the initial shot sequence and special effect nodes;

[0104] Verify the rationality of the timeline and automatically adjust the node order or insert fill-frame materials through the rule engine to generate a structured timeline;

[0105] Record the user's adjustment behavior through the interactive feedback module and update the parameters of the dynamic weight model to perform iterative optimization of the timeline.

[0106] The intelligent video timeline generation method proposed in this application realizes the full-process automation from material input to finished video output through a series of intelligent steps. First, through the multi-modal deep learning framework, joint feature extraction and semantic understanding are performed on the input media materials to generate structured metadata including shot breakdowns, face trajectories, scene semantics, and speech summaries. These metadata provide a solid data foundation for subsequent intelligent editing.

[0107] On this basis, the system triggers the cross-domain adaptation function according to the type of finished video selected by the user. If the required materials are stored on the edge nodes, the system will automatically initiate the migration operation of the source video, and verify the integrity and consistency of the data by calculating the MD5 value to ensure that the migrated materials can be accurately used for subsequent processing.

[0108] Next, based on the dynamic weight calculation model, the system quantifies the matching weights of template attributes and multi-modal features through the multi-head attention mechanism. This process can intelligently generate the initial shot sequence and special effect nodes to ensure that the generated timeline highly matches the user's needs.

[0109] To further improve the quality of the timeline, the system verifies the rationality of the timeline and automatically adjusts the node order or inserts fill-frame materials through the rule engine to generate the final structured timeline. This step ensures the clear logic and coherent content of the timeline.

[0110] Finally, through the interactive feedback module, the system records the user's adjustment behavior and uses this behavior data to update the parameters of the dynamic weight model. This closed-loop optimization mechanism enables the timeline to be continuously iteratively improved, and finally generates high-quality video content that meets the user's needs.

[0111] Through this method, this application realizes the intelligence, automation, and personalization of video editing, significantly improves the efficiency and quality of video production, reduces the professional skill threshold at the same time, and provides users with an efficient and flexible video production solution.

[0112] A dynamic weight calculation model for dynamically adjusting the weight of the logical timeline based on a time density index, where the time density index is calculated from shot entry and exit points, object segmentation, voice emotion, and semantic vectorized data;

[0113] The rule engine detects timeline conflicts and triggers automatic frame filling operations.

[0114] Dynamic weight calculation ensures the smoothness and coherence of video content while highlighting key plots and emotional climaxes. Dynamic weight calculation includes two main aspects: dynamically adjusting the weight of the logical timeline based on the time density index, and detecting timeline conflicts through the rule engine and triggering automatic frame filling operations.

[0115] The time density index is calculated by analyzing the shot entry and exit points, object segmentation, speech emotion, and semantic vectorization data. The shot entry and exit point data determines the start and end time of each shot, which is crucial for understanding the rhythm of the video. Object segmentation data identifies the position and movement of different objects in the video, which helps to understand the complexity and dynamics of the picture. Speech emotion data evaluates the intensity and changes of emotions in the audio, providing important information for the emotional expression of the video. Semantic vectorization data provides a deep understanding of the video content by converting video content into semantic vectors. Combining these data, the time density index can dynamically adjust the weight of each segment in the timeline to ensure that the rhythm and emotional expression of the video content are optimal.

[0116] On the other hand, the rule engine is used to detect conflicts in the timeline, such as jump cuts between shots, audio and video out of sync, etc. Once these conflicts are detected, the system triggers automatic frame filling operations to resolve these problems by inserting additional frames or adjusting existing frames. This mechanism ensures the final output quality of the video and avoids visual or auditory inconsistencies caused by timeline conflicts.

[0117] Through dynamic weight calculation, the intelligent video timeline generation system can provide an intelligent and automated video editing process, significantly improving the efficiency and quality of video production. This mechanism not only optimizes the presentation of video content, but also ensures that the final effect of the video meets the user's expectations through automated error correction and optimization functions.

[0118] The method further includes:

[0119] Preview the adjusted film effect in real time and conduct multi-dimensional review of sensitive content;

[0120] User review behaviors are recorded as feedback datasets and used to update parameters of the dynamic weight model.

[0121] In the intelligent video timeline generation system, after generating the structured timeline, the system further provides real-time preview and sensitive content review functions, and optimizes the dynamic weight model by recording users' review behaviors.

[0122] First, the system synchronously displays the adjusted finished video effect through the real-time preview unit. When editing the timeline, users can visually see the adjusted video effect in real time, including changes in the shot sequence, insertion of special effect nodes, synchronization of dubbing, etc. This immediate feedback mechanism significantly improves the user experience and editing efficiency, enabling users to quickly verify the editing effect and make adjustments.

[0123] Meanwhile, the system supports users to review the finished video from multiple dimensions through the sensitive content review unit. This includes content review to ensure the integrity and accuracy of the video content; effect review to evaluate whether the visual and audio effects of the video meet the expectations; theme review to verify whether the video theme is clear and consistent with the user's needs; and sensitive information review to detect whether there is inappropriate or sensitive content in the video. This multi-dimensional review mechanism helps users ensure the quality and compliance of the video.

[0124] In addition, the system records the users' behaviors during the review process as a feedback data set. These data reflect the users' preferences and adjustment intentions for the video content, providing valuable training data for the dynamic weight model. Through the reinforcement learning framework, the system dynamically updates the parameters of the dynamic weight model using this feedback data, thereby achieving iterative optimization of the timeline. This closed-loop optimization mechanism ensures that the system can continuously improve and generate high-quality video content that better meets the users' needs.

[0125] Through real-time preview and sensitive content review, combined with model optimization driven by user feedback, the intelligent video timeline generation system not only improves the efficiency of video production but also ensures the quality of the final finished video and user satisfaction. This intelligent and automated process provides an efficient and flexible solution for video content production.

[0126] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0127] First, through the multi-modal scene analysis module, the materials are automatically parsed, key features are extracted, and structured metadata is generated, reducing the workload of manual analysis. The intelligent timeline generation function is based on the dynamic weight calculation model and the rule engine, automatically generating and optimizing the timeline, avoiding the cumbersome process of manual construction. The real-time preview and feedback mechanism allows users to immediately see the adjustment effect, and combined with the reinforcement learning framework to dynamically update the model parameters, further improving the editing efficiency.

[0128] Second, multi-modal feature extraction and dynamic weight calculation ensure the intelligent adaptation of templates and materials, and the generated video content highly matches the user's needs. The time density index adjusts the logical timeline weight to ensure content coherence and rhythm fluency. The intelligent error correction feedback unit detects and resolves timeline conflicts, and the sensitive content review unit reviews the finished video from multiple dimensions to ensure content security and compliance, thus improving the professionalism and personalization of the video.

[0129] Third, users only need to upload materials and set simple parameters, and the system automatically completes the whole process from analysis to output without professional video editing skills. The visual editing interface supports simple operations such as dragging and cropping, enabling users to easily achieve personalized customization, reducing the operation difficulty of video editing, and enhancing the user experience.

[0130] Fourth, the cross-domain adaptation function automatically determines the storage location of materials, initiates migration when necessary, and ensures data integrity to adapt to different working environments. The multi-version management and backtracking function facilitate users to save and switch versions, enhancing the flexibility and controllability of creation. The system can flexibly handle various complex scenarios and meet the diverse needs of users.

[0131] This is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. An intelligent video timeline generation system, characterized in that: The system comprises: A multimodal scene analysis module is used to extract joint features from input media materials through a multimodal deep learning framework to obtain multimodal features and perform semantic understanding, wherein the media materials include video, audio, and text; The intelligent film setting module is used to trigger cross-domain adaptation according to the film type selected by the user, determine whether the material needs to be migrated, and call the pre-processing tool chain to complete the migration operation, and verify it by calculating the MD5 value; The dynamic weight calculation module is used to initialize the feature adaptation matrix according to the type of film selected by the user, quantify the matching weights of template attributes and multimodal features through the multi-head attention mechanism, and use the constructed dynamic weight calculation model to perform intelligent adaptation of templates and media assets; The structured timeline module is used to generate the initial shot sequence, special effect nodes and dubbing insertion points based on the dynamic weight calculation model to maximize user preference and content adaptability, and verify the rationality of the timeline through the rule engine, automatically adjust the node order or insert interpolation frame materials; The interactive feedback module is used to provide drag-and-drop editing of timeline nodes, weight parameter adjustment and multi-version comparison functions through a visual interface, record user modification behaviors to form a feedback data set, and dynamically update the parameters of the dynamic weight model through a reinforcement learning framework.

2. The intelligent video timeline generation system according to claim 1, characterized in that: The multimodal scene analysis module includes: The video processing unit is used to extract the temporal and spatial features of the video, identify the shot switching points and segment the independent shots through the target detection algorithm, and extract the shot duration, motion type and key frame features; Face recognition unit, used to detect face information in the video and mark face position, expression label and timestamp; An audio processing unit, which is used to perform speech recognition and speaker separation on the audio, and extract emotion labels and music rhythm features based on a pre-trained emotion classification network; The text processing unit is used to perform semantic analysis and keyword extraction on the text and perform cross-modal semantic alignment; The feature fusion unit is used to perform time axis alignment and weighted fusion of multimodal features of video, audio, and text in a unified vector space to generate structured metadata including shot storyboards, face trajectories, scene semantics, and voice summaries.

3. The intelligent video timeline generation system according to claim 1, characterized in that: The dynamic weight calculation module includes: A time density index calculation unit, used to dynamically adjust the weight of the logical timeline based on the shot entry and exit point data, object segmentation data, speech emotion recognition data and semantic vectorization data; Intelligent error correction feedback unit, which detects logical conflicts or missing materials in the timeline through the rule engine and triggers automatic frame filling operations.

4. The intelligent video timeline generation system according to claim 1, characterized in that: Both the structured timeline module and the interactive feedback module include: Real-time preview unit, used to synchronously display the adjusted film effect in the visual interface; The sensitive content review unit is used to support users in reviewing films from multiple dimensions including content, effects, themes and sensitive information, and to record user adjustment behaviors to build a feedback data set.

5. The intelligent video timeline generation system according to claim 2, characterized in that: The video processing unit extracts scene classification and object detection features through convolutional neural networks.

6. The intelligent video timeline generation system according to claim 2, characterized in that: The feature fusion unit performs weighted fusion of multimodal features through a cross-modal attention mechanism to generate unified vectorized data.

7. The intelligent video timeline generation system according to claim 1, characterized in that: The cross-domain adaptation function of the intelligent film setting module includes: Determine the storage location of the material required by the logical timeline. If the material is stored in an edge node, initiate the source code video migration. During the migration process, data is verified by calculating the MD5 value.

8. An intelligent video timeline generation method, characterized in that: The method is applied to the intelligent video timeline generation system of any one of claims 1 to 7, comprising: Through the multimodal deep learning framework, the input media materials are jointly extracted and semantically understood to generate structured metadata including shot splitting, face trajectory, scene semantics and voice summary; Trigger cross-domain adaptation based on the film type selected by the user, migrate the materials stored in the edge nodes and verify the data integrity; Based on the dynamic weight calculation model, the matching weights of template attributes and multimodal features are quantified through the multi-head attention mechanism to generate the initial shot sequence and special effect nodes; Verify the rationality of the timeline and automatically adjust the node order or insert interpolation materials through the rule engine to generate a structured timeline; The interactive feedback module records user adjustment behaviors and updates dynamic weight model parameters to perform iterative optimization of the timeline.

9. The intelligent video timeline generation method according to claim 8, characterized in that: The dynamic weight calculation model is used to dynamically adjust the logical timeline weight based on the time density index, and the time density index is calculated by the shot entry and exit points, object segmentation, voice emotion and semantic vectorized data; The rule engine detects timeline conflicts and triggers automatic frame filling operations.

10. The intelligent video timeline generation method according to claim 8, characterized in that: After generating the structured timeline, the method further includes: Preview the adjusted film effect in real time and conduct multi-dimensional review of sensitive content; User review behaviors are recorded as feedback datasets and used to update parameters of the dynamic weight model.