A video processing method and device based on multi-modal data, equipment and medium
By employing multimodal data analysis and collaborative processing technologies, we have solved the problems of low utilization of multimodal data, lack of emotional drive throughout the entire process, and poor narrative adaptability in video processing. This results in the generation of high-quality, logically coherent, and emotionally resonant video content that meets the needs of Vlog and short video scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-14
AI Technical Summary
Existing video processing technologies suffer from low multimodal data utilization, lack of emotion-driven approaches throughout the entire process, poor narrative and scene adaptability, and weak synergy between technology combinations. These issues result in fragmented video content, inconsistent logic, discontinuous emotional expression, and an inability to accurately match user viewing habits.
By analyzing multimodal data, we obtain sentiment time-series data, event segmentation results, topic clustering labels, and keyframe sequences, construct a narrative framework, extract highlight segments, and achieve deep collaborative processing of video, audio, and text using a multi-algorithm collaborative technical solution.
It improves the logical coherence, duration adaptability, and emotional resonance of video content, generates high-quality and smooth video content, reduces the cost of manual intervention, and adapts to the needs of Vlog and short video scenarios.
Smart Images

Figure CN121397299B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video processing technology, and in particular to a video processing method, apparatus, device, and medium based on multimodal data. Background Technology
[0002] With the popularization of short video and vlog content creation, users' demand for high-quality, story-driven video content is growing. However, existing video processing technologies generally have the following shortcomings when facing this demand:
[0003] (1) Low utilization of multimodal data: Current video processing methods mostly use a single data source for independent processing, such as analyzing the image based only on keyframes or selecting segments based only on event segmentation results. This processing method results in a lack of effective closed-loop correlation between different modal data, which fails to give full play to the collaborative value of multimodal data. As a result, the generated video content often presents fragmented and illogical problems, making it difficult to form a complete and smooth narrative whole.
[0004] (2) Emotion-driven approach not permeating the entire process: In the existing video creation process, the application of emotional data is often limited to a single stage, such as simply referencing emotional tendencies in the audio and music stage, without deeply linking it with the video content in key steps such as event selection, highlight extraction, and narrative design. This partial and fragmented application of emotion leads to a break in the emotional expression of the video, making it impossible to continuously and effectively guide the audience's emotional experience, ultimately affecting the emotional resonance and appeal of the video.
[0005] (3) Poor narrative and scene adaptability: Existing video narrative construction methods mostly adopt fixed narrative templates, such as a simple "beginning-middle-end" structure, which fail to fully consider the "3-5 minute golden viewing time" limit and the narrative logic requirements of "introduction, development, climax and conclusion" unique to short video formats such as Vlog. This narrative approach, which is not adapted to specific scenes, is prone to problems such as redundant or insufficient information in video content, and cannot accurately match the viewing habits and content expectations of the audience, thereby reducing the attractiveness and dissemination effect of the video.
[0006] (4) Weak synergy of technology combinations: In the key technical aspects of video generation, existing solutions often rely on a single algorithm for processing, such as using only the Latent Dirichlet Allocation (LDA) model for topic clustering, or only using keyframe scores for highlight segment selection. This single-algorithm application approach fails to form a multi-algorithm collaborative technical solution according to the complex needs of video creation, limiting the room for improvement in video generation quality and making it difficult to achieve complementary advantages and synergistic effects among various technical modules.
[0007] Therefore, there is an urgent need for a video processing method, device, equipment, and medium based on multimodal data to address the shortcomings of existing technologies. Summary of the Invention
[0008] The purpose of this invention is to propose a video processing method, apparatus, device, and medium based on multimodal data to solve the problems of multimodal data fragmentation, emotion-driven gaps, poor narrative adaptability, and weak technological synergy in existing video processing technologies. By integrating multi-dimensional information such as emotion, events, themes, and visuals, it constructs video content with narrative logic, realizing the automated and intelligent generation from raw video to story-based video, thereby improving video content quality and production efficiency.
[0009] In a first aspect, to achieve the above objectives, the present invention provides a video processing method based on multimodal data, comprising the following steps:
[0010] S1. Obtain multimodal data analysis results using the original multimodal data of the video to be processed;
[0011] S2. Perform event and narrative processing based on the multimodal data analysis results to obtain event processing results and narrative framework;
[0012] S3. Based on the multimodal data analysis results and the event processing results, extract highlight segments to obtain video highlight segments;
[0013] S4. Based on the multimodal data analysis results, the narrative framework, and the video highlight segments, obtain the video processing results;
[0014] The multimodal data analysis results include sentiment time-series data, event segmentation results, topic clustering labels, and keyframe sequences. The topic clustering labels include global topics and local subtopics.
[0015] Optionally, S1, using the multimodal raw data of the video to be processed, obtain the multimodal data analysis results;
[0016] Collect multimodal raw data of the video to be processed, including video frame sequences, audio signals, associated text information and metadata;
[0017] The multimodal raw data of the video to be processed is preprocessed to obtain standardized multimodal data;
[0018] Using the standardized multimodal data, a multimodal sentiment fusion algorithm is used to obtain sentiment time-series data;
[0019] Based on the standardized multimodal data, an image-audio collaborative segmentation strategy is adopted to obtain event segmentation results;
[0020] Semantic analysis is performed on the standardized multimodal data to obtain topic clustering labels;
[0021] Keyframes are extracted using the standardized multimodal data to obtain a keyframe sequence;
[0022] Based on the emotional time-series data, the event segmentation results, the topic clustering labels, and the keyframe sequence, multimodal data analysis results are obtained.
[0023] Optionally, S2, based on the multimodal data analysis results, perform event and narrative processing to obtain event processing results and a narrative framework, including:
[0024] The importance of the event segmentation results is evaluated based on the emotional time series data to obtain the event importance evaluation results;
[0025] High-value events are selected by combining the event importance assessment results with an emotional intensity threshold.
[0026] The high-value events are compared using a pre-trained language model to obtain the event processing results.
[0027] Based on the event processing results and the topic clustering labels, a joint model is used to perform correlation analysis to obtain the event correlation relationships;
[0028] Based on the event segmentation results and the event relationships, an event timeline is constructed.
[0029] Based on the target video length, a three-act narrative structure is constructed;
[0030] Based on the high-value events, the event timeline, and the three-act narrative structure, obtain the narrative framework;
[0031] The joint model includes a latent Dirichlet assignment model and a pre-trained language model.
[0032] Optionally, S3, based on the multimodal data analysis results and the event processing results, highlight clip extraction is performed to obtain video highlight clips, including:
[0033] Using the event processing results, the extraction range of the highlight fragment is determined;
[0034] Based on the multimodal data analysis results and the extraction range of the highlight fragments, a multidimensional evaluation is performed to obtain multidimensional evaluation results;
[0035] The multi-dimensional evaluation results are weighted and fused to obtain a comprehensive evaluation result;
[0036] Based on the extraction range of the highlight segments and the comprehensive evaluation results, the highlight segments of the video are obtained.
[0037] Optionally, based on the multimodal data analysis results and the extraction range of the highlight fragment, a multi-dimensional evaluation is performed to obtain multi-dimensional evaluation results, including:
[0038] By utilizing the extraction range of the highlight fragments, candidate fragments for multi-dimensional evaluation are determined;
[0039] Based on the keyframe sequence and the candidate segments evaluated in the multi-dimensional assessment, a content importance score is obtained to acquire the keyframe score.
[0040] Based on the emotional time-series data and the candidate segments evaluated in the multi-dimensional assessment, the slope of emotional change is obtained;
[0041] Based on the candidate segments evaluated in the multi-dimensional assessment, motion intensity and visual intersection are evaluated to obtain visual feature scores;
[0042] A multi-dimensional evaluation result is obtained based on the keyframe score, the slope of the emotion change, and the visual feature score.
[0043] Optionally, S4, based on the multimodal data analysis results, the narrative framework, and the video highlight segment, obtain the video processing results, including:
[0044] Based on the narrative framework and the event processing results, extract the core transition segments;
[0045] Based on the emotional time series data and the topic clustering tags, combined with the narrative framework, the video highlight segments and the core transition segments are integrated to obtain video footage;
[0046] Based on the emotional timing data and the video footage, combined with the narrative framework, a dynamic time adjustment algorithm is used to obtain the audio content.
[0047] Based on the topic clustering tags, the video highlight segments, and the core transition segments, combined with the narrative framework, the text content is obtained;
[0048] The video processing result is obtained using the video footage, the audio content, and the text content.
[0049] Optionally, the video processing result can be obtained using the video footage, the audio content, and the text content, including:
[0050] Based on the narrative framework, the video footage, audio content, and text content are calibrated to obtain optimized video footage, optimized audio content, and optimized text content, respectively.
[0051] Determine whether the timestamps corresponding to the optimized video, optimized audio, and optimized text are synchronized. If they are, perform the first operation; otherwise, perform the second operation.
[0052] The first operation is to integrate the optimized video footage, the optimized audio content, and the optimized text content to obtain an initial video processing result, and then perform the third operation.
[0053] The second operation is: determining whether the difference between the timestamps corresponding to the optimized video frame, the optimized audio content, and the optimized text content meets the deviation threshold; if so, the first operation is executed; otherwise, the fourth operation is executed.
[0054] The third operation is to remove redundant content from the initial video processing result and obtain the video processing result.
[0055] The fourth operation is to perform calibration processing on the video footage, the audio content, and the text content based on the narrative framework, and obtain optimized video footage, optimized audio content, and optimized text content respectively.
[0056] Secondly, to achieve the above objectives, the present invention provides a video processing device based on multimodal data, comprising: a data analysis module, an event and narrative processing module, a highlight extraction module, and a video processing module;
[0057] The data analysis module is used to obtain multimodal data analysis results using the multimodal raw data of the video to be processed;
[0058] The event and narrative processing module is used to process events and narratives based on the multimodal data analysis results, and to obtain event processing results and narrative frameworks.
[0059] The highlight extraction module is used to extract highlight segments based on the multimodal data analysis results and the event processing results, and to obtain video highlight segments.
[0060] The video processing module is used to obtain video processing results based on the multimodal data analysis results, the narrative framework, and the video highlight segments.
[0061] Thirdly, to achieve the above objectives, the present invention provides an electronic device, comprising: one or more processors; and a storage device having stored one or more programs thereon, which, when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation of the first aspect.
[0062] Fourthly, to achieve the above objectives, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by one or more processors, implements the method as described in any implementation of the first aspect.
[0063] Compared with the closest existing technology, the present invention has the following advantages:
[0064] This invention utilizes multimodal data acquisition and structured analysis techniques to transform raw data such as video frames, audio, and text into standardized analysis results such as emotional temporal data and event segmentation results. This effectively solves the problem of fragmented utilization of multimodal data in traditional technologies, maximizing the value of data linkage throughout the entire process. Through emotion-driven event assessment, LDA+BERT joint theme generation, and a three-act narrative structure design, this invention completes event processing and narrative framework construction, solving the problem of event-narrative disconnect and significantly improving the logical coherence and duration adaptability of video content. This invention employs high-value event range constraints and combines keyframe scoring, emotional changes, and multi-dimensional weighted filtering of visual features to extract highlight segments, avoiding the limitations of single feature extraction. This makes highlight segments more aligned with the core highlights and emotional resonance points of the video, enhancing content expressiveness. Based on image optimization, emotion-matched audio generation, subtitle synchronization, and multi-element temporal calibration integration, this invention achieves deep collaboration between video, audio, and text, solving the problem of poor integration of processing results. This significantly improves the visual experience, emotional delivery effect, and overall viewing smoothness of the final video. Simultaneously, the fully automated processing significantly reduces manual intervention costs and improves video creation efficiency. Attached Figure Description
[0065] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0066] Figure 1 This is a flowchart illustrating a video processing method based on multimodal data according to an embodiment of the present invention;
[0067] Figure 2 This is a schematic diagram of the structure of a video processing device based on multimodal data according to an embodiment of the present invention;
[0068] Figure 3 This is a schematic diagram of the structure of the electronic device proposed in an embodiment of the present invention. Detailed Implementation
[0069] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0070] The terminology used in the embodiments section of this invention is for the purpose of explaining specific embodiments of the invention only, and is not intended to limit the invention.
[0071] like Figure 1 As shown, this embodiment of the invention provides a video processing method based on multimodal data, including:
[0072] S1. Obtain multimodal data analysis results using the original multimodal data of the video to be processed;
[0073] The process involves collecting multimodal raw data from the video to be processed, and then preprocessing and analyzing it to obtain multimodal data analysis results. These results include sentiment time-series data, event segmentation results, topic clustering labels, and keyframe sequences. The core objective of this step is to transform unstructured raw multimodal data into standardized, reusable analysis results, providing data support for subsequent event processing, narrative construction, and highlight extraction.
[0074] S2. Perform event and narrative processing based on the multimodal data analysis results to obtain event processing results and narrative framework;
[0075] Based on the results of multimodal data analysis, the event segmentation results are first processed. This involves assessing the importance of events by combining sentiment time-series data and clustering similar events to obtain the event processing results. Secondly, a narrative framework is constructed based on the topic clustering labels and the event processing results: a joint model using the Latent Dirichlet Allocation (LDA) model and a pre-trained language model (BERT) is employed to generate topic logic. A three-act narrative structure is designed and adapted to the target duration, ensuring that events are arranged in an orderly manner according to the timeline and logical chain. The core purpose of this step is to transform fragmented events into a logical and structured narrative system, providing a clear direction for subsequent highlight extraction and content integration.
[0076] S3. Based on the multimodal data analysis results and the event processing results, extract highlight segments to obtain video highlight segments;
[0077] Based on multimodal data analysis results and combined with event processing results, highlight segments are extracted through multi-dimensional screening: Within the time range corresponding to high-value events, keyframe sequences are scored for content importance, and peak segments with "emotional change slope ≥ 0.6" in sentiment time series data are analyzed, along with visual feature analysis to select dynamic or focused scenes. The above multi-dimensional results are then weighted and ranked, with the top 30% of segments selected as video highlight segments. The core purpose of this step is to further filter out the segments that best reflect the core highlights and emotional resonance of the video from high-value events, thereby enhancing the final video's content expressiveness.
[0078] S4. Based on the multimodal data analysis results, the narrative framework, and the video highlight segments, obtain the video processing results;
[0079] This process integrates multimodal data analysis results, narrative frameworks, and video highlight clips for multi-element collaborative processing: optimizing video footage, generating matching audio based on emotional temporal data, and generating synchronized subtitles by combining topic tags and highlight clips; integrating the optimized video sequence, audio, and subtitles along the narrative framework's timeline, eliminating redundant content, and outputting a complete video processing result that conforms to narrative logic and emotional expression. The core objective of this step is to achieve deep collaboration between video, audio, and text, outputting a final product that conforms to narrative logic, emotional coherence, and provides an excellent viewing experience, completing the entire transformation from multimodal data to intelligent video.
[0080] In summary, steps S1 to S4 effectively solve the problems of multimodal data fragmentation, chaotic narrative logic, weak targeted highlight extraction, and poor audio-visual-text coordination in traditional video processing. They achieve deep data reuse and full-process automation, and the generated videos not only have a clear narrative structure and focus on core highlights, but also enhance user resonance through emotional integration throughout the entire process. At the same time, they significantly reduce the cost of manual intervention, improve generation efficiency compared to traditional editing methods, and can accurately adapt to the needs of scenarios such as Vlogs and short videos that require strong narrative and emotional transmission.
[0081] As one possible implementation, in the above embodiments, step S1 may specifically include the following steps:
[0082] S1-1. Acquire multimodal raw data of the video to be processed;
[0083] The scope of multimodal raw data contained in the video to be processed is clearly defined, specifically covering video frame sequences, audio signals, associated text information, and metadata. Video frame sequences are the continuous set of static images obtained after video decoding; each frame is image data containing pixel information, recording visual information such as human actions, scene environment, object shapes, and color distribution. Audio signals are the digitized sound wave data encapsulated in the video file, an audio stream with parameters such as sampling rate (e.g., 44.1kHz) and bit depth (e.g., 16-bit), containing sound information such as human voices, ambient sounds, and background music. Associated text information is unprocessed text information directly related to the video, including real-time transcription of speech recorded simultaneously during video shooting, original tags manually added by the photographer, unverified subtitles inherent in the video file, and shooting notes automatically recorded by the device. Metadata is the raw information describing video attributes, divided into technical metadata and content metadata: technical metadata includes original parameters generated by the device, such as shooting equipment model, video resolution, frame rate, encoding format, shooting timestamp, and file size; content metadata includes original content descriptions marked during shooting or associated with the device, such as shooting location, shooting object, and scene type.
[0084] After the data collection is completed, the raw data is classified and organized into four subsets: video, audio, text, and metadata. This classification clarifies the processing direction of different data types, provides a clear basis for subsequent targeted preprocessing, and ensures that subsequent processing steps can accurately match the characteristics of each data type.
[0085] S1-2. Preprocess the multimodal raw data of the video to be processed to obtain standardized multimodal data;
[0086] The multimodal raw data after this classification serves as input, and differentiated preprocessing operations are performed based on the characteristics of different data types. For video frame sequences, algorithms such as Gaussian filtering are used to remove image noise and improve frame clarity; for audio signals, spectral subtraction is used to eliminate background noise and retain effective sound information; for associated text information, operations such as format standardization, typo correction, and stop word removal are performed to achieve text standardization; for metadata, data integrity is verified and missing key information is supplemented, such as preliminary determination based on video footage when the scene type is not labeled.
[0087] Meanwhile, using the timestamps of video frames as a benchmark, audio signals, text information, and other data are aligned with the timeline of the video frames to ensure that all types of data correspond one-to-one in the time dimension. The final output is standardized multimodal data with a regular structure and synchronized time, laying the data foundation for subsequent feature analysis.
[0088] S1-3. Using the standardized multimodal data, obtain emotional time-series data through a multimodal emotion fusion algorithm;
[0089] This step relies on the obtained standardized multimodal data and focuses on the emotional dimension for analysis. First, speech and intonation features, such as speech rate, pitch, and volume changes, are extracted from the preprocessed audio signal. Then, emotional features, such as facial expressions, color saturation, and scene atmosphere, are extracted from the video frame sequence. Subsequently, a multimodal emotion fusion algorithm is used to weightedly fuse the emotional features extracted from the audio and video. Through a trained emotion recognition model, the emotional type of the video at different time points is determined, and the emotional intensity is quantified and output, with a value ranging from 0 to 1, where a higher value indicates a stronger emotion. Finally, emotional time-series data is formed, which includes the emotional intensity (0-1 points) and emotional type (such as "excitement," "calm," and "joy") that change over time. This data forms the basis for subsequent emotion-driven creation.
[0090] S1-4. Based on the standardized multimodal data, a video-audio collaborative segmentation strategy is adopted to obtain event segmentation results;
[0091] This step utilizes standardized data to achieve event-level segmentation of video content. First, scene transition detection algorithms (such as abrupt change detection based on inter-frame pixel differences and deep learning-based image content clustering) are used to detect abrupt changes in image features from the preprocessed video frame sequence. These abrupt changes include scene transitions, the appearance / disappearance of key objects, and significant changes in camera angle, serving as candidate points for image segmentation. Simultaneously, audio semantic segmentation algorithms are used from the preprocessed audio signal. Text content is extracted through speech recognition, semantic topic shifts are analyzed, and abrupt changes in audio features are detected, such as changes in the theme of the speech content and abrupt changes in background music style, serving as candidate points for audio segmentation. Combining the segmentation criteria of both image and audio, an image-audio collaborative segmentation strategy is adopted: when both image and audio segmentation candidate points appear at the same time point, they are directly identified as event boundaries; when only a single type of candidate point appears, the feature change trends of adjacent time points are combined to filter out the true event boundaries, determining the start and end timestamps of each independent event. Subsequently, the core information within each event time period is extracted: representative frames are selected from the video frames, key audio segments are extracted from the audio signal, and the core content description of the event is generated by fusion. Finally, an event segmentation result containing the event time range and content summary is formed, such as "Family dinner: 10:00-15:30", which clearly presents the event structure of the video and clarifies the start and end range of the event.
[0092] S1-5. Perform semantic analysis based on the standardized multimodal data to obtain topic clustering labels;
[0093] This step focuses on semantic theme analysis of the video content, using preprocessed related text information as the core input data. The standardized text information is divided into semantic units based on sentence punctuation, and meaningless interjections and repetitive content are removed before being input into the LDA topic clustering model. The model performs probabilistic clustering of semantic units by pre-setting the number of topics: first, it calculates the probability that each semantic unit belongs to a different topic; then, semantic units with a probability ≥ 0.6 are grouped into the same topic. Through topic keyword extraction, a global theme covering the entire video content is first determined, such as "family dinner"; then, based on the global theme, the text is further divided into different segments according to time sequence, and the above clustering process is repeated for each segment to further subdivide local sub-themes corresponding to the event, such as "food tasting" and "game interaction"; finally, topic clustering labels containing the global theme and local sub-themes are generated, clearly presenting the thematic hierarchy and core direction of the video content, facilitating subsequent narrative organization around the theme.
[0094] S1-6. Use the standardized multimodal data to extract keyframes and obtain a keyframe sequence;
[0095] For the preprocessed video frame sequence, inter-frame similarity calculation algorithms, such as histogram similarity and structural similarity (SSIM), are used to calculate the content similarity of adjacent frames. Frames with a similarity ≥ 85% are considered redundant and removed. For the remaining non-redundant frames, scores are assigned based on the weight of core elements (such as key figures, iconic actions, and important objects): frames featuring key figures (such as speakers or main characters) have a weight of 0.3; frames featuring iconic actions / objects (such as product demonstrations or ceremonial segments) have a weight of 0.3; frames corresponding to emotional peaks (such as cheers or smiles) have a weight of 0.2; and frames with high thematic relevance (such as scenes containing thematic keywords) have a weight of 0.2. The scores for each frame are calculated comprehensively, and the top 30% of frames that represent the key content of the video are selected to form a keyframe sequence, especially important source material for highlight segments. Each keyframe is bound to a corresponding timestamp and a core content description.
[0096] S1-7. Based on the emotional time series data, the event segmentation results, the topic clustering labels, and the keyframe sequence, obtain the multimodal data analysis results;
[0097] This approach involves multi-dimensional correlation between sentiment time-series data, event segmentation results, topic clustering tags, and keyframe sequences. In the time dimension, it ensures accurate matching of sentiment data, event information, topic tags, and keyframes at the same timestamp. In the content dimension, it strengthens data correlation through keyword association. The result is a structured multimodal data analysis, with a timeline as its core framework, integrating four main data categories: sentiment time-series, event segmentation, topic clustering, and keyframes. Each category includes specific content, time range, and relationships, providing comprehensive, collaborative, and directly accessible data support for subsequent video processing stages such as event processing, narrative framework construction, and highlight segment extraction.
[0098] In summary, steps S1-1 to S1-7 begin with the multimodal raw data of the video to be processed. First, the raw data is categorized by type and subjected to differentiated preprocessing to achieve data standardization and timeline alignment. Then, based on the standardized data, sentiment time-series data, event segmentation results, and topic clustering labels are generated, while keyframe sequences are extracted simultaneously. Finally, the four types of data are integrated and correlated according to time and content dimensions to form structured multimodal data analysis results. This process ensures data quality and temporal synchronization through classification preprocessing and overcomes the limitations of a single data dimension through multidimensional collaborative analysis. It not only solves the problems of fragmented and weakly correlated multimodal raw data, improving the accuracy of sentiment recognition, event segmentation, and topic clustering, but also provides comprehensive and collaborative data support for subsequent video event processing, narrative construction, and highlight extraction, effectively ensuring the logic and efficiency of the entire subsequent video processing workflow.
[0099] As one possible implementation, in the above embodiments, step S2 may specifically include the following steps:
[0100] S2-1. Based on the emotional time series data, evaluate the importance of the event segmentation results to obtain the event importance evaluation results;
[0101] Using sentiment time-series data and event segmentation results from multimodal data analysis as input, importance assessment is achieved through "sentiment-event association mapping." Specifically, the time interval of each event in the event segmentation results is first aligned with the time axis of the sentiment time-series data, and sentiment features within the corresponding time range of each event are extracted, including the mean sentiment intensity, the peak sentiment intensity, and the duration of sentiment intensity. Then, through weighted calculations, such as the mean sentiment intensity accounting for 40%, the peak sentiment intensity accounting for 30%, and the duration of sentiment intensity accounting for 30%, an importance score (0-1 point) for each event is obtained. Finally, an event importance assessment result containing "event identifier - event content - importance score" is formed, providing a quantitative basis for subsequent event selection.
[0102] S2-2. Based on the event importance assessment results and the emotional intensity threshold, high-value events are selected.
[0103] This step uses the event importance assessment results as the core basis and filters events by setting an importance threshold for emotional relevance. First, based on the emotional expression needs of the target video scenario (such as Vlog, short video), an emotional intensity threshold is set, usually 0.7, which can be dynamically adjusted according to the scenario. Second, all events in the event importance assessment results are traversed, retaining high-value events with emotional intensity ≥ the threshold and eliminating transitional or irrelevant events with emotional intensity < the threshold. Finally, the retained events are initially organized to form a high-value event set containing an event list, the core content of each event, and the original time interval, focusing on the core content in the video that is emotionally prominent and has high user attention.
[0104] S2-3. Use a pre-trained language model to perform content comparison on the high-value events and obtain the event processing results;
[0105] Using high-value events as input, this approach leverages the semantic understanding capabilities of pre-trained language models (such as BERT) to achieve event redundancy removal and optimization. First, the core content description of each high-value event is transformed into a vector representation. Second, the semantic similarity between event content vectors is calculated using the BERT model, setting a similarity threshold of 0.8. Duplicate or highly similar events are identified and merged, such as merging multiple "family photo" events into a "group photo" event. Simultaneously, the merged event content undergoes semantic refinement, supplementing key information such as the total duration and core scenarios of the merged events. Ultimately, this results in a compact, non-redundant event processing result containing a list of high-value events and their relationships (chronological order / logical causality), optimizing the compactness of the event content and improving the efficiency and logic of subsequent narrative construction.
[0106] S2-4. Based on the event processing results and the topic clustering labels, a joint model is used to perform correlation analysis to obtain the event correlation relationships;
[0107] This step takes the topic clustering labels from the time processing results and multimodal data analysis results as input, and uses a joint LDA+BERT model to achieve topic matching and association analysis. First, the LDA model is used to determine the global topic and topic keywords of the video; second, the BERT model is used to calculate the semantic similarity between each high-value event and the global topic keywords, and events unrelated to the global topic are eliminated, such as the "office work" event interspersed in a travel video; at the same time, based on topic association analysis, the logical relationship between high-value events is analyzed, and finally the event association relationship of "high-value event-global topic matching degree" and "logical association between events (causality / parallel / temporal)" is output to ensure the logical consistency between events and topics, and between events themselves.
[0108] S2-5. Based on the event segmentation results and the event relationships, construct an event timeline;
[0109] Using the original timestamps and event relationships from the event segmentation results as input, an event timeline that combines temporality and logic is constructed. First, the original start and end timestamps of each high-value event in the event segmentation results are extracted, and the events are initially arranged in chronological order. Second, combined with the logical relationships in the event relationships, the temporal order of the initially arranged events is optimized, and transition nodes between events are added, such as the connection time between the end of the "mountain climbing" event and the start of the "viewing from the mountaintop" event. Finally, with time as the horizontal axis, the start and end times, core content, and logical relationship markers between each high-value event are marked, forming a structured event timeline that intuitively presents the temporal flow and logical context of the events.
[0110] S2-6. Based on the target video length, construct a three-act narrative structure;
[0111] This step, constrained by the target video length (set according to the application scenario), such as the golden length of 3-5 minutes for vlogs and short videos, designs a three-act narrative framework (beginning-development-ending) that aligns with user viewing habits. First, allocate the time percentages of the three acts according to the target length, typically 10%-15% for the beginning, 60%-70% for development, and 15%-20% for the ending. For example, a 5-minute video would have a 0.5-0.75 minute beginning, 3-3.5 minutes development, and 0.75-1 minute ending. Second, clarify the functional positioning of each act: the beginning is responsible for scene introduction and background setup, such as the "departure scene" in a travel video; development focuses on the presentation of core events and emotional progression, such as high-value events like "mountain climbing and sightseeing" or "camping experience"; the ending serves as a conclusion and emotional sublimation, such as "travel reflections" or "group photo." Finally, a three-act narrative structure is formed, including "three-act time allocation - functional positioning of each act," providing a framework constraint for subsequent event allocation.
[0112] S2-7. Obtain the narrative framework based on the high-value events, the event timeline, and the three-act narrative structure;
[0113] This step involves collaborative matching using high-value events, an event timeline, and a three-act narrative structure as input. First, based on the chronological logic of the event timeline and the functional positioning of the three-act structure, high-value events are assigned to corresponding narrative acts, such as assigning "departure scene" to the beginning, "mountain climbing" and "camping" to the development, and "group photo" to the ending. Second, considering the duration proportion of each act, the duration of events allocated to each act is adjusted, such as compressing the duration of events at the beginning and extending the duration of core events in the development stage, ensuring that the duration of each act conforms to the preset proportion. Simultaneously, the narrative rhythm is optimized through the design of "introduction, development, climax, and conclusion" nodes, such as setting emotional climax nodes in the development stage. Finally, information such as "three-act structure - event allocation - duration adjustment - rhythm nodes" is integrated to form a complete narrative framework that combines chronological logic, thematic consistency, and duration adaptability, providing clear narrative guidance for subsequent video content integration.
[0114] In summary, steps S2-1 to S2-7, using sentiment time-series data, event segmentation results, and topic clustering labels from multimodal data as core inputs, systematically transform fragmented events into a structured narrative scheme through a progressive processing approach: "sentiment quantification to assess event importance → threshold screening to focus on high-value events → pre-trained model to merge redundant events → joint model to calibrate topic matching degree → combining time-series and logic to construct an event timeline → designing a three-act structure adapted to the target duration → integrating events and structure to form a narrative framework." This process solves the problem of ambiguous value judgments in traditional event processing by driving event screening with sentiment data; it avoids event redundancy and topic disconnect by optimizing event relevance and topic matching degree with pre-trained and joint models; and it constructs a narrative framework by combining time-series logic and a three-act structure, ensuring both the logic and rhythm of the narrative while adapting to the target video duration and user viewing habits. Ultimately, it provides accurate and efficient narrative guidance for subsequent video content integration, significantly improving the content coherence, emotional expression effectiveness, and scene adaptability of the generated video.
[0115] As one possible implementation, in the above embodiments, step S3 may specifically include the following steps:
[0116] S3-1. Using the event processing results, determine the extraction range of the highlight fragment;
[0117] The extraction boundary of highlight segments is determined based on the event processing results, which include a list of high-value events selected after importance assessment and the corresponding time interval information for each event. Since high-value events are key carriers of the core content and emotional highlights of the video, the extraction scope prioritizes the video event intervals corresponding to these high-value events. Specifically, the start and end timestamps of the high-value events recorded in the event processing results are mapped to the timeline of the video to be processed, clarifying the time boundary of the extraction scope. Simultaneously, the time intervals corresponding to low-value events (such as transitional scenes and irrelevant redundant content) marked in the event processing results are directly excluded to avoid invalid content interfering with highlight selection. Ultimately, a highlight segment extraction scope centered on the time intervals of high-value events is formed, ensuring that subsequent extraction work focuses on the core content area of the video.
[0118] S3-2. Based on the multimodal data analysis results and the extraction range of the highlight fragment, perform a multi-dimensional evaluation to obtain multi-dimensional evaluation results;
[0119] This step extracts candidate segments within the defined highlight segment extraction range. Combining the results of multimodal data analysis, the candidate segments are quantitatively evaluated from three dimensions: keyframe scoring (content importance), sentiment change analysis (sentiment peak), and visual feature analysis (motion intensity and visual focus). Finally, the independent evaluation score of each candidate segment in each dimension is obtained, forming a multi-dimensional evaluation result.
[0120] S3-3. The multi-dimensional evaluation results are weighted and fused to obtain a comprehensive evaluation result;
[0121] This step transforms multi-dimensional independent evaluation results into a unified quantitative indicator through weighted fusion. The core is to assign weights based on the contribution of each dimension to the "highlight attribute." First, the evaluation results of each dimension are standardized, such as mapping scores from different dimensions to a uniform 0-1 range to eliminate dimensional differences. Then, coefficients are assigned based on the weights of each dimension's influence on the "highlight segment"—weights are set based on video content characteristics and user focus, for example, keyframe score (reflecting content coreness) accounts for 30%, emotional change slope (reflecting emotional resonance) accounts for 30%, and visual feature score (reflecting visual appeal) accounts for 40%. Next, a linear weighted summation method is used to multiply the evaluation scores of each dimension of each candidate segment by their corresponding weights and then sum them to obtain the comprehensive evaluation score for that candidate segment. Finally, a comprehensive evaluation result is formed, with each candidate segment as a unit, including the comprehensive score of each segment, achieving a unified quantitative ranking of the "highlight degree" of candidate segments.
[0122] S3-4. Based on the extraction range of the highlight segments and the comprehensive evaluation results, obtain the video highlight segments;
[0123] First, based on a defined extraction scope, ensure that all selected segments correspond to high-value events, excluding low-value content outside the scope. Second, based on the comprehensive evaluation results, select segments from high to low based on their comprehensive scores—selection rules can be set, such as selecting the top 30% of high-scoring segments, or selecting segments with a comprehensive score ≥ 0.8. At the same time, verify the temporal continuity and content integrity of the selected segments to avoid content breaks caused by fragmentation of high-scoring segments. If necessary, merge adjacent high-scoring short segments. Finally, output a set of video highlight segments containing the time interval, content summary, and comprehensive score of each highlight segment, completing the entire highlight extraction process.
[0124] In summary, steps S3-1 to S3-4 first use the event processing results as a basis to define the time interval corresponding to high-value events as the scope for highlight segment extraction, avoiding ineffective extraction of low-value content. Then, within this scope, multimodal data analysis results are combined to conduct targeted evaluations from dimensions such as keyframe scoring, emotional changes, and visual features, obtaining multi-dimensional independent results. Subsequently, through standardization and weight allocation, the multi-dimensional results are weighted and fused into a comprehensive evaluation score and ranked. Finally, combining the extraction scope constraints and comprehensive score filtering, the video highlight segments are obtained. This process, through the progressive logic of "scope focusing - multi-dimensional evaluation - weighted fusion - precise filtering," effectively improves the accuracy and efficiency of highlight segment extraction, ensuring that the extracted highlight segments not only focus on the core content of the video but also consider emotional resonance and visual appeal, outputting high-quality core materials for subsequent video processing, thereby enhancing the content expressiveness and user viewing experience of the final video product.
[0125] As one possible implementation, in the above embodiments, step S3-2 may specifically include the following steps:
[0126] S3-2-1. Using the extraction range of the highlight fragment, determine the candidate fragments for multi-dimensional evaluation;
[0127] Using a pre-defined range for extracting highlight segments as boundaries, candidate segments for multi-dimensional evaluation are extracted from the video to be processed. Specifically, the start and end timestamps of the extraction range are first defined, such as the cake-cutting event: 12:00-13:30. Then, based on the temporal continuity and scene integrity of the video content, the video within this time interval is divided into several continuous segments. During the segmentation, features such as scene transitions and audio rhythm changes can be considered to avoid disrupting the integrity of a single action or scene. For example, the "cake-cutting" process can be divided into three continuous candidate segments: "preparing to raise the knife - the moment of cutting - dividing the cake." Finally, a set of all candidate segments within the extraction range that have independent content is formed, providing specific objects for subsequent multi-dimensional evaluation.
[0128] S3-2-2. Based on the keyframe sequence and the candidate segments evaluated by the multi-dimensional assessment, a content importance score is obtained to acquire the keyframe score.
[0129] Based on the keyframe sequences from the multimodal data analysis results, content importance scoring is conducted for each candidate segment. First, the keyframes corresponding to each candidate segment are located—through timestamp matching, all keyframes whose timestamps fall within the candidate segment's time interval are selected. Second, the content importance of these keyframes is assessed, with evaluation indicators including the salience of core elements within the frame and their relevance to the video's theme, using a 0-1 scoring system. Finally, the average or weighted average scores of all keyframes within a single candidate segment are taken, and this average is used as the keyframe score for that candidate segment, reflecting the core nature of the segment's content.
[0130] S3-2-3. Based on the emotional time series data and the candidate segments evaluated by the multi-dimensional method, obtain the emotional change slope;
[0131] Based on the sentiment time-series data from multimodal data analysis, the sentiment change slope for each candidate segment is calculated. First, according to the start and end timestamps of the candidate segments, sentiment intensity data points within the corresponding time interval are extracted from the sentiment time-series data. For example, the "cutting cake" candidate segment corresponds to 12:00-12:10, and the sentiment intensity value is extracted every 0.5 seconds within this interval. Second, a sentiment change curve is constructed with time as the horizontal axis and sentiment intensity as the vertical axis. The slope of this curve is calculated through linear fitting or differencing—the larger the absolute value of the slope, the more drastic the sentiment change within the candidate segment. For example, the sentiment intensity increases from 0.6 to 0.95 from "anticipation" to "excitement," with a slope of 0.82, indicating significant sentiment fluctuation. Finally, the calculated slope value is used as the sentiment change evaluation result for the candidate segment, i.e., the sentiment change slope.
[0132] S3-2-4. Based on the candidate segments evaluated in the multi-dimensional assessment, evaluate the motion intensity and visual intersection to obtain visual feature scores;
[0133] This step evaluates the visual appeal of candidate segments from two dimensions: "dynamic richness" and "visual focus," and integrates them into a visual feature score. For motion intensity evaluation, the optical flow density within the candidate segment is calculated—the vector magnitude of pixel motion between adjacent frames. A higher optical flow density value indicates richer dynamics; for example, the "family cheering" candidate segment has an optical flow density of 0.85, indicating strong dynamism. For visual focus evaluation, an attention heatmap algorithm is used to detect the visual focus area in each frame within the segment, scoring it based on the salience and stability of the focus area (0-1 points). Finally, the scores from both dimensions are weighted and summed according to preset weights, such as motion intensity accounting for 40% and visual focus accounting for 60%, to obtain the visual feature score for each candidate segment, comprehensively reflecting the visual appeal of the image.
[0134] S3-2-5. Obtain multi-dimensional evaluation results based on the keyframe score, the emotional change slope, and the visual feature score;
[0135] This step integrates the keyframe scores, sentiment change slopes, and visual feature scores obtained from the first three steps to form a structured, multi-dimensional evaluation result. First, the results for each dimension are standardized—the sentiment change slope is mapped to the 0-1 range using a normalization algorithm to ensure consistency with the dimensions of the keyframe scores and visual feature scores. Second, evaluation result entries are created for each candidate segment according to the dimensions of "keyframe score - sentiment change slope - visual feature score." Each entry includes the time interval of the candidate segment, the standardized scores for the three dimensions, and the original scores. Finally, a multi-dimensional evaluation result set is formed, based on candidate segments and covering content coreness (keyframe score), sentiment fluctuation (sentiment change slope), and visual appeal (visual feature score), providing a complete basis for subsequent weighted fusion and selection of highlight segments.
[0136] In summary, steps S3-2-1 to S3-2-5, based on the highlight segment extraction range, first extract candidate segments with independent content from the video to be processed within this range, achieving precise focus on the evaluation object. Then, for each candidate segment, a keyframe score reflecting the core content is calculated using the keyframe sequence; the slope of emotional change reflecting emotional fluctuation is obtained based on emotional temporal data; and visual feature scores representing visual appeal are obtained through motion intensity and visual focus analysis. Finally, the evaluation results of the three dimensions are standardized and integrated into a structured multi-dimensional evaluation result. This process, through the logic of "range constraint - multi-dimensional decomposition and evaluation - standardized integration," avoids ineffective evaluation of low-value content and comprehensively quantifies the value of segments from three key dimensions: content coreness, emotional resonance, and visual appeal. It solves the accuracy problem caused by the reliance on a single feature in traditional highlight extraction, significantly improving the comprehensiveness and accuracy of highlight segment evaluation. This provides a reliable and standardized evaluation basis for subsequent weighted screening of high-quality highlight segments that meet user needs, thereby ensuring the content expressiveness and user viewing experience of the final video processing result.
[0137] As one possible implementation, in the above embodiments, step S3 may specifically include the following steps:
[0138] S4-1. Based on the narrative framework and the event processing results, extract the core transition segments;
[0139] The narrative framework clarifies the overall structure of the video and the logical arrangement of each high-value event, while the event processing results include selected high-value events and low-value transitional events. During extraction, priority is given to selecting transitional content in the event processing results that has a direct logical connection with the high-value events. For example, between the two high-value events "Visiting Scenic Spot A" and "Food Experience," the transitional segment "On the way to the restaurant" in the event processing results is selected. At the same time, it is ensured that the duration of the extracted core transitional segments is appropriate to the narrative rhythm, usually 5-15 seconds per segment, and the content is related to the theme of adjacent high-value events. This ultimately forms a set of core transitional segments used to connect the highlight segments of the video and ensure narrative coherence.
[0140] S4-2. Based on the emotional time series data and the topic clustering labels, combined with the narrative framework, the video highlight segments and the core transition segments are integrated to obtain video footage;
[0141] Emotional timing data sets the tone for the visual atmosphere; for example, highlight clips with an emotion intensity ≥ 0.8 use highly saturated warm colors, while core transition clips with an emotion intensity < 0.5 use low-saturation soft colors. Thematic clustering labels unify the visual style; for example, under the theme "Travel - Natural Scenery," all clips are calibrated to a natural and fresh color tone. The narrative framework guides the technical processing of details: highlight clips are prioritized for spatiotemporal super-resolution adaptive reconstruction (upgrading resolution to 1080P) and image stabilization (eliminating handheld shooting shake), while core transition clips focus on transition adaptation, such as using "fade-in" and "fade-out" for transitions between high-value events and "slide transitions" for transitions within continuous scenes, ensuring that the transition effects match the narrative logic. Through the above integration and optimization, scattered highlight clips and core transition clips are transformed into a video sequence with a unified style, clear image quality, and smooth transitions.
[0142] S4-3. Based on the emotional time sequence data and the video footage, combined with the narrative framework, a dynamic time adjustment algorithm is used to obtain the audio content;
[0143] Emotional temporal data is directly mapped to audio parameters. Specifically, an emotional intensity ≥ 0.8 is matched with high-volume, fast-paced, energetic music, such as the highlight segment "Sunrise Cheers"; an emotional intensity between 0.5 and 0.8 is matched with medium-volume, medium-paced, soothing music, such as the core transition segment "Strolling Through the Scenic Area"; and an emotional intensity < 0.5 is matched with low-volume, slow-paced, gentle music, such as the transition segment "Night Rest". The rhythm of the video footage determines the audio details. By analyzing the editing rhythm of the footage, such as the rapid switching of highlight segments and the slow motion of transition segments, the music's tempo and melody fluctuations are adjusted. A Dynamic Time Adjustment (DTW) algorithm is used to achieve precise alignment between the audio rhythm and the intensity of the video motion. The narrative framework standardizes the overall audio style. For example, the audio for a "Travel Vlog" narrative is based on upbeat folk music, while the "Birthday Party" narrative is dominated by warm and cheerful melodies. High-value ambient sounds are extracted from multimodal data analysis results for sound effect enhancement, ultimately forming audio content that highly matches the emotional content, rhythm, and narrative style of the footage.
[0144] S4-4. Based on the topic clustering tags, the video highlight segments, and the core transition segments, combined with the narrative framework, obtain the text content;
[0145] Thematic clustering labels clearly define the text style direction. For example, under the theme of "family gathering," the text uses friendly, conversational language; while the theme of "academic conference" uses formal, written language. Highlight clips and key transition clips in the video provide text material. For highlight clips, core elements within the frame are extracted to generate core descriptions; for key transition clips, scene information is extracted to generate connecting text. The narrative framework ensures logical coherence. Following a "beginning-development-ending" narrative structure, each segment's text is given contextual relevance. For example, the transition clip text in the beginning section focuses on "scene introduction," while the highlight clip text in the development section focuses on "event details." Simultaneously, using the GPT-4 model, subtitle text is generated based on the above information, and the subtitle timestamps are aligned with the keyframe times of the corresponding clips (deviation ≤ 0.5s) to ensure that the text content both fits the visuals and conforms to the logical progression of the narrative, ultimately forming a complete subtitle text system.
[0146] S4-5. Using the video footage, the audio content, and the text content, obtain the video processing result;
[0147] This step integrates multiple elements to transform video footage, audio content, and text content into a complete video. Using the narrative timeline as a benchmark, the timestamps of the video footage, audio content, and text content are calibrated and optimized (deviation ≤ 0.3s). These three elements are then integrated according to the narrative logic to form a preliminary video. Redundant content is removed, and the video is adapted to the target duration before outputting the complete video processing result.
[0148] In summary, steps S4-1 to S4-5 of the video processing are guided by the narrative framework. First, core transition segments are extracted based on the event processing results to ensure narrative continuity. Then, by combining emotional temporal data and theme clustering tags from multimodal data, the highlight segments and core transition segments of the video are styled, their image quality is improved, and transitions are adapted. Subsequently, matching audio is generated based on emotion and visual rhythm, and logically coherent synchronous text is created around the theme and visual content. Finally, the optimized visuals, audio, and text are integrated according to narrative logic and timeline, and redundancy is eliminated to form a complete video processing result. This process achieves deep synergy between multimodal data and narrative needs. It not only solves the problems of fragmented visuals and disconnect between audio, visuals, and text in traditional processing, but also significantly improves the narrative coherence, audiovisual harmony, and content expressiveness of the video through emotion-driven multi-element adaptation and narrative logic. At the same time, the automated integration process reduces the cost of manual intervention and efficiently outputs high-quality video results that meet the needs of different scenarios (such as Vlogs and short videos).
[0149] As one possible implementation, in the above embodiments, step S4-5 may specifically include the following steps:
[0150] S4-5-1. Based on the narrative framework, the video frame, the audio content, and the text content are calibrated to obtain optimized video frame, optimized audio content, and optimized text content, respectively.
[0151] This step uses the narrative framework as the core calibration basis and conducts targeted optimization and calibration of video footage, audio content, and text content, as detailed below:
[0152] (1) For video footage, based on the style and rhythm of the narrative framework and timeline, the colors of the footage are calibrated a second time to unify the visual style, the duration of transition effects is adjusted to match the narrative rhythm, and super-resolution processing is performed on blurred frames to form optimized video footage.
[0153] (2) For audio content, combine the emotional direction of the narrative framework, such as a gentle beginning and an exciting development, calibrate the background music volume level, increase the volume of the highlight segments by 10%-15%, adjust the clarity of the ambient sound effects, and ensure that the audio style matches the narrative stage to obtain optimized audio content.
[0154] (3) For the text content, according to the logical context of the narrative framework, correct the relevance between the subtitle content and the context, such as supplementing the connecting subtitles for transitional segments, calibrating the font style and size of the subtitles to match the narrative scene, such as using a rounded font for a warm scene, and generating optimized text content.
[0155] By independently calibrating the three components, we ensure that the frame sequence of the optimized video, the waveform signal of the audio content, and the subtitle display time of the text content are completely synchronized, laying the foundation for subsequent synchronization and integration.
[0156] S4-5-2. Determine whether the timestamps corresponding to the optimized video frame, the optimized audio content, and the optimized text content are synchronized. If yes, proceed directly to S4-5-4; otherwise, proceed to S4-5-3.
[0157] This step is the initial assessment of timestamp synchronization, focusing on the timestamp matching degree of key nodes in the optimized video, audio, and text. Core time nodes in the narrative framework are selected, such as the start / end points of highlight clips, transition points, and core event trigger points. The timestamps of the corresponding frames in the optimized video, the waveform peak timestamps in the optimized audio, and the subtitle display / hide timestamps in the optimized text are extracted. The timestamps of the three at the same core node are compared to see if they are completely consistent. For example, in the highlight clip "Sunrise cheers at a scenic spot," the video frame timestamp is 1:30, the audio peak timestamp is 1:30, and the subtitle display timestamp is 1:30. If the timestamps of all core nodes are completely synchronized, the time coordination of the three meets the standard, and the integration stage can proceed directly. If any core node timestamp is mismatched, the next step, deviation threshold judgment, is performed.
[0158] S4-5-3. Determine whether the difference between the timestamps corresponding to the optimized video frame, the optimized audio content, and the optimized text content meets the deviation threshold. If yes, execute S4-5-4; otherwise, return to execute S4-5-1.
[0159] This step is a timestamp deviation tolerance check. A preset deviation threshold, typically set to ≤0.3s, is used to determine if the timestamp difference is within an acceptable range. The maximum difference between the three timestamps at the desynchronized core nodes is calculated. For example, if the video frame is 1:30, the audio is 1:30.2, and the subtitle is 1:30.1, the maximum difference is 0.2s. If this difference is ≤ the preset deviation threshold, it means the deviation has no significant impact on the audiovisual experience, and no recalibration is needed; the integration phase can proceed. If the difference is > the deviation threshold, it indicates that the time deviation will cause problems such as audio-visual desynchronization and subtitle misalignment. Recalibration needs to be performed based on this deviation data until the timestamp difference meets the tolerance requirements.
[0160] S4-5-4. Integrate the optimized video footage, the optimized audio content, and the optimized text content to obtain the initial video processing result;
[0161] Based on the timeline of the narrative framework, the frame sequence of optimized video footage is used as the basic carrier, and the audio track of optimized audio content is embedded to ensure that the background music, environmental sound effects and the rhythm of the video are matched. The subtitle track of optimized text content is also superimposed to ensure that the subtitles are displayed synchronously with the core elements of the video and the key information of the audio. At the same time, following the structural order of the narrative framework (beginning-development-ending), the segments are linked together according to the logical chain to form the initial video processing result containing complete video, audio and subtitles.
[0162] S4-5-5. Redundant content is removed from the initial video processing result to obtain the video processing result;
[0163] This step refines the initial video processing results to improve narrative coherence and content quality. Based on the core requirements of the narrative framework, three types of redundant content are selected and eliminated: first, segments that conflict with the narrative theme, such as irrelevant work footage inserted into a travel vlog; second, content with inconsistent emotions, such as chaotic transitions suddenly appearing in heartwarming scenes; and third, segments with excessive length, such as repetitive transitions exceeding the target duration or excessively long, informationless blank segments. The elimination process must maintain the integrity of the narrative logic, such as not deleting key transition segments, ensuring that the remaining content closely adheres to the narrative framework, ultimately outputting a complete video processing result that is appropriately sized, concise, coherent, and integrates audio, visuals, and text.
[0164] In summary, steps S4-5-1 to S4-5-5, through progressive processing, not only effectively solve the problems of audio-visual and text time misalignment and content disorder in traditional video integration, ensuring that the time synchronization deviation of the three is controlled within a reasonable range, but also significantly improve the narrative coherence and audiovisual coordination of the video; at the same time, the closed-loop calibration mechanism ensures the stability of processing quality, and the redundancy removal step further enhances the focus of content, ultimately achieving the technical effect of efficiently outputting high-quality video processing results that meet narrative needs and have stable quality.
[0165] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a video processing apparatus based on multimodal data, which is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0166] like Figure 2 As shown, a video processing device based on multimodal data in this embodiment includes: a data analysis module, an event and narrative processing module, a highlight extraction module, and a video processing module;
[0167] The data analysis module is used to obtain multimodal data analysis results using the multimodal raw data of the video to be processed;
[0168] The core function of this module is to transform the unstructured multimodal raw data of the video to be processed into standardized and reusable analysis results. Its input is the full-dimensional raw data of the video to be processed, specifically covering video frame sequences (including image texture, color distribution, and object motion trajectories), audio signals (including human voices, ambient sounds, rhythm and pitch features), associated text information (such as speech-to-text content and scene annotations during shooting), and metadata (such as shooting time, equipment parameters, and scene type). During processing, the module first performs preprocessing operations, including video frame denoising, audio noise reduction, and multimodal data timeline alignment to ensure consistency in timestamps for video, audio, and text; then, it generates four core results through specialized analysis algorithms:
[0169] (1) Emotional temporal data: Emotional temporal data with emotional intensity of 0-1 points and emotional type that change over time is generated by the fusion of "audio emotion recognition (voice tone analysis) + image emotion recognition (character expression, color atmosphere judgment)", such as "excitement" and "calm";
[0170] (2) Event segmentation results: By combining "scene switching detection + audio semantic segmentation", event segmentation results with clear start and end timestamps and core content descriptions for each independent event are generated;
[0171] (3) Topic clustering tags: Based on the LDA model, content semantic clustering generates topic clustering tags containing global topics and local subtopics;
[0172] (4) Keyframe sequence: The keyframe sequence containing representative frames with core characters and key actions is extracted and retained by “frame content similarity calculation + core element weight score”.
[0173] The final output of this module is structured multimodal data analysis results, which directly provide accurate data support for subsequent event and narrative processing, highlight extraction, and video processing modules, solving the problem of "fragmented and difficult-to-reuse" raw data.
[0174] The event and narrative processing module is used to process events and narratives based on the multimodal data analysis results, and to obtain event processing results and narrative frameworks.
[0175] This module takes the multimodal data analysis results output by the data analysis module as its sole input. Its core objective is to transform fragmented events into a logical and structured narrative system, and output the event processing results and narrative framework.
[0176] In the event processing stage, the module first filters all events in the event segmentation results based on the sentiment time series data—setting a sentiment intensity threshold (e.g., 0.7), retaining "high-value events" with sentiment intensity ≥ the threshold, and marking "low-value transitional events" with sentiment intensity < the threshold; then, it calculates the semantic similarity of event content through the BERT model, merges repeated or highly similar high-value events, and finally forms an event processing result containing "a list of high-value events and event relationships (chronological order / logical causality)".
[0177] In the narrative framework construction stage, the module first combines topic clustering tags and event processing results, and uses a joint "LDA+BERT" model to determine the topic logic. The LDA model locks in the global topic of the video, and the BERT model calibrates the semantic matching degree between high-value events and the global topic, eliminating events that are irrelevant to the topic. Then, based on user viewing habits, such as the 3-5 minute golden length of Vlogs and short videos, a "three-act narrative structure" (beginning-development-ending) is designed. The narrative rhythm is optimized by adjusting the length ratio of high-value events (core events account for ≥60%), and finally a narrative framework that combines logical coherence and length adaptability is formed.
[0178] The output of this module directly determines the core logic and presentation structure of subsequent video content, avoiding the problem of "chaotic and disorganized content" in videos.
[0179] The highlight extraction module is used to extract highlight segments based on the multimodal data analysis results and the event processing results, and to obtain video highlight segments.
[0180] This module takes the multimodal raw data from the data analysis module and the event processing results from the event and narrative processing module as dual inputs. Its core function is to accurately extract the most expressive and emotionally resonant highlight segments from the video. Its processing logic revolves around "range constraint + multi-dimensional evaluation": First, the "high-value event time interval" in the event processing results is used as the extraction range to directly exclude the interval where low-value events are located, reducing interference from invalid segments; Second, within this range, key features of multimodal data are used to conduct multi-dimensional evaluation: one is keyframe scoring, which calculates a score (0-1 point) based on the saliency of core elements within the frame (such as facial expressions and key props), and selects segments corresponding to keyframes with a score ≥0.85; the second is emotion change analysis, which calculates the "emotion change slope" of the emotion time series data within the interval, and retains emotional peak segments with a slope ≥0.6; the third is visual feature analysis, which calculates the motion intensity of the screen through optical flow density, and locates the visual focus area by combining attention heatmap, and selects segments with high motion intensity or prominent focus; Finally, the three types of evaluation results are weighted and fused according to "keyframe score 30% + emotion change slope 30% + visual feature score 40%", and the top 30% of segments are selected as video highlight segments according to the comprehensive score. The output of this module ensures that subsequent video processing results can focus on core highlights, enhancing the emotional resonance and content appeal for users.
[0181] The video processing module is used to obtain video processing results based on the multimodal data analysis results, the narrative framework, and the video highlight segments;
[0182] This module takes the multimodal data analysis results from the data analysis module, the narrative framework from the event and narrative processing module, and video highlight clips from the highlight extraction module as triple inputs. Its core function is to perform collaborative optimization and integration of "image-audio-text" to output a complete video processing result. The processing involves three steps:
[0183] The first step is multi-element specialized optimization—in terms of visual optimization, the video color style is unified based on theme clustering tags, spatiotemporal super-resolution adaptive reconstruction is performed on highlight clips, and transition effects are selected according to the relevance of events in the narrative framework; in terms of audio optimization, emotionally matched background music is generated by combining emotional temporal data, environmental sounds of high-value events are preserved and sound effects are enhanced, and the audio rhythm is aligned with the rhythm of the visual motion through the Dynamic Time Warping (DTW) algorithm; in terms of text optimization, subtitle text is generated based on theme tags and highlight clip content, using the GPT-4 model, and the subtitle timestamps are aligned with the keyframes of the highlight clips.
[0184] The second step is multi-element collaborative integration—integrating the optimized video sequence, audio track, and subtitle track according to the timeline of the narrative framework, ensuring that the timestamp deviation of the three is ≤0.3s.
[0185] The third step is to remove redundant content—deleting content that conflicts with the narrative framework or is emotionally inconsistent, ultimately generating a complete video that conforms to mainstream formats.
[0186] The output of this module marks a closed loop from "multimodal data input" to "intelligent video output", ensuring that the final video has logical coherence, emotional delivery, and visual appeal.
[0187] In summary, this device constitutes a collaborative video processing system encompassing "data analysis, logic construction, highlight selection, and final product integration," with each module progressing progressively and operating in a closed-loop data loop. The data analysis module transforms unstructured multimodal raw data into standardized analytical results, laying the data foundation; the event and narrative processing module constructs a logical event system and narrative framework based on the data, clarifying the content structure; the highlight extraction module combines data and event results to accurately select core highlight segments, focusing on content value; and the video processing module integrates previous results to complete the collaborative optimization and integration of visuals, audio, and text, outputting videos that combine narrative coherence, emotional delivery, and visual appeal, significantly improving the automation and efficiency of video processing.
[0188] In this embodiment, the specific processing of a video processing device based on multimodal data and the resulting technical effects can be referred to separately. Figure 1 The relevant descriptions of steps S1, S2, S3 and S4 in the corresponding embodiments will not be repeated here.
[0189] It should be noted that the implementation details and technical effects of each module and unit in the device provided in the embodiments of this disclosure can be referred to the description of other embodiments in this disclosure, and will not be repeated here.
[0190] The following is for reference. Figure 3 It shows a schematic diagram of the structure of a computer system 500 suitable for implementing the electronic device of the present disclosure. Figure 3 The computer system 500 shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0191] like Figure 3 As shown, the computer system 500 may include a processing unit 501, which can perform various appropriate actions and processes according to a program stored in ROM 502 or a program loaded into random access RAM 503 from storage device 508. RAM 503 also stores various programs and data required for the operation of the computer system 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. I / O interface 505 is also connected to bus 504.
[0192] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows computer system 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 A computer system 500 with various electronic devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0193] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0194] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0195] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0196] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the following functions: Figure 1 The embodiments shown and their alternative implementations illustrate a video processing method based on multimodal data.
[0197] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0198] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0199] The units or modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units or modules do not necessarily limit the unit itself; for example, an acquisition module can also be described as "acquiring preset prompts, including modality fusion prompts, attention mechanism prompts, and / or time-related prompts."
[0200] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
Claims
1. A video processing method based on multimodal data, characterized in that, include: S1. Obtain multimodal data analysis results using the original multimodal data of the video to be processed; S2. Perform event and narrative processing based on the multimodal data analysis results to obtain event processing results and narrative framework; S3. Based on the multimodal data analysis results and the event processing results, extract highlight segments to obtain video highlight segments, including: The importance of event segmentation results is evaluated based on sentiment time series data to obtain the event importance evaluation results; High-value events are selected by combining the event importance assessment results with an emotional intensity threshold. The high-value events are compared using a pre-trained language model to obtain the event processing results. Based on the event processing results and topic clustering tags, a joint model is used to perform correlation analysis to obtain the event correlation relationships; Based on the event segmentation results and the event relationships, an event timeline is constructed. Based on the target video length, a three-act narrative structure is constructed; Based on the high-value events, the event timeline, and the three-act narrative structure, obtain the narrative framework; The joint model includes a latent Dirichlet assignment model and a pre-trained language model. S4. Based on the multimodal data analysis results, the narrative framework, and the video highlight segments, obtain the video processing results; The multimodal data analysis results include sentiment time-series data, event segmentation results, topic clustering labels, and keyframe sequences. The topic clustering labels include global topics and local subtopics.
2. The video processing method based on multimodal data according to claim 1, characterized in that, S1. Using the multimodal raw data of the video to be processed, obtain the multimodal data analysis results, including: Collect multimodal raw data of the video to be processed, including video frame sequences, audio signals, associated text information and metadata; The multimodal raw data of the video to be processed is preprocessed to obtain standardized multimodal data; Using the standardized multimodal data, a multimodal sentiment fusion algorithm is used to obtain sentiment time-series data; Based on the standardized multimodal data, an image-audio collaborative segmentation strategy is adopted to obtain event segmentation results; Semantic analysis is performed on the standardized multimodal data to obtain topic clustering labels; Keyframes are extracted using the standardized multimodal data to obtain a keyframe sequence; Based on the emotional time-series data, the event segmentation results, the topic clustering labels, and the keyframe sequence, multimodal data analysis results are obtained.
3. The video processing method based on multimodal data according to claim 1, characterized in that, S3. Based on the multimodal data analysis results and the event processing results, extract highlight segments to obtain video highlight segments, including: Using the event processing results, the extraction range of the highlight fragment is determined; Based on the multimodal data analysis results and the extraction range of the highlight fragments, a multidimensional evaluation is performed to obtain multidimensional evaluation results; The multi-dimensional evaluation results are weighted and fused to obtain a comprehensive evaluation result; Based on the extraction range of the highlight segments and the comprehensive evaluation results, the highlight segments of the video are obtained.
4. The video processing method based on multimodal data according to claim 3, characterized in that, Based on the multimodal data analysis results and the extraction range of the highlight fragments, a multi-dimensional evaluation is performed to obtain multi-dimensional evaluation results, including: By utilizing the extraction range of the highlight fragments, candidate fragments for multi-dimensional evaluation are determined; Based on the keyframe sequence and the candidate segments evaluated in the multi-dimensional assessment, a content importance score is obtained to acquire the keyframe score. Based on the emotional time-series data and the candidate segments evaluated in the multi-dimensional assessment, the slope of emotional change is obtained; Based on the candidate segments evaluated in the multi-dimensional assessment, motion intensity and visual intersection are evaluated to obtain visual feature scores; A multi-dimensional evaluation result is obtained based on the keyframe score, the slope of the emotion change, and the visual feature score.
5. The video processing method based on multimodal data according to claim 1, characterized in that, S4. Based on the multimodal data analysis results, the narrative framework, and the video highlight segments, obtain the video processing results, including: Based on the narrative framework and the event processing results, extract the core transition segments; Based on the emotional time series data and the topic clustering tags, combined with the narrative framework, the video highlight segments and the core transition segments are integrated to obtain video footage; Based on the emotional timing data and the video footage, combined with the narrative framework, a dynamic time adjustment algorithm is used to obtain the audio content. Based on the topic clustering tags, the video highlight segments, and the core transition segments, combined with the narrative framework, the text content is obtained; The video processing result is obtained using the video footage, the audio content, and the text content.
6. The video processing method based on multimodal data according to claim 5, characterized in that, Using the video footage, the audio content, and the text content, the video processing result is obtained, including: Based on the narrative framework, the video footage, audio content, and text content are calibrated to obtain optimized video footage, optimized audio content, and optimized text content, respectively. Determine whether the timestamps corresponding to the optimized video, optimized audio, and optimized text are synchronized. If they are, perform the first operation; otherwise, perform the second operation. The first operation is to integrate the optimized video footage, the optimized audio content, and the optimized text content to obtain an initial video processing result, and then perform the third operation. The second operation is to determine whether the difference between the timestamps corresponding to the optimized video frame, the optimized audio content, and the optimized text content meets the deviation threshold. If yes, the first operation is executed; otherwise, the fourth operation is executed. The third operation is to remove redundant content from the initial video processing result and obtain the video processing result. The fourth operation is to perform calibration processing on the video footage, the audio content, and the text content based on the narrative framework, and obtain optimized video footage, optimized audio content, and optimized text content respectively.
7. A video processing apparatus based on multimodal data, employing the method as described in any one of claims 1-6, characterized in that, include: Data analysis module, event and narrative processing module, highlight extraction module, and video processing module; The data analysis module is used to obtain multimodal data analysis results using the multimodal raw data of the video to be processed; The event and narrative processing module is used to process events and narratives based on the multimodal data analysis results, and to obtain event processing results and narrative frameworks. The highlight extraction module is used to extract highlight segments based on the multimodal data analysis results and the event processing results, and to obtain video highlight segments. The video processing module is used to obtain video processing results based on the multimodal data analysis results, the narrative framework, and the video highlight segments.
8. An electronic device, characterized in that, include: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, It stores a computer program thereon, wherein the computer program, when executed by one or more processors, implements the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Video post-editing and video synthesis optimization method
CN116847123A
High-quality video content automatic generation method and related equipment
CN120050487A