Video processing method and device based on multi-modal data, equipment and medium
By using multimodal data analysis and constructing a three-act narrative framework, the problems of multimodal data fragmentation, emotional drive gaps, and poor narrative adaptability in video processing are solved, achieving efficient and intelligent generation of video content and improving the logical coherence and emotional resonance of the video.
Patent Information
- Application Number
- CN202511988023.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-12-26
AI Technical Summary
Existing video processing technologies suffer from low multimodal data utilization, lack of emotion-driven approaches throughout the entire process, poor narrative and scene adaptability, and weak synergy between technology combinations. These issues result in fragmented video content, inconsistent logic, discontinuous emotional expression, and an inability to accurately match user viewing habits.
By acquiring multimodal data and performing structured analysis, we obtain emotional time-series data, event segmentation results, topic clustering labels, and keyframe sequences, construct a three-act narrative framework, and combine highlight segment extraction and multi-element collaborative processing to achieve automated and intelligent generation of video content.
It improves the logical coherence and duration adaptability of video content, enhances the expressiveness and emotional resonance of the content, reduces the cost of manual intervention, and generates high-quality, smooth video content.
Smart Images

Figure CN121397299A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video processing, in particular to a video processing method and device based on multi-modal data, equipment and medium. BACKGROUND
[0002] With the popularity of short videos and Vlog content creation, users' demand for high-quality, story-based video content is growing. However, the existing video processing technology has the following shortcomings when facing this demand: (1) Low utilization rate of multi-modal data: Current video processing methods mostly use single data sources for independent processing, such as picture analysis based only on key frames or segment selection based only on event segmentation results. This processing method results in a lack of effective closed-loop correlation between different modal data, which cannot fully exploit the synergistic value of multi-modal data, thus making the generated video content often fragmented and logically incoherent, making it difficult to form a complete and smooth narrative whole.
[0003] (2) Emotion-driven not throughout the whole process: In the existing video creation process, the application of emotional data is often limited to a single link, such as simply referring to emotional tendencies in the audio music stage, without deep linkage with video content in key steps such as event filtering, highlight extraction, and narrative design. This partial and fragmented emotional application leads to a fault in the emotional expression of the video, which cannot continuously and effectively guide the audience's emotional experience, ultimately affecting the emotional resonance effect and appeal of the video.
[0004] (3) Poor adaptation of narrative and scene: The existing video narrative construction method mostly uses fixed narrative templates, such as the simple "beginning-middle-end" structure, without fully considering the "3-5 minute golden viewing time" restriction and "beginning, transition, and conclusion" narrative logic requirement of Vlog and other short video forms. This narrative method that does not adapt to specific scenes easily leads to redundancy or insufficient information in video content, which cannot accurately match the viewing habits and content expectations of the audience, thus reducing the attractiveness and dissemination effect of the video.
[0005] (4) Weak synergy of technology combination: In the key technical steps of video generation, existing solutions often rely on single algorithms for processing, such as using only the Latent Dirichlet Allocation (LDA) model for topic clustering or only using key frame scores for highlight segment filtering. This single algorithm application method cannot form a multi-algorithm synergistic technical solution according to the complex needs of video creation, limiting the space for improving video generation quality, and making it difficult to realize the complementary advantages and synergistic effect among technical modules.
[0006] Therefore, there is an urgent need for a video processing method, device, equipment and medium based on multi-modal data to solve the shortcomings in the prior art. SUMMARY
[0007] The present application aims to provide a video processing method and device based on multi-modal data, equipment and medium, to solve the problems of multi-modal data fragmentation, emotional driving fault, poor narrative adaptability and weak technical cooperation in existing video processing technology, by fusing multi-dimensional information such as emotion, event, theme and vision, constructing video content with narrative logic, realizing automatic and intelligent generation from original video to story video, and improving video content quality and production efficiency.
[0008] In a first aspect, to achieve the above-mentioned purpose, the present application provides a video processing method based on multi-modal data, comprising the following steps: S1, using multi-modal original data of a video to be processed, obtaining multi-modal data analysis results; S2, according to the multi-modal data analysis results, performing event and narrative processing to obtain event processing results and narrative framework; S3, based on the multi-modal data analysis results and the event processing results, extracting highlight segments to obtain video highlight segments; S4, according to the multi-modal data analysis results, the narrative framework and the video highlight segments, obtaining video processing results; Wherein, the multi-modal data analysis results include emotion time series data, event segmentation results, theme clustering labels and key frame sequences, and the theme clustering labels include global themes and local sub-themes.
[0009] Optionally, S1, using multi-modal original data of a video to be processed, obtaining multi-modal data analysis results; Collecting multi-modal original data of a video to be processed, the multi-modal original data including video frame sequences, audio signals, associated text information and metadata; Preprocessing the multi-modal original data of the video to be processed to obtain standardized multi-modal data; Using the standardized multi-modal data to obtain emotion time series data through a multi-modal emotion fusion algorithm; Based on the standardized multi-modal data, using a picture-audio collaborative segmentation strategy to obtain event segmentation results; According to the standardized multi-modal data, performing semantic analysis to obtain theme clustering labels; Using the standardized multi-modal data to extract key frames to obtain key frame sequences; According to the emotion time series data, the event segmentation results, the theme clustering labels and the key frame sequences, obtaining multi-modal data analysis results.
[0010] Optionally, S2, event and narrative processing is performed according to the multi-modal data analysis result, an event processing result and a narrative framework are obtained, and the event processing result is obtained by: Based on the emotion time series data, the importance of the event segmentation result is evaluated, and an event importance evaluation result is obtained; According to the event importance evaluation result, combined with the emotion intensity threshold, high-value events are screened and obtained; The content of the high-value event is compared by using a pre-trained language model, and an event processing result is obtained; Based on the event processing result and the theme clustering label, a joint model is used for correlation analysis, and an event correlation relationship is obtained; According to the event segmentation result combined with the event correlation relationship, an event timeline is constructed; Based on the target video duration, a three-act narrative structure is constructed; According to the high-value event, the event timeline and the three-act narrative structure, a narrative framework is obtained; The joint model includes a latent Dirichlet allocation model and a pre-trained language model.
[0011] Optionally, S3, based on the multi-modal data analysis result combined with the event processing result, a highlight segment is extracted, and a video highlight segment is obtained, including: The extraction range of the highlight segment is determined by using the event processing result; Based on the multi-modal data analysis result combined with the extraction range of the highlight segment, multi-dimensional evaluation is performed, and a multi-dimensional evaluation result is obtained; The multi-dimensional evaluation result is weighted and fused, and a comprehensive evaluation result is obtained; According to the extraction range of the highlight segment combined with the comprehensive evaluation result, a video highlight segment is obtained.
[0012] Optionally, based on the multi-modal data analysis result combined with the extraction range of the highlight segment, multi-dimensional evaluation is performed, and a multi-dimensional evaluation result is obtained, including: The extraction range of the highlight segment is determined by using the extraction range of the highlight segment; Based on the key frame sequence combined with the multi-dimensional evaluation candidate segment, content importance scoring is performed, and a key frame score is obtained; Based on the emotion time series data combined with the multi-dimensional evaluation candidate segment, an emotion change slope is obtained; According to the multi-dimensional evaluation candidate segment, motion intensity and visual intersection evaluation are performed, and a visual feature score is obtained; According to the key frame score, the emotion change slope and the visual feature score, a multi-dimensional evaluation result is obtained.
[0013] Optionally, S4, according to the multi-modal data analysis result, the narrative framework and the video highlight segment, acquires a video processing result, comprising: extracting a core transition segment according to the narrative framework and the event processing result; integrating the video highlight segment and the core transition segment based on the emotional time sequence data, the theme clustering label and the narrative framework, to acquire a video picture; acquiring an audio content according to the emotional time sequence data, the video picture and the narrative framework by using a dynamic time adjustment algorithm; acquiring a text content according to the theme clustering label, the video highlight segment, the core transition segment and the narrative framework; acquiring a video processing result by using the video picture, the audio content and the text content.
[0014] Optionally, acquiring a video processing result by using the video picture, the audio content and the text content, comprising: performing calibration processing on the video picture, the audio content and the text content based on the narrative framework, to respectively acquire an optimized video picture, an optimized audio content and an optimized text content; judging whether the time stamps corresponding to the optimized video picture, the optimized audio content and the optimized text content are synchronized, if yes, performing a first operation, otherwise, performing a second operation; wherein, the first operation is to integrate the optimized video picture, the optimized audio content and the optimized text content to acquire an initial video processing result, and perform a third operation; the second operation is to judge whether the difference of the time stamps corresponding to the optimized video picture, the optimized audio content and the optimized text content meets a deviation threshold, if yes, performing the first operation, otherwise, performing a fourth operation; the third operation is to perform redundant content elimination on the initial video processing result to acquire a video processing result; the fourth operation is to perform calibration processing on the video picture, the audio content and the text content based on the narrative framework, to respectively acquire an optimized video picture, an optimized audio content and an optimized text content.
[0015] The second aspect, in order to achieve the above object, the present application provides a kind of based on multi-modal data's video processing device, comprising: data analysis module, event and narrative processing module, highlight extraction module and video processing module; The data analysis module is used to acquire multi-modal data analysis result by using the multi-modal original data of video to be processed; The event and narrative processing module is used to process events and narratives based on the multimodal data analysis results, and to obtain event processing results and narrative frameworks. The highlight extraction module is used to extract highlight segments based on the multimodal data analysis results and the event processing results, and to obtain video highlight segments. The video processing module is used to obtain video processing results based on the multimodal data analysis results, the narrative framework, and the video highlight segments.
[0016] Thirdly, to achieve the above objectives, the present invention provides an electronic device, comprising: one or more processors; and a storage device having stored one or more programs thereon, which, when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation of the first aspect.
[0017] Fourthly, to achieve the above objectives, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by one or more processors, implements the method as described in any implementation of the first aspect.
[0018] Compared with the closest existing technology, the present invention has the following advantages: This invention utilizes multimodal data acquisition and structured analysis techniques to transform raw data such as video frames, audio, and text into standardized analysis results such as emotional temporal data and event segmentation results. This effectively solves the problem of fragmented utilization of multimodal data in traditional technologies, maximizing the value of data linkage throughout the entire process. Through emotion-driven event assessment, LDA+BERT joint theme generation, and a three-act narrative structure design, this invention completes event processing and narrative framework construction, solving the problem of event-narrative disconnect and significantly improving the logical coherence and duration adaptability of video content. This invention employs high-value event range constraints and combines keyframe scoring, emotional changes, and multi-dimensional weighted filtering of visual features to extract highlight segments, avoiding the limitations of single feature extraction. This makes highlight segments more aligned with the core highlights and emotional resonance points of the video, enhancing content expressiveness. Based on image optimization, emotion-matched audio generation, subtitle synchronization, and multi-element temporal calibration integration, this invention achieves deep collaboration between video, audio, and text, solving the problem of poor integration of processing results. This significantly improves the visual experience, emotional delivery effect, and overall viewing smoothness of the final video. Simultaneously, the fully automated processing significantly reduces manual intervention costs and improves video creation efficiency. Attached Figure Description
[0019] In order to more clearly illustrate the technical solutions in the specific embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or prior art description. Obviously, the drawings described below are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0020] Figure 1 A flow chart of a video processing method based on multi-modal data according to an embodiment of the present application; Figure 2 A structural schematic diagram of a video processing device based on multi-modal data according to an embodiment of the present application; Figure 3 A structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to make the objectives, technical solutions and advantages of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0022] The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0023] As shown in Figure 1 The present application provides a video processing method based on multi-modal data, comprising: S1, obtaining multi-modal data analysis results by using multi-modal original data of a video to be processed; The multi-modal original data of the video to be processed is collected, and through preprocessing and analysis, multi-modal data analysis results are obtained, which include emotional time series data, event segmentation results, theme clustering labels and key frame sequences. The core purpose of this step is to convert the unstructured original multi-modal data into standardized and reusable analysis results, providing data support for subsequent event processing, narrative construction and highlight extraction.
[0024] S2, performing event and narrative processing according to the multi-modal data analysis results to obtain event processing results and a narrative framework; Based on the multi-modal data analysis results, first, the event segmentation results are processed, that is, the importance of the event is evaluated in combination with the sentiment time series data, and similar events are clustered and merged to obtain the event processing results; second, a narrative framework is constructed based on the theme clustering label and the event processing results: a joint model of Latent Dirichlet Allocation (LDA) + pre-trained language model (BERT) is used to generate theme logic, a three-act narrative structure is designed and adapted to the target length, and the events are arranged in order along the time axis and the logical chain. The core purpose of this step is to convert fragmented events into a logical and structured narrative system, which provides a clear direction for subsequent highlight extraction and content integration.
[0025] S3, based on the multi-modal data analysis results and the event processing results, highlight segments are extracted to obtain video highlight segments; Based on the multi-modal data analysis results, in combination with the event processing results, highlight segments are extracted through multi-dimensional screening: in the time range corresponding to high-value events, content importance scores are given to key frame sequences, combined with peak segments with "emotional change slope ≥ 0.6" in the sentiment time series data, and dynamic rich or focus highlighted pictures are analyzed and screened; the above multi-dimensional results are weighted and sorted, and the top 30% of the segments are selected as video highlight segments. The core purpose of this step is to further filter out the segments that best reflect the core highlights and emotional resonance points of the video from the high-value events, thereby improving the content expressiveness of the final video.
[0026] S4, according to the multi-modal data analysis results, the narrative framework and the video highlight segments, a video processing result is obtained; Integrate multi-modal data analysis results, narrative framework and video highlight segments for multi-element collaborative processing: optimize the video pictures, generate matching audio based on the sentiment time series data, and generate synchronized subtitles in combination with the theme label and highlight segments; integrate the optimized video sequence, audio, and subtitles according to the time axis of the narrative framework, remove redundant content, and output a complete video processing result that conforms to the narrative logic and emotional expression. The core purpose of this step is to realize the deep collaboration of video, audio, and text, output a final product that conforms to the narrative logic, emotional coherence, and good viewing experience, and complete the whole process conversion from multi-modal data to intelligent video.
[0027] In summary, steps S1 to S4 effectively solve the problems of multi-modal data fragmentation, narrative logic confusion, weak highlight extraction, poor audio-visual-text collaboration in traditional video processing, realize data deep reuse and whole process automation, and generate a video that not only has a clear narrative structure and core highlight focus, but also can enhance user resonance through the whole process of emotion, while greatly reducing the cost of manual intervention, improving the generation efficiency compared to traditional editing methods, and accurately adapting to the scene needs of Vlog, short video and other scenes that require strong narrative and emotional transmission.
[0028] As a possible implementation, in the above embodiment, step S1 can specifically include the following steps: S1-1, collecting multi-modal raw data of the video to be processed; The multi-modal raw data category contained in the video to be processed is explicitly defined, specifically covering video frame sequences, audio signals, associated text information, and metadata. The video frame sequence is a set of continuous static pictures obtained after video decoding, each frame being image data containing pixel information, recording visual information such as character actions, scene environment, object shape, and color distribution in the picture. The audio signal is the digitized data of sound waves packaged in the video file, with sampling rate (e.g. 44.1 kHz) and bit depth (e.g. 16 bits) as parameters, containing sound information such as human voice, environmental sound, and background music. The associated text information is unprocessed text information directly related to the video, including real-time transcription text of voice recorded synchronously during video shooting, original labels manually added by the shooter, unverified subtitle text provided by the video file, and shooting notes automatically recorded by the device. Metadata is raw information describing video properties, divided into technical metadata and content metadata: technical metadata includes original parameters generated by the shooting device, such as device model, video resolution, frame rate, encoding format, shooting timestamp, and file size; content metadata includes shooting location, shooting object, and scene type, which are original content descriptions labeled during shooting or associated with the device.
[0029] After collection, the collected raw data is classified and organized into four data subsets: video, audio, text, and metadata. Through classification, the processing direction of different data is clarified, providing clear classification basis for subsequent targeted preprocessing, ensuring that subsequent processing links can accurately match the characteristics of various data.
[0030] S1-2, preprocessing the multi-modal raw data of the video to be processed to obtain standardized multi-modal data; The classified multi-modal raw data is input, and differential preprocessing operations are carried out according to the characteristics of different types of data. For video frame sequences, Gaussian filtering and other algorithms are used to remove picture noise and improve frame clarity; for audio signals, background noise is eliminated by spectral subtraction to retain valid sound information; for associated text information, format unification, error correction, and stop word deletion operations are performed to achieve text standardization; for metadata, data integrity is verified and missing key information is supplemented, such as preliminary determination of scene type based on video pictures when scene type is not labeled.
[0031] At the same time, the audio signal, text information, etc. are time-axis aligned with the video frame based on the timestamp of the video frame, ensuring that each type of data corresponds one-to-one in the time dimension, and finally outputting standardized multi-modal data with regular structure and time synchronization, laying a data foundation for subsequent feature analysis.
[0032] S1-3, obtaining emotion time series data by a multi-modal emotion fusion algorithm using the standardized multi-modal data; This step relies on the obtained standardized multi-modal data to carry out analysis on the emotion dimension. First, speech intonation features such as speech speed, tone, and volume changes are extracted from the pre-processed audio signal, and picture emotion features such as character facial expression, picture color saturation, and scene atmosphere are extracted from the video frame sequence. Then, a multi-modal emotion fusion algorithm is used to weight and fuse the emotion features extracted from the audio and picture, and a trained emotion recognition model is used to determine the emotion type of the video at different time points and quantitatively output the emotion intensity, with a value range of 0-1, and the higher the value, the stronger the emotion. Finally, emotion time series data is formed, which contains emotion intensity (0-1) and emotion type (such as "excited", "calm", "happy") changing with time, and is the basis for subsequent emotion-driven creation.
[0033] S1-4, obtaining event segmentation results based on the standardized multi-modal data using a picture-audio collaborative segmentation strategy; This step uses standardized data as support to realize event-level division of video content. First, scene change detection algorithms (such as sudden change detection based on inter-frame pixel difference, picture content clustering based on deep learning) are used to detect picture feature mutation points such as scene change, appearance / disappearance of core objects, and large changes in camera angle from the pre-processed video frame sequence, as picture segmentation candidate points. At the same time, audio semantic segmentation algorithms are used to extract text content through speech recognition and analyze semantic topic switching from the pre-processed audio signal, and audio feature mutation points such as speech content theme change and background music style mutation are detected as audio segmentation candidate points. Combining the picture and audio two-dimensional segmentation basis, a picture-audio collaborative segmentation strategy is adopted: when picture and audio segmentation candidate points appear at the same time point, it is directly determined as an event boundary; when only a single type of candidate point appears, the true event boundary is selected by combining the feature change trend of adjacent time points to determine the start and end time stamps of each independent event. Then, the core information in each event time period is extracted: representative frames are selected from the video frames, and key speech segments are extracted from the audio signal, and the event core content description is generated by fusion, finally forming an event segmentation result containing event time range and content summary, such as "family dinner: 10:00-15:30", clearly presenting the event structure of the video and clearly defining the start and end range of the event.
[0034] S1-5, performing semantic analysis according to the standardized multi-modal data to obtain theme clustering labels; This step focuses on semantic topic analysis of video content, taking pre-processed associated text information as the core input data. The standardized text information is divided into semantic units according to sentence punctuation, and after removing meaningless mood words and repeated content, it is input into the LDA topic clustering model. The model clusters the semantic units by probability based on the pre-set number of topics: first, calculate the probability of each semantic unit belonging to different topics, then classify semantic units with a probability ≥ 0.6 into the same topic, and through topic keyword extraction, first determine the global theme covering the entire video content, such as "family dinner"; then further divide the text into different paragraphs based on the global theme in chronological order, and repeat the above clustering process for each paragraph to subdivide the local sub-themes corresponding to the events, such as "dish tasting" and "game interaction"; finally, generate topic clustering labels containing global themes and local sub-themes, clearly presenting the theme hierarchy and core direction of the video content, facilitating subsequent organization of narratives around the theme.
[0035] S1-6, key frame extraction using the standardized multi-modal data to obtain a key frame sequence; For the pre-processed video frame sequence, inter-frame similarity calculation algorithms such as histogram similarity and structural similarity (SSIM) are used to calculate the content similarity of adjacent frames. When the similarity is ≥ 85%, it is determined as a redundant frame and removed. For the remaining non-redundant frames, score them according to the weight of core elements (such as key characters, iconic actions, important objects): frames with key characters (such as speakers, main characters) have a weight of 0.3, frames with iconic actions / objects (such as product demonstrations, ritual segments) have a weight of 0.3, frames corresponding to emotional peaks (such as cheers, smiling moments) have a weight of 0.2, and frames with high theme relevance (such as pictures containing theme keywords) have a weight of 0.2. Calculate the score of each frame, select the top 30% of frames that represent the key content of the video, and form a key frame sequence, especially important material sources for highlight clips. Each key frame is bound to a corresponding timestamp and core content description.
[0036] S1-7, obtaining a multi-modal data analysis result according to the emotional time series data, the event segmentation result, the topic clustering label, and the key frame sequence; The emotional time series data, event segmentation result, theme clustering label and key frame sequence are associated in multiple dimensions. In the time dimension, the emotional data, event information, theme label and key frame under the same timestamp are accurately matched. In the content dimension, the data association is strengthened through keyword association. Finally, a structured multi-modal data analysis result is formed. The result takes the time axis as the core context, integrates the emotional time series, event segmentation, theme clustering and key frame data, and each type of data contains specific content, time range and association relationship, providing comprehensive, collaborative and directly callable data support for subsequent video processing, narrative framework construction and highlight segment extraction.
[0037] In summary, steps S1-1 to S1-7 take the multi-modal raw data of the video to be processed as the starting point, first classify and differentially preprocess the raw data according to the type, realize data standardization and time axis alignment, then generate emotional time series data, event segmentation result and theme clustering label based on the standardized data, and synchronously extract the key frame sequence, and finally integrate the four types of data according to the time and content dimensions to form a structured multi-modal data analysis result. This process ensures data quality and time synchronization through classification preprocessing, and breaks through the limitations of single data dimension through multi-dimensional collaborative analysis. Not only does it solve the problem of fragmentation and weak association of multi-modal raw data, but also improves the accuracy of emotion recognition, event segmentation and theme clustering. It also provides comprehensive and collaborative data support for subsequent video event processing, narrative construction and highlight extraction, effectively ensuring the logic and efficiency of the whole video processing process.
[0038] As a possible implementation, in the above embodiment, step S2 can specifically include the following steps: S2-1, based on the emotional time series data, performing importance evaluation on the event segmentation result to obtain an event importance evaluation result; The emotional time series data and event segmentation result in the multi-modal data analysis result are taken as inputs to realize importance evaluation through "emotion-event association mapping". Specifically, the time interval of each event in the event segmentation result is aligned with the time axis of the emotional time series data, and the emotional features in the time range corresponding to each event are extracted, including the average emotional intensity, emotional peak value and emotional duration; then, through weighted calculation, such as 40% for the average emotional intensity, 30% for the emotional peak value and 30% for the emotional duration, the importance score (0-1) of each event is obtained, and finally the event importance evaluation result containing "event identifier-event content-importance score" is formed, providing a quantitative basis for subsequent event screening.
[0039] S2-2, screening according to the event importance evaluation result combined with an emotional intensity threshold to obtain a high-value event; This step takes the event importance evaluation result as the core basis, and realizes event screening by setting the importance threshold of emotional association. First, combined with the emotional expression needs of the target video scene (such as Vlog, short video), the emotional intensity threshold is set, usually 0.7, which can be dynamically adjusted according to the scene; secondly, all events in the event importance evaluation result are traversed, and high-value events with emotional intensity ≥ threshold are retained, and transitional and irrelevant events with emotional intensity < threshold are removed; finally, the retained events are preliminarily sorted to form a high-value event set containing an event list, the core content of each event, and the original time interval, focusing on the core content of the video with prominent emotions and high user attention.
[0040] S2-3, using a pre-trained language model to compare the content of the high-value events and obtain an event processing result; Taking high-value events as input, the semantic understanding ability of the pre-trained language model (such as BERT) is used to realize event deduplication and optimization. First, the core content description of each high-value event is converted into a vector representation; secondly, the semantic similarity between event content vectors is calculated by the BERT model, and a similarity threshold of 0.8 is set to identify and merge repeated or highly similar events, such as merging multiple “family photo” events into a “group photo” event; at the same time, the semantic extraction of the merged event content is carried out, and the total time length, core scene and other key information of the merged event are supplemented, and finally the event processing result containing the high-value event list and the event association relationship (time sequence / logical causality) is formed, which optimizes the compactness of the event content and improves the efficiency and logicality of subsequent narrative construction.
[0041] S2-4, based on the event processing result and the theme clustering label, using a joint model for correlation analysis to obtain an event association relationship; This step takes the time processing result and the theme clustering label in the multi-modal data analysis result as input, and realizes theme matching and correlation analysis by a joint model of LDA+BERT. First, the global theme and theme keywords of the video are determined by the LDA model; secondly, the semantic similarity between each high-value event and the global theme keywords is calculated by the BERT model, and events unrelated to the global theme are removed, such as “office work” events mixed in travel videos; at the same time, based on the theme association, the logical relationship between high-value events is analyzed, and finally the “high-value event-global theme matching degree” and the “event logical association (causality / parallelism / chronological order)” of the event association relationship are output, ensuring the logical consistency between events and themes, and between events.
[0042] S2-5, according to the event segmentation result combined with the event association relationship, constructing an event timeline; An event time axis with time sequence and logicality is constructed by taking the event original timestamp in the event segmentation result and the event association as input. First, the original start and end timestamps of each high-value event in the event segmentation result are extracted, and the events are preliminarily arranged in chronological order. Second, the logical relationship in the event association is combined to optimize the time sequence of the preliminarily arranged events, and the transition nodes between events are supplemented, such as the connection time from the end of the "mountain climbing" event to the beginning of the "viewing at the mountain top" event. Finally, the start and end time, core content and logical association between events of each high-value event are labeled with time as the horizontal axis to form a structured event time axis, which intuitively presents the time sequence and logical context of the events.
[0043] S2-6, based on the target video length, a three-act narrative structure is constructed; This step takes the target video length (set according to the application scenario) as a constraint, such as the 3-5 minute golden time length of Vlog and short video, and designs a three-act narrative framework (beginning-development-conclusion) that meets the user's viewing habits. First, the time length proportion of the three acts is allocated according to the target length, usually 10%-15% for the beginning, 60%-70% for the development, and 15%-20% for the conclusion, such as 0.5-0.75 minutes for the beginning, 3-3.5 minutes for the development, and 0.75-1 minute for the conclusion of a 5-minute video. Second, the functional positioning of each act is clarified: the beginning is responsible for scene introduction and background laying, such as the "departure scene" of a travel video; the development focuses on the presentation of core events and emotional progression, such as the high-value events of "mountain climbing and viewing" and "camping experience"; the conclusion undertakes the ending summary and emotional sublimation, such as "travel insights" and "group photo". Finally, a three-act narrative structure is formed, including "three-act time length allocation-function positioning of each act", which provides a framework constraint for subsequent event allocation.
[0044] S2-7, according to the high-value events, the event time axis and the three-act narrative structure, a narrative framework is obtained; This step takes the high-value events, the event time axis and the three-act narrative structure as input for collaborative matching. First, according to the time sequence logic of the event time axis and the functional positioning of the three-act structure, the high-value events are allocated to the corresponding narrative acts, such as "departure scene" to the beginning, "mountain climbing" and "camping" to the development, and "group photo" to the conclusion. Second, combined with the time length proportion of each act, the time length of the events allocated to each act is adjusted, such as compressing the time length of the beginning events and extending the time length of the core events in the development stage, to ensure that the time length of each act meets the preset proportion. At the same time, the narrative rhythm is optimized through "beginning, transition, transformation and combination" node design, such as setting an emotional climax node in the development stage. Finally, the information such as "three-act structure-event allocation-time length adjustment-rhythm node" is integrated to form a complete narrative framework with time sequence logic, theme consistency and time length adaptability, which provides a clear narrative guide for subsequent video content integration.
[0045] In summary, steps S2-1 to S2-7 take the sentiment time series data, event segmentation results, and theme clustering labels in the multi-modal data as core inputs, and through the progressive processing of “emotion quantification to evaluate event importance → threshold screening to focus on high-value events → pre-trained model to merge redundant events → joint model to calibrate theme matching degree → combination of time series and logic to construct event timeline → adaptation of target time length to design three-act structure → integration of events and structure to form narrative framework”, the conversion from fragmented events to structured narrative scheme is systematically completed. This process solves the problem of ambiguous value judgment in traditional event processing by driving event screening with emotional data; avoids event redundancy and theme disconnection by optimizing event relevance and theme matching degree with pre-trained models and joint models; and constructs a narrative framework by combining time series logic and three-act structure, which not only ensures the logicality and rhythm of the narrative, but also adapts to the target video length and user viewing habits, ultimately providing precise and efficient narrative guidance for subsequent video content integration, significantly improving the content coherence, emotional expression effectiveness, and scene adaptability of the generated video.
[0046] As a possible implementation, in the above embodiment, step S3 can specifically include the following steps: S3-1, determining the extraction range of the highlight segment by using the event processing result; The extraction range of the highlight segment is determined by locking the extraction boundaries of the highlight segment based on the event processing result, which includes a list of high-value events screened after importance evaluation and time interval information corresponding to each event. Since high-value events are key carriers of core content and emotional highlights of the video, the extraction range is preferentially focused on the video event interval corresponding to the high-value events. Specifically, by mapping the start and end timestamps of the high-value events in the event processing result to the time axis of the video to be processed, the time boundaries of the extraction range are determined. At the same time, the time intervals corresponding to low-value events (such as transitional scenes and irrelevant redundant content) marked in the event processing result are directly excluded to avoid interference from invalid content in the highlight screening, and finally the extraction range of the highlight segment is formed with the time interval of the high-value events as the core, ensuring that the subsequent extraction work focuses on the core content area of the video.
[0047] S3-2, performing multi-dimensional evaluation based on the multi-modal data analysis result and the extraction range of the highlight segment to obtain a multi-dimensional evaluation result; This step extracts the candidate segment within the determined extraction range of the highlight segment, and combines the multi-modal data analysis result to quantitatively evaluate the candidate segment from three dimensions of key frame score (content importance), emotional change analysis (emotional peak), and visual feature analysis (motion intensity and visual focus), and finally obtains independent evaluation scores of each candidate segment in each dimension to form a multi-dimensional evaluation result.
[0048] S3-3, weighting and fusing the multi-dimensional evaluation results to obtain a comprehensive evaluation result; This step converts the multi-dimensional independent evaluation results into a unified quantitative indicator through weighted fusion. The core is to assign weights according to the contribution of each dimension to the "highlight attribute". First, standardize the evaluation results of each dimension, such as mapping the scores of different dimensions to the 0-1 interval to eliminate dimensional differences; then, according to the influence weight distribution coefficient of each dimension on the "highlight segment", set the weight based on the video content characteristics and user focus, for example, key frame score (reflecting content core degree) accounts for 30%, emotional change slope (reflecting emotional resonance degree) accounts for 30%, visual feature score (reflecting visual attraction) accounts for 40%; then, using linear weighted summation, multiply each dimension evaluation score of each candidate segment by the corresponding weight and add them up to get the comprehensive evaluation score of the candidate segment; finally, form a comprehensive evaluation result containing the comprehensive score of each segment in units of candidate segments, realizing the unified quantitative sorting of the "highlight degree" of candidate segments.
[0049] S3-4, according to the extraction range of the highlight segment and the comprehensive evaluation result, obtaining the video highlight segment; First, based on the determined extraction range, ensure that the selected objects are all segments corresponding to high-value events, and exclude low-value content outside the range; second, according to the comprehensive evaluation result, select the segments from high to low according to the comprehensive score - you can set the selection rules, such as selecting the top 30% high-score segments, or selecting segments with a comprehensive score greater than or equal to 0.8; at the same time, check the time continuity and content integrity of the selected segments to avoid content fragmentation caused by high-score segment fragmentation, and if necessary, merge adjacent high-score short segments; finally, output the video highlight segment set containing the time interval, content summary and comprehensive score of each highlight segment, and complete the whole process of highlight extraction.
[0050] In summary, steps S3-1 to S3-4 first lock the time interval corresponding to high-value events as the extraction range of highlight segments based on the event processing result, avoiding invalid extraction of low-value content; then, in this range, combined with the multi-modal data analysis results, carry out targeted evaluation from the dimensions of key frame score, emotional change, visual feature, etc., to obtain multi-dimensional independent results; then, through standardization and weight distribution, the multi-dimensional results are weighted and fused into comprehensive evaluation scores and sorted; finally, combined with the extraction range constraint and comprehensive score selection, the video highlight segment is obtained. This process uses the progressive logic of "range focusing-multi-dimensional evaluation-weighted fusion-precise selection" to effectively improve the accuracy and efficiency of highlight segment extraction, ensuring that the extracted highlight segments focus on the core content of the video, and also consider emotional resonance and visual appeal, output high-quality core materials for subsequent video processing, and further improve the content performance and user viewing experience of the final video product.
[0051] As a possible implementation, in the above embodiment, step S3-2 can specifically include the following steps: S3-2-1, determining a candidate segment for multi-dimensional evaluation by using the extraction range of the highlight segment; With the predetermined highlight segment extraction range as the boundary, the candidate segment for multi-dimensional evaluation is split from the video to be processed. Specifically, the start and end timestamps of the extraction range are first determined, such as the cake cutting event: 12:00-13:30, and then the video in the time interval is split into several continuous segments according to the time continuity and scene integrity of the video content. The splitting can be combined with features such as scene change and audio rhythm change to avoid destroying the integrity of a single action or scene. For example, the cake cutting process is split into three continuous candidate segments: "holding knife preparation - cutting moment - dividing cake". Finally, a set of candidate segments located within the extraction range and having content independence is formed, providing specific objects for subsequent multi-dimensional evaluation.
[0052] S3-2-2, performing content importance scoring based on the key frame sequence and the candidate segment for multi-dimensional evaluation to obtain key frame scoring; Based on the key frame sequence in the multi-modal data analysis result, content importance scoring is performed for each candidate segment. First, the key frame corresponding to each candidate segment is located by timestamp matching to filter out all key frames whose time falls within the candidate segment interval. Second, content importance evaluation is performed on these key frames, and the evaluation indicators include the saliency of the core elements in the frame and the relevance of the video theme, and a 0-1 scoring system is used. Finally, the average or weighted average of the scores of all key frames in a single candidate segment is taken, and the average is taken as the key frame score in the candidate segment, reflecting the core degree of the segment content.
[0053] S3-2-3, obtaining the emotion change slope based on the emotion time series data and the candidate segment for multi-dimensional evaluation; Based on the emotion time series data in the multi-modal data analysis result, the emotion change slope of each candidate segment is calculated. First, according to the start and end timestamps of the candidate segment, the emotion intensity data points in the corresponding time interval are extracted from the emotion time series data, such as the "cake cutting candidate segment" corresponding to 12:00-12:10, and the emotion intensity value every 0.5 seconds in the interval is extracted. Second, the emotion change curve is constructed with time as the horizontal axis and emotion intensity as the vertical axis, and the slope of the curve is calculated by linear fitting or difference calculation. The greater the absolute value of the slope, the more intense the emotion change in the candidate segment, such as the emotion intensity from "expectation" to "excitement" rising from 0.6 to 0.95, with a slope of 0.82, indicating significant emotion fluctuation. Finally, the calculated slope value is taken as the emotion change evaluation result of the candidate segment, i.e. the emotion change slope.
[0054] S3-2-4, performing motion intensity and visual intersection evaluation on the candidate segment according to the multi-dimensional evaluation, and obtaining a visual feature score; This step evaluates and fuses the visual attraction of the candidate segment from two dimensions of "dynamic richness" and "visual focus degree" into a visual feature score. In the motion intensity evaluation, the optical flow density in the candidate segment, i.e., the vector size of the pixel motion between adjacent frames, is calculated. The higher the optical flow density value, the more dynamic the picture is, such as the optical flow density of the "family cheering" candidate segment being 0.85, which is strong in dynamic. In the visual focus evaluation, the attention heat map algorithm is used to detect the visual focus area of each frame in the segment, and the focus area is scored (0-1) according to its saliency and stability. Finally, the scores of the two dimensions are weighted and summed according to the preset weight, such as 40% for motion intensity and 60% for visual focus, to obtain the visual feature score of each candidate segment, which comprehensively reflects the visual attraction of the picture.
[0055] S3-2-5, obtaining a multi-dimensional evaluation result according to the key frame score, the emotional change slope and the visual feature score; This step integrates the key frame score, the emotional change slope and the visual feature score obtained in the previous three steps to form a structured multi-dimensional evaluation result. First, the results of each dimension are standardized - the emotional change slope is mapped to the 0-1 interval through the normalization algorithm to ensure that the dimensions are consistent with the key frame score and the visual feature score. Second, according to the dimension classification of "key frame score-emotional change slope-visual feature score", an evaluation result entry is established for each candidate segment, which contains the time interval of the candidate segment, the standardized scores and original scores of the three dimensions. Finally, a multi-dimensional evaluation result set is formed, which covers the content core degree (key frame score), emotional fluctuation degree (emotional change slope) and visual attraction (visual feature score) in units of candidate segments, providing a complete basis for subsequent weighted fusion and selection of highlight segments.
[0056] In summary, steps S3-2-1 to S3-2-5 are based on the extraction range of the highlight segment, and first, the range is used to split the candidate segment with content independence from the video to be processed, so as to accurately focus on the evaluation object; then, for each candidate segment, the key frame score reflecting the content core degree is calculated in combination with the key frame sequence, the emotional change slope reflecting the emotional fluctuation degree is obtained in combination with the emotional time sequence data, and the visual feature score reflecting the visual attraction is obtained through the motion intensity and visual focus analysis; finally, the evaluation results of the three dimensions are standardized and integrated into a structured multi-dimensional evaluation result. Through the logic of "range constraint-multi-dimensional disassembly evaluation-standardization integration", the process not only avoids invalid evaluation of low-value content, but also quantifies the segment value from three key dimensions of content core, emotional resonance and visual attraction, solves the problem of insufficient accuracy caused by the dependence of traditional highlight extraction on a single feature, significantly improves the comprehensiveness and accuracy of highlight segment evaluation, and provides reliable and standardized evaluation basis for subsequent weighted selection of high-quality highlight segments meeting user needs, thereby ensuring the content expressiveness of the final video processing result and the user viewing experience.
[0057] As a possible implementation, in the above embodiment, step S3 can specifically include the following steps: S4-1, extracting core transition segments according to the narrative framework and the event processing result; The narrative framework clearly defines the overall structure of the video and the arrangement logic of each high-value event, and the event processing result contains filtered high-value events and low-value transition events. When extracting, the transition content in the event processing result that has a direct logical association with the high-value event is preferentially selected, for example, between the two high-value events of "sightseeing at scenic spot A" and "food experience", the transition segment of "on the way to the food shop" in the event processing result is selected; at the same time, the length of the extracted core transition segment is adapted to the narrative rhythm, usually 5-15 seconds per segment, and the content is related to the theme of the adjacent high-value event, and finally a core transition segment set is formed to connect the video highlight segments and ensure the narrative coherence.
[0058] S4-2, integrating the video highlight segments and the core transition segments based on the emotional time sequence data, the theme clustering label and the narrative framework, to obtain a video picture; The emotional timing data sets the mood of the picture, such as high saturation warm color for high light segments with emotional intensity ≥ 0.8, and low saturation soft color for core transition segments with emotional intensity < 0.5; the theme clustering label unifies the visual style, such as under the theme of "travel-natural scenery", all segments are calibrated to natural and fresh color tone; the narrative framework guides the technical processing details, that is, high light segments are preferentially processed by spatio-temporal super-resolution adaptive reconstruction (resolution is increased to 1080P) and anti-shake processing (eliminating handheld shooting jitter), and core transition segments focus on transition adaptation, such as "fade-in and fade-out" for transitions between high-value events, and "slide transition" for transitions within continuous scenes, to ensure that the transition effect meets the needs of narrative logic connection. Through the above integration and optimization, the dispersed high light segments and core transition segments are transformed into video picture sequences with unified style, clear quality and smooth connection.
[0059] S4-3, according to the emotional timing data, the video picture and the narrative framework, an audio content is obtained by using a dynamic time adjustment algorithm; The emotional timing data directly maps the audio parameters, wherein high volume and fast rhythm music is matched when emotional intensity ≥ 0.8, such as high light segment "sunrise cheers"; medium volume and medium rhythm music is matched when emotional intensity is 0.5-0.8, such as core transition segment "walking in the scenic area"; low volume and slow rhythm music is matched when emotional intensity < 0.5, such as transition segment "night rest". The rhythm of the video picture determines the audio details, by analyzing the picture editing rhythm, such as fast switching of high light segments and slow shots of transition segments, adjusting the music rhythm and melody fluctuation, and using dynamic time adjustment (DTW) algorithm to realize the precise alignment of audio rhythm and picture motion intensity; the narrative framework standardizes the overall style of the audio, such as "travel Vlog" narrative corresponding to audio with light and lively folk songs as the base, and "birthday party" narrative with warm and cheerful melodies as the main melody, while extracting high-value environmental sound from multi-modal data analysis results for sound effect enhancement, and finally forming audio content highly matched with picture emotion, rhythm and narrative style.
[0060] S4-4, according to the theme clustering label, the video high light segment and the core transition segment combined with the narrative framework, a text content is obtained; The theme clustering label explicitly indicates the text style direction, such as the "family gathering" theme, in which the text adopts a warm and colloquial expression; and the "academic conference" theme, in which the text adopts a formal written expression. The video highlight segment and the core transition segment provide text materials. For the highlight segment, the core elements in the frame are extracted to generate core descriptions; and for the core transition segment, scene information is extracted to generate transition text. The narrative framework ensures the logical coherence of the text, and according to the narrative context of "beginning-development-conclusion", the context association is given to each segment text, such as the transition segment text in the beginning part focusing on "scene introduction", and the highlight segment text in the development part focusing on "event details". With the help of the GPT-4 model, the above information is combined to generate the subtitle text, and the subtitle timestamp is aligned with the key frame time of the corresponding segment (deviation ≤0.5s), ensuring that the text content not only fits the picture, but also conforms to the narrative logical progression, and finally forming a complete subtitle text system.
[0061] S4-5, obtaining a video processing result by using the video picture, the audio content and the text content; This step converts the video picture, audio content and text content into a complete video through the integration of multiple elements. Based on the narrative framework timeline, the timestamps of the video picture, audio content and text content are calibrated and optimized (deviation ≤0.3s), and the three are integrated in the order of narrative logic to form a preliminary video. After removing redundant content and adapting to the target length, the complete video processing result is output.
[0062] In summary, the video processing process of steps S4-1 to S4-5 takes the narrative framework as the core guide. First, the core transition segment is extracted based on the event processing result to ensure narrative continuity. Then, the video highlight segment and the core transition segment are unified in terms of picture style, improved in terms of picture quality, and adapted in terms of transition, based on the emotional time series data and the theme clustering label in the multi-modal data. Subsequently, the matching audio is generated based on the emotion and picture rhythm, and the synchronous text that is logically coherent is produced around the theme and picture content. Finally, the optimized picture, audio and text are integrated according to the narrative logic and timeline, and the redundant content is removed to form a complete video processing result. This process realizes the deep cooperation between multi-modal data and narrative requirements. It not only solves the problem of picture fragmentation and the disconnection between audio, picture and text in traditional processing, but also significantly improves the narrative coherence, audio-visual coordination and content expressiveness of the video through emotion-driven multi-element adaptation and narrative logic. At the same time, the process of automatic integration reduces the cost of manual intervention and efficiently outputs high-quality video results that meet the scene requirements (such as Vlog and short video).
[0063] As a possible implementation, in the above embodiment, step S4-5 can specifically include the following steps: S4-5-1, calibrate the video picture, the audio content and the text content based on the narrative framework, and obtain optimized video picture, optimized audio content and optimized text content respectively; This step takes the narrative framework as the core calibration basis and carries out targeted optimization calibration for video picture, audio content and text content, as follows: (1) For the video picture, the picture color is secondarily calibrated to unify the visual style according to the style keynote and time axis rhythm of the narrative framework, the transition effect duration is adjusted to adapt to the narrative rhythm, and the blurred frames are supplemented and super-resolution processed to form the optimized video picture; (2) For the audio content, the background music volume level is calibrated in combination with the emotional trend of the narrative framework, such as gentle at the beginning and intense in development, the volume is increased by 10-15% for highlight segments, the clarity of environmental sound effects is adjusted to ensure that the audio style matches the narrative stage, and the optimized audio content is obtained; (3) For the text content, the relevance of the subtitle content and the context is corrected according to the logical context of the narrative framework, such as supplementing the transitional segment of the subtitle, and the subtitle font style and size are calibrated to fit the narrative scene, such as using a round font for a warm scene, to generate optimized text content; Through independent calibration of the three, the frame sequence of the optimized video picture, the waveform signal of the audio content and the subtitle display time of the text content are completely synchronized, laying a foundation for subsequent synchronous integration.
[0064] S4-5-2, judge whether the timestamps corresponding to the optimized video picture, the optimized audio content and the optimized text content are synchronized, if yes, directly execute S4-5-4, otherwise, execute S4-5-3; This step is a preliminary judgment of timestamp synchronization, focusing on the timestamp matching degree of the key nodes of the optimized video, audio and text. Select the core time nodes in the narrative framework, such as the start / end point of the highlight segment, the transition switching point, the core event triggering point, etc., extract the timestamps of the frames corresponding to the optimized video picture, the timestamps of the waveform peaks corresponding to the optimized audio content, and the timestamps of the subtitle display / hide corresponding to the optimized text content, and compare whether the timestamps of the three at the same core node are completely consistent, such as the "scenic spot sunrise cheers" highlight segment, the video frame timestamp 1:30, the audio peak timestamp 1:30, and the subtitle display timestamp 1:30. If the timestamps of all core nodes are completely synchronized, it means that the time coordination of the three meets the standard and can directly enter the integration link; if there is any core node timestamp mismatch, it enters the next step of deviation threshold judgment.
[0065] S4-5-3, judge whether the difference value of the optimized video picture, the optimized audio content and the optimized text content corresponding timestamp meets the deviation threshold value, if yes, execute S4-5-4, otherwise, return to execute S4-5-1; This step is a timestamp deviation fault tolerance verification. By presetting a deviation threshold value, usually set to ≤0.3s, it is judged whether the timestamp difference value is within the acceptable range. The maximum difference value of the timestamps at the different synchronous core nodes is calculated, such as video frame 1:30, audio 1:30.2, and subtitle 1:30.1. The maximum difference value is 0.2s. If the difference value is ≤the preset deviation threshold value, it means that the deviation has no significant impact on the audio-visual experience, and there is no need to recalibrate. It can enter the integration link. If the difference value is >the deviation threshold value, it means that the time deviation will cause audio-visual out of sync, subtitle misplacement and other problems. It is necessary to recalibrate based on the deviation data this time until the timestamp difference value meets the fault tolerance requirements.
[0066] S4-5-4, integrate the optimized video picture, the optimized audio content and the optimized text content to obtain an initial video processing result; Based on the timeline of the narrative framework, the frame sequence of the optimized video picture is taken as the basic carrier, the audio track of the optimized audio content is embedded, the background music and environmental sound effects are matched with the picture rhythm, and the subtitle track of the optimized text content is superimposed to ensure that the subtitle and the picture core element and the audio key information are displayed synchronously. At the same time, the structure order (beginning-development-conclusion) of the narrative framework is followed, each segment is connected in logical chain, and the initial video processing result containing complete picture, audio and subtitle is formed.
[0067] S4-5-5, remove redundant content from the initial video processing result to obtain a video processing result; This step refines the content of the initial video processing result to improve the narrative compactness and content quality. According to the core needs of the narrative framework, three types of redundant content are screened and removed: first, the segments conflicting with the theme of the narrative, such as irrelevant work pictures inserted in travel Vlog; second, the content with inconsistent emotions, such as sudden appearance of messy transition segments in warm scenes; third, the parts with redundant time length, such as repeated transition pictures exceeding the target time length and long information blank segments. During the removal process, the integrity of the narrative logic needs to be preserved, such as not deleting key transition segments to ensure that the remaining content closely follows the narrative framework. Finally, the complete video processing result with adapted time length, refined content, coherent narrative and coordinated audio-visual-text is output.
[0068] To sum up, steps S4-5-1 to S4-5-5 are processed by progression, not only effectively solving the problems of time dislocation of audio, picture and text, chaotic content and the like in traditional video integration, ensuring that the time synchronization deviation of the three is controlled within a reasonable range, and significantly improving the narrative coherence and audio-visual coordination of the video; at the same time, the closed-loop calibration mechanism guarantees the stability of the processing quality, and the redundant elimination link further enhances the content focus, finally realizing the technical effect of efficiently outputting high-quality video processing results meeting the narrative requirements and stable in quality.
[0069] Further reference Figure 2 , as an implementation of the method shown in the above figures, the present disclosure provides one embodiment of a video processing device based on multi-modal data, which corresponds to the method embodiment shown in Figure 1 , and the device can be applied in various electronic devices.
[0070] As shown in Figure 2 , a video processing device based on multi-modal data in the embodiment comprises a data analysis module, an event and narrative processing module, a highlight extraction module and a video processing module. The data analysis module is configured to obtain multi-modal data analysis results by using multi-modal raw data of a video to be processed. The core function of the module is to convert the unstructured multi-modal raw data of the video to be processed into standardized and reusable analysis results. The input is the full-dimensional raw data of the video to be processed, specifically including video frame sequence (including picture texture, color distribution, object motion trajectory), audio signal (including human voice, environmental sound, rhythm and tone characteristics), associated text information (such as voice-to-text content, scene note tags when shooting) and metadata (such as shooting time, device parameters, scene type). In the processing process, the module first performs preprocessing operations, including video frame denoising, audio noise reduction and multi-modal data time axis alignment, to ensure that the timestamps of the video, audio and text are consistent; then four types of core results are generated through special analysis algorithms: (1) Emotional time series data: emotional time series data containing emotional intensity 0-1 points and emotional type changing with time are generated by fusion of "audio emotion recognition (voice tone analysis) + picture emotion recognition (person's expression, color atmosphere judgment)", such as "excitement" and "calmness"; (2) Event segmentation result: event segmentation result of clear start and end timestamps and core content description of each independent event is generated by combination of "picture scene switching detection + audio semantic segmentation"; (3) Theme clustering label: theme clustering label containing global theme and local sub-theme is generated based on LDA model content semantic clustering; (4) Keyframe sequence: The keyframe sequence containing representative frames with core characters and key actions is extracted by "frame content similarity calculation + core element weight scoring".
[0071] The final output of the module is a structured multi-modal data analysis result, which directly provides accurate data support for subsequent event and narrative processing, highlight extraction, and video processing modules, solving the problem of "fragmentation and difficulty in reuse" of raw data.
[0072] The event and narrative processing module is configured to perform event and narrative processing based on the multi-modal data analysis result, and obtain an event processing result and a narrative framework. The module takes the multi-modal data analysis result output by the data analysis module as the only input, and the core goal is to convert fragmented events into a logical and structured narrative system, outputting an event processing result and a narrative framework.
[0073] In the event processing link, the module first filters all events in the event segmentation result based on sentiment time series data - sets a sentiment intensity threshold (such as 0.7), retains "high-value events" with a sentiment intensity greater than or equal to the threshold, and marks "low-value transition events" with a sentiment intensity less than the threshold; then calculates the semantic similarity of event content through the BERT model, merges repeated or highly similar high-value events, and finally forms an event processing result containing a "high-value event list, event association relationship (time sequence / logical causality)".
[0074] In the narrative framework construction link, the module first combines the theme clustering label and the event processing result, and uses the "LDA + BERT" joint model to determine the theme logic, wherein the LDA model locks the global theme of the video, and the BERT model calibrates the semantic matching degree of high-value events and the global theme, and eliminates events unrelated to the theme; then based on user viewing habits such as Vlog and short video 3-5 minute golden time length, a "three-act narrative structure" (beginning-development-conclusion) is designed, the narrative rhythm is optimized by adjusting the length ratio of high-value events (core event ratio ≥ 60%), and finally a narrative framework with logical coherence and length adaptation is formed.
[0075] The output of the module directly determines the core logic and presentation structure of the subsequent video content, avoiding the problem of "content confusion and disorder".
[0076] The highlight extraction module is configured to extract a video highlight segment based on the multi-modal data analysis result and the event processing result. The module takes the multimodal raw data of the data analysis module and the event processing result of the event and narrative processing module as dual inputs, and the core function is to accurately extract the highlight segment with the most expressive and emotional from the video. Its processing logic revolves around "range constraint + multi-dimensional evaluation": first, the "high-value event time interval" in the event processing result is used as the extraction range, directly excluding the interval of low-value events to reduce invalid segment interference; second, within this range, the key features of multi-modal data are called to carry out multi-dimensional evaluation: one is key frame scoring, which calculates the score (0-1) based on the saliency of core elements (such as facial expressions of characters and key props) in the frame, and selects the segment corresponding to the key frame with a score ≥0.85; two is emotion change analysis, which calculates the "emotion change slope" of the emotion time series data in the interval, and retains the emotion peak segment with a slope ≥0.6; three is visual feature analysis, which calculates the picture motion intensity through optical flow density and locates the visual focus area through attention heat map, and selects the segment with high motion intensity or focus; finally, the three types of evaluation results are weighted and fused according to "key frame score 30% + emotion change slope 30% + visual feature score 40%", and the top 30% of the segments are selected as the video highlight segment according to the comprehensive score. The output of this module ensures that the subsequent video processing result can focus on the core highlights, improving the emotional resonance and content appeal of users when watching.
[0077] The video processing module is configured to obtain a video processing result according to the multi-modal data analysis result, the narrative framework, and the video highlight segment. The module takes the multimodal data analysis result of the data analysis module, the narrative framework of the event and narrative processing module, and the video highlight segment of the highlight extraction module as triple inputs, and the core function is to complete the collaborative optimization and integration of "picture-audio-text", and output the complete video processing result. Its processing process consists of three steps: The first step is multi-element special optimization - in picture optimization, the video color style is unified based on theme clustering labels, spatio-temporal super-resolution adaptive reconstruction is performed on the highlight segment, and transition effects are selected according to the event correlation in the narrative framework; in audio optimization, background music matching the emotion is generated combined with emotion time series data, the environment sound of high-value events is preserved and the sound effect is enhanced, and the audio rhythm is aligned with the picture motion rhythm through dynamic time warping (DTW) algorithm; in text optimization, based on theme labels and highlight segment content, GPT-4 model is used to generate subtitle text, and subtitle timestamp is aligned with key frame of highlight segment.
[0078] The second step is multi-element collaborative integration - according to the time axis of the narrative framework, the optimized video sequence, audio track, and subtitle track are integrated to ensure that the timestamp deviation of the three is ≤0.3s.
[0079] The third step is redundant content elimination, which deletes the content that conflicts with the narrative framework logic or is not emotionally coherent, and finally generates a complete video that meets the mainstream format.
[0080] The output of the module marks the full-process closed loop from "multimodal data input" to "intelligent video output", ensuring that the final video has logical coherence, emotional transmission and visual aesthetics.
[0081] In summary, the device constitutes a full-process collaborative video processing system of "data analysis-logic construction-highlight screening-finished product integration", with each module progressing layer by layer and data closed loop linkage. The data analysis module converts unstructured multimodal raw data into standardized analysis results, laying the data foundation; the event and narrative processing module builds a logical event system and narrative framework based on data, clarifying the content structure; the highlight extraction module accurately screens core highlight segments in combination with data and event results, focusing on content value; the video processing module integrates the previous results to complete the collaborative optimization and integration of pictures, audio and text, and outputs a video that has narrative coherence, emotional transmission and visual aesthetics, greatly improving the automation and efficiency of video processing.
[0082] In the embodiment, the specific processing of the video processing device based on multimodal data and the technical effects brought by the specific processing can be referred to Figure 1 The related descriptions of steps S1, S2, S3 and S4 in the corresponding embodiment will not be repeated here.
[0083] It should be noted that the implementation details and technical effects of each module and unit in the device provided by the embodiments of the present disclosure can be referred to the descriptions of other embodiments in the present disclosure, which will not be repeated here.
[0084] Reference will be made below to Figure 3 which shows a structural schematic diagram of a computer system 500 suitable for implementing the electronic device of the present disclosure. Figure 3 The computer system 500 shown is only an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0085] As shown in Figure 3 , the computer system 500 can include a processing device 501, which can perform various appropriate actions and processes according to programs stored in a ROM 502 or loaded from a storage device 508 to a random access RAM 503. In the RAM 503, various programs and data required for the operation of the computer system 500 are also stored. The processing device 501, the ROM 502 and the RAM 503 are connected to each other through a bus 504. An I / O interface 505 is also connected to the bus 504.
[0086] Generally, the following devices can be connected to the I / O interface 505: input devices 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, and the like; output devices 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 508 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 509. The communication devices 509 can allow the computer system 500 to communicate with other devices wirelessly or through wires to exchange data. Although Figure 3 The computer system 500 is shown with various devices, but it is understood that all of the illustrated devices are not required to implement or have the computer system. More or less devices can alternatively be implemented or have.
[0087] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 509, or installed from the storage devices 508, or installed from the ROM 502. When the computer program is executed by the processing devices 501, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.
[0088] It should be noted that the computer-readable medium in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, a RF (radio frequency) or the like, or any suitable combination of the above.
[0089] The computer-readable medium described above can be contained in the electronic device described above; or can exist separately and not be assembled into the electronic device.
[0090] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to implement the method for processing video based on multi-modal data as shown in the embodiments and optional implementation modes thereof. Figure 1 The embodiments and optional implementation modes shown in the above embodiments and optional implementation modes show a method for processing video based on multi-modal data.
[0091] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0092] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in some cases, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0093] The units or modules described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit or module does not constitute a limitation on the unit itself. For example, the acquisition module can also be described as "acquiring a preset prompt word, the preset prompt word including a modal fusion prompt word, an attention mechanism prompt word, and / or a time correlation prompt word".
[0094] The above description is merely that of the preferred embodiments of the present disclosure and a description of the technical principles of the present disclosure. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions with the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or their equivalent features without departing from the concept of the above disclosure. For example, the technical solutions formed by the mutual replacement of the above features and the technical features with similar functions disclosed in the present disclosure (but not limited to) without departing from the concept of the above disclosure.
Claims
1. A video processing method based on multimodal data, characterized in that, include: S1. Obtain multimodal data analysis results using the original multimodal data of the video to be processed; S2. Perform event and narrative processing based on the multimodal data analysis results to obtain event processing results and narrative framework; S3. Based on the multimodal data analysis results and the event processing results, extract highlight segments to obtain video highlight segments; S4. Based on the multimodal data analysis results, the narrative framework, and the video highlight segments, obtain the video processing results; The multimodal data analysis results include sentiment time-series data, event segmentation results, topic clustering labels, and keyframe sequences. The topic clustering labels include global topics and local subtopics.
2. The video processing method based on multimodal data according to claim 1, characterized in that, S1. Using the multimodal raw data of the video to be processed, obtain the multimodal data analysis results, including: Collect multimodal raw data of the video to be processed, including video frame sequences, audio signals, associated text information and metadata; The multimodal raw data of the video to be processed is preprocessed to obtain standardized multimodal data; Using the standardized multimodal data, a multimodal sentiment fusion algorithm is used to obtain sentiment time-series data; Based on the standardized multimodal data, an image-audio collaborative segmentation strategy is adopted to obtain event segmentation results; Semantic analysis is performed on the standardized multimodal data to obtain topic clustering labels; Keyframes are extracted using the standardized multimodal data to obtain a keyframe sequence; Based on the emotional time-series data, the event segmentation results, the topic clustering labels, and the keyframe sequence, multimodal data analysis results are obtained.
3. The video processing method based on multimodal data according to claim 1, characterized in that, S2. Based on the multimodal data analysis results, perform event and narrative processing to obtain event processing results and a narrative framework, including: The importance of the event segmentation results is evaluated based on the emotional time series data to obtain the event importance evaluation results; High-value events are selected by combining the event importance assessment results with an emotional intensity threshold. The high-value events are compared using a pre-trained language model to obtain the event processing results. Based on the event processing results and the topic clustering labels, a joint model is used to perform correlation analysis to obtain the event correlation relationships; Based on the event segmentation results and the event relationships, an event timeline is constructed. Based on the target video length, a three-act narrative structure is constructed. Based on the high-value events, the event timeline, and the three-act narrative structure, obtain the narrative framework; The joint model includes a latent Dirichlet assignment model and a pre-trained language model.
4. The video processing method based on multimodal data according to claim 1, characterized in that, S3. Based on the multimodal data analysis results and the event processing results, extract highlight segments to obtain video highlight segments, including: Using the event processing results, the extraction range of the highlight fragment is determined; Based on the multimodal data analysis results and the extraction range of the highlight fragments, a multidimensional evaluation is performed to obtain multidimensional evaluation results; The multi-dimensional evaluation results are weighted and fused to obtain a comprehensive evaluation result; Based on the extraction range of the highlight segments and the comprehensive evaluation results, the highlight segments of the video are obtained.
5. The video processing method based on multimodal data according to claim 4, characterized in that, Based on the multimodal data analysis results and the extraction range of the highlight fragments, a multi-dimensional evaluation is performed to obtain multi-dimensional evaluation results, including: By utilizing the extraction range of the highlight fragments, candidate fragments for multi-dimensional evaluation are determined; Based on the keyframe sequence and the candidate segments evaluated in the multi-dimensional assessment, a content importance score is obtained to acquire the keyframe score. Based on the emotional time-series data and the candidate segments evaluated in the multi-dimensional assessment, the slope of emotional change is obtained; Based on the candidate segments evaluated in the multi-dimensional assessment, motion intensity and visual intersection are evaluated to obtain visual feature scores; A multi-dimensional evaluation result is obtained based on the keyframe score, the slope of the emotion change, and the visual feature score.
6. The video processing method based on multimodal data according to claim 3, characterized in that, S4. Based on the multimodal data analysis results, the narrative framework, and the video highlight segments, obtain the video processing results, including: Based on the narrative framework and the event processing results, extract the core transition segments; Based on the emotional time series data and the topic clustering tags, combined with the narrative framework, the video highlight segments and the core transition segments are integrated to obtain video footage; Based on the emotional timing data and the video footage, combined with the narrative framework, a dynamic time adjustment algorithm is used to obtain the audio content. Based on the topic clustering tags, the video highlight segments, and the core transition segments, combined with the narrative framework, the text content is obtained; The video processing result is obtained using the video footage, the audio content, and the text content.
7. The video processing method based on multimodal data according to claim 6, characterized in that, Using the video footage, the audio content, and the text content, the video processing result is obtained, including: Based on the narrative framework, the video footage, audio content, and text content are calibrated to obtain optimized video footage, optimized audio content, and optimized text content, respectively. Determine whether the timestamps corresponding to the optimized video, optimized audio, and optimized text are synchronized. If they are, perform the first operation; otherwise, perform the second operation. The first operation is to integrate the optimized video footage, the optimized audio content, and the optimized text content to obtain an initial video processing result, and then perform the third operation. The second operation is: determining whether the difference between the timestamps corresponding to the optimized video frame, the optimized audio content, and the optimized text content meets the deviation threshold; if so, the first operation is executed; otherwise, the fourth operation is executed. The third operation is to remove redundant content from the initial video processing result and obtain the video processing result. The fourth operation is to perform calibration processing on the video footage, the audio content, and the text content based on the narrative framework, and obtain optimized video footage, optimized audio content, and optimized text content respectively.
8. A video processing apparatus based on multimodal data, employing the method as described in any one of claims 1-7, characterized in that, include: Data analysis module, event and narrative processing module, highlight extraction module, and video processing module; The data analysis module is used to obtain multimodal data analysis results using the multimodal raw data of the video to be processed; The event and narrative processing module is used to process events and narratives based on the multimodal data analysis results, and to obtain event processing results and narrative frameworks. The highlight extraction module is used to extract highlight segments based on the multimodal data analysis results and the event processing results, and to obtain video highlight segments. The video processing module is used to obtain video processing results based on the multimodal data analysis results, the narrative framework, and the video highlight segments.
9. An electronic device, characterized in that, include: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, It stores a computer program thereon, wherein the computer program, when executed by one or more processors, implements the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Video post-editing and video synthesis optimization method
CN116847123A
High-quality video content automatic generation method and related equipment
CN120050487A
Editing method and device of highlight short video, storage medium and electronic equipment
CN121194017A
Audio data selection for video matching using generative artificial intelligence model
US12444195B1
Method and systems for dynamically featuring items within the storyline context of a digital graphic narrative
US20250139895A1
Cited By
Intelligent video analysis system based on multi-mode AI
CN121982618A