Video content textualization method and system based on multi-modal fusion
Through the multimodal fusion video content textualization method, multiple information modalities in the video are dynamically identified and processed, which solves the problems of information fragmentation and poor scene adaptability in the single modality method, and realizes efficient and accurate video content extraction and summary generation.
Patent Information
- Application Number
- CN202510813149.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-26
AI Technical Summary
Existing video content extraction methods rely on single-modal information, resulting in information fragmentation, poor scene adaptability, redundancy and omissions, making it difficult to achieve cross-modal verification and reducing the reliability of the results.
A multimodal fusion video content textualization method is adopted to collaboratively process visual, audio and text information, dynamically identify effective modalities, adaptively adjust the processing flow, combine spatiotemporal alignment with semantic association, eliminate redundant information and enhance key content.
It improves the robustness and comprehensiveness of video content extraction, enhances processing efficiency and accuracy, significantly improves adaptability to complex scenarios, reduces invalid calculations, and enhances the robustness and resource utilization of the system.
Smart Images

Figure CN120708127A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimedia data processing, and in particular to a method and system for textualizing video content based on multimodal fusion. Background Art
[0002] With the explosive growth of multimedia data, efficient analysis of video content has become a crucial foundation for artificial intelligence applications. Traditional video content extraction methods typically rely on single-modal information, such as generating subtitles solely through speech recognition or generating descriptions through image frame sampling. These methods suffer from the following significant drawbacks: information fragmentation: single-modal processing fails to fully utilize the multi-source information of a video (such as subtitles, speech, and image content), resulting in one-sided extraction results; poor scene adaptability: existing methods struggle with complex scenarios. For example, when relying on speech recognition for videos without subtitles, they are susceptible to background noise interference, or low-quality images cause optical character recognition (OCR) failures; and redundancy and omission: random sampling strategies can easily result in missing key frames, while fixed-interval sampling can lead to redundant information, impacting the efficiency and accuracy of subsequent analysis. For example, when building video summarization systems, existing techniques may overlook key information in subtitles or fail to capture dynamic image content, resulting in incomplete summaries. Furthermore, the lack of multimodal information fusion makes cross-modal verification difficult, reducing the reliability of the results. Therefore, a textualization method that can adaptively integrate multimodal information and cover all video elements is urgently needed to improve the robustness and comprehensiveness of content extraction. Summary of the Invention
[0003] To overcome the shortcomings of existing technologies, this paper proposes a method and system for textualizing video content based on multimodal fusion. This method achieves high-coverage video content extraction by collaboratively processing visual, audio, and textual information from a video. The core of this invention lies in: a multimodal collaborative detection mechanism that dynamically determines the presence of valid modalities (such as subtitles, speech, and key frames) in a video and adaptively adjusts the processing flow; a cross-modal information fusion strategy that eliminates redundant information and enhances key content through spatiotemporal alignment and semantic association; and an adaptive sampling technique that dynamically adjusts the frame sampling interval based on content density, balancing processing efficiency and information integrity.
[0004] In order to solve the above technical problems, the present invention adopts the following technical solutions: A video content textualization method based on multimodal fusion includes the following steps: Step 1. Dynamically identify valid modal information in the video, including subtitle information detection, audio information detection, and key frame information sampling; Step 2. For the subtitle information, a region clustering-based OCR enhancement method is used to extract subtitles; for the audio information, a multi-engine collaborative transcription and weight fusion strategy is used to generate speech text; for the key frame information, a descriptive text is generated; Step 3. Temporally and spatially align the subtitles, speech text, descriptive text, and video timeline and perform semantic fusion. Adjust the key frame information sampling strategy based on the fusion result feedback to form adaptive feedback.
[0005] Furthermore, the step 1 includes: Subtitle information detection: locate the text area through edge feature analysis and morphological processing, and perform text recognition on the text area based on the PaddleOCR model; Audio information detection: Classify silence, human voice, and background music based on acoustic features, and use spectrum features and dynamic range to assist in judgment; Keyframe information sampling: adaptively set the sampling interval based on scene change intensity and motion vector analysis. Furthermore, in the subtitle information detection: Use color space conversion to enhance text contrast; Connect the broken text regions through morphological closing operation; Correct the tilt angle of text lines through Hough transform; Set a key threshold to verify the accuracy of subtitle extraction.
[0006] Furthermore, in the audio information detection: The silence segment is determined based on the energy threshold method, where the energy is continuously lower than the dynamic threshold and the duration exceeds 500ms; The human voice and background music are pre-trained in the SVM model. After the MFCC features are input, the speech probability and music probability are output, and the spectrum flatness and dynamic range are combined to assist in the judgment.
[0007] Furthermore, when human voice is detected, the voice-to-text process is triggered, and the background music segment is marked as a compressible or noise reduction area.
[0008] Furthermore, in the key frame information sampling: Dynamically adjust the current MSE threshold according to the MSE fluctuation of historical frames. When the difference between frames exceeds the corresponding MSE threshold, a key frame is forcibly inserted. The inter-frame differences are smoothed by exponentially weighted averaging to avoid instantaneous noise triggering invalid keyframes.
[0009] Furthermore, in step 2, the audio information is transcribed using a multi-engine collaborative transcription and weight fusion strategy to generate speech text, including: Splitting the audio information into silent segments; Parallel calls to the cloud API and local model to transcribe and output the transcribed text along with the corresponding confidence level and confidence weight. Calculate the context matching score and context matching weight of the transcribed text output by the cloud API and the local model respectively; Based on the language model, the cloud API and local model are transcribed and the output transcribed text is scored to obtain the language model score and language model score weight; A comprehensive weight is obtained by performing three-dimensional weighted fusion based on confidence weight, context matching weight, and language model score weight; For terms with transcription differences, high-weight results are selected based on the comprehensive weight.
[0010] Furthermore, the semantic fusion in step 3 includes: Generate text vectors using Transformer encoder; Build a semantic association graph based on graph neural network and mine semantic associations between multimodal texts through node aggregation; Key nodes are extracted from the semantic association graph to generate a final fused text.
[0011] Furthermore, the adaptive feedback includes: Preliminarily determine the key frame information sampling interval; Identify video segments in the fused text whose confidence level is below a threshold; The sampling density of the video segments whose confidence level is lower than the threshold is increased.
[0012] In another aspect, the present invention provides a video content textualization system based on multimodal fusion, comprising: Collaborative detection module. This module is used to dynamically identify valid modal information in the video, including subtitle information detection, audio information detection, and keyframe information sampling; Multimodal extraction module. It is used to extract subtitles from the subtitle information using an OCR enhancement method based on regional clustering; generate speech text from the audio information using a multi-engine collaborative transcription and weight fusion strategy; and generate descriptive text from the key frame information; Fusion Optimization Module. This module is used to spatially and temporally align the subtitles, speech text, descriptive text, and video timeline and perform semantic fusion. It also adjusts the keyframe information sampling strategy based on the fusion result feedback to form adaptive feedback.
[0013] Compared with the prior art, the present invention has the following beneficial effects: By dynamically fusing the multimodal information of subtitles, audio, and keyframes in a video, an OCR enhancement method based on regional clustering is used to improve the accuracy of subtitle extraction. Combining multi-engine collaborative transcription with a three-dimensional weighted fusion strategy (confidence, context matching, and language model score) achieves a high accuracy rate of 92% in speech transcription in noisy scenarios, effectively overcoming the problem of one-sided information caused by traditional single-modal processing. An adaptive feedback mechanism is used to dynamically optimize the keyframe sampling strategy, dynamically adjust the MSE threshold based on scene complexity, and reversely increase the sampling density of low-confidence segments through semantic fusion results, significantly improving adaptability to complex scenarios and addressing the redundancy and omission defects of fixed sampling strategies. Furthermore, spatiotemporal alignment is used to eliminate multi-source timing conflicts, and a semantic association graph is constructed with the help of a transformer encoder and a graph neural network to achieve cross-modal semantic fusion and redundant information deduplication, reducing ineffective computation by more than 30% while ensuring content integrity. Combining audio silence segment compression, a lightweight real-time processing module, and cloud-local heterogeneous computing power collaboration, the system's robustness, processing efficiency, and resource utilization are comprehensively improved to meet diverse needs from high-precision analysis to low-latency streaming media scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0015] Figure 1 This is the overall flow chart of the present invention, showing the collaborative process of multimodal detection, extraction and fusion.
[0016] Figure 2 : Flowchart of an embodiment of the present invention. DETAILED DESCRIPTION
[0017] To make the above-mentioned objects, features, and advantages of the present application more clearly understood, the specific embodiments of the present application are described in detail below with reference to the accompanying drawings. The following description sets forth many specific details to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the scope of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below. Example 1 like Figure 1 As shown, the present invention provides a method for textualizing video content based on multimodal fusion, comprising the following steps: Step 1. Dynamically identify valid modal information in the video, including subtitle information detection, audio information detection, and key frame information sampling; including: Subtitle information detection: locate the text area through edge feature analysis and morphological processing, and perform text recognition on the text area based on the PaddleOCR model; Audio information detection: Classify silence, human voice, and background music based on acoustic features, and use spectrum features and dynamic range to assist in judgment; Keyframe information sampling: adaptively set the sampling interval based on scene change intensity and motion vector analysis. In this embodiment, the goal of subtitle detection is to accurately identify the subtitle region in a video and extract the text content, while eliminating background interference. The implementation process first enhances contrast through image preprocessing: the video frame is converted to a grayscale image and a Gaussian filter (kernel size 5×5) is applied to reduce noise. The luminance channel (L channel) is then extracted in the LAB color space, leveraging its high contrast to highlight the subtitle region. Canny edge detection (threshold range 50-150) is used to locate edges in the processed grayscale image, and combined with horizontal projection analysis, the focus is placed on the bottom region of the image (the lower 10%-20% of the image height). Because text regions in an image may appear discontinuous (i.e., broken) after edge detection due to various factors (such as resolution, image quality, and complex background interference), a morphological closing operation (kernel size 15×3) is used to connect the broken text regions to form candidate boxes.
[0018] The PaddleOCR model is used in the text recognition stage. Adaptive binarization (block size 31×31) is performed on candidate regions to enhance contrast, and Hough transforms are used to correct for text line skew. Recognized text undergoes two verification steps: first, temporal continuity, requiring that the text in the same region appear in at least 8 out of 10 consecutive frames to eliminate brief interruptions; and second, semantic filtering, verifying the text's plausibility by including punctuation (such as ".!?") and common conjunctions ("then" and "but"). Key thresholds include color contrast (LAB spatial brightness difference >50), OCR confidence (>0.8), and region height range (bottom 10%-20%) to ensure accurate subtitle localization and extraction.
[0019] Audio detection aims to distinguish between silence, human voice, and background music, triggering appropriate processing logic. The audio signal is first framed (25ms per frame, 10ms step size) to extract short-term mean square energy (RMS), zero-crossing rate (ZCR), and MFCC features. Silence detection uses an energy threshold method: if the energy of 10 consecutive frames falls below a dynamic threshold (typically 1%-5% of the maximum energy, or -40dB), it is considered a silence segment. This duration must exceed 500ms to avoid misidentifying brief pauses. Speech and music classification is achieved using a pre-trained SVM model. MFCC features are input and output as speech and music probabilities: a speech probability > 0.7 indicates human voice, and a music probability > 0.6 indicates background music. To enhance robustness, spectral flatness (< 0.3 for music scenes) and dynamic range (typically 30-50dB for human voice) are used to aid in the classification process. Spectral flatness measures the flatness of the audio spectrum. In music scenes, the audio signal's spectrum is typically richer and more diverse, containing multiple frequency components. Its spectral flatness is typically less than 0.3. Spectral flatness can be crucial when the speech and music probabilities output by the SVM model are close to the judgment threshold, making it difficult to clearly determine the audio category. For example, if the speech probability is 0.68 (close to 0.7) and the music probability is 0.58 (close to 0.6), examining the spectral flatness will reveal that if it is less than 0.3, even if the speech probability is slightly higher, the characteristic of spectral flatness in music scenes suggests a higher probability of classification as background music. This is because spectral flatness reflects the distribution of the audio's frequency components. Low flatness indicates a rich and even distribution of frequency components, consistent with the spectral characteristics of music. The dynamic range of human voices is typically between 30 and 50 dB, a characteristic that can be used to further confirm whether an audio source is a human voice. Similarly, when the SVM model outputs a difficult probability judgment, detecting a dynamic range between 30 and 50 dB can lead to a higher probability of classification as a human voice, even if the speech probability does not meet the 0.7 judgment threshold. For example, when the probability of speech is 0.65 and the probability of music is 0.62, in this ambiguous situation, if the audio's dynamic range is 40dB, which falls within the dynamic range of human voice, it can help determine that the audio is a human voice. Because the dynamic range reflects the amplitude of changes in audio signal strength, the volume of human voices during normal speech typically fluctuates within this specific range. When a human voice is detected, the voice-to-text process is triggered, and the background music segment is marked as a region that can be compressed or de-noised.
[0020] The core goal of keyframe sampling is to dynamically adjust the interval between keyframes (I-frames) based on changes in video content, optimizing compression efficiency while maintaining video quality. The implementation process begins with scene change detection: the normalized mean square error (MSE) of consecutive frames is calculated to determine whether a scene has changed significantly. To account for varying scene complexity, the MSE threshold is dynamically adjusted. For simple scenes, such as those with a single subject and a simple background, the MSE threshold is set at 8%. For complex scenes with multiple moving objects and frequent changes in lighting and shadow, the MSE threshold is increased to 12%. When the inter-frame difference exceeds the threshold, it is determined to be a scene change or significant motion, and a keyframe is forcibly inserted. A dynamic adjustment mechanism is also introduced to adjust the current MSE threshold based on MSE fluctuations over the previous period of the video. The average and standard deviation of the MSE over the past N frames are calculated. A small standard deviation indicates a relatively stable scene, and the current MSE threshold is appropriately lowered. A large standard deviation indicates complex scene changes, and the MSE threshold is raised. For gradual changes (such as camera pans), optical flow analysis is used to analyze motion vectors. This strategy not only focuses on the proportion of pixels with motion amplitudes greater than a preset value (e.g., >30 pixels) but also calculates the distribution of motion vectors in different directions. In scenes with gradual changes, motion vector directions are relatively concentrated and their amplitudes show a gradual change. The video is divided into multiple subregions, and the relevant parameters of the motion vectors in each subregion are calculated. If the motion vectors in most subregions show a concentrated direction and gradually changing amplitudes, and if the proportion of pixels with motion amplitudes greater than a preset value (e.g., >30 pixels) exceeds 15%, the scene is considered gradual, triggering keyframe insertion. The dynamic adjustment strategy combines content complexity and coding constraints, comprehensively considering the results of scene change detection (MSE) and optical flow analysis of motion vectors, and assigns weights to different judgment factors. MSE reflects the change of the entire image and is assigned a weight of 0.6; the proportion of motion vector pixels reflects local motion and is assigned a weight of 0.4. This weighted calculation results in a comprehensive dynamic judgment index. When this index exceeds 0.7, the scene is considered high dynamic; when it is below 0.3, it is considered static. Weights are also adjusted based on video type. For videos primarily featuring human action (such as dance videos), the weight of the motion vector pixel ratio is increased to 0.5, and the MSE weight is adjusted to 0.5. For videos primarily featuring scene changes (such as those with changing scenery), the MSE weight is increased to 0.7, and the weight of the motion vector pixel ratio is adjusted to 0.3. For static scenes (such as conference presentations), the keyframe interval is extended to 10 seconds (corresponding to a GOP length of 300 frames, assuming a 30fps frame rate), while for highly dynamic scenes (such as sporting events), the interval is shortened to 1 second (with a GOP length of 30 frames). To balance real-time performance and quality, an adaptive learning mechanism is introduced: motion trends are predicted by analyzing historical frames. If the motion intensity increases for five consecutive frames (MSE increase >5% / frame), a keyframe is inserted in advance to avoid sudden freezes.Key thresholds include the MSE threshold for scene changes (dynamically adjusted based on scene complexity), the motion vector amplitude threshold (>30 pixels), and the motion pixel ratio threshold (≥15%). In actual deployments, these thresholds need to be dynamically calibrated based on encoder characteristics (such as H.264 / H.265). For example, in low bitrate mode, the MSE threshold can be appropriately lowered to 8% to increase keyframe density and prevent image blur caused by cumulative errors. Furthermore, a smooth transition strategy is implemented, using an exponentially weighted average (α = 0.25) to smooth inter-frame differences and prevent invalid keyframes triggered by transient noise (such as flash).
[0021] Step 2. For the subtitle information, a region clustering-based OCR enhancement method is used to extract subtitles; for the audio information, a multi-engine collaborative transcription and weight fusion strategy is used to generate speech text; for the key frame information, a descriptive text is generated; (1) Subtitle extraction: If valid subtitles are detected, an OCR enhancement method based on region clustering is used, which includes: spatial and temporal alignment of subtitle regions of consecutive frames to eliminate jitter interference; using multi-scale binarization (Adaptive Thresholding) to improve the recognition rate of low-contrast subtitles; and correcting OCR errors through semantic verification (N-gram language model).
[0022] (2) Speech transcription: If a valid human voice is detected, a fragmentation processing and multi-engine fusion strategy is adopted: the audio is divided into segments according to the silent segments, and the cloud API (Google Speech-to-Text) and the local model (Whisper) are used for transcription respectively; the final text is generated by weighted fusion of the confidence results. In the weighted fusion method of multi-engine speech transcription, the audio is first divided into multiple segments according to the silent segments, and each audio segment is transcribed using the cloud API (Google Speech-to-Text) and the local model (Whisper). The two engines will each output the transcribed text and the corresponding confidence. For example, for an audio segment, the transcribed text output by Google Speech-to-Text is "The weather is very good today" with a confidence of 0.8; the transcribed text output by Whisper is "The weather is pretty good today" with a confidence of 0.75. Next, the weight calculation is performed. First, the confidence weight (accounting for 60%) is calculated. Suppose the confidence of Google Speech-to-Text is , Whisper's confidence is , then the confidence weight is calculated as , , in the above example, , Meanwhile, calculate the context matching degree weight (accounting for 20%). Analyze the context matching degree of each transcription result with the transcribed texts of the previous and subsequent audio segments. Use the edit distance (Levenshtein distance) to measure. Calculate the sum of the edit distances between the transcribed text of the current segment and the transcribed texts of the previous and subsequent segments. The smaller the distance, the higher the matching degree. Assume the context matching degree score of the Google Speech-to-Text transcription result , and the context matching degree score of the Whisper transcription result , then the context matching degree weight is calculated as , , here , . Then calculate the language model score weight (accounting for 20%). Use a pre-trained language model (such as GPT, etc.) to score each transcription result. The higher the score, the more in line with natural language expression the transcription result is. Assume the language model score of the Google Speech-to-Text transcription result , and the language model score of the Whisper transcription result , then the language model score weight is calculated as , , here , . The final comprehensive weight of Google Speech-to-Text:
[0023] The final comprehensive weight of Whisper:
[0024] In the text fusion stage, compare the two transcribed texts word by word. For words at the same position, if the transcriptions of the two models are the same, directly include the word in the final fused text; if they are different, select according to the comprehensive weight. For example, at a certain position, Google Speech-to-Text transcribes as "happy", and Whisper transcribes as "glad". Since , then include "happy" in the final fused text. For special cases, if the transcribed text of a certain model is longer than the other, for the extra part, if the comprehensive weight of the model where it is located is greater than 0.6, and this part has a certain logical coherence with the context (which can be simply judged by the language model), then include the extra part in the final fused text. Through the above method, considering multiple factors for weighted fusion can obtain a more accurate and practical speech transcription result.
[0025] (3) Image description generation: Use the pre-trained vision-language model (BLIP-2) to generate description text for the sampled frames, and filter key descriptions based on the attention mechanism.
[0026] Step 3. Temporally and spatially align the subtitles, speech text, descriptive text, and video timeline and perform semantic fusion. Adjust the keyframe information sampling strategy based on the fusion result feedback to form adaptive feedback. This includes: (1) Spatiotemporal alignment: Match subtitles, voice text and video timeline to eliminate timing conflicts (2) Semantic Fusion: First, the multimodal text is vectorized using a Transformer encoder. The Transformer encoder is built based on the self-attention mechanism, and its architecture contains multiple identical encoding layers. Each encoding layer consists of two parts: a multi-head attention mechanism and a feedforward neural network, with layer normalization and residual connections interspersed in between. For the input multimodal text sequence, it first passes through the embedding layer to obtain the embedding vector, which is then input into the multi-head attention mechanism. The multi-head attention mechanism performs a linear transformation on the input to obtain the query vector, key vector, and value vector respectively. The attention score is calculated by dot product, then normalized by the softmax function, and finally multiplied with the value vector to obtain the attention output. The multi-head attention mechanism performs these operations multiple times in parallel, concatenates the results, and then outputs them through a linear transformation. After the multi-head attention mechanism, the output enters the feedforward neural network, which contains two fully connected layers with a ReLU activation function in between. In each encoding layer, the output of the multi-head attention mechanism and the feedforward neural network undergoes layer normalization and residual connections to accelerate training and avoid the gradient vanishing problem. By stacking multiple such encoding layers, the Transformer encoder can perform in-depth feature extraction and semantic understanding of multimodal text, converting multimodal text into vector representations with rich semantic information.
[0027] Next, a graph neural network (GNN) is used to construct a semantic association graph and extract key nodes. In semantic fusion scenarios, a graph consists of a set of nodes and edges. Nodes can represent words, phrases, or sentences in multimodal text, while edges represent the semantic associations between them. Taking a graph convolutional network (GCN) as an example, for each node, its initial feature vector can be the corresponding vector output by the Transformer encoder. Convolution operations are performed on the graph to aggregate feature information from the node's neighbors. During the convolution operation, node features are updated based on the node's neighbor set, node degree, and a learnable weight matrix and bias vector, combined with an activation function. Through multiple layers of graph convolution operations, a node gradually aggregates semantic information from its neighbors, capturing a wider range of semantic associations. Another common graph neural network architecture, the Graph Attention Network (GAT), introduces an attention mechanism to calculate weights between nodes, allowing the model to more flexibly focus on the importance of different neighboring nodes to the current node. Node features are updated by calculating an attention coefficient and then weightedly aggregating the neighboring node features. In the semantic fusion method of this invention, a GNN (such as GCN or GAT) is used to construct a semantic association graph. By updating and propagating node features, semantic associations between multimodal texts are mined. Finally, key nodes are extracted from the constructed semantic association graph. These key nodes represent the key semantic information in the multimodal text and are used to generate the final summary text. In this way, the Transformer-based encoder and the graph neural network (GNN) work together to achieve semantic fusion of multimodal texts, and then extract key nodes to generate the final summary text.
[0028] (3) Adaptive feedback: In step 1, the key frame sampling interval is preliminarily determined by calculating MSE, optical flow, and other methods, combined with content complexity and coding constraints. After the fusion results are obtained in the subsequent processing, if it is found that some segments have low confidence, adaptive feedback will increase the frame sampling density of these low-confidence segments. This is a correction and supplement to the key frame sampling strategy in step 1. The adjusted sampling strategy provides richer information for subsequent multimodal processing, forming a closed loop from sampling in step 1 to adaptive feedback optimization, thereby improving the accuracy of video content textualization. Preferred solution: (1) Subtitle detection optimization: When the video resolution is below the threshold, the super-resolution reconstruction module is enabled to pre-process the subtitle area. (2) Multi-language support: The language detection module is integrated into the OCR and speech recognition stages to automatically switch to the corresponding language model. (3) Real-time processing mode: A lightweight model (MobileNet + CTC) is used to achieve end-to-end real-time text conversion, suitable for streaming scenarios.
[0029] like Figure 2 As shown, this embodiment takes news video content extraction as an example and includes the following steps: Step 1: Input a 5-minute news video (including hard subtitles, clear audio and host screen).
[0030] Step 2: Perform multimodal detection on the video.
[0031] Subtitle detection: Video frames are first converted to grayscale images. A Gaussian filter with a 5×5 kernel is used to reduce noise. The L channel is extracted in the LAB color space to highlight the subtitle area. Canny edge detection (thresholds 50-150) is used to locate edges. Combined with horizontal projection analysis, the bottom 10%-20% of the frame is focused on. Morphological closing with a 15×3 kernel is used to connect disjointed text regions to form candidate boxes. The PaddleOCR model is used for text recognition. Adaptive binarization of the candidate regions using 31×31 blocks is performed to enhance contrast. Hough transform is used to correct text line skew. Recognized text undergoes two verification steps: first, the same text region must appear in at least 8 out of 10 consecutive frames. Second, the text is verified for legitimacy by using punctuation (such as ".!?") and common conjunctions ("then" and "but"). Testing shows that the subtitle region is concentrated at the bottom, with a text density of 0.8 (threshold 0.5), indicating the presence of valid subtitles.
[0032] Audio detection: The audio signal is framed (25ms per frame, 10ms step size) and short-term mean square energy (RMS), zero-crossing rate (ZCR), and MFCC features are extracted. Silence is detected using an energy threshold method. Silence is determined if the energy for 10 consecutive frames is below 1%-5% of the maximum energy (i.e., -40dB) and lasts for more than 500ms. A pre-trained SVM model is used to classify speech and music based on the probability values output by the MFCC features. Speech probabilities greater than 0.7 are considered human voices, and music probabilities greater than 0.6 are considered background music. Spectral flatness (music scene flatness <0.3) and dynamic range (human voice dynamic range 30-50dB) are also used to assist in this judgment. The audio signal-to-noise ratio (SNR) of the news video is 25dB, indicating that the speech is recognizable.
[0033] Keyframe sampling: The normalized mean square error (MSE) of consecutive frames is calculated. When the difference between frames is ≥10%, it is determined to be a scene change or significant motion, and a keyframe is inserted. For gradual changes (such as camera pans), optical flow is used to analyze motion vectors and count the percentage of pixels with a motion amplitude >30 pixels. If this percentage is ≥15%, a keyframe is inserted. Because news videos resemble static scenes (such as conference presentations), the keyframe interval is extended to 10 seconds (corresponding to a GOP length of 300 frames, assuming a 30fps frame rate). An adaptive learning mechanism is introduced to predict motion trends by analyzing historical frames. If the motion intensity increases for five consecutive frames (MSE increase >5% / frame), a keyframe is inserted in advance. Inter-frame differences are smoothed using exponentially weighted averaging (α = 0.25). Dynamic sampling is performed based on the motion amplitude (calculated by optical flow), with an initial sampling interval of 2 seconds.
[0034] Step 3: Extract information from the video.
[0035] Subtitle extraction: After valid subtitles are detected, an OCR enhancement method based on region clustering is applied. Temporal and spatial alignment of subtitle regions in consecutive frames eliminates jitter interference. Multi-scale binarization (Adaptive Thresholding) is used to improve the recognition rate of low-contrast subtitles. Semantic verification (N-gram language model) corrects OCR errors. Video Super Resolution (VSR) is used to upscale 720p videos to 1080p, improving OCR accuracy by 18%.
[0036] Speech transcription: Once valid speech is detected, a fragmented processing and multi-engine fusion strategy is employed. The audio is segmented into segments based on silence, and transcribed using both a cloud-based API (Google Speech-to-Text) and a local model (Whisper). The results are then weighted and fused, achieving an accuracy rate of 92%.
[0037] Image description generation: A pre-trained vision-language model (BLIP-2) is used to generate description text for sampled frames. Key descriptions are selected based on the attention mechanism. For example, a description such as "middle-aged man in a suit, with a news studio behind him" is generated for a close-up frame of the host.
[0038] Step 4: Fusion optimization.
[0039] Spatiotemporal alignment: Match subtitles and audio text to the video timeline to eliminate timing conflicts.
[0040] Semantic Fusion: We vectorized multimodal text using a Transformer-based encoder, removed duplicates using cosine similarity, and constructed a semantic association graph based on a graph neural network (GNN). We found that the duplication rate between subtitles and audio content exceeded 70%. After deduplication, we retained the audio text (which contains more complete information) and extracted the keywords "epidemic," "economic policy," and "expert interpretation" from the semantic association graph.
[0041] Generate summary: Based on the extracted keywords, combined with the theme and key information of the video, generate the final video summary and complete the textual processing of the video content.
[0042] The present invention provides a method for textualization of video content based on multimodal fusion. By using the method in this article, the textualization result of the video can be obtained by simply inputting a video in MP4 format.
[0043] Example 2 This embodiment provides a video content textualization system based on multimodal fusion, including: Collaborative detection module. This module is used to dynamically identify valid modal information in the video, including subtitle information detection, audio information detection, and keyframe information sampling; Multimodal extraction module. It is used to extract subtitles from the subtitle information using an OCR enhancement method based on regional clustering; generate speech text from the audio information using a multi-engine collaborative transcription and weight fusion strategy; and generate descriptive text from the key frame information; Fusion Optimization Module. This module is used to spatially and temporally align the subtitles, speech text, descriptive text, and video timeline and perform semantic fusion. It also adjusts the keyframe information sampling strategy based on the fusion result feedback to form adaptive feedback.
[0044] It should be understood that the parts not elaborated in detail in this specification belong to the prior art. The above description of the preferred embodiment is relatively detailed, but it cannot be regarded as limiting the scope of protection of the patent of the present invention. Under the guidance of the present invention, ordinary technicians in this field can also make substitutions or modifications without departing from the scope of protection of the claims of the present invention, which all fall within the scope of protection of the present invention. The scope of protection requested by the present invention shall be based on the attached claims.
[0045] It should be understood that parts not elaborated in detail in this specification belong to the prior art.
[0046] It should be understood that the above description of the preferred embodiment is relatively detailed and cannot be regarded as limiting the scope of protection of the patent of the present invention. Under the guidance of the present invention, ordinary technicians in this field can also make substitutions or modifications without departing from the scope of protection of the claims of the present invention, which all fall within the scope of protection of the present invention. The scope of protection requested by the present invention shall be based on the attached claims.
Claims
1. A video content textualization method based on multimodal fusion, characterized in that: The following steps are involved: Step 1. Dynamically identify valid modal information in the video, including subtitle information detection, audio information detection, and key frame information sampling; Step 2. Extracting subtitles from the subtitle information using an OCR enhancement method based on region clustering; For the audio information, a multi-engine collaborative transcription and weight fusion strategy is adopted to generate speech text; generating descriptive text for the key frame information; Step 3. Temporally and spatially aligning the subtitles, speech text, descriptive text, and video timeline and performing semantic fusion; The key frame information sampling strategy is adjusted according to the fusion result feedback to form adaptive feedback.
2. The method for converting video content into text based on multimodal fusion according to claim 1, characterized in that: The step 1 comprises: Subtitle information detection: locate the text area through edge feature analysis and morphological processing, and perform text recognition on the text area based on the PaddleOCR model; Audio information detection: Classify silence, human voice, and background music based on acoustic features, and use spectrum features and dynamic range to assist in judgment; Keyframe information sampling: Adaptively set the sampling interval based on scene change intensity and motion vector analysis.
3. The method for converting video content into text based on multimodal fusion according to claim 2, characterized in that: During the subtitle information detection: Use color space conversion to enhance text contrast; Connect the broken text regions through morphological closing operation; Correct the tilt angle of text lines through Hough transform; Set a key threshold to verify the accuracy of subtitle extraction.
4. The method for converting video content into text based on multimodal fusion according to claim 2, characterized in that: During the audio information detection: The silence segment is determined based on the energy threshold method, where the energy is continuously lower than the dynamic threshold and the duration exceeds 500ms; The human voice and background music are pre-trained in the SVM model. After the MFCC features are input, the speech probability and music probability are output, and the spectrum flatness and dynamic range are combined to assist in the judgment.
5. The method for converting video content into text based on multimodal fusion according to claim 4, characterized in that: When human voice is detected, the voice-to-text process is triggered, and the background music segment is marked as a compressible or noise reduction area.
6. The method for converting video content into text based on multimodal fusion according to claim 2, characterized in that: In the key frame information sampling: Dynamically adjust the current MSE threshold according to the MSE fluctuation of historical frames. When the difference between frames exceeds the corresponding MSE threshold, a key frame is forcibly inserted. The inter-frame differences are smoothed by exponentially weighted averaging to avoid instantaneous noise triggering invalid keyframes.
7. The method for converting video content into text based on multimodal fusion according to claim 1, characterized in that: In step 2, the audio information is transcribed using multiple engines and a weighted fusion strategy to generate speech text, including: Splitting the audio information into silent segments; Parallel calls to the cloud API and local model to transcribe and output the transcribed text along with the corresponding confidence level and confidence weight. Calculate the context matching score and context matching weight of the transcribed text output by the cloud API and the local model respectively; Based on the language model, the cloud API and local model are transcribed and the output transcribed text is scored to obtain the language model score and language model score weight; A comprehensive weight is obtained by performing three-dimensional weighted fusion based on confidence weight, context matching weight, and language model score weight; For terms with transcription differences, high-weight results are selected based on the comprehensive weight.
8. The method for converting video content into text based on multimodal fusion according to claim 1, characterized in that: The semantic fusion in step 3 includes: Generate text vectors using Transformer encoder; Build a semantic association graph based on graph neural network and mine semantic associations between multimodal texts through node aggregation; Key nodes are extracted from the semantic association graph to generate a final fused text.
9. The method for converting video content into text based on multimodal fusion according to claim 8, characterized in that: The adaptive feedback includes: Preliminarily determine the key frame information sampling interval; Identify video segments in the fused text whose confidence level is below a threshold; The sampling density of the video segments whose confidence level is lower than the threshold is increased.
10. A video content textualization system based on multimodal fusion, characterized in that: include: Collaborative detection module. This module is used to dynamically identify valid modal information in the video, including subtitle information detection, audio information detection, and keyframe information sampling; Multimodal extraction module. It is used to extract subtitles from the subtitle information using an OCR enhancement method based on regional clustering; and to generate speech text from the audio information using a multi-engine collaborative transcription and weight fusion strategy; generating descriptive text for the key frame information; A fusion optimization module is used to align the subtitles, speech text, descriptive text, and video timeline in time and space and perform semantic fusion; Adjusting the key frame information sampling strategy according to the fusion result feedback to form adaptive feedback; The video content textualization system based on multimodal fusion is used to execute the steps in the video content textualization method based on multimodal fusion according to any one of claims 1 to 9.
Citation Information
Cited By
Intelligent video content extraction and rapid positioning system based on multi-modal fusion and space-time perception
CN121353850A
Audio and video semantic enhancement processing method based on artificial intelligence
CN121438817A