Content extraction method and device, storage medium and electronic equipment
By extracting and fusing multimodal features and combining them with structured descriptive text based on global label data, the problem of incomplete information caused by single-modal processing in existing technologies is solved, and efficient management and retrieval of video data are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE INTERNET CO LTD
- Filing Date
- 2025-12-05
- Publication Date
- 2026-04-21
AI Technical Summary
Existing video content extraction technologies mainly rely on single-modal data (such as subtitles or transcribed text), which cannot fully capture the multi-dimensional features of video data. This results in fragmented and incomplete information in the extraction results, making it difficult to meet the needs of efficient management and retrieval in scenarios such as cloud storage.
By identifying the text, audio, and visual multimodal feature data of the video, aligning the feature dimensions of each modality by minimizing the intermodal matching cost, and performing feature fusion, the content extraction is finally achieved by combining the structured descriptive text of the global label data with a multimodal large model.
It improves the completeness and accuracy of content extraction results, supports precise content browsing, multi-dimensional intelligent retrieval, and automatic topic classification in scenarios such as cloud storage, and reduces user operation costs.
Smart Images

Figure CN121901447A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a content extraction method, apparatus, storage medium, and electronic device. Background Technology
[0002] In the digital age, video has become one of the core carriers of information dissemination and storage, especially in personal digital asset storage scenarios such as cloud storage, where video data exhibits characteristics such as massive growth, diverse types, and significant length variations. Video content extraction technology, as a key means of efficiently managing, retrieving, and utilizing video assets, is widely used in personal digital asset storage scenarios.
[0003] However, existing video content extraction technologies primarily focus on processing single-modal data, relying mainly on video subtitles or transcribed text to perform tasks. This single-modal data processing approach cannot fully capture the multi-dimensional features of video data, resulting in incomplete and fragmented extraction results. Summary of the Invention
[0004] In view of this, this application provides a content extraction method, apparatus, storage medium, and electronic device. By determining the text, audio, visual, and global tag multimodal feature data of the data to be processed (video), the method aligns the dimensions of each modal feature with the goal of minimizing the matching cost between modal data. Then, it fuses the multimodal features with unified dimensions and finally combines them with the structured descriptive text of the video's global tag data, utilizing a large multimodal model to achieve content extraction. This overcomes the limitations of existing single-modal processing, fully explores the complementary value of multimodal data, and improves the completeness and accuracy of the content extraction results.
[0005] Firstly, this application provides a content extraction method, including: Identify the multimodal feature data of the data to be processed; the multimodal feature data includes text feature data, audio feature data, and visual feature data.
[0006] With the goal of minimizing the matching cost between different modal feature data, the feature dimensions of text feature data, audio feature data, and visual feature data are aligned to obtain multimodal feature data with unified feature dimensions.
[0007] Feature fusion is performed on multimodal feature data with uniform feature dimensions to obtain multimodal fused features.
[0008] Using structured descriptive text containing multimodal fusion features and global labels of the data to be processed as input, the content extraction results are obtained by using a multimodal large model adapted to the structured descriptive text.
[0009] Secondly, this application provides a content extraction device, comprising: The determination module is configured to determine the multimodal feature data of the data to be processed; the multimodal feature data includes text feature data, audio feature data, and visual feature data.
[0010] The alignment module is configured to align the feature dimensions of the text feature data, the audio feature data, and the visual feature data with the goal of minimizing the matching cost between different modal feature data, so as to obtain multimodal feature data with unified feature dimensions.
[0011] The fusion module is configured to perform feature fusion on the multimodal feature data with the same feature dimensions to obtain multimodal fused features.
[0012] The output module is configured to take the multimodal fusion features and the structured descriptive text of the global labels of the data to be processed as input, and use the multimodal large model adapted to the structured descriptive text to obtain the content extraction result.
[0013] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0014] Fourthly, this application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect.
[0015] Fifthly, this application provides a computer program product having a computer program stored thereon, wherein the computer program product, when executed by a processor, implements the method described in the first aspect.
[0016] In view of the above embodiments, this application provides a content extraction method, apparatus, storage medium, and electronic device. Addressing the problem that existing video content extraction technologies rely on a single modality (such as subtitles or transcribed text), resulting in incomplete and fragmented information, this method first determines the text, audio, and visual multimodal feature data of the video to be processed, breaking the limitations of single-modal processing and fully covering the semantic, auditory, and visual multi-dimensional information of the video data. Then, aiming to minimize the matching cost between different modal features, it aligns the dimensions of each modal feature, solving the problem of ineffective feature association caused by differences in modal dimensions, laying the foundation for multimodal information fusion. Subsequently, it fuses multimodal features with unified dimensions to further explore the complementary value of text, audio, and visual data, avoiding extraction bias caused by the lack of single-modal information. Finally, the structured descriptive text of multimodal fusion features and global label data is used as input, and the processing results are obtained by using a multimodal large model. This ensures that the extracted results have both the comprehensiveness of multimodal data and conformity to the global semantics of the video, thereby significantly improving the completeness and accuracy of the content extraction results and effectively adapting to the needs of efficient management and retrieval of massive and diverse video assets in scenarios such as cloud storage.
[0017] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart illustrating a content extraction method provided in an embodiment of this application is shown; Figure 2 This illustration shows an example of a video frame deduplication process applicable to a content extraction method provided in this application embodiment; Figure 3 A flowchart illustrating a content extraction method provided in an embodiment of this application is shown; Figure 4 This illustration shows an example of video frame grouping applicable to a content extraction method provided in this application embodiment; Figure 5 An example diagram of feature compression applicable to a content extraction method provided in this application is shown; Figure 6 A flowchart illustrating a content extraction method provided in an embodiment of this application is shown; Figure 7 A flowchart illustrating a content extraction method provided in an embodiment of this application is shown; Figure 8 The diagram shows a structural example of a multimodal feature fusion model provided in an embodiment of this application. Figure 9 A flowchart illustrating a content extraction method provided in an embodiment of this application is shown; Figure 10 A flowchart illustrating a content extraction method provided in an embodiment of this application is shown; Figure 11 The illustration shows an example scenario applicable to the content extraction method provided in this application embodiment; Figure 12 The illustration shows an example scenario applicable to the content extraction method provided in this application embodiment; Figure 13 The illustration shows an example scenario applicable to the content extraction method provided in this application embodiment; Figure 14 This paper shows a structural example diagram of a content extraction system provided in an embodiment of this application; Figure 15 A schematic diagram of the structure of a content extraction device provided in an embodiment of this application is shown. Detailed Implementation
[0021] To facilitate the explanation of the embodiments of this application, the application scenarios related to the embodiments of this application are first introduced below, the problems existing in the prior art in these scenarios are explained, and then the solution and value of the solution of this application are explained.
[0022] The content extraction method provided in this application is primarily applied to the scenario of personal digital asset storage and management on cloud storage. Currently, cloud storage has become the core carrier for users to store video assets such as course videos, home videos, film clips, and work materials. Consequently, users have an urgent need for efficient management of these video assets. For example, users may want to quickly identify the core knowledge points of a one-hour course video to decide whether to watch it, accurately locate clips containing specific individuals from hundreds of home videos, or automatically organize scattered video files by theme such as travel or parties, thereby improving the utilization efficiency and management convenience of video assets.
[0023] However, existing technologies have significant drawbacks in the aforementioned scenarios, making it difficult to meet these needs: First, existing video content extraction technologies largely rely on a single modality—subtitles or transcribed text—failing to capture key visual information in video frames (such as formulas written on a blackboard in a course video) and emotional details in audio (such as the teacher's emphasis during explanations). This results in incomplete information in the extraction results, requiring users to repeatedly watch video clips to determine the content's value, leading to extremely low efficiency. Second, existing video content extraction technologies do not address the dimensional differences between text, audio, and visual features (e.g., the semantic vector dimension of text is 512, while the visual frame feature dimension is 1024). Multimodal information cannot be effectively correlated, preventing cloud storage from achieving multi-dimensional retrieval combining text descriptions and visual images (e.g., retrieving videos containing text about a sunset at the beach and images showing a red sunset). Furthermore, they cannot automatically classify content based on multimodal information (e.g., failing to categorize videos with audio containing "Happy Birthday" and images showing a cake as "Birthday Party"). Users still need to manually annotate and organize the content, resulting in high operational costs. Third, when processing long videos (such as 2-hour documentaries and 3-hour meeting recordings) that make up a large proportion of cloud storage, the system does not combine global video tags (such as the main scene of a documentary being the African savanna or the core topic of a meeting recording being project progress). Instead, it relies solely on local text fragments to generate extracted results, which often results in problems such as missing the core information in the second half of the summary and incorrect classification, leading to poor adaptability.
[0024] To address the aforementioned scenarios and the problems existing in current technologies, this application provides a content extraction method. First, the multimodal feature data of the data to be processed (video) is determined, specifically including text features (subtitles, speech-transcribed text), audio features (speech semantics, emotion features), and visual features (frame scenes, character actions). Next, aiming to minimize intermodal matching costs, the dimensions of the three types of features are aligned to eliminate collaboration barriers caused by dimensional differences. Subsequently, multimodal features with unified dimensions are fused using cross-modal attention mechanisms (e.g., associating visual features of blackboard formulas with audio features of teacher explanations to strengthen semantic connections between modalities). Finally, combined with the structured descriptive text of the video's global tags (e.g., the main knowledge point of the course video is calculus derivation, and the total duration is 60 minutes), the multimodal fused features and the structured descriptive text of the global tag data are input into a multimodal large model to generate a complete content extraction result, which includes video summaries with formula screenshots and key audio clips, as well as precise calculus course classification tags.
[0025] This method overcomes the information limitations of a single modality through multimodal feature extraction, dimensional alignment, and fusion. It also ensures the accuracy of long video processing through structured descriptions using global tags. Ultimately, it enables cloud storage to achieve precise content browsing, multi-dimensional intelligent retrieval, and automatic topic classification, significantly reducing user operating costs and improving the management and utilization efficiency of video assets.
[0026] It should be noted that the above-mentioned cloud-based personal digital asset storage and management scenario is merely an example. The content extraction method provided in this application can also be applied to other similar video content processing scenarios. For example, the core knowledge point extraction scenario of course videos on online education platforms helps learners quickly grasp the key points of the course. The segment summary generation scenario of film and television platforms, which recommends video highlights to users, and the core topic organization scenario of corporate meeting recordings, which helps employees efficiently review key meeting information, are all applicable to this application's solution, and will not be elaborated further here.
[0027] The embodiments of this application will now be described in more detail with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0028] This embodiment provides a content extraction method, such as Figure 1 As shown, the method includes: S101. Determine the multimodal feature data of the data to be processed.
[0029] The multimodal feature data includes text feature data, audio feature data, and visual feature data. For example, the data to be processed is video data uploaded by users to the cloud drive (such as movies, TV series, open courses, user-shot footage, etc.).
[0030] In some examples, after preprocessing and parsing the multimodal data in the video, the sensing signals of each modality are integrated, and the sensing signals of each modality are converted into feature data (vectors) through an algorithm model as input to the subsequent multimodal video model.
[0031] For example, text feature data F can be determined through the following steps a1 to a3. T .
[0032] a1. Multi-source text collection and integration.
[0033] Among them, multi-source texts include, for example, native subtitle text, audio transcribed text, and emotion recognition text.
[0034] For native subtitle text, multimedia processing tools (such as ffmpeg) can be used to extract native subtitle text from videos.
[0035] For audio-to-text transcription, the Fun ASR open-source speech recognition toolkit can be used to transcribe video audio streams and generate timestamped audio-to-text transcriptions.
[0036] For emotion recognition text, a general speech emotion representation model (such as Emotion2vec) can be used to perform emotion recognition on the sentences corresponding to the audio transcribed text, generating emotion recognition text with emotion tags (such as "happy" and "serious").
[0037] After obtaining the multi-source text, taking the timestamp as the core correlation dimension, compare the content of different text sources within the same time interval, remove duplicate text fragments and perform text integration.
[0038] a2. Pre-grouping of text structuring.
[0039] Exemplarily, split the integrated text into independent sentence units according to the semantic integrity of the sentences, assign a unique index to each sentence unit, and associate its corresponding timestamp (accurate to the second level), forming a structured data format of "index - timestamp - sentence content - emotion label".
[0040] For the scenario where the long video text data volume is large, an information gain algorithm can be additionally introduced to calculate the information contribution degree of each sentence unit, divide the sentence units with close semantic association and similar information gain into the same text group, implement pre-grouping processing of the text, and provide structured input for subsequent encoding.
[0041] a3. Encoding of text semantic features.
[0042] Exemplarily, perform word segmentation on the pre-grouped structured text (use word segmentation tools such as jieba in the Chinese scenario and NLTK in the English scenario), filter out meaningless stop words (such as "的", "了", "a", "the"), and retain the core semantic vocabulary.
[0043] Use the BLIP-2 multi-modal pre-training algorithm as the text encoder, input the segmented text sequence into the model, and through the text encoder module of the model (such as the semantic conversion structure combined with Q-Former and ViT), convert the text semantic information into a fixed-dimensional embedding vector sequence, and finally output the text feature F T 。
[0044] In some examples, the audio feature data F can be obtained through the following steps b1~b2 A 。
[0045] b1. Preprocessing of audio signals.
[0046] First, perform format standardization on the original audio stream separated from the video, such as uniformly converting it to the mp3 format to ensure compatibility for subsequent processing. For the possible noise (such as environmental noise, current noise) and signal attenuation problems in the audio, perform pre-emphasis (enhance the high-frequency signal intensity to compensate for the high-frequency loss during audio transmission) and noise removal (use algorithms such as spectral subtraction or wavelet transform to filter out the noise components) operations in sequence to optimize the audio signal quality and lay a foundation for subsequent feature extraction.
[0047] b2. Encoding of audio semantic features.
[0048] The preprocessed audio signal is input into the WavLM-large encoder. The time-domain and frequency-domain features of the audio are analyzed through the multi-layer Transformer structure of the model, generating a fixed-dimensional feature vector sequence that can represent the semantics of the audio (such as speech content and intonation changes). Finally, the audio feature F is output. A .
[0049] In some examples, the audio feature data F can be obtained through the following steps c1~c2. V .
[0050] c1. Video frame filtering and preprocessing.
[0051] Based on the total video duration and feature extraction requirements, a preset number of video frames are uniformly extracted (the number of frames n can be flexibly adjusted according to the video length, such as n=900 for long videos), forming a video frame set X={x1, x2, ..., xn}. The extracted video frames undergo standardization preprocessing, including size normalization (adjusting to a uniform resolution, such as 224×224) and pixel value normalization (mapping pixel values to the [0, 1] interval), eliminating the interference of inter-frame size and brightness differences on feature extraction.
[0052] c2. Visual semantic feature encoding.
[0053] The preprocessed video frames are input one by one into the CLIP-ViT-L model. The model's visual Transformer structure analyzes visual elements such as scenes, people, and targets within each frame, generating fixed-dimensional feature vectors that represent visual semantics. The feature vectors from all frames are then integrated to form a visual feature sequence F. V .
[0054] In some examples, the global label data G(X) is determined through the following steps d1 to d3.
[0055] d1. Local label extraction.
[0056] Based on the pre-grouping results of video frames (frame groups divided by visual depth scores), local time periods are determined according to the time intervals corresponding to the frame groups. The video frame extraction results (including scene, character, target, etc.) within each local time period are processed. First, duplicate information is removed by a text deduplication algorithm (such as duplicate content filtering based on cosine similarity). Then, the K-means clustering algorithm is used to classify the deduplicated information to generate local label features P(X) with timestamps. The label content specifically includes the theme, characters, scene, target, text content, character expressions, and behavioral content within the local time period.
[0057] d2. Global tag information aggregation.
[0058] Local label features P(X) are collected for all local time periods, and duplicate labels across local time periods (such as the same person appearing in multiple local time periods) are deduplicated a second time. The deduplicated local label information is then globally aggregated, and through statistical analysis (such as the frequency of character appearance and scene proportion analysis) and semantic clustering, the global label information G(X) of the entire video is extracted, specifically including core information such as the total number of characters, the main scene type, and the emotional change line of the characters.
[0059] d3. Tag information supplementation and mapping.
[0060] If the video to be processed is a resource already included in a public video library (confirmed by MD5 index matching), publicly available information about the video (such as the cast list and plot synopsis of a film or television drama, or the course outline of an open course) is retrieved from the internet. This retrieved information is then mapped and associated with the global label G(X) to supplement the detailed attributes of the label. If the video to be processed is private footage shot by a user, the face annotation library in the user's cloud drive album is accessed. Faces in the video frames are compared with faces in the annotation library using feature comparison (such as face similarity matching based on FaceNet) to determine the identities of the individuals and supplement the relationships between them (such as "father and son" or "colleagues"), enriching the information dimensions of the global label.
[0061] S102. With the goal of minimizing the matching cost between different modal feature data, align the feature dimensions between text feature data, audio feature data, and visual feature data to obtain multimodal feature data with unified feature dimensions.
[0062] In this context, matching cost refers to a metric used in the multimodal feature dimension alignment process to quantify the degree of semantic difference between different modal feature units (such as text feature unit tokens, audio feature unit tokens, and visual feature unit tokens). It measures the degree of "mismatch" between feature units in the semantic expression of two cross-modal features. A lower feature unit matching cost indicates a stronger semantic correlation and higher information consistency between the two cross-modal features. Conversely, a higher cost indicates a greater semantic difference and weaker information correlation.
[0063] For example, with minimizing inter-modal matching cost as the core objective, text feature F is selected. T As a dimensional benchmark (the benchmark modality can be flexibly adjusted because the text semantics better fit the needs of content extraction tasks), text F is achieved through three steps: "cost calculation - association matching - dimensional unification". T Audio F A Visual F VFeature dimension alignment. Specifically, focusing on the semantic similarity between modal features, cosine distance is used to quantify the matching cost of different modal feature units (such as text tokens, audio tokens, and visual tokens). Higher cosine similarity results in lower matching costs, and vice versa. Based on the Optimal Transport (OT) algorithm, a model is constructed to align each modality with the baseline modality (F...). T The matching relationship of audio F is determined by reducing computational complexity through a relaxation optimization strategy, and the inter-modal association weights that minimize the total matching cost are obtained, clarifying the semantic correspondence between feature units of different modalities. Based on the solved association weights, the audio F is analyzed. A Visual F V The feature dimensions are mapped and adjusted to be consistent with the baseline mode F. T Consistent dimensions ultimately yield multimodal feature data with unified dimensions and high semantic matching.
[0064] S103. Perform feature fusion on multimodal feature data with unified feature dimensions to obtain multimodal fused features.
[0065] For example, self-attention computation is performed separately for text, audio, and visual features after feature dimension unification. Weights are assigned to feature units within each modality to strengthen core semantic information (such as technical terms in text, key speech segments in audio, and core image patches in visuals) and weaken redundant information. Then, the enhanced single-modal features are converted into a unified temporal dimension. Cross-modal self-attention is used to capture the semantic associations of different modalities in the temporal dimension (such as text descriptions, audio narrations, and visual images within the same time period). Combined with an adapter layer to optimize features, temporal modal information complementarity is achieved. Finally, the temporally fused features are converted into a unified spatial dimension. Cross-modal self-attention is used again to mine the semantic matching of different modalities in the spatial dimension (such as action descriptions in text, human gestures in visuals, and ambient sounds in audio). Finally, the features are integrated to obtain multimodal fused features that combine the integrity and relevance of multimodal information.
[0066] S104. Using the structured descriptive text of the multimodal fusion features and the global labels of the data to be processed as input, the content extraction results are obtained by using the multimodal large model adapted to the structured descriptive text.
[0067] The multimodal large model adapted to structured descriptive text is optimized from a lightweight multimodal base model. Through fine-tuning, it accurately understands the structured information of global tags and outputs content extraction results that meet task requirements by combining multimodal fusion features. For example, the Deep seek Janus series models (such as Deep seekJanus-7B) can be used. This model natively supports multimodal inputs such as text and images, has a moderate number of parameters, and balances inference speed and semantic understanding capabilities, making it suitable for the efficient processing needs of long and short videos in cloud storage scenarios. For the structured descriptive text of global tags (including fixed fields of "basic attributes - content attributes - related attributes"), a dedicated prompt word template is designed to decompose the structured text into key information fields that the model can recognize (such as "[total number of characters: {}][main scene: {}][core theme: {}]"). Meanwhile, the pre-training parameters of the base model are fixed, and only the model's adapter layer is updated. Multiple rounds of fine-tuning are performed using labeled datasets (including video multimodal fusion features, global label structured text, and corresponding content extraction results) to enable the model to accurately associate structured fields with multimodal features.
[0068] In addition, similar multimodal large models with similar adaptation logic can also be selected, such as QwenVL-7B and InternVL-Base. They can all achieve adaptation to the structured description text of global tags through "basic model selection - structured text adaptation fine-tuning - Adapter layer optimization", which can meet the content extraction needs of different types of videos in cloud disk scenarios.
[0069] The content extraction method provided in the embodiments of this application will be described in detail below.
[0070] In the cloud-based video multimodal content extraction scenario applicable to the embodiments of this application, public video resources such as movies, TV series, and open courses uploaded by users to cloud drives often generate a large amount of redundant data due to repeated storage by multiple users. This not only occupies valuable cloud storage resources but also leads to the repeated execution of multimodal feature extraction and model inference processes for the same video, resulting in wasted computing resources and low processing efficiency. This application addresses the above problems from two dimensions: storage reuse and computational optimization, by constructing a public video library.
[0071] For example, a public video library focuses on frequently stored public video resources such as movies, TV series, and open courses. Through multi-channel data collection and structured storage, it forms a reusable video resource pool. The public video library may include the following core components: The video data layer collects video source files in mainstream formats such as MP4, AVI, MPEG, MOV, ASF, and FLASH through authorized platform access and compliant network resource collection, covering multiple categories such as film and television, education, and popular science, ensuring the integrity and compliance of the resources.
[0072] The basic information table creates a unique structured data record for each video segment. Fields include video type (such as movie, course, documentary), unique title, content description (information such as plot and knowledge points retrieved from authoritative sources such as Wikipedia), video duration, frame rate, and other basic attributes, supporting subsequent rapid matching and information retrieval.
[0073] An MD5 index table is used to calculate a unique MD5 digest value for each video segment (based on the MD5 Message DigestAlgorithm), establishing a mapping relationship between "video name - MD5 digest value". Leveraging the stability, speed, and excellent deduplication properties of the MD5 algorithm, unique video identification is achieved.
[0074] Through the collaborative design of the aforementioned video data layer, basic information table, and MD5 index table, a complete closed loop of "resource storage - information association - uniqueness verification" is formed, fundamentally solving the problems of storage redundancy and computational inefficiency of duplicate videos.
[0075] For example, after a user uploads a video to the cloud drive, the system performs deduplication and data reuse through the following steps e1~e2. Figure 2 As shown, an example diagram of a video frame deduplication process is used to illustrate the steps of deduplication judgment and data reuse.
[0076] e1. Extracting information from the video to be uploaded.
[0077] For example, for a user-uploaded video X, its basic attributes such as duration and frame rate are first obtained through a multimedia parsing tool, and then a unique identifier X for the video is calculated using the MD5 digest algorithm. MD5 .
[0078] e2. Multi-dimensional matching verification.
[0079] For example, the duration, frame rate, and X of video X are... MD5 The value is compared one by one with the video data table and MD5 index table of the public video library. If the data of the three are completely consistent, the video is determined to be a duplicate resource already existing in the library. The pre-stored video classification, summary, chapter summary, and other extraction results in the public video library are then used, without re-performing multimodal data processing and model inference. If any data is inconsistent, the video is determined to be new and proceeds to the subsequent multimodal data extraction and content extraction process.
[0080] In the cloud-based video multimodal content extraction scenario applicable to the embodiments of this application, video data often exhibits characteristics such as a large duration span (from a few minutes of home recording to several hours of course / conference recording), high inter-frame semantic redundancy (e.g., repeated frames in static scenes account for over 60%), and a surge in data volume (over 200GB of frame data per hour for 4K videos). Directly performing multimodal feature extraction and model inference on the original video frames would not only overload cloud computing resources (GPU / CPU) and cause processing time far exceeding user tolerance thresholds (e.g., 15-20 minutes for processing 1 hour of video), but might also lead to feature truncation due to multimodal model input length limitations (e.g., a visual token limit of 1024), resulting in the omission of crucial information from the latter half of long videos. This application provides a video frame compression method combining semantic-aware pre-grouping and contextual latent space compression, which reduces video frame data volume and computational overhead while preserving the core visual semantics of the video.
[0081] like Figure 3 As shown, the video frame compression method provided in this application includes the following steps S201 to S203. Wherein: S201. Based on the visual depth score, the video frames of the data to be processed are grouped to obtain video frame groups.
[0082] The visual depth score represents the degree of visual semantic difference between adjacent video frames.
[0083] For example, visual semantic abrupt change points (such as scene switching or changes in core elements of the screen) in a video can be located by the degree of visual semantic difference, and then the video can be segmented at the visual semantic abrupt change points to obtain multiple video frame groups (frame groups).
[0084] For example, taking a 1-minute family gathering video as an example, the first 30 seconds show an "indoor dining table scene (family members raising their glasses)," where the visual semantics of adjacent frames are highly similar (e.g., no significant changes in the positions of people or the table setting), and the calculated visual depth scores are all below a preset threshold. At the 31st second, the scene switches to an "outdoor courtyard scene (family taking photos)," and the visual semantic difference between the 30-second and 31-second frames increases sharply, with the corresponding visual depth score exceeding the threshold, thus being identified as a "visual semantic mutation point." Finally, using this mutation point as the boundary, the video frames are divided into two groups: Group 1 (the first 30 seconds, corresponding to the indoor scene) and Group 2 (the last 30 seconds, corresponding to the outdoor scene), achieving the grouping effect of "semantically coherent frame aggregation and semantically different frame segmentation."
[0085] S202. Based on the correlation features between multiple frame feature sequences contained in each video frame group and the visual features of each frame feature sequence, feature compression is performed to obtain visual feature data.
[0086] For example, for each group of video frames with coherent visual semantics, the temporal correlation of the frame feature sequence is first mined (such as the coherence of human actions in consecutive frames), and then the visual features of a single frame are combined for compression, which reduces the amount of data while preserving the semantic integrity of the frame group.
[0087] For example, the "indoor dining table scene frame group" (containing 30 frames, corresponding to the first 30 seconds) obtained from S201 is first divided into 5 frame feature sequences (each 6 frames is a subsequence). The Transformer encoder learns the temporal correlation features between frames within the subsequence (such as the action continuity of "raising a glass - clinking glasses - putting down the glass"). Then, at a ratio of "2 frames compressed into 1 feature", the visual features (extracted by the CLIP model) of each frame feature sequence are fused with the correlation features to generate 3 compressed features (6 frames - 3 compressed features). Finally, all the compressed features of this frame group are integrated to obtain visual feature data with a length of only 1 / 10 of the original frame group, which reduces the amount of data while completely preserving the core semantics of "raising a glass indoors".
[0088] S203. Integrate visual feature data, text feature data, audio feature data, and global label data to obtain multimodal feature data of the data to be processed.
[0089] For example, the visual feature data obtained after compressing video frames is associated and integrated with pre-processed text (such as subtitles, speech-transcribed text), audio (such as audio semantic features, emotion tags), and global tags (such as the main scene of the video "family gathering" and the total number of people "5") to form a multimodal feature set covering "visual-text-audio-semantic tags".
[0090] The video frame compression method provided in this application embodiment revolves around the core logic of "semantic grouping - contextual latent space compression," and achieves lightweight data volume while preserving the core visual semantics of the video through two major steps: S201 (video frame pre-grouping) and S202 (video frame feature compression). The video frame compression method provided in this application embodiment is further described below.
[0091] For example, the above S201 can also achieve semantic grouping of video frames from "feature extraction - similarity calculation - mutation point localization" through the following sub-steps S2011~S2013, ensuring that the grouping results are strongly correlated with the visual semantic coherence of the video. Wherein: S2011. Calculate the visual semantic feature similarity between a preset number of consecutive video frames in the data to be processed, and obtain multiple sets of similarity scores.
[0092] The similarity scores are arranged in chronological order of the video frames.
[0093] For example, taking a "3-minute family travel video (including two scenes: 'packing luggage indoors' and 'taking photos at an outdoor scenic spot')" as an example, firstly, a preset number of consecutive video frames are extracted according to the "duration adaptation principle". For example, the extraction is set to 90 frames (1 frame every 2 seconds), denoted as the frame set X={x1,x2,...,x... 90 Each frame corresponds to a unique temporal index (1 to 90).
[0094] frame x i (i is any value from 1 to 90) Input the CLIP-ViT-L pre-trained model (freeze model parameters, fine-tune weights using the ImageNet-21K dataset), and output the frame x through the model's visual encoder. i The 768-dimensional visual semantic feature vector f i The same processing is applied to the frame set to form a temporal visual feature sequence F. V ={f1,f2,...f 90 Then, using the cosine similarity formula, the visual semantic feature similarity S between each pair of adjacent frames (frame i and frame i+1) is calculated. i (Total 89), of which S i The value range is [0,1], and the closer the value is to 1, the more similar the visual semantics of adjacent frames. All S... i The video frames are sorted sequentially to form a similarity score sequence S={s1,s2,...s} 89 S2012. For any similarity score among multiple similarity scores, calculate the visual depth score by taking the maximum value among all temporally arranged similarity scores before any similarity score, the maximum value among all temporally arranged similarity scores after any similarity score, and any similarity score.
[0095] For example, the temporal similarity score sequence S={s1,s2,...,s...} generated in step S2011. n-1} (n is the total number of video frames extracted, s i The visual semantic similarity between the i-th frame and the (i+1)-th frame. Any similarity score s i Visual depth score d i Used to characterize "the degree of semantic abrupt change of this adjacent frame pair throughout the entire temporal sequence". When s i When the similarity score is significantly lower than the maximum value of all similarity scores before and after it, it indicates that the semantics between frames at that location have changed drastically. i This results in a larger value. Conversely, if s i The small difference between the maximum values before and after indicates that the semantics between frames are stable. i The value is relatively small.
[0096] Based on the above logic, the visual depth score di The calculation formula is defined as follows: ; in, Indicates taking s i The maximum value among all previous similarity scores. If s i If the first element of the sequence (i=1) has no prior similarity score, then... ; Indicates taking s i The maximum value among all subsequent similarity scores. If s i The last element of the sequence (i=n-1) If there is no post-similarity score, then ; 2×s i This indicates that s is strengthened through coefficient 2. i The difference between the maximum value before and after ensures that di at the semantic mutation location (point) can form a significant peak, which facilitates subsequent threshold judgment.
[0097] S2013. By using the preset semantic mutation judgment threshold and visual depth score, determine the visual semantic mutation points of the video frames included in the data to be processed, and divide the data to be processed into video frames according to the visual semantic mutation points to obtain video frame groups.
[0098] The semantic mutation judgment threshold δ is used to distinguish between peak values (mutation points) and normal values (non-mutation points) of the visual depth score. For example, the value of the semantic mutation judgment threshold δ can be determined based on the characteristics of the video scene and the distribution pattern of the depth score, so as to avoid missing mutation points due to an excessively high threshold or redundant grouping due to an excessively low threshold.
[0099] For example, such as Figure 4 As shown, for 6 video frames (denoted as x1-x6), the frames are quickly divided into 2 semantically coherent frame groups through steps S2011~S2013. The specific process is as follows: First, CLIP semantic features are extracted from each video frame. The cosine similarity of adjacent frame pairs (x1-x2), (x3-x4), and (x5-x6) is calculated, resulting in a similarity sequence S={s1,s2,s3}. Then, the depth score D={d1,d2,d3} corresponding to each similarity score is calculated using the formula. Finally, the visual depth score D is compared with a preset threshold δ (e.g., 0.6) (where d2 exceeds the threshold δ), ultimately dividing the 6 frames into two groups: the first group is x1-x4, and the second group is x5-x6.
[0100] For example, after obtaining the video frame group, for any semantically coherent video frame group, frame group feature compression is performed through sub-steps S2021 to S2024: S2021. Learn the temporal dependencies and visual semantic coherence between frame feature sequences through a visual model to construct a contextual latent space.
[0101] For example, the visual feature sequence F of a frame group of length L. V ={f1,f2,...,f L For example, let's take the visual feature sequence F as an example. V Input a lightweight visual model (such as a finely tuned version of InternVL-Base). The model learns inter-frame temporal dependencies (e.g., the progressive order of content from formula derivation to example calculation in blackboard writing) and visual semantic coherence (e.g., the logical association between blackboard text and teacher gestures between frames) through a self-attention mechanism. Ultimately, it constructs a contextual latent space covering the full semantic association of the frame group. This space accurately captures the dynamic relationships between elements within a frame, providing semantic support for subsequent feature fusion.
[0102] S2022. Based on the contextual latent space, the visual features of each frame feature sequence are fused with its corresponding associated features to obtain visual fusion features.
[0103] For example, in the aforementioned latent space, for the visual feature sequence F V Extract the original visual features (i.e., F) of each frame. V Each feature vector contains visual information such as "written text on the blackboard" and "person's actions." The model extracts the learned associated features (such as the logical association vectors of "formula derivation steps" between adjacent frames, and the correspondence vectors of "teacher's gestures and key points on the blackboard" within a frame group). Finally, through "feature concatenation + Layer Norm normalization," the original visual features of each frame are fused with their corresponding associated features to obtain a visual fusion feature sequence containing "visual details + semantic associations." This visual fusion feature sequence includes multiple visual fusion features.
[0104] S2023. Based on visual fusion features, each frame feature sequence is mapped to a fixed-dimensional compressed feature.
[0105] For example, based on the visual fusion feature sequence, compression is performed at a preset compression ratio (e.g., every 5 frames are compressed into 1 feature, compression ratio a=5). First, if the visual fusion feature sequence includes 10 frames of compressed features, it is divided into 2 frame feature sequences. Then, average pooling is performed on each frame feature sequence to extract the global semantic vector of the frame feature sequence, capturing the core visual semantics within the frame feature sequence. Finally, through the model's built-in learnable projection matrix, the global semantic vector of each frame feature sequence is mapped to a compressed feature of a fixed dimension (e.g., 768 dimensions), ultimately resulting in 2 compressed features (fa1, fa2), achieving a lightweight transformation from 10 frames to 2 frames.
[0106] S2024. Integrate the compressed features of fixed dimensions to obtain visual feature data.
[0107] For example, the temporal position of compressed features is determined by "alternating embedding of compressed features and visual markers", then effective compressed feature activation values are selected and retained, and finally, visual feature data with a unified dimension is formed by temporal arrangement and normalization.
[0108] like Figure 5 As shown, this application integrates fixed-dimensional compressed features through steps f1 to f4, transforming the multimodal features of video frame groups into structured visual feature data: f1. By alternating between fixed-dimensional compressed features and visual markers, fixed-dimensional compressed features are embedded into the visual marker sequence of video frame groups to form a hybrid coding sequence.
[0109] Visual markers are used to provide temporal location anchors for each fixed-dimensional compressed feature.
[0110] For example, Figure 5 This includes video frame group 1 (including x1-x4) and video frame group 2 (including x5 and x6). Video frame group 1 has two fixed-dimensional compression features (denoted as fa). 11 fa 12 Each compression feature corresponds to 2 original video frames. Video frame group 2 has 1 fixed-dimensional compression feature (denoted as fa2), corresponding to 2 original video frames.
[0111] Following the rule of "fixed-dimensional compressed features + alternating visual labels", fa 11 fa 12 The visual marker sequence embedded in video frame group 1 forms a hybrid coded sequence: {visual marker z1, visual marker z2, fa} 11 Visual marker z3, visual marker z4, fa 12 Similarly, the hybrid coding sequence for video frame group 2 is: {visual marker z5, visual marker z6, fa2}. Here, the visual markers provide temporal location anchors for each compressed feature, for example, fa... 11 By using "visual marker 1 - visual marker 2", the temporal intervals corresponding to x1-x2 are clearly defined, ensuring that the temporal correlation between compressed features and the original frame is not lost.
[0112] f2. Encode the hybrid coding sequence, discard the activation values corresponding to ordinary visual markers in the hybrid coding sequence according to the preset visual compression strategy, and retain the activation values of all fixed-dimensional compressed features.
[0113] Self-attention encoding is applied to the above round-encoded sequences to enhance the semantic relevance of compressed features (such as fa in video frame group 1). 11 with fa12 (Scene association). According to the preset visual compression strategy, the activation values corresponding to ordinary visual tags in the hybrid coding sequence are discarded, and only the activation values of all fixed-dimensional compressed features are retained.
[0114] f3. The activation values of all retained fixed-dimensional compressed features are arranged and integrated according to the temporal order of the video frame groups to form a feature matrix.
[0115] The activation values of the retained fixed-dimensional compressed features are arranged in temporal order according to the video frame group. For example, the activation values of fa in video frame group 1... 11 fa 12 Press "fa first" 11 After fa 12 The sequence of video frames is arranged sequentially; the fa2 of video frame group 2 is arranged separately. This results in a feature matrix: video frame group 1 corresponds to a submatrix of "2 rows × fixed dimensions", and video frame group 2 corresponds to a submatrix of "1 row × fixed dimensions". Figure 5 The arrangement of the "compressed features" on the right is consistent.
[0116] f4. Perform dimension normalization on the feature matrix to obtain visual feature data.
[0117] Layer Norm normalization is performed on the feature matrix. For example, the mean and variance of each row of features are calculated, and all activation values are mapped to a uniform numerical range. Finally, visual feature data with a uniform format is obtained, whose dimensions and numerical range are adapted to the input requirements of the subsequent multimodal feature alignment module, thus completing the integration process.
[0118] pass Figure 5 As can be seen, steps f1 to f4 realize the entire process from "hybrid sequence construction" to "normalized output", ensuring that fixed-dimensional compressed features are transformed into structured and standardized visual feature data while preserving temporal and semantic information, providing core visual input for the extraction of multimodal video content in the cloud.
[0119] In the cloud-based video multimodal content extraction scenario applicable to the embodiments of this application, multimodal data (text, audio, and visual) generally suffer from inconsistent feature dimensions (e.g., 512 dimensions for text, 1024 dimensions for audio, and 768 dimensions for visual) and loose semantic associations due to differences in acquisition methods (text from subtitle / speech transcription, audio from microphone acquisition, and visual from video frame sampling) and feature extraction models (BERT for text, Wav2Vec2.0 for audio, and CLIP for visual). This application uses text features as a benchmark and specifically aligns them with visual and audio features. Through the logic of "semantic filtering - weight allocation - dimension adjustment," with the core objective of "minimizing intermodal matching costs," it eliminates dimensional differences and semantic biases, ultimately obtaining multimodal feature data with unified dimensions and strong semantic associations.
[0120] In some embodiments, such as Figure 6 As shown, with the goal of minimizing the matching cost between different modal feature data, the feature dimensions of text feature data, audio feature data, and visual feature data are aligned to obtain multimodal feature data with unified feature dimensions, including the following steps S301~S303: S301. Determine the mapping weight of the visual feature unit based on the cosine distance between each text feature unit and the visual feature units within the feature matching range.
[0121] In this context, the sum of the mapping weights of multiple visual feature units within the feature matching range of the text feature unit is not a fixed value.
[0122] For example, suppose text features F have been extracted. T ∈L T ×d(L T =5, d=512, indicating that the text features include 5 text subjects (feature dimension is 512) and visual features F. V a ∈L V ×d(L V =10, d=768, indicating that the visual features include 10 keyframes, and the feature dimension is 768. First, based on the temporal correlation of multimodal data, a corresponding visual feature matching range is defined for each text feature unit, ensuring that each text feature unit is only associated with visual feature units within the same time period, avoiding invalid matching across time periods, and ultimately forming a matching relationship of "1 text feature unit corresponds to 1-3 visual feature units". Then, for each text feature unit, its feature vector is used to calculate the cosine distance with the vector of each visual feature unit within the matching range (a temporary dimension adaptation is performed before the calculation to ensure that the two can be used for similarity calculation), as shown in the following formula: ; in, It is the vector of the i-th visual feature unit (such as the 768-dimensional feature vector of the compressed keyframe).
[0123] This is the vector of the j-th text feature unit (e.g., the 512-dimensional feature vector of a text sentence).
[0124] pass The dot product measures the degree of overlap between vectors in a direction (the larger the dot product, the closer the directions are).
[0125] Let L2 be the L2 norm (vector length) of both.
[0126] Let be the cosine distance between the i-th visual feature unit and the j-th text feature unit. The smaller the cosine distance, the higher the semantic similarity. The larger the cosine distance, the greater the semantic difference.
[0127] Based on the Relaxed Optimal Transmission (OT) framework, the "complement of cosine distance (1-Cost)" is used as the basis for semantic similarity to perform weight normalization on visual feature units within the matching range. The higher the similarity of the visual feature unit, the greater the mapping weight. The sum of the weights of multiple visual feature units is not fixed (it changes dynamically with the number of matching units and semantic similarity; for example, when one text feature unit matches two visual feature units, the sum of weights may be 0.9 or 1.0, and when matching three, it may be 0.95). Finally, the mapping weight corresponding to each visual feature unit is obtained.
[0128] S302. Based on the mapping weight of each visual feature unit within the feature matching range of the text feature unit and the cosine distance with the text feature unit, determine the set of visual feature units whose mapping weight is greater than a preset threshold.
[0129] For example, a preset mapping weight threshold (e.g., 0.25-0.35) can be set based on the semantic density of the modal data. Visual feature units within the matching range of each text feature unit are filtered according to the preset mapping weight threshold. Visual feature units with mapping weights greater than the threshold are retained (these units have a strong semantic association with the text feature units and can be used as core objects for subsequent dimensional adjustments); visual feature units with mapping weights less than or equal to the threshold are removed (these units have a loose semantic association with the text feature units, and including them in the calculation will increase matching costs and affect alignment accuracy); if the weights of all visual feature units within the matching range of a certain text feature unit are lower than the threshold, the visual feature unit with the highest weight is retained to avoid situations where no visual feature unit can be matched. Through this filtering, each text feature unit corresponds to a "set of highly correlated visual feature units".
[0130] S303. Adjust the feature dimension of the visual feature data with the goal of minimizing the total matching cost between the visual feature unit set and the text feature unit, so that the feature dimension of the adjusted visual feature data is consistent with that of the text feature data.
[0131] For example, based on the mapping weights obtained in S301 and the set of visual feature units determined in S302, a vision-to-text transfer matrix M is constructed. V2T The matrix elements represent the matching strength between visual feature units and text feature units, and there is no total inflow constraint (supporting matching of one text feature unit with multiple visual feature units). Its transfer matrix M V2T As shown in the following formula: ; in, For the visual-to-text transfer matrix, matrix elements (i, j) represents the matching weight between the i-th visual feature unit and the j-th text feature unit; This refers to the length of the visual feature (i.e., the total number of visual feature units). For example, L... V =10); This can be represented as finding the cosine distance between the i-th visual feature unit and the i-th visual feature unit. The smallest text feature unit j (i.e., the text feature unit with the highest semantic similarity); The above formula ensures that each visual feature unit matches only the one text feature unit that is semantically closest to it, with a matching weight of . The matching weight with other text feature units is 0. This significantly reduces computational complexity while preserving the core semantic correspondence.
[0132] Next, a learnable dimensional transformation matrix (768 input dimensions, 512 output dimensions) is designed to map the 768-dimensional vector of the visual feature unit to a 512-dimensional vector. During the mapping process, the transformation matrix parameters are iteratively optimized using the "minimum total matching cost" as the loss function to ensure that the mapped visual features and text features are highly aligned at the semantic level.
[0133] Following the above logic, a dimensionality transformation is performed on all visual feature units, ultimately resulting in a visual feature F with a dimension of 512. V′ a ∈L′ V ×512(L′ V To ensure that the total number of visual feature units after filtering is completely consistent with the feature dimensions of text features, the visual-text modal feature dimensions are aligned.
[0134] In some embodiments, such as Figure 7 As shown, with the goal of minimizing the matching cost between different modal feature data, the feature dimensions of text feature data, audio feature data, and visual feature data are aligned to obtain multimodal feature data with unified feature dimensions, including the following steps S401~S403: S401. Determine the mapping weight of the audio feature unit based on the cosine distance between each text feature unit and the audio feature unit within the feature matching range.
[0135] In this context, the sum of the mapping weights of multiple audio feature units within the feature matching range of the text feature unit is not a fixed value.
[0136] For example, suppose text features F have been extracted. T ∈L T ×d(L T =5, d=512, indicating that the text features include 5 text bodies (feature dimension is 512) and audio features F. A ∈L A ×d(L A =8, d=1024, indicating that the audio features include 8 audio feature units with a feature dimension of 1024. First, based on the temporal correlation of multimodal data (e.g., the playback time of text sentences coincides with the recording time of audio speech segments), a corresponding audio feature matching range is defined for each text feature unit, ensuring that each text feature unit is only associated with audio feature units within the same time period, ultimately forming a matching relationship of "1 text feature unit corresponds to 1-2 audio feature units". Then, for each text feature unit, its feature vector is used to calculate the cosine distance with the vector of each audio feature unit within the matching range (a temporary dimension adaptation is performed before the calculation to allow similarity calculation between the two), as shown in the following formula: ; in, Let be the cosine distance between the i-th audio feature unit and the j-th text feature unit. The smaller the cosine distance, the higher the semantic similarity. The larger the cosine distance, the greater the semantic difference.
[0137] It is the vector of the i-th audio feature unit (such as the 1024-dimensional feature vector of audio).
[0138] This is the vector of the j-th text feature unit (e.g., the 512-dimensional feature vector of a text sentence).
[0139] pass The dot product measures the degree of overlap between vectors in a direction (the larger the dot product, the closer the directions are).
[0140] Let L2 be the L2 norm (vector length) of both.
[0141] Based on the Relaxed Optimal Transmission (OT) framework, the "complement of cosine distance (1-Cost)" is used as the basis for semantic similarity to normalize the weights of audio feature units within the matching range. Audio feature units with higher similarity have larger mapping weights, and the sum of the weights of multiple audio feature units is not fixed (it changes dynamically with the number of matching units and semantic similarity; for example, when one text feature unit matches two audio feature units, the sum of weights may be 0.9 or 1.0, and when matching three, it may be 0.95). Finally, the mapping weight corresponding to each audio feature unit is obtained.
[0142] S402. Based on the mapping weight of each audio feature unit within the feature matching range of the text feature unit, and the cosine distance with the text feature unit, determine the set of audio feature units whose mapping weight is greater than a preset threshold.
[0143] For example, a preset mapping weight threshold (e.g., 0.3) can be set based on the semantic density of the modal data. Audio feature units within the matching range of each text feature unit are filtered according to this preset mapping weight threshold. Audio feature units with mapping weights greater than the threshold are retained (these units have a strong semantic correlation with the text feature units and can be used as core objects for subsequent dimensional adjustments); audio feature units with mapping weights less than or equal to the threshold are removed (these units have a loose semantic correlation with the text feature units, and including them in the calculation will increase matching costs and affect alignment accuracy); if the weights of all audio feature units within the matching range of a certain text feature unit are lower than the threshold, the audio feature unit with the highest weight is retained to avoid situations where no audio feature unit can be matched. Through this filtering, each text feature unit corresponds to a "set of highly correlated audio feature units".
[0144] S403. Adjust the feature dimension of the audio feature data with the goal of minimizing the total matching cost between the audio feature unit set and the text feature unit, so that the feature dimension of the adjusted audio feature data is consistent with that of the text feature data.
[0145] For example, based on the mapping weights obtained in S401 and the set of audio feature units determined in S402, an audio-to-text transfer matrix M is constructed. A2T The matrix elements represent the matching strength between audio feature units and text feature units, and there is no total inflow constraint (supporting matching of one text feature unit with multiple audio feature units). Its transfer matrix M A2T As shown in the following formula: ; in, For audio-to-text transfer matrix, matrix elements (i, j) represents the matching weight between the i-th audio feature unit and the j-th text feature unit; This refers to the length of the audio feature (i.e., the total number of audio feature units). For example, L... A =8); This can be expressed as finding the cosine distance between the i-th audio feature unit and the i-th feature unit. The smallest text feature unit j (i.e. the text feature unit with the highest semantic similarity).
[0146] The above formula ensures that each audio feature unit is matched only with the one text feature unit that is semantically closest to it, with a matching weight of . The matching weight with other text feature units is 0. This significantly reduces computational complexity while preserving the core semantic correspondence.
[0147] Next, a learnable dimensionality transformation matrix (input dimension 1024, output dimension 512) is designed to map the 1024-dimensional vector of the audio feature unit to a 512-dimensional vector. During the mapping process, the transformation matrix parameters are iteratively optimized using the "minimum total matching cost" as the loss function to ensure that the mapped audio features and text features are highly aligned at the semantic level.
[0148] Following the above logic, a dimensionality transformation is performed on all audio feature units, ultimately resulting in an audio feature F with a dimension of 512. A′ ∈L′ A ×512(L′ A To ensure that the total number of audio feature units after filtering is completely consistent with the feature dimensions of text features, audio-text modal feature dimension alignment is achieved.
[0149] This application embodiment, after obtaining multimodal feature data with unified feature dimensions, achieves deep fusion through a process of "single-modal enhancement - temporal cross-modal fusion - spatial cross-modal fusion - feature integration". For example... Figure 8 The diagram illustrates the structure of a multimodal feature fusion model, which includes a unimodal self-attention module, a temporal cross-modal fusion module, and a spatial cross-modal fusion module. This model integrates the aligned text features F... T′ Visual features F V′ a and audio features F A′ As input, the data sequentially passes through a single-modal self-attention module, a temporal cross-modal module, and a spatial cross-modal module to achieve temporal and spatial feature fusion of multimodal data. Below, we will combine... Figure 9 The method example shown illustrates an implementation method for achieving deep fusion of multimodal features using a multimodal feature fusion model.
[0150] In some embodiments, such as Figure 9 As shown, the implementation method for achieving deep fusion of multimodal features using a multimodal feature fusion model includes the following steps S501~S504: S501. The feature importance weights of text feature data, audio feature data and visual feature data with unified feature dimensions are calculated separately through a single-modal self-attention mechanism, and the text enhancement features, audio enhancement features and visual enhancement features are obtained by weighting according to the feature importance weights.
[0151] like Figure 8 As shown, the text features F T′ Visual features F V′ a and audio features F A′ The input is fed into a unimodal self-attention module, which performs "normalization-self-attention (Att)-addition & normalization" operations. First, the text feature F is processed... T′ Visual features F V′ a and audio features F A′ Layer Norm normalization was performed on each feature to eliminate differences in numerical scale between features. Then, the text features F were processed separately. T′ Visual features F V′ a and audio features F A′ Perform unimodal self-attention computation. For text Att, focus on semantic associations within the text (e.g., weight allocation between "core argument" and "evidence"), and output the importance weights of text features. For visual Att, focus on inter-frame associations within the visual context (e.g., weight allocation between "keyframes" and "transition frames"), and output the importance weights of visual features. For audio Att, focus on speech segment associations within the audio context (e.g., weight allocation between "effective speech" and "silence"), and output the importance weights of audio features. Then, align the text features F... T′ Visual features F V′ a and audio features F A′ The text enhancement feature F is obtained by adding the features weighted by the self-attention weights and then normalizing them. T′′ Visual enhancement feature F V′′ a Audio enhancement feature F A′′ This enhances key information within a single modality.
[0152] S502. After dimensional transformation of the text enhancement features, audio enhancement features, and visual enhancement features, they are merged, and multimodal fusion features with temporal dimensions are obtained through temporal cross-modal self-attention mechanism and normalization processing.
[0153] like Figure 8 As shown, the text enhancement feature F T′′ Visual enhancement feature F V′′ a Audio enhancement feature F A′′ The input is fed into the temporal cross-modal fusion module, which performs cross-modal attention (TC-Att)-Adapter-addition & normalization to achieve multimodal fusion in the temporal dimension.
[0154] In some embodiments, S502 above includes the following steps S5021 to S5023: S5021. Perform dimensional transformation on the original dimensions of the text enhancement features, audio enhancement features, and visual enhancement features to obtain the dimensional transformed text temporal features, audio temporal features, and video temporal features.
[0155] For example, the temporal correlation between text and other modalities is captured through the text TC-Att (such as the correlation weight between the "00:05 second caption" and the audio and visual data at the same time). The text feature dimensions are adapted to the dimensions required for temporal fusion through the text Adapter (such as transforming from 512 dimensions to 256 dimensions).
[0156] Similarly, the visual and audio enhancement features are processed by "Visual TC-Att + Visual Adapter" and "Audio TC-Att + Audio Adapter" to obtain text temporal features, visual temporal features, and audio temporal features.
[0157] S5022. Merge the text temporal features, audio temporal features, and video temporal features according to the channel dimension to form temporal fusion input data.
[0158] For example, text, visual, and audio temporal features are concatenated along the channel dimension to form temporal fusion input data, providing a unified input format for cross-modal attention computation.
[0159] S5023. Perform feature optimization on the temporal fusion input data, and normalize the feature optimization results to obtain multimodal fusion features in the temporal dimension.
[0160] For example, the temporal fusion input data is fed into the temporal cross-modal self-attention module to calculate the association weights of different modal features in the temporal dimension. For instance, the temporal synchronization degree between the text temporal feature "instruction description at 00:15" and the visual temporal feature "operation screen at 00:15", and the audio temporal feature "voice instruction at 00:15", is calculated; features with higher synchronization degrees are assigned higher attention weights. Then, based on the learned association weights, the multimodal features in the temporal fusion input data are weighted and summed, giving features with strong temporal consistency higher representation strength (e.g., the feature values of "plot twist text at 00:30", "camera transition at 00:30", and "sudden change in background music at 00:30" are significantly improved after weighting), thus completing the temporal association fusion of multimodal temporal features.
[0161] Next, the weighted fused features are residually concatenated (added) with the original features of the temporal fusion input data. This preserves the original feature information while enhancing the gain from cross-modal temporal correlation, preventing the loss of basic temporal information during feature optimization. Finally, Layer Norm normalization is performed on the added features. By calculating the mean and variance of the feature matrix, all feature values are mapped to a uniform numerical range (e.g., [-1,1]), eliminating the distribution shift caused by time scale differences in different modal features and ensuring stable feature distribution.
[0162] S503. After performing inverse dimensional transformation on the temporal dimension multimodal fusion features, merge them, and obtain the spatial dimension multimodal fusion features through spatial cross-modal self-attention mechanism and normalization processing.
[0163] like Figure 8 As shown, the temporal multimodal fusion features are input into the spatial cross-modal fusion module. The spatial cross-modal fusion module performs "cross-modal attention (SC-Att)-Adapter-addition & normalization" to achieve spatial multimodal fusion.
[0164] In some embodiments, S503 above includes the following steps S5031 to S5033: S5031. Perform inverse dimensional transformation on the multimodal fusion features of the temporal dimension to obtain the text space features, audio space features and video space features after inverse dimensional transformation.
[0165] For example, text SC-Att captures the semantic associations between text and other modalities (such as the association weights between "'red' description" and the visual and auditory representations of red objects). A text adapter is used to inversely transform the temporally fused text feature dimensions back to the original spatial dimensions (e.g., restoring from 256 dimensions to 512 dimensions). Similarly, visual and audio temporal features are processed by "Visual SC-Att + Visual Adapter" and "Audio SC-Att + Audio Adapter" to obtain text spatial features, visual spatial features, and audio spatial features.
[0166] S5032. Merge the text spatial features, audio spatial features, and video spatial features according to the channel dimension to form spatial fusion input data.
[0167] For example, text, visual, and audio spatial features are concatenated along the channel dimension to form spatial fusion input data, providing a unified input format for spatial cross-modal attention computation.
[0168] S5033. Perform feature optimization on the spatial fusion input data, and normalize the feature optimization results to obtain multimodal fusion features in the spatial dimension.
[0169] For example, the spatial fusion input data is fed into the spatial cross-modal self-attention module to calculate the association weights of different modal features in the semantic space. For instance, the semantic matching degree between the text spatial feature "'red object' description" and the visual spatial feature "red object pixel block," and the audio spatial feature "'red' voice broadcast," is calculated. Feature pairs with higher matching degrees are assigned higher attention weights. Then, based on the learned association weights, the multimodal features in the spatial fusion input data are weighted and summed, so that features with strong semantic consistency receive higher representation strength (e.g., the feature values of "'happy' text emotion" and "smiley face image features" and "laughter audio features" are significantly improved after weighting), thus completing the semantic association fusion of multimodal spatial features.
[0170] Next, the weighted fused features are residually concatenated (added) with the original features of the spatially fused input data. This preserves the original feature information while enhancing the gains from cross-modal semantic associations, preventing the loss of basic semantics during feature optimization. Finally, Layer Norm normalization is performed on the added features. By calculating the mean and variance of the feature matrix, all feature values are mapped to a uniform numerical range (e.g., [-1, 1]), eliminating distribution shifts caused by differences in semantic expression between different modal features and ensuring stable feature distribution.
[0171] S504. Integrate the multimodal fusion features of the spatial dimension to obtain multimodal fusion features.
[0172] For example, global pooling and fully connected layer mapping are performed on the spatial dimension multimodal fusion features to integrate the temporal and spatial multimodal correlation information, resulting in the final multimodal fusion feature. This feature simultaneously includes the temporal synchronization and semantic consistency of text, vision, and audio, and can be directly used for cloud-based video multimodal content extraction tasks (such as key information extraction, sentiment analysis, etc.).
[0173] In conclusion, Figure 8 The three modules achieve a complete process of multimodal feature fusion from "single-modal independence" to "deep multimodal fusion" through hierarchical collaboration of "single-modal enhancement - temporal cross-modal association - spatial cross-modal association", providing high-quality fusion feature input for multimodal tasks.
[0174] In this embodiment, after acquiring multimodal fusion features, the multimodal fusion features are input into a pre-trained multimodal large model, accompanied by contextual prompts to guide the output. For example, in the "short video content extraction" scenario, while inputting the fusion features into the model, the prompt "Please extract the core theme, key visual objects, and audio keywords of the video based on the input features, and output them in the format 'Theme: XXX; Visual Object: XXX; Audio Keywords: XXX'" will be provided. The model will then rely on its pre-trained cross-modal understanding capabilities to quickly parse the multi-dimensional information in the fusion features and generate the corresponding content extraction results.
[0175] This application also provides a method for fine-tuning a multimodal large model to adapt to a specific scenario. The core of this method is to enable the model to accurately output structured descriptive text that meets the needs of the scenario by adjusting lightweight parameters.
[0176] In some embodiments, structured descriptive text of global labels from historical data and corresponding multimodal fusion features are used as training samples. By fixing the pre-trained parameters of the multimodal large model, only the parameters of the adapter layer are updated to perform multiple rounds of iterative optimization on the multimodal large model. Under the condition of satisfying the preset adjustment termination condition, a multimodal large model adapted to the structured descriptive text is obtained.
[0177] The preset adjustment termination conditions include, but are not limited to, the following conditions: Condition 1. In three consecutive iterations, the field matching accuracy between the model output and the labeled results is ≥95%; Condition 2. The number of iterations reaches the preset limit (e.g., 100 rounds); Condition 3. The loss function value is lower than the threshold (e.g., cross-entropy loss ≤ 0.03).
[0178] For example, by adjusting the parameters in a lightweight manner (updating only the Adapter layer), the multimodal large model can accurately output structured descriptive text that meets the needs of the scene. The overall adjustment process is as follows: using the "multimodal fusion features + structured descriptive text" of historical data as training samples, the model is guided to learn by combining prompt word templates, the pre-training parameters of the large model are fixed, and only the Adapter layer is iteratively optimized until the termination condition is met, and finally the scene-adapted model is obtained.
[0179] In some embodiments, such as Figure 10 As shown, fine-tuning the multimodal large model includes the following steps S601~S603: S601. Design prompt word templates that adapt to the input format of multimodal large models, based on the field characteristics of structured descriptive text.
[0180] For example, based on the fixed fields (such as topic, subject, attributes, etc.) contained in the structured description text, and in accordance with the natural language interaction logic that can be parsed by the multimodal large model, a prompt word template containing field names, filling requirements and output format is constructed. This template can not only clearly guide the model to focus on the information dimensions to be extracted, but also match the input specifications of the model, ensuring that the model can accurately understand the task objectives based on the template.
[0181] S602. Combine the structured descriptive text and historical multimodal fusion features from the training samples according to the prompt word template format, and input them into the multimodal large model containing the Adapter layer.
[0182] For example, the specific content of the structured descriptive text in the training samples is filled into the corresponding field positions of the prompt word template to form prompt words containing task guidance information. Then, the prompt words are associated and combined with the corresponding historical multimodal fusion features to form a data structure that meets the model input requirements. Subsequently, the combined overall data is input into the multimodal large model that integrates the Adapter layer to provide standardized input samples for model learning.
[0183] S603. Using the difference between the model output and the labeled results of historical data as the loss function, and under the premise of fixing the pre-training parameters of the multimodal large model, perform multiple rounds of iterative optimization by updating the parameters of the Adapter layer until the preset adjustment termination condition is met.
[0184] For example, the computational model calculates the deviation between the output generated from the input sample and the corresponding labeled result in historical data (such as the quantified value of missing fields or inconsistent content), and uses this deviation as the loss function. During training, the pre-training parameters of the multimodal large model are kept unchanged, and the parameters of the Adapter layer are adjusted only through the backpropagation algorithm to reduce the loss. After multiple iterations, when the loss function value reaches a preset threshold, or the number of iterations meets a set upper limit, or the model output accuracy is stable within the target range for multiple consecutive iterations, optimization stops, and model adaptation is completed.
[0185] After obtaining the content extraction results (including video classification tags, video summary, and video segmentation summary), this application can derive diversified intelligent applications adapted to cloud storage scenarios based on these core extraction results. The application scenarios are as follows: Scene 1 video overview, such as Figure 11 As shown, based on the video subtitles, text summaries, and segmented summaries extracted from the content, core information is displayed to users on the cloud drive video list page or preview page: keyword subtitle fragments, core summaries of up to 300 words, and key node summaries segmented semantically (such as "00:05-00:30 Project Background Introduction") are presented through structured layout. Users can quickly grasp the video theme, core highlights, and content structure without opening the entire video, efficiently judging the video's value to decide whether to watch it in its entirety. This improves the utilization efficiency of video assets and optimizes the user's browsing experience.
[0186] Scenario 2, video search, leverages extracted video category tags (such as "Education-Postgraduate Entrance Exam-Mathematics" and "Life-Travel-Island") and video summaries to construct a multi-dimensional search index based on "filename-content keywords-video type." Users can initiate searches by entering filename keywords, core video content descriptions (such as "explaining the definition of the limit of calculus"), and video types (such as "meeting recording"). The system uses the multi-dimensional index to achieve multi-path retrieval, quickly matching video resources that meet the user's needs, significantly improving the accuracy and efficiency of the search. Figure 12 As shown, users can retrieve a list of videos under the "Courses" category (including teaching videos 1 to 4) by entering "explaining the definition of the limit of calculus" in the search bar.
[0187] Scenario 3, intelligent recommendation, combines extracted video category tags, user viewing history, and behavioral preferences (such as viewing duration and collection / sharing actions) to build a personalized recommendation model. The model analyzes the matching degree between user preferences and video tags (e.g., if a user frequently watches "food tutorial" videos, content tagged with "food - home cooking" will be prioritized), pushing relevant or similar video resources to the user. This helps users discover more potentially interesting content, enriching the content consumption scenarios of the cloud drive and increasing user stickiness.
[0188] Scenario 4, Video Management, leverages the structured nature of video category tags to provide automated video organization capabilities for cloud storage. After a user uploads a video, the system automatically categorizes it into corresponding preset folders or adds unique tags based on extracted category tags (such as "Home Videos - 2024 Spring Festival" or "Work Materials - Product Launch"), eliminating the need for manual organization by the user. Figure 13 As shown, the system automatically organizes video recordings 1 to 6 into the "Home Videos" tag based on the extracted category tags.
[0189] Meanwhile, it supports batch operations based on tags, such as batch renaming (in the format of "tag + timestamp"), batch adding supplementary tags, and batch moving folders, which greatly reduces the operating cost for users to manage massive video assets and improves management efficiency.
[0190] Scenario 5, Video Creation, integrates content extraction results with basic video editing tools (editing, merging, adding subtitles, adding background music, etc.) to meet users' video creation needs. Users can quickly locate key segments to be edited based on video segmentation summaries, directly reuse extracted subtitles as video subtitle materials, and combine category tags to recommend suitable background music (such as upbeat music recommended under the "vlog-travel" tag). Through a closed-loop process of "extraction results - editing tools - sharing output," it provides users with a one-stop service from preprocessing creative materials to sharing finished products, lowering the creation threshold and enriching the functional scenarios of the cloud drive.
[0191] The following describes a content extraction system applicable to the content extraction method provided in this application, in conjunction with the above embodiments. For example... Figure 14 As shown, this content extraction system includes a public video library, a multimodal data processing module, a visual compression module, a video content extraction model, and a cloud-based intelligent video application. This system revolves around video data, using multi-module collaboration to process raw video into structured content. The specific data flow process is as follows: The public video library, as the data supplier, stores video resources to be analyzed, providing raw video input for the entire system.
[0192] The multimodal data processing module is used for multimodal decomposition and encoding of video. Specifically, it first performs audio preprocessing on the video's audio stream (such as noise reduction and audio format unification), and then uses a speech encoder to convert it into audio features F. A For video-related text (such as subtitles, video descriptions, etc.), text preprocessing (including word segmentation, noise reduction, etc.) is performed first, and then the text encoder encodes it into text features F. T The video frames undergo preprocessing (such as standardizing frame size and sampling keyframes), and then a visual encoder generates visual features F. VAt the same time, semantic tags and external related information can be added to visual features by means of "tag extraction", "key information extraction" and "web retrieval".
[0193] The visual compression module compresses visual features F. V By performing pre-grouping and visual compression operations, lightweight visual features F are obtained. V a Meanwhile, for audio features F A Text features F T OT local alignment is performed separately to achieve semantic-level association matching between audio, text, and visual features, resulting in aligned visual features F. V′ a Audio features F A′ Text features F T .
[0194] The video content extraction model consists of a multimodal feature fusion layer and a multimodal large model (VLM). The multimodal feature fusion layer processes visual features F... V′ a Audio features F A′ Text features F T The process involves three steps: "single-modal self-attention", "temporal cross-modal fusion", and "spatial cross-modal fusion" to obtain multimodal fusion features.
[0195] By using a multimodal large model (VLM) and combining it with prompt words to analyze multimodal fusion features, structured content extraction results such as "summary, classification, and chapter summary" are output.
[0196] The cloud drive video intelligent application applies the extracted structured results to scenarios such as video preview, video search, intelligent recommendation, video management, and video creation, realizing the value of multimodal content extraction in the cloud drive scenario.
[0197] Furthermore, this embodiment provides a content extraction device, such as... Figure 15 As shown, the device includes: a determining module 1510, an alignment module 1520, a fusion module 1530, and an output module 1540. Wherein: The determination module 1510 is configured to determine the multimodal feature data of the data to be processed; the multimodal feature data includes text feature data, audio feature data, visual feature data and global label data.
[0198] Alignment module 1520 is configured to align the feature dimensions of text feature data, audio feature data, and visual feature data with the goal of minimizing the matching cost between different modal feature data, so as to obtain multimodal feature data with unified feature dimensions.
[0199] The fusion module 1530 is configured to perform feature fusion on multimodal feature data with unified feature dimensions to obtain multimodal fused features.
[0200] Output module 1540 is configured to take structured descriptive text containing multimodal fusion features and global label data as input, and use a multimodal large model adapted to the structured descriptive text to obtain content extraction results.
[0201] In some embodiments, the determining module 1510 is further configured to group video frames of the data to be processed based on visual depth scores to obtain video frame groups; wherein the visual depth score characterizes the degree of visual semantic difference between adjacent video frames.
[0202] Visual feature data is obtained by compressing the features based on the correlation features between multiple frame feature sequences contained in each video frame group and the visual features of each frame feature sequence.
[0203] By integrating visual feature data, text feature data, audio feature data, and global label data, multimodal feature data of the data to be processed is obtained.
[0204] In some embodiments, the determining module 1510 is further configured to calculate the visual semantic feature similarity between a preset number of consecutive video frames in the data to be processed, and obtain multiple sets of similarity scores; the similarity scores are arranged in the temporal order of the video frames.
[0205] For any similarity score among multiple similarity scores, the visual depth score is calculated based on the maximum value among all temporally arranged similarity scores before any similarity score, the maximum value among all temporally arranged similarity scores after any similarity score, and the similarity score itself.
[0206] By using a preset semantic mutation judgment threshold and visual depth score, the visual semantic mutation points of the video frames included in the data to be processed are determined, and the video frames of the data to be processed are divided according to the visual semantic mutation points to obtain video frame groups.
[0207] In some embodiments, the determining module 1510 is further configured to construct a contextual latent space by learning the temporal dependencies and visual semantic coherence between frame feature sequences through a visual model.
[0208] Based on the contextual latent space, the visual features of each frame feature sequence are fused with its corresponding associated features to obtain visual fusion features.
[0209] Based on visual fusion features, each frame feature sequence is mapped to a fixed-dimensional compressed feature.
[0210] The compressed features of fixed dimensions are integrated to obtain visual feature data.
[0211] In some embodiments, the determining module 1510 is further configured to embed fixed-dimensional compressed features into a visual marker sequence of a video frame group in an alternating manner of fixed-dimensional compressed features and visual markers to form a hybrid coding sequence; wherein the visual markers are used to provide a temporal position anchor point for each fixed-dimensional compressed feature.
[0212] The hybrid coding sequence is encoded, and the activation values corresponding to ordinary visual tags in the hybrid coding sequence are discarded according to the preset visual compression strategy, while the activation values of all fixed-dimensional compressed features are retained.
[0213] The activation values of all retained fixed-dimensional compressed features are arranged and integrated according to the temporal order of video frame groups to form a feature matrix.
[0214] The feature matrix is normalized to obtain visual feature data.
[0215] In some embodiments, the alignment module 1520 is further configured to determine the mapping weight of a visual feature unit based on the cosine distance between each text feature unit and a visual feature unit within the feature matching range; wherein the sum of the mapping weights of multiple visual feature units within the feature matching range of the text feature unit is not a fixed value.
[0216] Based on the mapping weight of each visual feature unit within the feature matching range of the text feature unit, and the cosine distance with the text feature unit, a set of visual feature units with mapping weights greater than a preset threshold is determined.
[0217] The feature dimensions of the visual feature data are adjusted with the goal of minimizing the total matching cost between the visual feature unit set and the text feature unit set, so that the feature dimensions of the adjusted visual feature data and the text feature data are unified.
[0218] In some embodiments, the alignment module 1520 is further configured to determine the mapping weight of the audio feature unit based on the cosine distance between each text feature unit and the audio feature unit within the feature matching range; wherein the sum of the mapping weights of multiple audio feature units within the feature matching range of the text feature unit is not a fixed value.
[0219] Based on the mapping weight of each audio feature unit within the feature matching range of the text feature unit, and the cosine distance with the text feature unit, a set of audio feature units with mapping weights greater than a preset threshold is determined.
[0220] The feature dimensions of the audio feature data are adjusted with the goal of minimizing the total matching cost between the audio feature unit set and the text feature unit set, so that the feature dimensions of the adjusted audio feature data and the text feature data are unified.
[0221] In some embodiments, the fusion module 1530 is further configured to calculate the feature importance weights of text feature data, audio feature data and visual features with unified feature dimensions through a single-modal self-attention mechanism, and to obtain text enhancement features, audio enhancement features and visual enhancement features by weighting according to the feature importance weights.
[0222] After dimensional transformation of text enhancement features, audio enhancement features, and visual enhancement features, they are merged, and then multimodal fusion features with temporal dimensions are obtained through temporal cross-modal self-attention mechanism and normalization processing.
[0223] After performing an inverse dimensional transformation on the temporal multimodal fusion features, the features are merged, and then the spatial multimodal fusion features are obtained through a spatial cross-modal self-attention mechanism and normalization processing.
[0224] Multimodal fusion features are obtained by integrating the spatial dimension of multimodal fusion features.
[0225] In some embodiments, the fusion module 1530 is further configured to perform dimensional transformation on the original dimensions of the text enhancement features, audio enhancement features and visual enhancement features respectively, to obtain the dimensionally transformed text temporal features, audio temporal features and video temporal features.
[0226] Text time-series features, audio time-series features, and video time-series features are merged along the channel dimension to form time-series fusion input data.
[0227] Feature optimization is performed on the temporal fusion input data, and the feature optimization results are normalized to obtain multimodal fusion features in the temporal dimension.
[0228] In some embodiments, the fusion module 1530 is further configured to perform a dimensionality inverse transformation on the temporal dimension multimodal fusion features to obtain the inversely transformed text space features, audio space features, and video space features.
[0229] Text spatial features, audio spatial features, and video spatial features are merged along the channel dimension to form spatial fusion input data.
[0230] Feature optimization is performed on the spatial fusion input data, and the feature optimization results are normalized to obtain multimodal fusion features in the spatial dimension.
[0231] In some embodiments, the output module 1540 is further configured to use the structured description text of the global labels of historical data and the corresponding multimodal fusion features as training samples, and to perform multiple rounds of iterative optimization on the multimodal large model by fixing the pre-training parameters of the multimodal large model and only updating the adapter parameters.
[0232] Under the condition of satisfying the preset adjustment termination condition, a multimodal large model adapted to the structured descriptive text is obtained.
[0233] In some embodiments, the output module 1540 is further configured to design prompt word templates that adapt to the multimodal large model input format based on the field features of the structured descriptive text.
[0234] The structured descriptive text from the training samples and the historical multimodal fusion features are combined according to the prompt word template format and then input into the multimodal large model containing the Adapter layer.
[0235] Using the difference between the model output and the labeled results of historical data as the loss function, and with the pre-training parameters of the multimodal large model fixed, the parameters of the Adapter layer are updated for multiple rounds of iterative optimization until the preset adjustment termination condition is met.
[0236] It should be noted that other corresponding descriptions of the functional units involved in the content extraction device provided in this embodiment can be found in the description of the content extraction method in the above embodiments, and will not be repeated here.
[0237] Based on the content extraction method shown in the above embodiments, this embodiment also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method shown in the above embodiments.
[0238] Based on the methods shown in the above embodiments, this embodiment also provides a computer program product on which a computer program is stored, and when the computer program product is executed by a processor, it implements the methods shown in the above embodiments.
[0239] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0240] Based on the method shown in the above embodiments, and Figure 15 To achieve the above objectives, the present application also provides an electronic device, such as a terminal device, in the virtual device embodiment shown. The electronic device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the method shown in the above embodiment.
[0241] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0242] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0243] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0244] Through the above description of the implementation methods, those skilled in the art can clearly understand that this application can be implemented using software plus necessary general-purpose hardware platforms, or it can be implemented through hardware. Addressing the problem that existing video content extraction technologies rely on a single modality (such as subtitles or transcribed text), resulting in incomplete and fragmented information, this method first determines the text, audio, and visual multimodal feature data of the video to be processed, breaking the limitations of single-modal processing and fully covering the semantic, auditory, and visual multi-dimensional information of the video data. Then, aiming to minimize the matching cost between different modal features, it aligns the dimensions of each modal feature, solving the problem of ineffective feature association caused by differences in modal dimensions, laying the foundation for multimodal information fusion. Subsequently, it fuses multimodal features with unified dimensions to further explore the complementary value of text, audio, and visual data, avoiding extraction bias caused by the lack of single-modal information. Finally, the structured descriptive text of multimodal fusion features and global label data is used as input, and the processing results are obtained by using a multimodal large model. This ensures that the extracted results have both the comprehensiveness of multimodal data and conformity to the global semantics of the video, thereby significantly improving the completeness and accuracy of the content extraction results and effectively adapting to the needs of efficient management and retrieval of massive and diverse video assets in scenarios such as cloud storage.
[0245] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0246] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A content extraction method, characterized in that, include: Determine the multimodal feature data of the data to be processed; the multimodal feature data includes text feature data, audio feature data, and global label data; With the goal of minimizing the matching cost between different modal feature data, the feature dimensions of the text feature data, the audio feature data, and the visual feature data are aligned to obtain multimodal feature data with unified feature dimensions; Feature fusion is performed on the multimodal feature data with unified feature dimensions to obtain multimodal fused features; Using the multimodal fusion features and the structured descriptive text of the global label data as input, the content extraction result is obtained by using a multimodal large model adapted to the structured descriptive text.
2. The method according to claim 1, characterized in that, The multimodal feature data for determining the data to be processed includes: The video frames of the data to be processed are grouped based on the visual depth score to obtain video frame groups; wherein, the visual depth score represents the degree of visual semantic difference between adjacent video frames. Based on the correlation features between multiple frame feature sequences contained in each video frame group, and the visual features of each frame feature sequence, feature compression is performed to obtain the visual feature data; The visual feature data, text feature data, audio feature data, and global label data are integrated to obtain the multimodal feature data of the data to be processed.
3. The method according to claim 2, characterized in that, The video frames of the data to be processed are grouped based on visual depth scores to obtain video frame groups, including: Calculate the visual semantic feature similarity between a preset number of consecutive video frames in the data to be processed to obtain multiple sets of similarity scores; the similarity scores are arranged in the temporal order of the video frames; For any similarity score among multiple similarity scores, the visual depth score is calculated based on the maximum value among all temporally arranged similarity scores before the given similarity score, the maximum value among all temporally arranged similarity scores after the given similarity score, and the given similarity score. By using a preset semantic mutation judgment threshold and the visual depth score, the visual semantic mutation points of the video frames included in the data to be processed are determined, and the data to be processed is divided into video frames according to the visual semantic mutation points to obtain video frame groups.
4. The method according to claim 2, characterized in that, The visual feature data is obtained by performing feature compression based on the correlation features between the multiple frame feature sequences contained in each video frame group and the visual features of each frame feature sequence, including: By learning the temporal dependencies and visual semantic coherence between the frame feature sequences through a visual model, a contextual latent space is constructed. Based on the contextual latent space, the visual features of each frame feature sequence are fused with its corresponding associated features to obtain visual fusion features; Based on the visual fusion features, each frame feature sequence is mapped to a fixed-dimensional compressed feature; The visual feature data is obtained by integrating the compressed features of the fixed dimension.
5. The method according to claim 4, characterized in that, The process of integrating the compressed features of the fixed dimension to obtain the visual feature data includes: The fixed-dimensional compressed features are embedded into the visual tag sequence of the video frame group in an alternating manner with visual tags to form a hybrid coding sequence; wherein, the visual tags are used to provide temporal position anchors for each fixed-dimensional compressed feature; The hybrid coding sequence is encoded, and the activation values corresponding to the ordinary visual markers in the hybrid coding sequence are discarded according to a preset visual compression strategy, while the activation values of all the fixed-dimensional compressed features are retained. The activation values of all the retained fixed-dimensional compressed features are arranged and integrated according to the temporal order of the video frame group to form a feature matrix; The feature matrix is subjected to dimensionality normalization to obtain the visual feature data.
6. The method according to any one of claims 1-5, characterized in that, The text feature data includes multiple text feature units, the visual feature data includes multiple visual feature units, and the audio feature data includes multiple audio feature units. The goal is to minimize the matching cost between different modal feature data, aligning the feature dimensions of the text feature data, the audio feature data, and the visual feature data to obtain multimodal feature data with unified feature dimensions, including: The mapping weight of the visual feature unit is determined based on the cosine distance between each text feature unit and the visual feature units within the feature matching range; wherein the sum of the mapping weights of multiple visual feature units within the feature matching range of the text feature unit is not a fixed value. Based on the mapping weight of each visual feature unit within the feature matching range of the text feature unit, and the cosine distance with the text feature unit, a set of visual feature units whose mapping weight is greater than a preset threshold is determined. The feature dimension of the visual feature data is adjusted with the goal of minimizing the total matching cost between the visual feature unit set and the text feature unit, so that the feature dimension of the adjusted visual feature data is consistent with that of the text feature data.
7. The method according to claim 6, characterized in that, The method further includes: The mapping weight of the audio feature unit is determined based on the cosine distance between each text feature unit and the audio feature units within the feature matching range; wherein the sum of the mapping weights of multiple audio feature units within the feature matching range of the text feature unit is not a fixed value. Based on the mapping weight of each audio feature unit within the feature matching range of the text feature unit, and the cosine distance with the text feature unit, a set of audio feature units whose mapping weight is greater than a preset threshold is determined. The feature dimension of the audio feature data is adjusted with the goal of minimizing the total matching cost between the audio feature unit set and the text feature unit, so that the feature dimension of the adjusted audio feature data is consistent with that of the text feature data.
8. The method according to any one of claims 1-5, characterized in that, The process of fusing features from multimodal feature data with unified feature dimensions to obtain multimodal fused features includes: The feature importance weights of the text feature data, audio feature data and visual feature data with the same feature dimension are calculated by a single-modal self-attention mechanism, and then weighted according to the feature importance weights to obtain text enhancement features, audio enhancement features and visual enhancement features. The text enhancement features, audio enhancement features, and visual enhancement features are merged after dimensional transformation, and then multimodal fusion features with temporal dimensions are obtained through a temporal cross-modal self-attention mechanism and normalization processing. The temporal-dimensional multimodal fusion features are merged after undergoing inverse dimensional transformation, and then multimodal fusion features in the spatial dimension are obtained through spatial cross-modal self-attention mechanism and normalization processing. The multimodal fusion features are obtained by integrating the multimodal fusion features of the spatial dimension.
9. The method according to claim 8, characterized in that, The text enhancement features, audio enhancement features, and visual enhancement features are dimensionally transformed and then merged. Through a temporal cross-modal self-attention mechanism and normalization processing, a temporal-dimensional multimodal fusion feature is obtained, including: The original dimensions of the text enhancement features, audio enhancement features, and visual enhancement features are transformed to obtain the transformed text temporal features, audio temporal features, and video temporal features. The text temporal features, audio temporal features, and video temporal features are merged along the channel dimension to form temporal fusion input data; The temporal fusion input data is subjected to feature optimization, and the feature optimization results are normalized to obtain the multimodal fusion features of the temporal dimension.
10. The method according to claim 9, characterized in that, The process involves performing an inverse dimensional transformation on the temporal-dimensional multimodal fusion features, merging them, and then obtaining spatial-dimensional multimodal fusion features through a spatial cross-modal self-attention mechanism and normalization processing. This includes: The temporal-dimensional multimodal fusion features are subjected to inverse dimensional transformation to obtain the inverse dimensional transformation text space features, audio space features, and video space features; The text spatial features, audio spatial features, and video spatial features are merged along the channel dimension to form spatial fusion input data; The spatial fusion input data is subjected to feature optimization, and the feature optimization results are normalized to obtain the multimodal fusion features of the spatial dimension.
11. The method according to any one of claims 1-5, characterized in that, The method further includes: Using structured descriptive text of global labels from historical data and corresponding multimodal fusion features as training samples, the multimodal large model is iteratively optimized in multiple rounds by fixing the pre-training parameters of the multimodal large model and only updating the Adapter parameters. Under the condition of satisfying the preset adjustment termination condition, a multimodal large model adapted to the structured description text is obtained.
12. The method according to claim 11, characterized in that, The method uses structured descriptive text with global labels from historical data and corresponding multimodal fusion features as training samples. By fixing the pre-training parameters of the multimodal large model and only updating the parameters of the adapter layer, multiple rounds of iterative optimization are performed on the multimodal large model, including: Based on the field features of structured descriptive text, a prompt word template adapted to the input format of the multimodal large model is designed; The structured descriptive text and historical multimodal fusion features in the training samples are combined according to the prompt word template format and input into the multimodal large model containing the Adapter layer; Using the difference between the model output and the labeled results of the historical data as the loss function, and with the pre-training parameters of the multimodal large model fixed, the parameters of the Adapter layer are updated for multiple rounds of iterative optimization until the preset adjustment termination condition is met.
13. A content extraction device, characterized in that, include: The determination module is configured to determine the multimodal feature data of the data to be processed; the multimodal feature data includes text feature data, audio feature data, visual feature data, and global label data; The alignment module is configured to align the feature dimensions of the text feature data, the audio feature data, and the visual feature data with the goal of minimizing the matching cost between different modal feature data, so as to obtain multimodal feature data with unified feature dimensions; The fusion module is configured to perform feature fusion on the multimodal feature data with the same feature dimensions to obtain multimodal fused features; The output module is configured to take the structured descriptive text of the multimodal fusion features and the global label data as input, and use the multimodal large model adapted to the structured descriptive text to obtain the content extraction result.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 12.
15. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 12.