Multi-modal data processing method and device and storage medium
By converting multimodal data into data blocks and performing feature extraction and alignment, the problem of difficulty in integrating multimodal information in the knowledge base is solved, and the unified representation of multimodal features is realized, which improves the comprehensive understanding ability of the model.
Patent Information
- Application Number
- CN202510713503.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-26
AI Technical Summary
The existing technology is difficult to effectively integrate multimodal information in the knowledge base, especially the processing capacity of visual information such as images and videos is limited, which makes it difficult to directly fusion of heterogeneous data, affecting the comprehensive understanding effect of the model.
By obtaining multimodal data from the knowledge base, converting it into data blocks based on the correlation between modal contents, and feature extraction, alignment and combination are performed to form a unified multimodal feature representation.
It realizes effective integration of multimodal information, enhances the comprehensive understanding of multimodal data by artificial intelligence models, and solves the problem of fusion of heterogeneous data.
Smart Images

Figure CN120541274A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a multimodal data processing method, device, and storage medium. Background Art
[0002] With the development of artificial intelligence technology, the types of data stored in knowledge bases are becoming increasingly complex and diverse, including various heterogeneous data forms such as text documents, image files, video materials, audio recordings, etc. These different modal data have significant differences in feature representation methods, making it difficult for existing technologies to effectively integrate multimodal information in knowledge bases. Summary of the Invention
[0003] In view of this, the present disclosure provides a multimodal data processing method, device, and storage medium.
[0004] According to a first aspect of the present disclosure, a multimodal data processing method is provided, comprising:
[0005] Obtain multimodal data from the knowledge base;
[0006] Converting the multimodal data into at least two data blocks based on associations between different modal contents in the multimodal data;
[0007] Performing feature extraction and feature alignment on each of the data blocks to obtain at least one data feature group; the dimensions of the data features in the data feature group are the same;
[0008] Each of the data feature groups is combined to obtain multimodal features; the multimodal features represent the combing results of the knowledge base.
[0009] According to an embodiment of the present disclosure, converting the multimodal data into at least two data blocks based on the association relationship between different modal contents in the multimodal data includes:
[0010] Acquiring image content, voice content, and text content in the multimodal data;
[0011] Determining multiple frames of target images in the image content; wherein the difference between the target images in each frame is greater than a difference threshold;
[0012] Converting the speech content into text content;
[0013] For any frame of the target image, the target image and at least one text content in a corresponding time sequence form a data block.
[0014] According to an embodiment of the present disclosure, determining multiple frames of target images in the image content includes:
[0015] For any first image frame in the image content, performing a sliding window operation on the first image to obtain a second image;
[0016] Acquire a characteristic state of the second image in a color space; the characteristic state includes at least one of a hue state, a saturation state, and a brightness state;
[0017] Calculate the difference in feature states between any two adjacent frames of the second image to obtain the visual difference;
[0018] The latter second image frame of two adjacent second image frames whose visual difference is greater than the difference threshold is determined as the target image.
[0019] According to an embodiment of the present disclosure, the step of extracting and aligning features of each data block to obtain at least one data feature group includes:
[0020] Extracting features of different modal contents in each of the data blocks to obtain a first data feature set; the first data feature set includes text features and image features;
[0021] encoding the features in each of the first data feature sets based on associations between different modal contents in the data block to obtain a second data feature set;
[0022] Pruning is performed on the features in each of the second data feature sets to obtain a data feature group.
[0023] According to an embodiment of the present disclosure, encoding the features in each of the first data feature sets based on the association relationship between different modal contents in the data block to obtain the second data feature set includes:
[0024] For each image feature in the first data feature set, dividing the target image corresponding to the image feature into a plurality of image blocks, and encoding the image feature based on a positional relationship between the image blocks to obtain an encoded image feature;
[0025] For each text feature in the first data feature set, encoding the text feature based on the character order of the text content corresponding to the text feature to obtain an encoded text feature;
[0026] In the case where text content exists in the target image, encoding the text feature based on the position of the text content in the target image to obtain an encoded text feature;
[0027] The encoded image features and the encoded text features in each of the first data feature sets are used as a second data feature set.
[0028] According to an embodiment of the present disclosure, pruning the features in each of the second data feature sets to obtain a data feature group includes:
[0029] For any second data feature set, construct a correlation matrix between features in the second data feature set;
[0030] Combining the feature columns in the correlation matrix in sequence based on the order of correlation between the feature columns in the correlation matrix from strong to weak until the dimension of the feature columns in the correlation matrix is reduced to the first dimension;
[0031] Combining the characteristic rows in the correlation matrix in sequence based on the order of correlation between the characteristic rows in the correlation matrix from strong to weak until the dimension of the characteristic rows in the correlation matrix is reduced to the second dimension;
[0032] The second data feature set after dimensionality reduction is output as a data feature group.
[0033] According to an embodiment of the present disclosure, the data feature groups are combined to obtain multimodal features, including:
[0034] Splicing the data features in each of the data feature groups to obtain a target feature;
[0035] The target features are fused into the multimodal features.
[0036] According to an embodiment of the present disclosure, the method further includes:
[0037] The multimodal data is input into a fusion model to obtain multimodal features output by the fusion model.
[0038] A second aspect of the present disclosure provides a multimodal data processing device, comprising:
[0039] Acquisition module, used to obtain multimodal data;
[0040] a conversion module, configured to convert the multimodal data into at least two data blocks based on associations between different modal contents in the multimodal data;
[0041] an alignment module, configured to extract and align features of each of the data blocks to obtain at least one data feature group; the data features in the data feature group have the same dimension;
[0042] The combination module is used to combine each of the data feature groups to obtain multimodal features; the multimodal features represent the combing results of the knowledge base.
[0043] The third aspect of the present disclosure further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.
[0044] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0046] Figure 1 A flowchart of a multimodal data processing method provided by an embodiment of the present disclosure is schematically shown;
[0047] Figure 2 A schematic diagram of a multimodal feature fusion network provided by an embodiment of the present disclosure is shown;
[0048] Figure 3 The following schematically shows a structural block diagram of a multimodal data processing device according to an embodiment of the present disclosure;
[0049] Figure 4 A block diagram of an electronic device provided by an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0050] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0051] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0052] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0053] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0054] In the embodiments of this disclosure, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of all data involved (including, but not limited to, user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard the security of user personal information, network security, and national security.
[0055] The embodiments of the present disclosure provide a multimodal data processing method, device, and storage medium. Before introducing the technical solutions provided by the embodiments of the present disclosure, the related technologies involved in the present disclosure are first described.
[0056] In related technologies, multimodal data processing refers to the comprehensive processing technology of data in different forms including text, images, videos, audio, etc., which plays an increasingly important role in the fields of intelligent question answering, content understanding, knowledge retrieval, etc.
[0057] In the application of large-scale language models, Retrieval-augmented Generation (RAG) technology enhances the model's generative capabilities by retrieving relevant information from knowledge bases, effectively alleviating the model's "hallucination" and knowledge obsolescence issues. However, existing RAG is primarily designed for pure text data and has limited processing capabilities for visual information such as images and videos. This prevents the effective utilization of the vast amount of visual information in knowledge bases.
[0058] In multimodal data processing, data from different modalities have heterogeneous feature structures. For example, text data is typically represented as a sequence of word vectors, image data as a pixel matrix, and audio data as spectral features. This heterogeneity makes it difficult to directly and effectively fuse data from different modalities.
[0059] Attention mechanisms, a key technology in deep learning, enable models to focus on important parts of input data. In the field of multimodal processing, attention mechanisms can help models learn the relationships between different modalities and achieve cross-modal information alignment and fusion.
[0060] Despite progress in various areas, numerous challenges remain. The heterogeneity of data across different modalities remains a major obstacle to effective fusion. Related technologies lack effective cross-modal relationship modeling methods, hindering the full utilization of semantic connections between modalities. Important structural information is easily lost during feature alignment, impacting the ultimate understanding. Furthermore, related technologies lack end-to-end processing capabilities, often requiring multiple independent models to collaborate to complete tasks, increasing system complexity and processing costs.
[0061] The embodiments of the present disclosure provide a multimodal data processing method, device, and storage medium to solve the technical problems existing in the related art. The multimodal data processing method includes:
[0062] Obtain multimodal data from the knowledge base;
[0063] Based on the association relationship between different modal contents in the multimodal data, convert the multimodal data into at least two data blocks;
[0064] Performing feature extraction and feature alignment on each data block to obtain at least one data feature group; the dimensions of the data features in the data feature group are the same;
[0065] Each data feature group is combined separately to obtain multimodal features; the multimodal features represent the combing results of the knowledge base.
[0066] By implementing the embodiments of the present disclosure, multimodal data is converted into at least two data blocks based on the correlation between different modal contents in the multimodal data, effectively retaining and utilizing the intrinsic connection between cross-modal information; by performing feature extraction and feature alignment processing on each data block respectively, heterogeneous data that originally had differences in feature representation methods are converted into data feature groups with the same dimension, solving the problem in related technologies that different modal data are difficult to directly process due to feature heterogeneity; by combining each data feature group respectively to obtain multimodal features that represent the results of knowledge base sorting, effective integration of multimodal information in the knowledge base is achieved, forming a unified multimodal feature representation, and enhancing the comprehensive understanding ability of artificial intelligence-related learning models of multimodal data.
[0067] The multimodal data processing device in the embodiments of the present disclosure refers to a computing device with the ability to collect, process, and analyze multimodal data. This type of multimodal data processing device includes multimedia processing modules in mobile terminals such as smartphones, tablets, and laptops; data collection devices such as smart cameras, recording devices, and scanners; and new processing carriers such as the multi-sensory fusion processing units of augmented reality (AR) / virtual reality (VR) devices, the environmental perception systems of service robots, and the voice and image joint recognition modules of smart speakers. In addition, professional processing components such as the multi-sensor fusion systems of autonomous vehicles, the multimodal diagnostic and analysis devices of medical imaging equipment, the audio-visual joint recognition panels of industrial testing equipment, and vertical field processing units such as the multi-authentication systems in the financial sector, the smart classroom interactive analysis platforms in the education sector, and the multi-dimensional customer behavior analysis engines in retail scenarios also fall within the scope of the multimodal data processing device described in the embodiments of the present disclosure. At the same time, special scenario processing devices such as the smart home multimodal control center in the Internet of Things environment, the voice and vision fusion processor of the in-vehicle infotainment system, the multi-sensor data fusion system of avionics equipment, as well as distributed processing architectures such as the distributed multimodal computing interface of the server cluster and the multimedia data processing channel of the cloud-based intelligent analysis platform are all within the technical adaptation scope of the embodiments of the present disclosure.
[0068] It should be noted that the embodiments of the present disclosure do not limit the specific type of multimodal data processing device. The above examples are only illustrative descriptions. The technical solution of the present disclosure can be applied to any device that has multimodal data processing requirements and requires cross-modal information fusion and understanding.
[0069] The data processing device in the embodiments of the present disclosure refers to a terminal device with data parsing and output capabilities. This type of data processing device includes display modules for mobile terminals such as smartphones, tablets, and laptops; physical output devices such as printers, 3D projectors, and laser engravers; and new output carriers such as holographic presentation modules for augmented reality (AR) / virtual reality (VR) devices, interactive display screens for service robots, and tactile feedback units for smart bracelets. Furthermore, specialized output components such as ticket printing systems for self-service terminals, film imaging devices for medical imaging equipment, and visualization panels for industrial control equipment; vertical output units such as ATM receipt generation modules in the financial sector, electronic whiteboard writing systems for the education sector, and content rendering engines for digital signage in retail settings also fall within the scope of the data processing device described in the embodiments of the present disclosure. Furthermore, specialized output devices in IoT environments such as smart home control center display terminals, in-vehicle head-up displays (HUDs), and avionics flight instrument display systems, as well as distributed output architectures such as remote visualization interfaces for server clusters and streaming media output channels for cloud rendering terminals, are all within the technical adaptability of the embodiments of the present disclosure.
[0070] It should be noted that the embodiments of the present disclosure do not limit the specific type of data processing device. The above examples are only illustrative descriptions. The technical solution of the present disclosure can be applied to any device that needs to output data and may change the content presentation properties.
[0071] The multimodal data processing method provided by the embodiment of the present disclosure is described in detail below.
[0072] Figure 1 The figure schematically shows a flow chart of a multimodal data processing method provided by an embodiment of the present disclosure.
[0073] like Figure 1 As shown, the processing method of this embodiment may include operations S101 to S104.
[0074] Operation S101: obtaining multimodal data from a knowledge base;
[0075] Operation S102: converting the multimodal data into at least two data blocks based on association relationships between different modal contents in the multimodal data;
[0076] Operation S103: performing feature extraction and feature alignment on each data block to obtain at least one data feature group; the dimensions of the data features in the data feature group are the same;
[0077] In operation S104 , each data feature group is combined to obtain a multimodal feature; the multimodal feature represents the combing result of the knowledge base.
[0078] In operation S101, the knowledge base refers to a data set that stores structured or semi-structured information, which usually contains knowledge elements such as facts, concepts, rules and relationships in a specific field; illustratively, the knowledge base includes but is not limited to a comprehensive information storage system for different types of data such as text documents, image files, video materials, audio recordings, etc., which is used to provide data support for downstream tasks such as intelligent question answering, content retrieval, and knowledge reasoning.
[0079] Correspondingly, multimodal data refers to a data set that simultaneously contains two or more different representation forms. These representation forms can be different information carriers such as text, images, audio, video, etc., and can be understood in the embodiments of the present disclosure as a combination of heterogeneous data with intrinsic correlations extracted from a knowledge base; illustratively, multimodal data includes but is not limited to documents with picture descriptions, videos with subtitles, presentations with background music, and other complex information content, which is used to achieve more accurate semantic understanding and knowledge representation through cross-modal information fusion.
[0080] Specifically, the method of obtaining multimodal data from the knowledge base can be determined based on the storage structure and data type of the knowledge base;
[0081] In a feasible implementation, presentation files containing mixed text and image content can be identified by traversing the document index in the knowledge base, and the text paragraphs, embedded images, and page layout information therein can be extracted as associated multimodal data.
[0082] In another feasible implementation, relevant video materials can be retrieved from the knowledge base based on the user's query intention, and the video frame sequence and corresponding audio track information can be extracted through video parsing technology, and the visual content and audio content can be obtained as time-series associated multimodal data.
[0083] In another feasible implementation, technical manuals in PDF format can be exported in batches from the document management system of the knowledge base, and document parsing tools can be used to simultaneously extract different types of data elements such as plain text content, chart information, formula symbols, etc. to form a structured multimodal data set.
[0084] In another feasible implementation, audio lecture materials can be obtained through the API interface of the knowledge base, combined with the corresponding text records or subtitle files, and the voice signal features and text semantic information can be paired to form a multimodal data object with a time synchronization relationship.
[0085] In operation S102, the association relationship refers to the mutual correspondence or dependency relationship between data contents in different representation forms at the semantic, temporal, spatial or logical levels; in the embodiment of the present disclosure, the association relationship between different modal contents can be understood as the time synchronization relationship between image frames and audio content at corresponding moments, the semantic pointing relationship between charts and their explanatory texts in documents, the logical correspondence between video scene changes and voice content conversion, and other cross-modal intrinsic connections, which are used to determine the reasonable segmentation boundaries and organization methods of multimodal data to ensure that the semantic integrity and logical coherence of the original data are not destroyed during the conversion process.
[0086] Furthermore, a data block refers to a multimodal data organization unit formed based on a specific association relationship. Each data block contains different modal contents with a clear correspondence. In the embodiment of the present disclosure, it can be understood as a data set consisting of one or more associated image contents, text contents, audio contents, etc.
[0087] For example, a data block can be a key frame sequence corresponding to a complete scene in a video and the speech transcription text within its time period, or a topic page in a presentation containing background images, title text, and description text, etc., which are used as the basic processing units for subsequent feature extraction, alignment processing, and semantic understanding.
[0088] Specifically, the method of converting multimodal data into data blocks can be selected according to the characteristics of different modal contents and the types of association relationships.
[0089] In a feasible implementation, the key transition points in the video can be identified based on scene change detection technology of the video content, and the key frame images of each scene segment can be combined with the audio transcription text in the corresponding time window to form a graphic data block based on the scene.
[0090] In another feasible implementation, document structure analysis technology can be used to identify page boundaries and content layout in presentations or PDF documents, and different modal elements such as text paragraphs, charts, images, etc. in each page or chapter can be combined according to spatial positions and logical relationships to form multimodal data blocks based on pages.
[0091] In another feasible implementation, a semantic segmentation method can be used to divide long audio content into topics, and the speech transcription text of each topic paragraph can be time-aligned with the visual content (such as slides, presentation images, etc.) that appeared in the same time period to form audio and video data blocks based on themes.
[0092] In another feasible implementation, semantic recognition of text content can be performed based on natural language processing technology. When a reference relationship to an image, chart or other media content is detected in the text, the relevant text paragraphs are combined with the referenced visual content to form a cross-modal data block linked by a semantic reference relationship.
[0093] In operation S103, feature extraction refers to the process of extracting a numerical representation that can characterize the semantic information and structural features of the data content from the original multimodal data through calculation and transformation methods. In the embodiment of the present disclosure, it can be understood as using a deep learning network, a traditional machine learning algorithm or a special feature engineering technology to convert the image content in the data block into a visual feature vector, the text content into a text feature vector, and the audio content into an audio feature vector, so as to uniformly convert the original data of different modalities into a processable numerical feature representation.
[0094] Furthermore, due to the differences in dimension and distribution of features of different modalities, it is necessary to achieve a unified representation of features through feature alignment. Feature alignment refers to the process of adjusting multimodal features with different sources and large structural differences into a unified feature representation with the same dimensional specifications and similar distribution characteristics. In the embodiment of the present disclosure, it can be understood as unifying heterogeneous image features, text features, and audio features into vector representations of the same dimension through technical means such as position encoding, dimensionality transformation, and feature normalization, while maintaining the semantic correspondence and spatiotemporal correlation between the features of each modality, which is used to solve the heterogeneity problems of different modal features in terms of dimension, scale, distribution, etc., and ensure that the information of each modality can be effectively integrated in the subsequent feature combination process.
[0095] Based on the processing results of the above-mentioned feature extraction and feature alignment, a multimodal feature representation with a unified format is formed, namely a data feature group. A data feature group refers to a feature collection unit formed by organizing the aligned features of different modalities in the same data block according to their inherent correlation after feature extraction and feature alignment processing. In the embodiment of the present disclosure, it can be understood as a structured feature organization including aligned image feature vectors, text feature vectors, audio feature vectors, etc., wherein each feature vector maintains consistency in dimensional specifications and is related to each other in semantic content, and is used as the basic input unit of multimodal feature fusion to ensure that the complementary information between the modalities can be fully utilized during the fusion process, thereby improving the semantic accuracy of the final multimodal representation.
[0096] Among them, the same dimension of data features in the data feature group means that within the same data feature group, the feature vectors from all different modal sources have consistent specification requirements in the mathematical dimension. In the embodiment of the present disclosure, it can be understood as ensuring that image features, text features, audio features, etc. are adjusted to the same dimensional size before entering the fusion module through feature alignment processing, which is used to eliminate the differences in mathematical structure of different modal features, ensure the feasibility and effectiveness of subsequent feature fusion, splicing, calculation and other operations, and avoid calculation errors or information loss due to dimensional mismatch.
[0097] In operation S104, multimodal features refer to comprehensive feature vectors or feature matrices formed by fusing, integrating or jointly representing features of various modalities in different data feature groups. In the disclosed embodiment, it can be understood as a unified multimodal representation formed by combining image features, text features, audio features, etc. that have undergone feature extraction and alignment processing through splicing, weighted fusion, attention mechanism, etc. Multimodal features can simultaneously reflect the semantic information of each modality and the cross-modal association relationship, and are used to provide rich semantic feature inputs for downstream tasks such as intelligent question answering, content retrieval, and knowledge reasoning, thereby improving the model's understanding and processing capabilities of complex multimodal content.
[0098] Specifically, the method of combining the data feature groups can be selected according to application requirements.
[0099] In a feasible implementation, the data feature groups can be combined by feature splicing, and the image feature vectors, text feature vectors, and audio feature vectors in the same data block can be directly spliced in the feature dimension to form a higher-dimensional joint feature vector.
[0100] In another feasible implementation, the data feature groups can be weightedly fused based on the attention mechanism. By calculating the correlation weights between different modal features, the features of each modality are adaptively weighted and summed to form a fused feature representation that can highlight important modal information.
[0101] In another feasible implementation, a deep neural network can be used to perform nonlinear transformation and fusion on the data feature group, and each modal feature can be input into a multi-layer perceptron or transformer network to automatically discover the optimal feature combination pattern through end-to-end learning.
[0102] By adopting the embodiments of the present disclosure, multimodal data is converted into at least two data blocks based on the association relationship between different modal contents in the multimodal data, effectively retaining and utilizing the inherent connection of cross-modal information; by performing feature extraction and feature alignment processing on each data block separately, heterogeneous data that originally had different feature representation methods are converted into data feature groups with the same dimension, solving the problem in related technologies that different modal data are difficult to directly process due to feature heterogeneity; by combining each data feature group separately to obtain multimodal features that represent the results of knowledge base sorting, effective integration of multimodal information in the knowledge base is achieved, forming a unified multimodal feature representation, and enhancing the comprehensive understanding ability of artificial intelligence-related learning models of multimodal data.
[0103] In actual application scenarios, since the multimodal data obtained from the knowledge base often has complex temporal structures and heterogeneous characteristics, directly processing the overall data will face problems such as information redundancy, semantic fragmentation, and temporal misalignment between modalities.
[0104] To solve the above problem, based on the above embodiment, as an optional embodiment, operation S102 may further include the following operations:
[0105] Operation S201: acquiring image content, voice content, and text content in multimodal data;
[0106] Operation S202: determining multiple frames of target images in the image content; the difference between the target images in each frame is greater than a difference threshold;
[0107] Operation S203, converting the voice content into text content;
[0108] Operation S204 : for any frame of target image, the target image and at least one text content in a corresponding time sequence are combined into a data block.
[0109] In operation S201, image content refers to information carriers in multimodal data that exist in visual form, including but not limited to visual elements such as video frame sequences, static images, and charts; voice content refers to acoustic information in the form of audio signals in multimodal data, including but not limited to auditory elements such as speech recordings, background sound effects, and music clips; text content refers to language information in the form of text symbols in multimodal data, including but not limited to text elements such as subtitles, document paragraphs, and annotations.
[0110] In operation S202, the target image refers to a key frame image that can represent significant semantic changes or important visual information in the image content sequence. In the embodiment of the present disclosure, it can be understood as a representative image frame with scene transition, content update or semantic jump characteristics identified by visual difference calculation.
[0111] Among them, the difference threshold is a quantitative standard used to determine whether there are visual changes between adjacent image frames. By setting a reasonable threshold, meaningless changes caused by factors such as noise and slight jitter can be filtered out, ensuring that the screened target image truly reflects the substantial transformation of the content and avoiding repeated selection of similar frames, thereby providing a high-quality visual anchor for subsequent image and text alignment and data block construction.
[0112] In operation S203 , the voice content may be converted into text content through automatic speech recognition technology, thereby achieving a unified conversion from an auditory modality to a textual modality.
[0113] In operation S204, the selected target images are paired and combined with the text content within the corresponding time period based on the temporal correspondence to form a data block with clear semantic association. For any frame of the target image, by analyzing its timestamp in the original multimodal data, the text content paragraph corresponding to the target image in time is identified, including the speech transcription text at the same moment, related subtitle information, or context-related document text, etc., and the heterogeneous content with temporal correspondence or semantic pointing relationship is combined into a data block. The data block construction method based on dual temporal and semantic association ensures that the multimodal content within each data block logically supports each other and semantically complements each other.
[0114] By adopting the embodiments of the present disclosure, the problems of temporal alignment and semantic association in multimodal data processing are effectively solved, the interference of redundant information is avoided through target image screening, the unification of modal representation is achieved through speech-to-text conversion, and the logical integrity of cross-modal information is guaranteed by constructing data blocks based on temporal association.
[0115] In actual video content processing scenarios, raw video data often contains a large number of continuous and similar image frames with minimal visual differences between these frames. If all frames are processed without screening, this will lead to problems such as excessive redundant information, waste of computing resources, and low semantic information density. Furthermore, the presence of continuous similar frames will also affect the effectiveness and accuracy of subsequent multimodal feature fusion. To address the above issues, based on the above embodiment, as an optional embodiment, operation S202 may further include the following operations:
[0116] Operation S301: for any first image frame in the image content, perform a sliding window operation on the first image to obtain a second image;
[0117] Operation S302: Acquire a characteristic state of the second image in a color space; the characteristic state includes at least one of a hue state, a saturation state, and a brightness state;
[0118] Operation S303, calculating the difference in feature states between any two adjacent frames of the second image to obtain a visual difference;
[0119] In operation S304 , the latter second image frame of two adjacent second image frames whose visual difference is greater than a difference threshold is determined as a target image.
[0120] In operation S301, the sliding window operation refers to applying a filtering algorithm in a local area of the image, smoothing the image and reducing random noise introduced by factors such as device jitter, light changes, or compression artifacts by performing weighted averaging or other mathematical operations on the neighborhood pixel values. In the embodiment of the present disclosure, it can be understood as performing a weighted summation of the pixel values in the rectangular area around each pixel point of the first image according to the distance weight to generate a processed pixel value, thereby obtaining a second image with lower noise and more stable features.
[0121] For example, the specific formula may be as follows:
[0122] ;
[0123] Where, represents the visual state vector of the first image, Represents the value of the neighboring pixel in the (i, j) direction of the current pixel center, Represents the distance attenuation coefficient, which is used to control the effect of farther pixels on the current pixel, 0<λ<1; Represents the spatial distance, that is, the Euclidean distance between the current pixel and the neighboring pixel (i, j); Indicates the half-width of the neighborhood window, which is used to control the sliding window size and is usually 1 to 3.
[0124] In operation S302 , the second image is converted from the RGB color space to the HSV color space, and statistical features of the image in each channel are calculated respectively, so as to form a numerical representation that can represent the overall visual features of the image.
[0125] In operation S303, the visual difference can be calculated by using a weighted combination method, in which the hue state difference, saturation state difference and brightness state difference are linearly combined according to preset weights, where each state difference is calculated by a similarity measurement method such as Euclidean distance or cosine distance between corresponding feature vectors.
[0126] For example, the above process can be expressed by the following formula:
[0127] ;
[0128] Where, Indicates visual difference, Represents the hue component in the HSV color space, Represents the saturation component in HSV, Represents the V brightness component in HSV, Represent the corresponding weight coefficients respectively.
[0129] In operation S304, when the visual difference between two adjacent frames exceeds the threshold, it indicates that the subsequent frame has undergone significant visual changes relative to the previous frame, possibly corresponding to important semantic transition events such as scene changes, object movement, or content updates. Therefore, the second image in the subsequent frame is determined as the target image. The target image effectively summarizes the main visual content and semantic changes of the video, significantly reducing the number of images required for processing, and providing a high-quality visual anchor for subsequent image-text alignment and multimodal feature fusion.
[0130] For example, the above process can be expressed by the following formula:
[0131] ;
[0132] Where, if the result of the indicator function is 1, the current frame image is determined to be the target image; Indicates the basic threshold, which is used to control the sensitivity of scene switching. The larger the value, the stricter the control. Indicates the frame distance, that is, the number of frames between the current frame and the previous scene switching frame, used to reflect the time interval; The logarithm of the frame distance is used as a dynamic adjustment factor, so that the judgment threshold is gradually lowered over time (making it easier to trigger new scenes).
[0133] By adopting the embodiments of the present disclosure, the problem of excessive redundant frames in video data is effectively solved, the stability of image quality is improved through sliding window operations, a comprehensive description of the visual content of the image is achieved through color space feature state extraction, an objective inter-frame change metric is established through visual difference calculation, and finally, target images with important semantic value are screened out through a threshold judgment mechanism, which not only ensures the retention of key information but also greatly reduces the computational burden of data processing.
[0134] In actual multimodal data processing scenarios, due to the differences in feature representation, semantic density, structural complexity, etc. between data of different modalities, directly processing the original image content and text content will face problems such as inconsistent feature dimensions, difficult semantic alignment, and excessive computational complexity.
[0135] To solve the above problem, based on the above embodiment, as an optional embodiment, operation S103 may further include the following operations:
[0136] Operation S401: extracting features of different modal contents in each data block to obtain a first data feature set; the first data feature set includes text features and image features;
[0137] Operation S402: encoding features in each first data feature set based on associations between different modal contents in the data block to obtain a second data feature set;
[0138] In operation S403 , pruning is performed on the features in each second data feature set to obtain a data feature group.
[0139] In operation S401 , the first data feature set refers to a set of initial feature vectors obtained from different modal contents of a data block through a feature extraction algorithm, and includes a heterogeneous feature set of image feature vectors and text feature vectors.
[0140] Specifically, it is necessary to perform targeted feature extraction processing on the different modal contents in the data block to convert the original multimedia data into a computable numerical representation. For image content, deep learning models such as convolutional neural networks or visual transformers can be used to extract image features of the image. Image features can capture visual semantic information such as the edge, texture, shape, color distribution, etc. of the image to form a high-dimensional image feature vector. For text content, a pre-trained language model can be used for text encoding to convert the text sequence into a text feature vector that can represent semantic information. Through the above feature extraction process, the original image content and text content are respectively converted into corresponding image features and text features to form a first data feature set.
[0141] In operation S402, the second data feature set refers to a feature vector set that has been processed by association relationship encoding. In the embodiment of the present disclosure, it can be understood as an enhanced feature representation that incorporates cross-modal association information on the basis of the first data feature set. The data features in the second data feature set not only retain the original semantic information of each modality, but also increase the correspondence and structural information between modalities.
[0142] Specifically, encoding based on association relationships refers to the process of position encoding, semantic encoding or structural encoding of respective features based on the intrinsic connection between the image content and text content in the data block in terms of temporal, spatial, semantic and other dimensions.
[0143] For image features, position encoding can be performed based on the timestamp, spatial position or reference relationship between the image content and the text content in the original multimodal data, and the associated information can be incorporated into the image feature vector.
[0144] For text features, sequence encoding and semantic encoding can be performed based on the order of text content in the document, its correspondence with image content, or its hierarchical position in the overall semantic structure.
[0145] By encoding based on association relationships, a clear correspondence can be established between the originally independent image features and text features, forming a second data feature set with cross-modal semantic alignment capabilities.
[0146] In operation S403, pruning refers to the process of selecting the most representative and discriminative feature subset from the high-dimensional second data feature set through technical means such as feature selection, dimensionality reduction, and redundancy elimination, and adjusting it to a unified dimensional specification. In the embodiment of the present disclosure, based on methods such as feature importance evaluation and correlation analysis, the most valuable feature components for subsequent tasks can be identified and retained, while redundant or noise features can be eliminated, and features of different modalities can be unified into the same dimensional size through dimensionality transformation technology.
[0147] Correspondingly, the data feature group refers to the final feature representation unit formed after feature extraction, association coding and pruning processing. In the embodiment of the present disclosure, it can be understood as a multimodal feature set with unified dimensional specifications, retaining key semantic information, and eliminating redundant components, where each feature vector is compatible in mathematical structure and complements each other in semantic content.
[0148] By adopting the embodiments of the present disclosure, the problems of feature heterogeneity and dimensionality inconsistency in multimodal data processing are effectively solved. Through targeted feature extraction, the effective expression of semantic information of each modality is ensured. Through feature encoding based on association relationships, a cross-modal semantic alignment foundation is established. Through pruning processing, the dimensionality unification and redundancy elimination of features are achieved, forming a standardized data feature group.
[0149] In actual multimodal feature encoding processing scenarios, since the image features and text features in the first data feature set lack structured position information and sequence information, the feature representation cannot accurately reflect the spatial layout, temporal relationship and semantic structure of the original data, which seriously affects the accuracy of subsequent cross-modal alignment and fusion.
[0150] To solve the above problem, based on the above embodiment, as an optional embodiment, operation S402 may further include the following operations:
[0151] Operation S501: for each image feature in the first data feature set, divide the target image corresponding to the image feature into a plurality of image blocks, and encode the image feature based on the positional relationship between the image blocks to obtain an encoded image feature;
[0152] Operation S502: For each text feature in the first data feature set, encode the text feature based on the character order of the text content corresponding to the text feature to obtain an encoded text feature;
[0153] Operation S503, when the text content exists in the target image, encoding the text feature based on the position of the text content in the target image to obtain the encoded text feature;
[0154] In operation S504 , the encoded image features and the encoded text features in each first data feature set are used as a second data feature set.
[0155] Operation S501 may further include the following operations:
[0156] Operation S501a: dividing the target image into a plurality of image blocks according to a preset block division strategy; each image block has the same size specification;
[0157] Operation S501b: performing multi-scale feature extraction on each image block to obtain image block features;
[0158] Operation S501c: splicing and recombining the features of each small image block according to the original spatial arrangement relationship of the target image to obtain the encoded image features.
[0159] In operation S501a, the preset blocking strategy refers to an image segmentation method predetermined according to the image size and processing requirements. In the embodiment of the present disclosure, it can be understood as uniformly dividing the input target image according to a regular grid structure to form multiple small image blocks of the same size, so as to perform parallel feature extraction processing and subsequent spatial reconstruction operations.
[0160] Specifically, for a target image of size 3×w×h, where 3 represents the three RGB color channels, w represents the image width, and h represents the image height, a uniform grid segmentation strategy can be used to divide the image into (w / k)×(h / k) image blocks, where k represents the scale parameter of the block, which is used to control the size of each image block.
[0161] For example, the size of each image patch is 3×k×k. By setting an appropriate k value, we can ensure that each image patch contains sufficient local semantic information while maintaining spatial structural information. The total number of image patches obtained after segmentation is (w / k)×(h / k), and each image patch maintains its original spatial adjacency and arrangement order.
[0162] In operation S501b, multi-scale feature extraction refers to the process of obtaining feature representations at different levels of abstraction through a multi-layer convolutional neural network and fusing these features to form a comprehensive feature description. In the embodiment of the present disclosure, high-level semantic features and low-level detail features can be extracted for each image block separately, and then the feature information of different scales can be organically combined through feature fusion technology to form image block features that contain both detail information and semantic abstraction capabilities.
[0163] For example, small image blocks can be downsampled using a path-passing layer module. This refers to a downsampling structure that maintains feature channel connectivity. In the disclosed embodiments, this can be understood as reducing the spatial size of feature maps through strided convolution or pooling operations while preserving the original feature information through skip connections, thus avoiding the loss of important detail features during the downsampling process.
[0164] Then, a combination of convolutional pooling and activation functions is used to perform channel dimensionality reduction and nonlinear processing. Multiple layers of convolution are used to extract feature representations at different levels. Shallow convolutions primarily capture low-level visual features such as edges and textures, while deeper convolutions extract more abstract, high-level semantic features. Pooling is used to reduce feature dimensionality and enhance spatial invariance, while activation functions introduce nonlinear transformation capabilities.
[0165] Finally, the spatial pyramid pooling module performs pooling operations on the feature maps at various scales, concatenating the pooling results at different scales to form a fixed-dimensional feature representation that captures multi-scale spatial information. Through this multi-scale feature extraction process, each 3×k×k image patch can be converted into a d-dimensional image patch feature vector, where d represents the dimension of the feature vector.
[0166] In operation S501c, splicing and reorganization refers to the process of organizing and splicing the features of each image block in the original arrangement order according to the spatial position relationship of the image blocks in the original target image, so as to form an overall feature representation that maintains the spatial structure information. In the embodiment of the present disclosure, it can be understood as rearranging the (w / k)×(h / k) d-dimensional feature vectors that have been feature extracted according to their grid positions in the original image to form encoded image features with a spatial correspondence.
[0167] In operation S502, each text sentence in the structured document can be semantically encoded to obtain a sentence encoding vector; a serial position encoding is generated based on the sequential position of each sentence in the document; and the sentence encoding vector is fused with the corresponding serial position encoding to obtain a text feature vector containing contextual sequence information.
[0168] For example, the above process can be expressed by the following formula:
[0169] The sentence encoding vector is obtained by semantically embedding the text sentence and is expressed as:
[0170]
[0171] Where, represents the encoding vector of the i-th sentence, Represents the content of the i-th sentence, represents the sentence-level encoding function.
[0172] The generation formula of sequence position encoding is:
[0173]
[0174] Where, represents the positional encoding of the i-th sentence, Indicates the sentence number, Represents the position encoding function.
[0175] The calculation formula of the text feature vector containing context order information is:
[0176]
[0177] Where, The text feature vector representing the position of the fused sequence has an output dimension of m×d, where m represents the number of sentences and d represents the feature dimension.
[0178] In operation S503, text areas in the target image can be identified through text technology to obtain the bounding box coordinates and corresponding text content of each text area; word embedding processing can be performed on each text content to obtain a text semantic vector; a spatial position encoding vector can be generated based on the bounding box coordinates of each text area; the text semantic vector can be fused with the corresponding spatial position encoding vector to obtain a text feature vector of the fused visual position.
[0179] Among them, the bounding box coordinates refer to the spatial position parameters of the rectangular detection box output when the OCR technology recognizes the text area, and in the embodiment of the present disclosure, include the coordinates of the upper left corner (x1, y1) and the lower right corner (x2, y2) of the text area.
[0180] The text semantic vector is obtained by embedding the text content recognized by OCR. The formula is expressed as:
[0181]
[0182] Where, Represents the word vector encoding of the k-th OCR text content, Represents the k-th text content, and WordEmbedding represents the word embedding function.
[0183] The generation formula of the spatial position encoding vector is expressed as:
[0184]
[0185] Where, represents the position encoding vector generated based on the bounding box coordinates, and f(·) represents the position encoding function, which maps the four coordinate parameters into a high-dimensional vector representation.
[0186] The calculation formula of the text feature vector fused with visual position can be expressed as:
[0187]
[0188] Where, The text feature vector representing the fused visual position, represents the feature fusion operation, Represents a dimension resizing operation.
[0189] By adopting the embodiments of the present disclosure, the problems of missing structural information and ambiguous association relationships in multimodal feature encoding are effectively solved. The spatial perception ability of image features is enhanced through image block division and position encoding. The word order and logical relationship of text features are retained through text sorting encoding. The precise image-text spatial correspondence is established through the spatial position encoding of text in the image, and finally a second data feature set with complete structure and clear association is formed, which improves the feature expression quality of multimodal data and the accuracy of cross-modal alignment.
[0190] In actual multimodal feature processing scenarios, since the second data feature set often has a high feature dimension and complex feature distribution, directly using these high-dimensional features for subsequent multimodal fusion will face problems such as high computational complexity, huge storage overhead, and information noise caused by feature redundancy. At the same time, the differences in dimensional specifications of features from different modal sources will also affect the effectiveness and accuracy of feature combination. To solve the above problems, based on the above embodiment, as an optional embodiment, operation S403 can also specifically include the following operations:
[0191] Operation S601: for any second data feature set, construct a correlation matrix between features in the second data feature set;
[0192] Operation S602: combining the feature columns in the correlation matrix in sequence based on the order of correlation between the feature columns in the correlation matrix from strong to weak, until the dimension of the feature columns in the correlation matrix is reduced to the first dimension;
[0193] Operation S603, combining the characteristic rows in the correlation matrix in sequence based on the order of correlation between the characteristic rows in the correlation matrix from strong to weak, until the dimension of the characteristic rows in the correlation matrix is reduced to the second dimension;
[0194] Operation S604: output the second data feature set after dimensionality reduction as a data feature group.
[0195] In operation S601 , the correlation matrix refers to a two-dimensional numerical matrix that describes the degree of correlation between feature vectors in the second data feature set, and each element in the matrix represents the correlation strength between two corresponding features.
[0196] Specifically, for the second data feature set containing image features and text features, the feature vectors of each modality are first organized in a unified matrix format to form a feature matrix representation, and then a correlation matrix is constructed by calculating the correlation scores between the feature vectors pair by pair. This matrix can quantitatively reflect the dependency relationship and the degree of information overlap between different features.
[0197] In operation S602, the correlation score between each pair of feature columns in the correlation matrix is calculated, and then the feature column pairs with the strongest correlation are identified based on the correlation scores sorted from high to low. These highly correlated feature columns are merged into new feature columns through linear combination or nonlinear transformation. The above process is repeated until the total number of feature columns is reduced to a preset first dimension, where the first dimension is a target dimension value predetermined based on computing resource limitations and task accuracy requirements.
[0198] In operation S603, correlation analysis and combination processing are performed on the feature rows in the correlation matrix using a similar method to that used for feature column processing. Through appropriate weight assignment and feature fusion strategies, the information from multiple related feature rows is integrated into a single representative feature row. This process is repeated until the dimensionality of the feature rows is reduced to a preset second dimension, where the second dimension is also a preset target number of rows based on storage efficiency and processing performance.
[0199] By adopting the embodiments of the present disclosure, the problems of computational complexity and storage overhead in high-dimensional feature processing are effectively solved, an objective feature association metric is established through correlation matrix construction, dimensionality reduction with maximum information retention is achieved through the combination of feature columns and feature rows based on correlation strength, and a data feature group with a compact structure and rich semantics is formed through double dimensionality reduction processing, which not only improves the computational efficiency of subsequent processing, but also ensures the expression quality and fusion effect of multimodal information.
[0200] In actual multimodal feature combination scenarios, although each data feature group has achieved dimensional unification and redundancy elimination, these data feature groups still exist in a scattered form, lacking a unified feature representation and cross-modal semantic integration. Directly using the scattered data feature groups for downstream tasks will face problems such as incomplete feature expression, information fragmentation between modalities, and limited semantic understanding capabilities. At the same time, the complementary information and association relationships between different data feature groups cannot be fully utilized. To solve the above problems, based on the above embodiment, as an optional embodiment, operation S104 can also specifically include the following operations:
[0201] Operation S701: combining the data features in each data feature group to obtain a target feature;
[0202] In operation S702 , each target feature is fused into a multimodal feature.
[0203] In operation S701, the target feature refers to the feature representation formed by the splicing operation. In the embodiment of the present disclosure, it can be understood as a unified feature vector that simultaneously contains image semantic information, text semantic information, and cross-modal association information. The target feature establishes a connection relationship between modalities while maintaining the uniqueness of each modal feature.
[0204] In operation S702, multimodal features refer to comprehensive feature representations with a unified semantic space and enhanced expression capabilities formed through fusion processing. In the disclosed embodiment, they can be understood as not only containing the original semantic information of each modality, but also incorporating cross-modal association knowledge and high-level semantic expressions of contextual understanding. The multimodal features can provide rich semantic feature inputs for downstream tasks such as intelligent question answering, content retrieval, and knowledge reasoning.
[0205] Specifically, in the fusion process, the target features corresponding to each data block are first sent as input to the fusion model. The fusion model can adopt deep learning architectures such as multi-head attention mechanism, convolutional neural network, recurrent neural network or transformer network. Then, the fusion model automatically discovers the optimal feature combination method and weight distribution strategy by learning the mutual relationship and dependency pattern between the target features. In this process, it can capture the complementary information, correlation pattern and semantic correspondence between different target features. Finally, the fusion model integrates the scattered target features into a unified multimodal feature representation.
[0206] Figure 2 FIG. 1 shows a schematic diagram of a multimodal feature fusion network provided by an embodiment of the present disclosure. Figure 2 As shown in Figure 3, the fusion network adopts a layered processing architecture to achieve unified processing and deep fusion of different modal features.
[0207] In this fusion network, the input multimodal data features include image features 201, OCR text features 202, and plain text features 203. Image features 201 have a three-channel structure, corresponding to the red, green, and blue components of the RGB color space. Each channel carries the semantic information and spatial structure of the image in the corresponding color dimension. OCR text features 202 and plain text features 203 both use single-channel structures, focusing on expressing the spatial distribution characteristics of text in the image and the semantic logical relationships of the plain text, respectively. These three different modal features are spliced along the channel dimension to form a unified feature representation consisting of five channels.
[0208] The concatenated features are first processed by the VIT-patch module 204, which can be the patch processing mechanism of the Vision Transformer. It reorganizes the fused features into a patch sequence. Each patch contains local multimodal information. The patching process maintains the spatial structure and semantic integrity of the original features, while creating a suitable input format for subsequent attention mechanism processing. The processed patch sequence is fed into the stacked Transformer attention module 212, which adopts the stacked structure design of the typical Transformer architecture and contains K attention blocks of the same structure. Each attention block integrates the standard Transformer attention mechanism, including two core components: a multi-head attention mechanism and a feedforward neural network.
[0209] In the stacked attention module 212, each attention block processes input features sequentially according to a unified process flow. First, through the multi-head attention processing stage 208, the input features are mapped into a query vector, a key vector, and a value vector. Adaptive calculation of attention weights enables intelligent weighted fusion of features from different modalities. The multi-head attention 1 module 206 and the multi-head attention 2 module 205 represent multi-head attention mechanisms at different levels of the stacked structure, responsible for capturing local associations and global semantic dependencies, respectively. Subsequently, in the residual connection and layer normalization 209 process, the attention output is residually connected to the original input features. Layer normalization stabilizes the feature distribution, ensuring the stability and consistency of feature values during the fusion process. The features processed by the attention mechanism are further nonlinearly transformed through a feedforward neural network 210. This network uses a multi-layer perceptron structure to perform deep nonlinear mapping on the fused features, further improving the expressiveness and semantic abstraction of the features. The residual connection 211 mechanism is then applied again to fuse the output of the feedforward network with its input, completing the complete processing flow of a single attention block.
[0210] The final fused features are input into the classifier MLP213, which performs task-related mapping and output on the multimodal features according to the specific downstream task requirements. MLP213 supports various application scenarios including alignment tasks 214, filling tasks 215, and generation tasks 216, and achieves flexible adaptation to different task types through unified feature representation.
[0211] By adopting the embodiments of the present disclosure, the problems of feature dispersion and information fragmentation in multimodal data processing are effectively solved. The integration of different modal information in the same data block is achieved through data feature splicing, and a basic representation framework for multimodal features is established. Through target feature fusion, deep semantic integration and collaborative enhancement across data blocks are achieved, forming multimodal features with global semantic understanding capabilities, which not only ensures the complete preservation of each modal information, but also realizes the effective integration of cross-modal knowledge.
[0212] Based on the above embodiment, as an optional embodiment, the above method can also input multimodal data into the fusion model to obtain multimodal features output by the fusion model.
[0213] Exemplarily, the fusion model may include a knowledge conversion module, a feature alignment module, and a splicing and fusion module, wherein:
[0214] A knowledge conversion module, configured to convert the multimodal data into at least two data blocks based on associations between different modal contents in the multimodal data;
[0215] A feature alignment module is used to extract and align features of each data block to obtain at least one data feature group; the dimensions of the data features in the data feature group are the same;
[0216] The splicing and fusion module is used to combine each data feature group separately to obtain multimodal features; the multimodal features represent the combing results of the knowledge base.
[0217] The following describes the training process of the fusion model:
[0218] The fusion model uses a two-stage training strategy to ensure its multimodal understanding and task adaptability. The training process first pre-trains the splicing and fusion module using a multi-task joint loss. This pre-trained module is then fine-tuned end-to-end with other components to ultimately achieve a fusion model with multimodal processing capabilities.
[0219] During the first stage of pre-training, the disclosed embodiment designs three complementary self-supervised learning tasks to train the splicing and fusion module.
[0220] The feature comparison task enhances modality alignment by determining whether the fused features contain all modal information in the original data. Specifically, the data feature group is input into the splicing and fusion module to obtain the fused features, which are then compared with the original features or noise features. A multi-layer perceptron classifier is used to determine whether they match. Because the integrity of the information in the three modalities of image, text, and audio needs to be verified, this task is designed as a three-label multi-classification problem, and the loss function is:
[0221]
[0222] Where, represents feature contrast loss; Represents the binary cross entropy loss function, which combines Sigmoid activation and BCE cross entropy loss; represents a multilayer perceptron classifier; Represents the fusion feature vector output by the splicing and fusion module; represents the original modal eigenvector (positive sample case) or the random noise eigenvector (negative sample case); Represents the feature fusion operation, which can be concatenation, element-wise multiplication, or other fusion methods.
[0223] The text filling task enhances the ability to recover local details by predicting the obscured content. During training, text words and image patches are randomly masked, and the fusion model is allowed to predict the missing parts. The loss function is:
[0224]
[0225] Where, Represents the loss value of the text completion task; represents the mean square error loss function; Represents the similarity calculation function, which is used to measure the similarity between the predicted features and the real features; A feature vector representing the model's predictions; The feature vector corresponding to the real masked content; The true label value representing the similarity, usually a pre-calculated similarity score.
[0226] The sentence generation task is used to enhance the cross-modal context modeling capability. It predicts subsequent content by inputting previous text and image information and establishes temporal dependencies. The loss function is:
[0227]
[0228] Where, Represents the loss value of the sentence generation task; Represents the feature vector predicted based on the multimodal information at time t-1; Represents the feature vector corresponding to the true multimodal content at time t; the meanings of other parameters are the same as those defined in the text completion task.
[0229] The losses of the above three tasks are weighted and combined to form the final joint loss function:
[0230]
[0231] Where, represents the final multi-task joint loss; Represents the weight coefficient of the feature comparison task, which is used to control the training intensity of the modality alignment capability; Represents the weight coefficient of the text filling task, which is used to control the training intensity of the local detail restoration ability; Represents the weight coefficient of the sentence generation task, which is used to control the training intensity of the cross-modal context modeling ability; the specific values of the three weight coefficients are adjusted according to the task requirements and experimental results.
[0232] The joint fine-tuning process in the second phase coordinates the optimization of the pre-trained splicing and fusion module with the knowledge conversion module and feature alignment module. During the parameter initialization phase, the splicing and fusion module is initialized using the pre-trained parameters from the first phase, while the knowledge conversion module and feature alignment module use random initialization or related pre-trained weights. Joint fine-tuning adopts an end-to-end training strategy, updating the parameters of all modules simultaneously through the gradient backpropagation mechanism. However, a smaller learning rate is used for fine-tuning the pre-trained splicing and fusion module, while a normal learning rate is used for fully optimizing other modules. This differentiated learning rate setting not only protects the effective feature representation obtained through pre-training, but also allows each module to be adaptively adjusted according to specific task requirements.
[0233] During the joint fine-tuning process, we design task loss functions tailored to the specific downstream application tasks, such as cross-entropy loss for question-answering tasks and ranking loss for retrieval tasks. These loss functions are then combined with the multi-task joint loss for comprehensive optimization. The training process utilizes an early stopping mechanism and a learning rate decay strategy. Model convergence is determined by monitoring performance on the validation set. Training is terminated when performance no longer significantly improves, ensuring that the model is fully trained while avoiding overfitting.
[0234] Through this two-stage training strategy, the fusion model not only possesses powerful multimodal information understanding and alignment capabilities, but also adapts to the needs of diverse application scenarios. The pre-training phase ensures the model's fundamental understanding of multimodal data, while the joint fine-tuning phase establishes synergy between modules. Ultimately, this results in a high-performance multimodal fusion model with end-to-end processing capabilities, effectively supporting the application needs of a variety of downstream tasks, including intelligent question answering, content retrieval, and knowledge reasoning.
[0235] Figure 3 The figure schematically shows a structural block diagram of a multimodal data processing device according to an embodiment of the present disclosure.
[0236] The multimodal data processing device may include:
[0237] An acquisition module 301 is used to obtain multimodal data;
[0238] A conversion module 302 is configured to convert the multimodal data into at least two data blocks based on associations between different modal contents in the multimodal data;
[0239] An alignment module 303 is configured to extract and align features of each of the data blocks to obtain at least one data feature group; the data features in the data feature group have the same dimension;
[0240] The combining module 304 is configured to combine the data feature groups to obtain multimodal features. The multimodal features represent the combing results of the knowledge base.
[0241] Based on the above embodiment, as an optional embodiment, the conversion module 302 is further configured to:
[0242] Acquiring image content, voice content, and text content in the multimodal data;
[0243] Determining multiple frames of target images in the image content; wherein the difference between the target images in each frame is greater than a difference threshold;
[0244] Converting the speech content into text content;
[0245] For any frame of the target image, the target image and at least one text content in a corresponding time sequence form a data block.
[0246] Based on the above embodiment, as an optional embodiment, the conversion module 302 is further configured to:
[0247] For any first image frame in the image content, performing a sliding window operation on the first image to obtain a second image;
[0248] Acquire a characteristic state of the second image in a color space; the characteristic state includes at least one of a hue state, a saturation state, and a brightness state;
[0249] Calculate the difference in feature states between any two adjacent frames of the second image to obtain the visual difference;
[0250] The latter second image frame of two adjacent second image frames whose visual difference is greater than the difference threshold is determined as the target image.
[0251] Based on the above embodiment, as an optional embodiment, the alignment module 303 is further configured to:
[0252] Extracting features of different modal contents in each of the data blocks to obtain a first data feature set; the first data feature set includes text features and image features;
[0253] encoding the features in each of the first data feature sets based on associations between different modal contents in the data block to obtain a second data feature set;
[0254] Pruning is performed on the features in each of the second data feature sets to obtain a data feature group.
[0255] Based on the above embodiment, as an optional embodiment, the alignment module 303 is further configured to:
[0256] For each image feature in the first data feature set, dividing the target image corresponding to the image feature into a plurality of image blocks, and encoding the image feature based on a positional relationship between the image blocks to obtain an encoded image feature;
[0257] For each text feature in the first data feature set, encoding the text feature based on the character order of the text content corresponding to the text feature to obtain an encoded text feature;
[0258] In the case where text content exists in the target image, encoding the text feature based on the position of the text content in the target image to obtain an encoded text feature;
[0259] The encoded image features and the encoded text features in each of the first data feature sets are used as a second data feature set.
[0260] Based on the above embodiment, as an optional embodiment, the alignment module 303 is further configured to:
[0261] For any second data feature set, construct a correlation matrix between features in the second data feature set;
[0262] Combining the feature columns in the correlation matrix in sequence based on the order of correlation between the feature columns in the correlation matrix from strong to weak until the dimension of the feature columns in the correlation matrix is reduced to the first dimension;
[0263] Combining the characteristic rows in the correlation matrix in sequence based on the order of correlation between the characteristic rows in the correlation matrix from strong to weak until the dimension of the characteristic rows in the correlation matrix is reduced to the second dimension;
[0264] The second data feature set after dimensionality reduction is output as a data feature group.
[0265] Based on the above embodiment, as an optional embodiment, the combination module 304 is further configured to:
[0266] Splicing the data features in each of the data feature groups to obtain a target feature;
[0267] The target features are fused into the multimodal features.
[0268] According to the embodiments of the present invention, any number of modules, sub-modules, units, and sub-units, or at least part of the functions of any number of them, can be implemented in one module. According to the embodiments of the present invention, any one or more of the modules, sub-modules, units, and sub-units can be split into multiple modules for implementation. According to the embodiments of the present invention, any one or more of the modules, sub-modules, units, and sub-units can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented by hardware or firmware in any other reasonable way of integrating or packaging the circuit, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, according to the embodiments of the present invention, one or more of the modules, sub-modules, units, and sub-units can be at least partially implemented as a computer program module, which can perform the corresponding functions when the computer program module is executed.
[0269] For example, any number of acquisition module 301 and conversion module 302 can be combined into a single module / unit / sub-unit, or any one of these modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least part of the functionality of one or more of these modules / units / sub-units can be combined with at least part of the functionality of other modules / units / sub-units and implemented in a single module / unit / sub-unit. According to embodiments of the present disclosure, at least one of acquisition module 301 and conversion module 302 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or can be implemented in any one of software, hardware, and firmware, or any suitable combination of these. Alternatively, at least one of acquisition module 301 and conversion module 302 can be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.
[0270] Figure 4 A block diagram of an electronic device provided by an embodiment of the present disclosure is schematically shown.
[0271] like Figure 4 As shown, the electronic device 400 according to an embodiment of the present disclosure includes a processor 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage portion 408 into a random access memory (RAM) 403. The processor 401 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 401 may also include onboard memory for caching purposes. The processor 401 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present disclosure.
[0272] Various programs and data required for the operation of the electronic device 400 are stored in the RAM 403. The processor 401, ROM 402, and RAM 403 are connected to each other via a bus 404. The processor 401 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 402 and / or RAM 403. It should be noted that the programs may also be stored in one or more memories other than the ROM 402 and RAM 403. The processor 401 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0273] According to an embodiment of the present disclosure, electronic device 400 may further include an input / output (I / O) interface 405, which is also connected to bus 404. Electronic device 400 may also include one or more of the following components connected to I / O interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 408 including a hard disk; and a communication section 409 including a network interface card such as a LAN card or modem. Communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to I / O interface 405 as needed. Removable media 411, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 410 as needed, so that computer programs read from the removable media can be installed into storage section 408 as needed.
[0274] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.
[0275] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 402 and / or RAM 403 described above, and / or one or more memories other than ROM 402 and RAM 403.
[0276] The embodiments of the present disclosure are described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be used in combination to advantage. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. A multimodal data processing method, comprising: Obtain multimodal data from the knowledge base; Converting the multimodal data into at least two data blocks based on associations between different modal contents in the multimodal data; Performing feature extraction and feature alignment on each of the data blocks to obtain at least one data feature group; the dimensions of the data features in the data feature group are the same; Each of the data feature groups is combined to obtain multimodal features; the multimodal features represent the combing results of the knowledge base.
2. The method according to claim 1, wherein converting the multimodal data into at least two data blocks based on the association relationship between different modal contents in the multimodal data comprises: Acquiring image content, voice content, and text content in the multimodal data; Determining multiple frames of target images in the image content; The difference between the target images in each frame is greater than a difference threshold; Converting the speech content into text content; For any frame of the target image, the target image and at least one text content in a corresponding time sequence form a data block.
3. The method according to claim 2, wherein determining the multiple frames of target images in the image content comprises: For any first image frame in the image content, performing a sliding window operation on the first image to obtain a second image; Acquiring a characteristic state of the second image in a color space; The characteristic state includes at least one of a hue state, a saturation state, and a brightness state; Calculate the difference in feature states between any two adjacent frames of the second image to obtain the visual difference; The latter second image frame of two adjacent second image frames whose visual difference is greater than the difference threshold is determined as the target image.
4. The method according to claim 1, wherein the extracting and aligning features of each data block to obtain at least one data feature group comprises: Extracting features of different modal contents in each of the data blocks to obtain a first data feature set; The first data feature set includes text features and image features; encoding the features in each of the first data feature sets based on associations between different modal contents in the data block to obtain a second data feature set; Pruning is performed on the features in each of the second data feature sets to obtain a data feature group.
5. The method according to claim 4, wherein encoding the features in each of the first data feature sets based on the association relationship between different modal contents in the data block to obtain the second data feature set comprises: For each image feature in the first data feature set, dividing the target image corresponding to the image feature into a plurality of image blocks, and encoding the image feature based on a positional relationship between the image blocks to obtain an encoded image feature; For each text feature in the first data feature set, encoding the text feature based on the character order of the text content corresponding to the text feature to obtain an encoded text feature; In the case where text content exists in the target image, encoding the text feature based on the position of the text content in the target image to obtain an encoded text feature; The encoded image features and the encoded text features in each of the first data feature sets are used as a second data feature set.
6. The method according to claim 4, wherein pruning the features in each of the second data feature sets to obtain a data feature group comprises: For any second data feature set, construct a correlation matrix between features in the second data feature set; Combining the feature columns in the correlation matrix in sequence based on the order of correlation between the feature columns in the correlation matrix from strong to weak until the dimension of the feature columns in the correlation matrix is reduced to the first dimension; Combining the characteristic rows in the correlation matrix in sequence based on the order of correlation between the characteristic rows in the correlation matrix from strong to weak until the dimension of the characteristic rows in the correlation matrix is reduced to the second dimension; The second data feature set after dimensionality reduction is output as a data feature group.
7. The method according to claim 1, wherein the combining of the data feature groups to obtain multimodal features comprises: Splicing the data features in each of the data feature groups to obtain a target feature; The target features are fused into the multimodal features.
8. The method according to claim 1, further comprising: The multimodal data is input into a fusion model to obtain multimodal features output by the fusion model.
9. A multimodal data processing device, comprising: Acquisition module, used to obtain multimodal data; a conversion module, configured to convert the multimodal data into at least two data blocks based on associations between different modal contents in the multimodal data; an alignment module, configured to extract and align features of each of the data blocks to obtain at least one data feature group; the data features in the data feature group have the same dimension; The combination module is used to combine each of the data feature groups to obtain multimodal features; the multimodal features represent the combing results of the knowledge base.
10. A computer-readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Fine-grained video emotion content question and answer method and system based on multi-modal data
CN116226347A
Pre-training document model alignment optimization method based on comparative learning
CN116311323A
Football match comprehensive sports performance evaluation method and system
CN118552088A
Brassica oleracea knowledge dynamic expression and interaction method and system based on multi-modal fusion
CN119862954A
Cited By
PDF intelligent retrieval and generation method and system based on RAG
CN120994845A
RAG-based pdf intelligent retrieval and generation method and system
CN120994845B