Multi-modal data processing method and device, equipment and storage medium
By unifying the video and image modalities of multimodal data into text modalities and using large language models to process text features, the problems of performance degradation and high computational cost in multimodal data processing are solved, and efficient multimodal data processing is achieved.
Patent Information
- Application Number
- CN202411915913.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-06
AI Technical Summary
Multimodal large language models have problems of degradation in performance and high computational costs when processing longer videos, higher resolution images, and decisions based on more historical information.
By obtaining multimodal data and processing instructions, determining the data label and modal types, unifying the video and image modalities into text modalities, extracting and mapping features using visual encoder and feature mapper, and finally using a large language model to process text features and instructions to generate processing results.
The unified representation of data of different modalities is realized, which improves the efficiency of multimodal data processing, reduces the complexity brought about by heterogeneity between modalities, and reduces the consumption of computing resources when processing high-resolution images and long videos.
Smart Images

Figure CN119939211A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method and device for processing multimodal data, an electronic device, and a storage medium. Background Art
[0002] Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in processing multiple modal data such as text, images, and videos, but their application in multi-image understanding scenarios still faces challenges. In particular, when processing longer videos, higher-resolution images, and decisions based on more historical information, the model's context length extension and computational efficiency become key issues.
[0003] In related research, although some progress has been made in constructing long-context training data and optimizing training strategies, the problems of performance degradation and high computational cost when processing multimodal data have not been fully addressed. Summary of the invention
[0004] The embodiments of the present application provide a method for processing multimodal data to solve the problems of performance degradation and high computational cost in the process of multimodal data processing.
[0005] Correspondingly, an embodiment of the present application also provides a multimodal data processing device, an electronic device and a storage medium to ensure the implementation and application of the above method.
[0006] In order to solve the above problems, the present application discloses a method for processing multimodal data, which includes:
[0007] Acquire multimodal data and processing instructions corresponding to the multimodal data;
[0008] Determining data labels corresponding to the multimodal data;
[0009] Determining a data modality corresponding to the multimodal data according to a data label corresponding to the multimodal data; the data modality includes a video modality, an image modality, and a text modality;
[0010] Unifying the data modality corresponding to the multimodal data from the video modality and the image modality into the text modality to obtain text features corresponding to the multimodal data;
[0011] According to the text features corresponding to the multimodal data and the processing instructions, a processing result corresponding to the multimodal data is obtained.
[0012] Optionally, the data tags corresponding to the multimodal data include image tags and video tags, and determining the data modality corresponding to the multimodal data according to the data tags corresponding to the multimodal data includes:
[0013] If the data label corresponding to the multimodal data is the image label, determining that the data modality corresponding to the multimodal data is the image modality;
[0014] If the data tag corresponding to the multimodal data is the video tag, determining that the data modality corresponding to the multimodal data is the video modality;
[0015] If the data tag does not exist in the multimodal data, it is determined that the data modality corresponding to the multimodal data is the text modality.
[0016] Optionally, unifying the data modality corresponding to the multimodal data from the video modality and the image modality into the text modality to obtain text features corresponding to the multimodal data includes:
[0017] Build visual encoder and feature mapper;
[0018] If the data modality corresponding to the multimodal data is the image modality, extracting image features in the multimodal data using the visual encoder;
[0019] The feature mapper is used to map the image features into a text space corresponding to the text modality to obtain text features corresponding to the multimodal data.
[0020] Optionally, unifying the data modality corresponding to the multimodal data from the video modality and the image modality into the text modality to obtain text features corresponding to the multimodal data includes:
[0021] Build visual encoder and feature mapper;
[0022] If the data modality corresponding to the multimodal data is the video modality, splitting the multimodal data into a plurality of image data based on the video frame; the data modality corresponding to the image data is the image modality;
[0023] extracting image features from the multimodal data using the visual encoder;
[0024] The feature mapper is used to map the image features into a text space corresponding to the text modality to obtain text features corresponding to the multimodal data.
[0025] Optionally, the image feature has a corresponding dimension, the dimension including height and width, and after extracting the image feature in the multimodal data using the visual encoder, the method further includes:
[0026] Determine a pooling window for selecting a specific area on the image feature and a pooling step size for moving the pooling window on the image feature; the pooling window includes a window height and a window width, and the pooling step size includes a step height and a step width;
[0027] Calculate the target height according to the height, the window height and the step height;
[0028] Calculate the target width according to the width, the window width and the step width;
[0029] The target height is replaced by the height corresponding to the image feature, and the target width is replaced by the width corresponding to the image feature, to obtain a target dimension corresponding to the image feature; the target dimension is smaller than the dimension.
[0030] Optionally, obtaining a processing result corresponding to the multimodal data according to the text feature corresponding to the multimodal data and the processing instruction includes:
[0031] Build a large language model;
[0032] The large language model is used to process the text features corresponding to the multimodal data and the processing instructions to obtain a processing result corresponding to the multimodal data.
[0033] Optionally, the multimodal data includes image data whose data modality is the image modality, and the using the large language model to process the text features and the processing instructions corresponding to the multimodal data to obtain the processing results corresponding to the multimodal data includes:
[0034] If the processing instruction is a single image instruction, the single image instruction is used to infer text features corresponding to the image data to generate a first processing result for describing the image data.
[0035] Optionally, the multimodal data further includes video data whose data modality is the video modality, and the using the large language model to process the text features and the processing instructions corresponding to the multimodal data to obtain the processing results corresponding to the multimodal data includes:
[0036] If the processing instruction is a multi-image instruction, the multi-image instruction is used to infer text features corresponding to the video data, thereby generating a second processing result for describing the video data.
[0037] Optionally, the multimodal data further includes text data whose data modality is the text modality, and the adopting of the large language model to process the text features and the processing instructions corresponding to the multimodal data to obtain the processing results corresponding to the multimodal data includes:
[0038] If the processing instruction is a plain text instruction, the plain text instruction is used to infer the text features corresponding to the text data to generate a third processing result for describing the text data.
[0039] Optionally, the image data has a corresponding resolution, and the using the single image instruction to infer text features corresponding to the image data to generate a first processing result for describing the image data includes:
[0040] If the resolution corresponding to the image data exceeds the preset resolution, traversing from the upper left corner of the image data to the lower right corner of the image data, the image data is divided into a plurality of sub-image data;
[0041] Using the single image instruction to infer text features corresponding to the sub-image data, and generating a sub-processing result for describing the sub-image data;
[0042] The plurality of sub-processing results are combined to obtain the first processing result.
[0043] The embodiment of the present application also discloses a multimodal data processing device, the device comprising:
[0044] An acquisition module, used to acquire multimodal data and processing instructions corresponding to the multimodal data;
[0045] A label determination module, used to determine the data label corresponding to the multimodal data;
[0046] A modality determination module, configured to determine a data modality corresponding to the multimodal data according to a data label corresponding to the multimodal data; the data modality includes a video modality, an image modality, and a text modality;
[0047] A modality unification module, used to unify the data modality corresponding to the multimodal data from the video modality and the image modality into the text modality, and obtain text features corresponding to the multimodal data;
[0048] The data processing module is used to obtain the processing result corresponding to the multimodal data according to the text features corresponding to the multimodal data and the processing instructions.
[0049] An embodiment of the present application further discloses an electronic device, comprising: a processor; and a memory, on which executable code is stored. When the executable code is executed, the processor executes a method for processing multimodal data as described in any one of the embodiments of the present application.
[0050] The embodiments of the present application also disclose one or more machine-readable media on which executable codes are stored. When the executable codes are executed, a processor executes a method for processing multimodal data as described in any one of the embodiments of the present application.
[0051] Compared with the prior art, the embodiments of the present application have the following advantages:
[0052] In an embodiment of the present application, multimodal data and processing instructions corresponding to the multimodal data are obtained; data labels corresponding to the multimodal data are determined; data modalities corresponding to the multimodal data are determined according to data labels corresponding to the multimodal data; the data modalities include video modalities, image modalities, and text modalities; the data modalities corresponding to the multimodal data are unified from the video modality and the image modality to the text modality to obtain text features corresponding to the multimodal data; and processing results corresponding to the multimodal data are obtained according to the text features corresponding to the multimodal data and the processing instructions. By uniformly converting the video modality and the image modality into the text modality, a unified representation of data of different modalities is achieved, so that the model can process multimodal data more efficiently, reducing the complexity caused by the heterogeneity between modalities, and reducing the amount of calculation of the model when processing high-resolution images and long videos, thereby reducing the consumption of computing resources and improving computing efficiency. Moreover, with clear data labels, the model can more easily distinguish between text, images, and video frames, avoiding confusion between modalities. After the data modality is clear, it can be processed according to different modalities, optimizing computing resources and thus improving the accuracy of data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a flowchart of a multimodal data processing method embodiment of the present application;
[0054] Figure 2 It is a schematic diagram of the architecture of a large language model of this application;
[0055] Figure 3 It is a structural block diagram of an embodiment of a multimodal data processing device of the present application;
[0056] Figure 4 It is a schematic diagram of the structure of a device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0057] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0058] Reference Figure 1 , is a flowchart of a multimodal data processing method embodiment of the present application, comprising the following steps:
[0059] Step 101: Acquire multimodal data and processing instructions corresponding to the multimodal data;
[0060] In an embodiment of the present application, the model first needs to obtain input multimodal data (such as text, images, videos, etc.) and processing instructions given by the user or the system (such as generating descriptions, answering questions, etc.).
[0061] Step 102: Determine the data label corresponding to the multimodal data;
[0062] In the embodiment of the present application, in order to determine whether the realization and spatial dependency between images can be effectively distinguished in a multi-image scene, the system needs to classify the input multimodal data and determine the data label corresponding to each type of data. Specifically, the system uses special characters (such as , <t>and \n) to help the model distinguish different types of data and determine their corresponding data labels.
[0063] Exemplary, 1) Conventional single image and multi-image input: use Tags are used to mark image data to help the model distinguish between images and text tokens; 2) Video input: Add between different video frames <t>Label to indicate the temporal dependency between frames; 3) High-resolution image input: Use the "\n" label to distinguish the main image from its sub-images.
[0064] It can be understood that , <t>And "\n" are the data labels corresponding to the multimodal data in the embodiment of the present application.
[0065] Step 103: determining a data modality corresponding to the multimodal data according to a data tag corresponding to the multimodal data; the data modality includes a video modality, an image modality, and a text modality;
[0066] In an embodiment of the present application, multimodal data can be classified into specific modality types based on determined data labels so that subsequent processing can perform corresponding operations on different modalities.
[0067] For example, for a single image or multiple images, the system will Labels are used to mark image data, thereby determining that the data belongs to the "image modality".
[0068] For example, for video input, the system inserts <t>tags, thereby determining that this data belongs to the "video modality" and being able to identify the temporal order between frames.
[0069] For another example, for a high-resolution image, the system will divide the image into multiple sub-images and mark the spatial position relationship between the sub-images with the "\n" tag, thereby determining that the data belongs to the "image modality" and being able to identify the spatial dependency between the sub-images.
[0070] For example, for , <t>and "\n" tags, thereby determining that these data belong to "text mode".
[0071] Step 104: unify the data modality corresponding to the multimodal data from the video modality and the image modality into the text modality, and obtain text features corresponding to the multimodal data;
[0072] In an embodiment of the present application, for data in video modality and image modality, the system converts them into feature representations in the text embedding space through a visual encoder, thereby unifying data in all modalities into text features.
[0073] By unifying the video and image modalities into a text modality, the embodiment of the present application enables the system to utilize a large language model (LLM) to uniformly process data of all modalities, thereby simplifying the processing flow of multimodal data and improving processing efficiency.
[0074] Step 105: Obtain processing results corresponding to the multimodal data according to the text features corresponding to the multimodal data and the processing instructions.
[0075] In the embodiment of the present application, the system generates the final processing result (such as generating a description, answering a question, etc.) according to the text features obtained in step 104 and the processing instructions in step 101. The final processing of multimodal data is achieved, ensuring that the system can generate accurate results based on data and processing instructions of different modalities, and improving the system's multimodal understanding and response capabilities.
[0076] In an embodiment of the present application, multimodal data and processing instructions corresponding to the multimodal data are obtained; data labels corresponding to the multimodal data are determined; data modalities corresponding to the multimodal data are determined according to data labels corresponding to the multimodal data; the data modalities include video modalities, image modalities, and text modalities; the data modalities corresponding to the multimodal data are unified from the video modality and the image modality to the text modality to obtain text features corresponding to the multimodal data; and processing results corresponding to the multimodal data are obtained according to the text features corresponding to the multimodal data and the processing instructions. By uniformly converting the video modality and the image modality into the text modality, a unified representation of data of different modalities is achieved, so that the model can process multimodal data more efficiently, reducing the complexity caused by the heterogeneity between modalities, and reducing the amount of calculation of the model when processing high-resolution images and long videos, thereby reducing the consumption of computing resources and improving computing efficiency. Moreover, with clear data labels, the model can more easily distinguish between text, images, and video frames, avoiding confusion between modalities. After the data modality is clear, it can be processed according to different modalities, optimizing computing resources and thus improving the accuracy of data processing.
[0077] In one embodiment of the present application, the data tags corresponding to the multimodal data include image tags and video tags, and determining the data modality corresponding to the multimodal data according to the data tags corresponding to the multimodal data includes:
[0078] If the data label corresponding to the multimodal data is the image label, determining that the data modality corresponding to the multimodal data is the image modality;
[0079] If the data tag corresponding to the multimodal data is the video tag, determining that the data modality corresponding to the multimodal data is the video modality;
[0080] If the data tag does not exist in the multimodal data, it is determined that the data modality corresponding to the multimodal data is the text modality.
[0081] In the embodiment of the present application, the method for processing multimodal data determines the data modality through data tags. Specifically, the data tags include image tags and video tags, and the system determines the specific modality type of the data based on these tags.
[0082] Image tags and video tags are key identifiers used to label multimodal data types. These tags help the system distinguish between different types of data (images, videos, text) and provide a basis for subsequent modality classification. For example, image tags include and "\n", the video tag includes <t>.
[0083] If the multimodal data contains image labels ( or "\n"), the system will identify it as an image modality. The data of the image modality is usually the input of a single image or multiple images. The system needs to process the spatial dependencies between images. If the multimodal data contains video tags (such as <t>), the system will identify it as a video modality. The data of the video modality is usually a sequence of multiple frames, and the system needs to handle the temporal dependency between frames. If the multimodal data does not contain any image labels ( or "\n") or video tag ( <t>), the system will recognize it as a text modality. The data of the text modality is usually plain text input, and the system needs to process the semantics and context of the text.
[0084] For example, for the following input:
[0085] text="This is a description"
[0086] image1=" Embedding representation of image 1"
[0087] image2=" Embedding representation of image 2"
[0088] Output: input_sequence = f"{text}{image1}{image2}"
[0089] At this time, since the system recognizes that {image1} and {image2} in input_sequence contain label, so {image1} and {image2} are image modal data, while {text} does not contain a label, so {text} is text modal data, that is, the input is a text and two images.
[0090] For another example, for the following input:
[0091] frame1=" Embedded representation of frame 1"
[0092] frame2=" Embedded representation of frame 2"
[0093] frame3=" Embedded representation of frame 3"
[0094] frame4=" Embedded representation of frame 4"
[0095] frame5=" Embedded representation of frame 5"
[0096] Output: input_sequence_video=f"{frame1} <t>{frame2} <t><h2 style=";text-align:left;direction:ltr">{frame3}<h2 style=";text-align:left;direction:ltr"> <t>{frame4} <t>{frame5}”
[0097] At this time, since the system recognizes that {frame} in input_sequence_video contains Tags, so {frame1}, {frame2}, {frame3}, {frame4} and {frame5} are image modal data, but {frame} is separated by <t>Label connection, so the input is a video containing five frames of images.
[0098] For another example, for the following input:
[0099] sub_image1=" Embedded representation of sub-image 1"
[0100] sub_image2=" Embedded representation of sub-image 2"
[0101] sub_image3=" Embedded representation of sub-image 3"
[0102] sub_image4=" Embedded representation of sub-image 4"
[0103] Output: input_sequence_high_res=f"{sub_image1}{sub_image2}\n{sub_image3}{sub_image4}"
[0104] At this time, since the system recognizes that {sub_image} in input_sequence_high_res contains Tags, so {sub_image1}, {sub_image2}, {sub_image3} and {sub_image4} are image modal data, but {sub_image2} and {sub_image3} are connected by the "\n" tag, so the input is a high-resolution image.
[0105] By distinguishing data of different modalities, the system can better adapt to the complexity of multimodal data and improve the performance of the model in different tasks. For example, in a multi-image scene, the system can handle the spatial dependency between images; in a video scene, the system can handle the temporal dependency between frames.
[0106] In one embodiment of the present application, unifying the data modality corresponding to the multimodal data from the video modality and the image modality into the text modality to obtain text features corresponding to the multimodal data includes:
[0107] Build visual encoder and feature mapper;
[0108] If the data modality corresponding to the multimodal data is the image modality, extracting image features in the multimodal data using the visual encoder;
[0109] The feature mapper is used to map the image features into a text space corresponding to the text modality to obtain text features corresponding to the multimodal data.
[0110] In an embodiment of the present application, a visual encoder needs to be constructed to extract visual features in an image or video modality. The visual encoder is usually based on a deep learning model and can convert an image or video frame into a high-dimensional feature representation. Specifically, the embodiment of the present application uses CLIP (Contrastive Landscape-Image Pre-training, a multimodal pre-training model) as a visual encoder to encode visual information and obtain image features.
[0111] In addition to building a visual encoder, it is also necessary to build an additional feature mapper to map the image features extracted by the visual encoder to the text space corresponding to the text modality. The feature mapper is usually a multi-layer perceptron (MLP), which can convert visual features into text embedding representations suitable for processing by a large language model (LLM). Specifically, the embodiment of the present application uses a two-layer MLP as a feature mapper to map image features to the text space and obtain mapped text features.
[0112] After data processing by the visual encoder and feature mapper, the mapped text features can be directly input into the large language model together with the data of the text modality for processing.
[0113] In the embodiment of the present application, the “input_sequence” and “input_sequence_hi gh_res” mentioned above can be understood as “multimodal data” including an image modality as the data modality.
[0114] For example, for the multimodal data "input_sequence", the system will extract the image features of {image1} and {image2}, map the extracted image features to the text embedding space through a two-layer MLP feature mapper, and obtain the mapped text features.
[0115] For another example, for the multimodal data "input_sequence_high_res", the system extracts the image features of {sub_image1} to {sub_image4} and maps them into text features through the feature mapper.
[0116] The embodiment of the present application can efficiently extract image features through a visual encoder, and can quickly map image features to a text space through a feature mapper, and can unify multimodal data into text features of a text modality. The system can directly use a large language model to process data of all modalities, avoiding the design of different processing flows for different modalities, thereby reducing the consumption of computing resources and improving processing efficiency.
[0117] In one embodiment of the present application, in order to align the visual modality (image modality and video modality) with the text modality, the feature mapper is trained using datasets such as ALLaVA-Caption and ShareGPT4V to achieve single image alignment. These datasets contain about 600K high-quality image-caption pairs.
[0118] While training the feature mapper, the parameters of the visual encoder are frozen.
[0119] In one embodiment of the present application, unifying the data modality corresponding to the multimodal data from the video modality and the image modality into the text modality to obtain text features corresponding to the multimodal data includes:
[0120] Build visual encoder and feature mapper;
[0121] If the data modality corresponding to the multimodal data is the video modality, splitting the multimodal data into a plurality of image data based on the video frame; the data modality corresponding to the image data is the image modality;
[0122] extracting image features from the multimodal data using the visual encoder;
[0123] The feature mapper is used to map the image features into a text space corresponding to the text modality to obtain text features corresponding to the multimodal data.
[0124] In the embodiment of the present application, for multimodal data whose data modality is video modality, it is necessary to split the video into multiple images according to the video frame, and then obtain the text features corresponding to the multimodal data of video modality through the same processing process as the multimodal data of image modality. Specifically, the "input_sequence_video" mentioned above can be understood as "multimodal data" whose data modality is video modality.
[0125] For example, for the multimodal data "input_sequence_video", the system splits the video into images from {frame1} to {frame5}, extracts the image features of the images through the visual encoder, and maps them into text features through the feature mapper.
[0126] The embodiments of the present application can also quickly map image features in multimodal data of the video modality to the text space through a visual encoder and a feature mapper, and can unify the multimodal data into text features of the text modality. The system can directly use a large language model to process data of all modalities, avoiding the design of different processing flows for different modalities, thereby reducing the consumption of computing resources and improving processing efficiency.
[0127] In one embodiment of the present application, the image feature has a corresponding dimension, the dimension includes a height and a width, and after extracting the image feature in the multimodal data using the visual encoder, the method further includes:
[0128] Determine a pooling window for selecting a specific area on the image feature and a pooling step size for moving the pooling window on the image feature; the pooling window includes a window height and a window width, and the pooling step size includes a step height and a step width;
[0129] Calculate the target height according to the height, the window height and the step height;
[0130] Calculate the target width according to the width, the window width and the step width;
[0131] The target height is replaced by the height corresponding to the image feature, and the target width is replaced by the width corresponding to the image feature, to obtain a target dimension corresponding to the image feature; the target dimension is smaller than the dimension.
[0132] In an embodiment of the present application, before a feature mapper is used to map image features into a text space, a 2D pooling technique is used to process image features. The purpose of 2D pooling is to select a pooling window of a specific area, move the pooling window on the image features, select the maximum value or average value within the window, thereby reducing the dimension (height and width) of the image features, so as to reduce the consumption of computing resources and improve processing efficiency.
[0133] Image features usually have two dimensions: height and width. For example, the dimension of an image feature can be expressed as (H, W), where H is the height and W is the width.
[0134] The pooling window is used to select a specific area on the image feature, and its size is determined by the window height (window_height) and window width (window_width). The pooling stride is the step length of the pooling window moving on the image feature, and its size is determined by the stride height (stride_height) and stride width (stride_width).
[0135] The system needs to calculate the target height and target width based on the dimensions (height and width) of the image features and the size and stride of the pooling window.
[0136] The target height (target_height) is calculated by the following formula:
[0137] target_height=[(H-window_height) / (stride_height)]+1
[0138] The target width (target_width) is calculated by the following formula:
[0139] target_height=[(W-window_width) / (stride_width)]+1
[0140] The calculated target height (target_height) replaces the original height (H) of the image feature, and the calculated target width (target_width) replaces the original width (W) of the image feature.
[0141] The replaced image feature dimension is (target_height, target_width), which is smaller than the original dimension (H, W).
[0142] The embodiment of the present application can significantly reduce the dimension (height and width) of the image features through 2D pooling, thereby reducing the consumption of computing resources and improving the processing efficiency of multimodal data. For example, if the dimension of the original image feature is (224, 224), after 2D pooling, the target dimension may be reduced to (14, 14), which greatly reduces the amount of calculation for subsequent processing. Moreover, during the pooling process, although the dimension of the image is reduced, the relative position relationship between different areas in the image is still retained.
[0143] In one embodiment of the present application, obtaining a processing result corresponding to the multimodal data according to the text feature corresponding to the multimodal data and the processing instruction includes:
[0144] Build a large language model;
[0145] The large language model is used to process the text features corresponding to the multimodal data and the processing instructions to obtain a processing result corresponding to the multimodal data.
[0146] In an embodiment of the present application, the system processes the text features and processing instructions corresponding to the multimodal data by constructing a large language model (LLM), thereby obtaining the processing results corresponding to the multimodal data. The large language model is a model based on deep learning that can process and generate natural language text. Specifically, the large language model is used to process the text features and processing instructions corresponding to the multimodal data to generate the final processing results.
[0147] Reference Figure 2 , is a schematic diagram of the architecture of a large language model of the present application, which combines Transformer, Mamba (Mamba-MoE-layer and Mamba-layer) and state space model (SSM). The architecture ratio of the large language model is Transformer: Mamba: SSM = 1:8:3.
[0148] Mamba-MoE-layer: Mamba is an efficient sequence model that combines a mixture of experts (MoE) mechanism that can dynamically select different experts to process different input data. The MoE mechanism enhances the expressiveness and flexibility of the model by selecting the first two experts for each token.
[0149] Mamba-layer: The Mamba layer captures dynamic changes in time through recursive formulas and hidden states, enhancing the robustness and generalization ability of the model.
[0150] SSM (State Space Model): SSM captures dynamic changes in time through recursive formulas and hidden states, enhancing the robustness and generalization ability of the model.
[0151] Transformer-layer: The Transformer layer is a standard self-attention mechanism layer that can capture global dependencies in the input sequence.
[0152] Mamba-MoE-layer uses the Mixed Experts (MoE) mechanism, which enables the model to dynamically select different experts to process different input data, enhancing the model's expressiveness and flexibility and optimizing the use of computing resources. Transformer-laye captures the global dependencies in the input sequence through the self-attention mechanism, enhancing the model's expressiveness.
[0153] Mamba-layer and SSM capture dynamic changes in time through recursive formulas and hidden states, enhancing the robustness and generalization ability of the model. By combining Transformer, Mamba and SSM in a ratio of 1:8:3, the model can not only show better robustness and generalization ability in different tasks, but also optimize computational efficiency while maintaining high performance.
[0154] In the embodiment of the present application, the input multimodal data (such as text, image, video) is first extracted through a visual encoder (such as CLIP) to extract image features, and then mapped to the text embedding space through a mapping MLP. The mapped text features are processed through Mamba-MoE-layer, Mamba-layer, SSM and Transformer-layer to generate the final processing results, such as text descriptions, answering questions or executing instructions.
[0155] In the standard Transformer model, position embeddings are used to capture position information in the input sequence, because the self-attention mechanism itself does not contain position information. However, for some tasks, especially those that already capture temporal or sequential information in other ways, position embeddings may become redundant. In addition, the state space model (SSM) is already able to capture temporal dynamic changes and sequential information well through recursive formulas and hidden states. Therefore, position embeddings are no longer necessary in the SSM layer.
[0156] Therefore, in one embodiment of the present application, RMSNorm is used between layers to enhance normalization, thereby omitting the position embedding in the Transformer.
[0157] The embodiments of the present application can reduce the number of model parameters by omitting position embedding, thereby reducing the complexity and computational cost of the model. In addition, omitting position embedding can make the implementation of the model simpler and reduce potential errors and debugging difficulties.
[0158] The large language model in the embodiment of the present application integrates Query Attention (GQA) and SwiGLU activation function, which is similar to other large language models. Exemplarily, the total number of parameters of the model is 53 billion, and the total number of activation parameters in the reasoning process is 13 billion, where 53 billion and 13 billion are thresholds. In actual application, the activation parameter synthesis and the total number of parameters can be adjusted based on actual conditions.
[0159] In one embodiment of the present application, the multimodal data includes image data whose data modality is the image modality, and the using of the large language model to process the text features and the processing instructions corresponding to the multimodal data to obtain the processing results corresponding to the multimodal data includes:
[0160] If the processing instruction is a single image instruction, the single image instruction is used to infer text features corresponding to the image data to generate a first processing result for describing the image data.
[0161] Single-image instructions are processing instructions for a single image, such as "generate an image description" or "answer a question related to the image." Single-image instructions are instructions given by the user or the system to instruct the model on how to process a single image.
[0162] For example, for input image data, the processing instruction may be "generate image description", and the large language model will generate a text describing the image content based on the text features of the image data.
[0163] Exemplary, input: an image and a single image instruction "describe this image".
[0164] Processing: The model extracts image features and aligns them with text features, and combines them with instructions for reasoning.
[0165] Output: A description of the image, e.g. "This is a landscape photo, with mountains and a lake".
[0166] In one embodiment of the present application, in order to enable the model to better follow instructions and generate accurate descriptions or answers when processing a single image, the instructions for a single image (single image instructions) may be fine-tuned.
[0167] Use a dataset containing a single image and corresponding instructions (such as LLaVA-1.5 and Manti-Single, about 932k high-quality question-answer pairs) for training. For example, the dataset may contain images and corresponding description instructions (such as "generate image description") or question-answering instructions (such as "answer questions related to the image"); during training, the model will generate corresponding outputs (such as descriptions or answers) based on the instructions for a single image, and optimize the model parameters through backpropagation to make it perform better when processing a single image; in the fine-tuning stage, the model will gradually adapt to the instructions for a single image to ensure that accurate outputs can be generated when processing a single image.
[0168] During single image instruction fine-tuning, only the visual encoder is frozen and only the feature mapper and LLM parts are trained.
[0169] Through single-image instruction fine-tuning, the model can better understand the visual features of a single image and generate accurate descriptions or answers. In addition, fine-tuning enables the model to better follow the instructions of a single image and generate output that meets the requirements of the instructions.
[0170] In one embodiment of the present application, the multimodal data also includes video data whose data modality is the video modality, and the using of the large language model to process the text features and the processing instructions corresponding to the multimodal data to obtain the processing results corresponding to the multimodal data includes:
[0171] If the processing instruction is a multi-image instruction, the multi-image instruction is used to infer text features corresponding to the video data, thereby generating a second processing result for describing the video data.
[0172] Multi-image instructions refer to processing instructions for multiple images or video frames, such as "generate video description" or "answer questions related to the video".
[0173] When the processing instruction is a multi-image instruction, the system uses the multi-image instruction to infer text features corresponding to the video data and generates a second processing result for describing the video data.
[0174] For example, for the input video data, the processing instruction may be "generate video description", and the large language model will generate a text describing the video content based on the text features of the video data.
[0175] Exemplary, input: multiple frames of a video and a multi-image instruction "summarize the content of this video".
[0176] Processing: The model extracts visual features of each frame and passes <t>Marking the temporal dependencies between frames. The model aligns visual features with textual features and combines them for reasoning.
[0177] Output: A summary of the video, e.g. "This video shows a person climbing from the foot of a mountain to the top of a mountain."
[0178] In one embodiment of the present application, multi-image instructions may also be fine-tuned to improve the model's ability to understand multiple images or video frames and its ability to follow instructions.
[0179] First, the system is trained using a dataset containing multiple images or video frames and their corresponding instructions. The visual encoder extracts the image features of the video frames and maps them to the text embedding space. Then, the model gradually adapts to the instructions of multiple images or video frames in the multi-image instruction fine-tuning stage, and adopts a progressive training strategy, gradually transitioning from single image alignment and single image instruction fine-tuning to multi-image instruction fine-tuning to ensure that the model can handle multimodal long-context scenarios. During the training process, the model infers the text features of the video data based on the multi-image instructions, generates text describing the video content, and optimizes the model parameters through back propagation to make it perform better when processing multiple images or video frames. Through multi-image instruction fine-tuning, the model can generate high-quality text descriptions, answer questions or execute instructions, which improves the user experience and the practicality of the system, while enhancing the model's multimodal long-context capabilities and processing efficiency.
[0180] In one embodiment of the present application, the multimodal data also includes text data whose data modality is the text modality, and the large language model is used to process the text features and the processing instructions corresponding to the multimodal data to obtain the processing results corresponding to the multimodal data, including:
[0181] If the processing instruction is a plain text instruction, the plain text instruction is used to infer the text features corresponding to the text data to generate a third processing result for describing the text data.
[0182] Plain text instructions refer to processing instructions for plain text data, such as "generate a text description" or "answer questions related to the text."
[0183] When the processing instruction is a plain text instruction, the system uses the plain text instruction to infer text features corresponding to the text data, and generates a third processing result for describing the text data.
[0184] For example, for input text data, the processing instruction may be "generate text description", and the large language model will generate a text describing the content of the text based on the text data.
[0185] For example, input: a long text and a text instruction "summarize the main content of this text".
[0186] Processing: The model captures long-distance dependencies in long texts through the Transformer's self-attention mechanism and combines instructions for reasoning.
[0187] Output: A summary of a long text, such as "This text discusses the application of artificial intelligence in the medical field."
[0188] In one embodiment of the present application, it is also necessary to fine-tune the instructions for the plain text data so that the model can better follow the instructions and generate accurate descriptions or answers when processing the plain text data.
[0189] First, the system is trained using a data set containing plain text data and its corresponding instructions, and the plain text data is directly input into the large language model. Then, in the plain text instruction fine-tuning stage, the model gradually adapts to the instructions of the plain text data, adopts a progressive training strategy, and gradually transitions from basic pre-training to plain text instruction fine-tuning to ensure that the model can handle plain text instructions of different lengths. During the training process, the model infers the text features of the text data based on the plain text instructions, generates text that describes the text content, and optimizes the model parameters through back propagation to make it perform better when processing plain text data. Through plain text instruction fine-tuning, the model can generate high-quality text descriptions, answer questions, or execute instructions, which improves the user experience and the practicality of the system, while enhancing the model's plain text processing capabilities and processing efficiency.
[0190] It should be noted that in the fine-tuning process of plain text instructions, the data used is a dataset of 278k plain text entries from Evol-instruct-GPT4, WildChat, and LongAlign. In the fine-tuning process of multi-image instructions, the data used is 200K, 200K, and 50K data items sampled from Mantis, VideoChat2, and ShareGPT4Video, respectively.
[0191] In order to retain the model's single image understanding and plain text conversation capabilities, the additional 200K and 50K data items from the single image instruction fine-tuning and plain text instruction fine-tuning stages were used as the Replay part. In addition, in order to improve the model's ability to interpret complex single images (divided into multiple sub-images), the team sampled 50K data from the single image instruction fine-tuning stage, performed padding and segmentation, and divided the original image into sub-images of size 336x336 as the SubImage part.
[0192] In an embodiment of the present application, LLM is transformed into a multimodal long context model through plain text instruction fine-tuning, multimodal adaptation, single image alignment, single image instruction fine-tuning and multi-image instruction fine-tuning, thereby achieving single-modality and multi-modality adaptation.
[0193] In one embodiment of the present application, the image data has a corresponding resolution, and the using the single image instruction to infer text features corresponding to the image data to generate a first processing result for describing the image data includes:
[0194] If the resolution corresponding to the image data exceeds the preset resolution, traversing from the upper left corner of the image data to the lower right corner of the image data, the image data is divided into a plurality of sub-image data;
[0195] Using the single image instruction to infer text features corresponding to the sub-image data, and generating a sub-processing result for describing the sub-image data;
[0196] The plurality of sub-processing results are combined to obtain the first processing result.
[0197] The resolution of image data refers to the pixel size of the image, usually expressed as the product of height and width. For example, the resolution of an image can be 1920x1080 or 3840x2160. The preset resolution is a resolution threshold set by the system to determine whether image data needs to be segmented.
[0198] If the resolution of the image data exceeds the preset resolution, the system will split it into multiple sub-image data. Specifically, the system will traverse from the upper left corner of the image data to the lower right corner and split the image data into multiple sub-image data. For example, for a high-resolution image, the system can split it into multiple 336x336 sub-image data.
[0199] For each sub-image data, the system uses single image instructions to infer its corresponding text features and generate a sub-processing result for describing the sub-image data. For example, for the segmented sub-image data, the system generates a corresponding text description.
[0200] The system will merge the sub-processing results of the multiple sub-image data to generate a first processing result for describing the entire image data. For example, for the multiple sub-image data after segmentation, the system will merge their text descriptions into a complete image description.
[0201] By segmenting a high-resolution image into multiple sub-image data, the system can efficiently process each sub-image data, reduce the consumption of computing resources, and improve processing efficiency. The system can process image data of different resolutions, enhancing the adaptability of the model. By segmenting and fusing sub-image data, the system can simplify the processing flow of high-resolution images, ensure that each sub-image data can be processed correctly, and avoid confusion and errors.
[0202] In an embodiment of the present application, multimodal data and processing instructions corresponding to the multimodal data are obtained; data labels corresponding to the multimodal data are determined; data modalities corresponding to the multimodal data are determined according to data labels corresponding to the multimodal data; the data modalities include video modalities, image modalities, and text modalities; the data modalities corresponding to the multimodal data are unified from the video modality and the image modality to the text modality to obtain text features corresponding to the multimodal data; and processing results corresponding to the multimodal data are obtained according to the text features corresponding to the multimodal data and the processing instructions. By uniformly converting the video modality and the image modality into the text modality, a unified representation of data of different modalities is achieved, so that the model can process multimodal data more efficiently, reducing the complexity caused by the heterogeneity between modalities, and reducing the amount of calculation of the model when processing high-resolution images and long videos, thereby reducing the consumption of computing resources and improving computing efficiency. Moreover, with clear data labels, the model can more easily distinguish between text, images, and video frames, avoiding confusion between modalities. After the data modality is clear, it can be processed according to different modalities, optimizing computing resources and thus improving the accuracy of data processing.
[0203] It should be noted that, for the method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present application are not limited by the described order of actions, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.
[0204] On the basis of the above embodiments, this embodiment further provides a multimodal data processing device, which is applied to electronic devices such as terminal devices and servers.
[0205] Reference Figure 3 , shows a structural block diagram of an embodiment of a multimodal data processing device of the present application, which may specifically include the following modules:
[0206] An acquisition module 301 is used to acquire multimodal data and processing instructions corresponding to the multimodal data;
[0207] A label determination module 302, used to determine a data label corresponding to the multimodal data;
[0208] A modality determination module 303, configured to determine a data modality corresponding to the multimodal data according to a data tag corresponding to the multimodal data; the data modality includes a video modality, an image modality, and a text modality;
[0209] A modality unification module 304 is used to unify the data modality corresponding to the multimodal data from the video modality and the image modality into the text modality to obtain text features corresponding to the multimodal data;
[0210] The data processing module 305 is used to obtain the processing result corresponding to the multimodal data according to the text features corresponding to the multimodal data and the processing instructions.
[0211] The embodiment of the present application also provides a non-volatile readable storage medium, which stores one or more modules (programs). When the one or more modules are applied to a device, the device can execute instructions (instructions) of each method step in the embodiment of the present application.
[0212] The present application embodiment provides one or more machine-readable media on which instructions are stored, and when executed by one or more processors, an electronic device executes one or more of the methods described in the above embodiments. In the present application embodiment, the electronic device includes various types of devices such as terminal devices and servers (clusters).
[0213] The embodiments of the present disclosure may be implemented as a device configured as desired using any appropriate hardware, firmware, software, or any combination thereof, and the device may include electronic devices such as terminal devices and servers (clusters). Figure 4 An exemplary apparatus 400 that may be used to implement various embodiments described in this application is schematically illustrated.
[0214] For one embodiment, Figure 4 An exemplary apparatus 400 is shown having one or more processors 402, a control module (chip set) 404 coupled to at least one of the processor(s) 402, a memory 406 coupled to the control module 404, a non-volatile memory (NVM) / storage device 408 coupled to the control module 404, one or more input / output devices 410 coupled to the control module 404, and a network interface 412 coupled to the control module 404.
[0215] The processor 402 may include one or more single-core or multi-core processors, and the processor 402 may include any combination of general-purpose processors or special-purpose processors (such as graphics processors, application processors, baseband processors, etc.). In some embodiments, the device 400 can be used as a terminal device, server (cluster), etc. described in the embodiments of the present application.
[0216] In some embodiments, the apparatus 400 may include one or more computer-readable media (e.g., memory 406 or NVM / storage device 408) having instructions 414 and one or more processors 402 combined with the one or more computer-readable media and configured to execute the instructions 414 to implement a module to perform the actions described in the present disclosure.
[0217] For one embodiment, control module 404 may include any suitable interface controller to provide any suitable interface to at least one of processor(s) 402 and / or any suitable device or component in communication with control module 404 .
[0218] The control module 404 may include a memory controller module to provide an interface to the memory 406. The memory controller module may be a hardware module, a software module, and / or a firmware module.
[0219] The memory 406 may be used, for example, to load and store data and / or instructions 414 for the device 400. For one embodiment, the memory 406 may include any suitable volatile memory, such as a suitable DRAM. In some embodiments, the memory 406 may include a double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).
[0220] For one embodiment, control module 404 may include one or more input / output controllers to provide interfaces to NVM / storage device 408 and input / output device(s) 410 .
[0221] For example, NVM / storage 408 may be used to store data and / or instructions 414. NVM / storage 408 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact disk (CD) drives, and / or one or more digital versatile disk (DVD) drives).
[0222] NVM / storage device 408 may include storage resources that are physically part of the device on which apparatus 400 is installed, or it may be accessible to the device without being part of the device. For example, NVM / storage device 408 may be accessed via input / output device(s) 410 over a network.
[0223] (One or more) input / output devices 410 may provide an interface for the apparatus 400 to communicate with any other appropriate device, and the input / output device 410 may include a communication component, an audio component, a sensor component, etc. The network interface 412 may provide an interface for the apparatus 400 to communicate through one or more networks, and the apparatus 400 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, for example, accessing a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G, 5G, etc., or a combination thereof for wireless communication.
[0224] For one embodiment, at least one of the processor(s) 402 may be packaged together with the logic of one or more controllers (e.g., a memory controller module) of the control module 404. For one embodiment, at least one of the processor(s) 402 may be packaged together with the logic of one or more controllers of the control module 404 to form a system-in-package (SiP). For one embodiment, at least one of the processor(s) 402 may be integrated on the same die with the logic of one or more controllers of the control module 404. For one embodiment, at least one of the processor(s) 402 may be integrated on the same die with the logic of one or more controllers of the control module 404 to form a system-on-chip (SoC).
[0225] In various embodiments, the device 400 may be, but is not limited to, a terminal device such as a server, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, the device 400 may have more or fewer components and / or different architectures. For example, in some embodiments, the device 400 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touch screen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.
[0226] Among them, the main control chip can be used as a processor or control module in the detection device, sensor data, location information, etc. are stored in a memory or NVM / storage device, the sensor group can be used as an input / output device, and the communication interface may include a network interface.
[0227] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0228] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0229] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable multimodal data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable multimodal data processing terminal device generate instructions for implementing the processes in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0230] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable multimodal data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0231] These computer program instructions can also be loaded onto a computer or other programmable multimodal data processing terminal device, so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0232] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.
[0233] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "including a..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0234] The above is a detailed introduction to a multimodal data processing method and device, an electronic device and a storage medium provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.< / t> < / t> < / t> < / t> < / t> < / t> < / t> < / t> < / t> < / t> < / t> < / t> < / t> < / t>
Claims
1. A method for processing multimodal data, characterized in that: The method comprises: Acquire multimodal data and processing instructions corresponding to the multimodal data; Determining data labels corresponding to the multimodal data; Determining a data modality corresponding to the multimodal data according to a data label corresponding to the multimodal data; the data modality includes a video modality, an image modality, and a text modality; Unifying the data modality corresponding to the multimodal data from the video modality and the image modality into the text modality to obtain text features corresponding to the multimodal data; According to the text features corresponding to the multimodal data and the processing instructions, a processing result corresponding to the multimodal data is obtained.
2. The method according to claim 1, characterized in that The data tags corresponding to the multimodal data include image tags and video tags, and determining the data modality corresponding to the multimodal data according to the data tags corresponding to the multimodal data includes: If the data label corresponding to the multimodal data is the image label, determining that the data modality corresponding to the multimodal data is the image modality; If the data tag corresponding to the multimodal data is the video tag, determining that the data modality corresponding to the multimodal data is the video modality; If the data tag does not exist in the multimodal data, it is determined that the data modality corresponding to the multimodal data is the text modality.
3. The method according to claim 2, characterized in that The unifying the data modality corresponding to the multimodal data from the video modality and the image modality into the text modality to obtain text features corresponding to the multimodal data includes: Build visual encoder and feature mapper; If the data modality corresponding to the multimodal data is the image modality, extracting image features in the multimodal data using the visual encoder; The feature mapper is used to map the image features into a text space corresponding to the text modality to obtain text features corresponding to the multimodal data.
4. The method according to claim 2, characterized in that: The unifying the data modality corresponding to the multimodal data from the video modality and the image modality into the text modality to obtain text features corresponding to the multimodal data includes: Build visual encoder and feature mapper; If the data modality corresponding to the multimodal data is the video modality, splitting the multimodal data into a plurality of image data based on the video frame; the data modality corresponding to the image data is the image modality; extracting image features from the multimodal data using the visual encoder; The feature mapper is used to map the image features into a text space corresponding to the text modality to obtain text features corresponding to the multimodal data.
5. The method according to claim 3 or 4, characterized in that: The image features have corresponding dimensions, the dimensions including height and width, and after extracting the image features in the multimodal data using the visual encoder, the method further includes: Determine a pooling window for selecting a specific area on the image feature and a pooling step size for moving the pooling window on the image feature; the pooling window includes a window height and a window width, and the pooling step size includes a step height and a step width; Calculate the target height according to the height, the window height and the step height; Calculate the target width according to the width, the window width and the step width; The target height is replaced by the height corresponding to the image feature, and the target width is replaced by the width corresponding to the image feature, to obtain a target dimension corresponding to the image feature; the target dimension is smaller than the dimension.
6. The method according to claim 1, characterized in that Obtaining a processing result corresponding to the multimodal data according to the text features corresponding to the multimodal data and the processing instruction includes: Build large language models; The large language model is used to process the text features corresponding to the multimodal data and the processing instructions to obtain a processing result corresponding to the multimodal data.
7. The method according to claim 6, characterized in that The multimodal data includes image data whose data modality is the image modality, and the large language model is used to process text features and the processing instructions corresponding to the multimodal data to obtain a processing result corresponding to the multimodal data, including: If the processing instruction is a single image instruction, the single image instruction is used to infer text features corresponding to the image data to generate a first processing result for describing the image data.
8. The method according to claim 6, characterized in that The multimodal data also includes video data whose data modality is the video modality, and the large language model is used to process the text features and the processing instructions corresponding to the multimodal data to obtain the processing results corresponding to the multimodal data, including: If the processing instruction is a multi-image instruction, the multi-image instruction is used to infer text features corresponding to the video data, thereby generating a second processing result for describing the video data.
9. The method according to claim 6, characterized in that The multimodal data also includes text data whose data modality is the text modality, and the large language model is used to process the text features and the processing instructions corresponding to the multimodal data to obtain the processing results corresponding to the multimodal data, including: If the processing instruction is a plain text instruction, the plain text instruction is used to infer the text features corresponding to the text data to generate a third processing result for describing the text data.
10. The method according to claim 7, characterized in that The image data has a corresponding resolution, and the single image instruction is used to infer text features corresponding to the image data to generate a first processing result for describing the image data, including: If the resolution corresponding to the image data exceeds the preset resolution, traversing from the upper left corner of the image data to the lower right corner of the image data, the image data is divided into a plurality of sub-image data; Using the single image instruction to infer text features corresponding to the sub-image data, and generating a sub-processing result for describing the sub-image data; The plurality of sub-processing results are combined to obtain the first processing result.
11. A multimodal data processing device, characterized in that: The device comprises: An acquisition module, used to acquire multimodal data and processing instructions corresponding to the multimodal data; A label determination module, used to determine the data label corresponding to the multimodal data; A modality determination module, configured to determine a data modality corresponding to the multimodal data according to a data label corresponding to the multimodal data; the data modality includes a video modality, an image modality, and a text modality; A modality unification module, used to unify the data modality corresponding to the multimodal data from the video modality and the image modality into the text modality, and obtain text features corresponding to the multimodal data; The data processing module is used to obtain the processing result corresponding to the multimodal data according to the text features corresponding to the multimodal data and the processing instructions.
12. An electronic device, characterized in that: include: processor; and A memory having executable codes stored thereon, which, when executed, enables the processor to execute the method for processing multimodal data as described in any one of claims 1-10.
13. One or more machine-readable media having executable codes stored thereon, which, when executed, enable a processor to execute the multimodal data processing method according to any one of claims 1 to 10.