Multi-modal understanding large model reasoning method, device and equipment and storage medium
By using the tokens selector to compress and align the second modal data in the multimodal understanding big model, the problem of low inference efficiency of multimodal understanding big model is solved, and the optimization of computing resources and the improvement of inference efficiency is achieved.
Patent Information
- Application Number
- CN202510176007.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-07-18
AI Technical Summary
In the inference process, the existing multimodal understanding big model has a large amount of calculation and low inference efficiency when facing a large number of tokens, especially the increase in visual markers caused by video input, resulting in high computational cost.
The second modal data is compressed through the tokens selector, the target second tokens related to the first modal data is selected, and the connector is aligned and reasoned with the first tokens is performed, and the final inference is performed using a large language model.
The computing resource requirements and tokens length in the inference process are reduced, the inference efficiency is improved, redundancy and noise are reduced, and the correlation calculation rate is improved.
Smart Images

Figure CN120338093A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of large language models. Specifically, it relates to an inference method, device, equipment, and storage medium for a multi-modal understanding large model. Background Art
[0002] With the wide application of large models in online services, the visual understanding ability of large models has gradually become a key research field, and its importance is significant. On the one hand, it expands the dimension of human-computer interaction, enabling users to interact with large models on image and video content in multiple dimensions such as style analysis and emotion interpretation, greatly enriching the interaction form. On the other hand, in many fields such as medical image diagnosis, intelligent security, and autonomous driving, the visual understanding of large models can accurately extract and analyze visual information, promoting the intelligent transformation of related fields and being a core link towards general artificial intelligence.
[0003] Existing multi-modal understanding large models usually convert data such as pictures or videos into text-like token information. For example, LLaVA-1.5 encodes an image with a resolution of 336 into 576 tokens, while LLaVA-NeXT doubles the resolution, generating 2880 tokens. For video input, Video-LLaVA needs to process more visual tokens from multiple frames, which increases the inference calculation amount significantly and reduces the inference efficiency when the multi-modal understanding large model is faced with so many tokens during the inference process. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide an inference method, device, equipment, and storage medium for a multi-modal understanding large model to improve the inference efficiency of the multi-modal understanding large model.
[0005] In a first aspect, the embodiments of this application provide an inference method for a multi-modal understanding large model. The multi-modal understanding large model includes a multi-modal encoder, a token selector, a connector, and a large language model. The method includes:
[0006] Obtain first-modal data and second-modal data;
[0007] Encode the first-modal data using the first-modal encoder in the multi-modal encoder to obtain first tokens; encode the second-modal data using the second-modal encoder in the multi-modal encoder to obtain second tokens;
[0008] Select target second tokens that the first tokens are concerned about from the second tokens through the token selector;
[0009] Align the first tokens and the target second tokens through a connector;
[0010] Perform inference on the first tokens and the aligned target second tokens through a large language model to obtain an inference result.
[0011] In the embodiments of the present application, a tokens selector is used to compress the second tokens, reducing the length of the tokens for inference and the demand for computing resources, and improving the inference efficiency.
[0012] In any embodiment, selecting the target second tokens that the first tokens are concerned about from the second tokens through a tokens selector includes:
[0013] Compress the first tokens through a tokens selector to obtain compressed first tokens;
[0014] Calculate the correlation between the compressed first tokens and the second tokens;
[0015] Determine the target second tokens from the second tokens based on the correlation.
[0016] In the embodiments of the present application, by compressing the first tokens, the computational amount of subsequent correlation calculation is reduced, so that while maintaining key information, redundancy and noise are reduced, and the rate of correlation calculation is improved.
[0017] In any embodiment, compressing the first tokens through a tokens selector to obtain compressed first tokens includes:
[0018] Perform clustering analysis on the first tokens through a tokens selector to obtain multiple cluster centers, and determine the multiple cluster centers as the compressed first tokens.
[0019] In the embodiments of the present application, through clustering analysis of the first tokens, the original first tokens can be effectively divided into multiple categories, and each category is identified by a cluster center. In this way, a large number of original first tokens can be replaced by a small number of cluster centers, realizing data compression and dimensionality reduction, not only reducing the amount of data for subsequent processing, but also reducing the computational burden of the model.
[0020] In any embodiment, calculating the correlation between the compressed first tokens and the second tokens includes:
[0021] Calculate the attention scores between each compressed first token and each second token;
[0022] Determine relevance based on attention scores.
[0023] In the embodiments of the present application, by calculating the attention scores between each compressed first token and each second token, the relevance between them can be accurately captured, thereby providing a basis for subsequently selecting the target second tokens that the first tokens are concerned about from the second tokens.
[0024] In any embodiment, calculating the attention scores between each compressed first token and each second token includes:
[0025] According to the formula Calculate the attention scores between each first token and each second token;
[0026] where: Attention ij is the attention score between the i-th compressed first token and the j-th second token; C i is the i-th compressed first token; M j is the j-th second token; d is the preset feature dimension; the value range of i is 1, 2,..., n, where n is the number of compressed first tokens; the value range of j is 1, 2,..., m, where m is the number of second tokens.
[0027] In the present application, by calculating the attention scores between each compressed first token and each second token, the relevance between them can be accurately captured, thereby providing a basis for subsequently selecting the target second tokens that the first tokens are concerned about from the second tokens.
[0028] In any embodiment, determining the target second tokens from the second tokens based on the relevance includes:
[0029] Generate an attention matrix based on the attention scores; wherein, the row vectors of the attention matrix are the compressed first tokens, and the column vectors of the attention matrix are the second tokens;
[0030] Select a preset number of elements with the maximum attention scores from each row of the attention matrix; determine the elements as the relevant elements;
[0031] Take the second tokens corresponding to the elements as the target second tokens.
[0032] In the embodiments of the present application, by selecting a preset number of elements with the largest attention scores from each row of the attention matrix, the second tokens corresponding to the selected elements have the strongest correlation with the corresponding compressed first token, which helps to screen out the most critical information and provides strong support for subsequent tasks.
[0033] In any embodiment, after obtaining the second-modal data, the method further includes:
[0034] Segment the second-modal data to obtain multiple sub-blocks;
[0035] Encode the second-modal data using the second-modal encoder in the multi-modal encoder to obtain second tokens, including:
[0036] Input the multiple sub-blocks into the second-modal encoder to obtain the second tokens output by the second-modal encoder.
[0037] In the embodiments of the present application, segmenting the second-modal data can obtain smaller and more manageable sub-blocks, increasing the flexibility of data processing and enabling the encoder to more effectively process large-scale or complex data sets; inputting the multiple sub-blocks into the second-modal encoder, the second-modal encoder can process these sub-blocks in parallel, thereby improving the encoding efficiency.
[0038] In a second aspect, the embodiments of the present application provide an inference device for a multi-modal understanding large model. The multi-modal understanding large model includes a multi-modal encoder, a tokens selector, a connector, and a large language model; the device includes:
[0039] An acquisition module for acquiring first-modal data and second-modal data;
[0040] An encoding module for encoding the first-modal data using the first-modal encoder in the multi-modal encoder to obtain first tokens; encoding the second-modal data using the second-modal encoder in the multi-modal encoder to obtain second tokens;
[0041] A compression module for selecting target second tokens that the first tokens focus on from the second tokens through the tokens selector;
[0042] An alignment module for aligning the first tokens and the target second tokens through the connector;
[0043] An inference module for inferring the first tokens and the aligned target second tokens through the large language model to obtain an inference result.
[0044] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor, a memory, and a bus, where:
[0045] The processor and the memory communicate with each other through the bus;
[0046] The memory stores program instructions executable by the processor, and the processor can execute the method of the first aspect by invoking the program instructions.
[0047] In a fourth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium, including:
[0048] The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method of the first aspect.
[0049] In a fifth aspect, an embodiment of the present application provides a computer program product, including computer program instructions, which, when read and run by a processor, execute the method of the first aspect.
[0050] Other features and advantages of the present application will be described in the subsequent specification, and, in part, will be obvious from the specification, or can be understood by implementing the embodiments of the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 It is a schematic flow chart of a reasoning method for a multi-modal understanding large model provided by an embodiment of the present application;
[0053] Figure 2 It is a schematic diagram of the reasoning principle of a multi-modal understanding large model provided by an embodiment of the present application;
[0054] Figure 3 It is an attention heat map provided by an embodiment of the present application;
[0055] Figure 4 It is a schematic diagram of a processing method of a tokens selection module provided by an embodiment of the present application;
[0056] Figure 5 It is another schematic flow chart of a reasoning method for a multi-modal understanding large model provided by an embodiment of the present application;
[0057] Figure 6 Structural schematic diagram of an inference device for a multi-modal understanding large model provided by an embodiment of the present application;
[0058] Figure 7 Entity structure schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0059] The embodiments of the technical solutions of the present application will be described in detail below with reference to the accompanying drawings. The following embodiments are only used to illustrate the technical solutions of the present application more clearly, so they are only examples and cannot be used to limit the protection scope of the present application.
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above accompanying drawings are intended to cover non-exclusive inclusion.
[0061] In the description of the embodiments of the present application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order or primary-secondary relationship of the indicated technical features. In the description of the embodiments of the present application, "a plurality of" means more than two, unless otherwise clearly and specifically defined.
[0062] Referring to "embodiments" herein means that specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0063] In the description of the embodiments of the present application, the term "and / or" is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the front and rear associated objects.
[0064] In the description of the embodiments of the present application, the term "a plurality of" refers to more than two (including two). Similarly, "multiple groups" refers to more than two groups (including two groups), and "multiple pieces" refers to more than two pieces (including two pieces).
[0065] In the description of the embodiments of the present application, unless otherwise clearly specified and limited, technical terms such as "installation", "connection", "connection", "fixation", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific situations.
[0066] A multimodal understanding large model refers to an artificial intelligence model that can process and understand multiple types of data (such as text, images, audio, video, etc.).
[0067] With the wide application of large models in online services, the visual understanding ability of large models has gradually become a key research field with remarkable importance. On the one hand, it expands the dimension of human-computer interaction, enabling users to conduct multi-dimensional information interactions with large models on image and video content, such as style analysis and emotion interpretation, greatly enriching the interaction forms. On the other hand, in many fields such as medical image diagnosis, intelligent security, and autonomous driving, the visual understanding of large models can accurately extract and analyze visual information, promoting the intelligent transformation of related fields and being a core link towards general artificial intelligence.
[0068] Existing multimodal understanding large models usually convert picture or video data into text-like token information. For example, LLaVA-1.5 encodes an image with a resolution of 336 into 576 tokens, while LLaVA-NeXT doubles the resolution, generating 2880 tokens. For video input, Video-LLaVA needs to process more visual tokens from multiple frames, making the inference cost of visual language models (VLMs) prohibitively high. Optimizing the inference process of visual language models is an urgent task to enable their application in resource-constrained real-world scenarios.
[0069] Visual and video data are highly sparse, that is, the key information is sparsely distributed in their spatial and temporal dimensions. Due to the characteristics of visual imaging and the temporality of video, this makes data storage and transmission require compression, and it is difficult to extract features during analysis and understanding, and model training is easily interfered, affecting the efficiency and effect of model inference.
[0070] Based on this, the embodiments of the present application provide a reasoning method, device, equipment, and storage medium for a multimodal understanding large model. In this method, the second-modal data is compressed through a token selector, reducing the subsequent inference calculation amount and improving the inference efficiency.
[0071] It can be understood that the embodiments of the present application are applied to an electronic device on which a multimodal understanding large model is running. The electronic device can be a desktop computer, a laptop computer, a tablet computer, a vehicle-mounted computer, a mobile phone, a smart wearable device, etc.
[0072] Figure 1 It is a schematic flow chart of an inference method for a multimodal understanding large model provided by an embodiment of the present application. As Figure 1 shown, the method includes:
[0073] Step 101: Obtain first-modal data and second-modal data.
[0074] Among them, the first-modal data and the second-modal data are often different modal data. For example, the first-modal data can be text data, and the second-modal data can be picture data; the first-modal data can be picture data, and the second-modal data can be video data; the first-modal data can be audio data, and the second-modal data can be picture data. It should be noted that the first-modal data and the second-modal data can also be a combination of other modal data, and the embodiments of the present application do not make specific limitations in this regard. In addition, both the first-modal data and the second-modal data can include one or more modal data. For example, the first-modal data includes text data; the second-modal data can include picture data and video data; another example is that the first-modal data includes text data and picture data, and the second-modal data includes video data; still another example is that the first-modal data includes text data and picture data, and the second-modal data includes audio data and video data, etc. The first-modal data and the second-modal data can be input by the user into the multimodal understanding large model. For example, a multimodal understanding large model is running on an electronic device, and the multimodal understanding large model provides an input box through which the user can input the first-modal data and the second-modal data into the multimodal understanding large model.
[0075] Step 102: Encode the first-modal data using the first-modal encoder in the multimodal encoder to obtain first tokens; encode the second-modal data using the second-modal encoder in the multimodal encoder to obtain second tokens.
[0076] Among them, the multi-modal encoder refers to an encoder that integrates multiple modalities, such as: text encoder, visual encoder, speech encoder, etc. The text encoder can be Bidirectional Encoder Representations (BERT), GPT (Generative Pre-trained Transformer) series, etc. The visual encoder can be: Vision Transformer (VIT), Variational Autoencoder (VAE), etc. The speech encoder can be an encoder of the EVRC series, an encoder of the AMR series, an encoder of the MBE series, etc.
[0077] The electronic device can select the corresponding encoder for processing according to the input first-modal data and second-modal data. For example: if the first-modal data is text data, the text encoder is selected to encode the first-modal data to obtain the first tokens; if the second-modal data is image data, the visual encoder is selected to encode the second-modal data to obtain the second tokens. If the second-modal data includes image data and audio data, then the visual encoder can be used to encode the image data, and the audio encoder can be used to encode the audio data. It should be noted that since the first tokens will be used to compress the second tokens in the subsequent process of this application, the number of the first tokens is often smaller than the number of the second tokens.
[0078] Step 103: Select the target second tokens concerned by the first tokens from the second tokens through the tokens selector.
[0079] Since the number of the second tokens is large, in order to improve the inference efficiency, it is necessary to compress the second tokens. The main function of the tokens selector is to compress the second tokens. When using the tokens selector to compress the second tokens, it is not compressed randomly, but to select the target second tokens concerned by the first tokens from the second tokens. For example: the first-modal data is text data, and its specific text is "What color is the kitten in the picture?". The second-modal data is a picture with a kitten in the middle. Then when selecting the target second tokens from the second tokens, the second tokens corresponding to the kitten in the picture are used as the target second tokens. The advantage of doing this is that on the one hand, it realizes the compression of the second tokens, reduces the subsequent calculation amount, and improves the inference efficiency; on the other hand, it can ensure the inference accuracy as much as possible.
[0080] Step 104: Align the first tokens and the target second tokens through a connector.
[0081] Data in different modalities often contain complementary information. For example, in visual question answering tasks, images provide visual information, while the text in the question provides the key points and context to be focused on. By aligning this information, the model can form a more comprehensive understanding and thus answer questions more accurately. Alignment techniques can associate this information, enabling the model to comprehensively utilize the complementary information from different modalities and thereby provide a more accurate and comprehensive understanding and prediction.
[0082] Figure 2 This is a schematic diagram of the inference of a multi-modal understanding large model provided by an embodiment of the present application, as Figure 2 shown. The multi-modal understanding large model includes a multi-modal encoder, a tokens selection module, a connector, and a large language model. In an embodiment of the present application, the first modal data is text data and the second modal data is picture data as an example for description. The first modal data obtains the first tokens after passing through the text encoder, and the second modal data obtains the second tokens after passing through the picture encoder. The first tokens and the second tokens are input into the tokens selection module. The tokens selection module compresses the second tokens based on the first tokens and the second tokens to obtain the target second tokens, that is, the compressed second tokens. Then the target second tokens are input into the connector, which is used to align the target second tokens and the first tokens to obtain the aligned second tokens. The first tokens and the aligned second tokens are sent into the large language model. In an embodiment of the present application, the large language model selected is the Transformer Decoder model. The large language model infers the input first tokens and the aligned second tokens to obtain an inference result.
[0083] In an embodiment of the present application, the main role of the connector is to align the first tokens and the target second tokens. Among them, the connector can specifically be a CLIP connector or a Q-Former connector.
[0084] An embodiment of the present application takes the Q-Former connector as an example for introduction:
[0085] The core idea of the Q-Former connector is to improve the model's representation ability and information retrieval effect by introducing a query mechanism. It uses a set of learnable query vectors to extract visual features from a frozen visual model and align and fuse them with text information. Its working principle is as follows:
[0086] 1. Input Embedding:
[0087] Convert the input data (such as text, images, etc.) into a vector representation of a fixed dimension. For text data, common methods include Word Embedding and Contextual Embedding.
[0088] 2. Query Generation:
[0089] Generate query vectors for retrieval. These query vectors are parameters learned by the model and are used to extract key information from the input data.
[0090] 3. Interaction Layer:
[0091] Implement the interaction between the query vector and the input embedding vector to generate the final output representation. This is usually calculated by the Dot-Product Attention mechanism to compute the correlation between the query vector and the input embedding vector.
[0092] In applications such as BLIP2, the connection layer of Q-Former may contain multiple Transformer sub-modules, such as a learnable Query Encoder and a Text Encoder / Decoder. These sub-modules implement the interaction between Q vectors through self-attention and cross-attention layers, as well as the interaction between Q vectors and image features (I) and text (T).
[0093] 4. Representation Learning and Generative Learning:
[0094] In the BLIP2 framework, Q-Former first performs vision-language representation learning, forcing Q-Former to learn the visual representation most relevant to the text.
[0095] Then, it performs vision-to-language generative learning, connecting the output of Q-Former to a frozen large language model, forcing the visual representation learned by Q-Former to be interpretable and utilizable by the large language model.
[0096] Step 105: Perform inference on the first tokens and the aligned target second tokens through the large language model to obtain the inference result.
[0097] Large language models refer to various large language models with reasoning capabilities, such as: large language models of the GPT series, large language models of the BERT series, ERNIE Bot, Tongyi Qianwen, DeepSeek, etc. The large language model performs reasoning on the first tokens and the aligned target second tokens to obtain the reasoning result. It can be understood that for the large language model, there is no distinction in modalities. What is input into the large language model are all tokens, and the large language model can perform corresponding reasoning based on the input tokens.
[0098] In the encoding stage of the embodiments of the present application, dedicated encoders are used to process data of different modalities respectively, which can more effectively extract the key information of their respective modalities. Moreover, the tokens selector is used to filter out the information related to the first modality data, reducing the amount of data for subsequent processing, thereby improving the overall processing efficiency.
[0099] Based on the above embodiments, the tokens selector is used to select the target second tokens concerned by the first tokens from the second tokens, including:
[0100] The tokens selector compresses the first tokens to obtain the compressed first tokens;
[0101] Calculate the correlation between the compressed first tokens and the second tokens;
[0102] Determine the target second tokens from the second tokens based on the correlation.
[0103] In the specific implementation process, in order to reduce the amount of subsequent correlation calculation, the tokens selector compresses the first tokens to obtain the compressed first tokens. It can be understood that there are various methods for compressing the first tokens. For example: compression methods based on semantics, including semantic encoding and merging; compression methods based on algorithms, including linear mapping, downsampling, etc. It is also possible to perform clustering analysis on the first tokens and use the cluster center as the compressed first tokens.
[0104] After obtaining the first compressed tokens, calculate the correlation between the first compressed tokens and the second tokens. Specifically, the correlation between the first compressed tokens and each of the second tokens can be calculated, and then the target second tokens can be determined from the second tokens based on the correlation. It is also possible to calculate the correlation between each of the first compressed tokens and each of the second tokens, and then determine the target second tokens from the second tokens based on the correlation. Specifically, the second token with a relatively high correlation with the first compressed tokens can be used as the target second token. It should be noted that the first tokens include multiple first tokens; the second tokens include second tokens.
[0105] It should be noted that the purpose of calculating the correlation between the first compressed tokens and the second tokens is to find the tokens that the first compressed tokens are concerned about from the second tokens. For the tokens that the first compressed tokens are not concerned about, the embodiments of the present application consider them redundant and remove them. Thus, compression of the second tokens is achieved. The correlation can be determined by calculating the attention score between the first compressed tokens and the second tokens, and other correlation calculation methods can also be used, such as: cosine similarity, Pearson correlation sparsity, Euclidean distance, etc.
[0106] The embodiments of the present application compress the first tokens, reducing the computational complexity of subsequent correlation calculations, so as to reduce redundancy and noise while maintaining key information and improving the rate of correlation calculation.
[0107] Based on the above embodiments, the first tokens are compressed by a tokens selector to obtain the first compressed tokens, including:
[0108] Perform clustering analysis on the first tokens through a tokens selector to obtain multiple cluster centers, and determine the multiple cluster centers as the first compressed tokens.
[0109] In the specific implementation process, the embodiments of the present application provide a method for compressing the first tokens based on clustering analysis. Each first token can be regarded as a point in a high-dimensional space, and it is analyzed through a clustering algorithm. The number of cluster centers can be preset, for example, it can be 10, 20, etc. After clustering, multiple cluster centers can be obtained, and the cluster centers are used as the first compressed tokens.
[0110] Further, the embodiments of the present application may select the K-Means algorithm for clustering. It can be understood that other clustering algorithms may also be used, and the embodiments of the present application do not make specific limitations thereto. The method flow of K-Means clustering is as follows:
[0111] I. Initialization
[0112] Select a constant K, where this K value represents the number of clusters to be finally formed.
[0113] Randomly initialize K data points as the initial cluster centers. These cluster centers can be data points randomly selected from the data set (the first tokens), or other methods (such as K-Means++) can be used for more intelligent initialization to improve the convergence speed.
[0114] II. Assign data points to clusters
[0115] For each data point (the first token) in the data set, calculate its distance from each cluster center. Usually, the Euclidean distance is used as the metric standard, but other distance metric methods, such as the Manhattan distance, can also be used.
[0116] Assign each data point to the cluster center closest to it to form K clusters.
[0117] III. Update cluster centers
[0118] Calculate the feature mean of all data points in each cluster.
[0119] Take this mean as the new cluster center.
[0120] IV. Iteration
[0121] Repeat steps II and III until the cluster centers no longer change significantly (i.e., converge), or reach the preset maximum number of iterations.
[0122] During the iteration process, after each update of the cluster center, it is necessary to recalculate the distance from each data point to the new cluster center and reassign the data points to the closest cluster center.
[0123] V. Output results
[0124] When the algorithm stops iterating, output the final cluster centers and the cluster assignment results of the data points.
[0125] It should be noted that other algorithms can also be selected for clustering, such as hierarchical clustering, etc.
[0126] In the embodiment of the present application, through the clustering analysis of the first tokens, the original first tokens can be effectively divided into multiple categories, and each category is identified by a clustering center. In this way, a large number of original first tokens can be replaced by a small number of clustering centers, realizing data compression and dimensionality reduction, which not only reduces the amount of data for subsequent processing, but also reduces the computational burden of the model.
[0127] Based on the above embodiment, calculating the correlation between the compressed first tokens and the second tokens includes:
[0128] Calculating the attention score between each compressed first token and each second token;
[0129] Determining the correlation based on the attention score.
[0130] In a specific implementation process, when calculating the correlation between the compressed first tokens and the second tokens, the attention score between each compressed first token and each second token can be calculated. Among them, the calculation method of the attention score is as follows:
[0131]
[0132] Where: Attention ij is the attention score between the i-th compressed first token and the j-th second token; C i is the i-th compressed first token; M j is the j-th second token; d is the preset feature dimension; i takes values of 1, 2,..., n, where n is the number of compressed first tokens; j takes values of 1, 2,..., m, where m is the number of second tokens.
[0133] After obtaining the attention score between each compressed first token and each second token, the attention score is the correlation between the compressed first token and the second token.
[0134] In the embodiment of the present application, by calculating the attention score between each compressed first token and each second token, the correlation between them can be accurately captured, thereby providing a basis for subsequently selecting the target second tokens that the first tokens are concerned about from the second tokens.
[0135] Based on the above embodiment, determining the target second tokens from the second tokens based on the correlation includes:
[0136] Generate an attention matrix based on attention scores; wherein, the row vectors of the attention matrix are the first tokens after compression, and the column vectors of the attention matrix are the second tokens;
[0137] Select a preset number of elements with the largest attention scores from each row of the attention matrix; determine the elements as relevant elements;
[0138] Use the second tokens corresponding to the elements as the target second tokens.
[0139] In a specific implementation process, after obtaining the attention scores between each compressed first token and each second token, an attention matrix can be generated according to the attention scores. Each row vector in the attention matrix is used to represent the attention scores between a certain compressed first token and each second token. A preset number of elements with the largest attention scores can be selected from each row vector. For example, if the preset number is 2, then two elements can be selected from each row, and these two elements are used as relevant elements, so that an attention heatmap can be obtained, as Figure 3 shown.
[0140] According to the elements selected from the attention heatmap, the corresponding second tokens are the target second tokens.
[0141] In the embodiments of the present application, by selecting a preset number of elements with the largest attention scores from each row of the attention matrix, the second tokens corresponding to the selected elements have the strongest correlation with the corresponding compressed first token, which helps to screen out the most critical information and provides strong support for subsequent tasks.
[0142] Based on the above embodiments, after obtaining the second-modal data, the method further includes:
[0143] Segment the second-modal data to obtain multiple sub-blocks;
[0144] Encode the second-modal data using the second-modal encoder in the multi-modal encoder to obtain second tokens, including:
[0145] Input the multiple sub-blocks into the second-modal encoder to obtain the second tokens output by the second-modal encoder.
[0146] In a specific implementation process, for the sake of accelerating processing, after obtaining the second-modal data, the second-modal data can be segmented first to obtain multiple sub-blocks. Among them, the segmentation method can be preset according to the actual situation. For example: it can be segmented according to the preset size of the sub-blocks, or it can be segmented according to the number of sub-blocks, etc. Figure 4Schematic diagram of a method for processing a tokens selection module provided by an embodiment of the present application, as shown in Figure 4 Figure 1. The first modality data is text data, and after encoding, the first tokens are obtained. The second modality data is an image, and before inputting the second modality data into the tokens selection module, it is segmented through a Patch Embedding module to obtain multiple sub-blocks. Then, each sub-block is encoded by an image encoder to obtain the second tokens. The first tokens and the second tokens are input into the tokens selection module to obtain the target second tokens. It should be noted that the processing method of the tokens selection module can be referred to the above embodiment, and will not be elaborated here.
[0147] In the embodiment of the present application, the second modality data is segmented, smaller and more manageable sub-blocks can be obtained, increasing the flexibility of data processing, enabling the encoder to more effectively process large-scale or complex data sets; multiple sub-blocks are input into the second modality encoder, and the second modality encoder can process these sub-blocks in parallel, thereby improving the encoding efficiency.
[0148] Figure 5 Schematic diagram of a process of an inference method for another multi-modal understanding large model provided by an embodiment of the present application, as shown in Figure 5 Figure 2. The method includes:
[0149] Step 501: Obtain the first modality data and the second modality data; for example: the first modality data can be text data, and the second modality data can be image data. It can be understood that the second modality data can also be image data and video data.
[0150] Step 502: Segment the second modality data to obtain multiple sub-blocks.
[0151] Step 503: Respectively tokenize the first modality data and multiple sub-blocks by using corresponding encoders to obtain the first tokens and the second tokens.
[0152] Step 504: Compress the second tokens by using a tokens selector; the specific compression method can be referred to the above embodiment, and will not be elaborated here.
[0153] Step 505: Align the target second tokens with the first tokens by using a connector.
[0154] Step 506: Input the first tokens and the aligned second tokens into a large language model to obtain the inference result of the large language model.
[0155] It should be noted that the order of some of the above steps can be swapped or they can be executed in parallel. For example, step 502 and step 503 can be swapped or executed in parallel.
[0156] Figure 6 This is a schematic structural diagram of an inference device for a multi-modal understanding large model provided by an embodiment of the present application. The device can be a module, a program segment, or code on an electronic device. It should be understood that the device corresponds to the above Figure 1 method embodiment and can execute Figure 1 each step involved in the method embodiment. The specific functions of the device can be referred to the description above. To avoid repetition, the detailed description is appropriately omitted here. The device includes: an acquisition module 601, an encoding module 602, a compression module 603, an alignment module 604, and an inference module 605, where:
[0157] The acquisition module 601 is used to acquire first-modal data and second-modal data;
[0158] The encoding module 602 is used to encode the first-modal data by using the first-modal encoder in the multi-modal encoder to obtain first tokens; and encode the second-modal data by using the second-modal encoder in the multi-modal encoder to obtain second tokens;
[0159] The compression module 603 is used to select target second tokens that the first tokens are concerned about from the second tokens through a tokens selector;
[0160] The alignment module 604 is used to align the first tokens and the target second tokens through a connector;
[0161] The inference module 605 is used to perform inference on the first tokens and the aligned target second tokens through a large language model to obtain an inference result.
[0162] Based on the above embodiment, the compression module 603 is specifically used for:
[0163] Compress the first tokens through the tokens selector to obtain compressed first tokens;
[0164] Calculate the correlation between the compressed first tokens and the second tokens;
[0165] Determine the target second tokens from the second tokens based on the correlation.
[0166] Based on the above embodiment, the compression module 603 is specifically used for:
[0167] Performing clustering analysis on the first tokens through the tokens selector to obtain multiple clustering centers, and determining the multiple clustering centers as the first tokens after compression.
[0168] Based on the above embodiments, the compression module 603 is specifically configured to:
[0169] Calculating the attention scores between each first token after compression and each second token;
[0170] Determining the relevance based on the attention scores.
[0171] Based on the above embodiments, the compression module 603 is specifically configured to:
[0172] According to the formula Calculating the attention scores between each first token and each second token;
[0173] Where: Attention ij Is the attention score between the i-th first token after compression and the j-th second token; C i Is the i-th first token after compression; M j Is the j-th second token; d is the preset feature dimension; the value of i is 1, 2,..., n, where n is the number of the first tokens after compression; the value of j is 1, 2,..., m, where m is the number of the second tokens.
[0174] Based on the above embodiments, the compression module 603 is specifically configured to:
[0175] Generating an attention matrix based on the attention scores; wherein, the row vectors of the attention matrix are the first tokens after compression, and the column vectors of the attention matrix are the second tokens;
[0176] Selecting a preset number of elements with the largest attention scores from each row of the attention matrix; determining the elements as the relevant elements;
[0177] Taking the second tokens corresponding to the elements as the target second tokens.
[0178] Based on the above embodiments, the device further includes a segmentation module for:
[0179] Segmenting the second modality data to obtain multiple sub-blocks;
[0180] Correspondingly, the encoding module 602 is specifically configured to:
[0181] Input the multiple sub-blocks into the second modality encoder to obtain the second tokens output by the second modality encoder.
[0182] Figure 7 Schematic diagram of the physical structure of the electronic device provided by the embodiments of the present application, as Figure 7 shown, the electronic device includes: a processor 701, a memory 702, and a bus 703; wherein:
[0183] The processor 701 and the memory 702 communicate with each other through the bus 703;
[0184] The processor 701 is configured to call program instructions in the memory 702 to execute the methods provided in the above method embodiments, for example, including: obtaining first modality data and second modality data; encoding the first modality data by using the first modality encoder in the multi-modal encoder to obtain first tokens; encoding the second modality data by using the second modality encoder in the multi-modal encoder to obtain second tokens; selecting target second tokens concerned by the first tokens from the second tokens through the tokens selector; aligning the first tokens and the target second tokens through the connector; and performing inference on the first tokens and the aligned target second tokens through the large language model to obtain an inference result.
[0185] The processor 701 may be an integrated circuit chip with signal processing capabilities. The above processor 701 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0186] The memory 702 may include, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc.
[0187] This embodiment discloses a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above method embodiments. For example, it includes: obtaining first-modal data and second-modal data; encoding the first-modal data by using the first-modal encoder in the multi-modal encoder to obtain first tokens; encoding the second-modal data by using the second-modal encoder in the multi-modal encoder to obtain second tokens; selecting target second tokens concerned by the first tokens from the second tokens through the tokens selector; aligning the first tokens and the target second tokens through a connector; and performing inference on the first tokens and the aligned target second tokens through the large language model to obtain an inference result.
[0188] This embodiment provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions. The computer instructions cause the computer to execute the methods provided in the above method embodiments. For example, it includes: obtaining first-modal data and second-modal data; encoding the first-modal data by using the first-modal encoder in the multi-modal encoder to obtain first tokens; encoding the second-modal data by using the second-modal encoder in the multi-modal encoder to obtain second tokens; selecting target second tokens concerned by the first tokens from the second tokens through the tokens selector; aligning the first tokens and the target second tokens through a connector; and performing inference on the first tokens and the aligned target second tokens through the large language model to obtain an inference result.
[0189] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical or other form.
[0190] In addition, the units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0191] Furthermore, in each embodiment of the present application, the various functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0192] In this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0193] The above description is only for the embodiments of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An inference method for a multimodal understanding large model, characterized in that, The multi-modal understanding large model includes a multi-modal encoder, a tokens selector, a connector, and a large language model; the method includes: Obtain first-modal data and second-modal data; Encode the first-modal data using the first-modal encoder in the multi-modal encoder to obtain first tokens; encode the second-modal data using the second-modal encoder in the multi-modal encoder to obtain second tokens; Select, through the tokens selector, target second tokens that the first tokens are concerned about from the second tokens; Align the first tokens and the target second tokens through the connector; Perform inference on the first tokens and the aligned target second tokens through the large language model to obtain an inference result.
2. The method according to claim 1, wherein The step of selecting, through the tokens selector, target second tokens that the first tokens are concerned about from the second tokens includes: Compress the first tokens through the tokens selector to obtain compressed first tokens; Calculate the correlation between the compressed first tokens and the second tokens; Determine target second tokens from the second tokens based on the correlation.
3. The method according to claim 2, wherein The step of compressing the first tokens through the tokens selector to obtain compressed first tokens includes: Perform clustering analysis on the first tokens through the tokens selector to obtain multiple cluster centers, and determine the multiple cluster centers as the compressed first tokens.
4. The method according to claim 2, wherein The step of calculating the correlation between the compressed first tokens and the second tokens includes: Calculate the attention scores between each compressed first token and each second token; Determine the correlation based on the attention scores.
5. The method according to claim 4, characterized in that, The step of calculating the attention scores between each compressed first token and each second token includes: According to the formula calculate the attention scores between each first token and each second token; Where: Attention ij is the attention score between the i-th first token after compression and the j-th second token; C i is the i-th first token after compression; M j is the j-th second token; d is the preset feature dimension; i takes values of 1, 2, …, n, where n is the number of the first tokens after compression; j takes values of 1, 2, …, m, where m is the number of the second tokens.
6. The method according to claim 4, wherein The step of determining target second tokens from the second tokens based on the correlation includes: Generate an attention matrix based on the attention scores; wherein, the row vectors of the attention matrix are the compressed first tokens, and the column vectors of the attention matrix are the second tokens; Select a preset number of elements with the largest attention scores from each row of the attention matrix; determine the elements as the relevant elements; Take the second tokens corresponding to the elements as the target second tokens.
7. The method according to any one of claims 1-6, characterized in that After obtaining the second-modal data, the method further includes: Segment the second-modal data to obtain multiple sub-blocks; The step of encoding the second-modal data using the second-modal encoder in the multi-modal encoder to obtain second tokens includes: Input the multiple sub-blocks into the second-modal encoder to obtain the second tokens output by the second-modal encoder.
8. An inference device for a multi-modal understanding large model, characterized in that, The multimodal understanding large model includes a multimodal encoder, a tokens selector, a connector, and a large language model; the device includes: An acquisition module, configured to acquire first-modal data and second-modal data; An encoding module, configured to encode the first-modal data by using a first-modal encoder in the multimodal encoder to obtain first tokens; and encode the second-modal data by using a second-modal encoder in the multimodal encoder to obtain second tokens; A compression module, configured to select target second tokens concerned by the first tokens from the second tokens through the tokens selector; An alignment module, configured to align the first tokens and the target second tokens through the connector; An inference module, configured to perform inference on the first tokens and the aligned target second tokens through the large language model to obtain an inference result.
9. An electronic device, characterized in that, Includes: A processor, a memory, and a bus, where: The processor and the memory communicate with each other through the bus; The memory stores program instructions that can be executed by the processor, and the processor can execute the method according to any one of claims 1-7 by invoking the program instructions.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, and when the computer instructions are run by the computer, the computer executes the method according to any one of claims 1-7.
11. A computer program product, characterized in that, Includes computer program instructions, and when the computer program instructions are read and run by a processor, the method according to any one of claims 1-7 is executed.
Citation Information
Cited By
Video semantic token compression method, video identification method and electronic equipment
CN120881297A