Inference acceleration device and method for cross-modal information processing

By setting up transformation, central processing unit, synthesis and acceleration device in the inference acceleration device, decoupling the task recognition, sparse fusion and semantic alignment layers of the multimodal large language model, the problem of inference delay of the multimodal large language model is solved, and efficient calculation and model performance balance is achieved.

CN120197713AActive Publication Date: 2025-06-24NINGBO ORIENTAL UNIV OF TECH (TEMPORARY NAME)

Patent Information

Application Number
CN202510680375.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

The existing multimodal large language model has problems of inference delay and increased memory usage, and the existing solutions are difficult to effectively balance computing efficiency and model performance.

Method used

By setting up a conversion device, a central processor, a synthesis device and an acceleration device in the inference acceleration device, the conversion device visually encodes and projects multimodal instructions, the central processor executes a multimodal large language model algorithm, and the synthesis device converts the answer code into text and/or audio. The acceleration device implicitly expresses the category of the task recognition layer, the sparse fusion layer, and the semantic alignment layer according to the input token of the computing layer, and reduces cross-modal information interaction operations.

Benefits of technology

It realizes the acceleration of the operation of multimodal large language model algorithm, reduces the inference delay time, improves the model compression efficiency, alleviates the model occupancy rate, and effectively balances the computing efficiency and model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197713A_ABST
    Figure CN120197713A_ABST
Patent Text Reader

Abstract

The invention relates to a reasoning acceleration device and method for cross-modal information processing, and the method comprises the steps: carrying out the visual coding and projection of a multi-modal instruction of a user through setting a conversion device, so as to convert the multi-modal instruction into a coding format of the multi-modal instruction; the central processing unit executes a multi-mode large language model algorithm based on the coding format to obtain answer codes of the multi-mode instructions, and the obtained answer codes are converted by the arranged synthesis device and then output. On the basis, an acceleration device is arranged to decouple a task recognition layer, a sparse fusion layer and a semantic alignment layer from all operation layers of the multi-modal large language model algorithm according to implicit expression categories of tokens input by each operation layer in the multi-modal large language model algorithm; and then reducing cross-modal information interaction operation of the task identification layer according to a predetermined ratio, inputting redundant visual tokens of the sparse fusion layer and visual tokens of the semantic alignment layer, shortening reasoning delay time, and relieving a model occupancy rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer models, and more specifically, to an inference acceleration device and method for cross-modal information processing. Background Art

[0002] Currently, multi-modal large language models (MLLMs) connect visual, audio, etc. encoders with large language models (LLMs) to achieve cross-modal processing of large language models. However, such models face significant computational efficiency bottlenecks, specifically manifested in that the number of tokens generated by visual encoders is usually several times that of text tokens, resulting in a sharp increase in inference latency and video memory occupancy. Existing solutions, such as visual token clipping and fusion techniques, are difficult to effectively balance computational efficiency and model performance due to the lack of in-depth understanding of the internal cross-modal processing mechanism of multi-modal large language models. Summary of the Invention

[0003] The technical problem to be solved by the present invention is how to overcome the technical defect of inference latency existing in existing multi-modal large language models. To overcome the above defects of the prior art, the present invention provides an inference acceleration device and method for cross-modal information processing, specifically including an inference acceleration device for cross-modal information processing and an inference acceleration method for cross-modal information processing.

[0004] An inference acceleration device for cross-modal information processing provided by the present invention includes: A conversion device, configured to perform visual encoding and projection on a user's multi-modal instruction to convert it into an encoded format of the multi-modal instruction; A central processing unit, electrically connected to the conversion device, configured to execute a multi-modal large language model algorithm to obtain an answer encoding of the multi-modal instruction; A synthesis device, electrically connected to the central processing unit, configured to convert the answer encoding obtained by the central processing unit into text and / or audio and give an output; An acceleration device, electrically connected to the central processing unit, configured to decouple a task recognition layer, a sparse fusion layer, and a semantic alignment layer from all operation layers of the multi-modal large language model algorithm according to the implicit expression category of the input tokens of each operation layer in the multi-modal large language model algorithm, and then reduce the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer by a predetermined ratio to accelerate the operation of the multi-modal large language model algorithm.

[0005] The inference acceleration device for cross-modal information processing disclosed by the present invention can visually encode and project the user's multi-modal instructions through a conversion device to convert them into an encoding format of the multi-modal instructions. This encoding format is then received by the provided central processing unit, which executes the multi-modal large language model algorithm to obtain the answer encoding of the multi-modal instructions. The obtained answer encoding is then converted into text and / or audio by the provided synthesis device and output. On this basis, in order to accelerate the operation of the multi-modal large language model algorithm, an acceleration device is provided to decouple the task recognition layer, sparse fusion layer, and semantic alignment layer from all operation layers of the multi-modal large language model algorithm according to the implicit expression categories of the input tokens of each operation layer in the multi-modal large language model algorithm. Then, the cross-modal information interaction operations of the task recognition layer, redundant visual tokens of the input sparse fusion layer, and visual tokens of the input semantic alignment layer are reduced according to a predetermined ratio to accelerate the operation of the multi-modal large language model algorithm, reduce the inference latency time, decouple the model inference process of the multi-modal large language model algorithm into a task recognition layer, sparse fusion layer, and semantic alignment layer, provide universal support for the model optimization of the multi-modal large language model algorithm, improve the model compression efficiency, and then alleviate the model occupancy rate, effectively balancing the computing efficiency and model performance.

[0006] In a possible implementation manner, the multi-modal instructions include prompt text, text instructions, and an associated image bundled with the text instructions.

[0007] In a possible implementation manner, the conversion device includes: An image encoding module, electrically connected to the central processing unit, and configured to visually encode and project the associated image to convert it into an image encoding format adapted to the input of the multi-modal large language model algorithm; A tokenizer, electrically connected to the central processing unit, and configured to sequentially perform text segmentation and encoding mapping on the text instructions and the prompt text to convert them into a text encoding format adapted to the input of the multi-modal large language model algorithm; Thus, through the parallel processing of the image encoding module and the tokenizer, it is possible to convert the text instructions, associated image, and prompt text into a format adapted to the input format of the multi-modal large language model algorithm, further improving the operation efficiency of the multi-modal large language model algorithm and reducing the processing and inference latency time.

[0008] In a possible implementation manner, the image encoding module includes: An image encoder, configured to perform visual encoding of the associated image to obtain image encoding data; A projection device, electrically connected to the image encoder and the central processing unit at the same time, is configured to perform projection conversion on the image encoding data to obtain an image encoding format adapted to the input of the multi-modal large language model algorithm; In this solution, the associated image is encoded by the image encoder, and under the processing of the projection device, an image encoding format adapted to the input of the multi-modal large language model algorithm can be obtained, which can not only extract the key information of the image, but also further reduce the inference latency of the multi-modal large language model algorithm.

[0009] In a possible implementation, the acceleration device includes: A decoupling module, electrically connected to the central processing unit, is configured to decouple a task recognition layer, a sparse fusion layer, and a semantic alignment layer from all operation layers of the multi-modal large language model algorithm according to the implicit expression categories of the input tokens of each operation layer in the multi-modal large language model algorithm when the multi-modal large language model algorithm is initially executed; An execution module, electrically connected to the decoupling module, is configured to reduce the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer by a predetermined ratio when the multi-modal large language model algorithm is executed, so as to accelerate the operation of the multi-modal large language model algorithm; In this solution, by setting the decoupling module to decouple a task recognition layer, a sparse fusion layer, and a semantic alignment layer from all operation layers of the multi-modal large language model algorithm when the multi-modal large language model algorithm is initially executed, and by setting the execution module to reduce the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer by a predetermined ratio when the multi-modal large language model algorithm is executed for the second time and later, the acceleration of the operation of the multi-modal large language model algorithm is realized, the operation efficiency of the central processing unit for obtaining answers is improved, and the computational efficiency and model performance are effectively balanced.

[0010] In a possible implementation, the decoupling module is configured to perform the following steps: A1: Starting from the first operation layer in the multi-modal large language model algorithm, sequentially perform category determination on the implicit expressions of the input tokens of each operation layer based on the task recognition determination criterion to extract all the task recognition layers of the multi-modal large language model algorithm; A2: Sequentially perform category determination on the implicit expressions of the input tokens of each remaining operation layer in the multi-modal large language model algorithm based on the sparse fusion determination criterion to extract all the sparse fusion layers of the multi-modal large language model algorithm; A3: Based on the semantic alignment determination criteria, classify the implicit expressions of the input tokens of each remaining operation layer in the multi-modal large language model algorithm at this time to extract all the semantic alignment layers of the multi-modal large language model algorithm; This solution can ensure the decoupling of the task extraction layer, the sparse fusion layer, and the semantic alignment layer, improve the generalization ability, provide universal support for the model optimization of the multi-modal large language model algorithm, and further enhance the model compression efficiency.

[0011] In a possible implementation manner, the step A1 includes the following steps: A11: According to the running order of the multi-modal large language model algorithm, take the first operation layer in the multi-modal large language model algorithm as the current layer; A12: Project the implicit expression of the last input token of the current layer into the semantic space through the S-shaped activation function to obtain the semantic expression; A13: Determine whether the current layer is to summarize the text instructions input by the user according to the semantic expression, If so, execute the next step; If not, execute step A16; A14: Sequentially perform visual attention merging and attention mask extraction on the implicit expressions of multiple input tokens of the current layer to obtain the attention mask; A15: Determine whether cross-modal information fusion occurs according to the attention mask, If so, execute the next step; If not, take the next operation layer in the multi-modal large language model algorithm as the current layer, and then turn back to execute the step A12; A16: According to the running order of the multi-modal large language model algorithm, take all the operation layers from the first operation layer to the current layer in the multi-modal large language model algorithm as the task recognition layer, generate a fixed ratio for reducing the cross-modal information interaction operation of the task recognition layer, and transmit this fixed ratio to the execution module; This solution confirms the task recognition layer according to cross-modal information interaction and semantic expression, and generates a fixed ratio for reducing the cross-modal information interaction operation of the task recognition layer to ensure that the execution module removes an appropriate amount of redundant calculations when running in the task recognition layer, further improving the inference operation efficiency.

[0012] In a possible implementation manner, the step A2 includes the following steps: A21: According to the running order of the multi-modal large language model algorithm, take the next operation layer after the last task recognition layer in the multi-modal large language model algorithm as the current judgment layer; A22: Perform an attention masking operation on the implicit representations of the input tokens in the current judgment layer respectively to extract their respective attention masks; A23: Calculate the cosine similarity between the implicit representation of each input token and the attention mask respectively through the cosine similarity calculation formula to obtain their respective cosine similarities; A24: Calculate the distance between the implicit representation of each input token and the attention mask respectively through the distance calculation formula to obtain their respective mask distances; A25: Extract the implicit representations of all input tokens whose cosine similarity is greater than the first threshold and the mask distance is less than the second threshold, and record the numbers of all input units in the current judgment layer with these implicit representations of the input tokens as inputs, so as to use the implicit representations of all input tokens received by these input units hereafter as the redundant visual tokens of the current judgment layer; A26: Map the implicit representations of all input tokens in the current judgment layer through a feature function, so that the implicit representation of the input token corresponding to the image coding format is zero, and the rest remains unchanged, to obtain a feature mapping result; A27: Use the feature mapping result as the input of the current judgment layer, and judge whether the degree of decline in the model performance of the multi-modal large language model algorithm is not less than a specified value, If so, execute the next step; If not, take the next operation layer as the current judgment layer, and then turn back to execute the step A22; A28: Take all operation layers from the next operation layer of the last task recognition layer in the multi-modal large language model algorithm to the current judgment layer as the sparse fusion layer, generate a fixed ratio for reducing the redundant visual tokens input to the sparse fusion layer, and transmit this fixed ratio to the execution module; This solution decouples the sparse fusion layer, and at the same time, through the proposed dual-metric dynamic mask evaluation method (cosine similarity and distance) to quantify the perturbation degree of the output caused by the occlusion of a single visual token, overcomes the misjudgment problem caused by attention aggregation in the traditional attention weight method.

[0013] In a possible implementation manner, the step A3 includes the following steps: A31: According to the running order of the multi-modal large language model algorithm, take the next operation layer of the last sparse fusion layer in the multi-modal large language model algorithm as the current detection layer; A32: Project the implicit representation of the last input token in the current detection layer into the semantic space through the S-type activation function to obtain a projection expression; A33: Judge whether the current detection layer is to perform semantic alignment of input and output tokens according to the projection expression, If so, use the next operation layer in the multi-modal large language model algorithm as the current detection layer, and then loop back to execute step A32; If not, proceed to the next step; A34: According to the running order of the multi-modal large language model algorithm, use all operation layers from the next operation layer of the last sparse fusion layer to the current detection layer in the multi-modal large language model algorithm as the semantic alignment layer, generate a fixed ratio for reducing visual tokens input to the semantic alignment layer, and pass this fixed ratio to the execution module.

[0014] While decoupling the semantic alignment layer, the above solution further reduces the computational amount and cache occupancy during the model inference process of the multi-modal large language model algorithm by generating a fixed ratio for reducing visual tokens input to the semantic alignment layer, enabling the execution module to remove visual tokens.

[0015] Another technical solution of the present invention is to provide a method for accelerating inference of cross-modal information processing, including the following steps: S1: Use an acceleration device to decouple a task recognition layer, a sparse fusion layer, and a semantic alignment layer from all operation layers of the multi-modal large language model algorithm according to the implicit expression categories of input tokens of each operation layer executed by the central processing unit; S2: Prompt the user to input a multi-modal instruction, and use a conversion device to perform visual encoding and projection on the user's multi-modal instruction to obtain the encoded format of the multi-modal instruction; S3: Use the central processing unit to execute the multi-modal large language model algorithm to obtain the answer encoding of the multi-modal instruction, and when executing the multi-modal large language model algorithm, use the acceleration device to reduce the cross-modal information interaction operations of the task recognition layer, redundant visual tokens input to the sparse fusion layer, and visual tokens input to the semantic alignment layer by a predetermined ratio; S4: Use a synthesis device to convert the answer encoding obtained by the central processing unit into text and / or audio and give an output.

[0016] The method disclosed in this application first decouples the task recognition layer, the sparse fusion layer, and the semantic alignment layer from all the operation layers of the multi-modal large language model algorithm through an acceleration device according to the implicit expression categories of the input tokens of each operation layer in the multi-modal large language model algorithm. Subsequently, through a conversion device, the multi-modal instructions of the user are visually encoded and projected to be converted into the encoded format of the multi-modal instructions, and this encoded format is received by the central processing unit. The central processing unit executes the multi-modal large language model algorithm to obtain the answer encoding of the multi-modal instructions, and the obtained answer encoding is converted into text and / or audio through a synthesis device and output. When executing the multi-modal large language model algorithm, the acceleration device reduces the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens of the input sparse fusion layer, and the visual tokens of the input semantic alignment layer according to a predetermined ratio to accelerate the operation of the multi-modal large language model algorithm, reduce the inference delay time, and finally decouple the model inference process of the multi-modal large language model algorithm into the task recognition layer, the sparse fusion layer, and the semantic alignment layer, providing universal support for the model optimization of the multi-modal large language model algorithm, improving the model compression efficiency, thereby alleviating the model occupancy rate, and effectively balancing the computing efficiency and the model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 FIG. is a schematic structural diagram of an inference acceleration device for cross-modal information processing disclosed in an embodiment of this application; Figure 2 FIG. is a flowchart of the operation of the decoupling module disclosed in an embodiment of this application; Figure 3 FIG. is a flowchart corresponding to step A1 disclosed in an embodiment of this application; Figure 4 FIG. is a flowchart corresponding to step A2 disclosed in an embodiment of this application; Figure 5 FIG. is a flowchart corresponding to step A3 disclosed in an embodiment of this application; Figure 6 FIG. is a flowchart of the method disclosed in an embodiment of this application; Figure 7 FIG. is an execution effect diagram of the multi-modal large language model algorithm accelerated by the acceleration device disclosed in an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] First of all, those skilled in the art should understand that these embodiments are only used to explain the technical principles of the embodiments of this application and are not intended to limit the protection scope of the embodiments of this application. Those skilled in the art can adjust them as needed to adapt to specific application scenarios.

[0019] In the embodiments of the present application, unless otherwise clearly specified or limited, the electrical connection between the first feature and the second feature means that there is a transmission of electrical signals between the first feature and the second feature, that is, there is an electrical relationship, and the ways to achieve the transmission of electrical signals can be wire electrical connection, radio connection, electrical connection of electromagnetic media (such as semiconductors), communication implemented by channels, etc.

[0020] In the embodiments of the present application, unless otherwise clearly specified or limited, the distance between two objects should be understood according to the mathematical meaning of distance, that is, the so-called object and object distance refers to the mapping from any two objects in the set under discussion (such as the implicit expression of the input token and the attention mask in this embodiment) to the set of real numbers, and this mapping satisfies the following three requirements for any object , object and object simultaneously; (1) , and the equal sign holds if and only if ; (2) ; (3) . Specifically, it can be the Euclidean distance, or the distance induced by the inner product or norm. Those skilled in the art can adjust according to specific needs.

[0021] In the embodiments of the present application, unless otherwise clearly specified or limited, the first feature being "on" or "under" the second feature can be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature can be that the first feature is directly above or obliquely above the second feature, or merely indicates that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "beneath" and "underneath" the second feature can be that the first feature is directly below or obliquely below the second feature, or merely indicates that the horizontal height of the first feature is less than that of the second feature.

[0022] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] Referring to Figures 1 to 7 , an inference acceleration device for cross-modal information processing is disclosed in the embodiments of the present application. The structural schematic diagram of the inference acceleration device is as Figure 1 shown. The inference acceleration device includes a conversion device, a central processing unit, a synthesis device and an acceleration device. Among them, the central processing unit is electrically connected to the conversion device, the synthesis device is electrically connected to the central processing unit, and the acceleration device is electrically connected to the central processing unit.

[0024] In the inference acceleration device, the conversion device is configured to perform visual encoding and projection on the user's multimodal instructions to convert them into an encoded format of the multimodal instructions, where the multimodal instructions include prompt text, text instructions, and associated images bundled with the text instructions.

[0025] As Figure 1 shown, in this embodiment, the conversion device includes an image encoding module and a tokenizer, the tokenizer is electrically connected to the central processing unit, the image encoding module includes an image encoder and a projection device, and the projection device is electrically connected to both the image encoder and the central processing unit.

[0026] In the conversion device, the image encoding module is configured to perform visual encoding and projection on the associated image to convert it into an image encoding format adapted to the input of the multimodal large language model algorithm. Specifically, first, the image encoder in the image encoding module performs visual encoding of the associated image to obtain image encoding data. Second, the projection device in the image encoding module performs projection conversion on the image encoding data to obtain an image encoding format adapted to the input of the multimodal large language model algorithm. The tokenizer is configured to perform text segmentation and encoding mapping on the text instructions and prompt text in sequence to convert them into a text encoding format adapted to the input of the multimodal large language model algorithm. The image encoding format and the text encoding format are the set of input tokens of the multimodal large language model algorithm.

[0027] Please continue to refer to Figure 1 , in the inference acceleration device, the central processing unit is configured to execute the multimodal large language model algorithm to obtain the answer encoding of the multimodal instructions; the synthesis device is configured to convert the answer encoding obtained by the central processing unit into text and / or audio and give an output, and the conversion method is the prior art and will not be elaborated here.

[0028] Please continue to refer to Figure 1 , in the inference acceleration device, the acceleration device is configured to decouple the task recognition layer, the sparse fusion layer, and the semantic alignment layer from all the operation layers of the multimodal large language model algorithm according to the implicit expression categories of the input tokens of each operation layer in the multimodal large language model algorithm, and then reduce the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer by a predetermined ratio to accelerate the operation of the multimodal large language model algorithm. In this embodiment, the acceleration device includes a decoupling module and an execution module, the decoupling module is electrically connected to the central processing unit, and the execution module is electrically connected to the decoupling module.

[0029] Refer to Figure 2, in the acceleration device, the decoupling module is configured to, when initially executing the multi-modal large language model algorithm, decouple the task recognition layer, the sparse fusion layer, and the semantic alignment layer from all the operation layers of the multi-modal large language model algorithm according to the implicit expression categories of the input tokens of each operation layer in the multi-modal large language model algorithm.

[0030] As Figure 2 shown, in this embodiment, the decoupling module is configured to perform the following steps: A1: Starting from the first operation layer of the multi-modal large language model algorithm, sequentially perform category determination on the implicit expressions of the input tokens of each operation layer based on the task recognition determination criteria to extract all the task recognition layers of the multi-modal large language model algorithm.

[0031] Refer to Figure 3 , specifically, in this embodiment, the operation layers in the multi-modal large language model algorithm are divided according to the reception of the implicit expressions of the input tokens. The neuron group that can receive the implicit expression of one input token in the operation layer is regarded as an input unit, and the input units are numbered according to the reception time sequence.

[0032] In this embodiment, step A1 includes the following steps: A11: According to the running order of the multi-modal large language model algorithm, take the first operation layer of the multi-modal large language model algorithm as the current layer.

[0033] A12: Project the implicit expression of the last input token of the current layer into the semantic space through the S-shaped activation function to obtain the semantic expression. The S-shaped activation function in this embodiment is the Softmax activation function. The implicit expression of the input token received by the input unit with the last number of each operation layer is regarded as the implicit expression of the last input token.

[0034] A13: Judge whether the current layer is to summarize the task of the user input text instruction according to the semantic expression (which can be compared with the mapping result of the S-shaped activation function that standardly summarizes the task of the user input text instruction), if so, execute the next step; if not, execute step A16.

[0035] A14: Sequentially perform visual attention merging and attention mask extraction on the implicit expressions of the multiple input tokens of the current layer to obtain the attention mask.

[0036] A15: Judge whether cross-modal information fusion occurs according to the attention mask, if so, execute the next step; if not, take the next operation layer of the multi-modal large language model algorithm as the current layer, and then turn back to execute step A12.

[0037] A16: According to the running order of the multi-modal large language model algorithm, all operation layers from the first operation layer to the current layer in the multi-modal large language model algorithm are used as the task recognition layer, and a fixed ratio for reducing the cross-modal information interaction operations of the task recognition layer is generated. The fixed ratio in this embodiment is 1, and this fixed ratio is transmitted to the execution module.

[0038] A2: Based on the sparse fusion determination criterion, the implicit expressions of the input tokens of each remaining operation layer in the multi-modal large language model algorithm are sequentially subjected to category determination to extract all sparse fusion layers of the multi-modal large language model algorithm.

[0039] See Figure 4 , in this embodiment, step A2 includes the following steps: A21: According to the running order of the multi-modal large language model algorithm, the next operation layer after the last task recognition layer in the multi-modal large language model algorithm is used as the current judgment layer.

[0040] A22: Perform attention masking operations on the implicit expressions of the respective input tokens of the current judgment layer to extract their respective attention masks.

[0041] A23: Calculate the cosine similarity between the implicit expression of each input token and the attention mask respectively through the cosine similarity calculation formula to obtain their respective cosine similarities.

[0042] In this embodiment, the cosine similarity calculation formula is as follows: , In the formula, represents the implicit expression (vector) of the th input token; represents the attention mask of the th input token; represents the cosine similarity.

[0043] A24: Calculate the distance between the implicit expression of each input token and the attention mask respectively through the distance calculation formula to obtain their respective mask distances.

[0044] In this embodiment, the distance calculation formula is as follows: , In the formula, represents the distance induced by the 2-norm.

[0045] A25: Extract the implicit representations of all input tokens whose cosine similarity is greater than the first threshold and whose mask distance is less than the second threshold, record the numbers of all input units in the current judgment layer with the implicit representations of these input tokens as inputs, and use the implicit representations of all input tokens received by these input units hereafter as redundant visual tokens in the current judgment layer. In this embodiment, the first threshold is 0.999 and the second threshold is 0.2.

[0046] A26: Map the implicit representations of all input tokens in the current judgment layer through a feature function to set the implicit representations of the input tokens corresponding to the image encoding format to zero and keep the rest unchanged, obtaining a feature mapping result.

[0047] A27: Use the feature mapping result as the input of the current judgment layer and determine whether the degree of decline in the model performance of the multimodal large language model algorithm is not less than a specified value. If so, execute the next step; If not, according to the running order of the multimodal large language model algorithm, take the next operation layer as the current judgment layer, and then loop back to execute step A22.

[0048] A28: According to the running order of the multimodal large language model algorithm, take all operation layers from the next operation layer of the last task recognition layer in the multimodal large language model algorithm to the current judgment layer as the sparse fusion layer, generate a fixed ratio for reducing the redundant visual tokens in the input sparse fusion layer. In this embodiment, the set fixed ratio is 1, and this fixed ratio is passed to the execution module.

[0049] A3: Based on the semantic alignment determination criterion, sequentially perform category determination on the implicit representations of the input tokens of each remaining operation layer in the multimodal large language model algorithm at this time to extract all semantic alignment layers of the multimodal large language model algorithm.

[0050] See Figure 5 , in this embodiment, step A3 includes the following steps: A31: According to the running order of the multimodal large language model algorithm, take the next operation layer of the last sparse fusion layer in the multimodal large language model algorithm as the current detection layer.

[0051] A32: Project the implicit representation of the last input token in the current detection layer into the semantic space through the S-shaped activation function to obtain a projection expression.

[0052] A33: Determine whether the current detection layer performs semantic alignment of input and output tokens according to the projection expression. If so, take the next operation layer in the multimodal large language model algorithm as the current detection layer, and then loop back to execute the above step A32; If not, execute the next step.

[0053] A34: According to the running order of the multi-modal large language model algorithm, all operation layers from the next operation layer of the last sparse fusion layer to the current detection layer in the multi-modal large language model algorithm are used as the semantic alignment layer, and a fixed ratio for reducing the visual tokens of the input semantic alignment layer is generated. In this embodiment, the set fixed ratio is 1, and this fixed ratio is passed to the execution module.

[0054] In the acceleration device, the execution module is set to reduce the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens of the input sparse fusion layer, and the visual tokens of the input semantic alignment layer by a predetermined ratio when executing the multi-modal large language model algorithm, so as to accelerate the running of the multi-modal large language model algorithm.

[0055] See Figure 6 , the following will further disclose the usage method of the inference acceleration device for cross-modal information processing described in this embodiment. This method includes the following steps: S1: The acceleration device decouples the task recognition layer, the sparse fusion layer, and the semantic alignment layer from all operation layers of the multi-modal large language model algorithm according to the implicit expression categories of the input tokens of each operation layer executed by the central processing unit.

[0056] S2: Let the user input a multi-modal instruction to perform visual encoding and projection on the user's multi-modal instruction through the conversion device to obtain the encoded format of the multi-modal instruction.

[0057] S3: The central processing unit executes the multi-modal large language model algorithm to obtain the answer encoding of the multi-modal instruction. When executing the multi-modal large language model algorithm, the acceleration device reduces the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens of the input sparse fusion layer, and the visual tokens of the input semantic alignment layer by a predetermined ratio.

[0058] S4: The synthesis device converts the answer encoding obtained by the central processing unit into text and / or audio and gives the output.

[0059] See Figure 7, the technical effects of the inference acceleration device for cross-modal information processing described in this embodiment will be described in detail below. By applying the acceleration device in this embodiment to a central processing unit that separately executes three multi-modal large language model algorithms (parameter quantity 3B - 34B), namely LLaVA-1.5, mobileVLM, and Intern-VL, systematic tests are conducted covering benchmark tasks such as visual question answering (VQA-v2, GQA), fine-grained recognition (TextVQA, ST-VQA), and visual reasoning (ScienceQA). The experimental results show that the acceleration device in this embodiment maintains a performance greater than 99% while reducing the floating-point operation amount by 54.6% and the key-value cache by 98.9%. Figure 7 The running simulation effect of a question and answer is shown in Figure 7 . In the figure, the task recognition layer is abbreviated as the shallow layer, the sparse fusion layer is abbreviated as the middle layer, and the semantic alignment layer and the remaining operation layer are abbreviated as the deep layer.

[0060] For the inference acceleration device for cross-modal information processing disclosed in this embodiment, by setting a conversion device, the user's multi-modal instructions can be visually encoded and projected to be converted into the encoding format of multi-modal instructions. This encoding format will be received by the set central processing unit, and the central processing unit will execute the multi-modal large language model algorithm to obtain the answer encoding of the multi-modal instructions. The obtained answer encoding will be converted into text and / or audio by the set synthesis device and output. On this basis, by setting an acceleration device, the task recognition layer, the sparse fusion layer, and the semantic alignment layer are decoupled from all operation layers of the multi-modal large language model algorithm according to the implicit expression categories of the input tokens of each operation layer in the multi-modal large language model algorithm. Then, the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer are reduced according to a predetermined ratio to accelerate the operation of the multi-modal large language model algorithm and reduce the inference delay time. The model inference process of the multi-modal large language model algorithm can be decoupled into the task recognition layer, the sparse fusion layer, and the semantic alignment layer, providing universal support for the model optimization of the multi-modal large language model algorithm, improving the model compression efficiency, and further alleviating the model occupancy rate, effectively balancing the computing efficiency and the model performance.

[0061] In the description of the embodiments of the present application, it should be noted that in the description of the present application, terms such as "inner" and "outer" indicating the direction or positional relationship are based on the direction or positional relationship shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of the present application.

[0062] In the description of the present application, the description referring to terms such as "one embodiment", "some embodiments", "in this embodiment", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0063] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An inference acceleration device for cross-modal information processing, characterized in that Including: A conversion device, configured to visually encode and project a user's multimodal instruction to convert it into an encoded format of the multimodal instruction; A central processing unit, electrically connected to the conversion device, configured to execute a multimodal large language model algorithm to obtain an answer encoding of the multimodal instruction; A synthesis device, electrically connected to the central processing unit, configured to convert the answer encoding obtained by the central processing unit into text and / or audio and give an output; An acceleration device, electrically connected to the central processing unit, configured to decouple a task recognition layer, a sparse fusion layer, and a semantic alignment layer from all operation layers of the multimodal large language model algorithm according to the implicit expression category of the input tokens of each operation layer in the multimodal large language model algorithm, and then reduce the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer by a predetermined ratio to accelerate the operation of the multimodal large language model algorithm.

2. The inference acceleration device for cross-modal information processing according to claim 1, wherein The multimodal instruction includes a prompt text, a text instruction, and an associated image bundled with the text instruction.

3. The inference acceleration device for cross-modal information processing according to claim 2, wherein The conversion device includes: An image encoding module, electrically connected to the central processing unit, configured to visually encode and project the associated image to convert it into an image encoding format adapted to the input of the multimodal large language model algorithm; A tokenizer, electrically connected to the central processing unit, configured to sequentially perform text segmentation and encoding mapping on the text instruction and the prompt text to convert them into a text encoding format adapted to the input of the multimodal large language model algorithm.

4. The inference acceleration device for cross-modal information processing according to claim 3, wherein The image encoding module includes: An image encoder, configured to visually encode the associated image to obtain image encoding data; A projection device, electrically connected to both the image encoder and the central processing unit, configured to perform projection conversion on the image encoding data to obtain an image encoding format adapted to the input of the multimodal large language model algorithm.

5. The inference acceleration device for cross-modal information processing according to claim 3 or 4, characterized in that The acceleration device includes: A decoupling module, electrically connected to the central processing unit, configured to, when initially executing the multimodal large language model algorithm, decouple a task recognition layer, a sparse fusion layer, and a semantic alignment layer from all operation layers of the multimodal large language model algorithm according to the implicit expression category of the input tokens of each operation layer; An execution module, electrically connected to the decoupling module, configured to, when executing the multimodal large language model algorithm, reduce the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer by a predetermined ratio to accelerate the operation of the multimodal large language model algorithm.

6. The inference acceleration device for cross-modal information processing according to claim 5, wherein The decoupling module is configured to perform the following steps: A1: Starting from the first operation layer of the multimodal large language model algorithm, sequentially perform category determination on the implicit expression of the input tokens of each operation layer based on a task recognition determination criterion to extract all the task recognition layers of the multimodal large language model algorithm; A2: Classify the implicit expressions of the input tokens of each remaining operation layer in the multi-modal large language model algorithm in sequence based on the sparse fusion determination criterion, so as to extract all the sparse fusion layers of the multi-modal large language model algorithm; A3: Classify the implicit expressions of the input tokens of each remaining operation layer in the multi-modal large language model algorithm at this time in sequence based on the semantic alignment determination criterion, so as to extract all the semantic alignment layers of the multi-modal large language model algorithm.

7. The inference acceleration device for cross-modal information processing according to claim 6, wherein The step A1 includes the following steps: A11: According to the running order of the multi-modal large language model algorithm, take the first operation layer in the multi-modal large language model algorithm as the current layer; A12: Project the implicit expression of the last input token of the current layer into the semantic space through the S-shaped activation function to obtain the semantic expression; A13: Judge whether the current layer is to summarize the task of the user input text instruction according to the semantic expression, If so, execute the next step; If not, execute step A16; A14: Sequentially perform visual attention merging and attention mask extraction on the implicit expressions of multiple input tokens of the current layer to obtain the attention mask; A15: Judge whether cross-modal information fusion occurs according to the attention mask, If so, execute the next step; If not, take the next operation layer in the multi-modal large language model algorithm as the current layer, and then turn back to execute the step A12; A16: According to the running order of the multi-modal large language model algorithm, take all the operation layers from the first operation layer to the current layer in the multi-modal large language model algorithm as the task recognition layer, generate a fixed ratio for reducing the cross-modal information interaction operation of the task recognition layer, and transmit this fixed ratio to the execution module.

8. The inference acceleration device for cross-modal information processing according to claim 7, characterized in that, The step A2 includes the following steps: A21: According to the running order of the multi-modal large language model algorithm, take the next operation layer after the last task recognition layer in the multi-modal large language model algorithm as the current judgment layer; A22: Perform attention mask operations on the implicit expressions of the input tokens of the current judgment layer respectively to extract their respective attention masks; A23: Calculate the cosine similarity between the implicit expression of each input token and the attention mask respectively through the cosine similarity calculation formula to obtain their respective cosine similarities; A24: Calculate the distance between the implicit expression of each input token and the attention mask respectively through the distance calculation formula to obtain their respective mask distances; A25: Extract the implicit expressions of all input tokens whose cosine similarity is greater than the first threshold and the mask distance is less than the second threshold, record the numbers of all input units in the current judgment layer with these implicit expressions of the input tokens as inputs, so as to use the implicit expressions of all input tokens received by these input units hereafter as the redundant visual tokens of the current judgment layer; A26: Map the implicit expressions of all input tokens of the current judgment layer through the feature function, so that the implicit expression of the input token corresponding to the image coding format is zero, and the rest remains unchanged, to obtain the feature mapping result; A27: Use the feature mapping result as the input of the current judgment layer, and determine whether the degree of decline in the model performance of the multi-modal large language model algorithm is not less than a specified value. If so, proceed to the next step. If not, use the next operation layer as the current judgment layer, and then loop back to execute step A22. A28: Use all operation layers from the next operation layer of the last task recognition layer to the current judgment layer in the multi-modal large language model algorithm as the sparse fusion layer, generate a fixed ratio for reducing redundant visual tokens input to the sparse fusion layer, and transmit this fixed ratio to the execution module.

9. The inference acceleration device for cross-modal information processing according to claim 8, wherein Step A3 includes the following steps: A31: According to the running order of the multi-modal large language model algorithm, use the next operation layer of the last sparse fusion layer in the multi-modal large language model algorithm as the current detection layer. A32: Use the sigmoid activation function to project the implicit expression of the last input token of the current detection layer into the semantic space to obtain the projection expression. A33: Determine whether the current detection layer performs semantic alignment of input and output tokens based on the projection expression. If so, use the next operation layer in the multi-modal large language model algorithm as the current detection layer, and then loop back to execute step A32. If not, proceed to the next step. A34: According to the running order of the multi-modal large language model algorithm, use all operation layers from the next operation layer of the last sparse fusion layer to the current detection layer in the multi-modal large language model algorithm as the semantic alignment layer, generate a fixed ratio for reducing visual tokens input to the semantic alignment layer, and transmit this fixed ratio to the execution module.

10. An inference acceleration method for cross-modal information processing, characterized in that, The inference acceleration device for cross-modal information processing applicable to any one of claims 1-9 includes the following steps: S1: Use the acceleration device to decouple the task recognition layer, sparse fusion layer, and semantic alignment layer from all operation layers of the multi-modal large language model algorithm according to the implicit expression category of the input tokens of each operation layer executed by the central processing unit. S2: Let the user input a multi-modal instruction, and use the conversion device to perform visual encoding and projection on the user's multi-modal instruction to obtain the encoded format of the multi-modal instruction. S3: Use the central processing unit to execute the multi-modal large language model algorithm to obtain the answer encoding of the multi-modal instruction. When executing the multi-modal large language model algorithm, use the acceleration device to reduce the cross-modal information interaction operations of the task recognition layer, redundant visual tokens input to the sparse fusion layer, and visual tokens input to the semantic alignment layer by a predetermined ratio. S4: Use the synthesis device to convert the answer encoding obtained by the central processing unit into text and / or audio and give the output.

Citation Information

Patent Citations

  • Multi-modal emotion detection method and device, computer equipment and storage medium

    CN118656791A

  • Multi-modal large model-based traditional Chinese medicine tongue diagnosis analysis system and method

    CN118899077A

  • Multi-modal abstract method and system based on visual information fusion

    CN118964603A

  • Large and small model cooperative training method and device for multi-modal large language model

    CN119514645A

  • Methods and Systems for Interacting with Mobile Device

    US20190220471A1

Cited By

  • Multi-modal large language model reasoning method based on parallel visual token scheduling

    CN121189504A

  • A multi-modal large language model inference method based on parallel visual token scheduling

    CN121189504B

  • Layered vision Token compression method for multi-modal large language model

    CN121350590A

  • A hierarchical visual token compression method for multi-modal large language model

    CN121350590B