A device and method for accelerating inference in cross-modal information processing

By decoupling the task recognition layer, sparse fusion layer, and semantic alignment layer of the multimodal large language model, the operation of the multimodal large language model is optimized, the problems of inference latency and video memory occupancy are solved, and efficient cross-modal information processing is achieved.

CN120197713BActive Publication Date: 2025-10-03NINGBO ORIENTAL UNIV OF TECH (TEMPORARY NAME)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510680375.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-10-03
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

Existing large multimodal language models have problems with inference latency and video memory usage in cross-modal processing, and existing solutions find it difficult to effectively balance computing efficiency and model performance.

Method used

By setting up an inference acceleration device for cross-modal information processing, including a conversion device, a central processing unit, a synthesis device and an acceleration device, the task recognition layer, sparse fusion layer and semantic alignment layer of the multimodal large language model algorithm are decoupled, and redundant calculations are reduced to optimize model operation.

Benefits of technology

It shortens the inference delay time, improves the model compression efficiency, alleviates the computing resource usage, maintains high performance, and achieves a balance between computing efficiency and model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197713B_ABST
    Figure CN120197713B_ABST
Patent Text Reader

Abstract

The present invention relates to an inference acceleration device and method for cross-modal information processing. A conversion device is provided to visually encode and project a user's multimodal instructions into a multimodal instruction encoding format. A central processing unit then executes a multimodal large language model algorithm based on this encoding format to obtain an answer encoding for the multimodal instruction. The obtained answer encoding is then converted and output by a provided synthesis device. Furthermore, an acceleration device is provided to decouple the task recognition layer, sparse fusion layer, and semantic alignment layer from all operational layers of the multimodal large language model algorithm based on the implicit expression category of the input tokens of each operational layer in the multimodal large language model algorithm. This then reduces the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer by a predetermined ratio, thereby shortening the inference delay time and alleviating the model occupancy rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer model technology, and in particular to an inference acceleration device and method for cross-modal information processing. Background Art

[0002] Currently, multimodal large language models (MLLMs) achieve cross-modal processing by integrating encoders for visual, audio, and other functions with large language models (LLMs). However, these models face significant computational efficiency bottlenecks. Specifically, the number of tokens generated by the visual encoder is often several times that of text tokens, leading to a sharp increase in inference latency and graphics memory usage. Existing solutions, such as visual token pruning and fusion technologies, lack a deep understanding of the internal cross-modal processing mechanisms of multimodal large language models, making it difficult to effectively balance computational efficiency and model performance. Summary of the Invention

[0003] The technical problem to be solved by this invention is how to overcome the technical drawback of inference delay in existing large multimodal language models. To overcome the above-mentioned drawbacks of the existing technology, this invention provides an inference acceleration device and method for cross-modal information processing, specifically comprising an inference acceleration device and a method for cross-modal information processing.

[0004] The present invention provides an inference acceleration device for cross-modal information processing, comprising:

[0005] a conversion device configured to visually encode and project the user's multimodal instructions to convert them into an encoding format of the multimodal instructions;

[0006] a central processing unit, electrically connected to the conversion device, configured to execute a multimodal large language model algorithm to obtain an answer encoding for a multimodal instruction;

[0007] A synthesis device, electrically connected to the central processing unit, configured to convert the answer obtained by the central processing unit into text and / or audio and output it;

[0008] An acceleration device is electrically connected to the central processing unit and is configured to decouple the task identification layer, the sparse fusion layer, and the semantic alignment layer from all operation layers of the multimodal large language model algorithm according to the implicit expression category of the input token of each operation layer in the multimodal large language model algorithm, and then reduce the cross-modal information interaction operation of the task identification layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer according to a predetermined ratio to accelerate the operation of the multimodal large language model algorithm.

[0009] The inference acceleration device for cross-modal information processing disclosed in the present invention can visually encode and project the user's multimodal instructions by setting a conversion device to convert them into a coding format of multimodal instructions. This coding format will be received by the central processing unit set, and the central processing unit will execute the multimodal large language model algorithm to obtain the answer code of the multimodal instruction. The obtained answer code will be converted into text and / or audio by the synthesis device set and output. On this basis, in order to accelerate the operation of the multimodal large language model algorithm, an acceleration device is set to decouple the task recognition layer, sparse fusion layer and semantic alignment layer from all operation layers of the multimodal large language model algorithm according to the implicit expression category of the input tokens of each operation layer in the multimodal large language model algorithm, and then reduce the cross-modal information interaction operation of the task recognition layer, the redundant visual tokens of the input sparse fusion layer and the visual tokens of the input semantic alignment layer according to a predetermined ratio to accelerate the operation of the multimodal large language model algorithm and shorten the inference delay time. The model inference process of the multimodal large language model algorithm can be decoupled into the task recognition layer, the sparse fusion layer and the semantic alignment layer, providing universal support for the model optimization of the multimodal large language model algorithm, improving the model compression efficiency, thereby alleviating the model occupancy rate, and effectively balancing computing efficiency and model performance.

[0010] In a possible implementation, the multimodal instruction includes a prompt text, a text instruction, and an associated image bundled with the text instruction.

[0011] In one possible implementation, the conversion device includes:

[0012] an image encoding module, electrically connected to the central processing unit, and configured to visually encode and project the associated image to convert it into an image encoding format compatible with the input of the multimodal large language model algorithm;

[0013] a word segmenter, electrically connected to the central processing unit, configured to sequentially perform text segmentation and encoding mapping on the text instruction and the prompt text, so as to convert them into a text encoding format compatible with the input of the multimodal large language model algorithm;

[0014] Therefore, under the parallel processing of the image encoding module and the word segmenter, it is possible to convert text instructions, associated images and prompt text into a format compatible with the input format of the multimodal large language model algorithm, thereby further improving the operating efficiency of the multimodal large language model algorithm and reducing processing and inference delay time.

[0015] In a possible implementation, the image encoding module includes:

[0016] an image encoder configured to perform visual encoding of the associated image to obtain image encoding data;

[0017] a projection device, electrically connected to both the image encoder and the central processing unit, configured to perform a projection conversion on the image encoding data to obtain an image encoding format adapted to the input of the multimodal large language model algorithm;

[0018] This solution encodes the associated image through an image encoder, and under the processing of the projection device, it can obtain an image encoding format that is compatible with the input of the multimodal large language model algorithm. It can not only extract the key information of the image, but also further reduce the inference delay of the multimodal large language model algorithm.

[0019] In one possible implementation, the acceleration device includes:

[0020] a decoupling module, electrically connected to the central processing unit, and configured to decouple the task identification layer, the sparse fusion layer, and the semantic alignment layer from all operation layers of the multimodal large language model algorithm according to the implicit expression category of the input token of each operation layer in the multimodal large language model algorithm when the multimodal large language model algorithm is initially executed;

[0021] an execution module electrically connected to the decoupling module and configured to reduce, by a predetermined ratio, the cross-modal information interaction operation of the task recognition layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer when executing the multimodal large language model algorithm, so as to accelerate the operation of the multimodal large language model algorithm;

[0022] This solution sets a decoupling module to decouple the task recognition layer, sparse fusion layer, and semantic alignment layer from all computational layers of the multimodal large language model algorithm when the algorithm is first executed. It also sets an execution module to reduce the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer by a predetermined ratio when the algorithm is executed for the second or subsequent times. This accelerates the operation of the multimodal large language model algorithm, improves the efficiency of the central processor in obtaining answers, and effectively balances computing efficiency and model performance.

[0023] In one possible implementation, the decoupling module is configured to perform the following steps:

[0024] A1: Starting from the first operation layer in the multimodal large language model algorithm, the implicit expression of the input token of each operation layer is sequentially classified based on the task identification judgment standard to extract all the task identification layers of the multimodal large language model algorithm;

[0025] A2: Based on the sparse fusion judgment standard, the implicit expression of the input token of each currently remaining operation layer in the multimodal large language model algorithm is sequentially classified to extract all the sparse fusion layers of the multimodal large language model algorithm;

[0026] A3: Based on the semantic alignment judgment standard, the implicit expression of the input token of each remaining operation layer in the multimodal large language model algorithm is sequentially classified to extract all the semantic alignment layers of the multimodal large language model algorithm;

[0027] This solution can ensure the decoupling of the task extraction layer, sparse fusion layer, and semantic alignment layer, improve generalization capabilities, provide universal support for model optimization of multimodal large language model algorithms, and further improve model compression efficiency.

[0028] In a possible implementation, step A1 includes the following steps:

[0029] A11: According to the running order of the multimodal large language model algorithm, the first computing layer in the multimodal large language model algorithm is used as the current layer;

[0030] A12: Use the S-type activation function to project the implicit expression of the last input token of the current layer into the semantic space to obtain the semantic expression;

[0031] A13: Determine whether the current layer is a task summary of the text instructions input by the user based on the semantic expression.

[0032] If yes, proceed to the next step;

[0033] If not, proceed to step A16;

[0034] A14: Perform visual attention merging and attention mask extraction on the implicit representations of multiple input tokens of the current layer in sequence to obtain an attention mask.

[0035] A15: Determine whether cross-modal information fusion occurs based on the attention mask.

[0036] If yes, proceed to the next step;

[0037] If not, the next operation layer in the multimodal large language model algorithm is used as the current layer, and then the step A12 is executed again.

[0038] A16: In the order in which the multimodal large language model algorithm is run, all computation layers from the first computation layer to the current computation layer in the multimodal large language model algorithm are used as the task identification layer, and a fixed ratio for reducing cross-modal information interaction computations in the task identification layer is generated, and the fixed ratio is transmitted to the execution module.

[0039] This solution confirms the task identification layer based on cross-modal information interaction and semantic expression, and generates a fixed ratio to reduce the cross-modal information interaction operations in the task identification layer, ensuring that the execution module removes an appropriate amount of redundant calculations when running the task identification layer, further improving the efficiency of reasoning operations.

[0040] In a possible implementation, step A2 includes the following steps:

[0041] A21: According to the running order of the multimodal large language model algorithm, the next computing layer of the last task recognition layer in the multimodal large language model algorithm is used as the current judgment layer;

[0042] A22: Perform attention masking operations on the implicit representation of each input token of the current judgment layer to extract their respective attention masks;

[0043] A23: Calculate the cosine similarity between the implicit expression of each input token and the attention mask using the cosine similarity calculation formula to obtain their respective cosine similarities;

[0044] A24: Calculate the distance between the implicit expression of each input token and the attention mask using the distance calculation formula to obtain their respective mask distances;

[0045] A25: Extracting implicit expressions of all input tokens whose cosine similarity is greater than a first threshold and whose mask distance is less than a second threshold, recording the numbers of all input units in the current judgment layer that use the implicit expressions of these input tokens as input, and using the implicit expressions of all input tokens subsequently received by these input units as redundant visual tokens of the current judgment layer;

[0046] A26: Mapping the implicit expressions of all input tokens of the current judgment layer by the feature function so that the implicit expressions of the input tokens corresponding to the image coding format are zero and the others remain unchanged, thereby obtaining a feature mapping result;

[0047] A27: Using the feature mapping result as the input of the current judgment layer, it is determined whether the degree of degradation of the model performance of the multimodal large language model algorithm is not less than a specified value.

[0048] If yes, proceed to the next step;

[0049] If not, the next calculation layer is used as the current judgment layer, and then the step A22 is executed again;

[0050] A28: Using all computing layers from the next computing layer of the last task recognition layer to the current judgment layer in the multimodal large language model algorithm as the sparse fusion layer, generating a fixed ratio for reducing redundant visual tokens input into the sparse fusion layer, and transmitting the fixed ratio to the execution module;

[0051] While decoupling the sparse fusion layer, this scheme proposes a dual-metric dynamic mask evaluation method (cosine similarity and distance) to quantify the degree of perturbation of a single visual token on the output, overcoming the misjudgment problem caused by attention convergence in traditional attention weighting methods.

[0052] In a possible implementation, step A3 includes the following steps:

[0053] A31: According to the running order of the multimodal large language model algorithm, the next computing layer of the last sparse fusion layer in the multimodal large language model algorithm is used as the current detection layer;

[0054] A32: Use the S-type activation function to project the implicit expression of the last input token of the current detection layer into the semantic space to obtain the projected expression;

[0055] A33: Determine whether the current detection layer is performing semantic alignment of input and output tokens based on the projection expression.

[0056] If yes, the next computation layer in the multimodal large language model algorithm is used as the current detection layer, and the process then returns to step A32;

[0057] If not, proceed to the next step;

[0058] A34: According to the running order of the multimodal large language model algorithm, all the operation layers from the next operation layer of the last sparse fusion layer to the current detection layer in the multimodal large language model algorithm are used as the semantic alignment layer, and a fixed ratio of reducing the visual tokens input into the semantic alignment layer is generated, and this fixed ratio is transmitted to the execution module.

[0059] While decoupling the semantic alignment layer, the above solution generates a fixed ratio that reduces the visual tokens input to the semantic alignment layer, so that the execution module removes the visual tokens, thereby further reducing the computational complexity and cache usage during the model inference process of the multimodal large language model algorithm.

[0060] Another technical solution of the present invention is to provide a method for accelerating inference of cross-modal information processing, comprising the following steps:

[0061] S1: decoupling a task recognition layer, a sparse fusion layer, and a semantic alignment layer from all operation layers of the multimodal large language model algorithm executed by a central processing unit, using an acceleration device according to implicit expression categories of input tokens of each operation layer in the multimodal large language model algorithm;

[0062] S2: allowing the user to input a multimodal instruction so that the user's multimodal instruction is visually encoded and projected by the conversion device to obtain an encoding format of the multimodal instruction;

[0063] S3: executing the multimodal large language model algorithm by the central processor to obtain the answer code of the multimodal instruction, and when executing the multimodal large language model algorithm, reducing the cross-modal information interaction operation of the task recognition layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer by the acceleration device according to a predetermined ratio;

[0064] S4: The answer obtained by the central processor is encoded and converted into text and / or audio through a synthesis device, and outputted.

[0065] The method disclosed in the present application first uses an acceleration device to decouple the task recognition layer, sparse fusion layer and semantic alignment layer from all operation layers of the multimodal large language model algorithm according to the implicit expression category of the input token of each operation layer in the multimodal large language model algorithm. Then, a conversion device is used to visually encode and project the user's multimodal instructions to convert them into an encoding format of the multimodal instructions. This encoding format is received by the central processing unit, and the multimodal large language model algorithm is executed by the central processing unit to obtain the answer code of the multimodal instruction. The obtained answer code is converted into text and / or audio by a synthesis device and output. When executing the multimodal large language model algorithm, the acceleration device is used to reduce the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens of the input sparse fusion layer, and the visual tokens of the input semantic alignment layer according to a predetermined ratio, so as to accelerate the running of the multimodal large language model algorithm and shorten the inference delay time. Ultimately, the model inference process of the multimodal large language model algorithm can be decoupled into the task recognition layer, the sparse fusion layer, and the semantic alignment layer, providing universal support for the model optimization of the multimodal large language model algorithm, improving the model compression efficiency, thereby alleviating the model occupancy rate, and effectively balancing computing efficiency and model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 This is a schematic diagram of the structure of an inference acceleration device for cross-modal information processing disclosed in an embodiment of the present application;

[0067] Figure 2 This is a flowchart of the operation of the decoupling module disclosed in the embodiments of this application;

[0068] Figure 3 This is a flowchart corresponding to step A1 disclosed in the embodiments of this application;

[0069] Figure 4 This is a flowchart corresponding to step A2 disclosed in the embodiments of this application;

[0070] Figure 5 This is a flowchart corresponding to step A3 disclosed in the embodiments of this application;

[0071] Figure 6 A flow chart of the method disclosed in the embodiments of this application;

[0072] Figure 7 This is a diagram showing the execution effect of the multimodal large language model algorithm after being accelerated by the acceleration device disclosed in the embodiments of this application. DETAILED DESCRIPTION

[0073] First, those skilled in the art should understand that these embodiments are merely used to explain the technical principles of the embodiments of the present application and are not intended to limit the scope of protection of the embodiments of the present application. Those skilled in the art may adjust them as needed to suit specific application scenarios.

[0074] In the embodiments of the present application, unless otherwise clearly specified and limited, the electrical connection between the first feature and the second feature means that there is transmission of electrical signals between the first feature and the second feature, that is, there is an electrical relationship, and the way to achieve the transmission of electrical signals may be electrical connection of wires, radio connection, electrical connection of electromagnetic media (such as semiconductors), communication achieved by channels, etc.

[0075] In the embodiments of the present application, unless otherwise specified or limited, the distance between two objects should be understood according to the mathematical meaning of distance, that is, the so-called object with objects distance It refers to the mapping of any two objects in the set in question (such as the implicit expression of the input token and the attention mask in this embodiment) to the set of real numbers. This mapping is for any object in the set in question. , object and objects At the same time, the following three requirements must be met: (1) , if and only if When the equality sign holds; (2) ; (3) Specifically, it can be the Euclidean distance, or the distance induced by the inner product or norm, and those skilled in the art can adjust it according to specific needs.

[0076] In the embodiments of the present application, unless otherwise expressly specified or limited, a first feature being "above" or "below" a second feature may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Furthermore, a first feature being "above," "above," and "above" a second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. A first feature being "below," "below," and "below" a second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.

[0077] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0078] See also Figures 1 to 7 , the embodiment of the present application discloses an inference acceleration device for cross-modal information processing, the structural diagram of the inference acceleration device is as follows Figure 1 As shown, the inference acceleration device includes a conversion device, a central processing unit, a synthesis device and an acceleration device, wherein the central processing unit is electrically connected to the conversion device, the synthesis device is electrically connected to the central processing unit, and the acceleration device is electrically connected to the central processing unit.

[0079] In the inference acceleration device, the conversion device is configured to visually encode and project the user's multimodal instructions to convert them into an encoding format of the multimodal instructions, wherein the multimodal instructions include prompt text, text instructions, and associated images bundled with the text instructions.

[0080] like Figure 1 As shown, in this embodiment, the conversion device includes an image encoding module and a tokenizer, the tokenizer is electrically connected to the central processing unit, the image encoding module includes an image encoder and a projection device, and the projection device is electrically connected to the image encoder and the central processing unit at the same time.

[0081] In the conversion device, the image encoding module is configured to visually encode and project the associated image to convert it into an image encoding format that is compatible with the input of the multimodal large language model algorithm. Specifically, the image encoder in the image encoding module first visually encodes the associated image to obtain image encoding data. Secondly, the projection device in the image encoding module performs projection conversion on the image encoding data to obtain an image encoding format that is compatible with the input of the multimodal large language model algorithm. The word segmenter is configured to perform text segmentation and encoding mapping on the text instructions and prompt text in turn to convert them into a text encoding format that is compatible with the input of the multimodal large language model algorithm. The image encoding format and the text encoding format are the set of input tokens of the multimodal large language model algorithm.

[0082] Please continue to see Figure 1In the inference acceleration device, the central processing unit is configured to execute a multimodal large language model algorithm to obtain the answer code of the multimodal instruction; the synthesis device is configured to convert the answer code obtained by the central processing unit into text and / or audio and output it. The conversion method is the existing technology and will not be expanded here.

[0083] Please continue to see Figure 1 In this inference acceleration device, the acceleration device is configured to decouple the task identification layer, sparse fusion layer, and semantic alignment layer from all computational layers of the multimodal language model algorithm based on the implicit expression categories of the input tokens of each computational layer in the multimodal language model algorithm. The device then reduces the cross-modal information interaction operations of the task identification layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer by a predetermined ratio, thereby accelerating the execution of the multimodal language model algorithm. In this embodiment, the acceleration device includes a decoupling module and an execution module. The decoupling module is electrically connected to the central processing unit, and the execution module is electrically connected to the decoupling module.

[0084] See also Figure 2 In the acceleration device, the decoupling module is configured to decouple the task recognition layer, the sparse fusion layer, and the semantic alignment layer from all operation layers of the multimodal large language model algorithm according to the implicit expression category of the input token of each operation layer in the multimodal large language model algorithm when the multimodal large language model algorithm is executed for the first time.

[0085] like Figure 2 As shown, in this embodiment, the decoupling module is configured to perform the following steps:

[0086] A1: Starting from the first operation layer in the multimodal large language model algorithm, the implicit expression of the input token of each operation layer is classified based on the task identification judgment criteria to extract all task identification layers of the multimodal large language model algorithm.

[0087] See also Figure 3 Specifically, in this embodiment, each operation layer in the multimodal large language model algorithm is divided according to the reception of the implicit expression of the input token, and the neuron group in the operation layer that can receive the implicit expression of an input token is regarded as an input unit, and the input units are numbered according to the reception sequence.

[0088] In this embodiment, step A1 includes the following steps:

[0089] A11: According to the running order of the multimodal large language model algorithm, the first operation layer in the multimodal large language model algorithm is used as the current layer.

[0090] A12: The implicit expression of the last input token of the current layer is projected into the semantic space using a sigmoid activation function to obtain a semantic expression. The sigmoid activation function in this embodiment is a softmax activation function. The implicit expression of the input token received by the last input unit of each computation layer is considered to be the implicit expression of the last input token.

[0091] A13: Determine whether the current layer is summarizing the task of the text instructions input by the user based on the semantic expression (this can be compared with the mapping results of the standard S-type activation function for summarizing the task of the text instructions input by the user).

[0092] If yes, proceed to the next step;

[0093] If not, go to step A16.

[0094] A14: Perform visual attention merging and attention mask extraction on the implicit expressions of multiple input tokens of the current layer in sequence to obtain an attention mask.

[0095] A15: Determine whether cross-modal information fusion occurs based on the attention mask.

[0096] If yes, proceed to the next step;

[0097] If not, the next operation layer in the multimodal large language model algorithm is used as the current layer, and then the process returns to step A12.

[0098] A16: According to the running order of the multimodal large language model algorithm, all operation layers from the first operation layer to the current layer in the multimodal large language model algorithm are used as task identification layers, and a fixed ratio for reducing the cross-modal information interaction operation of the task identification layer is generated. The fixed ratio in this embodiment is 1, and this fixed ratio is transmitted to the execution module.

[0099] A2: Based on the sparse fusion judgment criteria, the implicit expression of the input token of each currently remaining operation layer in the multimodal large language model algorithm is classified in turn to extract all the sparse fusion layers of the multimodal large language model algorithm.

[0100] See also Figure 4 In this embodiment, step A2 includes the following steps:

[0101] A21: According to the running order of the multimodal large language model algorithm, the next operation layer of the last task recognition layer in the multimodal large language model algorithm is used as the current judgment layer.

[0102] A22: Perform attention mask operations on the implicit expressions of each input token of the current judgment layer to extract their respective attention masks.

[0103] A23: Calculate the cosine similarity between the implicit expression of each input token and the attention mask using the cosine similarity calculation formula to obtain their respective cosine similarities.

[0104] In this embodiment, the cosine similarity calculation formula is as follows:

[0105] ,

[0106] Where,

[0107] Representative implicit representation of the input tokens (vector);

[0108] Representative attention mask for each input token;

[0109] Represents cosine similarity.

[0110] A24: Calculate the distance between the implicit expression of each input token and the attention mask through the distance calculation formula to obtain their respective mask distances.

[0111] In this embodiment, the distance calculation formula is as follows:

[0112] ,

[0113] Where,

[0114] represents the distance induced by the 2-norm.

[0115] A25: Extract the implicit representations of all input tokens whose cosine similarity is greater than a first threshold and whose mask distance is less than a second threshold. Record the numbers of all input units in the current judgment layer that use the implicit representations of these input tokens as input, so that the implicit representations of all input tokens subsequently received by these input units are used as redundant visual tokens for the current judgment layer. In this embodiment, the first threshold is 0.999 and the second threshold is 0.2.

[0116] A26: Map the implicit expressions of all input tokens of the current judgment layer through the feature function so that the implicit expressions of the input tokens corresponding to the image encoding format are zero, and the rest remain unchanged to obtain the feature mapping result.

[0117] A27: Use the feature mapping result as the input of the current judgment layer to determine whether the performance degradation of the multimodal large language model algorithm is no less than the specified value.

[0118] If yes, proceed to the next step;

[0119] If not, the next operation layer is used as the current judgment layer according to the operation order of the multimodal large language model algorithm, and then the process returns to step A22.

[0120] A28: According to the running order of the multimodal large language model algorithm, all the operation layers from the next operation layer of the last task recognition layer to the current judgment layer in the multimodal large language model algorithm are used as sparse fusion layers, and a fixed ratio is generated to reduce the redundant visual tokens input to the sparse fusion layer. The fixed ratio set in this embodiment is 1, and this fixed ratio is transmitted to the execution module.

[0121] A3: Based on the semantic alignment judgment standard, the implicit expression of the input token of each remaining operation layer in the multimodal large language model algorithm is classified in turn to extract all the semantic alignment layers of the multimodal large language model algorithm.

[0122] See also Figure 5 In this embodiment, step A3 includes the following steps:

[0123] A31: According to the running order of the multimodal large language model algorithm, the next operation layer after the last sparse fusion layer in the multimodal large language model algorithm is used as the current detection layer.

[0124] A32: Use the S-type activation function to project the implicit expression of the last input token of the current detection layer into the semantic space to obtain the projected expression.

[0125] A33: Determine whether the current detection layer is performing semantic alignment of input and output tokens based on the projection expression.

[0126] If yes, the next computation layer in the multimodal large language model algorithm is used as the current detection layer, and the process then returns to step A32;

[0127] If not, proceed to the next step.

[0128] A34: According to the running order of the multimodal large language model algorithm, all the operation layers from the next operation layer of the last sparse fusion layer to the current detection layer in the multimodal large language model algorithm are used as semantic alignment layers, and a fixed ratio of visual tokens input to the semantic alignment layer is generated. In this embodiment, the fixed ratio is set to 1, and this fixed ratio is transmitted to the execution module.

[0129] In the acceleration device, the execution module is configured to reduce the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens of the input sparse fusion layer, and the visual tokens of the input semantic alignment layer by a predetermined ratio when executing the multimodal large language model algorithm, so as to accelerate the running of the multimodal large language model algorithm.

[0130] See also Figure 6The following further discloses a method for using the inference acceleration device for cross-modal information processing described in this embodiment, which includes the following steps:

[0131] S1: Decoupling the task recognition layer, sparse fusion layer, and semantic alignment layer from all operation layers of the multimodal large language model algorithm through an acceleration device based on the implicit expression category of the input token of each operation layer in the multimodal large language model algorithm executed by the central processing unit.

[0132] S2: The user is prompted to input a multimodal instruction so that the user's multimodal instruction is visually encoded and projected by a conversion device to obtain an encoding format of the multimodal instruction.

[0133] S3: Execute the multimodal large language model algorithm through the central processing unit to obtain the answer code of the multimodal instruction, and when executing the multimodal large language model algorithm, reduce the cross-modal information interaction operation of the task recognition layer, the redundant visual tokens of the input sparse fusion layer, and the visual tokens of the input semantic alignment layer by a predetermined ratio through the acceleration device.

[0134] S4: The answer obtained by the central processing unit is encoded and converted into text and / or audio through a synthesis device, and outputted.

[0135] See also Figure 7 The following describes in detail the technical effects of the cross-modal information processing inference acceleration device described in this embodiment. Systematic testing was conducted by applying the acceleration device in this embodiment to CPUs running three multimodal large language model algorithms (with 3B to 34B parameters): LLaVA-1.5, mobileVLM, and Intern-VL. The tests covered benchmark tasks such as image question answering (VQA-v2, GQA), fine-grained recognition (TextVQA, ST-VQA), and visual reasoning (ScienceQA). Experimental results show that the acceleration device in this embodiment maintains greater than 99% performance while reducing floating-point operations by 54.6% and key-value caching by 98.9%. Figure 7 The figure shows the running simulation effect of a question and answer. In the figure, the task recognition layer is abbreviated as the shallow layer, the sparse fusion layer is abbreviated as the middle layer, and the semantic alignment layer and the remaining operation layer are abbreviated as the deep layer.

[0136] The inference acceleration device for cross-modal information processing disclosed in this embodiment can visually encode and project the user's multimodal instructions to convert them into a coding format of multimodal instructions by setting a conversion device. This coding format will be received by the set central processing unit, and the central processing unit will execute the multimodal large language model algorithm to obtain the answer code of the multimodal instruction. The obtained answer code will be converted into text and / or audio by the set synthesis device and output. On this basis, an acceleration device is set up to decouple the task recognition layer, sparse fusion layer and semantic alignment layer from all the operation layers of the multimodal large language model algorithm according to the implicit expression category of the input token of each operation layer in the multimodal large language model algorithm, and then reduce the cross-modal information interaction operation of the task recognition layer, the redundant visual tokens of the input sparse fusion layer and the visual tokens of the input semantic alignment layer according to a predetermined ratio to accelerate the running of the multimodal large language model algorithm and shorten the inference delay time. The model inference process of the multimodal large language model algorithm can be decoupled into the task recognition layer, sparse fusion layer and semantic alignment layer, providing universal support for the model optimization of the multimodal large language model algorithm, improving the model compression efficiency, and then alleviating the model occupancy rate, effectively balancing computing efficiency and model performance.

[0137] In the description of the embodiments of the present application, it should be noted that in the description of the present application, terms such as "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present application.

[0138] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "in the present embodiment", "specific example", or "some examples" means that the specific features, mechanisms, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are mutually inconsistent.

[0139] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A device for accelerating inference in cross-modal information processing, characterized in that: include: A conversion device configured to visually encode and project a user's multimodal instruction to convert it into an encoded format of the multimodal instruction; the multimodal instruction includes a prompt text, a text instruction, and an associated image bundled with the text instruction; a central processing unit, electrically connected to the conversion device, configured to execute a multimodal large language model algorithm to obtain an answer encoding for a multimodal instruction; A synthesis device, electrically connected to the central processing unit, configured to convert the answer obtained by the central processing unit into text and / or audio and output it; an acceleration device, electrically connected to the central processing unit, configured to decouple the task identification layer, the sparse fusion layer, and the semantic alignment layer from all operation layers of the multimodal large language model algorithm according to the implicit expression category of the input tokens of each operation layer in the multimodal large language model algorithm, and then reduce the cross-modal information interaction operation of the task identification layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer according to a predetermined ratio, so as to accelerate the operation of the multimodal large language model algorithm; Wherein, the conversion device includes: an image encoding module, electrically connected to the central processing unit, and configured to visually encode and project the associated image to convert it into an image encoding format compatible with the input of the multimodal large language model algorithm; a word segmenter, electrically connected to the central processing unit, configured to sequentially perform text segmentation and encoding mapping on the text instruction and the prompt text, so as to convert them into a text encoding format compatible with the input of the multimodal large language model algorithm; In addition, according to the implicit expression category of the input token of each operation layer in the multimodal large language model algorithm, the process of decoupling the task identification layer, the sparse fusion layer and the semantic alignment layer from all operation layers of the multimodal large language model algorithm includes the following steps; A1: Starting from the first operation layer in the multimodal large language model algorithm, the implicit expression of the input token of each operation layer is sequentially classified based on the task identification judgment standard to extract all the task identification layers of the multimodal large language model algorithm; A2: Based on the sparse fusion judgment standard, the implicit expression of the input token of each currently remaining operation layer in the multimodal large language model algorithm is sequentially classified to extract all the sparse fusion layers of the multimodal large language model algorithm; A3: Based on the semantic alignment judgment standard, the implicit expression of the input token of each remaining operation layer in the multimodal large language model algorithm is sequentially classified to extract all the semantic alignment layers of the multimodal large language model algorithm; The step A1 comprises the following steps: A11: According to the running order of the multimodal large language model algorithm, the first computing layer in the multimodal large language model algorithm is used as the current layer; A12: Use the S-type activation function to project the implicit expression of the last input token of the current layer into the semantic space to obtain the semantic expression; A13: Determine whether the current layer is a task summary of the text instructions input by the user based on the semantic expression. If yes, proceed to the next step; If not, proceed to step A16; A14: Perform visual attention merging and attention mask extraction on the implicit representations of multiple input tokens of the current layer in sequence to obtain an attention mask. A15: Determine whether cross-modal information fusion occurs based on the attention mask. If yes, proceed to the next step; If not, the next operation layer in the multimodal large language model algorithm is used as the current layer, and then the step A12 is executed again. A16: In the order in which the multimodal large language model algorithm is run, all computation layers from the first computation layer to the current computation layer in the multimodal large language model algorithm are used as the task identification layer, and a fixed ratio for reducing cross-modal information interaction computations in the task identification layer is generated. The step A2 comprises the following steps: A21: According to the running order of the multimodal large language model algorithm, the next computing layer of the last task recognition layer in the multimodal large language model algorithm is used as the current judgment layer; A22: Perform attention masking operations on the implicit representation of each input token of the current judgment layer to extract their respective attention masks; A23: Calculate the cosine similarity between the implicit expression of each input token and the attention mask using the cosine similarity calculation formula to obtain their respective cosine similarities; A24: Calculate the distance between the implicit expression of each input token and the attention mask using the distance calculation formula to obtain their respective mask distances; A25: Extracting implicit expressions of all input tokens whose cosine similarity is greater than a first threshold and whose mask distance is less than a second threshold, recording the numbers of all input units in the current judgment layer that use the implicit expressions of these input tokens as input, and using the implicit expressions of all input tokens subsequently received by these input units as redundant visual tokens of the current judgment layer; A26: Mapping the implicit expressions of all input tokens of the current judgment layer by the feature function so that the implicit expressions of the input tokens corresponding to the image coding format are zero and the others remain unchanged, thereby obtaining a feature mapping result; A27: Using the feature mapping result as the input of the current judgment layer, it is determined whether the degree of degradation of the model performance of the multimodal large language model algorithm is not less than a specified value. If yes, proceed to the next step; If not, the next calculation layer is used as the current judgment layer, and then the step A22 is executed again; A28: Using all the computing layers from the next computing layer of the last task recognition layer to the current judgment layer in the multimodal large language model algorithm as the sparse fusion layer, and generating a fixed ratio for reducing redundant visual tokens input into the sparse fusion layer; Step A3 includes the following steps: A31: According to the running order of the multimodal large language model algorithm, the next computing layer of the last sparse fusion layer in the multimodal large language model algorithm is used as the current detection layer; A32: Use the S-type activation function to project the implicit expression of the last input token of the current detection layer into the semantic space to obtain the projected expression; A33: Determine whether the current detection layer is performing semantic alignment of input and output tokens based on the projection expression. If yes, the next computation layer in the multimodal large language model algorithm is used as the current detection layer, and the process then returns to step A32; If not, proceed to the next step; A34: According to the running order of the multimodal large language model algorithm, all the operation layers from the next operation layer of the last sparse fusion layer to the current detection layer in the multimodal large language model algorithm are used as the semantic alignment layer, and a fixed ratio for reducing the visual tokens input into the semantic alignment layer is generated.

2. The inference acceleration device for cross-modal information processing according to claim 1, characterized in that: The image encoding module includes: an image encoder configured to perform visual encoding of the associated image to obtain image encoding data; The projection device is electrically connected to the image encoder and the central processing unit, and is configured to perform projection conversion on the image coding data to obtain an image coding format adapted to the input of the multimodal large language model algorithm.

3. The inference acceleration device for cross-modal information processing according to claim 2, characterized in that: The acceleration device comprises: a decoupling module, electrically connected to the central processing unit, and configured to, when the multimodal large language model algorithm is initially executed, decouple the task recognition layer, the sparse fusion layer, and the semantic alignment layer from all computational layers of the multimodal large language model algorithm according to the implicit expression category of the input token of each computational layer in the multimodal large language model algorithm; An execution module is electrically connected to the decoupling module and is configured to reduce the cross-modal information interaction operations of the task recognition layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer by a predetermined ratio when executing the multimodal large language model algorithm, so as to accelerate the operation of the multimodal large language model algorithm.

4. A method for accelerating inference in cross-modal information processing, characterized in that: An inference acceleration device for cross-modal information processing applicable to any one of claims 1 to 3, comprising the following steps: S1: decoupling a task recognition layer, a sparse fusion layer, and a semantic alignment layer from all operation layers of the multimodal large language model algorithm executed by a central processing unit, using an acceleration device according to implicit expression categories of input tokens of each operation layer in the multimodal large language model algorithm; S2: allowing the user to input a multimodal instruction so that the user's multimodal instruction is visually encoded and projected by the conversion device to obtain an encoding format of the multimodal instruction; S3: executing the multimodal large language model algorithm by the central processor to obtain the answer code of the multimodal instruction, and when executing the multimodal large language model algorithm, reducing the cross-modal information interaction operation of the task recognition layer, the redundant visual tokens input to the sparse fusion layer, and the visual tokens input to the semantic alignment layer by the acceleration device according to a predetermined ratio; S4: The answer obtained by the central processor is encoded and converted into text and / or audio through a synthesis device, and outputted.

Citation Information

Patent Citations

  • Multi-modal emotion detection method and device, computer equipment and storage medium

    CN118656791A

  • Large and small model cooperative training method and device for multi-modal large language model

    CN119514645A