Multi-modal large model reasoning acceleration method
By performing feature analysis and modal complexity evaluation on multimodal large models, dynamically selecting the calculation depth and parameter quantity, using hierarchical fusion strategy and low-rank cross-modal attention calculation, the problems of wasted resources and increased delays in multimodal large models are solved, and more efficient hardware utilization and data throughput are achieved.
Patent Information
- Application Number
- CN202510339921.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
The existing multimodal large models waste serious computing resources and increase delays during inference. The existing technology has failed to effectively solve the problems of high computational complexity and high delays.
By performing feature analysis and modal complexity evaluation of the input multimodal data, the calculation depth and parameter quantity of the single-modal subnet are dynamically selected, and a hierarchical fusion strategy and low-rank cross-modal attention calculation are adopted, and the modal processing module is allocated in combination with hardware characteristics, and the results are synchronized through high-speed bus.
Reduces computational complexity, reduces memory usage, improves hardware utilization, reduces latency and redundant computing, and improves data throughput.
Smart Images

Figure CN120278264A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a method for accelerating the inference of a multimodal large model. Background Art
[0002] A multimodal large model is an artificial intelligence model that can simultaneously process multiple different types of data (such as text, images, voice, video, sensor data, etc.). Through deep neural network technology, it fuses and correlates information of different modalities, thereby achieving a more comprehensive and intelligent understanding and generation ability. The method for accelerating the inference of a multimodal large model aims to solve the problems of high computational complexity and large latency, and improve the response speed of the model in actual scenarios through technical optimization.
[0003] However, existing multimodal models use the same computational path for all input data during inference, resulting in waste of computational resources; or the fusion strategy is too complex, resulting in increased latency. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a method for accelerating the inference of a multimodal large model.
[0005] The technical solution adopted to solve the above technical problem is: A method for accelerating the inference of a multimodal large model, including the following specific steps:
[0006] Step 1: Perform feature analysis and modal complexity evaluation on the input multimodal data;
[0007] Step 2: Dynamically select the computational depth and the number of parameters of the unimodal sub-network according to the complexity;
[0008] Step 3: Adopt a hierarchical fusion strategy, perform low-rank cross-modal attention calculation on low-dimensional features, and implement cache sharing for high-dimensional features;
[0009] Step 4: Allocate modal processing modules based on hardware characteristics and synchronize the fusion results through a high-speed bus.
[0010] Through the above technical solution, by quantifying the modal complexity, combining the Gumbel-Softmax sampling to reduce redundant calculations, calculating the low-rank attention through formulas, the computational complexity is greatly reduced, the memory occupancy is reduced through cache sharing, and the hardware utilization rate is greatly improved through heterogeneous allocation.
[0011] Further, the dynamic selection is implemented through a differentiable architecture search (DARTS) pre-trained decision network, and the low-rank cross-modal attention calculation adopts Tucker decomposition to decompose the original weight matrix into a low-rank core matrix and modal-specific factor matrices.
[0012] Through the above technical solutions, the computational complexity can be reduced while the memory occupation is decreased.
[0013] Furthermore, the modal complexity evaluation adopts the following specific formula:
[0014] Define the modal complexity scoring function:
[0015]
[0016] where L m represents the text length, the number of audio frames, and the image resolution, α and β are learnable weights, and the normalization coefficient L max , R max is the maximum value of the training set statistics.
[0017] Through the above technical solutions, redundant calculations are reduced and the computational efficiency is improved.
[0018] Furthermore, the hierarchical cross-modal fusion acceleration adopts the following formula:
[0019] Low-rank cross-modal attention:
[0020] Perform Tucker decomposition on the attention matrix:
[0021]
[0022] where is the core tensor, r i <<d 原维度 , is the modality-specific factor matrix, and the computational complexity of attention is reduced from O(d 2 ) to O(r1r2r3 + ∑dr i ).
[0023] Through the above technical solutions, the latency of the attention module is reduced and the memory occupation is decreased.
[0024] Furthermore, the device allocation strategy adopts the following specific formula:
[0025] Define the device allocation strategy:
[0026]
[0027] θ GPU = 2.0, θ NPU = 1.5;
[0028] Cross-device transmission time model:
[0029]
[0030] Minimize the total latency through the greedy algorithm: min∑(Tcomp +T comm )。
[0031] Through the above technical solutions, the hardware utilization rate is maximized, the end-to-end latency is reduced, the data throughput is increased, duplicate calculations are reduced, and the number of memory accesses is decreased.
[0032] The beneficial effects of the present invention are as follows: By quantifying the modal complexity, combining the Gumbel-Softmax sampling to reduce redundant calculations, calculating the low-order attention through formulas, the computational complexity is greatly reduced, the memory occupancy is reduced through cache sharing, and the hardware utilization rate is greatly improved through heterogeneous allocation. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0035] As Figure 1 shown, an inference acceleration method for a multi-modal large model in this embodiment includes the following specific steps:
[0036] Step 1: Perform feature analysis and modal complexity evaluation on the input multi-modal data;
[0037] Step 2: Dynamically select the calculation depth and number of parameters of the single-modal sub-network according to the complexity;
[0038] Step 3: Adopt a hierarchical fusion strategy to perform low-rank cross-modal attention calculation on low-dimensional features and cache sharing on high-dimensional features;
[0039] Step 4: Allocate modal processing modules based on hardware characteristics and synchronize the fusion results through a high-speed bus.
[0040] By quantifying the modal complexity, combining the Gumbel-Softmax sampling to reduce redundant calculations, calculating the low-order attention through formulas, the computational complexity is greatly reduced, the memory occupancy is reduced through cache sharing, and the hardware utilization rate is greatly improved through heterogeneous allocation.
[0041] The dynamic selection is realized through a differentiable architecture search (DARTS) pre-trained decision network, and the low-rank cross-modal attention calculation adopts Tucker decomposition to decompose the original weight matrix into a low-rank core matrix and modal-specific factor matrices.
[0042] The computational complexity can be reduced while reducing the memory footprint.
[0043] The modal complexity evaluation adopts the following specific formula:
[0044] Define the modal complexity scoring function:
[0045]
[0046] Among them, L m represents the text length, the number of audio frames, and the image resolution. α and β are learnable weights, and the normalization coefficient L max , R max is the maximum value of the training set statistics.
[0047] Reduce redundant calculations and improve computational efficiency.
[0048] The hierarchical cross-modal fusion acceleration adopts the following formula:
[0049] Low-rank cross-modal attention:
[0050] Perform Tucker decomposition on the attention matrix:
[0051]
[0052] Among them, is the core tensor, r i <<d 原维度 , is the modality-specific factor matrix, and the attention computational complexity is reduced from O(d 2 ) to O(r1r2r3 + ∑dr i ).
[0053] Reduce the attention module latency and reduce the memory footprint.
[0054] The device allocation strategy adopts the following specific formula:
[0055] Define the device allocation strategy:
[0056]
[0057] θ GPU = 2.0, θ NPU = 1.5;
[0058] Cross-device transmission time model:
[0059]
[0060] Minimize the total latency through the greedy algorithm: min∑(T comp + T comm ).
[0061] Maximize hardware utilization, reduce end-to-end latency, increase data throughput, reduce redundant calculations, and decrease the number of memory accesses.
[0062] Computational optimization in the video question answering scenario:
[0063] The number of input video frames N = 30, with a resolution of 1920×1080. By performing Tucker decomposition on the attention matrix, the sampling is reduced to 480p, and the computational complexity is reduced from O(N×1920×1080×C) to O(N×480×270×C'), where C = 3 (RGB channels) and C' = 16 (MobileNetV3 compressed channels).
[0064] Effect comparison table
[0065]
[0066] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention.
Claims
1. An inference acceleration method for a multi-modal large model, characterized in that, It includes the following specific steps: Step 1: Conduct feature analysis and modality complexity assessment on the input multi-modal data; Step 2: Dynamically select the computational depth and number of parameters of the single-modal sub-network according to the complexity; Step 3: Adopt a hierarchical fusion strategy to perform low-rank cross-modal attention calculation on low-dimensional features and cache sharing on high-dimensional features; Step 4: Allocate modality processing modules based on hardware characteristics and synchronize the fusion results through a high-speed bus.
2. The inference acceleration method for a multi-modal large model according to claim 1, wherein The dynamic selection is realized through a differentiable architecture search (DARTS) pre-trained decision network, and the low-rank cross-modal attention calculation uses Tucker decomposition to decompose the original weight matrix into a low-rank core matrix and modality-specific factor matrices.
3. A method for accelerating the inference of a multi-modal large model according to claim 1, characterized in that, The modality complexity assessment uses the following specific formula: Define a modality complexity scoring function: Among them, L m represents the text length, the number of audio frames, and the image resolution. α and β are learnable weights, and the normalization coefficient L max , R max is the maximum value of the training set statistics.
4. A method for accelerating the inference of a multi-modal large model according to claim 1, characterized in that The hierarchical cross-modal fusion acceleration uses the following formula: Low-rank cross-modal attention: Perform Tucker decomposition on the attention matrix: Among them, is the core tensor, r i <<d 原维度 , is the modality-specific factor matrix, and the attention calculation complexity is reduced from O(d 2 ) to O(r1r2r3 + ∑dr i ).
5. A method for accelerating the inference of a multi-modal large model according to claim 1, characterized in that, The device allocation strategy uses the following specific formula: Define a device allocation strategy: θ GPU = 2.0, θ NPU = 1.5; Cross-device transmission time model: Minimize the total latency through the greedy algorithm: min∑(T comp +T comm ).