Multi-modal large model reasoning acceleration method

By performing feature analysis and modal complexity evaluation on multimodal large models, dynamically selecting the calculation depth and parameter quantity, using hierarchical fusion strategy and low-rank cross-modal attention calculation, the problems of wasted resources and increased delays in multimodal large models are solved, and more efficient hardware utilization and data throughput are achieved.

CN120278264APending Publication Date: 2025-07-08SHANGHAI XIAOGONGYI E-COMMERCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510339921.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing multimodal large models waste serious computing resources and increase delays during inference. The existing technology has failed to effectively solve the problems of high computational complexity and high delays.

Method used

By performing feature analysis and modal complexity evaluation of the input multimodal data, the calculation depth and parameter quantity of the single-modal subnet are dynamically selected, and a hierarchical fusion strategy and low-rank cross-modal attention calculation are adopted, and the modal processing module is allocated in combination with hardware characteristics, and the results are synchronized through high-speed bus.

Benefits of technology

Reduces computational complexity, reduces memory usage, improves hardware utilization, reduces latency and redundant computing, and improves data throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278264A_ABST
    Figure CN120278264A_ABST
Patent Text Reader

Abstract

The invention discloses a reasoning acceleration method for a multi-modal large model, and belongs to the technical field of artificial intelligence, the reasoning acceleration method for the multi-modal large model comprises the following specific steps: step 1, carrying out feature analysis and modal complexity evaluation on input multi-modal data; 2, dynamically selecting the calculation depth and parameter quantity of the single-mode sub-network according to the complexity; step 3, implementing low-rank cross-modal attention calculation on the low-dimensional features by adopting a hierarchical fusion strategy, and implementing cache sharing on the high-dimensional features; and 4, distributing modal processing modules based on hardware characteristics, and synchronizing a fusion result through a high-speed bus. According to the method, the complexity of the modal is quantified, the redundancy calculation is reduced in combination with Gumb l-Softmax sampling, the low-order attention is calculated through a formula, the calculation complexity is greatly reduced, the memory occupation is reduced through cache sharing, and the hardware utilization rate is greatly improved through heterogeneous allocation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a method for accelerating the inference of a multimodal large model. Background Art

[0002] A multimodal large model is an artificial intelligence model that can simultaneously process multiple different types of data (such as text, images, voice, video, sensor data, etc.). Through deep neural network technology, it fuses and correlates information of different modalities, thereby achieving a more comprehensive and intelligent understanding and generation ability. The method for accelerating the inference of a multimodal large model aims to solve the problems of high computational complexity and large latency, and improve the response speed of the model in actual scenarios through technical optimization.

[0003] However, existing multimodal models use the same computational path for all input data during inference, resulting in waste of computational resources; or the fusion strategy is too complex, resulting in increased latency. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a method for accelerating the inference of a multimodal large model.

[0005] The technical solution adopted to solve the above technical problem is: A method for accelerating the inference of a multimodal large model, including the following specific steps:

[0006] Step 1: Perform feature analysis and modal complexity evaluation on the input multimodal data;

[0007] Step 2: Dynamically select the computational depth and the number of parameters of the unimodal sub-network according to the complexity;

[0008] Step 3: Adopt a hierarchical fusion strategy, perform low-rank cross-modal attention calculation on low-dimensional features, and implement cache sharing for high-dimensional features;

[0009] Step 4: Allocate modal processing modules based on hardware characteristics and synchronize the fusion results through a high-speed bus.

[0010] Through the above technical solution, by quantifying the modal complexity, combining the Gumbel-Softmax sampling to reduce redundant calculations, calculating the low-rank attention through formulas, the computational complexity is greatly reduced, the memory occupancy is reduced through cache sharing, and the hardware utilization rate is greatly improved through heterogeneous allocation.

[0011] Further, the dynamic selection is implemented through a differentiable architecture search (DARTS) pre-trained decision network, and the low-rank cross-modal attention calculation adopts Tucker decomposition to decompose the original weight matrix into a low-rank core matrix and modal-specific factor matrices.

[0012] Through the above technical solutions, the computational complexity can be reduced while the memory occupation is decreased.

[0013] Furthermore, the modal complexity evaluation adopts the following specific formula:

[0014] Define the modal complexity scoring function:

[0015]

[0016] where L m represents the text length, the number of audio frames, and the image resolution, α and β are learnable weights, and the normalization coefficient L max , R max is the maximum value of the training set statistics.

[0017] Through the above technical solutions, redundant calculations are reduced and the computational efficiency is improved.

[0018] Furthermore, the hierarchical cross-modal fusion acceleration adopts the following formula:

[0019] Low-rank cross-modal attention:

[0020] Perform Tucker decomposition on the attention matrix:

[0021]

[0022] where is the core tensor, r i <<d 原维度 , is the modality-specific factor matrix, and the computational complexity of attention is reduced from O(d 2 ) to O(r1r2r3 + ∑dr i ).

[0023] Through the above technical solutions, the latency of the attention module is reduced and the memory occupation is decreased.

[0024] Furthermore, the device allocation strategy adopts the following specific formula:

[0025] Define the device allocation strategy:

[0026]

[0027] θ GPU = 2.0, θ NPU = 1.5;

[0028] Cross-device transmission time model:

[0029]

[0030] Minimize the total latency through the greedy algorithm: min∑(Tcomp +T comm )。

[0031] Through the above technical solutions, the hardware utilization rate is maximized, the end-to-end latency is reduced, the data throughput is increased, duplicate calculations are reduced, and the number of memory accesses is decreased.

[0032] The beneficial effects of the present invention are as follows: By quantifying the modal complexity, combining the Gumbel-Softmax sampling to reduce redundant calculations, calculating the low-order attention through formulas, the computational complexity is greatly reduced, the memory occupancy is reduced through cache sharing, and the hardware utilization rate is greatly improved through heterogeneous allocation. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0035] As Figure 1 shown, an inference acceleration method for a multi-modal large model in this embodiment includes the following specific steps:

[0036] Step 1: Perform feature analysis and modal complexity evaluation on the input multi-modal data;

[0037] Step 2: Dynamically select the calculation depth and number of parameters of the single-modal sub-network according to the complexity;

[0038] Step 3: Adopt a hierarchical fusion strategy to perform low-rank cross-modal attention calculation on low-dimensional features and cache sharing on high-dimensional features;

[0039] Step 4: Allocate modal processing modules based on hardware characteristics and synchronize the fusion results through a high-speed bus.

[0040] By quantifying the modal complexity, combining the Gumbel-Softmax sampling to reduce redundant calculations, calculating the low-order attention through formulas, the computational complexity is greatly reduced, the memory occupancy is reduced through cache sharing, and the hardware utilization rate is greatly improved through heterogeneous allocation.

[0041] The dynamic selection is realized through a differentiable architecture search (DARTS) pre-trained decision network, and the low-rank cross-modal attention calculation adopts Tucker decomposition to decompose the original weight matrix into a low-rank core matrix and modal-specific factor matrices.

[0042] The computational complexity can be reduced while reducing the memory footprint.

[0043] The modal complexity evaluation adopts the following specific formula:

[0044] Define the modal complexity scoring function:

[0045]

[0046] Among them, L m represents the text length, the number of audio frames, and the image resolution. α and β are learnable weights, and the normalization coefficient L max , R max is the maximum value of the training set statistics.

[0047] Reduce redundant calculations and improve computational efficiency.

[0048] The hierarchical cross-modal fusion acceleration adopts the following formula:

[0049] Low-rank cross-modal attention:

[0050] Perform Tucker decomposition on the attention matrix:

[0051]

[0052] Among them, is the core tensor, r i <<d 原维度 , is the modality-specific factor matrix, and the attention computational complexity is reduced from O(d 2 ) to O(r1r2r3 + ∑dr i ).

[0053] Reduce the attention module latency and reduce the memory footprint.

[0054] The device allocation strategy adopts the following specific formula:

[0055] Define the device allocation strategy:

[0056]

[0057] θ GPU = 2.0, θ NPU = 1.5;

[0058] Cross-device transmission time model:

[0059]

[0060] Minimize the total latency through the greedy algorithm: min∑(T comp + T comm ).

[0061] Maximize hardware utilization, reduce end-to-end latency, increase data throughput, reduce redundant calculations, and decrease the number of memory accesses.

[0062] Computational optimization in the video question answering scenario:

[0063] The number of input video frames N = 30, with a resolution of 1920×1080. By performing Tucker decomposition on the attention matrix, the sampling is reduced to 480p, and the computational complexity is reduced from O(N×1920×1080×C) to O(N×480×270×C'), where C = 3 (RGB channels) and C' = 16 (MobileNetV3 compressed channels).

[0064] Effect comparison table

[0065]

[0066] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention.

Claims

1. An inference acceleration method for a multi-modal large model, characterized in that, It includes the following specific steps: Step 1: Conduct feature analysis and modality complexity assessment on the input multi-modal data; Step 2: Dynamically select the computational depth and number of parameters of the single-modal sub-network according to the complexity; Step 3: Adopt a hierarchical fusion strategy to perform low-rank cross-modal attention calculation on low-dimensional features and cache sharing on high-dimensional features; Step 4: Allocate modality processing modules based on hardware characteristics and synchronize the fusion results through a high-speed bus.

2. The inference acceleration method for a multi-modal large model according to claim 1, wherein The dynamic selection is realized through a differentiable architecture search (DARTS) pre-trained decision network, and the low-rank cross-modal attention calculation uses Tucker decomposition to decompose the original weight matrix into a low-rank core matrix and modality-specific factor matrices.

3. A method for accelerating the inference of a multi-modal large model according to claim 1, characterized in that, The modality complexity assessment uses the following specific formula: Define a modality complexity scoring function: Among them, L m represents the text length, the number of audio frames, and the image resolution. α and β are learnable weights, and the normalization coefficient L max , R max is the maximum value of the training set statistics.

4. A method for accelerating the inference of a multi-modal large model according to claim 1, characterized in that The hierarchical cross-modal fusion acceleration uses the following formula: Low-rank cross-modal attention: Perform Tucker decomposition on the attention matrix: Among them, is the core tensor, r i <<d 原维度 , is the modality-specific factor matrix, and the attention calculation complexity is reduced from O(d 2 ) to O(r1r2r3 + ∑dr i ).

5. A method for accelerating the inference of a multi-modal large model according to claim 1, characterized in that, The device allocation strategy uses the following specific formula: Define a device allocation strategy: θ GPU = 2.0, θ NPU = 1.5; Cross-device transmission time model: Minimize the total latency through the greedy algorithm: min∑(T comp +T comm ).