Inference offload optimization method for edge-side multi-modal transformer model

CN117215795BActive Publication Date: 2026-09-25BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311263801.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-27
Publication Date
2026-09-25
Estimated Expiration
2043-09-27

AI Technical Summary

Technical Problem

然而,Transformer中独特的架构在推理卸载的过程中带来了一些新的挑战

Benefits of technology

[0065]本发明所述面向边缘侧多模态Transformer模型的推理卸载优化方法,通过利用在Transformer编码器内的多头注意力块中的多个注意力头的并行架构,将计算和传输并行,减少了基于Transfromer的多模态模型在边缘端异构设备推理卸载中由于其中不可分割计算节点引入的大量传输开销,从而最小化推理时间。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117215795B_ABST
    Figure CN117215795B_ABST
Patent Text Reader

Abstract

The application discloses an inference offloading optimization method for an edge-side multi-modal Transformer model. By utilizing the parallel architecture of multiple attention heads in the multi-head attention block in the Transformer encoder, the calculation and transmission are parallelized, and the large transmission overhead introduced by the indivisible calculation nodes in the multi-modal model based on the Transformer in the inference offloading of the edge-side heterogeneous device is reduced, so that the inference time is minimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of edge-side deep learning model technology, specifically involving an inference offloading optimization method for edge-side multimodal Transformer models. Background Technology

[0002] In real life, the information people encounter is mostly multimodal, such as text, images, sound, and video. For humans, understanding these different types of information and combining them to make decisions is natural, but for computer models, this is an extremely challenging task. The emergence of Transformer-based multimodal models (such as CLIP, ViLBERT, and VisualBERT) has enabled computer models to better understand and generate multimodal information, greatly expanding the application areas of artificial intelligence, such as visual question answering, automatic caption generation, social media analysis, online advertising, and autonomous driving.

[0003] While Transformer-based multimodal models offer numerous advantages and broad application potential, they typically involve large model sizes and high computational complexity. For some latency-critical applications running on resource-constrained edge devices, inference offloading optimization is still necessary to reduce inference latency while meeting memory constraints. However, the unique architecture of Transformers introduces new challenges during inference offloading.

[0004] Firstly, the Transformer contains numerous computational nodes that are difficult to split, leading to significant transmission overhead during inference offloading. These difficult-to-split nodes can only execute computations on a single device, making splitting and offloading impossible. In contrast, in traditional CNNs, all computational layers can typically be split and offloaded. During offloading, in CNNs, since most computational layers can be split, only necessary data needs to be transmitted after each inference layer. However, in the Transformer, these difficult-to-split nodes require synchronizing data from all devices onto a single node, resulting in substantial communication overhead.

[0005] Secondly, Transformer contains a large number of parallel architectures, such as independent and parallelizable attention heads, which is fundamentally different from CNN architecture. Existing CNN inference offloading methods do not consider this parallel architecture, and there is still a lot of room for optimization on this parallel architecture.

[0006] Therefore, traditional inference offloading methods suitable for CNNs are difficult to apply to inference offloading for Transformer-based multimodal models. For Transformers, existing optimization methods for inference offloading at the edge mainly rely on pipelined parallelism to parallelize communication and computation between different sample inferences; however, this method does not actually reduce the time for inference per sample. Some studies combine pipelined parallelism and tensor parallelism to reduce training time; however, these methods primarily target training and are performed in homogeneous, powerful GPU data centers.

[0007] Based on the aforementioned technical problems in the existing technology, this invention proposes an inference offloading optimization method for edge-side multimodal Transformer models. Summary of the Invention

[0008] The purpose of this invention is to address the shortcomings of existing technologies by synergistically optimizing model segmentation and system resource allocation for mobile computer vision inference tasks, and finding the optimal block combination among all blocks to ensure model inference accuracy while minimizing inference time.

[0009] The present invention adopts the following technical solution:

[0010] This invention provides an inference offloading optimization method for edge-side multimodal Transformer models, comprising:

[0011] Step 1: By analyzing the Transformer-based multimodal model, information such as the computation node graph, input / output shapes, and memory usage is obtained;

[0012] Step 2: Calculate the optimal segmentation and unloading strategy using the segmentation and unloading optimization algorithm;

[0013] Step 3: The computing tasks of each computing node are split and unloaded according to the split and unload strategy.

[0014] Furthermore, in step 1, the inference experiment includes preparation in the offline phase and inference in the inference phase based on the Transformer encoder in the Transformer multimodal model.

[0015] Further, step 2 includes: based on the computation node graph, the input and output shapes of each node, memory usage, and device resource information, obtaining the partitioning and offloading strategy of each node on each device that minimizes latency while satisfying the device memory constraint.

[0016] Further, step 1 includes:

[0017] Step 1.1: Configure the experiment, determine the Transformer-based multimodal model and its computation graph, and configure the equipment;

[0018] Step 1.2: Each device establishes a TCP long connection between each other through the communication module, and analyzes the uplink and downlink bandwidth between each pair;

[0019] Step 1.3: The main module notifies the offline analysis module of each device through the communication module to analyze the available memory of each device;

[0020] Step 1.4: The main module analyzes the multimodal model to obtain the computation node graph and its input and output shapes in the multi-head attention block of the Transformer encoder, as well as the memory usage of the weights corresponding to each node, and numbers the nodes.

[0021] Step 1.5: The main module notifies the offline analysis modules of each device through the communication module to analyze the corresponding mapping between the execution time and input of each node on the current device. The offline analysis modules of each device send the analyzed results to the main module through the communication module.

[0022] Furthermore, step 3 includes:

[0023] Step 3.1: The main module sends the segmentation and unloading strategy and the corresponding weights to the scheduling module of the corresponding device. Each device scheduling module creates multiple task modules according to the segmentation strategy.

[0024] Step 3.2: Load the weights of the corresponding nodes in the next multi-head self-attention block into the device's memory or video memory for each task module;

[0025] Step 3.3: The main module calculates the embedded part, and then sends the output of the embedded part to each device according to the segmentation and offloading strategy of the first node.

[0026] Step 3.4: Each device receives input data and sends the corresponding range of inputs to the task modules of all devices according to the segmentation and unloading strategy;

[0027] Step 3.5: After receiving the data, each task module puts it into the buffer inside the module. If the data in the buffer meets the input range specified for the current task module in the split unloading strategy, the executable flag is marked as True; otherwise, it is marked as False.

[0028] Step 3.6: Each device scheduling module determines the executable flag of the task module corresponding to the current computing node. If the executable flag is False, it waits for execution and returns to step 3.4. If the executable flag is True, it executes directly and then sends the output of the corresponding range to the corresponding task module of the corresponding device according to the segmentation strategy.

[0029] Step 3.7: Determine if there is a next node. If yes, set the next node as the current node and go to step 3.6. If not, continue to step 3.8.

[0030] Step 3.8, continue with step 3.4, until the current multi-head self-attention block has been executed on all nodes of all devices;

[0031] Step 3.9: Each device, according to the partitioning and unloading strategy of the first node on each device, sends the partial execution results of the last node in the computation graph on each device to the data within the required input range of the first node on each device to the corresponding device.

[0032] Step 3.10: Each task module loads the corresponding weights of the corresponding nodes in the next multi-head self-attention block and continues with step 1.11 until all multi-head self-attention blocks in the current Transformer encoder have been executed and the results are transmitted back to the main module.

[0033] Step 3.11: The main module continues to calculate the remaining part.

[0034] Further, step 2 includes:

[0035] Step 2.1: The partitioning and unloading optimization problem is formalized into a constrained optimization problem;

[0036] Step 2.2: Solve the problem using a genetic algorithm.

[0037] Furthermore, in step 2.1, each computing node generates a segmentation strategy for the output, with the segmentation dimension being the penultimate dimension of the computing node's output.

[0038] Furthermore, in step 2.1:

[0039] Suppose there are M computing nodes and D edge heterogeneous devices for partitioning and offloading. The size of the penultimate dimension of the output from the m-th computing node is H. m The output dimension of the m-th computation node is (H m W m The m-th computing node computes on the i-th device in the second-to-last dimension, where the range is... The output of the node is splitting strategy as follows:

[0040]

[0041] in:

[0042]

[0043]

[0044] m = 1, 2, ... M……(4);

[0045] If m is an indivisible node, then:

[0046]

[0047] The m-th computing node computes on the i-th device, with a range of [value] in the penultimate dimension. The output of , where i∈[1,D], let the set of nodes that depend on node m be m. next For each node h∈m next The segmented offloading range on the i-th device is For node h, the amount of data transmitted from the m-th computing node to the i-th device after the j-th device completes its computation is:

[0048]

[0049] Let a m,j,i,h If we want to indicate whether the m-th node, after completing its computation on the i-th device, should transmit data to the h-th computation node on the j-th device, then:

[0050]

[0051] Let w be the memory occupied by the m-th computing node. m The available memory on the i-th device is A. i ,have:

[0052]

[0053] in:

[0054]

[0055] Let T m,i This represents the time when the m-th node completes its computation on the i-th device, where all nodes that the m-th node depends on are m prev , p∈m prev Then the completion time of the m-th node on the i-th device is:

[0056]

[0057] Among them Bj,i Let T be the bandwidth between the j-th device and the i-th device. m-1,i Let a be the time when the next computing node begins after the previous node has been computed on device i. p,j,i,m ·(T p,j +p p,j,i,m / B j,i The time it takes for the nodes that the m-th computing node depends on to send data to the current i-th device after completion is reached. This represents the function relating the input of the m-th node analyzed in the offline phase to the computation time, i.e., the input range of the m-th node on the i-th device is... The time it took to input the data;

[0058] Send the results of the current multi-head attention block on each device to the necessary data that matches the input range of the first node of the next multi-head attention block on each device, and start the computation of the next multi-head attention block. The total time of the current multi-head attention block on the i-th device is:

[0059]

[0060] Completion time of the last node on each device:

[0061]

[0062] st(2)(3)(4)(5)(8).

[0063] Further, in step 2.2, the encoding method is set to real number encoding, the mutation operator is selected as differential mutation, the differential mutation scaling factor is set to 0.5, the recombination crossover method is selected as binomial distribution crossover, the crossover probability is set to 0.5, the fitness function is formula (12), the constraints are formulas (2)(3)(4)(5)(8), the population size is set to 100, and the solution is obtained through the standard genetic algorithm process.

[0064] The beneficial effects of this invention are:

[0065] The inference offloading optimization method for edge-side multimodal Transformer models described in this invention utilizes the parallel architecture of multiple attention heads in the multi-head attention block within the Transformer encoder to parallelize computation and transmission, thereby reducing the large transmission overhead introduced by the indivisible computation nodes in the inference offloading of Transformer-based multimodal models on heterogeneous devices at the edge and minimizing inference time. Attached Figure Description

[0066] Figure 1This is a schematic diagram of the Transformer encoder architecture in the Transformer multimodal model according to an embodiment of the present invention;

[0067] Figure 2 This is a schematic diagram of the Transformer encoder segmentation and unloading process in the Transformer multimodal model according to an embodiment of the present invention;

[0068] Figure 3 This is a schematic diagram of the system architecture described in an embodiment of the present invention;

[0069] Figure 4 This is a schematic diagram of the steps of the method described in the embodiment of the present invention. Detailed Implementation

[0070] To better understand the above-mentioned objectives, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.

[0071] Example

[0072] Unless otherwise defined, the technical or scientific terms used in the following embodiments shall have the ordinary meaning as understood by one of ordinary skill in the art to which this disclosure pertains.

[0073] in:

[0074] like Figure 1 As shown, the computational node: In this invention, each step of the computation in the multi-head attention block of the Transformer encoder is regarded as a computational node. Generally, all the rectangles in the multi-head attention block are computational nodes. In this invention, a multi-head attention block has eleven types of nodes: start, add, layer_norm, q, k, v, softmax, attn, concat, linear, and mlp. These are respectively the start input placeholder node, the addition node, the layer normalization node, the node for calculating the Q matrix, the node for calculating the K matrix, the node for calculating the V matrix, the node for calculating attention, the node for calculating the multiplication of attention and V, the combination node, the linear layer node, and the multilayer perceptron layer node.

[0075] Splitable Compute Node: A node in which a computation task can be split, including add, layer_norm, q, k, v, attn, linear, and mlp;

[0076] Parallel Nodes: Different nodes that can be placed on different devices and executed simultaneously, including different headers for q, k, v, softmax, and attn;

[0077] Non-splitable Compute Node: A node that can only be executed on one device, including softmax and concat;

[0078] Sequential Node: A node that can only be executed independently and cannot be executed in parallel with other nodes, including add, layernorm, concat, linear, and mlp;

[0079] Computational Graph: A computational graph describes the relationships between computational nodes. Each model corresponds to a directed acyclic computational graph. The vertices in the graph are computational nodes, and the edges in the graph indicate whether two nodes are dependent. If node B uses the output of node A as its input, then node B is said to be dependent on node A.

[0080] Edge: In this invention, the edge is a resource-constrained edge mobile device in the Internet of Things, used to execute specific sub-tasks;

[0081] Offline Stage: In this invention, the offline stage is mainly responsible for analyzing the time function of the computation nodes in the Transformer encoder, that is, analyzing the correlation function between the execution time and the input on various heterogeneous edge devices. Based on the optimization algorithm proposed in this invention, a specific task strategy is generated, and the relevant weights are sent to the corresponding devices for loading according to the task strategy.

[0082] Inference Stage: In this invention, the inference stage is mainly based on the segmentation and offloading strategy obtained in the offline stage. It performs calculations based on the given input data and the computation node graph of the Transformer encoder. During the calculation process, the computation tasks corresponding to each computation node are segmented and offloaded based on the available resources of the mobile edge device, thereby completing the computation task of the entire original model.

[0083] Experiment: One experiment in this invention includes the preparation and inference phases of the Transformer encoder in the offline phase based on the Transformer multimodal model.

[0084] Figure 2This describes a portion of the computational process in a Transformer-based multimodal encoder where one head of a multi-head self-attention block is split and offloaded across three devices (Raspberry Pi and two Jetson TX2s). The `layer_norm` node corresponds to the layer normalization operation, the `q` and `k` nodes linearly transform the input into Q and K, and the `softmax` node is an indivisible node that can only compute attention on one device using Q and K. The results computed on other devices must be transmitted to the device where the `softmax` node resides, which introduces significant communication overhead. Each solid rectangle marked with a number represents input or output data, while the dashed rectangle marked with a number represents data that is missing from the current input and needs to be transmitted from other devices to this device.

[0085] The segmentation offloading algorithm described in this embodiment firstly segments the input into three parts according to the offloading strategy and transmits them to three devices respectively; secondly, it performs layer normalization through the layer_norm node; thirdly, it calculates Q through node 1 after the layer normalization (for simplicity, the segmentation offloading strategy for q and k followed in this embodiment is consistent with layer_norm, and no transmission is required); fourthly, it calculates K through node k1 and transmits Q calculated in the previous step to the TX2 device in the middle (only the Raspberry Pi and the Jetson TX2 device on the right need to be transmitted); fifthly, it transmits K to the Jetson TX2 device in the middle; sixthly, it calculates attention on the Jetson TX2 device in the middle, and then the subsequent nodes follow the same principle.

[0086] The fourth step embodies the parallelism of computation and transmission. It is just one node. There are a large number of parallel nodes in the multi-head attention block, such as q, k, v inside all heads, and any two nodes between different heads. Any parallel node can hide the transmission overhead in this way, thereby reducing the overall communication overhead during the segmentation and unloading process, and thus shortening the overall latency of segmentation and unloading.

[0087] Specifically, the inference offloading optimization method for edge-side multimodal Transformer models includes:

[0088] Step 1: By analyzing the Transformer-based multimodal model, information such as its computational node graph, input / output shape, and memory usage is obtained;

[0089] Step 1.1: Configure the experiment, determine the Transformer-based multimodal model and its computation graph, and configure the equipment;

[0090] Step 1.2: Each device establishes a TCP long connection between each other through the communication module, and analyzes the uplink and downlink bandwidth between each pair;

[0091] Step 1.3: The main module notifies the offline analysis module of each device through the communication module to analyze the available memory of each device;

[0092] Step 1.4: The main module analyzes the multimodal model to obtain the computation node graph and its input and output shapes in the multi-head attention block of the Transformer encoder, as well as the memory usage of the weights corresponding to each node, and numbers the nodes.

[0093] Step 1.5: The main module notifies the offline analysis modules of each device through the communication module to analyze the corresponding mapping between the execution time and input of each node on the current device. The offline analysis modules of each device send the analyzed results to the main module through the communication module.

[0094] Step 2: Calculate the optimal segmentation and unloading strategy using the segmentation and unloading optimization algorithm;

[0095] Step 2.1: Formalize the scenario as a constrained optimization problem. Specifically, assume there are M computing nodes and D edge heterogeneous devices for segmentation and offloading. The size of the penultimate dimension of the output from the m-th computing node is H. m The node splitting strategy is as follows:

[0096]

[0097] in:

[0098]

[0099]

[0100] m = 1, 2, ... M……(4);

[0101] If m is an indivisible node, then:

[0102]

[0103] The m-th computing node computes on the i-th device, with a range of [value] in the penultimate dimension. The output of , where i∈[1,D], let the set of nodes that depend on node m be m. next For each node h∈m next The segmented offloading range on the i-th device is For node h, the amount of data that the m-th computing node needs to transmit to the i-th device after completing the computation on the j-th device is:

[0104]

[0105] Let a m,j,i,h If we want to indicate whether the m-th node, after completing its computation on the i-th device, should transmit data to the h-th computation node on the j-th device, then:

[0106]

[0107] Let w be the memory occupied by the m-th computing node. m The available memory on the i-th device is A. i ,have:

[0108]

[0109] in:

[0110]

[0111] Let T m,i This represents the time when the m-th node completes its computation on the i-th device, where all nodes that the m-th node depends on are m prev , p∈m prev Then the completion time of the m-th node on the i-th device is:

[0112]

[0113] Among them B j,i Let T be the bandwidth between the j-th device and the i-th device. m-1,i Let a be the time when the previous node has just been computed on device i and the next node is being computed. p,j,i,m ·(T p,j +p p,j,i,m / B j,i () represents the time when the nodes dependent on the m-th computing node finish sending data to the current i-th device after completion. This represents the function relating the input of the m-th node analyzed in the offline phase to the computation time, i.e., the input range of the m-th node on the i-th device is... The time it took to input the data;

[0114] The computation of the next multi-head attention block can only begin after the results of the current multi-head attention block on each device are sent to the necessary data that conforms to the input range of the first node of the next multi-head attention block on each device. That is, the total time of the current multi-head attention block on the i-th device is:

[0115]

[0116] Completion time of the last node on each device:

[0117]

[0118] st(2)(3)(4)(5)(8);

[0119] Step 2.2: Solve the problem using a genetic algorithm. Set the encoding method to real number encoding, select differential mutation as the mutation operator, set the differential mutation scaling factor to 0.5, select binomial crossover as the recombination and crossover method, set the crossover probability to 0.5, use formula (12) as the fitness function, use formulas (2)(3)(4)(5)(8) as the constraints, set the population size to 100, and solve the problem using the standard genetic algorithm process.

[0120] Step 3: The computing tasks of each computing node are divided and unloaded for execution according to the division and unloading strategy;

[0121] Step 3.1: The main module sends the segmentation and unloading strategy and the corresponding weights to the scheduling module of the corresponding device. Each device scheduling module creates multiple task modules according to the segmentation strategy.

[0122] Step 3.2: Load the weights of the corresponding nodes in the next multi-head self-attention block into the device's memory or video memory for each task module;

[0123] Step 3.3: The main module calculates the embedded part, and then sends the output of the embedded part to each device according to the segmentation and offloading strategy of the first node.

[0124] Step 3.4: Each device receives input data and sends the corresponding range of inputs to the task modules of all devices according to the segmentation and unloading strategy;

[0125] Step 3.5: After receiving the data, each task module puts it into the buffer inside the module. If the data in the buffer meets the input range specified for the current task module in the split unloading strategy, the executable flag is marked as True; otherwise, it is marked as False.

[0126] Step 3.6: Each device scheduling module determines the executable flag of the task module corresponding to the current computing node. If the executable flag is False, it waits for execution and returns to step 3.4. If the executable flag is True, it executes directly and then sends the output of the corresponding range to the corresponding task module of the corresponding device according to the segmentation strategy.

[0127] Step 3.7: Determine if there is a next node. If yes, set the next node as the current node and go to step 3.6. If not, continue to step 3.8.

[0128] Step 3.8, continue with step 3.4, until the current multi-head self-attention block has been executed on all nodes of all devices;

[0129] Step 3.9: According to the partitioning and unloading strategy of the first node on each device, each device sends the partial execution results of the last node in the computation graph (if it has been unloaded to the device) on each device to the data of the first node in each device within the required input range to the corresponding device.

[0130] Step 3.10: Each task module loads the corresponding weights of the corresponding nodes in the next multi-head self-attention block and continues with step 1.11 until all multi-head self-attention blocks in the current Transformer encoder have been executed and the results are transmitted back to the main module.

[0131] Step 3.11: The main module continues to calculate the remaining part.

[0132] In step 1 of the above embodiment, the inference experiment includes preparation in the offline phase and inference in the inference phase based on the Transformer encoder in the Transformer multimodal model.

[0133] In step 2.1 of the above embodiment, each computing node generates a segmentation strategy for the output, with the segmentation dimension being the penultimate dimension of the computing node's output.

[0134] like Figure 1 and Figure 3 As shown, this is an inference offloading optimization system for edge-side multimodal Transformer models, including:

[0135] The computation node module corresponds to each computation node in the Transformer Encoder computation graph and is used to perform calculations based on the input (divided into 10 different computation node sub-modules: start, layer_norm, q, k, softmax, v, attn, concat, linear, mlp, and add; the start node is just a placeholder node representing the initial input).

[0136] The task module is responsible for receiving data for a single task on the device, performing calculations through the internal computing node module, and sending the results.

[0137] The scheduling module is used to receive policies on the device, create all task modules, and schedule their execution.

[0138] The communication module is used for data transmission with other devices;

[0139] The offline analysis module is used to analyze the correspondence between the execution time of each computing node on the current device and the input.

[0140] The main module is used to call the offline analysis modules on each device to perform analysis during the offline phase, collect all analysis results, generate the segmentation strategy for each computing node through optimization algorithms based on the analysis results, and send the corresponding part to the scheduling module of each device.

[0141] In this embodiment, an edge-side inference offloading optimization device is disclosed, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements an edge-side multimodal Transformer model inference offloading optimization method.

[0142] In this embodiment, a computer-readable storage medium is disclosed, which stores a computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute an inference offloading optimization method for an edge-side multimodal Transformer model.

[0143] like Figure 4 As shown, given specific data, the edge-oriented inference offloading optimization method described in this embodiment is further illustrated. Assuming an inference experiment is conducted, the initial configuration includes: a Transformer-based multimodal model, test data, and a set of edge devices. The method includes:

[0144] Step 101: Configure the experiment. The goal is to use a CLIP (ViT-L / 14) visual Transformer encoder based on a multimodal Transformer model. Test data will be a single image of size [1, 3, 224, 224] and all possible categories (e.g., a photo of a bird) containing a fixed cue word. There are three edge mobile devices configured as follows:

[0145] Two NVIDIA Jetson TX2 processors, each with an ARM Cortex-A57 @ 2.0GHz CPU and 8GB of RAM; one Raspberry Pi 4B processor, each with a Quad core Cortex-A72 (ARM v8) 64-bit @ 1.5GHz CPU and 4GB of RAM. The edge devices are numbered as follows: 0 is the Raspberry Pi 4B, and 1 and 2 are the NVIDIA Jetson TX2 processors.

[0146] Step 102: Enable the communication module on the three edge devices, establish TCP long connections between each pair of the three devices, and test the bandwidth between each pair of the three edge devices by sending test data to confirm that the bandwidth is 1000Mbps.

[0147] Step 103: The main module obtains the available memory status of each edge device [101,99,50];

[0148] Step 104: Input CLIP (ViT-L / 14) into the main module for analysis to obtain its computation graph, the dependencies between nodes, and the memory usage of each node. Its visual Transformer encoder has 16 attention heads and a total of 87 computation nodes.

[0149] Node 0 is the start node, node 1 is the layer_norm node, nodes 2, 3, 4, 5, and 6 are the q, k, softmax, v, and attn of the first attention head, respectively, nodes 7-81 are the nodes corresponding to the remaining 15 attention heads, with every 5 nodes corresponding to one attention head, node 82 is a concat calculation node, node 83 is a linear node, node 84 is an add node, node 85 is a layer_norm node, node 86 is an MLP node, and node 87 is an add node.

[0150] The 0th node has no dependencies, the 1st node has no dependencies, the 2nd, 3rd, and 5th nodes depend on the 1st node, the 4th node depends on the 2nd and 3rd nodes, the 6th node depends on the 5th node, the dependencies between every 5 nodes from the 7th to the 81st node are the same as those between the 2nd and 6th nodes, the 82nd node depends on the [6,11,16,21,26,31,36,41,46,51,56,61,66,71,76,81]th node, the 83rd node depends on the 82nd node, the 84th node depends on the 83rd node and the 0th node, the 85th node depends on the 84th node, the 86th node depends on the 85th node, and the 87th node depends on the 85th and 86th nodes.

[0151] The image is input into the visual Transformer encoder, and the output shape of each computation node is analyzed. The input and output shapes of the 0th and 1st nodes are [257, 1024]. The input and output ranges of every 5 nodes from the 2nd to the 81st nodes are [257, 64], [257, 64], [257, 257], [257, 64], [257, 64], and the output range of the 82nd to the 87th nodes is [257, 1024].

[0152] The corresponding memory sizes are as follows: 0 and 1 nodes occupy 0, 2 and 3 nodes occupy 0.125M, 4th node occupies 0M, 5th node occupies 0.125M, 6th node occupies 0M, and for nodes 7-81, the memory size of every 5 nodes is the same as that of nodes 2-6. 82nd, 84th, 85th, and 87th nodes occupy 0M, 83rd node occupies 2M, and 86th node occupies 16M.

[0153] Step 105: The main module is placed on the first edge device. The average computation time of each of the seven types of nodes on the three edge devices is obtained by performing 100 calculations on each input range using random data within the maximum input range of the node. This is to obtain the mapping between all input ranges and their execution time. Then, the results on the second and third devices are transmitted to the first edge device.

[0154] Step 106: Calculate the segmentation and offloading strategy for each node on each device using an optimization algorithm. The 0th and 1st nodes are [0,0,144,257], the 2nd node is [0,35,143,257], the 3rd node is [0,143,143,257], the 4th node is [0,35,143,257], ..., and the 87th node is [0,35,143,257].

[0155] Step 107: The main module sends the weights of the corresponding nodes in the first multi-head self-attention block to each device through the communication module, sends the weights of the [2,3,…]th node in each multi-head attention block to the first device, sends the weights of the [1,2,4,…]th node in each multi-head attention block to the second device, sends the weights of the [1,2,3,…]th node in each multi-head attention block to the third device, and then, according to the task module that creates the corresponding node;

[0156] Step 108: Load the weights of the corresponding computing nodes in the first multi-head attention block through the task modules on each device. Load the weights into memory for device 1, and load them into video memory for devices 2 and 3. Set the task module corresponding to computing node 1 as the current task module.

[0157] Step 109: Calculate the embedded part of the input in the main module, and send the [0,144) part of the result to the task module corresponding to the first node on the second device, and send the [144,257) part of the result to the task module corresponding to the first node on the third device;

[0158] Step 110: After receiving the data in the [0,144) part, device 2 sends the data to the task module corresponding to computing node 1. After receiving the data in the [144,257) part, device 3 sends the data to the task module corresponding to computing node 1.

[0159] Step 111: The task module corresponding to computing node 1 on device 2 puts the received data into the buffer and sets the executable flag to True. The task module corresponding to the first node in device 3 puts the received data into the buffer and sets the executable flag to True.

[0160] Step 111: Devices 2 and 3 both start executing the first node. Upon finding the executable flag to be True, they execute immediately. After the second device finishes executing the first node, it sets the executable flag of the second node on its device to True and sets the current task module to the task module corresponding to the second node. It transmits the output of part [0,35) to the first device and sends [143,144) to the third device. While transmitting, it executes the task module corresponding to the second node. After the third device finishes executing the first node, it sets the task module corresponding to the second node as the current task module of its device. This task module receives the data of part [143,144) transmitted from the second device, sets the executable flag to True, and then executes the second node. This process continues until all three devices have executed the 87th node.

[0161] Step 112: The three devices begin loading the weights of the second multi-head attention block into the computing node modules of the corresponding task modules, while releasing the previous weights;

[0162] Step 113: Begin reasoning for the second multi-head attention block, and so on, until all multi-head attention blocks have been reasoned out;

[0163] Step 114: After the first device finishes executing node 87, it sends the output data of part [0,35) to the second device. After the third device finishes executing node 87, it sends the data of part [143,144) to the second device.

[0164] Step 115: Transfer the results of the last multi-head attention block on the three devices to the main module to perform the remaining calculations.

[0165] This invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the invention as claimed. The scope of protection of this invention is defined by the appended claims.

Claims

1. A method for inference offloading optimization of edge-side multimodal Transformer models, characterized in that, include: Step 1: By analyzing the Transformer-based multimodal model, information such as the computation node graph, input / output shapes, and memory usage is obtained; Step 2: Calculate the optimal segmentation and unloading strategy using the segmentation and unloading optimization algorithm; Step 3 utilizes the parallel architecture of multiple attention heads in the multi-head attention block within the Transformer encoder to parallelize computation and transmission. The computational tasks of each computing node are then segmented and offloaded according to this segmentation and offloading strategy, including: Step 3.1: The main module sends the segmentation and unloading strategy and the corresponding weights to the scheduling module of the corresponding device. Each device scheduling module creates multiple task modules according to the segmentation strategy. Step 3.2: Load the weights of the corresponding nodes in the next multi-head self-attention block into the device's memory or video memory for each task module; Step 3.3: The main module calculates the embedded part, and then sends the output of the embedded part to each device according to the segmentation and offloading strategy of the first node. Step 3.4: Each device receives input data and sends the corresponding range of inputs to the task modules of all devices according to the segmentation and unloading strategy; Step 3.5: After receiving the data, each task module puts it into the buffer inside the module. If the data in the buffer meets the input range specified for the current task module in the split unloading strategy, the executable flag is marked as True; otherwise, it is marked as False. Step 3.6: Each device scheduling module determines the executable flag of the task module corresponding to the current computing node. If the executable flag is False, it waits for execution and returns to step 3.

4. If the executable flag is True, it executes directly and then sends the output of the corresponding range to the corresponding task module of the corresponding device according to the segmentation strategy. Step 3.7: Determine if there is a next node. If yes, set the next node as the current node and go to step 3.

6. If not, continue to step 3.

8. Step 3.8, continue with step 3.4, until the current multi-head self-attention block has been executed on all nodes of all devices; Step 3.9: Each device, according to the partitioning and unloading strategy of the first node on each device, sends the partial execution results of the last node in the computation graph on each device to the data within the required input range of the first node on each device to the corresponding device. Step 3.10: Each task module loads the corresponding weights of the corresponding nodes in the next multi-head self-attention block, and continues with step 3.11 until all multi-head self-attention blocks in the current Transformer encoder have been executed, and transmits the results back to the main module. Step 3.11: The main module continues to calculate the remaining part.

2. The inference offloading optimization method for edge-side multimodal Transformer models according to claim 1, characterized in that, Step 1 includes: Step 1.1: Configure the experiment, determine the Transformer-based multimodal model and its computation graph, and configure the equipment; Step 1.2: Each device establishes a TCP long connection between each other through the communication module, and analyzes the uplink and downlink bandwidth between each pair; Step 1.3: The main module notifies the offline analysis module of each device through the communication module to analyze the available memory of each device; Step 1.4: The main module analyzes the multimodal model to obtain the computation node graph and its input and output shapes in the multi-head attention block of the Transformer encoder, as well as the memory usage of the weights corresponding to each node, and numbers the nodes. Step 1.5: The main module notifies the offline analysis modules of each device through the communication module to analyze the corresponding mapping between the execution time and input of each node on the current device. The offline analysis modules of each device send the analyzed results to the main module through the communication module.

3. The inference offloading optimization method for edge-side multimodal Transformer models according to claim 1, characterized in that, Step 2 includes: Step 2.1: The partitioning and unloading optimization problem is formalized into a constrained optimization problem; Step 2.2: Solve the problem using a genetic algorithm.

4. The inference offloading optimization method for edge-side multimodal Transformer models according to claim 3, characterized in that, In step 2.1, each computing node generates a segmentation strategy for the output, with the segmentation dimension being the penultimate dimension of the computing node's output.

5. The inference offloading optimization method for edge-side multimodal Transformer models according to claim 3 or 4, characterized in that, In step 2.1: Suppose there are M computing nodes and D edge heterogeneous devices for partitioning and offloading. The size of the second-to-last dimension output by the m-th computing node is... The output dimension of the m-th computation node is The m-th computing node computes on the i-th device in the second-to-last dimension with a range of... The output of the node is splitting strategy as follows: ……(1); in: ……(2); ……(3); ……(4); If m is an indivisible node, then: ……(5); The m-th computing node computes on the i-th device, with a range of [value] in the penultimate dimension. The output of, where Let the set of nodes that depend on node m be... For each node The segmented offloading range on the i-th device is For node h, the amount of data transmitted from the m-th computing node to the i-th device after the j-th device completes its computation is: ……(6); make If we want to indicate whether the m-th node, after completing its computation on the i-th device, should transmit data to the h-th computation node on the j-th device, then: ……(7); Let the memory occupied by the m-th computing node be... The available memory on the i-th device is ,have: ……(8); in: ……(9); make This represents the time when the m-th node completes its computation on the i-th device, where all nodes that the m-th node depends on are... , Then the completion time of the m-th node on the i-th device is: ……(10); in Let be the bandwidth between the j-th device and the i-th device. This is the time when the computation of the previous node on device i has just finished and the computation of the next node begins. Let be the time when the nodes dependent on the m-th computing node finish sending data to the current i-th device. This represents the function relating the input of the m-th node analyzed in the offline phase to the computation time, i.e., the input range of the m-th node on the i-th device is... The time it took to input the data; Send the results of the current multi-head attention block on each device to the necessary data that matches the input range of the first node of the next multi-head attention block on each device, and start the computation of the next multi-head attention block. The total time of the current multi-head attention block on the i-th device is: ……(11); This represents the time when the Mth node completes its computation on the i-th device, and the completion time of the last node on each device: ……(12); st (2) (3) (4) (5) (8).

6. The inference offloading optimization method for edge-side multimodal Transformer models according to claim 4, characterized in that, In step 2.2, the encoding method is set to real number encoding, the mutation operator is selected as differential mutation, the differential mutation scaling factor is set to 0.5, the recombination crossover method is selected as binomial distribution crossover, the crossover probability is set to 0.5, the fitness function is formula (12), the constraints are formula (2) (3) (4) (5) (8), the population size is set to 100, and the solution is obtained through the genetic algorithm process.

Citation Information

Patent Citations

  • Load decomposition method and system based on deep learning attention mechanism

    CN115659148A

  • Transform model segmentation method and device in mobile edge environment

    CN116033492A