A large model perception segmentation multi-terminal device cooperative reasoning method and system

CN121365740BActive Publication Date: 2026-09-29CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511607599.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-09-29
Estimated Expiration
2045-11-05

AI Technical Summary

Technical Problem

[0006]本发明提供了一种大模型感知切分的多终端设备协同推理方法及系统,以解决现有的协同推理框架存在定位的问题

Benefits of technology

本发明提供的大模型感知切分的多终端设备协同推理方法,通过算子感知的细粒度划分策略,针对不同自注意力架构MHA、MQA、GQA,以及前馈网络(FFN)和输出投影等不同模块的计算–通信特性,自动选择最优的行列或词表维度切分方式,并结合设备异构性进行负载均衡映射;同时,通过同步开销最小化执行机制,构建轻量子图、采用主机中心化聚合协调、并深度融合计算–通信重叠与动态负载调度,显著隐藏跨设备通信延迟。本发明不仅支持超出单设备内存容量的大模型协同推理,还在真实边缘集群中实现高达5.4倍的推理加速,显著优于传统单机推理或管道并行基线,适用于隐私敏感、低延迟要求的本地化智能助手、AR/VR交互、边缘AI服务等实际应用场景,使百亿参数级大语言模型真正具备在云外边缘环境高效运行的能力。能够有效解决现有并行推理方案在边缘场景下面临的同步开销大、通信冗余高、算子划分粗粒度等核心问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121365740B_ABST
    Figure CN121365740B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of model reasoning, and discloses a large model perception segmentation multi-terminal device cooperative reasoning method, which comprises the following steps: performing fine-grained analysis on the calculation graph of a target large language model, identifying the types of various operators and the calculation-communication characteristics thereof; automatically generating a differentiated weight division strategy according to the types of the operators, memory occupation, calculation intensity and attention mechanism variants; mapping the divided weight fragments and calculation tasks to multiple edge devices to construct minimum execution subgraphs of the devices; in the reasoning process, a host centralized aggregation coordination and calculation-communication overlap mechanism is adopted to minimize the synchronization cost; self-recurrence reasoning is performed, output is generated token by token, and the subsequent reasoning efficiency is optimized through a dynamic weight caching mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of model reasoning technology, and in particular to a multi-terminal device collaborative reasoning method and system for large model perception segmentation. Background Technology

[0002] Large language models, due to their powerful language understanding and generation capabilities, are gradually being deployed to edge devices to support local intelligent applications with privacy protection and low-latency response. However, limited by the limited memory and computing resources of edge devices such as mobile phones, embedded development boards, and IoT nodes, current mainstream large models are difficult to run efficiently on a single device, especially for models with billions of parameters, which often cannot be loaded due to insufficient memory.

[0003] To overcome the resource bottleneck of single devices, researchers have attempted to collaborate multiple edge devices to jointly execute LLM inference tasks. Currently, mainstream collaborative inference techniques mainly draw on parallel strategies used in data center scenarios, primarily falling into two categories: 1) pipeline parallelism, which splits the model layer by layer and distributes it across different devices, passing intermediate activation values ​​layer by layer during forward propagation; 2) tensor parallelism, which splits the weight matrix within a single layer by row or column, with each device computing its portion of the results in parallel before aggregation. These methods perform well in GPU clusters and batch inference scenarios, but face significant challenges in resource-constrained edge environments where autoregressive decoding is performed in a token-by-token manner.

[0004] Specifically, existing technologies suffer from the following key problems: First, pipeline parallelism suffers from severe device idleness during the token-by-token decoding stage. Since each device must wait for the output of the previous layer before it can start computing, the utilization rate of computing resources is low, making it difficult to reduce single-token latency. Second, tensor parallelism introduces excessive synchronization overhead on edge devices. Each token generation requires multiple fine-grained communications and synchronizations between devices. However, the limited bandwidth and low efficiency of the edge network communication stack mean that the synchronization time under normal circumstances will be greater than the local computing time, thus becoming a performance bottleneck. Finally, existing methods generally adopt a uniform, operator-independent partitioning strategy, ignoring the significant differences in computation-communication characteristics and KV cache semantics (MHA / MQA / GQA) between different modules in the Transformer architecture (such as self-attention, feedforward network, and output projection). This leads to unbalanced load, excessive communication volume, and even damage to the correctness of the attention mechanism after partitioning.

[0005] To address the aforementioned issues, some studies have attempted to introduce optimization techniques such as hybrid parallelism or computation-communication overlap, but these have not fundamentally solved the synchronization-dominated problem in token-by-token inference and lack adaptive support for different attention variants. Therefore, a novel collaborative inference framework is urgently needed to achieve efficient, scalable, and low-latency large language model inference on resource-constrained edge device clusters. Summary of the Invention

[0006] This invention provides a multi-terminal device collaborative reasoning method and system for large model perception segmentation to solve the localization problem in existing collaborative reasoning frameworks.

[0007] To achieve the above objectives, the present invention employs the following technical solution: In a first aspect, the present invention provides a multi-terminal device collaborative reasoning method for large-model perceptual segmentation, comprising: Fine-grained analysis of the computation graph of the target large language model is performed to identify the types of operators and their computation-communication characteristics; Based on the type of operator, memory usage, computational intensity, and attention mechanism variants, a differentiated weighting strategy is automatically generated; The weighted shards and computational tasks are mapped to multiple edge devices to construct the minimum execution subgraph for each device. During the inference process, a host-centralized aggregation coordination and computation-communication overlap mechanism is adopted to minimize synchronization overhead; Perform autoregressive inference, generate outputs per token, and optimize subsequent inference efficiency through a dynamic weight caching mechanism.

[0008] Optionally, the weighting strategy includes: The weights of Query, Key, and Value in the self-attention module are either column-splitting or shared and copied according to the attention mechanism type. The up-projection and gated projection weights in the feedforward network are split into columns, and the down-projection weights are split into rows. The output projection layer weights are column-segmented according to the vocabulary dimension; Lightweight operators are marked as locally reserved and executed only on the host device.

[0009] Optionally, the attention mechanism variants include multi-head attention mechanism, multi-query attention mechanism, and grouped query attention mechanism, with the mechanism classification methods as follows: For the multi-head attention mechanism, the Q, K, and V weights are all column-segmented, and the output projection weights are row-segmented. For the multi-query attention mechanism, the K and V weights are fully copied to each device, and only the Q weight is column-splittered. For the grouped query attention mechanism, the K and V heads are grouped and shared according to the ratio of the number of devices to the number of head groups, while the Q head is evenly distributed.

[0010] Optionally, the construction of the minimum execution subgraph includes: Generate an execution graph for each device that contains only its local operators; The synchronization operation nodes are explicitly marked in the diagram, including AllReduce and AllGather; Eliminate lightweight operators that do not participate in distributed execution to avoid unnecessary communication dependencies.

[0011] Optionally, the computation-communication overlap mechanism includes: The host starts local computation immediately after distributing input, without waiting for a response from the worker node; Use a pre-allocated buffer and an I / O thread pool to asynchronously receive partial results from worker nodes; Once the results are collected, the aggregation operation is triggered immediately and the next stage of calculation is advanced.

[0012] Optionally, the method further includes: After the initial inference is completed, each device caches weighted shards based on memory conditions; In subsequent inference, the host retrieves and skips cached weight transfers by weight identifier.

[0013] In a second aspect, embodiments of this application provide a multi-terminal device collaborative reasoning system for implementing large model-aware segmentation as described in any of the first aspects, comprising: The model operator-aware partitioning module is used to parse the model structure and generate partitioning strategies; The device-aware load balancing mapping module is used to map task shards to heterogeneous device clusters; The synchronization overhead minimization inference module is used to perform inference and optimize communication synchronization; The collaborative device weight caching module is used to manage the caching and reuse of weight shards.

[0014] Optionally, the model operator-aware partitioning module includes: The model weight analysis submodule is used to identify separable weights and local retention operators; An adaptive partitioning submodule for attention mechanisms is used to dynamically generate partitioning templates based on attention types. The feedforward network collaborative partitioning submodule is used to perform complementary row and column partitioning of the FFN weights; The output projection vocabulary segmentation module is used to segment the output layer according to the vocabulary dimension; The Lightweight Operator Local Retention Submodule is used to mark and retain lightweight operators for execution on the host machine.

[0015] Optionally, the synchronization overhead minimization inference module includes: The minimized subgraph generation submodule is used to construct the device local execution graph; The host-centralized aggregation and coordination submodule is used to centrally process AllReduce and AllGather operations; The compute-communication overlapped execution submodule is used to implement asynchronous distribution and parallel computation; The collaborative device weight cache submodule is used to manage the storage and retrieval of weight cache.

[0016] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method as described in any one of the first aspects.

[0017] Beneficial effects: This invention provides a multi-terminal device collaborative inference method for large model-aware segmentation. Through a fine-grained partitioning strategy based on operator awareness, it automatically selects the optimal row / column or vocabulary dimension segmentation method based on the computation-communication characteristics of different self-attention architectures (MHA, MQA, GQA), feedforward networks (FFN), and output projections, and performs load balancing mapping in conjunction with device heterogeneity. Simultaneously, through a synchronization overhead minimization execution mechanism, it constructs a lightweight quantum graph, employs host-centric aggregation coordination, and deeply integrates computation-communication overlap and dynamic load scheduling, significantly hiding cross-device communication latency. This invention not only supports large model collaborative inference exceeding the memory capacity of a single device but also achieves up to 5.4 times inference acceleration in real-world edge clusters, significantly outperforming traditional single-machine inference or pipelined parallel baselines. It is suitable for practical application scenarios such as privacy-sensitive, low-latency localized intelligent assistants, AR / VR interaction, and edge AI services, enabling large language models with billions of parameters to truly run efficiently in cloud-based edge environments. It effectively solves the core problems faced by existing parallel inference solutions in edge scenarios, such as high synchronization overhead, high communication redundancy, and coarse-grained operator partitioning. Attached Figure Description

[0018] Figure 1 This is a flowchart of multi-terminal device collaborative reasoning for large model perception segmentation, a preferred embodiment of the present invention. Figure 2 This is a framework diagram of the multi-device collaborative reasoning system designed for this invention; Figure 3 The flowchart for model weight partitioning and collaborative reasoning designed for this invention; Figure 4 This is a performance comparison chart of the I / O and computational overlap proposed in this invention and the traditional I / O and computational synchronization. Figure 5A comparison of inference latency between deploying and collaboratively inferring an LLAMA3-120b Q2_K quantization model using eight Raspberry Pi 5 machines and pipelined parallel in a preferred embodiment of the present invention. Detailed Implementation

[0019] The technical solution of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms "an" or "a," and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms "connected" or "linked," and similar terms, are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up," "down," "left," "right," etc., are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship also changes accordingly.

[0021] Please see Figure 1 This application provides a multi-terminal device collaborative reasoning method for large model perception segmentation, including: Fine-grained analysis of the computation graph of the target large language model is performed to identify the types of operators and their computation-communication characteristics; Based on the type of operator, memory usage, computational intensity, and attention mechanism variants, a differentiated weighting strategy is automatically generated; The weighted shards and computational tasks are mapped to multiple edge devices to construct the minimum execution subgraph for each device. During the inference process, a host-centralized aggregation coordination and computation-communication overlap mechanism is adopted to minimize synchronization overhead; Perform autoregressive inference, generate outputs per token, and optimize subsequent inference efficiency through a dynamic weight caching mechanism.

[0022] Optionally, the weighting strategy includes: The weights of Query, Key, and Value in the self-attention module are either column-splitting or shared and copied according to the attention mechanism type. The up-projection and gated projection weights in the feedforward network are split into columns, and the down-projection weights are split into rows. The output projection layer weights are column-segmented according to the vocabulary dimension; Lightweight operators are marked as locally reserved and executed only on the host device.

[0023] In the above embodiments, during the model loading phase, all weights of the LLM are traversed to identify all divisible weight types, including self-attention, feedforward networks (FFN), and output projection, and their computational intensity, memory usage, and intermediate tensor size are evaluated. Based on preset thresholds such as the proportion of single operator latency or the proportion of parameters, computationally intensive operators are marked as "requires collaborative execution," while lightweight operators such as TokenEmbedding, LayerNorm, and rotation position encoding are marked as "locally retained," avoiding the introduction of unnecessary cross-device communication.

[0024] Optionally, the attention mechanism variants include multi-head attention mechanism, multi-query attention mechanism, and grouped query attention mechanism, with the mechanism classification methods as follows: For the multi-head attention mechanism, the Q, K, and V weights are all column-segmented, and the output projection weights are row-segmented. For the multi-query attention mechanism, the K and V weights are fully copied to each device, and only the Q weight is column-splittered. For the grouped query attention mechanism, the K and V heads are grouped and shared according to the ratio of the number of devices to the number of head groups, while the Q head is evenly distributed.

[0025] In the above embodiments, the model's metadata is first read to automatically identify whether the current model uses a key-value sharing strategy (MHA, MQA, or GQA). Then, a corresponding partitioning template is dynamically generated based on the identification results. For MHA (Multi-head Attention), the weight matrices of Query (Q), Key (K), and Value (V) are all column-splittered, so that each device can independently calculate a portion of the attention head;

[0026] The attention output projection matrix is ​​row-segmented, allowing local outputs to be summed and aggregated via AllReduce.

[0027] During the calculation, each device first calculates the attention result for its corresponding attention head separately:

[0028]

[0029] Then, the output of the attention head is calculated along with the corresponding output projection weights to obtain the output of the attention head:

[0030] For MQA (Multi-Query Attention), the shared key-value weights are fully copied to all devices, and only the Q weights are column-splittered to maintain semantic correctness and reduce communication volume. For GQA (Grouped Query Attention), the key / value headers are shared within the device group according to the ratio of the number of devices to the number of header groups, while the Q headers are evenly distributed to achieve a balance between computation and communication.

[0031] Optionally, the construction of the minimum execution subgraph includes: Generate an execution graph for each device that contains only its local operators; The synchronization operation nodes are explicitly marked in the diagram, including AllReduce and AllGather; Eliminate lightweight operators that do not participate in distributed execution to avoid unnecessary communication dependencies.

[0032] In the above embodiments, regarding the FFN contained in Three sub-operators, employing a complementary segmentation strategy: and Column partitioning is used to enable each device to independently calculate high-dimensional intermediate activations; By employing row partitioning, the local downprojection results can be directly summed and aggregated into a complete output through AllReduce. This design avoids cross-device broadcasting during intermediate activation, compressing the communication volume to a single reduction operation and significantly reducing synchronization overhead.

[0033] Optionally, the computation-communication overlap mechanism includes: The host starts local computation immediately after distributing input, without waiting for a response from the worker node; Use a pre-allocated buffer and an I / O thread pool to asynchronously receive partial results from worker nodes; Once the results are collected, the aggregation operation is triggered immediately and the next stage of calculation is advanced.

[0034] In the above embodiment, the output projection layer at the end of the model is specifically optimized. Since the vocabulary size is typically much larger than the model's hidden dimensions, the weights in this layer consume significant memory. Specifically, this submodule outputs the projection matrix. Divide into N non-overlapping sub-matrices along the column direction (i.e., vocabulary dimension). Each device holds only the weights corresponding to a portion of the vocabulary. During inference, each device computes local logits in parallel:

[0035] Subsequently, all devices concatenate the local logits along the vocabulary dimension using the AllGather operation to form a complete logits vector for final token prediction. This design distributes the output layer memory overhead across all devices, avoiding single-device memory bottlenecks, while the communication volume is only the size of the logits vector (not the weights themselves), significantly outperforming the full copy strategy.

[0036] Optionally, the method further includes: After the initial inference is completed, each device caches weighted shards based on memory conditions; In subsequent inference, the host retrieves and skips cached weight transfers by weight identifier.

[0037] In the above embodiments, lightweight operators that contribute minimally to overall latency but would introduce significant communication overhead if executed in a distributed manner are identified and retained, are forced to execute locally on the host device.

[0038] The specific operators retained include: Token Embedding: Although its weight matrix size is comparable to the output projection, its forward computation is only performed once in the prefill stage, and there is no KV cache dependency in the decode stage. It can be fully loaded into the host memory without splitting. Layer Normalization: The calculation is element-wise normalization, with a very small number of parameters (only scaling and offset vectors), and requires the complete input tensor. Distributed execution requires additional AllGather, which is not worthwhile. Rotation Position Encoding (RoPE): This is a deterministic function with no parameters, negligible computational overhead, and no communication cost when executed locally.

[0039] By marking the above operators as "locally retained", the system removes them from the distributed execution plan when building the device subgraph, and they are executed only by the host, thereby eliminating unnecessary cross-device synchronization points.

[0040] After completing the partitioning templates for each operator, the weighted shards and computation tasks are mapped to the actual edge device clusters participating in the collaboration. The core of this approach lies in jointly considering device heterogeneity (memory, computing power, bandwidth) and operator computation-communication characteristics to achieve global load balancing.

[0041] This application also provides a multi-terminal device collaborative reasoning system for large model perceptual segmentation, including: The model operator-aware partitioning module is used to parse the model structure and generate partitioning strategies; The device-aware load balancing mapping module is used to map task shards to heterogeneous device clusters; The synchronization overhead minimization inference module is used to perform inference and optimize communication synchronization; The collaborative device weight caching module is used to manage the caching and reuse of weight shards.

[0042] Optionally, the model operator-aware partitioning module includes: The model weight analysis submodule is used to identify separable weights and local retention operators; An adaptive partitioning submodule for attention mechanisms is used to dynamically generate partitioning templates based on attention types. The feedforward network collaborative partitioning submodule is used to perform complementary row and column partitioning of the FFN weights; The output projection vocabulary segmentation module is used to segment the output layer according to the vocabulary dimension; The Lightweight Operator Local Retention Submodule is used to mark and retain lightweight operators for execution on the host machine.

[0043] Optionally, the synchronization overhead minimization inference module includes: The minimized subgraph generation submodule is used to construct the device local execution graph; The host-centralized aggregation and coordination submodule is used to centrally process AllReduce and AllGather operations; The compute-communication overlapped execution submodule is used to implement asynchronous distribution and parallel computation; The collaborative device weight cache submodule is used to manage the storage and retrieval of weight cache.

[0044] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any one of the methods for multi-terminal device collaborative reasoning in large model perceptual segmentation.

[0045] Specifically, in the existing embodiment, the deployment and running of the LLAMA3-120b Q2_K quantization model (total model size approximately 46.2GB) on eight Raspberry Pi 5 development boards (equipped with 8 GB RAM and quad-core ARM Cortex-A76 CPUs) is used as an example to illustrate the specific implementation steps of the present invention. The entire implementation process is divided into two stages: offline planning and online inference. The overall workflow is as follows: Figure 2 As shown.

[0046] I. Offline Planning Phase Step 1: Operator-aware model weight partitioning Before loading the model, the model operator-aware partitioning module of this invention first parses the computation graph structure of LLaMA3-120B to identify the types of operators and their key-value sharing strategies. The model uses Grouped Query Attention (GQA), with each layer containing 8 key-value headers. With a total of 8 devices, each device is allocated exactly one key-value header.

[0047] Self-attention module: Query weight per layer Column-wise splitting is used, with each device allocated an 8192×1024 submatrix; Key and Value Weights Due to the GQA sharing mechanism, the data is evenly divided according to the number of KV heads, with each device allocated an 8192×128 sub-matrix; Attention output projection weights Row-wise splitting is used, with each device allocated a submatrix of 1024×8192.

[0048] Feedforward Network (FFN): This model FFN does not contain bias terms and includes upprojection. Gated projection (All are 8192×28672) and lower projection

[0049] This invention relates to and Using column partitioning, for Row partitioning is used, and each device obtains weighted fragments of 8192×3584, 8192×3584 and 3584×8192 respectively.

[0050] Output projection layer: Vocabulary projection matrix The data is segmented by column dimension, and each device is assigned an 8192×4000 submatrix.

[0051] Locally retained operators: Lightweight operators such as Token Embedding and LayerNorm are executed entirely on the coordinating host because they have low computational overhead and require complete input, without being split.

[0052] The size and file offset of all weighted slices are pre-calculated during the planning phase. For example, the third device... The offset is 8192×1024×2×element_size. Since the total size of the model (46.2 GB) far exceeds the memory of a single device (8 GB), the host reads each tensor fragment as needed during loading via the system call read, and distributes them to the corresponding device or keeps them locally via RPC.

[0053] Step 2: Construct a minimal device subgraph and a collaborative computing graph After weight partitioning and transmission are completed, the system enters the computation graph construction phase. The core of this invention lies in generating a minimal execution subgraph for each device, containing only the operator nodes that it actually needs to execute and the necessary synchronization primitives, thereby avoiding the high latency and memory overhead caused by full graph broadcasting in traditional RPC frameworks.

[0054] The specific build process is as follows: 1. Host subgraph construction: The subgraph of the coordinating host (i.e., the designated master device) contains two types of nodes: Local exclusive operators: Lightweight operators such as Token Embedding, all LayerNorms, and RoPE rotation position encoding are stored directly on the host because they do not require distributed execution; Aggregation Nodes: Insert synchronization primitives after each Self-Attention, FFN, and Output Projection module—insert AllReduce nodes for the outputs of Self-Attention and FFN (for summing and aggregating local results), and insert AllGather nodes for Output Projection (for concatenating local logits).

[0055] 2. Construction of the working node subgraph: The subgraph for each worker node (i.e., the 7 Raspberry Pis participating in the collaboration) only contains the computation nodes corresponding to its assigned weight: Self-Attention Slice: Receives hidden state input from the host, performs local Q / K / V projection (according to GQA rules, K / V is single head, Q is partial head), and calculates local attention output; FFN Fragments: Execute Locally , Column segmentation calculation, then through Line splitting yields local FFN output; Output Projection: Perform local vocabulary projection on the FFN output to generate logits fragments with dimensions [1, 4000]. All subgraphs do not contain any aggregation logic; they only return the computation results to the host via RPC.

[0056] 3. Explicit modeling using synchronization primitives: The system explicitly marks two types of synchronization operations in the computation graph: AllReduce: Used for additivity intermediate results (such as Self-Attention output, FFN output). Each device calculates a local partial result, and the host sums and broadcasts the complete result. AllGather: Used for results that cannot be added but can be concatenated (such as Output Projection logits). Each device returns local logits, and the host concatenates them along the vocabulary dimension to form a complete logits of [1, 32000].

[0057] 4. Align dependencies with pipelines: All subgraphs strictly follow the Transformer layer order: Embedding → LayerNorm → Self-Attention → LayerNorm → FFN → LayerNorm → … → Output Projection.

[0058] The inputs to each layer's Self-Attention and FFN come from the complete output of the previous layer (already aggregated in the preceding AllReduce), ensuring semantic correctness. The execution flow of the computation graph is as follows: Figure 3 As shown.

[0059] Step 3: Collaborative Reasoning Execution and Computation – Communication Overlap The inference phase employs an autoregressive approach, executing token-by-token. The processing flow for each token is as follows, fully embodying the synchronization minimization and computation-communication overlap mechanism of this invention: 1. Host-based local preprocessing: The host first performs Token Embedding and initial LayerNorm to generate the input tensor for the first layer of Self-Attention.

[0060] 2. Asynchronous dispatch and parallel startup: like Figure 4 As shown, the host distributes X to all 7 worker nodes in parallel via non-blocking RPC, and immediately starts forward computation of the local Self-Attention slice (without waiting for worker nodes to confirm receipt), realizing "computation upon sending".

[0061] 3. Parallel computing of worker nodes: Self-Attention slice calculation (including Q / K / V projection, attention score, and local output); FFN slice calculation (up / gate / down projection); Output Projection is calculated in slices (e.g., if the current layer is the last layer). After the calculation is complete, the results (such as local attention output, local FFN output, or local logits) are returned to the host via RPC.

[0062] 4. Asynchronous result collection and aggregation: The host maintains an I / O thread pool, where each thread is responsible for listening to the return data stream from a worker node. When a partial result arrives, the thread writes it directly to a pre-allocated aggregate buffer, avoiding latency jitter introduced by dynamic memory allocation.

[0063] Once all Self-Attention local results have arrived, the host executes AllReduce to obtain the complete attention output; similarly, the FFN output is also aggregated through AllReduce; finally, the local logits of Output Projection are combined into complete logits through AllGather (concatenation).

[0064] Through the above mechanism, this invention successfully ran LLaMA3-120B-Q2 (46.2GB) on 8 Raspberry Pi 5 machines. Using the pipeline parallelism built into llama.cpp as the baseline, the inference latency per token was significantly lower than the baseline (measured at 4.6× speedup), verifying the efficiency and feasibility of the collaborative inference framework in resource-constrained edge environments.

[0065] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A multi-terminal device collaborative reasoning method for large-scale model perception segmentation, characterized in that, include: Fine-grained analysis of the computation graph of the target large language model is performed to identify the types of operators and their computation-communication characteristics; Based on the type of operator, memory usage, computational intensity, and attention mechanism variants, a differentiated weighting strategy is automatically generated; The weighted shards and computational tasks are mapped to multiple edge devices to construct the minimum execution subgraph for each device. During the inference process, a host-centralized aggregation coordination and computation-communication overlap mechanism is adopted to minimize synchronization overhead; Perform autoregressive inference, generate outputs for each token, and optimize subsequent inference efficiency through a dynamic weight caching mechanism; The weighting strategy includes: The weights of Query, Key, and Value in the self-attention module are either column-splitting or shared and copied according to the attention mechanism type. The up-projection and gated projection weights in the feedforward network are split into columns, and the down-projection weights are split into rows. The output projection layer weights are column-segmented according to the vocabulary dimension; Lightweight operators are marked as locally reserved and executed only on the host device; these lightweight operators include: TokenEmbedding, LayerNorm, and rotation position encoding. The attention mechanism variants include multi-head attention mechanism, multi-query attention mechanism, and grouped query attention mechanism, and the mechanism classification methods are as follows: For the multi-head attention mechanism, the Q, K, and V weights are all column-segmented, and the output projection weights are row-segmented. For the multi-query attention mechanism, the K and V weights are fully copied to each device, and only the Q weight is column-splittered. For the grouped query attention mechanism, the K and V heads are grouped and shared according to the ratio of the number of devices to the number of head groups, while the Q head is evenly distributed. The construction of the minimum execution subgraph includes: Generate an execution graph for each device that contains only its local operators; The synchronization operation nodes are explicitly marked in the diagram, including AllReduce and AllGather; Eliminate lightweight operators that do not participate in distributed execution to avoid unnecessary communication dependencies; The computation-communication overlap mechanism includes: The host starts local computation immediately after distributing input, without waiting for a response from the worker node; Use a pre-allocated buffer and an I / O thread pool to asynchronously receive partial results from worker nodes; Once the results are collected, the aggregation operation is triggered immediately and the next stage of calculation is advanced.

2. The multi-terminal device collaborative reasoning method for large model perception segmentation according to claim 1, characterized in that, The method further includes: After the initial inference is completed, each device caches weighted shards based on memory conditions; In subsequent inference, the host retrieves and skips cached weight transfers by weight identifier.

3. A multi-terminal device collaborative reasoning system for implementing large model perceptual segmentation as described in any one of claims 1 to 2, characterized in that, include: The model operator-aware partitioning module is used to parse the model structure and generate partitioning strategies; The device-aware load balancing mapping module is used to map task shards to a heterogeneous device cluster; The synchronization overhead minimization inference module is used to perform inference and optimize communication synchronization; The collaborative device weight caching module is used to manage the caching and reuse of weight shards.

4. The multi-terminal device collaborative reasoning system for large model perception segmentation according to claim 3, characterized in that, The model operator-aware partitioning module includes: The model weight analysis submodule is used to identify separable weights and local retention operators; An adaptive partitioning submodule for attention mechanisms is used to dynamically generate partitioning templates based on attention types. The feedforward network collaborative partitioning submodule is used to perform complementary row and column partitioning of the FFN weights; The output projection vocabulary segmentation module is used to segment the output layer according to the vocabulary dimension; The lightweight operator local retention submodule is used to mark and retain lightweight operators for execution on the host machine.

5. The multi-terminal device collaborative reasoning system for large model perception segmentation according to claim 3, characterized in that, The synchronization overhead minimization inference module includes: The minimized subgraph generation submodule is used to construct the device local execution graph; The host-centralized aggregation and coordination submodule is used to centrally process AllReduce and AllGather operations; The compute-communication overlapped execution submodule is used to implement asynchronous distribution and parallel computation; The collaborative device weight cache submodule is used to manage the storage and retrieval of weight cache.

Citation Information

Patent Citations

  • Large-scale language model reasoning optimization method based on token fusion

    CN118761468A

  • Large language model reasoning acceleration method and system based on GQA characteristic optimization and application

    CN120892179A