Large model reasoning optimization method and device, storage medium and electronic equipment
By using greedy and maximum flow algorithms to optimize data transmission paths in large model streaming networks, and combining them with a key-value hierarchical caching mechanism, the transmission bottleneck and uneven resource load issues of large models in distributed inference in the banking industry are resolved, thereby improving inference efficiency and system stability.
Patent Information
- Application Number
- CN202511151950.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-25
AI Technical Summary
Large models face challenges in distributed inference within the banking industry, including high KV cache transmission latency, mismatch between memory usage and transmission efficiency, and uneven resource load, leading to low inference efficiency.
A greedy algorithm is used to determine the optimal global transmission path in a pre-built large model flow network, and the maximum flow algorithm is combined for dynamic traffic allocation to optimize data transmission and resource scheduling. At the same time, a key-value hierarchical caching mechanism is used to accelerate the inference process.
It significantly improves the efficiency of large model inference, reduces cross-machine transmission latency, increases memory utilization, achieves load balancing, and ensures system stability and response speed.
Smart Images

Figure CN121009992A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a large model inference optimization method and device, a storage medium and an electronic device. BACKGROUND
[0002] At present, large models (LLM) have shown significant value in intelligent customer service, risk analysis and other scenarios in the banking industry.
[0003] However, the distributed inference efficiency of large models is limited by the KV Cache transmission bottleneck. The banking industry generally uses a Prefill and Decode separated inference architecture to improve throughput, but in the multi-machine collaborative scenario, it faces problems such as high KV Cache delay, mismatch between memory usage and transmission efficiency, and uneven resource load, which leads to low inference efficiency of large models and reduces the reliability of large model services.
[0004] Therefore, how to improve the inference efficiency of large models has become a technical problem that needs to be solved by those skilled in the art. SUMMARY
[0005] In view of the above problems, the present application provides a large model inference optimization method, device, storage medium and electronic device to overcome the above problems or at least partially solve the above problems, and the technical solutions are as follows:
[0006] A large model inference optimization method, comprising:
[0007] obtaining an input sequence to be processed;
[0008] determining an optimal global transmission path for processing the input sequence in a pre-constructed large model flow network using a greedy algorithm, wherein the large model flow network is composed of each pre-filled node, each decoding node of the large model, and an edge connecting the pre-filled node and the decoding node;
[0009] In the case where the input sequence is input to the large model, the pre-filled node and the decoding node in the optimal global transmission path are used for calculation and inference until the large model outputs an inference result corresponding to the input sequence.
[0010] Optionally, the capacity of each edge in the large model flow network is configured as the maximum transmission bandwidth between the pre-filled node and the decoding node connected by the edge, and the node load of each edge is configured as the occupied data flow between the pre-filled node and the decoding node connected by the edge. The method for determining the optimal global transmission path for processing the input sequence in the pre-constructed large model flow network using a greedy algorithm comprises:
[0011] In the pre-constructed large model flow network, the current demand flow of each pre-filled node is obtained;
[0012] In order to the current demand flow from high to low, the greedy algorithm is used to select the edge with the maximum current remaining capacity and the minimum current node load as the transmission path for each pre-filled node in turn;
[0013] In the case of selecting any transmission path, the maximum flow algorithm is used to dynamically allocate flow to the transmission path, and the remaining capacity and node load of each edge in the large model flow network are updated in real time;
[0014] The transmission path of each pre-filled node is used as the optimal global transmission path for processing the input sequence.
[0015] Optionally, the method further comprises:
[0016] The node load of each edge in the large model flow network is monitored in real time, and in the case that the node load of any edge exceeds the preset load threshold, the transmission path and flow allocation of the pre-filled node corresponding to the edge are readjusted using the greedy algorithm to realize the load balancing of the large model flow network.
[0017] Optionally, before the optimal global transmission path for processing the input sequence is determined in the pre-constructed large model flow network using the greedy algorithm, the method further comprises:
[0018] Each pre-filled node and each decoding node of the large model are obtained, wherein the pre-filled node and the decoding node represent a machine or a processing unit;
[0019] The edges connecting each pre-filled node and each decoding node are constructed, and the capacity of each edge is configured as the maximum transmission bandwidth between the pre-filled node and the decoding node connected by the edge, and the node load of each edge is configured as the occupied data flow between the pre-filled node and the decoding node connected by the edge, to obtain the constructed large model flow network.
[0020] Optionally, the method further comprises:
[0021] In the inference stage of the decoding node, a key-value hierarchical cache mechanism is used for inference acceleration to transmit the cache data in multiple cache layers to the decoding node in the order of cache layer priority.
[0022] Optionally, the method further comprises:
[0023] A hash value is generated for the cache data of the cache layer;
[0024] Before transmitting the cached data of the cache layer to the decoding node, the cached data is checked for consistency using the hash value. If the check passes, the step of transmitting the cached data of the cache layer to the decoding node is executed.
[0025] Optionally, the number of cache layers in the key-value hierarchical caching mechanism is related to the computation graph of the large model, and the cache size of each cache layer is related to the sequence length of the input sequence.
[0026] A large-scale model inference optimization device includes: an input sequence acquisition unit, an optimal global transmission path determination unit, and an inference result acquisition unit.
[0027] The input sequence obtaining unit is used to obtain the input sequence to be processed;
[0028] The optimal global transmission path determination unit is used to determine the optimal global transmission path for processing the input sequence in a pre-constructed large model stream network using a greedy algorithm. The large model stream network consists of each pre-filled node, each decoding node, and edges connecting the pre-filled nodes and the decoding nodes of the large model.
[0029] The reasoning result acquisition unit is used to perform calculation and reasoning via the pre-filled node and the decoding node in the optimal global transmission path when the input sequence is input to the large model, until the large model outputs a reasoning result corresponding to the input sequence.
[0030] A computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the large model inference optimization method described above.
[0031] An electronic device includes at least one processor, at least one memory connected to the processor, and a bus; wherein the processor and the memory communicate with each other via the bus; the processor is used to call program instructions in the memory to execute the large model inference optimization method.
[0032] By employing the above technical solution, this invention provides a large-model inference optimization method, apparatus, storage medium, and electronic device. The method obtains the input sequence to be processed; utilizes a greedy algorithm to determine the optimal global transmission path for processing the input sequence within a pre-constructed large-model stream network, wherein the large-model stream network consists of pre-filled nodes, decoded nodes, and edges connecting the pre-filled and decoded nodes of the large model; when the input sequence is input to the large model, computational inference is performed via the pre-filled and decoded nodes in the optimal global transmission path until the large model outputs the inference result corresponding to the input sequence. This invention effectively optimizes data transmission and computational scheduling between pre-filled and decoded nodes by dynamically determining the optimal global transmission path for input sequence processing within the large-model stream network based on a greedy algorithm, thereby improving resource utilization and response speed during large-model inference and achieving a significant improvement in the efficiency of large-model inference.
[0033] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0034] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0035] Figure 1 The diagram illustrates a flowchart of one implementation of the large model inference optimization method provided in this invention.
[0036] Figure 2 A flowchart illustrating another implementation of the large model inference optimization method provided in this invention is shown.
[0037] Figure 3 A schematic diagram of the structure of the large model inference optimization device provided in an embodiment of the present invention is shown;
[0038] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present invention is shown. Detailed Implementation
[0039] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0040] With the rapid development of artificial intelligence technology, Large Language Models (LLMs) are increasingly being used in the banking industry, especially in key business areas such as intelligent customer service, risk analysis, and automated financial statement generation, demonstrating strong application potential. Through its ability to understand and generate large amounts of financial text data, LLMs significantly improve the intelligence level and service quality of banking operations. However, limited by the computational scale and complexity of the large models themselves, inference efficiency and resource utilization have become key factors affecting system performance.
[0041] Currently, the banking industry widely relies on distributed computing platforms to support large-scale model inference and data processing. To adapt to the characteristics of large model inference, a distributed inference architecture that separates Prefill and Decode is widely adopted. This technology significantly improves the model's throughput and scalability by separating the model's prefill and decode phases. However, in cross-machine distributed inference, especially in multi-machine environments, there is a significant performance bottleneck in the transmission of critical key-value cache data, limiting the overall improvement in inference efficiency.
[0042] The technical solutions currently used in banking scenarios face the following main challenges:
[0043] 1. KV Cache has high transmission latency.
[0044] Due to limitations in network bandwidth and latency across multiple machines, the cross-node transmission efficiency of KV Cache is low, becoming a bottleneck in distributed inference. This problem is particularly prominent in banking applications with extremely high real-time requirements, such as instant response in intelligent customer service and real-time risk assessment, where transmission latency directly impacts business experience and system response speed.
[0045] 2. Mismatch between video memory usage and transfer efficiency
[0046] Banking operations typically involve large-scale customer data and long text inference, leading to a sharp increase in GPU memory requirements. If KVCache management is not effectively optimized, it can easily cause out-of-memory (OOM) problems or resource waste, thereby affecting inference stability and cost efficiency.
[0047] 3. Uneven resource load
[0048] Because bank computing platforms often employ hybrid architectures with diverse and complex distribution of computing resources, load balancing and scheduling are challenging. Currently, resource scheduling largely relies on manual experience, lacking intelligent and dynamic adjustment mechanisms. This is especially pronounced in prefill / decode separation architectures, where uneven load distribution is more pronounced, leading to overloaded resources on some nodes while other nodes remain idle, thus reducing overall system performance.
[0049] To address the aforementioned challenges, designing efficient KV Cache transmission optimization schemes and achieving refined resource management and dynamic load balancing have become key technical challenges in improving the performance and service quality of large-scale model inference in the banking industry.
[0050] Based on this, the large-model inference optimization method provided in this embodiment of the invention significantly reduces cross-machine transmission latency and improves inference speed and memory utilization by more than 30% through hierarchical KV cache management and greedy algorithm optimization. Simultaneously, intelligent resource allocation and real-time load balancing mechanisms ensure the stable operation of the large-model inference system under high-concurrency scenarios, avoiding node overload. Dynamic allocation of hierarchical cache effectively reduces memory usage, avoids OOM (Out of Memory) issues, and improves system stability. Furthermore, this embodiment of the invention supports automatic node expansion, which can efficiently handle a large number of concurrent tasks, ensuring rapid response and processing efficiency. Overall, the scheme based on hierarchical KV cache management and graph algorithm scheduling optimization effectively solves the data transmission bottleneck and memory usage problems in cross-machine inference, optimizes resource allocation and load balancing, and significantly improves the efficiency, real-time performance, stability, and reliability of the bank's large-model inference system, thereby providing more efficient, accurate, and stable technical support for the bank's intelligent customer service, risk assessment, credit approval, and other businesses.
[0051] like Figure 1 The diagram shows a flowchart of one implementation of the large model inference optimization method provided in this invention. The method may include:
[0052] S100, Obtain the input sequence to be processed.
[0053] In this context, the input sequence refers to a set of ordered data elements in the large-scale model inference task. Arranged in a specific order, the input sequence serves as the input information for the large model, driving it to perform calculations and inferences through pre-filled nodes and decoding nodes, thereby generating the corresponding inference results. Large-scale model inference tasks can be those requiring large-scale model inference in key business areas such as intelligent customer service, risk analysis, and automatic financial statement generation. For example, in intelligent customer service scenarios, the input sequence might be the user's dialogue text or query content. In risk analysis, the input sequence might contain relevant financial data, transaction records, or market indicators. In automatic financial statement generation, the input sequence might be raw financial data entries or historical accounting information. These input sequences, organized in a specific order, serve as the input to the large model, driving its computational inference process through pre-filled nodes and decoding nodes, ultimately generating accurate output results tailored to specific business needs.
[0054] S110. Using a greedy algorithm, determine the optimal global transmission path for processing the input sequence in a pre-constructed large model stream network. The large model stream network consists of each pre-filled node, each decoding node, and edges connecting the pre-filled nodes and the decoding nodes of the large model.
[0055] Among them, the greedy algorithm is an algorithm design strategy for quickly determining the approximate optimal path for processing input sequences in large model flow networks, so as to reduce data transmission latency and system load when processing input sequences.
[0056] Among them, the large model flow network is a graph-based network structure used to represent data transmission paths during large model inference. This network consists of multiple nodes (including pre-filled and decoded nodes) and edges connecting them. Edges have capacity constraints, and data flow between nodes follows flow conservation and capacity constraints. The flow network model helps the system efficiently schedule and manage cross-node data transmission, optimizing inference performance.
[0057] The optimal global transmission path refers to the path from the input node (pre-filled node) to the output node (decoding node) in a large model flow network. This path takes into account multiple factors such as transmission delay, bandwidth, and load to ensure that the overall transmission efficiency reaches the global optimum. Even when multiple paths are available, this path has the best performance indicators.
[0058] Among them, prefill nodes are a type of node in large model flow networks, used to preload or cache some of the state information required for model computation (such as KV Cache), in order to prepare for the inference computation of subsequent decoding nodes.
[0059] Among them, the Decode node is a type of node in the large model flow network, which is used to perform the decoding calculation of the model based on the input sequence and the state information provided by the pre-filled nodes, and gradually generate the final inference result.
[0060] In the large model flow network, the directed edges connecting the pre-filled nodes and the decoding nodes represent data transmission paths. Each edge has a certain capacity (representing transmission bandwidth or resource constraints), and data flow must satisfy capacity constraints and flow conservation principles. The weights of the edges (such as latency and bandwidth) affect path selection and become an important basis for the greedy algorithm to calculate the optimal path.
[0061] Optionally, the capacity of each edge in the large model flow network is configured as the maximum transmission bandwidth between the pre-filled nodes and the decoding nodes connected by the edge, and the node load of each edge is configured as the occupied data traffic between the pre-filled nodes and the decoding nodes connected by the edge.
[0062] Specifically, embodiments of the present invention can start from pre-filled nodes to prepare for selecting a transmission path. At each decision step, a greedy algorithm selects the edge with the lowest transmission delay, highest bandwidth, or lowest system load from the currently available edges as the optimal next-hop path. This greedy algorithm selection process is repeated to progressively construct a complete transmission path from the pre-filled node to the decoding node. Ultimately, a global transmission path with optimal overall performance that satisfies capacity and flow constraints is obtained, used for efficient processing of inference computations of input sequences.
[0063] S120. When the input sequence is input into the large model, computation and inference are performed through the pre-filled nodes and decoding nodes in the optimal global transmission path until the large model outputs the inference result corresponding to the input sequence.
[0064] Specifically, when an input sequence is input into a large model, the input sequence is first transmitted to a pre-filled node in the large model's streaming network. The pre-filled node performs preliminary processing and state preparation on the input data. Based on the input sequence and model state, the pre-filled node performs pre-computation and transmits the results along the optimal global transmission path to the subsequent decoding nodes. The decoding nodes receive data from the pre-filled nodes and the previous state data, perform the core decoding and inference computations, and gradually generate inference outputs corresponding to the input sequence. The inference computation process flows and processes sequentially between nodes along a pre-determined optimal path, ensuring efficient transmission and computational collaboration. After continuous computation by the pre-filled and decoding nodes, the model finally outputs the inference result corresponding to the input sequence.
[0065] This invention provides a large-model inference optimization method, which includes: obtaining an input sequence to be processed; using a greedy algorithm to determine the optimal global transmission path for processing the input sequence in a pre-constructed large-model flow network, wherein the large-model flow network consists of pre-filled nodes, decoded nodes, and edges connecting the pre-filled nodes and decoded nodes of the large model; when the input sequence is input into the large model, computational inference is performed through the pre-filled nodes and decoded nodes in the optimal global transmission path until the large model outputs an inference result corresponding to the input sequence. This invention effectively optimizes data transmission and computational scheduling between pre-filled nodes and decoded nodes by dynamically determining the optimal global transmission path for input sequence processing in the large-model flow network based on a greedy algorithm, thereby improving resource utilization and response speed in the large-model inference process and achieving a significant improvement in the efficiency of large-model inference.
[0066] Optional, based on Figure 1 The method shown is as follows: Figure 2 As shown, this is a flowchart illustrating another implementation of the large model inference optimization method provided in this embodiment of the invention. Step S110 may include:
[0067] S200. In the pre-built large model flow network, obtain the current demand flow of each pre-filled node.
[0068] In this context, demand flow refers to the data flow that the pre-filled node needs to process and transmit in the current inference task. Demand flow can include the size of the input sequence data transmitted, the size of the state data (such as KVCache) generated during the pre-filling process, and the size of the intermediate result data generated during computation.
[0069] Because the banking industry relies on distributed computing platforms to support large-scale model inference and data processing, some pre-filled nodes bear more data loading and state preparation tasks when processing input sequences during large model inference, resulting in a large data transmission and computation load. Therefore, it is necessary to identify the current traffic demand of each pre-filled node, which will help with subsequent path selection and traffic allocation.
[0070] This invention can estimate the current traffic demand of a pre-filled node based on the input sequence and the size of its buffer state during the pre-filling process. This invention can also collect the transmission data rate and traffic statistics of the pre-filled node in real time at the network layer to obtain the current traffic demand of the pre-filled node.
[0071] S210. According to the current demand flow from high to low, use a greedy algorithm to select the edge with the largest remaining capacity and the smallest load of the current node as the transmission path for each pre-filled node.
[0072] Specifically, in this embodiment of the invention, for the pre-filled node with the highest traffic demand, a greedy algorithm is used to prioritize selecting the edge with the largest remaining capacity (bandwidth) and the smallest current load from all available transmission edges as its transmission path. After determining the transmission path, the remaining capacity and load status of that edge in the large model flow network are updated. Subsequently, for each remaining pre-filled node, the same greedy strategy is used to select transmission paths according to its current traffic demand, that is, prioritizing the edge with the largest remaining capacity and the smallest load, ensuring that each pre-filled node can obtain the best transmission resource allocation, and ultimately achieving efficient utilization and load balancing of overall network resources.
[0073] S220. Given any transmission path, dynamically allocate traffic to the transmission path using the maximum flow algorithm, and update the remaining capacity and node load of each edge in the large model flow network in real time.
[0074] The Maximum Flow Algorithm is a class of algorithms used to solve the maximum flow problem in flowing networks. In a given flowing network, which consists of nodes and directed edges with capacity constraints, the goal is to find the maximum possible flow distribution from the source node to the sink node, maximizing the total flow that the network can transmit without exceeding the edge capacity constraints. Maximum flow algorithms can include the Edmonds-Karp algorithm and the Dinic algorithm.
[0075] In this embodiment of the invention, after determining any transmission path, traffic on that path is dynamically allocated based on the maximum flow algorithm to ensure maximum transmission efficiency. After each traffic allocation, the remaining capacity of each edge on that path and the load status of related nodes are updated in real time to reflect the current resource usage. In this way, the traffic distribution in large-scale model flow networks can be continuously adjusted and optimized, achieving reasonable management and dynamic scheduling of network resources.
[0076] This invention, based on the specific transmission path determined by the greedy algorithm, utilizes the maximum flow algorithm to precisely allocate traffic along that path. By finding augmenting paths, the maximum flow algorithm dynamically adjusts the traffic allocation for each edge, ensuring maximum transmission within edge capacity constraints. After each traffic allocation, the remaining capacity of the edges and the node load in the path are updated in real time, providing accurate network state information for the next round of path selection and traffic allocation.
[0077] This invention employs a combination of greedy and maximum flow algorithms (the greedy algorithm quickly selects paths, while the maximum flow algorithm finely allocates traffic and updates resource status). This allows for dynamic adaptation to traffic demands and resource changes in large-scale model networks, achieving overall network load balancing and maximizing resource utilization. Consequently, it promotes the rational planning of transmission paths and dynamic adjustment of traffic in large-scale model flow networks, thereby improving the inference efficiency of large-scale models.
[0078] S230. The transmission path of each pre-filled node is taken as the optimal global transmission path for processing the input sequence.
[0079] In this embodiment of the invention, edge selection is performed greedily according to the demand flow from high to low, prioritizing the edge with the largest remaining capacity and the smallest node load, ensuring load balancing and sufficient resources in the transmission path. Subsequently, the maximum flow algorithm is combined to dynamically allocate traffic and update resources in real time on the selected path, ensuring optimal scheduling of network traffic and effective satisfaction of capacity constraints. This organically integrates the transmission paths of each pre-filled node into a globally optimal path, which not only improves the computational efficiency and transmission performance in the large model inference process, but also enhances the system's adaptability to dynamic load changes, thereby helping to improve the efficiency of large model inference.
[0080] Understandably, when the input sequence is transmitted from the pre-filled node to the decoding node in each round, the system first selects a transmission path using a greedy algorithm based on the current resource status and node load of the large model flow network. Then, it dynamically allocates traffic to the selected path using a maximum flow algorithm. After each traffic allocation, the system updates the remaining capacity and node load information of each edge in the flow network in real time, and based on the updated network state, it recalculates the optimal global transmission path for the next round using a combination of the greedy algorithm and the maximum flow algorithm, thereby continuously ensuring the global optimality of the transmission path and the efficient utilization of network resources.
[0081] Optionally, in the above Figure 2 Based on one or more corresponding embodiments, in another optional embodiment provided by the present invention, the method may further include:
[0082] The node load of each edge in the large model flow network is monitored in real time. If the node load of any edge exceeds the preset load threshold, a greedy algorithm is used to readjust the transmission path and traffic distribution of the pre-filled nodes corresponding to the edge in order to achieve load balancing of the large model flow network.
[0083] Specifically, in this embodiment of the invention, when the node load (the load of the decoding node connected to the edge) on any side exceeds a preset load threshold, a path and traffic readjustment mechanism based on a greedy algorithm is triggered. This mechanism dynamically replans the transmission paths of the corresponding pre-filled nodes by prioritizing edges with larger remaining capacity and lower node loads, and rationally allocates traffic using a maximum flow algorithm, thereby achieving a balanced distribution of the overall network load. Simultaneously, after each traffic adjustment, the capacity and node load information of each edge in the large-scale model flow network are updated in real time, ensuring that subsequent path selections are based on the latest network resource status.
[0084] This invention, through timely identification of load anomalies when the node load on any side of a large model flow network exceeds a preset threshold, and by reapplying a greedy algorithm to readjust and redistribute the transmission path and traffic of the pre-filled node corresponding to that side, effectively avoids excessive load and path congestion, ensuring a uniform distribution of load among nodes in the large model flow network, thereby helping to improve the inference efficiency of large models.
[0085] Optionally, in the above Figure 2 Based on one or more corresponding embodiments, in another optional embodiment provided by the present invention, before step S110, the method may further include:
[0086] Obtain the pre-filled nodes and decoder nodes of the large model, where each pre-filled node and decoder node represents a machine or processing unit. Construct edges connecting each pre-filled node and each decoder node, and configure the capacity of each edge to the maximum transmission bandwidth between the pre-filled node and the decoder node connected by the edge. Configure the node load of each edge to the occupied data traffic between the pre-filled node and the decoder node connected by the edge, thus obtaining the constructed large model streaming network.
[0087] In this context, "machine" refers to the physical or virtual computing devices that participate in large-scale model inference computations, typically including hardware entities such as servers, compute nodes, or workstations. These machines undertake the computational tasks of pre-filling nodes and decoding nodes, performing computations and data transmissions at different stages of the model.
[0088] A processing unit (CPU) refers to a functional unit within a machine that performs specific computational or data processing tasks. It can be a central processing unit (CPU), a graphics processing unit (GPU), or other dedicated processor. Each processing unit corresponds to a pre-filling node or a decoding node, which handles the pre-filling or decoding computations assigned to that node and communicates with other nodes.
[0089] Before using a greedy algorithm to determine the optimal global transmission path for the input sequence, this invention obtains a complete and accurate large-model flow network that reflects the current network state. This allows the greedy algorithm to selectively allocate paths based on the real and dynamically updated network resource conditions, choosing the edge with the largest remaining capacity and the smallest load. This ensures that the selection of transmission paths maximizes bandwidth utilization and load balancing, effectively avoiding network bottlenecks and congestion, improving the efficiency and stability of input sequence inference processing, and ultimately achieving efficient and globally optimal computational resource scheduling and data transmission management, significantly enhancing the performance and response speed during large-model inference.
[0090] Optionally, in the above Figure 1 Based on one or more corresponding embodiments, in another optional embodiment provided by the present invention, the method may further include:
[0091] During the inference phase of the decoding node, a key-value hierarchical caching mechanism is used to accelerate inference, so as to transfer cached data from multiple cache layers to the decoding node according to the priority of the cache layers.
[0092] The key-value hierarchical caching mechanism is a method for hierarchical management and prioritized transmission of cached key and value data during model inference. By dividing cached data into multiple levels and selectively transmitting cached data according to hierarchical priority, it improves inference efficiency, reduces memory consumption, and ensures the timely availability of critical data.
[0093] In this embodiment of the invention, the KV Cache is divided into different levels, with each cache layer containing intermediate computation results or partial cached data for a specific stage. Different cache layers reflect different computation depths or computation steps in the model inference process, and higher-priority cache layers typically contain later or more granular cached information.
[0094] The cached data, stored in the KV Cache, is intermediate result data used to accelerate model inference. It includes key and corresponding value information, i.e., the state information required for model computation. Depending on its cache level, the cached data undertakes different priority transmission tasks to ensure data consistency and efficiency during inference.
[0095] Optionally, the number of cache layers in the key-value hierarchical caching mechanism is related to the computation graph of the large model, and the cache size of each cache layer is related to the sequence length of the input sequence.
[0096] In deep learning models, a computation graph is a structured graph representing the computation process using a directed acyclic graph (DAG). The computation graph abstracts the various computational units of a large model (such as neural network layers, activation functions, matrix operations, etc.) as nodes, and the data flow relationships between different computational units are connected by directed edges, reflecting the order and dependencies of computation. Through the computation graph, the overall computational flow of a large model from input to output can be clearly described. For large models, the hierarchical structure of the computation graph directly affects the number of cache layers in the key-value hierarchical caching mechanism, because each computational layer corresponds to one cache layer.
[0097] In the key-value hierarchical caching mechanism provided in this invention, the number of caching layers is determined by the computation graph structure of the large model, with different computation layers corresponding to different numbers of caching layers. The cache size of each caching layer depends on the length of the input sequence; the longer the sequence, the larger the amount of cached data at the corresponding level. For example, assuming there are L layers of KV cache, each caching layer... The cache size is determined by this layer's index. This is determined in conjunction with the length of the input sequence. To speed up inference, the first sequence is transmitted first. Layer cache data (in < L, so that the decoding node can obtain the key calculation results as early as possible and reduce the waiting time.
[0098] In the inference stage of the decoding node, this embodiment of the invention adopts a key-value hierarchical caching mechanism, which divides the KV Cache into multiple levels and transmits cached data layer by layer according to the priority order of the cache layers. This can ensure that critical and computationally intensive intermediate results are available in a timely manner, reduce data transmission latency and memory consumption, effectively avoid inference bottlenecks caused by waiting for complete cached data, and thus significantly improve the inference speed of large models.
[0099] Optionally, embodiments of the present invention may also generate hash values for the cached data in the cache layer. Before transmitting the cached data from the cache layer to the decoding node, the hash value is used to perform a consistency check on the cached data. If the check passes, the step of transmitting the cached data from the cache layer to the decoding node is executed.
[0100] Specifically, in this embodiment of the invention, before transmitting the cached data of the cache layer to the decoding node, a hash value is first generated for the data of the cache layer to uniquely identify its content. A consistency check is performed by comparing the hash values before and after transmission to ensure that the cached data has not changed or been corrupted during transmission. Only when the check result shows a consistency condition (such as...) is the cached data transmitted to the decoding node. Only when the conditions are met will the operation of transferring the cached data of the cache layer to the decoding node be performed, thereby ensuring the integrity and correctness of the cached data.
[0101] Although the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous.
[0102] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0103] Corresponding to the above method embodiments, this invention also provides a large model inference optimization device, the structure of which is as follows: Figure 3 As shown, it may include: an input sequence acquisition unit 10, an optimal global transmission path determination unit 20, and an inference result acquisition unit 30.
[0104] Input sequence acquisition unit 10 is used to acquire the input sequence to be processed.
[0105] The optimal global transmission path determination unit 20 is used to determine the optimal global transmission path for processing the input sequence in a pre-constructed large model stream network using a greedy algorithm. The large model stream network consists of each pre-filled node, each decoding node, and edges connecting the pre-filled nodes and the decoding nodes of the large model.
[0106] The reasoning result acquisition unit 30 is used to perform computational reasoning through pre-filled nodes and decoding nodes in the optimal global transmission path when the input sequence is input into the large model, until the large model outputs the reasoning result corresponding to the input sequence.
[0107] Optionally, the capacity of each edge in the large model flow network is configured as the maximum transmission bandwidth between the pre-filled nodes and the decoding nodes connected by the edge, and the node load of each edge is configured as the occupied data traffic between the pre-filled nodes and the decoding nodes connected by the edge.
[0108] Optionally, the optimal global transmission path determination unit 20 can be used to obtain the current demand flow of each pre-filled node in the pre-constructed large model flow network; according to the current demand flow in descending order, use a greedy algorithm to select the edge with the largest current remaining capacity and the smallest current node load as the transmission path for each pre-filled node; when any transmission path is selected, use the maximum flow algorithm to dynamically allocate flow to the transmission path, and update the remaining capacity and node load of each edge in the large model flow network in real time; and use the transmission path of each pre-filled node as the optimal global transmission path for processing the input sequence.
[0109] Optionally, the large model inference optimization device may also include a load monitoring unit.
[0110] The load monitoring unit is used to monitor the node load of each edge in the large model flow network in real time. When the node load of any edge exceeds the preset load threshold, a greedy algorithm is used to readjust the transmission path and traffic distribution of the pre-filled nodes corresponding to the edge in order to achieve load balancing of the large model flow network.
[0111] Optionally, the large model inference optimization apparatus may also include: a large model flow network building unit.
[0112] The large model stream network construction unit, used by the optimal global transmission path determination unit 20, uses a greedy algorithm to determine the optimal global transmission path for processing the input sequence in the pre-constructed large model stream network. Before this, it obtains each pre-filled node and each decoding node of the large model, where each pre-filled node and decoding node represents a machine or processing unit. It constructs edges connecting each pre-filled node and each decoding node, configures the capacity of each edge to the maximum transmission bandwidth between the pre-filled node and the decoding node connected by the edge, and configures the node load of each edge to the occupied data flow between the pre-filled node and the decoding node connected by the edge, thus obtaining the constructed large model stream network.
[0113] Optionally, the large model inference optimization device may also include an inference acceleration unit.
[0114] The inference acceleration unit is used to accelerate inference during the inference phase of the decoding node using a key-value hierarchical caching mechanism, so as to transfer cached data from multiple cache layers to the decoding node according to the priority of the cache layers.
[0115] Optionally, the large model inference optimization device may also include a hash verification unit.
[0116] The hash verification unit is used to generate hash values for the cached data in the cache layer. Before transmitting the cached data from the cache layer to the decoding node, the hash value is used to verify the consistency of the cached data. If the verification passes, the cached data from the cache layer is transmitted to the decoding node.
[0117] Optionally, the number of cache layers in the key-value hierarchical caching mechanism is related to the computation graph of the large model, and the cache size of each cache layer is related to the sequence length of the input sequence.
[0118] This invention provides a large-model inference optimization device, which is used to: obtain an input sequence to be processed; determine the optimal global transmission path for processing the input sequence in a pre-constructed large-model flow network using a greedy algorithm, wherein the large-model flow network consists of pre-filled nodes, decoded nodes, and edges connecting the pre-filled nodes and decoded nodes of the large model; when the input sequence is input to the large model, computational inference is performed through the pre-filled nodes and decoded nodes in the optimal global transmission path until the large model outputs an inference result corresponding to the input sequence. This invention effectively optimizes data transmission and computational scheduling between pre-filled nodes and decoded nodes by dynamically determining the optimal global transmission path for input sequence processing in the large-model flow network based on a greedy algorithm, thereby improving resource utilization and response speed in the large-model inference process and achieving a significant improvement in the efficiency of large-model inference.
[0119] Regarding the apparatus in the above embodiments, the specific manner in which each unit performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0120] The large model inference optimization device includes a processor and a memory. The input sequence acquisition unit 10, the optimal global transmission path determination unit 20, and the inference result acquisition unit 30 are all stored as program units in the memory. The processor executes the program units stored in the memory to realize the corresponding functions.
[0121] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured. By adjusting kernel parameters, the optimal global transmission path for input sequence processing in large model streaming networks can be dynamically determined based on a greedy algorithm. This effectively optimizes data transmission and computation scheduling between pre-filled nodes and decoding nodes, improving the inference efficiency of large models.
[0122] This invention provides a computer-readable storage medium storing a program that, when executed by a processor, implements the large model inference optimization method.
[0123] This invention provides a processor for running a program, wherein the program executes the large model inference optimization method during runtime.
[0124] like Figure 4As shown, this embodiment of the invention provides an electronic device 1000, which includes at least one processor 1001, at least one memory 1002 connected to the processor 1001, and a bus 1003. The processor 1001 and the memory 1002 communicate with each other via the bus 1003. The processor 1001 is used to call program instructions in the memory 1002 to execute the aforementioned large model inference optimization method. The electronic device in this document can be a server, PC, PAD, mobile phone, etc.
[0125] The present invention also provides a computer program product that, when executed on an electronic device, is suitable for executing a program that initializes a large model inference optimization method step.
[0126] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0127] In a typical configuration, an electronic device includes one or more processors (CPUs), memory, and a bus. The electronic device may also include input / output interfaces, network interfaces, etc.
[0128] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM, and memory includes at least one memory chip. Memory is an example of computer-readable media.
[0129] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0130] In the description of this invention, it should be understood that if the terms "upper", "lower", "front", "rear", "left" and "right" are used to indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the position or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.
[0131] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0132] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0133] The above are merely embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the present invention.
Claims
1. A large-scale model inference optimization method, characterized in that, include: Obtain the input sequence to be processed; Using a greedy algorithm, the optimal global transmission path for processing the input sequence is determined in a pre-constructed large model stream network, wherein the large model stream network consists of each pre-filled node of the large model, each decoding node, and edges connecting the pre-filled nodes and the decoding nodes. When the input sequence is input to the large model, computational inference is performed via the pre-filled nodes and the decoding nodes in the optimal global transmission path until the large model outputs an inference result corresponding to the input sequence.
2. The method according to claim 1, characterized in that, In the large model stream network, the capacity of each edge is configured as the maximum transmission bandwidth between the pre-filled node and the decoding node connected by the edge, and the node load of each edge is configured as the occupied data traffic between the pre-filled node and the decoding node connected by the edge. The step of using a greedy algorithm to determine the optimal global transmission path for processing the input sequence in the pre-constructed large model stream network includes: In the pre-built large model flow network, obtain the current demand flow of each pre-filled node; According to the current demand flow from high to low, the greedy algorithm is used to select the edge with the largest remaining capacity and the smallest load of the current node as the transmission path for each of the pre-filled nodes. If any of the transmission paths is selected, the traffic is dynamically allocated to the transmission path using the maximum flow algorithm, and the remaining capacity and node load of each edge in the large model flow network are updated in real time. The transmission path of each of the pre-filled nodes is used as the optimal global transmission path for processing the input sequence.
3. The method according to claim 2, characterized in that, Also includes: The node load of each edge in the large model flow network is monitored in real time. If the node load of any edge exceeds a preset load threshold, a greedy algorithm is used to readjust the transmission path and traffic allocation of the pre-filled node corresponding to the edge in order to achieve load balancing of the large model flow network.
4. The method according to claim 2, characterized in that, Before determining the optimal global transmission path for processing the input sequence in a pre-constructed large model stream network using a greedy algorithm, the method further includes: Obtain each pre-filled node and each decoded node of the large model, wherein the pre-filled node and the decoded node represent a machine or processing unit; Construct edges connecting each of the pre-filled nodes and each of the decoded nodes, configure the capacity of each edge as the maximum transmission bandwidth between the pre-filled node and the decoded node connected by the edge, and configure the node load of each edge as the occupied data traffic between the pre-filled node and the decoded node connected by the edge, thereby obtaining the constructed large model stream network.
5. The method according to claim 1, characterized in that, Also includes: During the inference phase of the decoding node, a key-value hierarchical caching mechanism is used to accelerate inference, so as to transfer cached data from multiple cache layers to the decoding node according to the cache layer priority order.
6. The method according to claim 5, characterized in that, Also includes: Generate hash values for the cached data in the cache layer; Before transmitting the cached data of the cache layer to the decoding node, the cached data is checked for consistency using the hash value. If the check passes, the step of transmitting the cached data of the cache layer to the decoding node is executed.
7. The method according to claim 5, characterized in that, The number of cache layers in the key-value hierarchical caching mechanism is related to the computation graph of the large model, and the cache size of each cache layer is related to the sequence length of the input sequence.
8. A large-scale model inference optimization device, characterized in that, include: The system includes an input sequence acquisition unit, an optimal global transmission path determination unit, and an inference result acquisition unit. The input sequence obtaining unit is used to obtain the input sequence to be processed; The optimal global transmission path determination unit is used to determine the optimal global transmission path for processing the input sequence in a pre-constructed large model stream network using a greedy algorithm. The large model stream network consists of each pre-filled node, each decoding node, and edges connecting the pre-filled nodes and the decoding nodes of the large model. The reasoning result acquisition unit is used to perform calculation and reasoning via the pre-filled node and the decoding node in the optimal global transmission path when the input sequence is input to the large model, until the large model outputs a reasoning result corresponding to the input sequence.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the large model inference optimization method as described in any one of claims 1 to 7.
10. An electronic device, characterized in that, The electronic device includes at least one processor, at least one memory connected to the processor, and a bus; wherein the processor and the memory communicate with each other through the bus; the processor is used to call program instructions in the memory to execute the large model inference optimization method as described in any one of claims 1 to 7.