Reasoning system, reasoning method, and computer program product

CN122838006APending Publication Date: 2026-09-29HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510392157.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-29
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]由于不同推理节点在推理过程中可能会产生相同的Block,即同一个索引键对应的Block可能存在于不同推理节点中,因此,持有该索引键对应的Block的不同推理节点会先后向中心节点上报,中心节点每收到一个推理节点上报的关于该索引键的位置信息,就需要刷新一次全局索引,导致全局索引的刷新次数过多,中心节点的计算压力和通信压力过大

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838006A_ABST
    Figure CN122838006A_ABST
Patent Text Reader

Abstract

This application provides an inference system, inference method, and computer program product. The inference system includes M inference nodes and at least two proxy nodes. The M inference nodes perform inference operations and each stores multiple data blocks generated during the inference operation. Each data block corresponds to a key-value pair. The first proxy node among the at least two proxy nodes is used to obtain the first key-value pair of the first data block stored in N inference nodes among the M inference nodes and the position information of the first data block in the N inference nodes, where N is less than M. The first proxy node is also used to send the first key-value pair and the N position information corresponding to the N inference nodes to the central node. Based on the solution of this application, the computational and communication pressure on the central node can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a reasoning system, reasoning method and computer program product. Background Technology

[0002] Key-value caching (KV Cache) is a caching technique used to improve the inference speed of large language models. Its basic idea is to cache the intermediate results generated during the model inference process so that the cached intermediate results can be directly used in subsequent inference processes without repeated calculations, thereby accelerating the model inference process.

[0003] like Figure 1 As shown, current inference systems have incorporated KV Cache technology. Each inference node caches intermediate results locally during model inference. For ease of management and use, these cached intermediate results are typically divided into data blocks of a certain size. Each block is identified by an index key, and blocks with different contents have different index keys. Each inference node reports the index keys and location information of its held blocks to the central node. The central node refreshes the global index based on the information reported by all inference nodes. The global index includes the index keys and location information of blocks in all inference nodes, reflecting the block distribution of the entire system. When the central node receives an inference request, it can determine the appropriate inference node based on the global index and then schedule the inference request to the appropriate inference node for inference.

[0004] Since different inference nodes may generate the same block during the inference process, that is, the block corresponding to the same index key may exist in different inference nodes, different inference nodes holding the block corresponding to the index key will report to the central node one after another. Every time the central node receives the location information of the index key reported by an inference node, it needs to refresh the global index once, resulting in too many refreshes of the global index and too much computational and communication pressure on the central node. Summary of the Invention

[0005] This application provides an inference system, inference method, and computer program product that can reduce the computational and communication pressure on the central node.

[0006] In a first aspect, this application provides an inference system, including M inference nodes and at least two proxy nodes. The M inference nodes perform inference operations and each stores multiple data blocks generated during the inference operation, each data block corresponding to a key-value pair. A first proxy node among the at least two proxy nodes is used to obtain the first key-value pair of the first data block stored in N inference nodes among the M inference nodes, and the position information of the first data block in the N inference nodes, where N is less than M. The first proxy node is also used to send the first key-value pair and the N position information corresponding to the aforementioned N inference nodes to a central node.

[0007] In the above scheme, N inference nodes store a first data block, and the key-value information corresponding to the first data block is the first key-value information. The first proxy node obtains the location information of the first data block in the N inference nodes, thus obtaining N location information corresponding to the N inference nodes. Then, the first proxy node sends the first key-value information and the aforementioned N location information to the central node. The central node then collects the location information of the first data block in the N inference nodes, enabling it to globally schedule inference requests based on this information. Compared to each of the N inference nodes sending its stored first data block location information to the central node, this scheme eliminates the need for the central node to communicate with each of the N inference nodes separately to obtain N location information. Instead, the central node collects the location information of the first data block in the N inference nodes all at once from the first proxy node. The central node can record the first key-value information and N location information of the first data block together, avoiding repeated refreshing of the information recorded by the central node, thereby reducing the communication and computational pressure on the central node.

[0008] Based on the first aspect, in a possible implementation, the central node is used to add the first key-value information and N location information to the global information, the global information including the key-value information and location information of the data blocks stored in the M inference nodes; the first agent node is also used to add the first key-value information and N location information to the first fragment corresponding to the first agent node, the first fragment including part of the information in the global information.

[0009] In the above scheme, the central node manages global information, which includes the key-value information and location information of data blocks in M ​​inference nodes. When the central node receives the first key-value information and N location information sent by the first agent node, it can add the received information to the global information, thereby refreshing (updating) the global information. This allows the central node to globally schedule inference requests based on the global information; that is, the central node selects a suitable inference node for the inference request based on the global information and schedules the inference request to the appropriate inference node for inference. In addition, the first agent node manages a first shard, which includes a portion of the global information. The first agent node can add the acquired first key-value information and N location information to the first shard to support the central node or inference nodes in obtaining the required information from the first agent node's first shard.

[0010] For example, when a central node loses global information due to a failure, since the first shard includes a portion of the global information, the central node can obtain the information from the first proxy node. This allows the central node to quickly recover the global information, improving its fault recovery efficiency and reducing its cost. As another example, when an inference node needs to retrieve a first data block from other inference nodes, it can send a data retrieval request carrying the first key-value information of the first data block to the first proxy node. The first proxy node then retrieves the location information of the first data block from the first shard within at least one of the N inference nodes based on the data retrieval request, and sends this location information to the inference node so that the inference node can retrieve the first data block from it.

[0011] Based on the first aspect, in a possible implementation, the central node is used to determine whether the first key-value information already exists in the central node after receiving the first key-value information. If it exists, the central node determines the second proxy node that has reported the first key-value information among at least two proxy nodes, and notifies the first proxy node to send the first key-value information and N location information to the second proxy node. The first proxy node is also used to send the acquired first key-value information and N location information to the second proxy node according to the notification, and delete the first key-value information and N location information recorded by the first proxy node.

[0012] In the above scheme, after receiving the first key-value information sent by the first proxy node, the central node checks whether the first key-value information already exists on the central node. If it does, it means that other proxy nodes have already reported the first key-value information to the central node before the first proxy node. Both the first proxy node and other proxy nodes have recorded the first key-value information and its corresponding location information, resulting in duplicate information and wasted storage resources. To save storage resources, the central node determines that the second proxy node is one that reported the first key-value information before the first proxy node. Then, the central node instructs the first proxy node to send the first key-value information and N location information it obtained to the second proxy node, and instructs the first proxy node to delete the recorded first key-value information and N location information to achieve deduplication. When the second proxy node receives the first key-value information and N location information sent by the first proxy node, it records these N location information and the correspondence between the first key-value information and these N location information. At this point, the corresponding location information of the first data block is all gathered at the second proxy node.

[0013] Based on the first aspect, in a possible implementation, the first inference node among the N inference nodes corresponds to the first proxy node, and the second inference node among the N inference nodes corresponds to the second proxy node among at least two proxy nodes; the first proxy node obtains the first key value information of the first data block stored in the N inference nodes and the position information of the first data block in the N inference nodes, specifically including: receiving the first key value information of the first data block and the first position information of the first data block in the first inference node sent by the first proxy node; receiving the first key value information and the second position information of the first data block in the second inference node sent by the second proxy node; and recording the correspondence between the first key value information and the first position information and the second position information.

[0014] In the above scheme, inference nodes and proxy nodes have a corresponding relationship. An inference node sends the key-value information and location information of the data block stored in its proxy node to the corresponding proxy node. Specifically, the first inference node corresponds to the first proxy node, and the second inference node corresponds to the second proxy node. Therefore, the first inference node sends the first key-value information and the first location information of the first data block in its first inference node to the first proxy node, and the second inference node sends the first key-value information and the second location information of the first data block in its second inference node to the second proxy node. To allow the corresponding location information of the first data block to be aggregated at the first proxy node, the second proxy node can send the first key-value information and the second location information to the first proxy node. In other words, of the N location information obtained by the first agent node, some are obtained from the corresponding inference node, and some are obtained from other agent nodes. Once the first agent node obtains these N location information, it can send them to the central node in a unified manner, instead of having multiple agent nodes or N communication nodes send these N location information to the central node, which helps to reduce the communication and processing pressure on the central node.

[0015] Based on the first aspect, in a possible implementation, the first proxy node is further configured to send the fragmentation information of the first fragment to the central node; the central node is further configured to, when instructing the first inference node among the M inference nodes to perform the first inference operation, carry the fragmentation information of the first fragment to which the first data block generated by the first inference operation belongs and the information of the first proxy node corresponding to the first fragment; the first inference node is further configured to, after generating the first data block, send the first key value information of the first data block and the position information of the first data block in the first inference node to the first proxy node according to the fragmentation information of the first fragment and the information of the first proxy node.

[0016] In the above scheme, the first proxy node not only sends the acquired first key-value information and N location information to the central node, but also sends the fragmentation information of the first fragment corresponding to the first proxy node to the central node, indicating to the central node that the first fragment of the first proxy node records the first key-value information. When the central node instructs the first inference node to perform the first inference operation (the first data block generated by the first inference operation), it can also send the fragmentation information of the first fragment to which the first data block belongs (i.e., the first fragment records the first key-value information of the first data block) and the information of the first proxy node to the first inference node, instructing the first inference node to send the first key-value information and location information of the generated first data block to the first proxy node, so that the corresponding location information of the first data block can be aggregated in the first fragment of the first proxy node. Based on the fragmentation information of the first fragment and the information of the first proxy node, the first inference node can directly send the first key-value information and location information of the first data block generated by the first inference node to the first proxy node, without needing other proxy nodes to forward it to the first proxy node.

[0017] Based on the first aspect, in a possible implementation, the first agent node manages the first shard, and the first shard is set with key-value features of key-value information managed by the first agent node; the first inference node among the M inference nodes is used to determine the first key-value feature corresponding to the first key-value information based on the first key-value information of the first data block after the first data block is generated, determine the first shard managed by the first agent node based on the first key-value feature, and send the first key-value information and the position information of the first data block in the first inference node to the first agent node.

[0018] In the above scheme, the first shard is configured with key-value features for key-value information managed by the first proxy node. For any key-value information, the following rule can be used to determine whether the key-value information is managed by the first proxy node: if the key-value feature corresponding to the key-value information is the key-value feature set by the first shard, then it can be determined that the key-value information is managed by the first proxy node, and the key-value information belongs to the first shard corresponding to the first proxy node. Based on the above rule, the first inference node determines that the first key-value information belongs to the first shard based on the first key-value feature of the first key-value information of the first data block it generates. Therefore, the first key-value information and the position information of the first data block in the first inference node can be directly sent to the first proxy node without the need for other proxy nodes to forward it to the first proxy node, nor is it necessary for the central node to send the shard information of the first shard and the information of the first proxy node to the first inference node. The entire implementation process is relatively simple.

[0019] Based on the first aspect, in a possible implementation, the first inference node among the M inference nodes is used to send a data acquisition request to the first proxy node, the data acquisition request including first key-value information; the first proxy node is also used to send the location information of the first data block in at least one of the N inference nodes to the first inference node according to the data acquisition request; the first inference node is also used to acquire the first data block from at least one inference node according to the location information of the first data block in at least one inference node.

[0020] In the above scheme, since the first shard of the first proxy node records the first key-value information of the first data block and the N location information of the first data block among the N inference nodes, when an inference node wants to obtain the first data block, it can send a data acquisition request carrying the first key-value information to the first proxy node. Then, the first proxy node, based on the first key-value information in the data acquisition request, obtains the location information of the first data block in at least one of the N inference nodes from the first shard and sends the location information of the first data block in the at least one inference node to that inference node. The at least one inference node may include one or more inference nodes. When an inference node receives the location information of the first data block in the at least one inference node, it can choose to obtain the first data block from one of the at least one inference nodes. Compared to the inference node sending a data acquisition request to the central node to obtain the location information of the first data block recorded in the global information, the inference node in the above scheme only needs to interact with the proxy node and does not need to interact with the central node, which can reduce the communication pressure on the central node.

[0021] Secondly, this application provides an inference method for an inference system comprising M inference nodes and at least two proxy nodes. The method comprises: the M inference nodes performing inference operations and each storing multiple data blocks generated during the inference operations, each data block corresponding to a key-value information; a first proxy node among the at least two proxy nodes acquiring the first key-value information of the first data block stored in N inference nodes among the M inference nodes and the position information of the first data block in the N inference nodes, where N is less than M; and the first proxy node sending the first key-value information and the N position information corresponding to the N inference nodes to a central node.

[0022] Based on the second aspect, in a possible implementation, the central node adds the first key-value information and N location information to the global information, which includes the key-value information and location information of the data blocks stored in the M inference nodes; the first agent node adds the first key-value information and N location information to the first shard corresponding to the first agent node, which includes some information from the global information.

[0023] Based on the second aspect, in a possible implementation, after receiving the first key-value information, the central node determines whether the first key-value information already exists in the central node. If it does, it identifies the second proxy node among at least two proxy nodes that has reported the first key-value information and notifies the first proxy node to send the first key-value information and N location information to the second proxy node. The first proxy node sends the acquired first key-value information and N location information to the second proxy node according to the notification and deletes the first key-value information and N location information recorded by the first proxy node.

[0024] Based on the second aspect, in a possible implementation, the first inference node among the N inference nodes corresponds to the first proxy node, and the second inference node among the N inference nodes corresponds to the second proxy node among at least two proxy nodes. The first proxy node obtains the first key value information and the position information of the first data block stored in the N inference nodes among the M inference nodes, including: receiving the first key value information and the first position information of the first data block in the first inference node sent by the first proxy node; receiving the first key value information and the second position information of the first data block in the second inference node sent by the second proxy node; and recording the correspondence between the first key value information and the first position information and the second position information.

[0025] Based on the second aspect, in a possible implementation, the first proxy node sends the fragmentation information of the first fragment to the central node; when the central node instructs the first inference node among the M inference nodes to perform the first inference operation, it carries the fragmentation information of the first fragment to which the first data block generated by the first inference operation belongs and the information of the first proxy node corresponding to the first fragment; after generating the first data block, the first inference node sends the first key value information of the first data block and the position information of the first data block in the first inference node to the first proxy node according to the fragmentation information of the first fragment and the information of the first proxy node.

[0026] Based on the second aspect, in a possible implementation, the first agent node manages the first shard, and the first shard is set with key value features of key value information managed by the first agent node. After generating the first data block, the first inference node among the M inference nodes determines the first key value feature corresponding to the first key value information based on the first key value information of the first data block, determines the first shard managed by the first agent node based on the first key value feature, and sends the first key value information and the location information of the first data block in the first inference node to the first agent node.

[0027] Based on the second aspect, in a possible implementation, the first inference node among the M inference nodes sends a data acquisition request to the first proxy node, the data acquisition request including first key-value information; the first proxy node sends the location information of the first data block in at least one of the N inference nodes to the first inference node according to the data acquisition request; the first inference node acquires the first data block from at least one inference node according to the location information of the first data block in at least one inference node.

[0028] Thirdly, this application provides a computer program product, including instructions that, when executed by an inference system, cause the inference system to perform the method of the second aspect or any possible implementation thereof. Attached Figure Description

[0029] Figure 1 This is an architecture diagram of a reasoning system in a related technology.

[0030] Figure 2 This is an architecture diagram of an inference system provided in an embodiment of this application;

[0031] Figure 3 This is an architecture diagram of another inference system provided in an embodiment of this application;

[0032] Figure 4 This is an interactive flowchart provided in an embodiment of this application;

[0033] Figure 5 This is a schematic diagram illustrating the positional relationship between inference nodes provided in an embodiment of this application;

[0034] Figure 6 This is a flowchart illustrating a reasoning method provided in an embodiment of this application. Detailed Implementation

[0035] Please see Figure 2 , Figure 2 This is an architecture diagram of an inference system provided in an embodiment of this application, including a central node, at least two agent nodes and M inference nodes, where M is a positive integer greater than 1.

[0036] An inference node is an executor within an inference system. It deploys an inference model, such as a Large Language Model (LLM) based on self-attention or other models. Inference nodes also possess usable computational and storage resources, enabling them to perform complete inference operations on inference prompts based on the inference model and these resources to obtain inference results. Different inference nodes can run in parallel, and different inference nodes have different identifiers (IDs). The aforementioned computational resources may include a central processing unit (CPU), a graphics processing unit (GPU), a neural network processing unit (NPU), a tensor processing unit (TPU), etc. This application does not specifically limit the type or number of processors. Among these, processors other than the CPU are typically implemented as accelerator cards. An accelerator card is a piece of hardware used to accelerate computation, including a processor and memory, and supports flexible plug-and-play functionality on the interface of a computing device. Alternatively, the processors can be packaged as chips, which can be installed inside a computing device. The aforementioned storage resources may include high-bandwidth memory (HBM), dynamic random access memory (DRAM), solid-state disk (SSD), or other types of storage media, and this application does not specifically limit them.

[0037] Optionally, one or more inference nodes can be deployed on a single computing device, with the computing and storage resources of the device allocated to these nodes. Each inference node can perform inference operations based on the allocated computing and storage resources. Alternatively, an inference node can be deployed across multiple computing devices, performing inference operations based on the computing and storage resources of these multiple devices. (About...) Figure 2 The number M of inference nodes in the inference system is not specifically limited in this application. The type of computing device is also not specifically limited in this application; for example, it can be a server, a desktop computer, or other physical device with computing and storage capabilities.

[0038] Optional, Figure 2In the inference system, inference nodes and agent nodes can be deployed on different computing devices, in which case communication between the inference nodes and agent nodes needs to be across computing devices; alternatively, inference nodes and agent nodes can be deployed on the same computing device. This application does not impose a specific limit on the number of agent nodes; for example, the number of agent nodes can be less than the number of inference nodes, M.

[0039] Optionally, the central node and the agent node can be deployed on different computing devices, in which case communication between the central node and the agent node needs to cross computing devices; or, the central node and the agent node can be deployed on the same computing device, in which case the central node and the agent node communicate within the same computing device, and the communication distance is relatively short. For information on the types of computing devices, please refer to the previous section, which will not be repeated here.

[0040] To accelerate the inference process, the inference node employs a storage-for-computation technique, exemplified by KV Cache. Specifically, during the inference operation for a given inference request, the inference node can store the intermediate results. These stored intermediate results can then be directly used (reused) in subsequent inference operations for that request, avoiding redundant computation and thus accelerating the overall inference process. The stored intermediate results can be reused not only in subsequent inference operations for the current request but also in inference operations for other requests, which may be performed by the same inference node or by other inference nodes. The content of these intermediate results is not specifically limited in this application. For example, when the inference model used by the inference node is based on an attention mechanism, the stored intermediate results include the K-vector and V-vector mapped to each token in the token sequence of the inference request. The K-vector and V-vector of each token are information used to calculate the attention score during the inference operation, and these stored K-vectors and V-vectors of the tokens are referred to as KV Cache data. In this context, a token is the basic unit (smallest processing unit) for text processing by the model. It can be a word, character, sub-word, punctuation mark, or character, depending on the design and training method of the model. This application does not impose any specific limitations on it.

[0041] To facilitate management and reuse, intermediate results stored by inference nodes are typically divided into fixed-size data blocks, with each block serving as the basic unit of reuse (i.e., the smallest reuse granularity is one block). Each block corresponds to a key-value pair. For example, when the intermediate results stored by the inference node are KV Cache data, the KV Cache data can be divided into fixed-size blocks, commonly referred to as KV Cache Blocks or KV Blocks. Each KV Block includes KV Cache data corresponding to multiple consecutive tokens (this application does not limit the number of tokens here). Each KV Block has corresponding key-value pairs, which can be index keys or other information that can identify the corresponding KV Block. The key-value pairs are used to identify the content of the KV Block. KV Blocks with the same key-value pairs generally have the same KV Cache data, and KV Blocks with different key-value pairs have different KV Cache data. This application does not specifically limit the size of the blocks and can set it according to usage requirements. This application also does not specifically limit how the key-value pairs of a block are determined. For example, for a KV Block, the prefix hash values ​​of the multiple tokens corresponding to the KV Block can be calculated, and the prefix hash values ​​can be used as the key-value information of the KV Block.

[0042] The central node is the entry point for inference requests to enter the inference system. For example, inference requests generated by clients or other systems can be sent to the central node in the inference system. The central node is responsible for the global scheduling of inference requests, that is, scheduling the received inference requests to appropriate inference nodes, and then the inference nodes perform inference operations on the inference requests to obtain inference results. The block distribution on the M inference nodes in the inference system is an important factor for the central node to consider when performing global scheduling. There may also be other factors to consider, such as the load of the inference nodes, etc., which are not specifically limited in this application.

[0043] To facilitate global scheduling based on the block distribution across the entire system, a central node maintains global information, which includes the key-value information and location information of blocks in M ​​inference nodes. When the global information is constructed in the form of an index, it can be called a global index, which includes multiple key-value information and the value corresponding to each key-value information. Each key-value information and its corresponding value form a key-value pair. The value corresponding to each key-value information includes the location information of the block corresponding to that key-value information. The location information includes information about the inference node that stores the block corresponding to that key-value information (such as its ID), and may also include the type of storage medium used by the corresponding inference node to store the block corresponding to that key-value information. It may also include the identifier of the storage area of ​​the block corresponding to that key-value information in the corresponding storage medium. The identifier of the storage area can be the physical address of the storage area or information that can be mapped to the physical address of the storage area. This application does not specifically limit this. When the central node needs to schedule inference requests, it can generate some key-value information based on the token sequence of the inference request. The central node can then query global information to determine which inference nodes the corresponding key-value information is located on. It may also further determine which storage media the corresponding inference node uses to store the aforementioned block and the identifier of the storage area where the block is located. Then, based on the query results and the scheduling algorithm (this application does not specifically limit this), the central node selects a suitable inference node for the inference request and assigns the inference request to the selected inference node for inference.

[0044] The above global information is updated (refreshed) by the central node based on the key-value information and location information obtained and sent by the proxy node. The following uses five possible implementation methods as examples to introduce how the proxy node obtains and sends key-value information and location information to the central node.

[0045] In the first possible implementation, at least two agent nodes in the inference system each manage a corresponding shard. The shard records the location and key-value information of the Block corresponding to the key-value information managed by the respective agent node. Since the key-value information managed by different agent nodes is different, each shard includes a portion of the global information. Each shard is set with key-value characteristics for the key-value information managed by the corresponding agent node, and each shard can have one or more key-value characteristics. For each inference node, a Block is generated during inference operations. The inference node can determine the key-value characteristics corresponding to the Block based on its key-value information, then determine the shard to which the Block's key-value information belongs based on the key-value characteristics, and finally send the Block's key-value information and its location information within that agent node to the agent node corresponding to the aforementioned shard. After receiving the key-value information and corresponding location information sent by the inference node, the agent node adds the key-value information and location information to the shard managed by the agent node, and sends the key-value information and location information to the central node so that the central node can add the key-value information and location information to the global information, thereby realizing the update of the global information.

[0046] This application does not impose specific limitations on how to determine the key-value characteristics of key-value information. For example, a hash algorithm can be used to calculate the hash value of the key-value information, and then the hash value can be used as the key-value characteristic of the key-value information. This application does not impose specific limitations on the choice of hash algorithm.

[0047] This application does not specify which key-value information a proxy node manages. Taking hash values ​​as an example, different hash values ​​can be set for shards of different proxy nodes. Each shard can have multiple hash values ​​set. Each hash value set for a shard is the hash value of the key-value information managed by the proxy node corresponding to that shard. That is, by setting the corresponding hash value for the shard corresponding to the proxy node, it is possible to indirectly specify which key-value information the proxy node can manage. The key-value information managed by the proxy node belongs to the shard corresponding to the proxy node.

[0048] It should be understood that in the first possible implementation, there is no strict correspondence between the inference node and the proxy node. The inference node does not send the key-value information and location information of the block it generates to a fixed proxy node. Instead, it determines which proxy node to send the information to based on the key-value characteristics of the block's key-value information: if the key-value characteristics of the block's key-value information are the key-value characteristics set by a shard managed by a certain proxy node, then it means that the key-value information of the block is managed by that proxy node and belongs to the shard managed by that proxy node. Therefore, the inference node sends the key-value information and location information of the block to that proxy node.

[0049] When a single inference node generates multiple blocks, and the inference node determines, based on the aforementioned rules, that the key-value information of these blocks is managed by different proxy nodes, then the inference node will send the key-value information and location information of these blocks to the corresponding proxy nodes. In other words, each proxy node will only receive the location and key-value information of the blocks corresponding to the key-value information managed by itself from the inference node, and will not receive the location and key-value information of blocks corresponding to key-value information managed by other proxy nodes, nor will it need to communicate with other proxy nodes. Proxy nodes can directly add the key-value information and location information sent by the inference node to their corresponding shards without needing to determine whether the key-value information sent by the inference node is managed by themselves before adding it, thus simplifying the operation of the proxy nodes. Each inference node sends the key-value information obtained by itself (i.e., the key-value information managed by itself) and the corresponding location information to the central node, and the key-value information sent by different inference nodes to the central node is not duplicated, so that the central node can update the global information.

[0050] For example, such as Figure 2As shown, proxy node 1 manages shard 1, and proxy node 2 manages shard 2. Each shard has key-value features for the key-value information managed by the corresponding proxy node, and each shard has multiple key-value features. Assume that inference node A generates Block1 and Block2 during inference operations. Inference node A calculates a first key-value feature based on the first key-value information of Block1 and a second key-value feature based on the second key-value information of Block2. The first key-value feature is one of the key-value features set by shard 1 managed by proxy node 1, and the second key-value feature is one of the key-value features set by shard 1 managed by proxy node 2. Therefore, inference node A sends the first key-value information of Block1 and its location information within inference node A to proxy node 1, and sends the second key-value information of Block2 and its location information within inference node A to proxy node 2. Assuming that inference node C also generates Block1 when performing inference operations, inference node C calculates the first key value feature based on the first key value information of Block1. The first key value feature is one of the key value features set by shard 1 managed by proxy node 1. Therefore, inference node C sends the first key value information of Block1 and the location information of Block1 in inference node C to proxy node 1.

[0051] When proxy node 1 receives the first key-value information and location information of Block 1 sent by inference node A, and also receives the first key-value information and location information of Block 1 sent by inference node C, proxy node 1 adds the first key-value information and the two location information of Block 1 in inference nodes A and C to shard 1. Then, proxy node 1 sends the first key-value information and the two location information to the central node. The central node adds the received first key-value information and the two location information to the global information, thereby refreshing the global information. Subsequently, by querying the global information, the central node can determine that both inference nodes A and C store the data block (i.e., Block 1) corresponding to the first key-value information. It should be understood that if inference node A and inference node C send the first key value information and location information of their respective Block 1 to the central node, the central node will receive the above information sent by inference nodes A and C in a sequential order. The central node needs to refresh the global information twice, that is, refresh the global information once when it receives the first key value information and location information sent by inference node A, and refresh the global information once when it receives the first key value information and location information sent by inference node C. Too many global information refreshes will lead to a large computational burden on the central node. In this example, the central node only needs to refresh the global information once based on the first key value information and the above two location information sent by proxy node 1, which can reduce the number of global information refreshes and thus reduce the computational and communication burden on the central node.

[0052] When proxy node 2 receives the second key-value information and location information of Block 2 from inference node A, proxy node 2 adds the second key-value information and the location information of Block 2 in inference node A to shard 2. Then, proxy node 2 sends the second key-value information and the location information of Block 2 in inference node A to the central node. The central node adds the received second key-value information and the location information of Block 2 in inference node A to the global information, thereby refreshing the global information. Subsequently, the central node can determine that the data block (i.e., Block 2) corresponding to the second key-value information is stored in inference node A by querying the global information.

[0053] Optionally, the shards of the proxy nodes can be constructed in an indexed manner. Each shard includes key-value information managed by the corresponding proxy node and the corresponding value. Each key-value pair consists of a key and a corresponding value. The value corresponding to each key-value pair includes the location information of the Block corresponding to that key-value pair. The location information includes information about the inference node storing the Block (such as the ID of the inference node), and may also include the type of storage medium used by the inference node to store the Block. It may also include the identifier of the storage area of ​​the Block in the corresponding storage medium. The identifier of the storage area can be the physical address of the storage area or information that can be mapped to the physical address of the storage area; this application does not specifically limit this. For ease of description, this application uniformly uses the form "XX→XX" to represent key-value pairs. The element to the left of the arrow represents the key (key-value information) of the key-value pair, and the element to the right of the arrow represents the value of the key-value pair.

[0054] For example, assuming IndexKey1 is the key-value information of Block1, and both inference node A and inference node C have Block1, the shard of the proxy node managing this key-value information can record the following key-value pair: IndexKey1 → {Location information of Block1 in inference node A, Location information of Block1 in inference node C}. Correspondingly, the global information of the central node can record the following key-value pair: IndexKey1 → {Information of the proxy node managing IndexKey1, {Location information of Block1 in inference node A, Location information of Block1 in inference node C}}.

[0055] Optionally, when an inference node sends the key-value information and corresponding location information managed by the proxy node to the proxy node, it can do so in a full or incremental manner. Full sending means that the inference node sends the location information and key-value information of all blocks in its own inference node that contain key-value information managed by the proxy node to the proxy node each time. Incremental sending means that the inference node only sends the location information and key-value information of blocks in its own inference node that have changed since the last transmission (such as being added or deleted) and contain key-value information managed by the proxy node, thereby reducing the amount of communication data between the inference node and the proxy node and lowering communication overhead.

[0056] Optionally, the inference node can periodically or irregularly send the key-value information and corresponding location information managed by the proxy node to the proxy node. For example, the inference node can send the key-value information and corresponding location information managed by the proxy node to the proxy node at fixed time intervals. The above time interval can be set according to actual usage needs, such as 10 minutes, 1 hour, or 2 hours, etc., and this application does not make a specific limitation on it. Furthermore, the inference node will only send the key-value information and corresponding location information managed by the proxy node to the proxy node when it generates a Block corresponding to the key-value information managed by the proxy node. If the inference node does not generate a Block corresponding to the key-value information managed by the proxy node within a certain period, it will not send the key-value information and corresponding location information to the proxy node during that period.

[0057] Optionally, when a proxy node sends the key-value information and corresponding location information of the shards it manages to the central node, it can do so incrementally or in full. Full sending means that the proxy node sends all the key-value information and location information of the shards it manages to the central node each time. Incremental sending means that the proxy node only sends the location information and corresponding key-value information of the shards it manages that have changed compared to the last time it sent to the central node each time.

[0058] As introduced above, each agent node in the inference system manages its own corresponding shard. Each shard includes a portion of the global information of the central node. Therefore, when the central node loses global information due to a failure, it can obtain information from each agent node's shards and quickly restore the global information based on the information obtained from each shard, resulting in high recovery efficiency.

[0059] For example, such as Figure 2As shown, assuming the central node loses global information due to a failure and restart, to recover the global information, the central node can send information retrieval instructions to proxy nodes 1 and 2, instructing them to send the information from their respective managed shards to the central node. Proxy node 1 sends the information from shard 1 (including key-value information and corresponding location information) to the central node according to the instructions, and proxy node 2 sends the information from shard 2 to the central node according to the instructions. Shards 1 and 2 can be considered as a copy of the global information, but this copy is distributed across the two proxy nodes in the form of shards. The central node records the received information from shard 1 and shard 2, thereby recovering the global information. Since the key-value information managed by proxy nodes 1 and 2 is only a part of the key-value information in the global information, the amount of key-value information managed by each proxy node is relatively small. Even if a proxy node loses information from a shard due to a failure, it can recover the information by recollecting the location information of the Blocks corresponding to the key-value information it manages in the inference node. Therefore, the failure recovery speed of proxy nodes is fast.

[0060] In the second possible implementation, at least two proxy nodes in the inference system correspond to M inference nodes, with each proxy node corresponding to at least one inference node, and the inference nodes corresponding to different proxy nodes do not overlap. For each proxy node, the proxy node obtains the key-value information and location information of the Block stored in the corresponding inference node. This obtaining can refer to the proxy node actively pulling the key-value information and location information of the Block stored in the corresponding inference node, or it can refer to the inference node actively sending the key-value information and location information of the Block stored in its own inference node to the corresponding proxy node. In other words, the object that triggers the inference node to send the key-value information and location information of the Block stored in its own inference node to the corresponding proxy node can be the inference node (as the sender) or the proxy node (as the receiver), and this application does not specifically limit this. After the proxy node obtains the key-value information and location information of the Block stored in the corresponding inference node, the proxy node records the key-value information and location information and sends the obtained key-value information and location information to the central node. When the central node receives the key-value information and location information sent by the agent node, it can add the key-value information and location information to the global information to refresh the global information.

[0061] Optionally, the proxy node may retrieve the key-value information and location information of the Block stored in the corresponding inference node periodically or irregularly. For example, the proxy node may retrieve the key-value information and location information of the Block stored in the corresponding inference node at fixed time intervals. The time interval can be set according to actual usage needs, such as 10 minutes, 1 hour, or 2 hours, etc., and this application does not make a specific limitation on this. Alternatively, the inference node may only send the key-value information and location information of the Block stored in its own inference node to the corresponding proxy node when the Block it stores changes (such as adding or deleting a Block). Therefore, the proxy node's retrieval of key-value information and location information from the corresponding inference node is irregular.

[0062] Optionally, when the proxy node obtains the key-value information and location information of the Block stored in the corresponding inference node, it can do so in a full or incremental manner. The full method means that the proxy node obtains the location information and key-value information of all Blocks stored in the corresponding inference node each time. The incremental method means that the proxy node only obtains the key-value information and location information of the Blocks that have changed (such as added or deleted) from the corresponding proxy node each time, so as to reduce the amount of communication data between the inference node and the proxy node and reduce communication overhead.

[0063] It should be understood that in the second possible implementation, the proxy node and the inference node have a corresponding relationship. The inference node only needs to send the key-value information and location information of the block it generates to the corresponding proxy node, and does not need to send them to other proxy nodes. Therefore, before sending the key-value information and location information of the block it generates, the inference node does not need to determine which proxy node manages the key-value information, which simplifies the operation of the inference node.

[0064] Because different proxy nodes may generate blocks with the same key-value information, and each inference node only sends the key-value information and location information of the block stored in its own inference node to the corresponding proxy node, different proxy nodes may obtain and record the same key-value information and corresponding location information, meaning that there may be duplication between the information recorded by different proxy nodes. To save storage space, it is necessary to deduplicate the information recorded by different proxy nodes. This deduplication can be performed using a central node. When the central node receives the key-value information and corresponding location information sent by a proxy node, it determines whether the key-value information already exists in the central node. If it does, it means that another proxy node has already sent the key-value information and corresponding location information to the central node before that proxy node. The central node can then notify that proxy node to send its recorded key-value information and corresponding location information to the aforementioned other proxy node. When the proxy node receives the notification from the central node, it sends its recorded key-value information and corresponding location information to the aforementioned other proxy node according to the notification, and deletes its recorded key-value information and corresponding location information, thus achieving deduplication. When the other proxy node receives the key-value information and corresponding location information recorded by the other proxy node, it can record the key-value information and corresponding location information, so that all the key-value information and corresponding location information are gathered in the other proxy node.

[0065] For example, such as Figure 3As shown, assuming proxy node 1 corresponds to inference nodes A and B, and proxy node 2 corresponds to inference nodes C and D, each inference node sends the key-value information and location information of the Block stored in its inference node to its corresponding proxy node. When inference node A performs an inference operation, it generates Block1 and Block2. Then, inference node A sends the key-value information of these two Blocks and their location information within inference node A to its corresponding proxy node 1. Proxy node 1 records the key-value information and their location information within inference node A in shard 1 and sends these information to the central node. The central node determines whether the key-value information of the two Blocks mentioned above exists in the global information. Here, it is assumed that the key-value information of the two Blocks does not yet exist in the global information. Therefore, the central node adds the key-value information of the two Blocks and their position information in inference node A to the global information, and records the correspondence between the key-value information of the two Blocks and the information of agent node 1 (such as the ID of agent node), so as to indicate that agent node 1 is the first agent node to send the key-value information of the two Blocks to the central node, and the key-value information of the two Blocks is managed by agent node 1.

[0066] Then, during the inference operation, inference node C generates Block2 and Block5. Inference node C sends the key-value information of these two blocks and their location information within inference node C to the corresponding proxy node 2. Proxy node 2 records the key-value information of these two blocks and their location information within inference node C in shard 2, and then sends this information to the central node. The central node determines whether the key-value information of Block2 and Block5 sent by proxy node 2 already exists in the central node. Since the key-value information of Block5 does not exist in the central node, the central node adds the key-value information of Block5 sent by proxy node 2 and its location information within inference node C to the global information, and records the correspondence between the key-value information of Block5 and the information of proxy node 2, indicating that proxy node 2 is the first proxy node to send the key-value information of Block5 to the central node. Since the key-value information of Block2 already exists in the central node, and the central node records the correspondence between the key-value information of Block2 and the information of proxy node 1, the central node determines that proxy node 1 has already sent the key-value information of Block2 before proxy node 2. Then the central node notifies proxy node 2 to send the key-value information of Block2 and the corresponding location information it has recorded to proxy node 1.

[0067] When proxy node 2 receives the aforementioned notification, it sends the key-value information of Block 2 recorded in shard 2 and the location information of Block 2 in inference node C to proxy node 1, and deletes the recorded key-value information of Block 2 and the location information of Block 2 in inference node C, thereby saving storage resources for proxy node 2. Proxy node 1 records the key-value information of Block 2 and the location information of Block 2 in inference node C sent by proxy node 2 in shard 1, realizing the refresh of shard 1. At this time, only shard 1 records the key-value information of Block 2 and the corresponding location information; other shards do not have the key-value information of Block 2 and the corresponding location information. All the location information of Block 2 is gathered in shard 1, including the location information of Block 2 in inference node A and the location information of Block 2 in inference node C.

[0068] It should be noted that, in the example above, when the central node receives the key-value information of Block2 and its location information in proxy node C from proxy node 2, the central node can add this information to the global information. Thus, the global information includes the location information of Block2 in inference node A and the location information of Block2 in inference node C. In other words, although proxy node 2 is not the first proxy node to send the key-value information of Block2 to the central node, the central node can still update the global information based on the location information of Block2 sent by proxy node 2.

[0069] Alternatively, the central node may not add the location information of Block2 in proxy node C sent by proxy node 2 to the global information. After proxy node 2 sends the key-value information of Block2 and its location information in inference node C to proxy node 1, proxy node 1 sends the key-value information of Block2 and its location information in inference node C to the central node. Then, the central node adds the location information of Block2 in inference node C sent by proxy node 1 to the global information. In other words, since proxy node 2 is not the first proxy node to send the key-value information of Block2 to the central node, but proxy node 1 is the first proxy node to send the key-value information of Block2 to the central node, and the key-value information of Block2 belongs to shard 1 of proxy node 1, the central node will not update the global information based on the location information of Block2 sent by proxy node 2, but only based on the location information of Block2 sent by proxy node 1.

[0070] In the third possible implementation, at least two agent nodes in the inference system each manage a corresponding shard. The shard records the location and key-value information of the Block corresponding to the key-value information managed by the respective agent node. The key-value information managed by different agent nodes is different; therefore, each shard includes a portion of the global information. Each shard is set with key-value characteristics for the key-value information managed by the corresponding agent node, and each shard can have one or more key-value characteristics. Furthermore, at least two agent nodes in the inference system correspond to M inference nodes, with each agent node corresponding to at least one inference node, and the inference nodes corresponding to different agent nodes do not overlap. For each agent node, the agent node obtains the key-value information and location information of the Block stored in the corresponding inference node. When a proxy node obtains the key-value information and location information of a Block stored in the corresponding inference node, it calculates the key-value characteristics of the key-value information to determine whether these characteristics are set by the shard managed by the proxy node. If so, it means the key-value information of the Block obtained by the proxy node is managed by the proxy node, and the proxy node records the obtained key-value information and corresponding location information in its shard and sends it to the central node so that the central node can update the global information. If not, it means the key-value information of the Block obtained by the proxy node is not managed by the proxy node, but by another proxy node, and the proxy node sends the obtained key-value information and corresponding location information to the other proxy node. The other proxy node records the key-value information and corresponding location information of the Block sent by the proxy node in its corresponding shard and sends it to the central node so that the central node can update the global information.

[0071] It should be understood that in the third possible implementation, the proxy nodes and inference nodes have a corresponding relationship. The inference node only sends the key-value information and location information of the Block stored in its own inference node to the corresponding proxy node, and does not send them to other proxy nodes. The operation of the inference node is relatively simple. Since different proxy nodes may generate Blocks with the same key-value information, and each inference node only sends the key-value information and location information of the Block stored in its own inference node to the corresponding proxy node, the proxy node will receive not only the key-value information and corresponding location information managed by its own proxy node, but also the key-value information and corresponding location information managed by other proxy nodes from the corresponding inference node. The key-value information and corresponding location information managed by its own proxy node should belong to the shard corresponding to its own proxy node, while the key-value information and corresponding location information managed by other proxy nodes should belong to the shard corresponding to those other proxy nodes. Therefore, the own proxy node can forward the key-value information and corresponding location information managed by other proxy nodes to those other proxy nodes, enabling those other proxy nodes to add the aforementioned key-value information and key-value features to their corresponding shards.

[0072] Based on the aforementioned forwarding rules, each proxy node can receive key-value information and corresponding location information managed by itself from the corresponding inference node, and may also receive key-value information and corresponding location information managed by itself from other proxy nodes. It then records the key-value information and corresponding location information managed by itself into its corresponding shard and sends it to the central node so that the central node can update the global information. Since the key-value information sent by each proxy node to the proxy node is managed by itself, and the key-value information managed by different proxy nodes does not overlap, the key-value information obtained by the central node from different proxy nodes will not be duplicated. Therefore, the central node does not need to deduplicate the information sent by different proxy nodes. The central node can directly add the information (including key-value information and location information) sent by each proxy node to the global information, thus reducing the computational burden on the central node when updating the global information.

[0073] For example, such as Figure 3As shown, assuming proxy node 1 corresponds to inference nodes A and B, and proxy node 2 corresponds to inference nodes C and D, each inference node sends the key-value information and location information of the Block stored in its inference node to its corresponding proxy node. Inference node A generates Block1 and Block2 during inference operations, and then sends the key-value information of these two Blocks, along with their location information within inference node A, to its corresponding proxy node 1. When proxy node 1 receives the information from inference node A, it calculates a first key-value feature based on the first key-value information of Block1 and a second key-value feature based on the second key-value information of Block2. Then, it determines the fragment to which the first and second key-value information belong based on these two features. Here, it is assumed that the first key value feature is one of the key value features set by shard 1 managed by proxy node 1, and the second key value feature is one of the key value features set by shard 2 managed by proxy node 2. Therefore, the first key value information belongs to shard 1, and the second key value feature belongs to shard 2. Then, proxy node 1 records the first key value information of Block 1 and the location information of Block 1 in inference node A in shard 1, sends the first key value information and the location information of Block 1 in inference node A to the central node, and forwards the second key value information of Block 2 and the location information of Block 2 in inference node A to proxy node 2.

[0074] Agent node 2 receives the second key-value information forwarded by agent node 1 and the location information of Block 2 in inference node A. It also receives the second key-value information and the location information of Block 2 in inference node C sent by the inference node C corresponding to agent node 2. Then, agent node 2 records the correspondence between the second key-value information and the two location information of Block 2 in inference node A and inference node C in shard 2, and sends the second key-value information and the two location information to the central node so that the central node can update the global information.

[0075] In the fourth possible implementation, at least two proxy nodes in the inference system correspond to M inference nodes, with each proxy node corresponding to at least one inference node, and the inference nodes corresponding to different proxy nodes do not overlap. When the central node receives an inference request, it can generate some key-value information based on the token sequence of the inference request, and determine whether each key-value information already exists in the central node. If the key-value information already exists in the central node, the central node determines the proxy node that has reported the key-value information, and the shard managed by the proxy node that has reported the key-value information records the key-value information. After determining the target inference node for the inference request based on global information and the scheduling algorithm, the central node sends the inference request to the target inference node to instruct the target inference node to perform the inference operation on the inference request, and sends the shard information (such as shard ID) of the shard to which the key-value information already exists in the central node belongs, as well as the information of the proxy node corresponding to the shard. When the target inference node performs the inference operation on the inference request, it generates a Block corresponding to the aforementioned key-value information. For any of these key-value information, if the central node has already sent the shard information of the shard to which the key-value information belongs and the information of the proxy node corresponding to the shard to the target inference node, then the target inference node can send the key-value information and the location information of the block corresponding to the key-value information in the target inference node to the aforementioned proxy node based on the aforementioned shard information and proxy node information. However, if the central node has not sent the shard information of the shard to which the key-value information belongs and the information of the proxy node corresponding to the shard to the target inference node, then the target inference node will send the key-value information and the location information of the block corresponding to the key-value information in the target inference node to the proxy node corresponding to the target inference node.

[0076] For any proxy node, when it receives key-value information and the location information of the corresponding Block in the target inference node from the target inference node, if the key-value information already exists in the shard managed by the proxy node, it means that the Block corresponding to the key-value information is not generated for the first time in the entire system, and the proxy node has previously recorded the key-value information and the location information of the corresponding Block in the inference node (which may be the target inference node or other inference nodes). Therefore, the proxy node adds the location information of the corresponding Block in the target inference node sent by the target inference node to the shard, and sends the key-value information and the location information of the corresponding Block in the target inference node to the central node so that the central node can update the global information. If the key-value information does not exist in the shard managed by the agent node, it means that the Block corresponding to the key-value information is generated for the first time in the entire system. Therefore, the agent node adds the key-value information sent by the target inference node and the location information of the Block corresponding to the key-value information in the target inference node to the shard, and sends the key-value information and the location information of the Block corresponding to the key-value information in the target inference node to the central node.

[0077] When the central node receives key-value information and corresponding location information from a proxy node, it determines whether the key-value information exists on the central node. If it exists, it means that the proxy node has previously reported the key-value information, and the central node has recorded the correspondence between the key-value information, the proxy node's information, and the shard information of the shards managed by the proxy node. Therefore, the central node only needs to add the location information sent by the proxy node this time to the global information. If the key-value information does not exist on the central node, it means that the proxy node was the first to report the key-value information to the central node. Therefore, the central node adds the key-value information and corresponding location information sent by the proxy node to the global information and records the correspondence between the key-value information, the proxy node's information, and the shard information of the proxy node, indicating that the key-value information belongs to the shards managed by the proxy node.

[0078] For example, such as Figure 2As shown, assume inference nodes A and B correspond to proxy node 1, and inference nodes C and D correspond to proxy node 2. When the central node receives a certain inference request, it calculates the first key-value information and the second key-value information based on the token sequence of the inference request. The central node determines that the first key-value information exists in the global information but the second key-value information does not exist. The central node also records the correspondence between the first key-value information and proxy node 1 and shard 1. This indicates that the block corresponding to the first key-value information has already been generated in the entire system, and shard 1 of proxy node 1 records the first key-value information, but the block corresponding to the second key-value information has not yet been generated in the entire system. Assuming the target inference node determined by the central node for this inference request is inference node A, the central node sends the inference request, the shard information of shard 1 corresponding to the first key-value information, and the information of proxy node 1 to inference node A, instructing inference node A to perform an inference operation on the inference request. The central node also sends the location information and key-value information of the block generated during the inference operation to proxy node 1 for recording in shard 1.

[0079] When inference node A executes the inference operation on the inference request, it generates a Block corresponding to the first key-value information and a Block corresponding to the second key-value information. Then, based on the shard information of shard 1 corresponding to the first key-value information sent by the central node and the information of proxy node 1, inference node A sends the location information of the first key-value information and the Block corresponding to the first key-value information within inference node A to proxy node 1. However, inference node A has not received the shard information of the shard corresponding to the second key-value information from the central node. This indicates that inference node A is the first inference node to generate the Block corresponding to the second key-value information, and currently no shard of any proxy node records the second key-value information; that is, it is not yet determined which shard the second key-value information belongs to. Therefore, inference node A can send the second key-value information and the location information of the Block corresponding to the second key-value information within inference node A to the proxy node (i.e., proxy node 1) corresponding to inference node A.

[0080] When proxy node 1 receives the first key-value information and the location information of the corresponding block in inference node A (denoted as the first location information) sent by inference node A, since the first key-value information already exists in shard 1 of proxy node 1, proxy node 1 only needs to add the first location information to shard 1 and record the correspondence between the first key-value information and the first location information. When it receives the second key-value information and the location information of the corresponding block in inference node A (denoted as the second location information) sent by inference node A, since the second key-value information does not exist in shard 1 of proxy node 1, proxy node 1 needs to add both the second key-value information and the second location information to shard 1 and record the correspondence between the second key-value information and the second location information. At this time, the second key-value information is determined to belong to shard 1.

[0081] Agent node 1 sends the first key-value information and first location information, and the second key-value information and second location information to the central node. The central node determines that the first key-value information already exists in the global information, so it only needs to add the first location information to the global information and record the correspondence between the first key-value information and the first location information. The central node determines that the second key-value information does not exist in the global information, so it adds both the second key-value information and the second location information to the global information, records the correspondence between the second key-value information and the second location information, and records the correspondence between the second key-value information and the shard information of shard 1 and the information of agent node 1, indicating that the second key-value information belongs to shard 1 of agent node 1.

[0082] Subsequently, the central node receives another inference request. The central node determines that the target inference node for this request is inference node D, and that the key-value information calculated from the token sequence of this inference request includes second key-value information. Therefore, the central node sends this inference request to inference node D, along with the shard information of shard 1 to which the second key-value information belongs, and the information of the proxy node 1 corresponding to shard 1. When inference node D performs the inference operation on this request, it generates a block corresponding to the second key-value information. Then, based on the shard information of shard 1 and the information of proxy node 1, inference node D sends the second key-value information and the position information of the block corresponding to the second key-value information within inference node D (denoted as the third position information) to proxy node 1. Proxy node 1 receives the second key-value information and the third position information sent by inference node D. Since the second key-value information already exists in shard 1, proxy node 1 only needs to add the third position information to shard 1, record the correspondence between the second key-value information and the third position information, and then send the second key-value information and the third position information to the central node so that the central node can update the global information.

[0083] In the fifth possible implementation, at least two proxy nodes in the inference system correspond to M inference nodes, with each proxy node corresponding to at least one inference node, and the inference nodes corresponding to different proxy nodes do not overlap. When the central node receives an inference request, it can generate some key-value information based on the token sequence of the inference request, and determine whether each key-value information already exists in the central node. If the key-value information already exists in the central node, the central node determines the proxy node that has reported the key-value information, and the shard managed by the proxy node that has reported the key-value information records the key-value information. After determining the target inference node for the inference request based on global information and the scheduling algorithm, the central node sends the inference request to the target inference node to instruct the target inference node to perform the inference operation for the inference request, and sends the shard information (such as shard ID) of the shard to which the key-value information already exists in the central node belongs, as well as the information of the proxy node corresponding to the shard, to the target inference node. When the target inference node performs inference operations on the inference request, it generates Blocks corresponding to the aforementioned key-value information. The target inference node sends the key-value information and location information of the generated Blocks to the corresponding proxy node, and forwards the sharding information sent by the central node and the information of the proxy node to the corresponding proxy node.

[0084] When the proxy node corresponding to the target inference node receives the key-value information and location information sent by the target inference node, the proxy node determines whether the target inference node has also sent the shard information of the shard to which the key-value information belongs and the information of the proxy node corresponding to the shard. For any key-value information received by the proxy node from the target inference node, if the target inference node has not sent the shard information of the shard to which the key-value information belongs and the information of the proxy node corresponding to the shard, it means that the target inference node is the first inference node in the inference system to generate the Block corresponding to the key-value information. Currently, no proxy node's shard records the key-value information, and no proxy node has reported the key-value information and its corresponding location to the central node. Therefore, the proxy node corresponding to the target inference node can record the key-value information and the location information of the Block corresponding to the key-value information in the target inference node in the corresponding shard. The key-value information then belongs to that shard, and the proxy node sends the key-value information and the location information of the Block corresponding to the key-value information in the target inference node to the central node.

[0085] For any key-value information received by the proxy node from the target inference node, if the target inference node sends the shard information of the shard to which the key-value information belongs and the information of the proxy node corresponding to the shard, it indicates that an inference node has previously generated the Block corresponding to the key-value information, and the key-value information is recorded in the shard corresponding to the aforementioned shard information. Further, if the shard corresponding to the aforementioned shard information is a shard managed by the inference node, the proxy node can record the location information of the Block corresponding to the key-value information in the target inference node within the managed shard, and send the key-value information and its location information to the central node. If the shard corresponding to the aforementioned shard information is a shard managed by another inference node, the proxy node, based on the aforementioned shard information and the information of the proxy node, sends the key-value information and its location information to the other inference node. The other proxy node records the key-value information and location information of the Block sent by the proxy node in the shard it manages, and sends the key-value information and location information of the Block to the central node.

[0086] When the central node receives key-value information and location information from a proxy node, it determines whether the key-value information already exists on the central node. If it does, it means the proxy node has previously reported the key-value information, and the central node has recorded the correspondence between the key-value information, the proxy node's information, and the shard information managed by the proxy node. Therefore, the central node only needs to add the location information sent by the proxy node this time to the global information. If the key-value information does not exist on the central node, it means the proxy node was the first to report the key-value information to the central node. Therefore, the central node adds the key-value information and corresponding location information sent by the proxy node to the global information and records the correspondence between the key-value information, the proxy node's information, and the shard information of the proxy node's shards.

[0087] For example, such as Figure 3As shown, assume that inference nodes A and B correspond to proxy node 1, and inference nodes C and D correspond to proxy node 2. When the central node receives a certain inference request, it calculates the first key value information, the second key value information, and the third key value information based on the token sequence of the inference request. The central node determines that the first key value information and the second key value information already exist in the global information, but the third key value information does not exist. The central node records the correspondence between the first key value information and the information of proxy node 1 and the shard information of shard 1, and also records the correspondence between the second key value information and the information of proxy node 2 and the shard information of shard 2. This indicates that the inference system has already generated the Block corresponding to the first key value information and the second key value information, but has not yet generated the Block corresponding to the third key value information. The first key value information is recorded in shard 1 of proxy node 1, and the second key value information is recorded in shard 2 of proxy node 2. Assuming the central node is the target inference node determined by the inference request and is inference node A, the central node sends the inference request, the shard information of shard 1 corresponding to the first key value information and the information of proxy node 1, the shard information of shard 2 corresponding to the second key value information and the information of proxy node 2 to inference node A.

[0088] When inference node A executes the inference operation on the inference request, it generates the Block corresponding to the first key-value information, the Block corresponding to the second key-value information, and the Block corresponding to the third key-value information. Then, inference node A sends the following information to the corresponding proxy node 1: the first key-value information, the first position information of the Block corresponding to the first key-value information in inference node A, the second key-value information, the second position information of the Block corresponding to the second key-value information in inference node A, the third key-value information, the third position information of the Block corresponding to the third key-value information in inference node A, the shard information of shard 1 to which the first key-value information belongs, the information of proxy node 1 that manages shard 1, the shard information of shard 2 to which the second key-value information belongs, and the information of proxy node 2 that manages shard 2.

[0089] When proxy node 1 receives the aforementioned information sent by inference node A, proxy node 1 determines, based on the shard information of shard 1 to which the first key-value information belongs and the information of proxy node 1, that the first key-value information already exists in shard 1 of proxy node 1. Therefore, proxy node 1 can record the aforementioned first location information sent by inference node A in shard 1 and record the correspondence between the first key-value information and the first location information. Based on the shard information of shard 2 to which the second key-value information sent by inference node A belongs and the information of proxy node 2, proxy node 1 determines that the second key-value information already exists in shard 2 of proxy node 2. Therefore, proxy node 1 forwards the second key-value information and the second location information sent by inference node A to proxy node 2. Since proxy node 1 has not received the shard information of the shard to which the third key-value information belongs from inference node A, proxy node 1 determines that no shard currently records the third key-value information. Proxy node 1 can add the third key-value information and the third location information to shard 1. At this time, the third key-value information belongs to shard 1, and the correspondence between the third key-value information and the third location information is recorded.

[0090] Agent Node 1 sends the aforementioned first key-value information and first location information, and third key-value information and third key-value information to the central node. The central node determines that the first key-value information already exists in the global information, so it only needs to add the first location information to the global information and record the correspondence between the first key-value information and the first location information. The central node determines that the third key-value information does not exist in the global information, so the central node adds both the third key-value information and the third location information to the global information, records the correspondence between the third key-value information and the third location information, and also records the correspondence between the third key-value information and the information of shard 1 and agent Node 1, to indicate that the third key-value information has been recorded in shard 1 of agent Node 1, and that the third key-value information belongs to shard 1 managed by agent Node 1.

[0091] When proxy node 2 receives the second key-value information and the second location information sent by proxy node 1, since the second key-value information already exists in shard 2, proxy node 2 only needs to add the second location information to shard 2 and record the correspondence between the second key-value information and the second location information. Then, proxy node 2 sends the second key-value information and the second location information to the central node. The central node determines that the second key-value information already exists in the global information, so it only needs to add the second key-value information to the global information and record the correspondence between the second key-value information and the second location information.

[0092] As discussed in the previous sections on possible implementations, each proxy node can manage its corresponding shards. Each shard records key-value information and the location information of the corresponding Block within one or more inference nodes. The central node's global information includes the key-value information and location information of Blocks in M ​​inference nodes. Therefore, shards can be considered to include a portion of the information in the central node's global information. In practical applications, for the same key-value information, the location information recorded in a shard can be richer (more detailed) than the location information recorded in the global information. That is, the global information records coarse-grained location information, while the shards record fine-grained location information.

[0093] Taking the first key-value information as an example, the first key-value information can be any or a specific key-value information. The shard to which the first key-value information belongs includes the first key-value information and its corresponding fine-grained location information. The fine-grained location information includes the information of the inference node where the block corresponding to the first key-value information is located (such as the ID of the inference node), the type of storage medium used by the corresponding inference node to store the block corresponding to the first key-value information, and the identifier of the storage area of ​​the block corresponding to the first key-value information in the corresponding storage medium. The identifier of the storage area can be the physical address of the storage area or information that can be mapped to the physical address of the storage area. The global information includes the first key-value information and its corresponding coarse-grained location information. The coarse-grained location information includes the information of the inference node where the block corresponding to the first key-value information is located, and may also include the type of storage medium used by the corresponding inference node to store the block corresponding to the first key-value information. In other words, although the proxy node corresponding to the aforementioned shard obtains and records the fine-grained location information corresponding to the first key-value information in the shard, this proxy node only sends a portion of the fine-grained location information to the central node. Therefore, the central node only records this portion of the information corresponding to the first key-value information in the global information; this portion is the coarse-grained location information corresponding to the first key-value information. This allows the central node to have a global understanding of the block distribution in the entire inference system, avoids the central node recording too much information, reduces the computational and storage pressure on the central node, and also reduces the communication overhead between the proxy node and the central node. Since the central node does not have the identifier of the storage area where the block corresponding to the first key-value information is located, the central node can instruct the inference node to obtain the identifier of the aforementioned storage area from the proxy node corresponding to the shard to which the first key-value information belongs, so that the proxy node can obtain the block corresponding to the first key-value information from that storage area based on the identifier of the aforementioned storage area.

[0094] For example, see Figure 4 , Figure 4This is an interactive flowchart provided in an embodiment of the present application, which includes the following steps 1 to 7.

[0095] Step 1: The central node obtains the inference request.

[0096] This application does not specifically limit the source of the inference request; for example, it may be sent from a client or other system to the inference system.

[0097] Step 2: The central node determines the target inference node, the key-value information of the target Block, and the target proxy node for the inference request based on the global information.

[0098] The aforementioned target inference node refers to the proxy node determined by the central node based on global information and a scheduling algorithm, which is suitable for executing the inference request. The central node will allocate the inference request to the target inference node for inference. This application does not specifically limit the aforementioned scheduling algorithm. The global information includes the key-value information and coarse-grained location information of the target block. The coarse-grained location information includes the ID of the inference node storing the target block, and may also include the type of storage medium used by the corresponding inference node to store the target block. The central node also records the correspondence between the key-value information, the shard to which the key-value information belongs, and the proxy node corresponding to the shard.

[0099] The aforementioned target block refers to the block that the target inference node needs to obtain from other inference nodes so that it can be reused during the inference operation performed by the target inference node on the inference request, thereby reducing the amount of computation. This application does not specifically limit the method by which the central node determines the target block. For example, the central node can calculate some key-value information based on the token sequence of the inference request, and use some or all of the blocks corresponding to these key-value information that do not exist in the target inference node as the target block. There may be one or more target blocks, and the proxy node corresponding to the shard to which the key-value information of the target block belongs is the target proxy node; therefore, there may be one or more target proxy nodes. For ease of description, the following steps will use proxy node 2 as the target proxy node and inference node D as the target inference node.

[0100] Step 3: The central node sends an inference request, the ID of the target proxy node, and the key-value information of the target block to the target inference node.

[0101] Step 4: The target inference node sends the key-value information of the target block to the target agent node.

[0102] Step 5: The target agent node sends the fine-grained location information of the target block to the target inference node.

[0103] Specifically, when the target inference node receives the inference request, the target proxy node's ID, and the target block's key-value information from the central node, it can send the target block's key-value information to the target proxy node based on the proxy node's ID. When the target proxy node receives this information, it can search within the managed shards to obtain the target block's fine-grained location information. This fine-grained location information includes the ID of the inference node storing the target block, the type of storage medium used by the inference node to store the target block, and the physical address of the target block's storage area within the corresponding storage medium.

[0104] Optionally, when the target proxy node determines through sharding that the number of proxy nodes holding the target block is greater than 1, the fine-grained location information of the target block sent by the target proxy node to the target inference node may include the IDs of all proxy nodes storing the target block. In this case, the target inference node decides for itself which proxy node to obtain the target block from. Alternatively, the fine-grained location information of the target block sent by the target proxy node to the target inference node may include the IDs of some proxy nodes storing the target block. When the number of the aforementioned partial proxy nodes is greater than 1, the target inference node still decides for itself which proxy node to obtain the target block from. When the number of the aforementioned partial proxy nodes is equal to 1, it is equivalent to the target proxy node selecting an inference node for the target inference node. In this case, the target inference node can only obtain the target block from the selected inference node.

[0105] It should be noted that steps 4 and 5 above involve the target inference node directly sending the key-value information of the target block to the target proxy node, and then the target proxy node feeding back the fine-grained location information of the target block to the target inference node. In practical applications, the target inference node can also indirectly obtain the fine-grained location information of the target block from the target proxy node through other proxy nodes. For example, the target proxy node can first send the key-value information of the target block and the ID of the target proxy node to its corresponding proxy node. Then, the corresponding proxy node sends the key-value information of the target block to the target proxy node based on the ID of the target proxy node. After the target proxy node feeds back the fine-grained location information of the target block to its corresponding proxy node, the corresponding proxy node then forwards the fine-grained location information of the target block to the target inference node.

[0106] Step 6: The target inference node obtains the target Block from other inference nodes based on the acquired fine-grained location information.

[0107] For any target block, if the fine-grained location information of the target block obtained by the target inference node from the target proxy node contains only the ID of one inference node and the identifier of the target block's storage area within that inference node, then the target inference node can only obtain the target block from the aforementioned storage area within that inference node based on this location information. If the fine-grained location information of the target block obtained by the target inference node from the target proxy node includes the IDs of K inference nodes and the identifier of the target block's storage area within each of these K inference nodes, where K is a positive integer greater than 1, then the target inference node can independently select one of these K inference nodes and then obtain the target block from the corresponding storage area of ​​the selected inference node.

[0108] Optionally, when there are multiple target blocks, the target inference node can obtain different target blocks in parallel from different inference nodes based on the fine-grained location information of each target block. This fully utilizes the communication bandwidth between inference nodes, enabling the target inference node to quickly obtain all target blocks and thus complete the inference operation for the inference request as quickly as possible. For example, Figure 4 As shown, assume that Block 3 in inference node B and Block 5 in inference node C are the target inference nodes ( Figure 4 If the inference node D requires two target blocks, then inference node D can obtain Block 3 from inference node B and Block 5 from inference node C in parallel, based on the fine-grained location information of Block 3 and Block 5 obtained from proxy node 2. After the data of Block 3 and Block 5 is transmitted to inference node D, inference node D can reuse the data of these two blocks in the inference operation, thereby reducing the amount of inference computation and improving the overall inference speed.

[0109] Step 7: The target inference node performs inference operations on the inference request based on the obtained target Block.

[0110] Regarding the execution order of steps 6 and 7 above, this application does not impose specific limitations; they can be executed sequentially or in parallel. For example, the target inference node may first obtain all target blocks from other inference nodes, and then start performing inference operations on the inference request based on all the obtained target blocks; or, the target inference node may first initiate the inference operation on the inference request, and then obtain target blocks from other inference nodes during the execution of the inference operation, and then perform the subsequent inference process of the inference request based on the obtained target blocks.

[0111] It should be understood that after the target inference node completes the inference operation for the inference request in step 7, the Blocks stored in the target inference node may change. For example, the target inference node may generate and store new Blocks during the inference operation, or it may store target Blocks obtained from other inference nodes. Therefore, after completing the inference operation, the target inference node can send the key-value information and fine-grained location information of the Blocks within the target inference node to the proxy node so that the proxy node can update the shards it manages. Regarding how the target inference node sends the key-value information and fine-grained location information of the Blocks within the target proxy node to the proxy node, please refer to the methods described above for each inference node to send the key-value information and location information of the Blocks in its own inference node to the proxy node in the various possible implementations; these will not be elaborated here. Regarding how the proxy node updates the shards it manages, please also refer to the relevant descriptions in the various possible implementations described above; these will not be elaborated here.

[0112] Optionally, for each inference node, the inference node can count the number of times it uses the corresponding Block for each key-value information within a certain period. The number of uses is the sum of a first usage count and a second usage count. The first usage count is the number of times the inference node uses the Block corresponding to the key-value information stored within the inference node, and the second usage count is the number of times the inference node uses the Block corresponding to the key-value information obtained from other inference nodes. It should be understood that whether using the Block stored within the inference node or using the Block obtained from other inference nodes for inference, it can reduce the amount of inference computation and improve the overall inference speed to some extent. This application does not specifically limit the length of the aforementioned period; for example, it could be 1 hour, 1 day, or 1 week. Then, the inference node sends the counted usage counts of the Block corresponding to each key-value information to the proxy node. Regarding which proxy node the inference node sends the usage count of the corresponding Block for each key-value information to, refer to the methods described above for determining which proxy node to send the generated Block's location information to. In other words, for any key-value information, if the inference node sends the location information of the Block corresponding to that key-value information within the inference node to a certain proxy node, then the inference node can also send the usage count of the Block corresponding to that key-value information to that proxy node.

[0113] Taking the first key-value information as an example, when the proxy node obtains the number of times the corresponding Block of the first key-value information is used by M inference nodes within a certain period, it can directly record it in the shard of the proxy node. Alternatively, the proxy node can sum or average the number of times the corresponding Block of the first key-value information is used by the M inference nodes to obtain the total number of times / average number of times the corresponding Block of the first key-value information is used, and then record the total number of times / average number of times the corresponding Block of the first key-value information is used in the shard. Subsequently, the proxy node can manage the corresponding Block of the first key-value information based on the total number of times / average number of times of use recorded in the shard. This management can include deleting, copying, or migrating the corresponding Block of the first key-value information.

[0114] For example, when the number of times each inference node uses the Block corresponding to the first key-value information is less than a first threshold, it is considered that the number of times all inference nodes use the Block corresponding to the first key-value information is low. In this case, the agent node can send a management instruction to some inference nodes that have the Block corresponding to the first key-value information. This management instruction is used to instruct the inference nodes to delete the Block corresponding to the first key-value information stored in the inference node, so as to save the storage space of the inference node. The aforementioned first threshold can be set according to actual usage needs, and this application does not specifically limit it.

[0115] For example, if the first inference node uses the Block corresponding to the first key-value information less than a second threshold, and the second inference node uses the Block corresponding to the first key-value information more than a third threshold (the second threshold being less than the third threshold), then it is considered that the first inference node rarely uses the Block corresponding to the first key-value information, while the second inference node uses the Block corresponding to the first key-value information more frequently. Furthermore, if the proxy node determines that the second inference node does not store the Block corresponding to the first key-value information, the proxy node can issue a management instruction to the first inference node. This management instruction instructs the first inference node to migrate the Block corresponding to the first key-value information to the second inference node, so that the second inference node can use the Block corresponding to the first key-value information and save storage space on the first inference node. The aforementioned second and third thresholds can be set according to usage needs, and this application does not specifically limit them. The first agent node executes the aforementioned management instructions to migrate the Block corresponding to the first key-value information from the first agent node to the second inference node. This allows the second inference node to directly use the Block corresponding to the first key-value information stored in the second inference node during subsequent inference processes, without needing to obtain the Block corresponding to the first key-value information from other inference nodes. This helps to improve the overall inference speed of the second inference node.

[0116] For example, suppose the first inference node uses the Block corresponding to the first key-value information more than the fourth threshold. This indicates that the first inference node uses the Block corresponding to the first key-value information relatively frequently. Furthermore, if the proxy node determines that only HBM stores the Block corresponding to the first key-value information in the first inference node, the first proxy node can issue a management instruction to the first inference node. This instruction instructs the first inference node to copy the Block corresponding to the first key-value information from HBM to other storage media within the first inference node, such as DRAM or SSD. It should be understood that HBM has limited storage space, and the Block stored in HBM may be released at any time. To ensure that the first proxy node can still quickly use the Block corresponding to the first key-value information later, the proxy node instructs the first proxy node to perform the above copy. After the first proxy node completes the above copy, even if the Block corresponding to the first key-value information in HBM is released, the first inference node can still use the Block corresponding to the first key-value information in other storage media within the first inference node.

[0117] Optionally, the proxy node also includes the positional relationship between inference nodes, and the proxy node can manage the Blocks in the inference nodes according to the above positional relationship.

[0118] For example, such as Figure 5 As shown, inference node A and inference node B are both deployed on computing device 1. Each inference node A and inference node B can use a portion of the computing and storage resources in computing device 1. The computing resources include CPUs and other processors (such as GPUs, DPUs, or TPUs), and the storage resources include memory (such as HBM) and other storage media (such as DRAM or SSDs). This application does not specifically limit the specific storage resources available to inference node A. The blocks owned by inference node A are stored in the storage resources available to inference node A, and the blocks owned by inference node B are stored in the storage resources available to inference node B. Inference node C and inference node D are both deployed on computing device 2. Each inference node C and inference node D can use a portion of the computing and storage resources in computing device 2. The blocks owned by inference node C are stored in the storage resources available to inference node C, and the blocks owned by inference node D are stored in the storage resources available to inference node D.

[0119] The proxy nodes have positional relationships with the inference nodes. This application does not specifically limit the source of these positional relationships; for example, they may originate from the central node or other locations, or they may be configured by the user. These positional relationships can indicate whether the inference nodes are on the same computing device, and may also indicate the distance between the computing devices where the inference nodes reside (such as physical distance or data transmission distance). Assuming that inference node A has a Block corresponding to the first key-value information in its HBM, and given the limited storage space of HBM, the data stored in HBM is easily released, inference node A requests proxy node 1 to copy the Block corresponding to the first key-value information to another storage medium of proxy node A. Proxy node 1 can then determine the location relationship between the inference nodes: if the location relationship between inference node A and another inference node meets a set condition, and the other inference node has a Block corresponding to the first key-value information in its other storage medium (excluding HBM), then inference node A can quickly obtain the Block corresponding to the first key-value information from the other inference node. In this case, proxy node 1 can instruct the proxy node not to copy the Block corresponding to the first key-value information to avoid creating too many copies of the Block, thus saving storage space for inference node A. The set condition can be that the two inference nodes are located in the same computing device, or that the distance between the computing devices of the two inference nodes is less than a distance threshold. The distance threshold can be set according to requirements, and this application does not specifically limit it.

[0120] Based on the reasoning system described above, the reasoning method provided in this application is introduced below.

[0121] Please see Figure 6 , Figure 6 This is a flowchart illustrating a reasoning method provided in an embodiment of this application, including steps S601 to S604.

[0122] S601, M inference nodes perform inference operations and each stores multiple data blocks generated during the inference operation, with each data block corresponding to a key-value information.

[0123] S602. The first proxy node among at least two proxy nodes obtains the first key value information of the first data block stored in N inference nodes among M inference nodes and the position information of the first data block in N inference nodes, where N is less than M.

[0124] The aforementioned M inference nodes and at least two proxy nodes constitute the inference system, which can be referred to in the previous section for details. The at least two proxy nodes are used to obtain and send the key-value information and location information of the Blocks stored in the inference nodes to the central node. Information regarding data blocks, key-value information, and location information can be found in the previous section and will not be repeated here.

[0125] The first data block mentioned above is a data block generated by N inference nodes out of M inference nodes when performing inference operations. Each of the N inference nodes stores the first data block, and the key-value information corresponding to the first data block is the first key-value information. The first proxy node mentioned above is one of at least two proxy nodes in the inference system. The first proxy node obtains the position information of the first data block among the N inference nodes, thereby obtaining N position information corresponding to the N inference nodes. Regarding how the proxy node obtains the key-value information and position information of the Block in the inference node, please refer to the relevant descriptions in the five possible implementation methods introduced above; they will not be repeated here.

[0126] Optionally, the first proxy node manages the first shard. The first shard is configured with key-value characteristics of the key-value information managed by the first proxy node. For details on the key-value characteristics, please refer to the previous description; they will not be repeated here. After generating the first data block, the first inference node among the M inference nodes determines the first key-value characteristics corresponding to the first key-value information based on the first key-value information of the first data block. Then, based on the first key-value characteristics, it determines the first shard managed by the first proxy node corresponding to the first key-value information and sends the first key-value information and the location information of the first data block within the first inference node to the first proxy node.

[0127] In the above-mentioned optional scheme, the first shard is configured with key-value characteristics of key-value information managed by the first proxy node. For any key-value information, the following rule can be used to determine whether the key-value information is managed by the first proxy node: if the key-value characteristic corresponding to the key-value information is the key-value characteristic set by the first shard, then it can be determined that the key-value information is managed by the first proxy node, and the key-value information belongs to the first shard corresponding to the first proxy node. Based on the above rule, the first inference node determines that the first key-value information belongs to the first shard according to the first key-value characteristic of the first key-value information of the first data block it generates. Therefore, the first key-value information and the position information of the first data block in the first inference node can be directly sent to the first proxy node without the need for other proxy nodes to forward it to the first proxy node, nor is it necessary for the central node to send the shard information of the first shard and the information of the first proxy node to the first inference node. The entire implementation process is relatively simple. Similarly, other inference nodes besides the first inference node can also determine, based on the above rules, whether the key-value information of their generated Block belongs to the first fragment of the first proxy node. If so, they can send the key-value information and location information of the generated Block to the first proxy node. For details, please refer to the relevant description in the first possible implementation method introduced above, which will not be repeated here.

[0128] S603, the first agent node sends the first key value information and the N location information corresponding to the N inference nodes to the central node.

[0129] Optionally, in addition to sending the acquired first key-value information and the aforementioned N location information to the central node, the first proxy node can also add the first key-value information and the aforementioned N location information to its corresponding first shard. The first shard includes a portion of the central node's global information. Besides the first proxy node, other proxy nodes in the inference system can also manage their corresponding shards. Shards are used to record the partial or global key-value information and corresponding location information acquired by the corresponding proxy node. For details on the shards corresponding to proxy nodes, please refer to the previous description; they will not be repeated here. It should be understood that by managing the first shard and adding the acquired key-value information and corresponding location information to the first shard, the first proxy node can support the central node or inference node in obtaining the required information from the first shard of the first proxy node.

[0130] For example, when the central node loses global information due to a failure, since the first shard contains a portion of the global information, the central node can retrieve the information from the first proxy node. Similarly, the central node can retrieve the information from other proxy nodes for their respective shards, enabling the central node to quickly recover the global information. This helps improve the central node's fault recovery efficiency and reduce its cost. For details on how proxy nodes update shards and how the central node updates global information, please refer to the descriptions of the possible implementation methods introduced earlier; they will not be repeated here.

[0131] For example, when an inference node needs to retrieve a first data block from other inference nodes, it can send a data retrieval request carrying the first key-value information of the first data block to a first proxy node. Then, the first proxy node retrieves the location information of the first data block in at least one of the N inference nodes from the first shard according to the data retrieval request, and sends this location information to the first inference node. This allows the inference node to retrieve the first data block from the first inference node based on its location information. After obtaining the first data block, the inference node can use it as needed during inference operations to reduce computational load and improve overall inference speed.

[0132] Optionally, the first inference node among the M inference nodes can send a data retrieval request to the first proxy node. The data retrieval request includes first key-value information. Then, the first proxy node sends the location information of the first data block in at least one of the N inference nodes to the first inference node according to the data retrieval request. The first inference node can then retrieve the first data block from the at least one inference node based on the location information of the first data block in the at least one inference node sent by the first proxy node.

[0133] The aforementioned at least one inference node may include one or more inference nodes. When an inference node receives the location information of the first data block within the aforementioned at least one inference node, it can choose to retrieve the first data block from one of the aforementioned at least one inference node. Compared to the inference node sending a data retrieval request to the central node to obtain the location information of the first data block recorded in the global information, the inference node in the above scheme only needs to interact with the proxy node and does not need to interact with the central node, which can reduce the communication pressure on the central node.

[0134] Optionally, the first inference node among the N inference nodes corresponds to the first proxy node, and the second inference node among the N inference nodes corresponds to the second proxy node among at least two proxy nodes; the first proxy node obtains the first key-value information of the first data block stored in the N inference nodes and the position information of the first data block in the N inference nodes, specifically including: receiving the first key-value information of the first data block and the first position information of the first data block in the first inference node sent by the first proxy node; receiving the first key-value information and the second position information of the first data block in the second inference node sent by the second proxy node; and recording the correspondence between the first key-value information and the first position information and the second position information.

[0135] In the above-mentioned optional schemes, inference nodes and proxy nodes have a corresponding relationship. An inference node sends the key-value information and location information of the data block stored in its proxy node to the corresponding proxy node. Specifically, the first inference node corresponds to the first proxy node, and the second inference node corresponds to the second proxy node. Thus, the first inference node sends the first key-value information and the first location information of the first data block in its first inference node to the first proxy node, and the second inference node sends the first key-value information and the second location information of the first data block in its second inference node to the second proxy node. To aggregate the corresponding location information of the first data block at the first proxy node, the second proxy node can send the first key-value information and the second location information to the first proxy node. That is, of the N location information obtained by the first proxy node, some are obtained from the corresponding inference node, and some are obtained from other proxy nodes. For details on how a proxy node obtains key-value information and corresponding location information from inference nodes and other proxy nodes, please refer to the relevant descriptions in the first, third, or fifth possible implementation methods introduced above; these will not be repeated here.

[0136] Once the first agent node obtains these N location information points, it can send them to the central node in a unified manner, instead of having multiple agent nodes or N communication nodes send these N location information points to the central node. This helps reduce the communication and processing pressure on the central node.

[0137] Optionally, the first proxy node can send the fragmentation information of the first fragment to the central node, so that when the central node instructs the first inference node among the M inference nodes to perform the first inference operation, it can carry the fragmentation information of the first fragment to which the first data block generated by the first inference operation belongs, as well as the information of the first proxy node corresponding to the first fragment. After the first inference node generates the first data block, it can send the first key-value information of the first data block and the position information of the first data block in the first inference node to the first proxy node according to the fragmentation information of the first fragment and the information of the first proxy node.

[0138] In the above-mentioned optional scheme, the first proxy node not only sends the acquired first key-value information and N location information to the central node, but also sends the fragment information of the first fragment corresponding to the first proxy node to the central node, so as to indicate to the central node that the first key-value information is recorded in the first fragment of the first proxy node, indicating that the first key-value information belongs to the first fragment. When the central node instructs the first inference node to perform the first inference operation (the first data block generated by the first inference operation), it can send the fragment information of the first fragment to which the first data block belongs (i.e., the first fragment records the first key-value information of the first data block) and the information of the first proxy node to the first inference node, so as to instruct the first inference node to send the first key-value information and location information of the first data block generated during the subsequent inference operation to the first proxy node, thereby making the corresponding location information of the first data block converge in the first fragment of the first proxy node.

[0139] Based on the fragmentation information of the first fragment sent by the central node and the information of the first proxy node, the first inference node can directly send the first key-value information and location information of the first data block generated by the first inference node to the first proxy node, without needing other proxy nodes to forward it to the first proxy node. This reduces the number of forwardings and communication overhead, allowing the first key-value information and location information of the first data block generated by the first inference node to be transmitted to the first proxy node more quickly. In turn, the first proxy node can transmit the first key-value information and location information of the first data block generated by the first inference node to the central node more quickly, resulting in higher information collection efficiency.

[0140] S604. The central node adds the first key value information and N location information to the global information. The global information includes the key value information and location information of the data blocks stored in the M inference nodes.

[0141] Optionally, after receiving the first key-value information sent by the first proxy node, the central node determines whether the first key-value information already exists in the central node. If it exists, it identifies the second proxy node among at least two proxy nodes that has reported the first key-value information, and notifies the first proxy node to send the first key-value information and N location information to the second proxy node. Then, the first proxy node sends the acquired first key-value information and N location information to the second proxy node according to the above notification, and deletes the first key-value information and N location information recorded by the first proxy node.

[0142] In the above-mentioned optional scheme, when the central node receives the first key-value information and N location information sent by the first proxy node, it does not immediately add the first key-value information and N location information to the global information. Instead, it first checks whether the first key-value information already exists in the central node. If it does, it means that before the first proxy node, other proxy nodes have already reported the first key-value information to the central node, and both the first proxy node and other proxy nodes have recorded the first key-value information and the corresponding location information. This results in duplicate information recorded by the first proxy node and other proxy nodes, which leads to a waste of storage resources. To save storage resources, the central node determines that the second proxy node is another proxy node that has reported the first key-value information before the first proxy node. Then, the central node notifies the first proxy node to send the first key-value information and N location information it has obtained to the second proxy node, and instructs the first proxy node to delete the recorded first key-value information and N location information to achieve information deduplication. For deduplication, please refer to the relevant description in the second possible implementation method introduced above, which will not be repeated here.

[0143] When the second proxy node receives the first key-value information and N location information sent by the first proxy node, it can record these N location information and the correspondence between the first key-value information and these N location information. At this point, the corresponding location information of the first data block is all gathered at the second proxy node. Subsequently, the second proxy node sends the first key-value information and these N location information to the central node, and only then does the central node add these N location information to the global information.

[0144] In summary, the inference system provided in this application includes M inference nodes, at least two proxy nodes, and a central node. The first proxy node among the at least two proxy nodes obtains the first key value information of the first data block stored in N inference nodes among the M inference nodes and the location information of the first data block in the N inference nodes, and sends the first key value information and the N location information corresponding to the N inference nodes to the central node, so that the central node can collect the location information of the first data block in the N inference nodes, so as to perform global scheduling of inference requests based on the collected information. Compared to the approach where each of the N inference nodes sends the location information of the first data block to the central node, the proposed solution does not require the central node to communicate with each of the N inference nodes separately to obtain N location information. Instead, the central node collects the location information of the first data block across the N inference nodes from the first proxy node all at once. Essentially, the first proxy node aggregates the location information of the first data block across the N inference nodes (obtaining N location information) before sending it to the central node. The central node can then record the first key-value information and the N location information of the first data block together, avoiding repeated refreshing of the information recorded in the central node, thereby reducing the communication and computational pressure on the central node.

[0145] Because at least two proxy nodes in the inference system replace the M inference nodes in sending the key-value information and location information of the Blocks stored in those M inference nodes to the central node, the proxy nodes essentially act as communication proxies for the inference nodes. This avoids excessive communication pressure on the central node caused by all inference nodes sending their respective Block location and key-value information to the central node. Furthermore, when an inference node needs to retrieve a Block stored in another inference node, it can query the Block's location information from the proxy nodes instead of querying the central node, further reducing the communication pressure on the central node. When the central node loses global information due to a failure, it can quickly recover the global information by retrieving information from the shards of each proxy node.

[0146] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc. The computer-readable storage medium includes instructions that, when executed on an inference system, cause the inference system to perform... Figure 6 The operational steps in the reasoning method.

[0147] This application also provides a computer program product, which may be a software or program product containing instructions capable of running on a computing device or stored on any usable medium. When the instructions in the aforementioned computer program product are executed on an inference system, the inference system can be caused to perform... Figure 6 The operational steps in the reasoning method.

[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.

Claims

1. A reasoning system, characterized in that, include: There are M inference nodes, which are used to perform inference operations. Each node stores multiple data blocks generated during the inference operation, and each data block corresponds to a key-value information. At least two proxy nodes, wherein the first proxy node among the at least two proxy nodes is used to obtain the first key value information of the first data block stored in N inference nodes among the M inference nodes and the position information of the first data block in the N inference nodes, wherein N is less than M; The first proxy node is also used to send the first key value information and the N location information corresponding to the N inference nodes to the central node.

2. The system according to claim 1, characterized in that, The central node is used to add the first key value information and the N location information to the global information, and the global information includes the key value information and location information of the data blocks stored in the M inference nodes; The first proxy node is also used to add the first key value information and the N location information to the first shard corresponding to the first proxy node, and the first shard includes some information in the global information.

3. The system according to claim 1 or 2, characterized in that, The central node is used to determine whether the first key value information already exists in the central node after receiving the first key value information. If it exists, it determines the second proxy node among the at least two proxy nodes that has reported the first key value information, and notifies the first proxy node to send the first key value information and the N location information to the second proxy node. The first proxy node is further configured to send the acquired first key value information and the N location information to the second proxy node according to the notification, and delete the first key value information and the N location information recorded by the first proxy node.

4. The system according to claim 1 or 2, characterized in that, The first inference node among the N inference nodes corresponds to the first proxy node, and the second inference node among the N inference nodes corresponds to the second proxy node among the at least two proxy nodes; When the first proxy node is used to obtain the first key-value information of the first data block stored in N inference nodes out of the M inference nodes, it is specifically used for: Receive the first key value information of the first data block and the first position information of the first data block in the first inference node sent by the first agent node; Receive the first key value information and the second position information of the first data block in the second inference node sent by the second agent node; Record the correspondence between the first key value information and the first and second position information.

5. The system according to claim 2, characterized in that, The first proxy node is also used to send the fragmentation information of the first fragment to the central node; The central node is also used to carry the fragmentation information of the first fragment to which the first data block generated by the execution of the first inference operation belongs and the information of the first proxy node corresponding to the first fragment when instructing the first inference node among the M inference nodes to perform the first inference operation. The first inference node is further configured to, after generating the first data block, send the first key information of the first data block and the position information of the first data block in the first inference node to the first agent node according to the fragmentation information of the first fragment and the information of the first agent node.

6. The system according to claim 1 or 2, characterized in that, The first proxy node manages the first shard, and the first shard is set with key-value features of key-value information managed by the first proxy node; The first inference node among the M inference nodes is used to, after generating the first data block, determine the first key value feature corresponding to the first key value information based on the first key value information of the first data block, determine the first fragment managed by the first agent node based on the first key value feature, and send the first key value information and the location information of the first data block in the first inference node to the first agent node.

7. The system according to claim 2, characterized in that, The first inference node among the M inference nodes is used to send a data acquisition request to the first agent node, and the data acquisition request includes the first key-value information; The first proxy node is also configured to send the location information of the first data block in at least one of the N inference nodes to the first inference node according to the data acquisition request; The first inference node is further configured to obtain the first data block from the at least one inference node based on the location information of the first data block in the at least one inference node.

8. A reasoning method, characterized in that, The method is used in an inference system, the inference system comprising M inference nodes and at least two agent nodes, the method comprising: The M inference nodes perform inference operations and each stores multiple data blocks generated during the inference operation, with each data block corresponding to a key-value information; The first proxy node among the at least two proxy nodes obtains the first key value information of the first data block stored in N inference nodes among the M inference nodes and the position information of the first data block in the N inference nodes, where N is less than M; The first agent node sends the first key value information and the N location information corresponding to the N inference nodes to the central node.

9. The method according to claim 8, characterized in that, The method further includes: The central node adds the first key-value information and the N location information to the global information, and the global information includes the key-value information and location information of the data blocks stored in the M inference nodes; The first proxy node adds the first key value information and the N location information to the first shard corresponding to the first proxy node, and the first shard includes some information from the global information.

10. The method according to claim 8 or 9, characterized in that, The method further includes: After receiving the first key value information, the central node determines whether the first key value information already exists in the central node. If it does, it determines the second proxy node among the at least two proxy nodes that has reported the first key value information, and notifies the first proxy node to send the first key value information and the N location information to the second proxy node. The first proxy node sends the first key value information and the N location information obtained to the second proxy node according to the notification, and deletes the first key value information and the N location information recorded by the first proxy node.

11. The method according to claim 8 or 9, characterized in that, The first inference node among the N inference nodes corresponds to the first proxy node, and the second inference node among the N inference nodes corresponds to the second proxy node among the at least two proxy nodes. The first proxy node obtains the first key-value information of the first data block stored in N inference nodes out of the M inference nodes, and the position information of the first data block in the N inference nodes, including: Receive the first key value information of the first data block and the first position information of the first data block in the first inference node sent by the first agent node; Receive the first key value information and the second position information of the first data block in the second inference node sent by the second agent node; Record the correspondence between the first key value information and the first and second position information.

12. The method according to claim 9, characterized in that, The method further includes: The first proxy node sends the fragmentation information of the first fragment to the central node; When the central node instructs the first inference node among the M inference nodes to perform the first inference operation, it carries the fragmentation information of the first fragment to which the first data block generated by the first inference operation belongs, and the information of the first proxy node corresponding to the first fragment. After generating the first data block, the first inference node sends the first key information of the first data block and the location information of the first data block in the first inference node to the first agent node according to the fragmentation information of the first fragment and the information of the first agent node.

13. The method according to claim 8 or 9, characterized in that, The first proxy node manages the first shard, and the first shard is configured with key-value characteristics of key-value information managed by the first proxy node. The method further includes: After generating the first data block, the first inference node among the M inference nodes determines the first key value feature corresponding to the first key value information based on the first key value information of the first data block, determines the first fragment managed by the first agent node based on the first key value feature, and sends the first key value information and the location information of the first data block in the first inference node to the first agent node.

14. The method according to claim 9, characterized in that, The method further includes: The first inference node among the M inference nodes sends a data acquisition request to the first agent node, and the data acquisition request includes the first key-value information. The first agent node sends the location information of the first data block in at least one of the N inference nodes to the first inference node according to the data acquisition request; The first inference node obtains the first data block from the at least one inference node based on the location information of the first data block in the at least one inference node.

15. A computer program product, characterized in that, Includes instructions that, when executed by the inference system, cause the inference system to perform the method as described in any one of claims 8-14.