Task reasoning method and device and electronic equipment

By building a shared memory pool in the distributed inference architecture, the problem of cross-node transmission latency is solved, enabling efficient data sharing and inference task execution, and improving the efficiency and accuracy of task processing.

CN121684007APending Publication Date: 2026-03-17XFUSION DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511591427.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

In a distributed inference architecture, the data migration process across nodes results in slow response times and low inference task efficiency.

Method used

By building a shared memory pool between the first and second nodes, retrieval results can be stored and retrieved directly in the shared memory pool, avoiding network transmission delays across nodes and achieving efficient data sharing between nodes.

Benefits of technology

It significantly improves the overall execution efficiency and accuracy of reasoning tasks, reduces the waste of system resources, and enhances the real-time performance of task processing and overall reasoning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121684007A_ABST
    Figure CN121684007A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a task reasoning method and device and electronic equipment. The method is applied to a distributed reasoning architecture, the distributed reasoning architecture comprises a first node, a second node and a shared memory pool, and the method comprises the steps that the first node obtains at least one first to-be-reasoned task; the first node retrieves the first to-be-reasoned task to obtain at least one first retrieval result related to the first to-be-reasoned task; the first node stores the at least one first retrieval result in a shared memory pool; the second node reads at least one first retrieval result from the shared memory pool; and the second node takes the at least one first to-be-reasoned task and the at least one first retrieval result corresponding to the first to-be-reasoned task as input of a model, and executes the first to-be-reasoned task by utilizing the model. By adopting the mode, the delay problem caused by cross-node transmission can be avoided, so that the overall efficiency of the reasoning task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a task reasoning method, apparatus, and computing device. Background Technology

[0002] Model reasoning tasks are widely used in intelligent scenarios such as online customer service and smart assistant business applications, where the accuracy and timeliness of reasoning results are extremely important. However, once a model is trained, its knowledge becomes fixed and it cannot acquire new information after the training data expires. When faced with dynamically changing real-world needs, it is prone to generating outdated or fictitious content. To address this, retrieval-augmented generation (RAG) technology is introduced. This technology can retrieve external knowledge bases in real time during the reasoning process, providing the model with the latest and most reliable information support, thereby effectively improving the accuracy of the task's reasoning results.

[0003] However, in actual deployments, the inference process of RAG models typically runs on heterogeneous computing architectures to handle massive data and computational loads. The raw data of the knowledge database is persistently stored on the storage devices of the retrieval nodes, while the model's inference computation is performed by the inference nodes. In collaborative inference tasks, knowledge retrieval must first be completed on the retrieval node side, and then the retrieval results are transmitted to the inference node side via a bus for further processing. This cross-node data migration process leads to slow response times, resulting in low inference task efficiency. Summary of the Invention

[0004] This application provides a task reasoning method, apparatus, and computing device. By utilizing shared memory to establish a communication connection between a first node and a second node, the latency problem caused by cross-node transmission is effectively avoided, thereby improving the overall execution efficiency of the reasoning task.

[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions: In a first aspect, embodiments of this application provide a task reasoning method applied to a distributed reasoning architecture, the distributed reasoning architecture including a first node, a second node, and a shared memory pool. The method includes: the first node acquiring a first task to be reasoned; the first node retrieving the first task to be reasoned and acquiring at least one first retrieval result related to the first task to be reasoned; the first node storing the at least one first retrieval result in the shared memory pool; wherein, the shared memory pool is a memory area accessible to both the first node and the second node; the second node reading at least one first retrieval result from the shared memory pool; and the second node using the first task to be reasoned and the at least one first retrieval result corresponding to the first task to be reasoned as input to a first model, and executing the first task to be reasoned using the first model.

[0006] Based on this scheme, the first retrieval data obtained by the first node is directly written into the shared memory pool, and the subsequent second node reads directly from the shared memory pool. In this way, data sharing between nodes is realized through the shared memory pool, avoiding the network transmission overhead and latency in traditional cross-node communication, effectively eliminating the data handling burden of the first node, reducing the waste of system resources, and thus significantly improving the overall execution efficiency of the inference task.

[0007] In one possible implementation, the first node searches for a first reasoning task and obtains at least one first search result related to the first reasoning task, including: the first node calculates the initial similarity between the first reasoning task and at least one knowledge fragment in a preset vector knowledge base; the first node determines at least one target similarity from the at least one initial similarity based on the preset similarity; and the first node determines at least one first search result related to the first reasoning task based on the at least one target similarity.

[0008] Based on this scheme, a similarity retrieval mechanism is used to obtain knowledge fragments related to the reasoning task from the vector database, i.e., the first retrieval result. Since the knowledge fragments contain the latest real-time factual information, they can provide more accurate, timely and reliable knowledge support for the subsequent generation of responses by the first big language model, thereby effectively improving the accuracy of the first model's reasoning and the overall processing quality of the task.

[0009] In one possible implementation, before the second node reads at least one first retrieval result from the shared memory pool, the method further includes: the first node recording the storage state of the first retrieval result and generating a first storage address corresponding to the first retrieval result; and the first node sending the first storage address.

[0010] Based on this solution, by recording the storage status of the retrieval results and introducing a notification mechanism, efficient collaboration between the first and second nodes is achieved. When the second node receives a notification message containing the storage address, it can immediately and accurately read the required data from the shared memory pool. This eliminates the need for the second node to poll or repeatedly check the shared memory pool, effectively reducing the consumption of system resources and waiting time. It ensures that the second node obtains the required data in a timely and accurate manner, improving the real-time performance of task processing and the overall inference efficiency.

[0011] In one possible implementation, the second node reads at least one first retrieval result from the shared memory pool, including: the second node reads at least one first retrieval result from the shared memory pool based on a first storage address.

[0012] Based on this scheme, the second node obtains the target storage address of the first search result in the shared memory pool and directly reads the corresponding search result based on that address. This achieves precise data access location, avoids invalid scanning or polling operations, further improves data reading efficiency and system response speed, enhances the real-time performance of inter-node collaboration, and improves the real-time performance of task processing and overall inference efficiency.

[0013] In one possible implementation, the second node takes at least one first inference task and at least one first retrieval result corresponding to the first inference task as input to the first model, and uses the first model to execute the first inference task, including: the second node processes the first inference task and at least one first retrieval result corresponding to the first inference task to obtain an input prompt sequence corresponding to the first inference task; the second node processes the input prompt sequence using the first model to obtain at least one key-value cache; wherein, the at least one key-value cache is a set of key-value pairs generated at each attention layer for each input token in the input prompt sequence when the first model encodes the input prompt sequence during the task inference process; the second node stores the at least one key-value cache in a shared memory pool, and executes the first inference task based on the read key-value cache during the regression decoding stage.

[0014] Based on this scheme, the second node concatenates the first inference task with its corresponding retrieval results into an input prompt sequence. The sequence is then encoded using the first neural network model to generate a corresponding key-value cache, which is stored in the shared memory pool. In the subsequent decoding stage, the inference task is executed directly based on this key-value cache. This direct reading of the key-value cache from shared memory effectively improves the efficiency and resource utilization of the first model's inference.

[0015] In one possible implementation, after the first inference task is completed, the method further includes: releasing the shared storage space of the first inference task in the shared memory pool; wherein the shared storage space includes the storage space occupied by the first retrieval result and key-value cache corresponding to the first inference task.

[0016] Based on this scheme, after the first inference task is completed, the storage space it occupies in the shared memory pool is released promptly, namely the first search result and key-value cache corresponding to the first inference task. This effectively reclaims resources occupied by temporary data, avoids memory redundancy, and improves the utilization rate of shared memory.

[0017] In one possible implementation, after the first inference task is completed, the method further includes: a first node acquiring a second inference task; the second inference task being generated based on subsequent user input to the first inference task; the first node retrieving the second inference task to obtain at least one second retrieval result related to the second inference task; the first node storing at least one second retrieval result in a shared memory pool; the second node reading at least one second retrieval result from the shared memory pool; and the second node using the second inference task and at least one second retrieval result corresponding to the second inference task as input to a first model, and executing the second inference task using the first model.

[0018] Based on this scheme, after the first inference task is completed, the first node obtains the second inference task generated based on the user's feedback on the previous task, retrieves relevant second retrieval results from it, and stores the results in a shared memory pool. The second node reads the second retrieval results from the pool, combines them with the second inference task, and inputs them into the first model to execute subsequent inference. In this way, if the first inference task is unsatisfactory, multiple rounds of continuous inference can be performed, and relevant data can be transferred through the shared memory pool, improving the inference efficiency in complex interaction scenarios.

[0019] In one possible implementation, the first node stores at least one first retrieval result in a shared memory pool, including: the first node stores at least one first retrieval result in the shared memory pool via an interconnection protocol; wherein the interconnection protocol includes the CXL.mem protocol, the OpenCAPI protocol, or the Gen-Z protocol.

[0020] Based on this solution, by adopting a high-performance interconnection protocol, efficient and low-latency data transmission between nodes and the shared memory pool is achieved, thereby improving the overall data access speed and communication efficiency of the system.

[0021] Secondly, embodiments of this application provide a task reasoning apparatus, comprising: an acquisition module configured to acquire a first task to be reasoned; the acquisition module is further configured to search the first task to be reasoned and acquire at least one first search result related to the first task to be reasoned; a storage module configured to store at least one first search result in a shared memory pool; wherein the shared memory pool is a memory area that supports access by a first node and a second node; a reading module configured to read at least one first search result from the shared memory pool; and a reasoning module configured to use the first task to be reasoned and at least one first search result corresponding to the first task to be reasoned as input to a first model, and execute the first task to be reasoned using the first model.

[0022] In one possible implementation, the acquisition module is specifically configured to: calculate the initial similarity between the first reasoning task and at least one knowledge fragment in a preset vector knowledge base; determine at least one target similarity from the at least one initial similarity based on the preset similarity; and determine at least one first retrieval result related to the first reasoning task based on the at least one target similarity.

[0023] In one possible implementation, a generation module is also included, which is configured to: record the storage status of the first search result, generate the first storage address corresponding to the first search result, and send the first storage address.

[0024] In one possible implementation, the reading module is specifically configured to read at least one first retrieval result from the shared memory pool based on a first storage address.

[0025] In one possible implementation, the inference module is specifically configured to: process a first task to be inferred and at least one first retrieval result corresponding to the first task to be inferred to obtain an input prompt sequence corresponding to the first task to be inferred; process the input prompt sequence using a first model to obtain at least one key-value cache; wherein, the at least one key-value cache is a set of key-value pairs generated at each attention layer for each input token in the input prompt sequence when the first model encodes the input prompt sequence during the task inference process; store the at least one key-value cache in a shared memory pool, and execute the first task to be inferred based on the read key-value cache during the regression decoding stage.

[0026] In one possible implementation, a deletion module is also included, which is configured to: release the shared storage space of the first task to be inferred in the shared memory pool; wherein the shared storage space includes the storage space occupied by the first retrieval result and key-value cache corresponding to the first task to be inferred.

[0027] In one possible implementation, the acquisition module is specifically configured to: acquire a second inference task; the second inference task is an inference task generated based on subsequent user input to the first inference task; the acquisition module is further specifically configured to: search the second inference task and acquire at least one second search result related to the second inference task; the storage module is specifically configured to: store at least one second search result in a shared memory pool; the reading module is specifically configured to: read at least one second search result from the shared memory pool; and the inference module is specifically configured to: use the second inference task and at least one second search result corresponding to the second inference task as input to a first model, and execute the second inference task using the first model.

[0028] In one possible implementation, the storage module is specifically configured to store at least one first retrieval result to a shared memory pool via an interconnect protocol; wherein the interconnect protocol includes the CXL.mem protocol, the OpenCAPI protocol, or the Gen-Z protocol.

[0029] Thirdly, embodiments of this application also provide a computing device, including: a processor and a memory; the processor and the memory are coupled; the memory is used to store program instructions; the processor is used to execute the program instructions to perform the method as described in any of the first aspects above.

[0030] Fourthly, embodiments of this application provide a chip for performing the methods described in any of the first aspects above.

[0031] Fifthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a computer, implement the method as described in any of the first aspects.

[0032] In a sixth aspect, embodiments of this application provide a program product including a computer program that, when executed by a processor, implements the method as described in any of the first aspects. Attached Figure Description

[0033] Figure 1 This is a scenario framework diagram for task reasoning provided in an embodiment of this application; Figure 2 This is a first flowchart illustrating a task reasoning method provided in an embodiment of this application; Figure 3 This is a first structural schematic diagram of a task reasoning method provided in an embodiment of this application; Figure 4 This is a second flowchart illustrating a task reasoning method provided in an embodiment of this application; Figure 5 This is a schematic diagram of an interface for obtaining a first inference task provided in an embodiment of this application; Figure 6 This is a schematic diagram of a process for obtaining a first search result provided in an embodiment of this application; Figure 7 This is a schematic diagram of a training method for a second model provided in an embodiment of this application; Figure 8 This is a third flowchart illustrating a task reasoning method provided in an embodiment of this application; Figure 9 This is a fourth flowchart illustrating a task reasoning method provided in an embodiment of this application; Figure 10This is a schematic diagram of a task reasoning device provided in an embodiment of this application; Figure 11 This is a schematic diagram of a computing device provided in an embodiment of this application. Detailed Implementation

[0034] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. To facilitate a clear description of the technical solutions of the embodiments of this application, the use of terms such as "first," "second," etc., in the embodiments of this application is for illustrative purposes and to distinguish the objects being described. There is no particular order between them, nor does it indicate a specific limitation on the number of devices in the embodiments of this application, and they do not constitute any limitation on the embodiments of this application.

[0035] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of this application.

[0036] It should be noted that many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited to the specific embodiments disclosed below.

[0037] The following explanations of the technical terms mentioned in the embodiments of this application are provided to facilitate understanding by those skilled in the art.

[0038] Large language models (LLMs) are neural network models based on deep learning techniques with a large number of parameters (usually hundreds of millions to hundreds of billions or even more). By pre-training on massive amounts of text data, they have the ability to understand and generate natural language and can be applied to various natural language processing and reasoning tasks such as text generation, question answering, translation, and summarization.

[0039] Retrieval-augmented generation (RAG) is a technique that combines information retrieval and language generation to improve the accuracy and reliability of large language models (LLMs) in tasks such as question answering, dialogue, and content generation. For example, it retrieves document fragments or factual information relevant to the input question or context from an external knowledge base, and then feeds these retrieval results, along with the original input, as prompts into the language model to guide it in generating more accurate and evidence-based responses.

[0040] Key-value cache (KV cache) refers to a model framework where each input token generates a corresponding key and value vector during the attention mechanism. To avoid repeatedly calculating the attention results of historical tokens when generating each new token, the model caches the key and value of already processed tokens, forming a key-value cache.

[0041] Compute Express Link (CXL) is a high-performance, low-latency open interconnect technology primarily used to improve communication efficiency between processors and components such as accelerators and memory expansion devices. The processor, as the main control core, is responsible for task scheduling and data coordination; it typically refers to a central processing unit (CPU) or a graphics processing unit (GPU). Accelerators are intelligent devices with dedicated computing capabilities connected via the CXL interface, such as neural processing units (NPUs), data processing units (DPUs), and field-programmable gate arrays (FPGAs).

[0042] CXL shared memory is a memory sharing mechanism based on the Compute Fast Interconnect protocol that supports efficient, low-latency, and cache-consistent sharing across nodes. In this embodiment, a shared memory pool is used as an example for illustration.

[0043] The embodiments of this application will now be described with reference to the accompanying drawings.

[0044] In scenarios where models are used for task inference, the large number of parameters and complex network structures of the models themselves often mean that the resources of a single computing node are insufficient to meet the requirements for their complete deployment and efficient execution. Therefore, large models must be distributed across multiple computing nodes to run collaboratively. This technical architecture, which completes a single inference task through multi-node collaboration, is called distributed inference.

[0045] Figure 1 This is a scenario framework diagram of distributed task reasoning provided in an embodiment of this application.

[0046] like Figure 1As shown, the distributed inference framework of this application embodiment is illustrated using two computing nodes as an example, namely computing node A and computing node B. During the execution of the inference task, the inference request Q is first input to computing node A. This node uses its internal CPU / GPU, AI accelerator, and inference module to complete part of the computational tasks for the inference request Q (such as the initial stage of model forward propagation). Computing node B can then continue to execute subsequent inference steps (such as autoregressive decoding) and finally generate a complete output, realizing resource sharing and efficient collaborative processing.

[0047] The limitations of models that rely on fixed training knowledge make them ill-equipped to effectively address dynamically changing information demands. To solve this problem, retrieval-enhanced generation is introduced. By retrieving external knowledge bases in real time during the inference phase, the model is provided with accurate contextual information, thereby constructing a collaborative working model of "understanding-retrieval-generation".

[0048] Specifically, under this architecture, this embodiment proposes a task reasoning method applicable to processing nodes composed of heterogeneous computing units. The first node is used to obtain at least one first task to be reasoned; and performs a retrieval operation for each task to obtain the corresponding first retrieval result, and then stores the first retrieval result in a shared memory pool for the second node to read as needed. After obtaining the original task request and the corresponding retrieval content, the second node sends them together as input to the first model to complete complex reasoning calculations.

[0049] It should be noted that the first or second node can choose to cooperate with computing accelerators such as CPU, GPU and NPU, and the specific type is not limited, as long as the functional division and collaborative access are met.

[0050] In summary, by utilizing shared memory to build a communication connection between the first and second nodes, the latency issues caused by cross-node transmission are effectively avoided, thereby improving the overall execution efficiency of the inference task.

[0051] The following section, with reference to the accompanying diagram, details the specific implementation process of the task reasoning method, using the CPU as the first node and the GPU as the second node as an example.

[0052] Figure 2 This is a first flowchart illustrating a task reasoning method provided in an embodiment of this application.

[0053] Figure 3 This is the first structural schematic diagram of a task reasoning method provided in the embodiments of this application.

[0054] like Figure 2 and Figure 3 As shown, the task reasoning method includes the following steps: S1: The first node obtains the first task to be reasoned.

[0055] The first inference task refers to one or more original task requests received by the first node in the initial stage, awaiting inference processing. These tasks are typically initiated by users, such as natural language tasks that require the participation of a large language model (LLM), such as question answering, text generation, and summary extraction. The initial stage is the initial phase of the model's forward propagation, specifically including input encoding, embedding representation generation, and forward computation of some network layers. The first node is responsible for the initial processing of the first inference task received initially.

[0056] Figure 4 This is a second flowchart illustrating a task reasoning method provided in an embodiment of this application.

[0057] In one implementation, the first node in step S1 can obtain the first inference task in various ways, such as user interaction, system call, or automated process. For example, it can receive user input through the UI module of the terminal module, or receive external requests through a standard API interface, or have the scheduling module automatically generate tasks according to a preset period. No specific limitation is made here. The following description uses UI module input as an example in this embodiment.

[0058] like Figure 4 As shown, step S1 includes steps S1001-S1002.

[0059] S1001: The UI module of the terminal device responds to the user's input operation and obtains the first task to be inferred.

[0060] Figure 5 This is a schematic diagram of an interface for obtaining a first inference task provided in an embodiment of this application.

[0061] like Figure 5 As shown in (a), the terminal device responds to the user's input operation (such as entering "locallhost / ........html") when opening the inference interface in browser page 1, and displays as shown in (a). Figure 5 The first interface 2 is shown in (b) of the diagram. The first interface 2 includes an input box 21 for the user to input a reasoning request, i.e., to submit a reasoning task.

[0062] For example, the terminal device responds to the first reasoning task A, which is "What is the reason for Einstein winning the Nobel Prize?", entered by the user in the input box 21 of the first interface 2.

[0063] S1002: The terminal device sends the first inference task to the first node.

[0064] Continuing with the example above, such as Figure 3As shown, the terminal device (UI module) sends the first reasoning task A, "What is the reason for Einstein winning the Nobel Prize?", to the retrieval module of the first node (computing node - CPU) of the computing device through the API interface.

[0065] S2: The first node searches for the first task to be reasoned and obtains at least one first search result related to the first task to be reasoned.

[0066] The method of obtaining at least one first retrieval result related to the first reasoning task can be by similarity calculation. The implementation method is explained in detail below.

[0067] In one implementation, at least one first retrieval result related to the first reasoning task can be calculated using similarity. The following section combines... Figure 6 An example is provided.

[0068] Figure 6 This is a schematic diagram of a process for obtaining a first search result provided in an embodiment of this application.

[0069] like Figure 6 As shown, step S2 includes steps S21-S23.

[0070] S21: The first node calculates the initial similarity between the first reasoning task and at least one knowledge fragment in the preset vector knowledge base.

[0071] A vector knowledge base is a collection of knowledge pre-built and stored in a system. Each knowledge fragment in the vector knowledge base (such as a document paragraph, question-answer pair, product description, etc.) is processed by a specific embedding model and transformed into a vector representation in a high-dimensional vector space. Here, high-dimensional vector space refers to a semantic space with higher dimensions, typically with vector dimensions of 768, 1024, or 1536. These dimensions are determined by the hidden layer size of the embedding model used (such as BERT, Sentence-BERT, OpenAI Embeddings), effectively capturing the deep semantic features of the text. For example, taking a vector knowledge base containing five knowledge fragments, such as vector Q1 corresponding to knowledge fragment Q, vector W2 corresponding to knowledge fragment W, vector E3 corresponding to knowledge fragment E, vector R3 corresponding to knowledge fragment R, and vector T3 corresponding to knowledge fragment T.

[0072] Continue to combine Figure 6 As shown, step S21 includes steps S211-S212.

[0073] S211: The first node extracts features from the first task to be reasoned, and obtains the task features.

[0074] The retrieval module of the first node is used to obtain the retrieval results of the task to be reasoned, and feature vectors can be extracted using a lightweight model. For example, this lightweight model can be deployed directly inside the first node, or it can be deployed independently outside the first node, allowing the retrieval module to initiate encoding requests through a standard interface.

[0075] In one example, taking the integrated model within the retrieval module as an example, we will continue to combine... Figure 4 As shown, step S211 includes step S1003.

[0076] S1003: The retrieval module of the first node takes the first reasoning task as the input of the pre-trained second model and uses the second model to obtain the task features corresponding to the first reasoning task.

[0077] The second model can be either a neural network model or a machine learning model; no specific limitation is made here. The following explanation uses a second neural network learning model as an example. For instance, the second neural network model can be a feature extraction model.

[0078] For example, the feature extraction model can be a bidirectional encoder representation from transformers (BERT), a sentence-bidirectional encoder representation from transformers (Sentence-BERT), or another text encoder.

[0079] In this context, a pre-trained feature extraction model refers to a neural network model that has been trained on a large-scale dataset before use. To ensure the generalization ability of the feature extraction model, it is necessary to train the model on a large amount of data.

[0080] Figure 7 This is a schematic diagram of a training method for a second neural network model provided in an embodiment of this application.

[0081] like Figure 7 As shown, the training method for the second neural network model includes the following steps: S701: Obtain the training set.

[0082] The training set includes at least one sample reasoning task and the target task features corresponding to each sample reasoning task.

[0083] S702: Train the second neural network model by using at least one sample reasoning task as input and at least one target task feature as output.

[0084] In one example, at least one sample inference task is input into the original neural network model. Using the original neural network model, the predicted task features corresponding to each sample inference task are output. Furthermore, using a preset loss function, the loss value between each predicted task feature and each target task feature is calculated. Based on the loss value, the original neural network model is continuously adjusted. When the loss value is less than a preset threshold, it means that the initial neural network model has been trained and a trained second neural network model is obtained.

[0085] Continue to combine Figure 3 As shown, after the second neural network model is trained, the retrieval module of the first node inputs the first inference task A into the pre-trained second neural network model to obtain task features A1.

[0086] S212: The first node calculates the similarity based on the task features and at least one knowledge fragment in the vector knowledge base to obtain at least one initial similarity.

[0087] The vector knowledge base can be built into the first node, specifically deployed in the storage device of the first node, or deployed on other nodes outside the first node, without any specific restrictions.

[0088] In one example, taking the deployment of a vector knowledge base in the retrieval module as an example, we will continue to combine... Figure 4 As shown, step S212 includes step S1004.

[0089] S1004: The retrieval module of the first node calculates the similarity between the task features and each knowledge fragment in the vector knowledge base to obtain at least one initial similarity.

[0090] Similarity can be determined using formula (1): Formula (1); in, It is the feature vector of the "first task to be reasoned", that is, the task feature. It is a vector of a knowledge fragment in a vector knowledge base.

[0091] Continuing with the example above, using formula (1), the retrieval module of the first node determines the initial similarity between task feature A1 and the five knowledge fragments in the vector knowledge base. For example, the initial similarity SIM1 between task feature A1 and vector Q1 corresponding to knowledge fragment Q is 0.9; the initial similarity SIM2 between task feature A1 and vector W1 corresponding to knowledge fragment W is 0.8; the initial similarity SIM3 between task feature A1 and vector E1 corresponding to knowledge fragment E is 0.5; the initial similarity SIM4 between task feature A1 and vector R1 corresponding to knowledge fragment R is 0.2; and the initial similarity SIM5 between task feature A1 and vector T1 corresponding to knowledge fragment T is 0.1.

[0092] S22: The first node determines at least one target similarity from at least one initial similarity based on a preset similarity.

[0093] In one example, continue combining Figure 4 As shown, step S22 includes step S1005.

[0094] S1005: The retrieval module of the first node determines at least one target similarity greater than the preset similarity from at least one initial similarity.

[0095] The preset similarity is set according to actual needs. The preset similarity can be 0.5, 0.6, or other values. There is no single limitation here.

[0096] Continuing with the example above, taking a preset similarity of 0.5 as an example, the initial similarity SIM1 and initial similarity SIM2 are determined from the above five initial similarities to meet the conditions, that is, the target similarity SIM1' and target similarity SIM2'.

[0097] S23: The first node determines at least one first retrieval result related to the first reasoning task based on at least one target similarity.

[0098] In one example, continue combining Figure 4 As shown, step S23 includes step S1006.

[0099] S1006: The retrieval module of the first node determines at least one first retrieval result related to the first reasoning task based on at least one target similarity.

[0100] Continuing with the example above, the retrieval module of the first node determines two first retrieval results related to the first reasoning task based on the two target similarities, namely target similarity SIM1' and target similarity SIM2'. These are first retrieval result S1 and first retrieval result S2. First retrieval result S1 represents knowledge fragment Q, and first retrieval result S2 represents knowledge fragment W. For example, first retrieval result S1: Albert Einstein received the 1921 Nobel Prize in Physics for his discovery of the photoelectric effect, not for relativity. First retrieval result S2: Official Nobel Prize records show that Einstein's award citation explicitly points to the theoretical explanation of the photoelectric effect.

[0101] S3: The first node stores at least one first retrieval result into the shared memory pool.

[0102] The shared memory pool is a memory area that can be accessed by both the first and second nodes.

[0103] In one implementation, the combination continues. Figure 4 As shown, step S3 includes step S1007.

[0104] S1007: The retrieval module of the first node stores at least one first retrieval result into the shared memory pool through the interconnection protocol.

[0105] The interconnect protocols include the CXL.mem protocol, the Open Coherent Accelerator Processor Interface (OpenCAPI) protocol, and the Generation Z Protocol (Gen-Z) protocol. The CXL.mem protocol is one of the three sub-protocols in the Fast Interconnect Standard for Computing, allowing host CPUs to efficiently access device-side memory. The OpenCAPI protocol is a processor interconnect standard that supports memory semantics and bidirectional cache coherency, allowing CPUs and accelerators to access each other's memory space as efficiently as accessing local memory. The Gen-Z protocol is an open, memory-semantic-based high-speed interconnect architecture and communication protocol designed to enable byte-level direct access and sharing of computing resources such as processors, memory, storage, and accelerators through a low-latency, high-bandwidth interconnect network.

[0106] Continuing with the example above, such as Figure 3 As shown, the two initial search results are transmitted via the CXL Switch device through the CXL.mem protocol and written to the shared memory pool. The CXL Switch device acts as a bridge between the first node and the shared memory pool, enabling dynamic connection and data forwarding, thereby achieving efficient, low-latency, and high-bandwidth memory access capabilities in a heterogeneous computing architecture.

[0107] S4: The first node records the storage status of the first search result and generates the first storage address of the first search result.

[0108] In one implementation, continue to combine Figure 4 As shown, step S4 includes steps S1008-S1011.

[0109] S1008: The retrieval module of the first node sends the target information to the data consistency and scheduling engine of the first node.

[0110] The target information includes the storage status of the first search result, the data length of the first inference task, and the first storage address of the first search result. The first storage address is the valid address returned when the cache is not locked. The storage status of the first search result includes "write complete" and "writing in progress".

[0111] S1009: The data consistency and scheduling engine of the first node records the storage status of the first retrieval result based on the target information.

[0112] S1010: The first node's data consistency and the scheduling engine send the first storage address of the first retrieval result to the second node.

[0113] It should be noted that in inference task scenarios, the second node typically exists as multiple instances. In this case, the first node needs to work in conjunction with the central scheduling engine (not shown in the diagram). The central scheduling engine makes task allocation decisions and notifies the first node's data consistency and scheduling engine, instructing it to forward the task to the designated second node, thereby achieving efficient resource scheduling and task distribution.

[0114] S5: The second node retrieves the first search result based on the first storage address.

[0115] Figure 8 This is a third flowchart illustrating a task reasoning method provided in an embodiment of this application.

[0116] Continue to combine Figure 8 As shown, step S5 includes step S51.

[0117] S51: The second node reads at least one first retrieval result from the shared memory pool based on the first storage address.

[0118] In one example, continue combining Figure 4 As shown, step S51 includes step S1011.

[0119] S1011: The second node obtains the first retrieval result from the shared memory pool based on the first storage address through the interconnection protocol.

[0120] like Figure 3 As shown, for example, the interaction between the second node and the shared memory pool requires a CXL Switch device as a bridge to obtain the first retrieval result corresponding to the first storage address from the shared memory pool.

[0121] S6: The second node takes at least one first reasoning task and at least one first retrieval result corresponding to the first reasoning task as input to the first model, and uses the first model to execute the first reasoning task.

[0122] In order to distinguish it from the aforementioned lightweight neural network model, the first model used by the second node can be defined as the first model, which is usually a large language model with strong reasoning ability.

[0123] In one example, continue combining Figure 4 As shown, step S6 includes step S1012.

[0124] S1012: The second node takes at least one first reasoning task and at least one first retrieval result corresponding to the first reasoning task as input to the second model, and uses the second model to execute the first reasoning task.

[0125] Continue to combine Figure 8 As shown, step S6 includes steps S61-S63.

[0126] S61: The second node processes the first task to be reasoned and at least one first retrieval result corresponding to the first task to be reasoned, and obtains the input prompt sequence corresponding to the first task to be reasoned.

[0127] For example, the search results are combined with relevant information about the task to be reasoned to construct an input prompt sequence for processing by the neural network model. Specifically, continuing the example above, the constructed input prompt sequence is: "Question: What is the reason Einstein received the Nobel Prize? - Relevant information: Albert Einstein received the 1921 Nobel Prize in Physics for his discovery of the photoelectric effect, not for relativity. Official Nobel Prize records show that Einstein's award reason clearly points to the theoretical explanation of the photoelectric effect - Please answer the question based on the above information." S62: The second node uses the second model to process the input prompt sequence and obtain at least one key-value cache.

[0128] The execution of the inference task includes a pre-filling phase and an autoregressive decoding phase.

[0129] The pre-filling stage refers to the process by which the second model fully encodes the received input prompt sequence in one go. Specifically, after the input prompt sequence is segmented into multiple input tokens, the model uses these tokens as parallel inputs to the complete sequence, performs attention calculations layer by layer during the forward propagation process, and generates corresponding key and value vectors for each input token at each attention layer.

[0130] Since the second model typically contains multiple stacked attention modules, such as 12 layers, each layer independently performs feature transformation and context modeling on the input sequence. In the first layer, based on the embedding representation of the input token, a first set of key and value vectors is generated for 6 tokens. In the second layer, using the output of the previous layer as input, new key and value vectors are generated again for these 6 tokens. This process continues, with each layer independently performing feature transformation and context modeling, continuously generating key and value vectors for the corresponding layer. Finally, the key and value vectors generated by all layers are organized hierarchically according to the network structure, forming a multi-layered key-value cache. In other words, the network hierarchy organization means that the key-value cache is stored and managed hierarchically according to the physical hierarchy of the attention layers in the neural network model. The key and value vectors generated by each attention module are independently stored as a cache unit and are not mixed with other layers. At least one key-value cache in step S62 is generated during the pre-filling stage of task inference when the second model performs forward computation on the input prompt sequence. It contains the key and value vectors generated by each attention layer for each input token in the sequence, forming a set of key-value pairs organized hierarchically according to the network structure.

[0131] For example, suppose the input prompt "What is the reason for Einstein winning the Nobel Prize?" is segmented into 6 input tokens. In the prefill stage of the neural network model, the second node encodes these 6 tokens. In each layer of the attention mechanism, the model calculates and generates corresponding key and value vectors for these 6 tokens respectively. These vectors are organized into key-value caches layer by layer.

[0132] S63: The second node stores at least one key-value cache in the shared memory pool and performs the first inference task based on the read key-value cache during the regression decoding phase.

[0133] The Autoregressive Decoding Phase, also known as the autoregressive decoding phase, refers to the process by which the model gradually generates the output sequence after the input cue sequence has been fully encoded in the pre-filling phase. This process is executed in an autoregressive manner, meaning that only one output token is generated at each step, and the generated token is used as the input for the next step, repeating this process until a complete response is generated or the termination condition is met. As shown in the previous example, in the pre-filling phase, the key-value cache generated for the input cue "What is the reason Einstein won the Nobel Prize?"—that is, the key and value vectors calculated by each attention layer for the six input tokens—is written to the shared memory pool by the second node. In the subsequent autoregressive decoding phase, the model does not need to recalculate the attention representations of these input tokens; instead, it directly reads and reuses the cached key and value information from the shared memory pool, effectively avoiding the large amount of redundant computation caused by repeatedly processing the input cue sequence in each decoding step, thus significantly improving inference efficiency.

[0134] After the first inference task is completed, the shared storage space occupied by the task in the shared memory pool is released according to the following process, which will be further explained in conjunction with step S7 below.

[0135] S7: After the first task to be inferred is completed, release the shared storage space of the first task to be inferred in the shared memory pool.

[0136] The shared storage space includes the storage space occupied by the first retrieval result and key-value cache corresponding to the first inference task.

[0137] In one example, the second node sends a memory release request to the resource management module of the shared memory pool. The request includes the identification information of the memory region to be released, which includes the first search result and the starting address and length of the key-value cache. Subsequently, the resource management module verifies the execution status of the task, confirming that the task has been completed and no other nodes are accessing the relevant data. After successful verification, the resource management module marks the memory region as "free" and adds it to the free memory linked list for subsequent task allocation. Finally, the resource management module returns a release success response to the second node, completing the reclamation of the shared storage space occupied by the task.

[0138] In summary, by utilizing shared memory to build a communication connection between the first and second nodes, the latency issues caused by cross-node transmission are effectively avoided, thereby improving the overall execution efficiency of the inference task.

[0139] Corresponding to the aforementioned task reasoning method embodiments, after the initial reasoning task is completed and preliminary results are generated, if the user is not satisfied with the results or has a need for further optimization, secondary or even multiple iterative reasoning is also supported. The following takes secondary reasoning as an example, and the following steps are performed on the basis of the aforementioned step S7: Figure 9 This is the fourth flowchart of a task reasoning method provided in an embodiment of this application.

[0140] As shown in Figure 9, further optimization of the reasoning method includes the following steps: S8: The first node obtains at least one second task to be inferred.

[0141] The second reasoning task is generated based on the user's subsequent input to the first reasoning task.

[0142] Step S8 can be referred to in similar detail to step S1 above, and will not be repeated here.

[0143] S9: The first node searches for the second task to be reasoned and obtains at least one second search result related to the second task to be reasoned.

[0144] Step S9 can be referred to the specific content of step S2 above, and will not be repeated here.

[0145] S10: The first node stores at least one second retrieval result into the shared memory pool.

[0146] S11: The first node records the storage status of the second search result and generates the second storage address of the second search result.

[0147] S12: The second node obtains the second storage address of the second search result.

[0148] S13: The second node reads at least one second retrieval result based on the second storage address.

[0149] S14: The second node takes at least one second reasoning task and at least one second retrieval result corresponding to the second reasoning task as input to the first model, and uses the first model to execute the second reasoning task.

[0150] S15: After the second task to be inferred is completed, release the shared storage space of the second task to be inferred in the shared memory pool.

[0151] The specific details of steps S10-S15 can be found in similar details to steps S3-S7 above, and will not be repeated here.

[0152] In summary, when the initial inference result fails to meet the expected requirements, a second inference can be performed by using shared memory to build a communication connection between the first and second nodes. This effectively avoids the latency issues caused by cross-node transmission and thus improves the overall execution efficiency of the inference task.

[0153] It should be noted that the above embodiments are illustrated by using a CPU node as the first node and a GPU node as the second node. However, this solution is not limited to this. In actual deployment, the hardware types of the first and second nodes can be flexibly configured according to specific scenarios: for example, the first node can be a GPU node and the second node a CPU node, or both can be CPU nodes, or both can be GPU nodes, or the first node can be a heterogeneous computing node including both CPU and GPU, and the second node can also be a heterogeneous computing node including both CPU and GPU. This solution does not limit this and can be adapted according to the characteristics of the computing task, resource availability, and performance requirements. Corresponding to the aforementioned embodiments of the task inference method, this application also provides embodiments of a task inference method apparatus.

[0154] Figure 10 This is a schematic diagram of a task reasoning device provided in an embodiment of this application.

[0155] like Figure 10 As shown, the task reasoning device 1000 includes: an acquisition module 1001, a storage module 1002, a reading module 1003, a reasoning module 1004, a generation module 1005, and a deletion module 1006.

[0156] The acquisition module 1001 is configured to acquire a first task to be inferred; the acquisition module 1001 is also configured to search the first task to be inferred and acquire at least one first search result related to the first task to be inferred; the storage module 1002 is configured to store at least one first search result in a shared memory pool; wherein, the shared memory pool is a unified memory area that supports access by the first node and the second node; the reading module 1003 is configured to read at least one first search result from the shared memory pool; the inference module 1004 is configured to use the first task to be inferred and at least one first search result corresponding to the first task to be inferred as input to the first model, and use the first model to execute the first task to be inferred.

[0157] In one possible implementation, the acquisition module 1001 is specifically configured to: calculate the initial similarity between the first reasoning task and at least one knowledge fragment in a preset vector knowledge base; determine at least one target similarity from the at least one initial similarity based on the preset similarity; and determine at least one first retrieval result related to the first reasoning task based on the at least one target similarity.

[0158] In one possible implementation, the generation module 1005 is further configured to: record the storage status of the first search result, generate the first storage address corresponding to the first search result, and send the first storage address.

[0159] In one possible implementation, the reading module 1003 is specifically configured to read at least one first retrieval result from the shared memory pool based on a first storage address.

[0160] In one possible implementation, the inference module 1004 is specifically configured to: process the first task to be inferred and at least one first retrieval result corresponding to the first task to be inferred to obtain an input prompt sequence corresponding to the first task to be inferred; process the input prompt sequence using a neural network model to obtain at least one key-value cache; wherein, the at least one key-value cache is a set of key-value pairs generated at each attention layer for each input token in the input prompt sequence when the neural network model encodes the input prompt sequence during the task inference process; store the at least one key-value cache in a shared memory pool, and execute the first task to be inferred based on the read key-value cache during the regression decoding stage.

[0161] In one possible implementation, a deletion module 1006 is also included, configured to: release the shared storage space of the first task to be inferred in the shared memory pool; wherein the shared storage space includes the storage space occupied by the first retrieval result and key-value cache corresponding to the first task to be inferred.

[0162] In one possible implementation, the acquisition module 1001 is specifically configured to: acquire a second inference task; the second inference task is an inference task generated based on subsequent user input to the first inference task; the acquisition module 1002 is further specifically configured to: search the second inference task and acquire at least one second search result related to the second inference task; the storage module 1003 is specifically configured to: store at least one second search result in a shared memory pool; the reading module 1004 is specifically configured to: read at least one second search result from the shared memory pool; and the inference module 1005 is specifically configured to: use the second inference task and at least one second search result corresponding to the second inference task as input to a first model, and execute the second inference task using the first model.

[0163] In one possible implementation, the storage module 1003 is specifically configured to store at least one first retrieval result to a shared memory pool via an interconnect protocol; wherein the interconnect protocol includes the CXL.mem protocol, the OpenCAPI protocol, or the Gen-Z protocol.

[0164] Figure 11 is a schematic diagram of a computing device provided in an embodiment of this application.

[0165] like Figure 11 As shown, the computing device 1100 includes a processor 1101 and a memory 1102. Exemplarily, the computing device 1100 may also include a communications interface 1103 and a communications bus 1104.

[0166] The processor 1101, memory 1102, and communication interface 1103 communicate with each other via communication bus 1104. The communication interface 1103 may include a transmitter and receiver for communicating with other devices or communication networks. It can be a wired interface (port), such as a fiber distributed data interface (FDDI) or a gigabit Ethernet interface (GE).

[0167] In some embodiments, the processor 1101 is used to execute program 1105, specifically performing the relevant steps in the above-described inference task execution method embodiments. Specifically, program 1105 may include program code, which includes computer-executable instructions.

[0168] For example, processor 1101 may be a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement some embodiments of this application. Computing device 1100 may include one or more processors, which may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs. The CPU may be a single-core CPU or a multi-core CPU.

[0169] In some embodiments, memory 1102 is used to store program 1105. Memory 1102 may include high-speed random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device.

[0170] Specifically, program 1105 can be called by processor 1101 to cause computing device 1100 to perform unit test code generation operations.

[0171] Some embodiments of this application provide a computer-readable storage medium storing at least one executable instruction that, when executed on a computing device 1100, causes the computing device 1000 to perform the task reasoning method described in the above embodiments.

[0172] For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device.

[0173] Some embodiments of this application provide a chip system applied to a server. The chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via lines. The interface circuits are used to receive signals from the server's memory and send signals to the processors, the signals including computer instructions stored in the memory. When the server processor executes the computer instructions, the server performs various steps in the task reasoning method shown in the above-described method embodiments.

[0174] The beneficial effects that the readable storage medium provided in some embodiments of this application can achieve can be referred to the beneficial effects in the corresponding task reasoning and execution methods provided above, and will not be repeated here.

[0175] It should be noted that, in this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0176] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0177] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0178] For the purposes of this specification, "computer-readable medium" can mean any means that can contain, store, communicate, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0179] More specific examples of computer-readable media (a non-exhaustive list) include the following: electrical connections having one or more wires (electronic devices), portable computer disks (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM).

[0180] Furthermore, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory. It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof.

[0181] In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc. The above embodiments are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made based on the technical solutions of this application should be included within the scope of protection of this application.

Claims

1. A task inference method applied to a distributed inference architecture, the distributed inference architecture comprising a first node, a second node and a shared memory pool, characterized in that, Comprise: The first node obtains a first to-be-reasoned task; The first node retrieves the first to-be-reasoned task and obtains at least one first retrieval result related to the first to-be-reasoned task; The first node stores the at least one first retrieval result to the shared memory pool; wherein the shared memory pool is a memory area supporting access by the first node and the second node; The second node reads the at least one first retrieval result from the shared memory pool; The second node takes the first to-be-reasoned task and the at least one first retrieval result corresponding to the first to-be-reasoned task as input of a first model, and executes the first to-be-reasoned task by using the first model.

2. The task reasoning method of claim 1, wherein, The first node retrieves the first to-be-reasoned task and obtains at least one first retrieval result related to the first to-be-reasoned task, comprising: The first node calculates the initial similarity of the first to-be-reasoned task and at least one knowledge segment in a preset vector knowledge base; The first node determines at least one target similarity from at least one of the initial similarities based on a preset similarity; The first node determines the at least one first retrieval result related to the first to-be-reasoned task based on the at least one target similarity.

3. The task reasoning method according to claim 1 or 2, characterized in that, Before the second node reads the at least one first retrieval result from the shared memory pool, further comprising: The first node records the storage state of the first retrieval result and generates a first storage address corresponding to the first retrieval result; The first node sends the first storage address.

4. The task reasoning method of claim 3, wherein, The second node reads the at least one first retrieval result from the shared memory pool, comprising: The second node reads the at least one first retrieval result from the shared memory pool based on the first storage address.

5. The task reasoning method according to any one of claims 1-4, characterized in that, The second node takes the at least one first to-be-reasoned task and the at least one first retrieval result corresponding to the first to-be-reasoned task as input of a first model, and executes the first to-be-reasoned task by using the first model, comprising: The second node processes the first to-be-reasoned task and the at least one first retrieval result corresponding to the first to-be-reasoned task to obtain an input prompt sequence corresponding to the first to-be-reasoned task; The second node processes the input prompt sequence by using the first model to obtain at least one key-value cache; wherein the at least one key-value cache is a key-value pair set generated for each input token in the input prompt sequence at each attention layer during task reasoning by the first model encoding the input prompt sequence; The second node stores the at least one key-value cache in the shared memory pool, and executes the first to-be-reasoned task based on the read key-value cache in the regression decoding stage.

6. The task reasoning method according to any one of claims 1-5, wherein, After the first to-be-reasoned task is executed, further comprising: Release shared storage space of the first to-be-reasoned task in the shared memory pool; wherein the shared storage space includes storage space occupied by the first search result corresponding to the first to-be-reasoned task and the key-value cache.

7. The task reasoning method according to any one of claims 1-4, characterized in that, After the first to-be-reasoned task is executed, the method further includes: The first node obtains a second to-be-reasoned task; the second to-be-reasoned task is a reasoning task generated based on subsequent input of a user to the first to-be-reasoned task; The first node performs search on the second to-be-reasoned task to obtain at least one second search result related to the second to-be-reasoned task; The first node stores the at least one second search result to the shared memory pool; The second node reads the at least one second search result from the shared memory pool; The second node takes the second to-be-reasoned task and the at least one second search result corresponding to the second to-be-reasoned task as input of the first model, and executes the second to-be-reasoned task by using the first model.

8. The task reasoning method according to any one of claims 1-7, characterized in that, The first node stores the at least one first search result to the shared memory pool, including: The first node stores the at least one first search result to the shared memory pool by using an interconnection protocol; wherein the interconnection protocol includes a CXL.mem protocol, an OpenCAPI protocol, or a Gen-Z protocol.

9. A task inference apparatus characterized by comprising: including: An obtaining module configured to obtain at least one first to-be-reasoned task; The obtaining module is further configured to perform search on the first to-be-reasoned task to obtain at least one first search result related to the first to-be-reasoned task; A storage module configured to store the at least one first search result to a shared memory pool; wherein the shared memory pool is a unified memory area supporting access of a first node and a second node; A reading module configured to read the at least one first search result from the shared memory pool; A reasoning module configured to take the at least one first to-be-reasoned task and the at least one first search result corresponding to the first to-be-reasoned task as input of a first model, and execute the first to-be-reasoned task by using the first model.

10. A computing device, comprising: The computing device includes a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, the computer program code includes computer instructions, when the processor is used to execute the computer instructions, to make the computing device execute the task reasoning method in any one of claims 1 to 8.