Language Model Retrieval Inference System and Method Based on High-Speed Interconnection
The high-speed interconnect architecture with adaptive partitioning optimizes resource allocation and reduces memory migrations, addressing inefficiencies in existing language model inference systems by enhancing memory bandwidth and computation efficiency.
Patent Information
- Application Number
- CN202510389049.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-31
AI Technical Summary
In the inference system of distributed language model, memory capacity and bandwidth cannot meet the needs, resulting in frequent cross-host data migration, increasing performance overhead, and memory resource scheduling lacks flexibility and intelligence, and cannot be dynamically adjusted according to tasks.
The language model retrieval inference system based on high-speed intercommunication is adopted, and the switch and adaptive partition module are connected through high-speed intercommunication, and the memory pool sharing with high-bandwidth and low-latency are realized, computing resources are dynamically allocated, cross-host memory migration is reduced, and memory sharing and resource usage efficiency is optimized.
It improves the processing efficiency of language model retrieval and inference tasks, supports the efficient operation of large-scale models and data, reduces the performance losses caused by memory migration, optimizes the efficiency of computing resources, and avoids the problems of insufficient resources or excessively affecting the execution efficiency of inference tasks.
Smart Images

Figure CN119902900B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a language model retrieval and inference system and method based on high-speed interconnection. Background Art
[0002] Currently, for the inference application of language models, a multi-host distributed system architecture is usually adopted, and the work of the application is simultaneously distributed to multiple hosts for simultaneous processing. Multiple GPUs (Graphic Processing Units) are configured in each host and connected to their respective GPUs through a PCIe (Peripheral Component Interconnect Express) switch for accelerating computing tasks. Among them, the GPU video memory is usually used to store model parameters and intermediate calculation results. However, due to the limited capacity of the GPU video memory, when the data processed by the GPU exceeds the video memory, some computing loads need to be allocated to the host's DDR (Double Data Rate) memory or NVMe (Non-Volatile Memory Express) storage device. At this time, the data needs to be first sent back to the host and then migrated to the storage unit of the host (such as DDR or NVMe), and the host CPU processes the data in the storage unit to ensure the continuation of the computing task.
[0003] Although the prior art has made certain progress in implementing a distributed language model inference system, there are still some significant drawbacks: the memory capacity and bandwidth cannot meet the requirements; large-scale models and computing tasks often require frequent memory migration across hosts, which brings a large performance overhead; at the same time, the memory resource scheduling of related technologies usually relies on static configuration or simple load balancing algorithms, lacking intelligence and flexibility, and unable to adjust the memory allocation object in real time according to the dynamic requirements of tasks. Summary of the Invention
[0004] This application provides a language model retrieval and inference system and method based on high-speed interconnection, so as to at least solve the problems that in the implementation of distributed language model inference in related technologies, the memory capacity and bandwidth cannot meet the requirements, the performance overhead caused by continuous data migration across hosts during computing tasks, and the inability to adjust the memory allocation object in real time according to the dynamic requirements of tasks.
[0005] The present application provides a language model retrieval and inference system based on high-speed interconnection. The system includes: a host, a high-speed interconnection switch, a high-speed interconnection memory pool, and an image processor. The high-speed interconnection memory pool includes near-memory computing memory and double data rate memory. The high-speed interconnection switch contains an adaptive partitioning module. Among them, the host is connected to the high-speed interconnection memory pool and the image processor respectively through the high-speed interconnection switch.
[0006] The host is configured to receive a first prompt word, start a retrieval phase based on the first prompt word, perform vectorization processing and database retrieval on the first prompt word, generate a second prompt word, and store the second prompt word in the double data rate memory through the high-speed interconnection switch.
[0007] The host is further configured to extract the second prompt word from the double data rate memory through the high-speed interconnection switch, perform inference on the second prompt word, and use the adaptive partitioning module in the connected high-speed interconnection switch to determine the task assignment object in the inference phase.
[0008] The adaptive partitioning module is configured to obtain the computing power of all step tasks executed by the second prompt word in the inference phase, determine the assignment object of the step tasks based on the computing power, and assign the step tasks to the corresponding assignment objects for task calculation, where the assignment objects are near-memory computing memory or image processors.
[0009] The present application provides a language model retrieval and inference method based on high-speed interconnection technology. The method is applied to a language model retrieval and inference system based on high-speed interconnection. The method includes:
[0010] Receiving a first prompt word, starting a retrieval phase based on the first prompt word, performing vectorization processing and database retrieval on the first prompt word, and generating a second prompt word;
[0011] Extracting the second prompt word and performing inference on the second prompt word;
[0012] Obtaining the computing power of all step tasks executed by the second prompt word in the inference phase, and determining the assignment object of the step tasks based on the computing power, where the assignment object is used to perform task calculation on the step tasks.
[0013] The present application provides a language model retrieval and inference device based on high-speed interconnection technology. The device includes:
[0014] A receiving module, configured to receive a first prompt word, start a retrieval phase based on the first prompt word, perform vectorization processing and database retrieval on the first prompt word, and generate a second prompt word;
[0015] An extraction module, configured to extract the second prompt word and perform inference on the second prompt word;
[0016] An acquisition module, configured to acquire the computing power of all step tasks executed by the second prompt word in the inference stage, and determine the allocation object of the step tasks based on the computing power, where the allocation object is used to perform task calculations on the step tasks.
[0017] The present application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing any one of the above language model retrieval inference methods based on high-speed interconnection technology when executing the computer program.
[0018] The present application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, any one of the above language model retrieval inference methods based on high-speed interconnection technology is implemented.
[0019] The present application provides a computer program product, including computer instructions for causing a computer to execute any one of the above language model retrieval inference methods based on high-speed interconnection technology.
[0020] Through this application, due to the designed language model retrieval and inference system based on high-speed interconnection, including: a host, a high-speed interconnection switch, a high-speed interconnection memory pool, and an image processor. The high-speed interconnection memory pool includes near-memory computing memory and double data rate memory. The high-speed interconnection switch contains an adaptive partitioning module. Among them, the host is connected to the high-speed interconnection memory pool and the image processor respectively through the high-speed interconnection switch. Through this system architecture, multiple computing nodes (i.e., the host and the image processor) can share a high-bandwidth and low-latency memory pool through the high-speed interconnection memory pool, thereby providing higher memory bandwidth, improving the processing efficiency of language model retrieval and inference tasks, and supporting the efficient operation of large-scale models and data. At the same time, due to the language model retrieval and inference system based on high-speed interconnection provided by the embodiments of this application, the host receives the first prompt word, and based on the first prompt word, starts the retrieval stage, performs vectorization processing and database retrieval on the first prompt word, generates the second prompt word, and stores the second prompt word in the double data rate memory through the high-speed interconnection switch. The host can also extract the second prompt word from the double data rate memory through the high-speed interconnection switch, perform inference on the second prompt word, and use the adaptive partitioning module in the connected high-speed interconnection switch to determine the task assignment object in the inference stage. Using the adaptive partitioning module in the high-speed interconnection switch, obtain the computing power of all step tasks executed by the second prompt word in the inference stage, and based on the computing power, determine the assignment object of the step task, and assign the step task to the corresponding assignment object for task calculation, where the assignment object is near-memory computing memory or an image processor. In this way, since the generated second prompt word is stored in the double data rate memory through the high-speed interconnection switch, when subsequent data interaction is performed, there is no need to pass through the host, and only need to extract data from the double data rate memory, reducing the need for memory migration, optimizing memory sharing, especially reducing cross-host memory migration through the high-speed interconnection memory pool, thereby reducing the performance loss caused by memory migration. In addition, by performing task calculation on the computing power generated by each step task in the inference stage and determining the assignment object of the step task based on the computing power, dynamic allocation of resources is achieved, optimizing the use efficiency of computing resources, and avoiding the problem of affecting the execution efficiency of inference tasks due to excessive or insufficient resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1Schematic diagram of a language model retrieval and inference system based on high-speed interconnection provided by an embodiment of the present application;
[0023] Figure 2 Schematic diagram of a memory access path provided by an embodiment of the present application;
[0024] Figure 3 Schematic flowchart of a method for language model retrieval and inference based on high-speed interconnection provided by an embodiment of the present application;
[0025] Figure 4 Block diagram of a structure of a language model retrieval and inference device based on high-speed interconnection technology provided by an embodiment of the present application;
[0026] Figure 5 Schematic diagram of a hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0028] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0029] In order to enable those skilled in the art of this technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0030] With the rapid development of artificial intelligence and deep learning technologies, large language models and retrieval-augmented generation systems have achieved remarkable results in the field of natural language processing. These applications require efficient computing resources and large-scale memory support to achieve an efficient inference process. However, traditional memory architectures and computing resource configurations cannot meet these needs, limiting the efficiency and scale of model inference. Especially for complex large language models and retrieval-augmented generation, traditional memory management technologies can no longer meet their requirements for high bandwidth, low latency, and large-scale memory access, and factors such as memory capacity, bandwidth, and access latency often become performance bottlenecks. Therefore, the Compute Express Link (CXL) protocol and distributed computing architectures have been proposed to break through these bottlenecks and achieve more efficient inference computing.
[0031] However, although CXL technology and distributed architectures provide new means for language model inference, existing technologies have limitations in the following key aspects:
[0032] a. Bottlenecks in memory capacity and bandwidth
[0033] Large language models and retrieval-augmented generation applications typically need to process models and data in the hundreds of gigabytes or even larger scale. Traditional memory architectures cannot provide sufficient memory capacity and bandwidth to support the efficient operation of these applications. Even with the acceleration of Graphics Processing Units (GPUs), the capacity and bandwidth of GPU video memory often become performance bottlenecks. Existing technologies mainly rely on frequent data transfer between GPU video memory and host memory, resulting in increased memory bandwidth pressure and thus affecting inference efficiency.
[0034] b. Overhead of memory migration
[0035] Large-scale models and complex computing tasks require frequent data migration between different memory regions (such as CPU memory, GPU video memory, and Proximity Computing Memory (PNM)). However, this memory migration incurs significant overhead, especially in cross-host memory sharing and data exchange. Memory migration in existing technologies is usually synchronous, resulting in increased waiting time for computing tasks and reduced overall system throughput. Inflexible memory management also prevents the system from fully leveraging the maximum efficiency of resources under high-concurrency tasks.
[0036] c. Insufficient compute offloading and resource scheduling
[0037] Existing computing architectures often rely on fixed resource configurations, such as forcibly offloading computing tasks to GPUs or CPUs, without dynamically adjusting computing resources according to task requirements. This fixed resource scheduling method cannot flexibly select the best computing resources based on characteristics such as the arithmetic intensity and memory access pattern of tasks, resulting in low processing efficiency of computing loads. Especially when facing dynamically changing workloads, optimal computing resource allocation cannot be achieved.
[0038] d. Management and scheduling complexity of the memory pool
[0039] The proposal of the CXL protocol provides a new direction for memory pool management, but existing CXL memory pool technologies still have certain deficiencies in dynamic management and scheduling. Current memory pool management lacks sufficient intelligence and automation. The allocation and adjustment of memory resources mainly rely on manual configuration or simple static rules, and the advantages of CXL technology in memory sharing and dynamic management cannot be fully utilized. Especially in multi-host systems, how to reasonably configure and schedule memory resources to avoid memory waste and excessive competition remains an urgent problem to be solved.
[0040] To solve the above problems, according to the embodiments of the present application, an embodiment of a language model retrieval and inference system based on high-speed interconnection is provided, as Figure 1 shown, Figure 1 including: a host, a high-speed interconnection switch, a high-speed interconnection memory pool, and an image processor. The high-speed interconnection memory pool includes near-memory computing memory and double data rate memory. The high-speed interconnection switch contains an adaptive partitioning module; wherein, the host is respectively connected to the high-speed interconnection memory pool and the image processor through the high-speed interconnection switch;
[0041] The host is used to receive the first prompt word, and based on the first prompt word, start the retrieval stage, perform vectorization processing and database retrieval on the first prompt word, generate the second prompt word, and store the second prompt word in the double data rate memory through the high-speed interconnection switch;
[0042] The host is further used to extract the second prompt word from the double data rate memory through the high-speed interconnection switch, perform inference on the second prompt word, and use the adaptive partitioning module in the connected high-speed interconnection switch to determine the task allocation object in the inference stage;
[0043] The adaptive partitioning module is used to obtain the computing power of all step tasks executed by the second prompt word in the inference stage, and based on the computing power, determine the allocation object of the step tasks, and allocate the step tasks to the corresponding allocation objects for task calculation, where the allocation object is the near-memory computing memory or the image processor.
[0044] Optionally, the language model retrieval and inference system based on high-speed interconnection proposed in the embodiments of the present application Figure 1 consists of the following components: 4 host servers (Host 1, Host 2, Host 3, Host 4), 4 high-speed interconnection switches (i.e., the CXL switches in the figure), 6 image processors (i.e., the GPUs in the figure), and a high-speed interconnection memory pool (i.e., the CXL memory pool in the figure). Near the high-speed interconnection switch end (short physical link) in the high-speed interconnection memory pool, there is near-memory computing memory (i.e., the PNM memory in the figure), and far from the high-speed interconnection switch end (long physical link), there is double data rate memory (i.e., the DDR memory in the figure). In addition, an adaptive partitioning module is included in the high-speed interconnection switch to achieve resource allocation through this adaptive partitioning module.
[0045] It should be noted that the language model retrieval and inference system based on high-speed interconnection in the embodiments of the present application is not limited to Figure 1 the 4 hosts, 4 high-speed interconnection switches, and 6 image processors shown in
[0046] Furthermore, it can be concluded from Figure 1 that the host is respectively connected to the high-speed interconnection memory pool and the image processor through the high-speed interconnection switch. At the same time, double data rate memory is provided in the high-speed interconnection memory pool for data storage.
[0047] In the language model retrieval and inference task, it is jointly completed by four hosts. The language model retrieval and inference task is divided into three stages: RAG (Retrieval-augmented Generation) stage, Prefill stage, and Decode stage.
[0048] RAG stage: The user can input the first prompt word from any host, and multiple hosts are supported to input the first prompt word simultaneously. The host starts the retrieval stage based on the output first prompt word, performs vectorization processing and database retrieval on the first prompt word, generates the second prompt word, and stores the second prompt word in the double data rate memory through the high-speed interconnection switch.
[0049] The Prefill stage of inference is responsible for two hosts (e.g., Host 1 and Host 2). Through a high-speed interconnection switch, it extracts the second prompt word from the double data rate memory and performs inference on the second prompt word: embedding and encoding, layer normalization, QKV mapping, QK matrix calculation, Softmax, SV matrix calculation, output projection, residual connection, layer normalization, fully connected layer calculation, activation function calculation, fully connected layer calculation, residual connection, linear layer, Softmax function, recording the KV address and the value of the first token (i.e., the letter token). Each host collects the second prompt word every unit time (taking one second as an example) and performs model inference using the method of continuous batch processing.
[0050] The Decode stage of inference is responsible for another two hosts (e.g., Host 3 and Host 4). The Decode task process is similar to that of Prefill. Host 3 and Host 4 will obtain the addresses of the KV value and the first token value from Host 1 and Host 2 through RDMA (remote direct memory access) and perform decoding. After that, through the same inference process as the Prefill stage, the inference result is finally obtained.
[0051] In both the Prefill stage and the Decode stage, the following step tasks will be passed through: embedding and encoding, layer normalization, QKV mapping, QK matrix calculation, Softmax, SV matrix calculation, output projection, residual connection, layer normalization, fully connected layer calculation, activation function calculation, fully connected layer calculation, residual connection, linear layer, Softmax function. The computing power of these step tasks is obtained by using the adaptive partition module in the connected high-speed interconnection switch, and the current computing task is allocated to the graphics processor or the near-memory computing memory based on the computing power. Operations with lower arithmetic intensity will generate higher data transmission overheads, so the task steps with lower arithmetic intensity are allocated to the near-memory computing memory for execution, and the task steps with higher arithmetic intensity are performed by the graphics processor.
[0052] Preferably, when the step task needs to be calculated in the near-memory computing memory, the calculation is preferably placed in the near-memory computing memory closer to the host server to reduce latency.
[0053] In the embodiments of the present application, through the design of the language based on high-speed interconnection, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the protection scope of the present application.
[0054] It should be noted that in the description of this application, the terms "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0055] In order to enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.
[0056] With the rapid development of artificial intelligence and deep learning technologies, large language models and retrieval augmented generation systems have achieved remarkable results in the field of natural language processing. These applications require efficient computing resources and large-scale memory support to achieve an efficient inference process. However, traditional memory architectures and computing resource configurations cannot meet these requirements, restricting the efficiency and scale of model inference. Especially for complex large language models and retrieval augmented generation, traditional memory management technologies can no longer meet their requirements for high bandwidth, low latency, and large-scale memory access, and factors such as memory capacity, bandwidth, and access latency often become performance bottlenecks. Therefore, the Compute Express Link (CXL) protocol and distributed computing architectures have been proposed to break through these bottlenecks and achieve more efficient inference computing.
[0057] However, although CXL technology and distributed architectures provide new means for language model inference, the prior art has limitations in the following key aspects:
[0058] a. Bottlenecks in memory capacity and bandwidth
[0059] Large language models and retrieval augmented generation applications usually need to process models and data in the hundreds of gigabytes or even larger scale. Traditional memory architectures cannot provide sufficient memory capacity and bandwidth to support the efficient operation of these applications. Even with the acceleration of a Graphics Processing Unit (GPU), the capacity and bandwidth of the GPU video memory often become performance bottlenecks. The prior art mainly relies on frequent data transfer between the GPU video memory and the host memory, resulting in increased pressure on the memory bandwidth and thus affecting the inference efficiency.
[0060] b. Overhead of memory migration
[0061] Large-scale models and complex computing tasks often require frequent data migration between different memory regions, such as CPU memory, GPU video memory, and PNM (Proximity Computing memory). However, this memory migration incurs significant overhead, especially in cross-host memory sharing and data exchange. Memory migration in existing technologies is usually synchronous, resulting in increased waiting time for computing tasks and reduced throughput of the overall system. The lack of flexibility in memory management also prevents the system from fully leveraging the maximum efficiency of resources under high-concurrency tasks.
[0062] c. Insufficient computing offloading and resource scheduling
[0063] Existing computing architectures often rely on fixed resource configurations, such as forcibly offloading computing tasks to GPUs or CPUs without dynamically adjusting computing resources according to task requirements. This fixed resource scheduling method cannot flexibly select the optimal computing resources based on characteristics such as the arithmetic intensity and memory access pattern of the task, resulting in low processing efficiency of the computing load. Especially when facing dynamically changing workloads, it is impossible to achieve optimal computing resource allocation.
[0064] d. Management and scheduling complexity of memory pools
[0065] The proposed CXL protocol provides a new direction for memory pool management, but existing CXL memory pool technologies still have certain deficiencies in dynamic management and scheduling. Current memory pool management lacks sufficient intelligence and automation. The allocation and adjustment of memory resources mainly rely on manual configuration or simple static rules, failing to fully leverage the advantages of CXL technology in memory sharing and dynamic management. Especially in multi-host systems, how to reasonably configure and schedule memory resources to avoid memory waste and excessive competition remains an urgent problem to be solved.
[0066] To solve the above problems, according to the embodiments of the present application, there is provided an embodiment of a language model retrieval and inference system based on high-speed interconnection, as Figure 1 shown, Figure 1 including: a host, a high-speed interconnection switch, a high-speed interconnection memory pool, and an image processor. The high-speed interconnection memory pool includes proximity computing memory and double data rate memory. The high-speed interconnection switch contains an adaptive partitioning module. Among them, the host is connected to the high-speed interconnection memory pool and the image processor respectively through the high-speed interconnection switch;
[0067] The host is used to receive a first prompt word, start a retrieval phase based on the first prompt word, perform vectorization processing and database retrieval on the first prompt word, generate a second prompt word, and store the second prompt word in the double data rate memory through the high-speed interconnection switch;
[0068] The host is also used to extract the second prompt word from the double data rate memory through the high-speed interconnect switch, perform reasoning on the second prompt word, and use the adaptive partitioning module in the connected high-speed interconnect switch to determine the task assignment object in the reasoning stage;
[0069] The adaptive partitioning module is used to obtain the computing power of all step tasks executed by the second prompt word in the reasoning stage, determine the assignment object of the step tasks based on the computing power, and assign the step tasks to the corresponding assignment objects for task calculation, where the assignment object is the near-memory computing memory or the graphics processor.
[0070] Optionally, the language model retrieval and reasoning system based on high-speed interconnect proposed in the embodiments of the present application Figure 1 consists of the following components: 4 host servers (Host 1, Host 2, Host 3, Host 4), 4 high-speed interconnect switches (i.e., the CXL switches in the figure), 6 graphics processors (i.e., the GPUs in the figure), and a high-speed interconnect memory pool (i.e., the CXL memory pool in the figure). The high-speed interconnect memory pool is equipped with near-memory computing memory (i.e., the PNM memory in the figure) near the high-speed interconnect switch end (short physical link) and double data rate memory (i.e., the DDR memory in the figure) away from the high-speed interconnect switch end (long physical link). In addition, an adaptive partitioning module is included in the high-speed interconnect switch to achieve resource allocation through this adaptive partitioning module.
[0071] It should be noted that the language model retrieval and reasoning system based on high-speed interconnect in the embodiments of the present application is not limited to Figure 1 the 4 hosts, 4 high-speed interconnect switches, and 6 graphics processors shown in
[0072] Furthermore, it can be obtained that the host is connected to the high-speed interconnect memory pool and the graphics processor respectively through the high-speed interconnect switch. At the same time, double data rate memory is provided in the high-speed interconnect memory pool for data storage. Figure 1 In the language model retrieval and reasoning task, it is jointly completed by four hosts. The language model retrieval and reasoning task is divided into a RAG (Retrieval-augmented Generation) stage, a Prefill stage, and a Decode stage.
[0073]
[0074] RAG Phase: The user can input the first prompt from any host, and multiple hosts are supported to input the first prompt simultaneously. The host starts the retrieval phase based on the output first prompt, performs vectorization processing and database retrieval on the first prompt, generates the second prompt, and stores the second prompt in the double data rate memory through a high-speed interconnection switch.
[0075] The Prefill phase of inference is responsible by two hosts (such as Host 1 and Host 2). Through the high-speed interconnection switch, the second prompt is extracted from the double data rate memory, and the second prompt is inferred: embedding and encoding, layer normalization, QKV mapping, QK matrix calculation, Softmax, SV matrix calculation, output projection, residual connection, layer normalization, fully connected layer calculation, activation function calculation, fully connected layer calculation, residual connection, linear layer, Softmax function, record the KV address and the value of the first token (i.e., letter token). The host collects the second prompt every unit time (taking one second as an example), and uses the method of continuous batch processing for model inference.
[0076] The Decode phase of inference is responsible by another two hosts (such as Host 3 and Host 4). The Decode task process is similar to Prefill. Host 3 and Host 4 will obtain the addresses of the KV values and the first token value from Host 1 and Host 2 through RDMA (remote direct memory access) and perform decoding. After that, through the same inference process as the Prefill phase, the inference result is finally obtained.
[0077] In both the Prefill phase and the Decode phase, the following step tasks will be passed through: embedding and encoding, layer normalization, QKV mapping, QK matrix calculation, Softmax, SV matrix calculation, output projection, residual connection, layer normalization, fully connected layer calculation, activation function calculation, fully connected layer calculation, residual connection, linear layer, Softmax function. The computing power of these step tasks is obtained by using the adaptive partitioning module in the connected high-speed interconnection switch, and the current computing task is allocated to the graphics processor or the near-memory computing memory based on the computing power. Operations with lower arithmetic intensity will generate higher data transmission overhead, so the task steps with lower arithmetic intensity are allocated to the near-memory computing memory for execution, and the task steps with higher arithmetic intensity are performed by the graphics processor.
[0078] Preferably, when the step task needs to be calculated in the near-memory computing memory, the calculation is preferably placed in the near-memory computing memory close to the host server to reduce latency.
[0079] In the embodiment of the present application, by designing a language model retrieval and inference system based on high-speed interconnection, the system includes: a host, a high-speed interconnection switch, a high-speed interconnection memory pool, and an image processor. The high-speed interconnection memory pool includes near-memory computing memory and double data rate memory. The high-speed interconnection switch contains an adaptive partitioning module. Among them, the host is connected to the high-speed interconnection memory pool and the image processor respectively through the high-speed interconnection switch. Through this system architecture, multiple computing nodes (i.e., the host and the image processor) can share a high-bandwidth and low-latency memory pool through the high-speed interconnection memory pool, thereby providing higher memory bandwidth and improving the processing efficiency of language model retrieval and inference tasks, and supporting the efficient operation of large-scale models and data. At the same time, since the language model retrieval and inference system based on high-speed interconnection provided by the embodiment of the present application uses the host to receive the first prompt word, and based on the first prompt word, starts the retrieval stage, vectorizes the first prompt word and performs database retrieval to generate the second prompt word, and stores the second prompt word in the double data rate memory through the high-speed interconnection switch. The host can also extract the second prompt word from the double data rate memory through the high-speed interconnection switch, perform inference on the second prompt word, and use the adaptive partitioning module in the connected high-speed interconnection switch to determine the task assignment object in the inference stage. Using the adaptive partitioning module in the high-speed interconnection switch, obtain the computing power of all step tasks executed by the second prompt word in the inference stage, and based on the computing power, determine the assignment object of the step task, and assign the step task to the corresponding assignment object for task calculation, where the assignment object is near-memory computing memory or an image processor. In this way, since the generated second prompt word is stored in the double data rate memory through the high-speed interconnection switch, when subsequent data interaction is performed, there is no need to pass through the host, and only need to extract data from the double data rate memory, reducing the need for memory migration, optimizing memory sharing, especially reducing cross-host memory migration through the high-speed interconnection memory pool, thereby reducing the performance loss caused by memory migration. In addition, by performing task calculation on the computing power generated by each step task in the inference stage and determining the assignment object of the step task based on the computing power, dynamic allocation of resources is achieved, the use efficiency of computing resources is optimized, and the problem of affecting the execution efficiency of inference tasks due to excessive or insufficient resources is avoided.
[0080] In some optional implementation manners, the double data rate memory includes a non-fixed memory pool and a fixed memory pool. Among them, the non-fixed memory pool is used to store intermediate parameters generated in the inference stage, and the fixed memory pool is used to store fixed parameters in the retrieval stage and the inference stage, as well as the storage addresses of the intermediate parameters and the fixed parameters.
[0081] The high-speed interconnection memory pool further includes a memory migration module.
[0082] A memory migration module for storing the first prompt word and the second prompt word into a non-fixed memory pool through a high-speed interconnection switch, migrating the vector value obtained by an allocation object calculating a step task from local memory to the non-fixed memory pool through the high-speed interconnection switch, storing intermediate parameters into the non-fixed memory pool through the high-speed interconnection switch, and storing fixed parameters and storage addresses into a fixed memory pool through the high-speed interconnection switch.
[0083] Optionally, in Figure 1 , the double data rate memory includes a non-fixed memory pool and a fixed memory pool. It should be noted that there is no fixed location for the non-fixed memory pool and the fixed memory pool here. It only loads fixed parameters such as the weights of the language model, fixed parameters in the retrieval stage and the inference stage into the fixed memory pool, and loads intermediate parameters generated in the inference stage, such as KV values in the Prefill stage, the first token value, and the inference result obtained at the end of the inference in the Decode stage, into the non-fixed memory pool to partition the storage of parameters.
[0084] Of course, the non-fixed memory pool also includes variable parameters such as the language model program, the RAG program, the first prompt word, and the second prompt word. In addition, the fixed memory pool also stores the storage addresses of all parameters in the non-fixed memory pool and the fixed memory pool.
[0085] After the double data rate memory is provided with a non-fixed memory pool and a fixed memory pool, based on Figure 1 the memory migration module in the high-speed interconnection memory pool stores the first prompt word and the second prompt word into the non-fixed memory pool through the high-speed interconnection switch, migrates the vector value obtained by an allocation object calculating a step task from PNM memory or the local memory of the GPU to the non-fixed memory pool through the high-speed interconnection switch, stores intermediate parameters into the non-fixed memory pool through the high-speed interconnection switch, and stores fixed parameters and storage addresses into the fixed memory pool through the high-speed interconnection switch.
[0086] In the embodiment of the present application, the memory migration module directly stores data into the double data rate memory through the high-speed interconnection switch without passing through the host, reducing data migration and replication between hosts, greatly reducing the performance loss caused by memory migration, and improving the overall inference efficiency of the system.
[0087] In some optional embodiments, the high-speed interconnection memory pool further includes a high-speed interconnection protocol module, and the high-speed interconnection memory pool further includes a memory configuration module;
[0088] A memory configuration module, configured to initiate an access request from a host or an image processor to a non-fixed memory pool and / or a fixed memory pool in a double data rate memory via a high-speed interconnection protocol module;
[0089] The memory configuration module is further configured to initiate an access request from a near-memory computing memory to a non-fixed memory pool and / or a fixed memory pool in a double data rate memory via a high-speed interconnection protocol module.
[0090] Optionally, as Figure 1 , the high-speed interconnection memory pool further includes a high-speed interconnection protocol module (i.e., the CXL protocol module in Figure 1 ), and the high-speed interconnection memory pool further includes a memory configuration module. Among them, the memory configuration module is configured to initiate access requests from a host or an image processor to both the non-fixed memory pool and the fixed memory pool in the double data rate memory via the high-speed interconnection protocol module, or initiate an access request to either the non-fixed memory pool or the fixed memory pool. As in Figure 2 the data transmission channel 1, the host or the image processor communicates with the high-speed interconnection protocol module through the data transmission channel 1, and then initiates an access request to the non-fixed memory pool and / or the fixed memory pool in the double data rate memory through the high-speed interconnection protocol module.
[0091] The memory configuration module is further configured to initiate access requests from a near-memory computing memory to both the non-fixed memory pool and the fixed memory pool in the double data rate memory via the high-speed interconnection protocol module, or initiate an access request to either the non-fixed memory pool or the fixed memory pool. As in Figure 2 the data transmission channels 2 and 3. The near-memory computing memory first communicates with the high-speed interconnection protocol module through the data transmission channel 2, and then the memory configuration module connects the memory access request to the non-fixed memory pool and / or the fixed memory pool in the double data rate memory through the data transmission channel 3.
[0092] Furthermore, the number of data transmission channels 1, 2, and 3 connecting the high-speed interconnection protocol module to the high-speed interconnection memory pool is limited. For example, there is 1 data transmission channel 1, but the memory configuration module based on the memory pool enables the host and the image processor to access the double data rate memory at any physical address; similarly, there is 1 data transmission channel 2 and 1 data transmission channel 3, but the memory configuration module based on the memory pool enables the near-memory computing memory to access the double data rate memory at any physical address.
[0093] In the embodiments of the present application, the memory configuration module further supports the dynamic connection of the double data rate memory at any physical address in the host, the image processor, and the near-memory computing memory to the high-speed interconnection memory pool, improving the utilization efficiency of memory resources, avoiding resource waste and bottlenecks, and optimizing system performance.
[0094] In this embodiment, a language model retrieval and inference method based on high-speed interconnection technology is provided, which can be used in the above-mentioned language model retrieval and inference system based on high-speed interconnection, such as Figure 3 As shown, the process includes the following steps:
[0095] Step S301, receive the first prompt word, and based on the first prompt word, start the retrieval stage, perform vectorization processing and database retrieval on the first prompt word, and generate the second prompt word.
[0096] Optionally, the host (which can be multiple hosts, such as host 1, host 2, host 3, host 4) included in the language model retrieval and inference system based on high-speed interconnection will receive the first prompt word input by the user through the input interface, and then store the first prompt word in the high-speed interconnection memory pool of the above system, and then send a message to the host to indicate the start of the retrieval augmented generation (i.e., RAG) stage.
[0097] When the host receives the "start RAG stage" message, it will start the RAG task, perform vectorization processing and database retrieval on the first prompt word, and then generate the second prompt word.
[0098] After that, the second prompt word can be put into the non-fixed memory pool of the double data rate memory in the language model retrieval and inference system based on high-speed interconnection.
[0099] Step S302, extract the second prompt word and perform inference on the second prompt word.
[0100] Optionally, then the host can collect the second prompt word and perform inference on the second prompt word. Among them, the inference stage includes two parts, the Prefill stage and the Decode stage. Further, in the embodiments of the present application, host 1 and host 2 can be responsible for the Prefill stage, and host 1 and host 2 collect the second prompt word from the non-fixed memory pool every unit time interval (such as one second). The continuous batch processing method is adopted, that is, multiple second prompt words are concentrated and processed together to improve the inference efficiency.
[0101] Then start the inference: embedding and encoding: perform embedding processing on the collected second prompt word, convert the text into an encoding form suitable for model processing, and further extract the semantic features of the text.
[0102] Layer Normalization: Normalize the data to make the data distribution more stable, which helps the training and inference of the model, and improves the convergence speed and performance of the model.
[0103] QKV Mapping: Map the processed data into Query (query vector), Key (key vector), and Value (value vector), which is an important step in many attention mechanism-based models (such as Transformer).
[0104] QK Matrix Calculation: Calculate the matrix product of Query and Key to measure the similarity between different elements, thereby determining the importance of each element in the current context.
[0105] Softmax Function: Apply the Softmax function to the result of the QK matrix calculation to convert it into a probability distribution for determining the weight of each element.
[0106] SV Matrix Calculation: Calculate the matrix product of the Softmax result and Value to obtain the final weighted result.
[0107] Output Projection: Perform a projection transformation on the result of the SV matrix calculation to convert it into a suitable dimension and form.
[0108] Residual Connection: Introduce a residual connection to add the original input to the output after the above processing, which helps to solve the problems of gradient disappearance and gradient explosion in deep neural networks, and improves the training effect and generalization ability of the model.
[0109] Layer Normalization (again): Perform layer normalization again to further stabilize the data distribution.
[0110] Fully Connected Layer Calculation: Perform a linear transformation on the data through a fully connected layer to extract higher-level feature representations.
[0111] Activation Function Calculation: Apply an activation function to the output of the fully connected layer for non-linear transformation to increase the expressive power of the model.
[0112] Fully Connected Layer Calculation (again): Perform another fully connected layer calculation to further process the data.
[0113] Residual Connection (again): Use the residual connection again to enhance the performance of the model.
[0114] Linear Layer: Perform a final linear transformation on the data through a linear layer.
[0115] Softmax (again): Apply the Softmax function to obtain the final probability distribution for generating prediction results.
[0116] Record the KV address and the first token value: After the inference is completed, store the KV values (i.e., the Key value and the Value value) and the first token value generated during the calculation process in the non-fixed memory pool for use in the subsequent Decode stage.
[0117] The host 3 and the host 4 are responsible for the Decode stage. The host 3 and the host 4 will obtain the addresses of the KV values and the first token value from the host 1 and the host 2 through RDMA and perform decoding operations to generate the final inference result. After the inference is completed, store the generated inference result in the non-fixed memory pool to complete the entire inference process.
[0118] Step S303, obtain the computing power of all step tasks executed by the second prompt word in the inference stage, and determine the allocation object of the step tasks based on the computing power, where the allocation object is used to perform task calculations on the step tasks.
[0119] Optionally, the language model retrieval inference system based on high-speed interconnection calculates the computing power of the step tasks involved in the above Prefill stage and Decode stage (such as embedding and encoding, layer normalization, etc.), and determines the allocation object of the step tasks based on the obtained computing power, so that the allocation object performs task calculations on the corresponding step tasks. Among them, the step tasks with higher arithmetic intensity are calculated by the image processor, and the step tasks with higher arithmetic intensity are calculated by the near-memory computing memory.
[0120] In addition, when calculating the step tasks with higher arithmetic intensity in the near-memory computing memory, preferentially place the calculation in the near-memory computing memory close to the host server to reduce latency.
[0121] The embodiment of the present application optimizes the use efficiency of computing resources by performing task calculations on the computing power generated by each step task in the inference stage, and determines the allocation object of the step tasks based on the computing power, avoiding the problem of affecting the execution efficiency of the inference task due to excessive or insufficient resources.
[0122] In some alternative embodiments, receiving the first prompt word and starting the retrieval stage based on the first prompt word includes:
[0123] When receiving the first prompt word, obtain the first storage address corresponding to the first prompt word;
[0124] Determine the acquisition status of the first prompt word based on the first storage address;
[0125] When the acquisition status meets the preset conditions, start the retrieval stage.
[0126] Optionally, when the host receives the first prompt word, it starts the RAG stage and obtains the first storage address corresponding to the first prompt word at the same time.
[0127] Multiple hosts can jointly perform the RAG task. Each host determines the collection status of the first prompt word according to the first storage address, and then decides whether to start the retrieval phase based on the specific situation of the collection status. For example, the collection status is determined by setting a flag bit, and a specific flag bit is set in the data structure or memory area storing the relevant data of the first prompt word to represent the collection status. For example, a binary bit is used to indicate whether the first prompt word has been collected, where 0 indicates not collected and 1 indicates collected. The host determines the collection status of the first prompt word by reading the value of this flag bit.
[0128] The collection status can also be determined by comparing timestamps. Timestamp information, including the write time of the first prompt word and the expected collection completion time, etc., is recorded for the storage address of each first prompt word. The host determines the collection status of the first prompt word by comparing the current time with the timestamp. For example, if the current time exceeds the expected collection completion time and the first prompt word is still not marked as collected, it may indicate an abnormality in the collection; or by comparing the write time and the current time, it is judged whether the set collection interval time has been reached to decide whether the collection operation can be performed to obtain the collection status of the first prompt word.
[0129] The collection status can also be determined by means of a message queue or an event mechanism, using the message queue or the event mechanism to transmit information about the collection status of the first prompt word. When the first prompt word is stored in the first storage address, a corresponding message is sent or an event is triggered to inform the host that the status of the first prompt word has been updated. The host obtains the collection status of the first prompt word by listening to the message queue or the event.
[0130] In some alternative embodiments, determining the collection status of the first prompt word based on the first storage address includes:
[0131] Obtaining a first read status of the status register of the first storage address and a second read status of the status register of the next storage address of the first storage address;
[0132] Based on the first read status and the second read status, determining the collection status of the first prompt word.
[0133] Optionally, when performing the RAG task, the host first checks whether the status register of the first storage address is to be read. If it is to be read, it reads the first prompt word and changes the status register to read. If the status register is read, the host first checks whether the status register of the next storage address is to be read. If the status register is end, it means that all the first prompt words for reasoning have been collected at this time, and the retrieval phase is started. See Table 1 for details.
[0134] As shown in Table 1
[0135]
[0136] In the embodiment of the present application, the host determines the acquisition result of the first prompt word by checking the status of the status register, without the need for complex addressing and data retrieval operations such as accessing the file system or database, and can obtain the acquisition status of the first prompt word at an extremely fast speed, improving the operation efficiency of the system.
[0137] In some alternative embodiments, vectorizing the first prompt word and performing database retrieval to generate a second prompt word includes:
[0138] Converting the first prompt word into a first vector;
[0139] Inputting the first vector into the database for retrieval to obtain a second vector associated with the first vector;
[0140] Integrating the first vector and the second vector to obtain a second prompt word.
[0141] Optionally, the host can read the previously stored first prompt word from the non-fixed memory pool of the double data rate memory. This is the starting step of the RAG task to obtain the input data for subsequent processing.
[0142] Convert the read first prompt word into vector form. In natural language processing, specific algorithms or models (such as word embedding models) are usually used to convert text into numerical vectors so that the computer can process and analyze this text information, preparing for subsequent retrieval in the vector database.
[0143] Use the vectorized first prompt word to perform retrieval in the vector database to find relevant information. The vector database can quickly find other vectors similar to the input vector, and these vectors may correspond to relevant text or knowledge, thereby obtaining supplementary information related to the first prompt word.
[0144] Integrate the retrieved relevant information with the first prompt word and reconstruct the context to generate a second prompt message.
[0145] In some alternative embodiments, obtaining the computing power of all step tasks performed by the second prompt word in the inference stage includes:
[0146] Obtaining the amount of computation and memory access corresponding to the step task;
[0147] Based on the ratio of the amount of computation to the memory access amount, obtain the computing power of the step task.
[0148] Optionally, some non-matrix operations (such as embedding and encoding, layer normalization, activation function calculation, and residual connection) exhibit low arithmetic intensity. In the embodiments of the present application, the computing power of embedding and encoding, layer normalization, activation function calculation, and residual connection will be directly recognized as low and allocated to the PNM memory for calculation. For other step tasks in the inference phase, when calculating the computing power, it is necessary to obtain the corresponding computational amount and memory access amount, and then based on the ratio of the computational amount to the memory access amount, the computing power of the step task is obtained.
[0149] It should be noted that the computational amount and memory access amount in the step task are obtained by fusing the batch size, sequence length, and hidden layer size in the step task.
[0150] For example: Computational amount of QKV mapping in Prefill = 3×2×B×L×S 2 ; Memory access amount of QKV mapping in Prefill = B×L×S + 3×S 2 . Where B is the batch size, L is the input sequence length, and S is the hidden layer size.
[0151] Therefore, the calculation of the computing power of the step task is as follows:
[0152] QKV mapping in Prefill: I = 3×2×B×L×S 2 / (B×L×S + 3×S 2 )
[0153] QK matrix calculation or SV matrix calculation in Prefill: I = 2×B×L 2 ×S / (2×B×L×S)
[0154] Output projection in Prefill: I = 2×B×L×S 2 / (B×L×S + S 2 )
[0155] Fully connected layer calculation in Prefill: I = 4×B×L×S 2 / (B×L×S + 4×S 2 )
[0156] QKV mapping in Decode: I = 3×2×B×L×S 2 / (B×S + 3×S 2 )
[0157] QK matrix calculation or SV matrix calculation in Decode: I = 2×B×L 2 ×S / (2×B×S)
[0158] Output projection in Decode: I = 2×B×L×S 2 / (1 + B×S + S 2 )
[0159] Calculation of the fully connected layer in Decode: I = 4×B×L×S 2 / (B×S + 4×S 2 )
[0160] In some optional embodiments, determining the assignment object of the step task based on computing power includes:
[0161] When the computing power is less than or equal to the preset value, the step task is assigned to the first assignment object, where the first assignment object is used to process tasks with an arithmetic intensity of the first intensity;
[0162] When the computing power is greater than the preset value, the step task is assigned to the second assignment object, where the second assignment object is used to process tasks with an arithmetic intensity of the second intensity.
[0163] Optionally, when the arithmetic intensity of the step task satisfies the following formula, the task will be assigned to the PNM memory closer to the host for calculation, that is, the PNM memory is used to process tasks with an arithmetic intensity of the first intensity (smaller intensity); if the formula is not satisfied, the task will be assigned to the GPU for calculation, that is, the GPU is used to process tasks with an arithmetic intensity of the second intensity (greater intensity).
[0164] Formula: I ≤ d×(1 / BW cxl −1 / BW pnm ) / (1 / TH pnm −1 / TH gpu )
[0165] Where d×(1 / BW cxl −1 / BW pnm ) / (1 / TH pnm −1 / TH gpu ) is the preset value, d is the size of the operand data type (in bytes), BW cxl and BW pnm respectively represent the bandwidth of CXL and the bandwidth of the PNM memory (in GB / s), TH pnm and TH gpu respectively represent the computing throughput of the PNM memory and the GPU for a given data type.
[0166] The embodiments of the present application can determine the allocation target (GPU or PNM memory) in real time according to the arithmetic intensity of the step tasks. This dynamic resource allocation mechanism ensures that tasks with low arithmetic intensity are offloaded to the PNM memory closer to the host to reduce memory read / write latency, while tasks with higher arithmetic intensity are processed by the GPU. By optimizing the task offloading strategy, the advantages of different computing resources are fully utilized, thereby improving the computing efficiency and reducing resource idleness.
[0167] In some alternative embodiments, the method further includes:
[0168] Obtain a data log, where the data log includes the first storage address of the first prompt word, the second storage address of the second prompt word, the third storage address of the intermediate parameters obtained in the inference stage, and the fourth storage address of the fixed parameters in the retrieval stage and the inference stage;
[0169] Obtain the adjustment operations performed on the data stored in the first storage address, the second storage address, the third storage address, and the fourth storage address to obtain the adjusted data and the adjusted first storage address, the adjusted second storage address, the adjusted third storage address, and the adjusted fourth storage address;
[0170] Update the data log based on the adjusted data and the adjusted first storage address, the adjusted second storage address, the adjusted third storage address, and the adjusted fourth storage address, and share the data using the updated data log.
[0171] Optionally, after enabling the inference application, the first storage address of the first prompt word, the second storage address of the second prompt word, the third storage address of the intermediate parameters obtained in the inference stage, and the fourth storage address of the fixed parameters in the retrieval stage and the inference stage that are stored in the non-fixed memory pool will remain in the fixed memory pool. These data are called the data log and are jointly maintained by four hosts and the high-speed interconnected memory pool.
[0172] When the host and the high-speed interconnected memory pool perform adjustment operations such as addition, deletion, and modification on the data stored in the non-fixed memory pool and the fixed memory pool, after the operation is completed, the modification method, data type, and the adjusted data and the adjusted first storage address, the adjusted second storage address, the adjusted third storage address, and the adjusted fourth storage address will be stored in the data log to update the data log. Other hosts can access the task outputs of other hosts by obtaining the data log, thereby achieving data sharing.
[0173] In addition, since the KV values generated by the GPU or PNM memory of the host responsible for the Prefill stage occupy too much memory, they can be directly transmitted to the non-fixed memory pool of the high-speed interconnection memory pool through the high-speed interconnection switch, and the data log is updated. When new KV values are added to the data log, the host responsible for the Decode stage can obtain the KV values and token positions through the data log to start the Decode stage task.
[0174] In the embodiment of the present application, by migrating the intermediate parameters (such as KV values) generated during the inference process to the non-fixed memory pool in real time and updating the data log according to the progress of the inference task, the real-time sharing and access of data are ensured, efficient cooperation between multiple hosts is realized, frequent data transmission between multiple hosts is avoided, thereby effectively reducing the data transmission delay in distributed computing, and thus accelerating the overall inference process.
[0175] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0176] In this embodiment, a language model retrieval and inference device is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0177] This embodiment provides a language model retrieval and inference device based on high-speed interconnection technology, as Figure 4 , this device includes:
[0178] A receiving module 401, configured to receive a first prompt word, and based on the first prompt word, start a retrieval stage, perform vectorization processing and database retrieval on the first prompt word, and generate a second prompt word;
[0179] An extraction module 402, configured to extract the second prompt word and perform inference on the second prompt word;
[0180] An obtaining module 403, configured to obtain the computing power of all step tasks executed by the second prompt word in the inference stage, and determine the assignment object of the step tasks based on the computing power, where the assignment object is used to perform task calculations on the step tasks.
[0181] In some alternative embodiments, the receiving module 401 is configured to, when receiving a first prompt word, obtain a first storage address corresponding to the first prompt word; determine the acquisition status of the first prompt word based on the first storage address; and start a retrieval phase when the acquisition status meets a preset condition.
[0182] In some alternative embodiments, the receiving module 401 is further configured to obtain a first reading status of the status register of the first storage address and a second reading status of the status register of the next storage address of the first storage address; and determine the acquisition status of the first prompt word based on the first reading status and the second reading status.
[0183] In some alternative embodiments, the receiving module 401 is further configured to determine that the acquisition status is that the first prompt word has been acquired and start a retrieval phase when the first reading status is read and the second reading status is read.
[0184] In some alternative embodiments, the receiving module 401 is further configured to convert the first prompt word into a first vector; input the first vector into a database for retrieval to obtain a second vector associated with the first vector; and integrate the first vector and the second vector to obtain a second prompt word.
[0185] In some alternative embodiments, the obtaining module 403 is configured to obtain the computation amount and the memory access amount corresponding to a step task; and obtain the computing power of the step task based on the ratio of the computation amount to the memory access amount.
[0186] In some alternative embodiments, the obtaining module 403 is further configured to obtain the batch size, the sequence length, and the hidden layer size in the step task; and obtain the computation amount and the memory access amount based on the batch size, the sequence length, and the hidden layer size.
[0187] In some alternative embodiments, the obtaining module 403 is further configured to, when the computing power is less than or equal to a preset value, allocate the step task to a first allocation object, where the first allocation object is used to process tasks with an arithmetic intensity of a first intensity; and when the computing power is greater than the preset value, allocate the step task to a second allocation object, where the second allocation object is used to process tasks with an arithmetic intensity of a second intensity.
[0188] In some alternative embodiments, the device further includes: obtaining a data log, where the data log contains a first storage address of a first prompt, a second storage address of a second prompt, a third storage address of intermediate parameters obtained in the inference phase, and a fourth storage address of fixed parameters in the retrieval phase and the inference phase; obtaining adjustment operations performed on the data stored in the first storage address, the second storage address, the third storage address, and the fourth storage address to obtain adjusted data and adjusted first, second, third, and fourth storage addresses; using the adjusted data and the adjusted first, second, third, and fourth storage addresses to update the data log, and sharing data using the updated data log.
[0189] For the description of the features in the corresponding embodiments of the language model retrieval and inference device based on high-speed interconnection technology, reference can be made to the relevant description in the corresponding embodiments of the language model retrieval and inference method based on high-speed interconnection technology, which will not be elaborated here one by one.
[0190] An embodiment of the present application also provides an electronic device, as Figure 5 shown, including a memory 10 and a processor 20. A computer program is stored in the memory 10, and the processor 20 is configured to run the computer program to execute the steps in any one of the above embodiments of the language model retrieval and inference method based on high-speed interconnection technology.
[0191] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above embodiments of the language model retrieval and inference method based on high-speed interconnection technology when running.
[0192] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.
[0193] An embodiment of the present application also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above embodiments of the language model retrieval and inference method based on high-speed interconnection technology.
[0194] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the language model retrieval and reasoning method based on high-speed interconnection technology are implemented.
[0195] Those skilled in the art can further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0196] The above has introduced in detail a language model retrieval and reasoning system and method based on high-speed interconnection technology provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A language model retrieval and inference system based on high-speed interconnection, characterized in that The system includes: a host, a high-speed interconnected switch, a high-speed interconnected memory pool, an image processor, and a memory migration module. The high-speed interconnected memory pool includes a near-memory computing memory, a double data rate memory, and the memory migration module. The high-speed interconnected switch includes an adaptive partitioning module. Among them, the host is connected to the high-speed interconnected memory pool and the image processor respectively through the high-speed interconnected switch. The host is configured to receive a first prompt word, start a retrieval phase based on the first prompt word, perform vectorization processing and database retrieval on the first prompt word, generate a second prompt word, and store the second prompt word in the double data rate memory through the high-speed interconnected switch. The double data rate memory includes a non-fixed memory pool and a fixed memory pool. The host is further configured to extract the second prompt word from the double data rate memory through the high-speed interconnected switch, perform reasoning on the second prompt word, and use the adaptive partitioning module in the connected high-speed interconnected switch to determine the task assignment object in the reasoning phase. The memory migration module is configured to store the first prompt word and the second prompt word in the non-fixed memory pool through the high-speed interconnected switch, migrate the vector value obtained by the assignment object after calculating the step tasks from the PNM memory or the local memory of the GPU to the non-fixed memory pool through the high-speed interconnected switch, and store the intermediate parameters and variable parameters generated in the reasoning phase in the non-fixed memory pool through the high-speed interconnected switch. Store the weights of the language model, the fixed parameters in the retrieval phase and the reasoning phase, and the storage addresses of all parameters in the non-fixed memory pool and the fixed memory pool in the fixed memory pool through the high-speed interconnected switch. The adaptive partitioning module is configured to obtain the computing power of all step tasks executed by the second prompt word in the reasoning phase, determine the assignment object of the step tasks based on the computing power, and assign the step tasks to the corresponding assignment object for task calculation. The assignment object is the near-memory computing memory or the image processor.
2. The system according to claim 1, wherein The double data rate memory includes a non-fixed memory pool and a fixed memory pool. The non-fixed memory pool is used to store the intermediate parameters generated in the reasoning phase, and the fixed memory pool is used to store the fixed parameters in the retrieval phase and the reasoning phase, as well as the storage addresses of the intermediate parameters and the fixed parameters. The high-speed interconnected memory pool further includes a memory migration module. The memory migration module is used to store the first prompt word and the second prompt word into the non-fixed memory pool through the high-speed interconnection switch, migrate the vector value obtained by the allocation object from the local memory to the non-fixed memory pool through the high-speed interconnection switch after calculating the step task, store the intermediate parameter into the non-fixed memory pool through the high-speed interconnection switch, and store the fixed parameter and the storage address into the fixed memory pool through the high-speed interconnection switch.
3. The system according to claim 1, wherein The high-speed interconnection memory pool further includes a high-speed interconnection protocol module, and the high-speed interconnection memory pool further includes a memory configuration module; The memory configuration module is used to initiate an access request to the non-fixed memory pool and / or the fixed memory pool in the double data rate memory by the host or the image processor through the high-speed interconnection protocol module; The memory configuration module is further used to initiate an access request to the non-fixed memory pool and / or the fixed memory pool in the double data rate memory by the near-memory computing memory through the high-speed interconnection protocol module.
4. A language model retrieval and inference method based on high-speed interconnection technology, characterized in that, The method is applied to a language model retrieval and inference system based on high-speed interconnection, and the method includes: Receiving a first prompt word, and based on the first prompt word, starting a retrieval phase, performing vectorization processing and database retrieval on the first prompt word, generating a second prompt word, and putting the second prompt word into the non-fixed memory pool of the double data rate memory in the language model retrieval and inference system based on high-speed interconnection; Adopting a continuous batch processing method, collecting the second prompt word from the non-fixed memory pool once every unit time interval, extracting the second prompt word, and performing inference on the second prompt word, and storing the generated inference result into the non-fixed memory pool; Obtaining the computing power of all step tasks executed in the inference phase of the second prompt word, and determining the allocation object of the step task based on the computing power, where the allocation object is used to perform task calculation on the step task.
5. The method according to claim 4, characterized in that The receiving the first prompt word and starting the retrieval phase based on the first prompt word includes: When receiving the first prompt word, obtaining the first storage address corresponding to the first prompt word; Determining the acquisition status of the first prompt word based on the first storage address; When the acquisition status meets a preset condition, starting the retrieval phase.
6. The method according to claim 5, wherein The determining the acquisition status of the first prompt word based on the first storage address includes: Obtaining the first reading status of the status register of the first storage address and the second reading status of the status register of the next storage address of the first storage address; Determining the acquisition status of the first prompt word based on the first reading status and the second reading status.
7. The method according to claim 6, wherein The starting the retrieval phase when the acquisition status meets a preset condition includes: When the first reading status is read and the second reading status is read, determining that the acquisition status is that the first prompt word has been acquired, and starting the retrieval phase.
8. The method according to claim 4, wherein Performing vectorization processing and database retrieval on the first prompt word to generate a second prompt word includes: Converting the first prompt word into a first vector; Inputting the first vector into the database for retrieval to obtain a second vector associated with the first vector; Integrating the first vector and the second vector to obtain the second prompt word.
9. The method according to claim 4, wherein Obtaining the computing power for all step tasks executed by the second prompt word during the inference stage, including: Obtaining the computational amount and memory access amount corresponding to the step task; Based on the ratio of the computational amount to the memory access amount, obtaining the computing power of the step task.
10. The method according to claim 9, wherein Obtaining the computational amount and memory access amount corresponding to the step task includes: Obtaining the batch size, sequence length, and hidden layer size in the step task; Based on the batch size, the sequence length, and the hidden layer size, obtaining the computational amount and the memory access amount.
11. The method according to claim 4, characterized in that Determining the allocation object of the step task based on the computing power includes: In the case where the computing power is less than or equal to a preset value, allocating the step task to a first allocation object, where the first allocation object is used to process tasks with an arithmetic intensity of a first intensity; In the case where the computing power is greater than the preset value, allocating the step task to a second allocation object, where the second allocation object is used to process tasks with an arithmetic intensity of a second intensity.
12. The method according to claim 4, wherein The method further includes: Obtaining a data log, where the data log contains a first storage address of the first prompt word, a second storage address of the second prompt word, a third storage address of intermediate parameters obtained during the inference stage, and a fourth storage address of fixed parameters during the retrieval stage and the inference stage; Obtaining adjustment operations performed on the data stored in the first storage address, the second storage address, the third storage address, and the fourth storage address to obtain adjusted data and adjusted first, second, third, and fourth storage addresses; Based on the adjusted data and the adjusted first, second, third, and fourth storage addresses, updating the data log and sharing data using the updated data log.
13. An electronic device, characterized in that, Including: A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the language model retrieval and inference method based on high-speed interconnection technology according to any one of claims 4 to 12.
14. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores computer instructions for causing a computer to execute the language model retrieval and inference method based on high-speed interconnection technology according to any one of claims 4 to 12.
15. A computer program product, characterized in that, Including computer instructions for causing a computer to execute the language model retrieval and inference method based on high-speed interconnection technology according to any one of claims 4 to 12.
Citation Information
Patent Citations
Rapid reasoning method, device and system for large language model of smart phone
CN118446321A
Data caching method, system, product, equipment and storage medium
CN119027300A