A computing architecture and method

By introducing a shared memory system into the computing architecture, managing and scheduling computing resources, the delay problem caused by the frequent flow of data between storage media is solved, and more efficient computing performance and user experience is achieved.

CN119356881BActive Publication Date: 2025-05-23STORAGEX TECH INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411910452.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-23
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

In the existing computing architecture, the number of data flows between storage media leads to a large delay in data transmission, affecting system performance and user experience.

Method used

Using a computing architecture, including a computing management system, a shared memory system and a computing cluster, the shared memory system manages and schedules computing resources to reduce the number of data flows between storage media.

Benefits of technology

It reduces data transmission delay, improves system performance and user experience, saves network infrastructure construction costs, and supports efficient integration of heterogeneous computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119356881B_ABST
    Figure CN119356881B_ABST
Patent Text Reader

Abstract

The present application relates to a computing architecture and method, which is applied in the field of cluster computing, including a computing management system for acquiring and analyzing tasks and managing and scheduling computing resources required for executing tasks; a shared memory system is connected in communication with the computing management system, which is used to manage the public memory required for executing tasks and the data generated during the execution of tasks; a computing cluster is connected in communication with the shared memory system, which is used to receive tasks, calculate tasks and output calculation results; the computing management system includes a management node, which is used to control the computing management system to directly access the shared memory system; the computing cluster includes computing nodes, which are used to control the computing cluster to directly access the shared memory system. The technical effects of the present application are: reducing access delay, simplifying the system, reducing data flow links, reducing costs, and improving computing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cluster computing technology, and in particular to a computing architecture and method. Background Art

[0002] As AI computing migrates to large models, large models are different from traditional AI models. The large scale of large models, the large KV cache space required, and the long computing cycle bring many challenges to the computing architecture design of large models.

[0003] In the existing computing architecture, a computing cluster retrieves a model from a database and loads it into a computing device for execution. Usually, the model in the storage device is first called, and then sent to the computing device for model loading. The data flow diagram of the model is as follows: Figure 1 shown.

[0004] In the existing model computing architecture, the storage end is a cluster device. The number of times data flows between storage media will increase, resulting in increased data transmission delay, which may cause the entire system to be stuck, resulting in a poor user experience. Summary of the invention

[0005] In order to help solve the problem of large data transmission delay caused by the frequent transfer of data between storage media in existing computing architectures, the present application provides a computing architecture and method.

[0006] In a first aspect, the present application provides a computing architecture, which adopts the following technical solution: the computing architecture includes a computing management system, a shared memory system and a computing cluster;

[0007] The computing management system is used to acquire and analyze tasks and manage and schedule computing resources required to execute the tasks;

[0008] The shared memory system is in communication with the computing management system and is used to manage the public memory required for executing the task and the data generated during the execution of the task;

[0009] The computing cluster is communicatively connected with the shared memory system, and is used to receive the task, calculate the task and output the calculation result;

[0010] The computing management system includes a management node, and the management node is used to control the computing management system to directly access the shared memory system; the computing cluster includes computing nodes, and the computing nodes are used to control the computing cluster to directly access the shared memory system.

[0011] In a specific implementation scheme, the management node includes a CPU; the CPU is used to manage data calculations in the computing management system.

[0012] In a specific implementation scheme, the computing node includes an accelerator, and the accelerator is used to accelerate the computing of the task; wherein data is directly transmitted between the accelerator and the shared memory system.

[0013] In a specific possible implementation scheme, the shared memory system includes a storage subsystem, a data routing unit, and a shared memory pool;

[0014] The storage subsystem is used to store the model and model parameters used to perform the task;

[0015] The data routing unit is in communication connection with the storage subsystem, and is used to provide a communication network between the shared memory system and the computing management system, and between the shared memory system and the computing cluster;

[0016] The shared memory pool is in communication connection with the data routing unit and is used for storing and managing the data required during the execution of the task and the data generated during the execution of the task.

[0017] In a specific possible implementation scheme, the shared memory pool includes a shared memory access management module and a memory medium module;

[0018] The shared memory access management module is used to address the memory medium module;

[0019] The memory medium module is in communication connection with the shared memory access management module, and is used for storing the data required during the task execution and the data generated during the task execution.

[0020] In a specific possible implementation scheme, the shared memory access management unit includes an access port and an interleaving management submodule;

[0021] The access port is used to provide a port for the shared memory system to interact with the computing management system and computing cluster for data;

[0022] The interleaving management submodule is used to address the memory medium module using a partial interleaving mode to generate a shared memory unified address;

[0023] The calculation method of the shared memory unified address includes:

[0024] N*C port + (A / x)*k*x + M*x + (A%x);

[0025] C port = C memory * k;

[0026] Where N represents the Nth access port, C port Indicates the memory capacity corresponding to each access port, A indicates each memory medium address, C memory represents the capacity of each memory medium, k represents the number of memory media, x represents the size of the alternating arrangement of memory media, and M represents the shared memory unified address corresponding to the Mth memory medium in the access port.

[0027] In a second aspect, the present application provides a computing method, which adopts the following technical solution: the method is applied to a computing architecture, the computing architecture includes a shared memory system, the shared memory system includes a shared memory pool, and the method includes:

[0028] Obtaining tasks to be processed and analyzing computing resources required for the tasks to be processed;

[0029] Querying free computing resources and allocating the free computing resources to the tasks to be processed;

[0030] Reading the address of the computing model required by the task to be processed from the shared memory system, and importing the computing model into the shared memory pool;

[0031] The computing model is used to call the computing resources to execute the task to be processed according to the type of the task to be processed, and the computing result is output.

[0032] In a specific possible implementation scheme, the shared memory pool includes a model management table and a model parameter buffer area, the model parameter buffer area is used to cache the parameters of the used computing models, and the model management table is used to record the used computing models;

[0033] The importing the computing model into the shared memory pool comprises:

[0034] Searching the model management table and determining whether the model management table contains the calculation model;

[0035] If the model management table includes the calculation model, directly reading the parameters corresponding to the calculation model from the model parameter buffer area, and importing the calculation model into the shared memory pool;

[0036] If the model management table does not include the calculation model, the required calculation model is requested from the shared memory system, the requested calculation model is initialized and the calculation model is imported into the shared memory pool.

[0037] In a specific embodiment, the method further comprises:

[0038] When importing the computing model into the shared memory pool, monitoring the capacity of the model parameter buffer area;

[0039] If the capacity of the model parameter cache exceeds a preset threshold, determining whether there is a computing model whose usage cycle has been exhausted in the model parameter cache;

[0040] If there is a calculation model whose usage cycle is exhausted, determining whether the access frequency of the calculation model whose usage cycle is exhausted is greater than a preset value;

[0041] If the access frequency of the computing model whose usage cycle is exhausted is not greater than the preset value, the computing model whose usage cycle is exhausted and whose access frequency is not greater than the preset value will be released.

[0042] In a specific possible implementation scheme, the computing architecture includes a first accelerator and a second accelerator, the task to be processed includes a collaborative task and a migration task, and if the task to be processed is a collaborative task, then according to the type of the task to be processed, calling the computing resource to execute the task to be processed, and outputting the computing result includes:

[0043] Determine whether the to-be-processed tasks include preset historical pre-filled tasks;

[0044] If the to-be-processed task does not include a preset historical pre-filled task, controlling the second accelerator to directly process the to-be-processed task and generate a calculation result;

[0045] If the to-be-processed task includes a preset historical pre-filled task, then importing the historical KV cache data corresponding to the historical pre-filled task;

[0046] Controlling the first accelerator to calculate the historical KV cache data and generate calculated KV cache data;

[0047] caching the calculated KV cache data into the shared memory pool, and importing the calculated KV cache data into the memory of the second accelerator;

[0048] The second accelerator is controlled to read the calculated KV cache data from the memory of the second accelerator, and the tasks to be processed other than the historical pre-filled tasks are calculated according to the calculated KV cache data to generate results of the tasks to be processed.

[0049] In a specific implementation scheme, the computing architecture includes a third accelerator and a fourth accelerator, the task to be processed includes a collaborative task and a migration task, and if the task to be processed is a migration task, then according to the type of the task to be processed, calling the computing resource to execute the task to be processed, and outputting the computing result includes:

[0050] Controlling the third accelerator to calculate the task to be processed, generating pre-migration KV cache data, and writing the pre-migration KV cache data into the memory of the third accelerator;

[0051] In response to the data migration instruction, import the pre-migration KV cache data in the memory of the third accelerator into the shared memory pool, and at the same time, control the third accelerator to continue calculating the pending task to generate post-migration KV cache data, and write the post-migration KV cache data into the shared memory pool;

[0052] Importing the pre-migration KV cache data and the post-migration KV cache data into the memory of the fourth accelerator;

[0053] Based on the pre-migration KV cache data and the post-migration KV cache data, the fourth accelerator is controlled to continue executing the task to be processed and generate a result of the task to be processed.

[0054] In summary, this application has the following beneficial technical effects:

[0055] 1. Short response delay and better user experience;

[0056] 2. Flexible deployment, wide support range, and easy to use;

[0057] 3. The system is simplified, data transfer links are reduced, and the cost of network infrastructure construction is greatly saved;

[0058] 4. Efficiently integrate heterogeneous computing resources, high resource utilization, and reduce construction costs;

[0059] 5. Flexible deployment and easy to expand;

[0060] 6. The architecture is highly adaptable and can support access to different computing devices. It is also easy to upgrade and can be quickly deployed when computing devices are replaced. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is a schematic diagram used to illustrate the data flow of the existing model computing architecture;

[0062] Figure 2 This is an overall framework diagram of the computing architecture in the embodiment of the present application;

[0063] Figure 3 is a schematic diagram of a management node in an embodiment of the present application;

[0064] Figure 4 is a schematic diagram of a computing node in an embodiment of the present application;

[0065] Figure 5 is a schematic diagram of a shared memory system in an embodiment of the present application;

[0066] Figure 6 is a schematic diagram of shared memory access management in an embodiment of the present application;

[0067] Figure 7 is a schematic diagram of a unified addressing method in an embodiment of the present application;

[0068] Figure 8 is a flow chart of the calculation method in the embodiment of the present application;

[0069] Fig. 9 is a schematic diagram of model parameter management in an embodiment of the present application;

[0070] Fig.10 is a schematic diagram of a collaborative task processing process in an embodiment of the present application;

[0071] Fig.11 is a schematic diagram of a migration task processing process in an embodiment of the present application;

[0072] Fig.12 It is a schematic diagram of the complete large model calculation process in the embodiment of the present application. DETAILED DESCRIPTION

[0073] The following combination Figure 1-Figure 12 This application is described in further detail.

[0074] The embodiment of the present application discloses a computing architecture, which can simplify the computing architecture, reduce data flow links, and greatly save the cost of network infrastructure construction. With the migration of AI computing to large models, large models are different from traditional AI models. The large scale of large models, the large KV cache space required, and the long computing cycle bring many challenges to the computing architecture design of large models.

[0075] First, in the existing computing architecture, a computing cluster retrieves a model from the database and loads it into the computing device for execution. Usually, the model in the storage device is first called, and then sent to the computing device for model loading. The data flow diagram of the model is as follows: Figure 1 In the existing model computing architecture, the storage end is a cluster device. The number of data transfers between storage media will increase, which will increase the data transmission delay, which may cause the entire system to be stuck, resulting in a poor user experience.

[0076] Secondly, the computing clusters in data centers provide different computing tasks for different users, and a storage system is needed for the large models that calculate these tasks. Currently, a large number of end users use models that are retrained based on different pre-trained models combined with industry or institution-specific data. There are many models and a large number of parameters for each model. In extreme cases, each user may have multiple models. At the same time, in a cluster containing heterogeneous computing devices, deployment on different computing devices has different requirements for model parameter formats and segmentation methods. For cost and feasibility considerations, it is difficult to store models locally on each computing device, and a dedicated storage service subsystem is required. In clusters of different sizes, the storage service device that constitutes the subsystem may be a server or a service device cluster. At present, traditional methods usually pre-build parameter formats suitable for each type of computing device for each model, but this will dramatically increase the storage space requirements. In addition, each model is stored and compressed in a unified form, and is decompressed and converted to a data format in real time when called, and then imported. Since the model parameter scale itself is in the billions, some models may reach hundreds of billions, which greatly increases the initialization delay.

[0077] In addition, the large model calculation generation stage is mostly continuous calculation, and a large amount of KV cache data needs to be cached during the calculation process. In the case of insufficient computing device resources, data and task migration is required. Taking the medium-scale parameter model llama3.1-70B as an example, the pre-trained basic model supports 128K context length, the K and V cache formats are BF16 (2 bytes), the model structure is 80 layers, each layer has 8 groups of K and V heads, and each head dimension is 128. The capacity required for the full space of KV cache data is 128*2 bytes*2*8*80*128K, about 40GB; in the conversation, the conversation length is uncertain, and the initial division will not reserve 40GB space to cache KV cache data. Under the support of multi-task computing, each task will initially allocate a smaller storage space; but as the conversation continues, the cache becomes longer, and some tasks need to continuously allocate new memory space to store KV cache data, which will lead to insufficient storage space of the computing device, and then it is necessary to migrate the calculated KV cache to other devices for continued computing. At the same time, due to the continuity of the conversation, it is impossible to predict the duration of each conversation task. The task load between different computing devices is unbalanced, and task migration is also needed to balance the load and improve task responsiveness. The migration of large amounts of data between the memories of different devices requires a high bandwidth, and at the same time, there are many storage media transfer links. The existing computing architecture is difficult to support the efficient migration of large amounts of data and large tasks between different devices.

[0078] Finally, when different types of devices perform collaborative computing, different hardware resources often cannot directly exchange data due to different architectures, and often need to be transferred through the CPU, resulting in low transmission efficiency, limited computing performance, and the inability of hardware to perform at its best, which is reflected in hardware resource redundancy and directly manifested as an increase in total cost of ownership (TCO). In order to help improve the efficiency of large model processing tasks and reduce response delays, this application provides a computing architecture.

[0079] Reference Figure 2 ,The computing architecture includes a computing management system, a shared memory system and a computing cluster, where a computing cluster can contain several ,clusters.

[0080] The computing management system is used to obtain and analyze tasks and manage and schedule the computing resources required to execute tasks; wherein, the computing management system includes a user service subsystem and a resource scheduling subsystem. The user service subsystem is used to obtain tasks, analyze the resources required for tasks and apply for resources from the resource scheduling subsystem. The resource scheduling subsystem queries the available resources according to the request, allocates resources and locks them, and returns the relevant description of resource allocation to the user service subsystem. In addition, the resource scheduling subsystem can also manage and schedule computing resources in the entire computing architecture. The user service subsystem and the resource scheduling subsystem are connected in communication. The shared memory system is connected in communication with the computing management system to manage the public memory required for executing tasks and the data generated during the execution of tasks; the computing cluster is connected in communication with the shared memory system to receive tasks, calculate tasks and output calculation results.

[0081] The computing management system includes a management node, which is used to control the computing management system to directly access the shared memory system; the management node is usually represented as a server in the computing management system, see Figure 2 Each server in the user service subsystem and resource scheduling subsystem can be regarded as a management node; in addition, there may also be management nodes in the computing cluster, for example, refer to Figure 2 The acceleration device includes a management node composed of a computing management / CPU. In a computing cluster, the management node may be in the form of a server or a minimum system.

[0082] Reference Figure 3, the management node includes a CPU, and the CPU is mainly used to manage data calculations in the computing management system. Specifically, the CPU can complete task response, resource management and scheduling, computing management, data management and other operations according to the different functions of the management node, while configuring only the local memory with the minimum capacity for system operation or even not configuring external memory and relying only on on-chip storage. It should be noted that different management nodes all include a CPU, but the functions may be different. In the embodiments of the present application, less or even no memory needs to be configured locally; secondly, the CPU does not need to touch a large amount of data to be calculated when performing management, thereby improving efficiency.

[0083] In addition, the management node may also include a hard disk design. The hard disk is used to store data required during task execution and data generated during task execution. Data can be directly transmitted between the hard disk and the shared memory system. Usually, the CPU of the management node controls the hard disk to directly access the shared memory system. Data can interact directly between the hard disk and the shared memory system without passing through the CPU. When multiple management nodes work together or management nodes of different systems communicate, the number of data flows can be reduced and the increase in memory capacity demand and data synchronization overhead caused by maintaining the same data or forms can be reduced. In addition, if the management node requires a larger cache, or the calculation process requires high-speed and frequent memory access, the memory can also be configured separately. For example, refer to Figure 3 In the figure, the memory in the dotted box indicates that it can be not set. In the case of requiring a larger cache, the user can set a dedicated memory. It should be noted that in actual applications, the hard disk is not a necessary part, and the user can design whether the hard disk is needed according to the actual memory requirements.

[0084] The computing cluster includes computing nodes, which are used to control the computing cluster to directly access the shared memory system. Figure 4 , the computing node includes an accelerator, which is used to accelerate the calculation of the task. The accelerator contains a dedicated high-performance memory inside, and data is directly transmitted between the accelerator and the shared memory system. It should be noted that the computing node usually includes one or more computing accelerators, each of which can communicate directly with the shared memory system; in addition, each accelerator in the figure can be composed of only one accelerator, or it can be an accelerator device composed of a group of accelerators. During the calculation process, each accelerator often requires a large cache and has high requirements for cache access speed. Therefore, the corresponding high-performance memory (such as HBM, GDDR, etc.) has been configured in the design of the accelerator; at the same time, due to the requirements for data block integrity, data preloading, etc., the accelerator uses a configured dedicated high-performance memory. Therefore, the data path of the accelerator is different from that of the management node, and data interaction with the shared memory system is carried out through the configured dedicated high-performance memory.

[0085] Through the solution of this application, the computing management system and the computing cluster, the computing devices / computing accelerators within the computing cluster, or the computing clusters exchange business data through the shared memory system to achieve low-latency transmission. The computing accelerators at the bottom of the computing cluster can directly access the shared memory, thereby shortening the data interaction path and minimizing the number of data transfers between storage media, thereby improving computing efficiency.

[0086] Reference Figure 5 The shared memory system includes a storage subsystem, a data routing unit and a shared memory pool. The storage subsystem is used to store the model and the parameters of the model used to execute the task. The data routing unit is connected to the storage subsystem in communication, and is used to provide a communication network between the shared memory system and the computing management system, and between the shared memory system and the computing cluster. In addition, the data routing can provide a many-to-many access interface, forwarding the access command and data to the correct port according to the address, and the number of ports facing the shared memory pool and the external port of the data routing does not have to be consistent. The shared memory pool is connected to the data routing unit in communication, and is used to store and manage the data required during the task execution and the data generated during the task execution.

[0087] The shared memory pool includes a shared memory access management module and a memory medium module; the shared memory access management module is used to address the memory medium module, refer to Figure 6 , shared memory access management can provide a memory controller that controls memory media access in parallel and realizes the unified addressing of shared memory; the memory medium module is connected to the shared memory access management module in communication, and is used to store the data required during the task execution and the data generated during the task execution. The shared memory access management unit includes an access port and an interleaving management submodule. The access port is used to provide a port for the shared memory system to interact with the computing management system and the computing cluster for data. The interleaving management submodule is used to address the memory medium module using a partial interleaving mode to generate a shared memory unified address; in the addressing and use of the memory medium, the full interleaving mode is not used, but the partial interleaving mode is used. The number of memory media for each partial interleaving addressing is determined by the bandwidth of a single access port of the shared memory system, and both sides should match.

[0088] The calculation method of the shared memory unified address can be understood as the process of mapping the memory medium address to the shared memory address. The controllers and memory media correspond one to one, and the number of controllers is the same as the number of memory media. Assume that the capacity of each memory medium is C memory , each interleaving management contains k controllers and s ports, then the total capacity C total =C memory * k* s, the memory capacity corresponding to each port can be expressed as C port = C memory* k. For addressing, the total space address is 0~C total -1, the address space corresponding to the Nth (0~s-1) port is N*C port ~(N+1)*C port -1; within the port, if the width of a single memory medium is x bytes, the corresponding addresses of the memory media in the system are arranged alternately according to x B, then the address A of the Mth (0~k-1)th memory medium in the port is mapped to the shared memory address can be expressed as:

[0089] N*C port + (A / x)*k*x + M*x + (A%x);

[0090] Where N represents the Nth access port, C port Indicates the memory capacity corresponding to each access port, A indicates each memory medium address, C memory represents the capacity of each memory medium, k represents the number of memory media, x represents the size of the alternating arrangement of memory media, and M represents the shared memory unified address corresponding to the Mth memory medium in the access port.

[0091] Assuming that the bit width of a single memory medium is 64 bits (8 bytes), and x is 8, the corresponding addresses of the memory media in the system are arranged alternately in 8 bytes. The shared memory address mapped to the address A of the Mth (0~k-1)th memory medium in the port can be expressed as:

[0092] N*C port + (A / 8)*k*8 + M*8 + (A%8).

[0093] It should be noted that the shared memory access interface can be CXL or RDMA. Using the CXL interface can achieve lower latency, and using the RDMA interface can make the entire system larger and have a higher performance ceiling. The shared memory access management part can use a group of ports as the smallest unit, which can be implemented by one or more SoCs, ASICs, one or more FPGAs, and one or more CPUs.

[0094] In the present application scheme, by integrating the storage subsystem into the shared memory system, it is convenient to unify the data access interface and address, reduce the complexity, and also achieve the reduction of the delay of high-frequency data access.

[0095] It should be noted that in the embodiments of the present application, in addition to accessing shared memory, some processors such as accelerators or servers in the computing architecture also have dedicated memory inside. For example, the accelerator contains dedicated high-performance memory. In order to distinguish shared memory from dedicated memory and reduce the address interaction overhead between different processors, a unified addressing method is designed in the embodiments of the present application to support memory access for each processor. The memory address uses 64 bits, and the compilation rules refer to Figure 7 As shown, the bit arrangement is {0, 1, ..., 63}. A flag bit of 0 indicates access to local memory, and a flag bit of 1 indicates access to shared memory. Each processor can be configured with up to 2 local 63 The addressing space is about 8192PB. Similarly, the maximum shared memory space is 8192PB. The address space can meet the memory capacity requirements of any current system.

[0096] Based on the above computing architecture, an embodiment of the present application also discloses a computing method.

[0097] Reference Figure 8 , the method comprises the following steps:

[0098] S10, obtaining tasks to be processed and analyzing computing resources required for the tasks to be processed.

[0099] Specifically, the computing architecture obtains the tasks to be processed through the user service subsystem under the computing management system, analyzes the computing resources required for the tasks through the user service subsystem, and applies for computing resources from the resource scheduling subsystem under the computing management system.

[0100] S20, querying the available computing resources, and allocating the available computing resources to the tasks to be processed.

[0101] Specifically, the resource scheduling subsystem under the computing management system receives resource applications, queries available computing resources, allocates corresponding computing resources to pending tasks and locks computing resources; the resource scheduling subsystem returns the corresponding description of the allocated resources to the user service subsystem.

[0102] S30, reading the address of the calculation model required by the task to be processed from the shared memory system, and importing the calculation model into the shared memory pool.

[0103] Specifically, the large model required to execute the task is stored in the shared memory system. The computing architecture reads the address of the computing model required to execute the task to be processed from the shared memory system, obtains the computing model from the corresponding address, and imports the computing model into the shared memory pool so that the computing management system and the computing cluster can interact with the shared memory system for data to complete the task.

[0104] S40, using the computing model, according to the type of the task to be processed, calling computing resources to execute the task to be processed, and outputting the computing result.

[0105] Specifically, in the process of executing a computing task, data will be exchanged between computing devices to complete the task and generate results. After importing the computing model, the allocated computing resources will be called according to different task types, and data will be exchanged between computing devices to complete the task; among them, the task types mainly include collaborative tasks and migration tasks. Collaborative tasks can be understood as completing a task together between different computing devices or different computing clusters; migration tasks can be understood as migrating tasks from one computing device to another computing device for continued execution. This mainly occurs in scenarios such as insufficient dedicated memory resources of the current accelerator and imbalanced computing load that has affected computing performance. At this time, it is necessary to migrate to another accelerator to solve the problems of insufficient memory and imbalanced load.

[0106] In the present application scheme, the computing method based on the computing architecture can efficiently integrate computing resources, improve resource utilization, reduce response delay, and improve computing efficiency; in addition, during the entire computing process, the data flow process is reduced, which greatly saves the cost of network infrastructure construction while improving computing efficiency, thereby improving user experience.

[0107] In one embodiment, the shared memory pool includes a model management table and a model parameter buffer area, the model parameter buffer area is used to cache the parameters of the used computing model, and the model management table is used to record the used computing model; the method of importing the computing model into the shared memory pool can be specifically performed as follows:

[0108] First, search the model management table and determine whether the model management table contains the calculation model; if the model management table contains the calculation model, directly read the parameters corresponding to the calculation model from the model parameter buffer, and import the calculation model into the shared memory pool; if the model management table does not contain the calculation model, request the required calculation model from the shared memory system, initialize the requested calculation model and import the calculation model into the shared memory pool.

[0109] Reference Fig. 9 In the shared memory pool, the table cache area and the model parameter cache area are divided according to the system configuration. The table cache area is used to store the quick lookup table (such as the hash table). The quick lookup table is also the model management table, which records the used models. Check from the quick lookup table whether the required model is in the cache area. If it is in the cache area, it can be directly obtained and imported into the shared memory pool. If the required model is not in the cache area, it is necessary to apply for the model from the storage subsystem in the shared memory system and import it into the shared memory pool.

[0110] Considering that the capacity of the divided model parameter cache area has an upper limit, when the capacity exceeds the upper limit, it is difficult to store new model parameters in the model cache area. Therefore, the calculation method can also perform the following steps:

[0111] When importing a computing model into a shared memory pool, the capacity of the model parameter cache is monitored; if the capacity of the model parameter cache exceeds a preset threshold, that is, when the capacity is close to exhaustion, some models need to be released or recovered. Specifically, first determine whether there is a computing model whose usage cycle has been exhausted in the model parameter cache; if there is a computing model whose usage cycle has been exhausted, determine whether the access frequency of the computing model whose usage cycle has been exhausted is greater than a preset value; if the access frequency of the computing model whose usage cycle has been exhausted is not greater than the preset value, then release the computing model whose usage cycle has been exhausted and whose access frequency is not greater than the preset value.

[0112] When applying to load a model at the beginning, the use cycle of the model data will be set first. The use cycle can also be expressed as the life cycle, which can be understood as the minimum existence time of the model data. The minimum existence time of the model has been set when the model is first applied. When the capacity is close to exhaustion, that is, when the capacity is greater than the preset threshold, the release or recycling target is selected from the model data that has been exhausted in the life cycle. If there is no life cycle exhausted model at present, the system will be issued a reminder of insufficient cache capacity. The system can expand the cache area or wait for a certain delay before trying to recycle memory; if there is a model data that has been exhausted in the life cycle, the access frequency is set to a threshold. Model data greater than the set threshold is not allowed to be released or recycled, and models with lower access frequencies are recycled first.

[0113] In addition, according to the recycling rules, life cycle and heat management are required. Heat refers to the frequency of model access. The higher the model access frequency, the higher the heat. The less the model access frequency, the lower the heat. For the life cycle, when applying for model loading, set the minimum existence time of the hot data. The background timer will update the life cycle regularly and subtract the delay difference between the two updates. For heat management, it can be updated on a periodic basis. The number of times it is accessed within the period is added to the heat value at the end of the period. If it is not accessed within the period, a fixed value will be subtracted at the end of the heat. The fixed value can be set based on experience according to the actual needs of different application scenarios.

[0114] It should be noted that the embodiments of the present application support computing on heterogeneous devices, and the hard disk of the storage subsystem stores model parameters in a unified format. Therefore, the target operating hardware platform must be specified during loading, and multiple platforms can be specified. During loading, they will be converted into the specified format and stored in the memory, which can reduce the delay caused by re-reading and converting the format for storage.

[0115] In the solution of the present application, the model parameters are managed by setting the model cache area and the table cache area, thereby reducing the access delay of the model parameters and improving the calculation efficiency.

[0116] In one embodiment, referring to Fig.10 The computing architecture includes a first accelerator and a second accelerator. The tasks to be processed include collaborative tasks and migration tasks. If the tasks to be processed are collaborative tasks, then according to the types of the tasks to be processed, the computing resources are called to execute the tasks to be processed, and the method of outputting the computing results can be specifically implemented as follows:

[0117] First, determine whether the pending task contains the preset historical prefill task, which is also called the historical prefill task. In many computing scenarios of large models, the same prefill task is often used, such as customer service consultation, legal interpretation, article analysis, etc. By retaining the common prefill task, it can be directly imported during the next calculation, which can effectively reduce the overall computing delay. If the pending task does not contain the preset historical prefill task, the second accelerator is controlled to directly process the pending task and generate the calculation result.

[0118] If the task to be processed includes a preset historical pre-filled task, then the historical KV cache data corresponding to the historical pre-filled task is imported; the first accelerator is controlled to calculate the historical KV cache data and generate the calculated KV cache data; the calculated KV cache data is cached to the shared memory pool, and the calculated KV cache data is imported into the memory of the second accelerator. Afterwards, the second accelerator is controlled to read the calculated KV cache data from the memory of the second accelerator, and the tasks other than the historical pre-filled tasks in the task to be processed are calculated according to the calculated KV cache data, and the results of the task to be processed are generated. In an embodiment of the present application, a collaborative task refers to a task of collaborative computing of heterogeneous devices, that is, the same task needs to be completed by computing devices in different computing clusters in collaborative computing to generate the results corresponding to the task.

[0119] The second accelerator may be a computing resource pool. When the prefill task starts, no specific second accelerator is assigned to perform the result generation task. When the prefill calculation is finished or near the end, the computing resources are assigned to the second accelerator that executes the task, taking into account the goals of computing efficiency, load balancing, and energy saving. In addition, the first accelerator and the second accelerator can be a single accelerator or a group of accelerators on hardware. Considering that some models are large and require multiple accelerators to calculate together, data transmission between similar accelerators is generally through dedicated point-to-point connections, not shared memory.

[0120] It should be noted that the KV cache data management of the embodiment of the present application supports fast retrieval and data background management. Fast retrieval is achieved through two tables in shared memory, the KV cache data record table and the KV cache hot data list. The KV cache data record table mainly records whether there is historical calculation of the current prefill, supports data storage with different structures across applications, and performs hash operations on the embedding sequence of the prefill in combination with the model description, records the hash value and uses it as a query target, thereby saving storage space and avoiding as much as possible the inability to unify keyword retrieval due to differences in different prefill sources, data structures, etc.; the KV cache hot data list mainly marks the KV cache data that has been imported into the shared memory pool for fast import during query. In addition, not all prefills will be recorded historically. Only prefills that appear at the beginning of the sequence and are highly repeatable will be recorded. These prefills are usually some public issues or public reference materials.

[0121] In the present application, the historical prefill task is run on the first accelerator, which is usually a computationally intensive accelerator, which can make full use of the matrix multiplication computing resources of the computationally intensive accelerator and provide higher prefill performance; the result generation task is run on the second accelerator, which is usually a bandwidth-first computing accelerator, which makes full use of the dedicated memory bandwidth to minimize the delay in the generation phase. The combination of two heterogeneous accelerators can efficiently complete tasks, improve computing efficiency, and enhance user experience.

[0122] In one embodiment, referring to Fig.11 The computing architecture includes a third accelerator and a fourth accelerator. The tasks to be processed include collaborative tasks and migration tasks. If the tasks to be processed are migration tasks, then according to the types of the tasks to be processed, the computing resources are called to execute the tasks to be processed and the computing results are output. The specific execution is as follows:

[0123] First, control the third accelerator to calculate the pending tasks, generate pre-migration KV cache data, and write the pre-migration KV cache data into the memory of the third accelerator. In response to the data migration instruction, import the pre-migration KV cache data in the memory of the third accelerator into the shared memory pool. At the same time, control the third accelerator to continue calculating the pending tasks to generate post-migration KV cache data, and write the post-migration KV cache data into the shared memory pool; import the pre-migration KV cache data and post-migration KV cache data into the memory of the fourth accelerator; based on the pre-migration KV cache data and post-migration KV cache data, control the fourth accelerator to continue to execute the pending tasks and generate the results of the pending tasks, and the newly generated KV cache data is only stored in the memory of the fourth accelerator. Among them, the third accelerator and the fourth accelerator in the embodiment of the present application are accelerators of the same type, or are located in the same computing cluster.

[0124] In the present application scheme, when the current accelerator dedicated memory is insufficient, migration is performed to obtain more memory resources to support the completion of the task. The resource scheduling system can allocate corresponding computing resources according to the actual needs of the architecture calculation to avoid as much as possible the problem of insufficient bandwidth that may occur during the migration process.

[0125] Reference Fig.12 , which is a complete flowchart of large model calculation for the embodiment of the present application. First, the task is obtained and the resources required for the task are calculated through the user service subsystem, and the computing resources are applied to the resource scheduling subsystem. The resource scheduling subsystem queries the available resources and allocates them to the tasks to be processed, and returns the resource allocation description to the user service subsystem. The user service subsystem reads the address of the large model required to process the task from the memory sharing system, obtains the calculation model through the address, and imports the calculation model into the memory sharing pool of the memory sharing system. The computing resource receives the task, and the shared memory system first determines whether there is a historical prefill task. If so, the corresponding historical KV cache data is imported. The computing resource calculates the task through data interaction between different computing devices, generates and returns the result to the user service subsystem. In addition, during the task calculation process, if the calculated KV cache data has preservation value, the KV cache data with preservation value is stored in the shared memory system for storage.

[0126] Figure 8 FIG. 1 is a flow chart of a calculation method in one embodiment. It should be understood that although Figure 8 The steps in the flowchart are shown in sequence as indicated by the arrows, but the steps are not necessarily executed in the order indicated by the arrows; unless otherwise specified in this document, there is no strict order restriction for the execution of the steps, and the steps may be executed in other orders; and Figure 8 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0127] This specific embodiment is merely an explanation of the present invention and is not a limitation of the present invention. After reading this specification, those skilled in the art may make non-creative modifications to the present embodiment as needed. However, as long as they are within the scope of the claims of the present invention, they are protected by the patent law.

Claims

1. A computing system, characterized in that: include: A computing management system for acquiring and analyzing tasks and managing and scheduling the computing resources needed to execute the tasks; A shared memory system used to manage the public memory required to execute tasks and the data generated during the execution of tasks; A computing cluster used to receive tasks, calculate tasks, and output calculation results; The computing management system includes a management node for controlling the computing management system to directly access the shared memory system; the computing cluster includes a computing node for controlling the computing cluster to directly access the shared memory system; The shared memory system includes a storage subsystem for storing the model and parameters of the model used in executing the task, a data routing unit for providing a communication network between the shared memory system and the computing management system, and between the shared memory system and the computing cluster, and a shared memory pool for storing and managing data required during the task execution and data generated during the task execution; The shared memory pool includes a shared memory access management module for addressing the memory medium module, and a memory medium module for storing data required during task execution and data generated during task execution; The shared memory access management module includes an access port for providing a port for data interaction between the shared memory system and the computing management system and the computing cluster, and an interleaving management submodule for addressing the memory medium module in a partial interleaving mode to generate a shared memory unified address; The calculation methods of the shared memory unified address include: Where N is the Nth access port, C port is the memory capacity corresponding to each access port, A is each memory medium address, C is memory is the capacity of each memory medium, k is the number of memory media, x is the size of the alternating arrangement of the memory media, and M is the shared memory unified address corresponding to the Mth memory medium in the access port.

2. The computing system according to claim 1, characterized in that: The management node includes a CPU; the CPU is used to manage data calculations in the computing management system.

3. The computing system according to claim 1, wherein: The computing node includes an accelerator, and the accelerator is used to accelerate the computing of the task; wherein data is directly transmitted between the accelerator and the shared memory system.

4. A computing method, applied to the computing system according to any one of claims 1 to 3, characterized in that: The computing system includes a shared memory system, the shared memory system includes a shared memory pool, and the method includes: Obtaining tasks to be processed and analyzing computing resources required for the tasks to be processed; Querying free computing resources and allocating the free computing resources to the tasks to be processed; Reading the address of the computing model required by the task to be processed from the shared memory system, and importing the computing model into the shared memory pool; The computing model is used to call the computing resources to execute the task to be processed according to the type of the task to be processed, and the computing result is output.

5. The calculation method according to claim 4, characterized in that: The shared memory pool includes a model management table and a model parameter buffer area, the model parameter buffer area is used to cache the parameters of the used computing models, and the model management table is used to record the used computing models; The importing the computing model into the shared memory pool comprises: Searching the model management table and determining whether the model management table contains the calculation model; If the model management table includes the calculation model, directly reading the parameters corresponding to the calculation model from the model parameter buffer area, and importing the calculation model into the shared memory pool; If the model management table does not include the calculation model, the required calculation model is requested from the shared memory system, the requested calculation model is initialized and the calculation model is imported into the shared memory pool.

6. The calculation method according to claim 5, characterized in that: The method further comprises: When importing the computing model into the shared memory pool, monitoring the capacity of the model parameter buffer area; If the capacity of the model parameter cache exceeds a preset threshold, determining whether there is a computing model whose usage cycle has been exhausted in the model parameter cache; If there is a calculation model whose usage cycle is exhausted, determining whether the access frequency of the calculation model whose usage cycle is exhausted is greater than a preset value; If the access frequency of the computing model whose usage cycle is exhausted is not greater than the preset value, the computing model whose usage cycle is exhausted and whose access frequency is not greater than the preset value will be released.

7. The calculation method according to claim 4, characterized in that: The computing system includes a first accelerator and a second accelerator, the to-be-processed task includes a collaborative task and a migration task, and if the to-be-processed task is a collaborative task, calling the computing resource to execute the to-be-processed task according to the type of the to-be-processed task, and outputting the computing result includes: Determine whether the to-be-processed tasks include preset historical pre-filled tasks; If the to-be-processed task does not include a preset historical pre-filled task, controlling the second accelerator to directly process the to-be-processed task and generate a calculation result; If the to-be-processed task includes a preset historical pre-filled task, then importing the historical KV cache data corresponding to the historical pre-filled task; Controlling the first accelerator to calculate the historical KV cache data and generate calculated KV cache data; caching the calculated KV cache data into the shared memory pool, and importing the calculated KV cache data into the memory of the second accelerator; The second accelerator is controlled to read the calculated KV cache data from the memory of the second accelerator, and the tasks to be processed other than the historical pre-filled tasks are calculated according to the calculated KV cache data to generate results of the tasks to be processed.

8. The calculation method according to claim 4, characterized in that: The computing system includes a third accelerator and a fourth accelerator, the to-be-processed task includes a collaborative task and a migration task, and if the to-be-processed task is a migration task, calling the computing resource to execute the to-be-processed task according to the type of the to-be-processed task, and outputting the computing result includes: Controlling the third accelerator to calculate the task to be processed, generating pre-migration KV cache data, and writing the pre-migration KV cache data into the memory of the third accelerator; In response to the data migration instruction, import the pre-migration KV cache data in the memory of the third accelerator into the shared memory pool, and at the same time, control the third accelerator to continue calculating the pending task to generate post-migration KV cache data, and write the post-migration KV cache data into the shared memory pool; Importing the pre-migration KV cache data and the post-migration KV cache data into the memory of the fourth accelerator; Based on the pre-migration KV cache data and the post-migration KV cache data, the fourth accelerator is controlled to continue executing the task to be processed and generate a result of the task to be processed.

Citation Information

Patent Citations

  • Parallel processing interleaver applied in orthogonal frequency division multiplexing transmission system

    CN105577325A

  • Model training method, computing device and system

    CN118278540A