A data processing system, method, device, medium and program product

By introducing logical processing devices into the data processing system and determining the execution end of the computing operation based on available computing power, the problem of lag in the computing device during model deployment is solved, and the model training and inference efficiency is improved.

CN119887498BActive Publication Date: 2025-05-27SHANDONG HAILIANG INFORMATION TECH RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510361498.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-05-27
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

Existing computing devices are prone to lag during model deployment, and cannot fully utilize the model's performance, resulting in reduced model training and inference efficiency.

Method used

By introducing logical processing devices into the data processing system, the target model weight data and key-value pair data corresponding to the model processing task are stored, and the execution ends of each computing operation are determined based on the available computing power of the host, logic processing device and computing device, thereby optimizing the computing efficiency of the model.

Benefits of technology

This solution not only reduces the pressure of the operation of computing devices and storing the model, but also selects the appropriate execution end for each computing operation in the model, improves the computing efficiency of the model, and gives full play to the performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119887498B_ABST
    Figure CN119887498B_ABST
Patent Text Reader

Abstract

The present application discloses a data processing system, method, device, medium and program product, which relates to the field of computer technology. The present application stores the weight data of the target model corresponding to the model processing task and the key-value pair data to be processed in a logical processing device, and the execution end of each operation in the target model can be: a host, a logical processing device or a computing device. Therefore, not only can the pressure on the computing device to run and store the model be reduced, but also a suitable execution end can be selected for each operation in the model to improve the model operation efficiency. The execution end only needs to load the weight sub-data and key-value pair sub-data required for the currently executed operation from the logical processing device, and further accelerates the model operation efficiency through a small amount of data loading. Therefore, the present application can give full play to the performance of the model and improve the model training and inference efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a data processing system, method, device, medium, and program product. Background Art

[0002] Currently, as models become more and more complex, higher requirements are imposed on computing devices such as GPUs (Graphics Processing Units). Many computing devices experience lag during model deployment, unable to fully utilize the performance of the model, which may reduce the efficiency of model training and inference.

[0003] Therefore, how to improve the efficiency of model training and inference is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a data processing system, method, device, medium, and program product to improve the efficiency of model training and inference.

[0005] In a first aspect, this application provides a data processing system, including: a host, a first switching device, a logic processing device, a second switching device, and a computing device;

[0006] The host is connected to the logic processing device through the first switching device and to the computing device through the second switching device;

[0007] The host is configured to send model processing tasks to the computing device and the logic processing device;

[0008] The logic processing device is configured to store the weight data of the target model corresponding to the model processing task and the key-value pair data to be processed; determine the execution end of each operation in the target model according to the available computing power of the host, the logic processing device, and the computing device; the execution end is: the host, the logic processing device, or the computing device;

[0009] The execution end is configured to load the weight sub-data and key-value pair sub-data required for the currently executed operation from the logic processing device and store the operation result of the currently executed operation to the logic processing device.

[0010] In a second aspect, this application provides a data processing method, which is applied to a logic processing device and includes:

[0011] Store the weight data of the target model corresponding to the model processing task sent by the host and the key-value pair data to be processed;

[0012] Determine the execution end of each operation in the target model according to the available computing power of the host, the logical processing device, and the computing device connected to the host, so that the execution end loads the weight sub-data and key-value pair sub-data required for the currently executed operation from the logical processing device, and stores the operation result of the currently executed operation in the logical processing device; the execution end is: the host, the logical processing device, or the computing device;

[0013] Wherein, the host is connected to the logical processing device through a first switching device and connected to the computing device through a second switching device; the host sends the model processing task to the computing device and the logical processing device.

[0014] In a third aspect, the present application provides an electronic device, including:

[0015] A memory for storing a computer program;

[0016] A processor for executing the computer program to implement the data processing method disclosed above.

[0017] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the data processing method disclosed above.

[0018] In a fifth aspect, the present application provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the data processing method disclosed above are implemented.

[0019] Through the present application, the weight data of the target model corresponding to the model processing task and the key-value pair data to be processed are stored in the logical processing device, and the execution end of each operation in the target model can be: the host, the logical processing device, or the computing device. Therefore, not only can the pressure on the computing device to run and store the model be reduced, but also a suitable execution end can be selected for each operation in the model to improve the model operation efficiency. The execution end only needs to load the weight sub-data and key-value pair sub-data required for the currently executed operation from the logical processing device, and further accelerates the model operation efficiency through a small amount of data loading. Therefore, the present application can give full play to the performance of the model and improve the model training and inference efficiency.

[0020] Correspondingly, a data processing method, medium, and program product provided by the present application also have the above technical effects. Description of the Drawings

[0021] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0022] Figure 1 Schematic diagram of a data processing system disclosed in the present application;

[0023] Figure 2 Another schematic diagram of a data processing system disclosed in the present application;

[0024] Figure 3 Schematic diagram of a logic processing device disclosed in the present application;

[0025] Figure 4 Flowchart of a data processing method disclosed in the present application;

[0026] Figure 5 Schematic diagram of an electronic device disclosed in the present application;

[0027] Figure 6 Structural diagram of a server provided by the present application;

[0028] Figure 7 Structural diagram of a terminal provided by the present application. Detailed implementation manners

[0029] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some, rather than all, embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0030] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0031] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0032] See Figure 1As shown in the figure, an embodiment of the present application discloses a data processing system, including: a host, a first switching device, a logic processing device, a second switching device, and a computing device; the host is connected to the logic processing device through the first switching device and connected to the computing device through the second switching device. Among them, the logic processing device can be implemented based on FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuits). The computing device can be a GPU, TPU (Tensor Processing Unit), NPU (Neural Processing Unit), FPGA, ASIC, etc. The first switching device and the second switching device can be the same or different, that is: the first switching device and the second switching device can use the same protocol to enable the host to communicate with the logic processing device and enable the host to communicate with the computing device; they can also use different protocols to enable the host to communicate with the logic processing device and enable the host to communicate with the computing device. The first switching device can be connected to a plurality of logic processing devices and a plurality of extended memory devices, and the second switching device can be connected to a plurality of computing devices.

[0033] In this embodiment, the host is used to send model processing tasks to the computing device and the logic processing device; the logic processing device is used to store the weight data of the target model corresponding to the model processing task and the key-value pair data to be processed; according to the available computing power of the host, the logic processing device, and the computing device, determine the execution end of each operation operation in the target model; the execution end is: the host, the logic processing device, or the computing device; the execution end is used to load the weight sub-data and key-value pair sub-data required for the currently executed operation operation from the logic processing device, and store the operation result of the currently executed operation operation to the logic processing device. Among them, the weight sub-data can be a subset of the weight data of the target model, and the key-value pair sub-data can be a subset of the key-value pair data to be processed by the target model.

[0034] In this embodiment, the weight data of the target model corresponding to the model processing task and the key-value pair data to be processed are stored in the logic processing device, and the execution end of each operation operation in the target model can be: the host, the logic processing device, or the computing device. Therefore, not only can the pressure on the computing device to run and store the model be reduced, but also a suitable execution end can be selected for each operation operation in the model, improving the model operation efficiency. The execution end only needs to load the weight sub-data and key-value pair sub-data required for the currently executed operation operation from the logic processing device, and further accelerates the model operation efficiency through a small amount of data loading.

[0035] To further improve the model operation efficiency, during the execution of any operation, the execution end can load and cache the weight sub-data and key-value pair sub-data corresponding to the next operation of the currently executed operation from the logical processing device, so that after the currently executed operation is completed, the next operation can be directly executed based on the loaded data. Specifically, a ping-pong cache can be set at the execution end to implement a design of loading while executing, that is: execute the previous operation in one buffer area of the ping-pong cache, and at the same time load the weight sub-data and key-value pair sub-data corresponding to the next operation in another buffer area of the ping-pong cache, thereby saving data loading time.

[0036] In one implementation, the host is also connected to the extended memory device through the first switching device; the extended memory device is used to store weight data and / or key-value pair data when the storage space of the logical processing device is less than a preset threshold; the preset threshold is greater than the data volume size of the weight data and / or the data volume size of the key-value pair data. The extended memory device can be any type of storage device. In one implementation, the host and the first switching device, the first switching device and the logical processing device, and the first switching device and the extended memory device are connected through a cache coherence protocol. The cache coherence protocol such as the CXL (Compute Express Link) protocol can provide higher data throughput and lower latency.

[0037] In one implementation, the host and the second switching device, and the second switching device and the computing device are connected through the Peripheral Component Interconnect express (PCIe).

[0038] It should be noted that the process for the logical processing device to determine each operation in the target model includes: the logical processing device determines the execution memory size of a single operation, and splits the model operation logic in the target model into each operation according to the execution memory size; or the logical processing device divides the model operation logic in the target model according to the number of model layers to obtain M operation sub-logics; and divides each operation sub-logic into N operations. To facilitate the prefetching of the weight sub-data and key-value pair sub-data required for executing any operation, the logical processing device can store the weight data and key-value pair data in partitions according to each operation, that is: the weight sub-data and key-value pair sub-data required for each operation are stored in different storage areas in the logical processing device.

[0039] In this embodiment, in order to improve the data interaction efficiency between the logic processing device and the computing device, between the host and the logic processing device, and between the host and the computing device, the data interaction between the logic processing device and the computing device, between the host and the logic processing device, and between the host and the computing device can utilize the memory direct access technology. Memory direct access technologies include DMA (Data Memory Access, Direct Memory Access), RDMA (Remote Direct Memory Access, Remote Direct Memory Access), etc.

[0040] In one implementation, the execution end is used to locally store the operation result of the previous operation in two consecutive operation operations when both of them are executed locally. That is: the host, the logic processing device, and the computing device can adopt the In-place technology to save memory and save the time of memory allocation and release.

[0041] In one implementation, the logic processing device includes: a memory coherence engine for implementing the coherence tracking logic between local memory and external memory; a first-in-first-out buffer for sequentially buffering the weight sub-data prefetched from the weight data and the key-value pair sub-data prefetched from the key-value pair data; a computing module for determining the execution end of each operation operation in the target model according to the available computing power of the host, the logic processing device, and the computing device; executing the received operation operation; and a memory for storing the weight data and the key-value pair data.

[0042] The data processing system provided in this embodiment can include multiple hosts, and each host can connect its downstream devices (such as computing devices, logic processing devices, etc.) according to the Figure 1 shown structure; at the same time, different hosts can be connected through any type of switching device. Connecting each device according to this logic can enable the entire system to fully exert the performance of the model and improve the model training and inference efficiency.

[0043] Please refer to Figure 2 , another data processing system includes: a host, a CXL switch, and a PCIE switch; multiple CXL-FPGAs (i.e., logic processing devices) and a CXL extended memory device are connected under each CXL interface of the CXL switch (i.e., the first switching device); multiple GPU computing devices are connected under the PCIE switch (i.e., the second switching device).

[0044] The CXL memory device is a memory expansion device based on the CXL protocol, which is used to solve the problem of insufficient memory of GPUs and hosts. The CXL-FPGA is a device that supports the CXL protocol and has computing capabilities, specifically including: a device coherency engine (dcoh), a first-in-first-out buffer: I / O FIFO, a computing module, and a DDR memory. For details, please refer to Figure 3 . The computing module mainly completes the parallelizable forward computing tasks in the inference task, and the I / O FIFO is responsible for prefetching inference data / model weights; the memory coherency engine is used for the coherency tracking logic of device memory and host memory, and the DDR memory is used to cache key-value pair data, model data, etc. required for computing.

[0045] In this embodiment, the software architecture reduces the use of memory from the dimensions of loading and computing. The computing is mainly based on GPUs, supplemented by CXL-FPGAs and hosts. Since the weight parameters of large models are large, in this embodiment, the large models are stored in the host memory, the device memory of the CXL-FPGA, or the CXL extended memory device.

[0046] Considering that the model is stored in the CXL extended memory device and the cross-media reading speed of the GPU is slow, in this embodiment, to solve the problem of slow data loading and reading, the model weight data is loaded in segments. First, according to the number of layers L of the model, the model is split into layers [1, L]; correspondingly, a memory space of size 1 / L is allocated for the model in the GPU to load the model layer by layer, which can significantly reduce the memory and time required for each load.

[0047] Considering that the model weights of size 1 / N may still cause GPU memory overflow, the model can be split into M×N segments, where N is the approximate number of segments by which the model is segmented according to the weight size of each layer. Assuming that the total weight data size of the model is Q, the pingpang cache size of the GPU is set to (Q / MN)×2. Thus, during the execution of any segment of the model on the GPU, the pingpang cache loads the weight sub-data and key-value pair sub-data corresponding to the subsequent operation of the currently executed operation, minimizing the memory occupancy of the model by the GPU and improving the model processing efficiency at the same time. For example, when executing the a-th segment of the second layer of the model, the model weights of the [1,a) segment of the third layer can be considered for loading, so that the memory copy can be performed while computing, hiding the copy time into the computing time. Moreover, the GPU Direct technology can be used during loading to directly transfer data between the GPU and other devices, reducing CPU intervention and data transfer latency, and further optimizing the speed and efficiency of data transfer. Also, since the weight sizes of each layer of the model are the same, there is no need for repeated allocation and release, and only the weights of the previous layer need to be overwritten during loading. For the weights of the post-processing layer Hiddenlayer in the model, since their memory occupancy is relatively small, they only need to be allocated when in use.

[0048] For the inference of large language models, the kv cache data (i.e., key-value pair data) will increase linearly with the number of input questions, the length of the input, and the length of the inference output text. Assuming that 10 questions are inferred simultaneously, and the total token length of the input and output of each question is 128×1024, then for this inference, the kv cache data alone will exceed 200G. To solve the problem of excessive kv cache data, in this embodiment, the kv cache data is allocated and stored in the device memory of the CXL-FPGA or the CXL extended memory device, and only the current kv cache data required for GPU operation and the corresponding operation results are allocated in the GPU video memory. Combining the above model layering, only a kv cache data space of size 1 / L needs to be allocated for the model in the GPU memory, significantly reducing the memory occupancy required for each loading.

[0049] When the GPU is in the calculation process of any layer, the GPU stores the current kv cache data in the GPU video memory, and then reads the kv cache data required for the next layer of calculation from the device memory of the CXL-FPGA or the CXL extended memory device, and stores it in the GPU video memory through the GPU Direct technology. It should be noted that when the large model is inferring, one token is inferred each time, and one token represents one character / word. Each time of inference will append the kv cache data of the current layer of the current inference to the kv cache data of the previous time, that is, complete the update operation of the kv cache data; the subsequent calculation needs to retrieve the corresponding data from the updated kv cache data. Similarly, the next inference needs to use all the current and previous kv cache data.

[0050] Considering that the intermediate result dimensions of each layer are the same, there is no need to repeatedly allocate and release, and only the intermediate results of the previous layer of calculation need to be overwritten, thus reducing the memory overhead. At the same time, the In-place technology is adopted to complete the result overwrite of some operations locally. For example: after the matrix dot product calculation is completed, the calculation result can be cached locally, so as to use this calculation result to continue to perform operations such as matrix addition, thus reducing the memory allocation and release time and saving the memory overhead.

[0051] For each operation in the model, the CXL-FPGA or the host can calculate the computational complexity of each operation, and allocate the operations according to the computational complexity and the computing power of each device. For example: if the computational complexities of three operations are a, b, and c from large to small, and the available computing powers of the GPU, CXL-FPGA, and host are A, B, and C from large to small, then the greater the available computing power of the device, the higher the computational complexity of the allocated operation. If a / A > b / B > c / C, then the three operations are allocated to 3 devices respectively; if a / A < c / C and a / A > b / B, it means that the computing power of device C is much smaller than that of device A, then the operations with computational complexities of a and c are calculated by device A; similarly, if a / A < b / B, then all three operations are calculated by device A. Moreover, non-interfering operations can be calculated in parallel by different devices, improving the hardware utilization rate and thus improving the model inference efficiency. For example: the matrix multiplication operations of three matrices can be performed on the GPU, CXL-FPGA, and host respectively, and then the CXL-FPGA saves the results.

[0052] In this embodiment, by storing the weights, intermediate results, and kv cache data during non-settlement outside the GPU, a larger-scale model can be realized to run on limited GPU hardware resources.

[0053] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. A data processing method provided in an embodiment of the present application is introduced below, and a data processing method described below can be cross-referenced with other embodiments described herein.

[0054] See also Figure 4 As shown, the embodiment of the present application discloses a data processing method, which is applied to a logic processing device, including:

[0055] S401. The weight data of the target model and the key-value pair data to be processed corresponding to the model processing task sent by the storage host.

[0056] S402. Determine the execution end of each operation in the target model according to the available computing power of the host, the logic processing device, and the computing device connected to the host, so that the execution end loads the weight sub-data and key-value pair sub-data required for the currently executed operation from the logic processing device, and stores the operation result of the currently executed operation to the logic processing device.

[0057] In this embodiment, the execution end is: a host, a logic processing device or a computing device; the host is connected to the logic processing device through a first switching device, and is connected to the computing device through a second switching device; the host sends the model processing task to the computing device and the logic processing device. In one embodiment, the host is also connected to an extended memory device through a first switching device; the extended memory device is used to store weight data and / or key-value pair data when the storage space of the logic processing device is less than a preset threshold; the preset threshold is greater than the data volume of the weight data and / or the data volume of the key-value pair data. In one embodiment, the host and the first switching device, the first switching device and the logic processing device, and the first switching device and the extended memory device are connected through a cache consistency protocol. In one embodiment, the host and the second switching device, and the second switching device and the computing device are connected through a high-speed serial computer expansion bus standard.

[0058] In one implementation, the logic processing device can load and cache the weight sub-data and key-value pair sub-data corresponding to the next operation of the currently executed operation from the logic processing device during the execution of any operation.

[0059] In one implementation, the logic processing device can determine the execution memory size of a single computing operation, and split the model computing logic in the target model into individual computing operations according to the execution memory size.

[0060] In one embodiment, the logic processing device can divide the model operation logic in the target model according to the number of model layers to obtain M operator logics; and divide each operator logic into N operation operations.

[0061] In one embodiment, the logic processing device can partition and store the weight data and key-value pair data according to each operation operation.

[0062] In one embodiment, data interaction is performed between the logic processing device and the computing device, between the host and the logic processing device, and between the host and the computing device by using the direct memory access technology.

[0063] In one embodiment, when two consecutive operation operations are both executed locally, the logic processing device can locally store the operation result of the previous operation operation in the two consecutive operation operations.

[0064] In one embodiment, the logic processing device includes: a memory coherence engine for implementing the coherence tracking logic between the local memory and the external memory; a first-in first-out buffer for caching the weight sub-data prefetched from the weight data and the key-value pair sub-data prefetched from the key-value pair data in sequence; a computing module for determining the execution end of each operation operation in the target model according to the available computing power of the host, the logic processing device, and the computing device; executing the received operation operation; and a memory for storing the weight data and the key-value pair data.

[0065] Among them, for the more specific working processes of each module and unit in this embodiment, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.

[0066] It can be seen that in this embodiment, the weight data of the target model corresponding to the model processing task and the key-value pair data to be processed are stored in the logic processing device, and the execution end of each operation operation in the target model can be: the host, the logic processing device, or the computing device. Therefore, not only can the pressure on the computing device to run and store the model be reduced, but also a suitable execution end can be selected for each operation operation in the model, improving the model operation efficiency. The execution end only needs to load the weight sub-data and the key-value pair sub-data required for the currently executed operation operation from the logic processing device, and further accelerates the model operation efficiency through a small amount of data loading. Therefore, this application can give full play to the performance of the model and improve the model training and inference efficiency.

[0067] Next, an electronic device provided by an embodiment of the present application is introduced. The electronic device described below can be mutually referred to with other embodiments described herein. The electronic device described in this embodiment can be any functional module or device mentioned above.

[0068] See Figure 5As shown in the figure, an embodiment of the present application discloses an electronic device, including:

[0069] A memory 501 for storing a computer program;

[0070] A processor 502 for executing the computer program to implement the method disclosed in any of the above embodiments.

[0071] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: storing the weight data of the target model corresponding to the model processing task sent by the host and the key-value pair data to be processed; determining the execution end of each operation in the target model according to the available computing power of the host, the logical processing device, and the computing devices connected to the host.

[0072] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: loading the weight sub-data and key-value pair sub-data required for the currently executed operation from the logical processing device, and storing the operation result of the currently executed operation to the logical processing device.

[0073] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: storing the weight data and / or key-value pair data when the storage space of the logical processing device is less than a preset threshold; the preset threshold is greater than the data volume size of the weight data and / or the data volume size of the key-value pair data.

[0074] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: during the execution of any operation, loading and caching the weight sub-data and key-value pair sub-data corresponding to the subsequent operation of the currently executed operation from the logical processing device.

[0075] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: determining the execution memory size of a single operation, and splitting the model operation logic in the target model into each operation according to the execution memory size.

[0076] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: dividing the model operation logic in the target model according to the model layers to obtain M operation sub-logics; dividing each operation sub-logic into N operations.

[0077] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: storing the weight data and key-value pair data in partitions according to each operation.

[0078] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: when two consecutive arithmetic operations are both executed locally, locally store the operation result of the previous arithmetic operation in the two consecutive arithmetic operations.

[0079] Furthermore, an embodiment of the present application also provides an electronic device. Among them, the above-mentioned electronic device can be either a Figure 6 server as shown, or a Figure 7 terminal as shown. Figure 6 and Figure 7 are both structural diagrams of electronic devices shown according to an exemplary embodiment, and the content in the figure cannot be regarded as any limitation on the scope of use of the present application.

[0080] Figure 6 This is a schematic structural diagram of a server provided by an embodiment of the present application. The server may specifically include: at least one processor, at least one memory, a power supply, a communication interface, an input / output interface, and a communication bus. Among them, the memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the relevant steps in the data processing disclosed in any of the foregoing embodiments.

[0081] In this embodiment, the power supply is used to provide operating voltage for each hardware device on the server; the communication interface can create a data processing channel between the server and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0082] In addition, as a carrier for resource storage, the memory can be a read-only memory, a random access memory, a disk, or an optical disc, etc., and the resources stored thereon include an operating system, a computer program, and data, etc., and the storage method can be temporary storage or permanent storage.

[0083] Among them, the operating system is used to manage and control each hardware device and computer program on the server to implement the operation and processing of data in the memory by the processor, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the data processing method disclosed in any of the foregoing embodiments, the computer program can further include a computer program that can be used to complete other specific tasks. In addition to data such as update information of the application program, the data can also include data such as developer information of the application program.

[0084] Figure 7A schematic structural diagram of a terminal provided by an embodiment of the present application. The terminal may specifically include, but is not limited to, a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.

[0085] Generally, the terminal in this embodiment includes a processor and a memory.

[0086] Among them, the processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA, PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor may be integrated with a GPU, and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, which is used to process computational operations related to machine learning.

[0087] The memory may include one or more computer non-volatile storage media, and the computer non-volatile storage media may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory is at least used to store the following computer programs. After the computer programs are loaded and executed by the processor, the relevant steps in the data processing method executed by the terminal side disclosed in any of the foregoing embodiments can be implemented. In addition, the resources stored in the memory may also include an operating system and data, etc., and the storage method may be transient storage or permanent storage. Among them, the operating system may include Windows, Unix, Linux, etc. The data may include, but is not limited to, update information of application programs.

[0088] In some embodiments, the terminal may further include a display screen, an input / output interface, a communication interface, sensors, a power supply, and a communication bus.

[0089] Those skilled in the art can understand that Figure 7 the structure shown in

[0090] The following introduces a computer-readable storage medium provided by an embodiment of the present application. The computer-readable storage medium described below can be referred to in conjunction with other embodiments described herein.

[0091] A computer-readable storage medium is used to store a computer program. When the computer program is executed by a processor, it implements the data processing method disclosed in the foregoing embodiments. The computer-readable storage medium is a non-volatile computer-readable storage medium. As a carrier for storing resources, it can be a read-only memory, a random access memory, a magnetic disk, an optical disc, etc. The resources stored thereon include an operating system, a computer program, data, etc. The storage method can be transient storage or permanent storage.

[0092] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.

[0093] In this embodiment, when the processor executes the computer program stored in the computer-readable storage medium, the following steps can be specifically implemented: storing the weight data of the target model corresponding to the model processing task sent by the storage host and the key-value pair data to be processed; determining the execution end of each operation in the target model according to the available computing power of the host, the logical processing device, and the computing devices connected to the host.

[0094] In this embodiment, when the processor executes the computer program stored in the computer-readable storage medium, the following steps can be specifically implemented: loading the weight sub-data and key-value pair sub-data required for the currently executed operation from the logical processing device, and storing the operation result of the currently executed operation to the logical processing device.

[0095] In this embodiment, when the processor executes the computer program stored in the computer-readable storage medium, the following steps can be specifically implemented: when the storage space of the logical processing device is less than a preset threshold, storing the weight data and / or the key-value pair data; the preset threshold is greater than the data volume of the weight data and / or the data volume of the key-value pair data.

[0096] In this embodiment, when the processor executes the computer program stored in the computer-readable storage medium, the following steps can be specifically implemented: during the execution of any operation, loading and caching the weight sub-data and key-value pair sub-data corresponding to the subsequent operation of the currently executed operation from the logical processing device.

[0097] In this embodiment, when the processor executes the computer program stored in the computer-readable storage medium, the following steps may be specifically implemented: determining the execution memory size of a single arithmetic operation, and splitting the model arithmetic logic in the target model into respective arithmetic operations according to the execution memory size.

[0098] In this embodiment, when the processor executes the computer program stored in the computer-readable storage medium, the following steps may be specifically implemented: dividing the model arithmetic logic in the target model according to the number of model layers to obtain M sub-arithmetic logics; and dividing each sub-arithmetic logic into N arithmetic operations.

[0099] In this embodiment, when the processor executes the computer program stored in the computer-readable storage medium, the following steps may be specifically implemented: storing the weight data and key-value pair data in partitions according to each arithmetic operation.

[0100] In this embodiment, when the processor executes the computer program stored in the computer-readable storage medium, the following steps may be specifically implemented: when two consecutive arithmetic operations are both executed locally, locally storing the operation result of the previous arithmetic operation in the two consecutive arithmetic operations.

[0101] Next, a computer program product provided by an embodiment of the present application will be introduced. A computer program product described below may be referred to each other with other embodiments described herein.

[0102] A computer program product includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the data processing method disclosed above are implemented.

[0103] Another embodiment of the present application further provides a computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps in any of the above embodiments are implemented.

[0104] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments may be referred to each other.

[0105] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0106] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of non-volatile storage medium well known in the art.

[0107] Specific examples are used in this article to illustrate the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A data processing system, characterized in that: include: A host, a first switching device, a logic processing device, a second switching device, and a computing device; The host is connected to the logic processing device via the first switching device, and is connected to the computing device via the second switching device; The host is used to send the model processing task to the computing device and the logic processing device; The logic processing device is used to store the weight data of the target model corresponding to the model processing task and the key-value pair data to be processed; Determine the execution end of each operation in the target model according to the available computing power of the host, the logic processing device and the computing device; the execution end is: the host, the logic processing device or the computing device; The execution end is used to load the weight sub-data and key-value pair sub-data required for the currently executed operation from the logic processing device, and store the operation result of the currently executed operation to the logic processing device.

2. The data processing system according to claim 1, characterized in that The execution end is used to load and cache the weight sub-data and key-value pair sub-data corresponding to the next operation of the currently executed operation from the logic processing device during the execution of any operation.

3. The data processing system according to claim 1, characterized in that: The host is also connected to an extended memory device via the first switching device; The extended memory device is used to store the weight data and / or the key-value pair data when the storage space of the logic processing device is less than a preset threshold; The preset threshold is greater than the data volume of the weight data and / or the data volume of the key-value pair data.

4. The data processing system according to claim 3, characterized in that: The host and the first switching device, the first switching device and the logic processing device, and the first switching device and the extended memory device are connected via a cache coherence protocol.

5. The data processing system according to claim 1, characterized in that: The host and the second switching device, and the second switching device and the computing device are connected via a high-speed serial computer expansion bus standard.

6. The data processing system according to claim 1, characterized in that: The logic processing device is used to determine the execution memory size of a single operation, and split the model operation logic in the target model into individual operation operations according to the execution memory size.

7. The data processing system according to claim 1, characterized in that: The logic processing device is used to divide the model operation logic in the target model according to the number of model layers to obtain M operator logics; and divide each operator logic into N operation operations.

8. The data processing system according to claim 1, characterized in that: The logic processing device is used to partition and store the weight data and the key-value pair data according to various calculation operations.

9. The data processing system according to claim 1, characterized in that: Data is exchanged between the logic processing device and the computing device, between the host and the logic processing device, and between the host and the computing device using a memory direct access technology.

10. The data processing system according to claim 1, characterized in that: The execution end is used to locally store the operation result of the previous operation in two consecutive operation operations when both of the two consecutive operation operations are executed locally.

11. The data processing system according to any one of claims 1 to 10, characterized in that: The logic processing device comprises: Memory consistency engine, used to implement consistency tracking logic between local memory and external memory; A first-in-first-out buffer, used for sequentially caching weight sub-data pre-fetched from the weight data and key-value pair sub-data pre-fetched from the key-value pair data; A computing module, configured to determine an execution end of each computing operation in the target model according to the available computing power of the host, the logic processing device and the computing device; and execute the received computing operation; A memory is used to store the weight data and the key-value pair data.

12. A data processing method, characterized in that: Applicable to logic processing devices, including: The weight data of the target model and the key-value pair data to be processed corresponding to the model processing task sent by the storage host; Determine the execution end of each operation in the target model according to the available computing power of the host, the logic processing device and the computing device connected to the host, so that the execution end loads the weight sub-data and key-value pair sub-data required for the currently executed operation from the logic processing device, and stores the operation result of the currently executed operation to the logic processing device; the execution end is: the host, the logic processing device or the computing device; The host is connected to the logic processing device via a first switching device and connected to the computing device via a second switching device; the host sends the model processing task to the computing device and the logic processing device.

13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program to implement the method according to claim 12.

14. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein the computer program implements the method according to claim 12 when executed by a processor.

15. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to claim 12 is implemented.

Citation Information

Patent Citations

  • New energy charging pile power supply intelligent management method and system, electronic equipment and storage medium

    CN118410955A

  • Inference method and system, computer equipment and storage medium

    CN119378681A