A multi-stage caching method for accelerating AI data processing
By introducing a multi-stage caching method into the AI data processing pipeline to cache the original data and preprocessed data, the problems of limited acceleration effect of data processing and insufficient compatibility across hardware platforms in the prior art are solved, and more efficient data processing and better hardware compatibility are achieved.
Patent Information
- Application Number
- CN202510228500.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The failure of existing AI data processing technologies to effectively cache preprocessed data, resulting in limited acceleration of data processing, and most technologies are limited to GPUs, lacking compatibility across hardware platforms.
A multi-stage caching method is proposed. By building a data processing pipeline, setting up a cache client, a cache server and a shared space, the cache data set includes the original data and preprocessed data, using the gRPC framework and the shared space for information exchange, and node verification and optimization of the graph structure is performed to generate and execute graphs for data processing.
This method not only caches raw data, but also caches preprocessed data, reducing I/O operations and data preprocessing time, improving data processing efficiency, and is suitable for multiple hardware platforms, enhancing the compatibility and ease of use of data processing.
Smart Images

Figure CN119782008B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence data processing, and particularly relates to a multi-stage caching method for accelerating AI data processing. Background Art
[0002] In the model training of artificial intelligence and deep learning, AI data processing is an important pre-stage. Therefore, how to improve the speed of AI data processing has become a practical need.
[0003] Existing caching technologies such as Quiver and ICACHE use caching strategies that focus on the data loading layer, call the GPU to cache the original data, retain important data, and evict unimportant data, thereby improving data processing efficiency.
[0004] However, the existing technologies only consider caching the original data and do not consider caching the preprocessed data. Moreover, the time occupied by data preprocessing in the data loading stage is often very high. Therefore, the existing technologies have limited acceleration effects on data processing, and most of the existing technologies are only used on the GPU and lack compatibility across multiple hardware platforms. Therefore, the present invention provides a multi-stage caching method for accelerating AI data processing to solve the problems existing in the prior art. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-stage caching method for accelerating AI data processing to solve the following technical problems:
[0006] The existing caching technologies in the data processing process do not consider caching the preprocessed data, resulting in limited acceleration effects in data processing. Moreover, the existing technologies are often limited to use on the GPU and lack sufficient cross-hardware platform compatibility.
[0007] The purpose of the present invention can be achieved through the following technical solutions:
[0008] A multi-stage caching method for accelerating AI data processing includes the following steps:
[0009] Construct a data processing pipeline through a top-level interface, and at the same time, insert caching operations into the data processing pipeline in advance and mark the positions.
[0010] Set up a cache client, a cache server, and a shared space. The data processing pipeline communicates through the cache client and the cache server. The cache server is used to store the cache dataset, and the cache dataset includes the loaded original data and the preprocessed data. Information exchange between the cache client and the cache server is carried out using the gRPC framework and the shared space.
[0011] Convert the constructed data processing pipeline into a graph structure, and perform node verification on the graph structure. After the verification passes, divide the cached data set into a mapped data set and an unmapped data set according to whether the cached data set supports full mapping to memory. For different data sets, use corresponding methods to transform the graph structure, generate an execution graph, and optimize duplicate operators;
[0012] Start the data processing pipeline and perform data processing according to the generated execution graph. When the data pipeline performs cached data writing and cached data reading of the cached data set, call the cache client to send request information to the cache server. The cache server receives the request information and performs corresponding operations, and returns corresponding information to the cache client after completing the operations.
[0013] As a further solution of the present invention: when using the gRPC framework between the cache client and the cache server, set variables rq and reply for each request sent by the cache client to the cache server. The rq is used to record the information that needs to be sent to the cache server, and the reply is used to obtain the response information sent from the cache server.
[0014] As a further solution of the present invention: the specific process of verifying the graph structure is as follows:
[0015] Call the cache verification passed function, start from the starting node of the graph structure, traverse all nodes of the entire graph structure along the data processing path, and separately filter out the nodes for batch processing, connection, filtering, skipping, fetching, and compression operations. When there is no pre-inserted cache operation in the subsequent data processing path of any node, it is determined that the verification passes; otherwise, it is determined that the verification fails, and an error report is output after the traversal is completed to remind the user to modify the insertion position of the cache operation.
[0016] As a further solution of the present invention: after determining that the verification passes, the specific process of transforming the graph structure is as follows:
[0017] When the cached data set is a mapped data set, create a cache merge node in the graph structure and insert it on the map node as the parent node of the map node. Obtain a sampler from the data loading node, create a cache lookup node according to the sampler, and use the cache lookup node to replace the sampler to find and return the data cached in the cache merge node. Call the Build() function of all nodes in the graph structure to sequentially construct the operators corresponding to each node at runtime, generate an execution graph, and mark the operator corresponding to the cache merge node as a cache merge operator and the operator corresponding to the cache lookup node as a cache lookup operator;
[0018] When the cached data set is an unmappable data set, obtain the sampler corresponding to the data loading node that belongs to the leaf node within the graph structure, create a cache node according to the sampler, replace the sampler with the cache node, insert the cache node onto the map node as the parent node of the map node, call the Build() function of all nodes in the graph structure, sequentially construct the operators corresponding to each node at runtime to generate an execution graph, and mark the operator corresponding to the cache node as a cache operator.
[0019] As a further solution of the present invention: The specific process of optimizing duplicate operators for the graph structure is as follows:
[0020] Judge the type of the cached data set. When the cached data set is a mappable data set, obtain the position of the duplicate operator in the operation queue. If the duplicate operator is before the cache lookup operator and the cache merge operator, set the duplicate count of the cache lookup operator according to the duplicate count set in the duplicate operator; otherwise, set the duplicate count of the cache lookup operator according to the duplicate count set in the cache merge operator.
[0021] When the cached data set is an unmappable data set, mark the data processed by repeating the duplicate operator as a cache object, and traverse the graph structure from top to bottom. During the traversal process, add the cache objects to the cache node stack one by one. When the duplicate optimization mechanism traverses the graph structure from bottom to top, take out the corresponding cache objects from the cache node stack one by one, and set the duplicate count of the cache objects to 1.
[0022] As a further solution of the present invention: The specific running processes of the cache merge operator and the cache lookup operator are as follows:
[0023] When the data loading operator performs a data loading operation, provide the sampler to the cache lookup operator to obtain the ID of the data row. After the cache lookup operator obtains the ID, start N child threads in the main thread, where N is a positive integer. After collecting a preset number of IDs, the main thread distributes the IDs to the child threads through an internal queue. The child threads batch send cache query requests to the cache server. If the ID of the cache query request hits in the cache server, send the data stream corresponding to the ID to the cache merge operator; otherwise, the cache lookup operator sends the unhit ID to the data loading operator, and the data loading operator reads and parses the data stream again, and at the same time sends the data stream with placeholders to the cache merge operator.
[0024] The cache merge operator continuously waits and receives the data stream, merges and processes the data stream corresponding to the hit ID and stores it in the cache server. For the data stream containing placeholders, replace the placeholders with actual data, and then store the complete data stream in the cache server after the replacement is completed.
[0025] As a further solution of the present invention: The specific operation process of the cache operator is as follows:
[0026] The cache operator starts M sub-threads in the main thread, where M is a positive integer. Before each sub-thread successfully stores data into the cache server and receives the return information, it remains in a blocked state; when the data storage tasks of all sub-threads are completed and the return information is obtained, the cache operator switches to the acquisition state and starts to read the data stored in the cache server through the cache operator.
[0027] As a further solution of the present invention: The specific processes of the data processing pipeline for cache writing and cache reading of the cache dataset are as follows:
[0028] When performing cache writing, the data processing pipeline calls the cache client to submit an asynchronous buffer request to the gRPC queue. After receiving the request, the gRPC queue returns a shared memory address. Subsequently, the cache client starts to continuously write the cache dataset into the shared memory. When the shared memory is full, the buffer client submits a cache write request to the gRPC queue, and the gRPC thread submits the cache write request to the request queue of the cache server. The working threads of the cache server write the cache dataset from the shared memory into the memory of the cache server in sequence according to the request queue;
[0029] When performing cache reading, the data processing pipeline calls the cache client to submit a data row request to the gRPC queue. After receiving the request, the gRPC queue locates the position of the data row in the memory of the cache server, allocates shared memory for the data row, and simultaneously submits a data extraction request to the request queue of the cache server. The working threads of the cache server write the cache dataset from the memory of the cache server into the shared memory. When the cache dataset is successfully read into the shared memory, the gRPC queue returns the shared memory address corresponding to the data to the cache client. The cache client copies the cache dataset from the shared memory to the target position in the data processing pipeline and submits a request to release the shared memory to the gRPC queue, and the gRPC thread completes the memory release.
[0030] As a further solution of the present invention: During the cache writing and cache reading operations of the cache dataset by the data processing pipeline, it also includes optimizing the memory management. The specific optimization process is as follows:
[0031] When optimizing the memory management, based on the NUMA framework, according to the ratio of the memory space of the cache server to the total system memory space, a cache memory pool perceived by the NUMA framework is configured. Each cache memory pool consists of memory blocks, the number of memory blocks is set to the number of CPUs, and the memory space for the memory blocks is allocated through specific NUMA nodes.
[0032] Advantages of the present invention:
[0033] The present invention proposes a caching method that can penetrate the entire data processing flow, enabling flexible caching at various stages of data processing. It can cache not only the original data but also the preprocessed data. By caching during the preprocessing stage, the present invention can reduce both I / O operations and data preprocessing operations simultaneously, improving the efficiency of the data processing stage, thus effectively preventing data processing from becoming a bottleneck in deep neural network training. Moreover, the present invention designs multiple automatic optimization modules for the data processing pipeline after users use the cached data, including automatic verification and transformation of the data pipeline graph structure, optimizing data transmission using shared memory, simple cache sharing, etc. At the same time, the present invention also provides direct API support for multiple popular datasets in multiple scenarios, has good hardware compatibility, and users can directly use the designed interfaces without having to design them themselves, improving the usability and user experience of the present invention. Brief Description of the Drawings
[0034] The present invention will be further described below with reference to the accompanying drawings.
[0035] Figure 1 is a schematic flowchart of the present invention;
[0036] Figure 2 is a schematic flowchart of cache writing and cache reading during data processing;
[0037] Figure 3 is a schematic flowchart of graph structure conversion for a mapped dataset;
[0038] Figure 4 is a schematic flowchart of generating graph structure conversion for an unmapped dataset. Detailed Embodiments
[0039] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] Please refer to Figures 1-4 as shown, the present invention is a multi-stage caching method for accelerating AI data processing, including the following steps:
[0041] The user first enables the cache server in the terminal, then creates a session and obtains the corresponding ID, and then uses the session ID to create a cache instance ds.DatasetCache(), which needs to be passed in when using the cache later. The user then customizes the image preprocessing operation Transform, which will act on the dataset.
[0042] In the context of cache acceleration technology design, the APIs of data loading operators and map operators have been updated. If the user chooses to cache in the data loading phase, the cache instance is passed to the "cache" flag of Cifar10Dataset(). If the user chooses to cache in the data preprocessing phase, the cache instance is passed to the "cache" flag of map(). At this point, the user has completed the construction of the data processing pipeline and the insertion of the cache operator.
[0043] Set up a cache client, cache server and shared space. The user cannot communicate directly with the cache server through the cache client. You only need to insert a cache operator at the specified position in the data processing pipeline. The cache operator will automatically call the cache client interface. The cache server is used to store cache data sets, which include loaded original data and preprocessed data. The cache client and cache server use the gRPC framework and shared space to exchange information. Each request is sent in protobuf format, which is a data format for efficiently serializing structured data.
[0044] When data tensors need to be transferred, they must first be serialized. We call different functions to serialize the tensor and its corresponding header information. The header information includes the data size, an ID that uniquely identifies the tensor row, and the size of the header itself. After serializing the tensor, we write it to shared memory to optimize the transfer process.
[0045] Convert the constructed data processing pipeline into a graph structure, and perform node verification on the graph structure. Call the CacheValidationPass function, starting from the starting node of the graph structure, traverse all nodes of the entire graph structure along the data processing path, and separately filter out the nodes for batch, concat, filter, skip, take, and zip operations. When there is no pre-inserted cache operation in the subsequent data processing path of any node, it is judged that the verification passes; otherwise, it is judged that the verification fails, and an error report is output after the traversal is completed to remind the user to modify the insertion position of the cache operation. Verification is particularly important when using the map operator with random data augmentation, when caching operators are nested (for example, caching is used in both the data loading and preprocessing stages), or when a Repeat operator is added after the caching operator in a mapable dataset.
[0046] After the verification passes, divide the cached dataset into a mapable dataset and a non-mapable dataset according to whether it supports full mapping into memory, and adopt corresponding methods to transform the graph structure for different datasets.
[0047] When the cached dataset is a mapable dataset, create a CacheMergeNode in the graph structure and insert it above the map node as the parent node of the map node. Obtain the sampler from the data loading node, create a CacheLookupNode according to the sampler, and use the CacheLookupNode to replace the sampler for looking up and returning the data cached in the CacheMergeNode. Call the Build() function of all nodes in the graph structure to sequentially construct the operators corresponding to each node at runtime, generate an execution graph, and mark the operator corresponding to the cache node as the CacheLookupOperator, and the operator corresponding to the cache lookup node as the CheLookupOpeacrator;
[0048] When the cached dataset is a non-mapable dataset, obtain the sampler corresponding to the data loading node belonging to the leaf node in the graph structure, create a CacheNode according to the sampler, and use the CacheNode to replace the sampler. Insert the CacheNode above the map node as the parent node of the map node. Call the Build() function of all nodes in the graph structure to sequentially construct the operators corresponding to each node at runtime, generate an execution graph, and mark the operator corresponding to the cache node as the CacheOperator.
[0049] It should be noted that CacheLookupOperator, CheLookupOpeacrator, and CacheOperator are all operators designed by us for caching operations. Only one or two of them are enabled according to different data sets and are used to perform caching tasks in two stages of the data processing pipeline.
[0050] The present invention also considers the optimization of repeated operators. We construct a RepeatPass to optimize the repeated operators in the graph structure containing caching operators. Since there are different processes for mappable and non-mappable data sets in cache processing, two different methods are adopted to optimize the repeated operators in their respective data processing pipelines. The specific optimization process is as follows:
[0051] For mappable data sets, if the repeated operator is inserted before the caching operator, when RepeatPass accesses the cache lookup operator, it does not immediately set the repetition count; instead, it saves its shared pointer to the cache lookup member variable. Later, when accessing the Repeat operator, the repetition count of the cache lookup operator is set according to the repetition count specified in the Repeat operator to ensure consistency. On the contrary, if the repeated operator is inserted after the caching operator, RepeatPass also does not set the repetition count when accessing the cache lookup operator; it only sets the same repetition count after accessing the cache merge operator;
[0052] For non-mappable data sets, we focus on the case where the repeated operator is inserted after the caching operator. Specifically, we regard the cache object for which the repeated operator performs repeated operations as a caching operator. When RepeatPass traverses the graph structure from top to bottom, each caching operator is added to the cache node stack. Subsequently, when RepeatPass traverses the graph structure from bottom to top, if the caching operator is accessed again, the previously cached operator is taken out of the stack and its repetition count is set to 1. In this way, the caching operator will only execute once to completely store the data in the cache server.
[0053] For the cache merge operator (CheLookupOpeacrator) and the cache lookup operator (CacheLookupOperator), their specific operation processes are as follows:
[0054] When the data set loading operator is changed to perform indexing and encoding of data, the main thread of the cache lookup operator starts retrieving data row IDs from its sampler. After obtaining the required data rows from the sampler, the cache lookup operator is responsible for requesting the corresponding data rows from the cache server and submitting the data IDs with cache misses to the data loading operator of the leaf node for re-reading and parsing. Then, it sends the data stream containing placeholders to the cache merging operator. The cache merging operator continuously waits for the data stream with cache misses to be filled, replaces the placeholders with actual data, and finally passes the complete data to the upper-layer operator;
[0055] The present invention further optimizes the above steps. After the main thread retrieves the data row IDs from the sampler, the cache lookup operator does not directly distribute each data row ID to the worker threads, which then send requests to the cache server. This is because, by default, the list of data row IDs distributed to each worker thread contains only one element, which results in poor performance as each worker thread only sends a query request for one data row ID. Therefore, the main thread of the cache lookup operator starts several additional worker threads, called Prefetchers. After collecting a specified number of data row IDs, the main thread sends the list containing these IDs to each Prefetcher through an internal queue, thus allowing the Prefetcher to send a query request to the cache server;
[0056] The cache merging operator starts three groups of worker threads during runtime: cache hit stream threads, cache miss stream threads, and cleanup threads. The cache hit stream worker threads repeatedly call the cache lookup operator to retrieve data. For the two tasks of caching the original data and the preprocessed data in the cache, the cache miss stream threads will repeatedly call the data loading operator or the map operator of the leaf node to process the ID data with cache misses. Then, they asynchronously send data write requests to store the data on the cache server. The cleanup threads are responsible for waiting for the data write requests to complete and checking the results returned by the cache server.
[0057] For the CacheOperator, its specific running mode is as follows:
[0058] First, the construction phase is executed, and then the fetch phase is executed. During the construction phase, the main thread of the cache operator enables multiple worker threads, which will be blocked while waiting for each worker thread to write all data rows to the cache server. Once all data is cached on the server, the pipeline switches to the fetch phase, at which time the cache operator requests the corresponding data rows from the cache server;
[0059] Since all data rows have been cached on the server, the query hit rate of the cache server is 100%. It should be noted that in distributed training, multiple data processing pipelines may share the same cache service, but only one selected pipeline will enter the build phase, build the cache service and cache all data on the cache server, while other pipelines will only block and continuously query the current phase of the cache server until the build phase is completed.
[0060] After the data processing pipeline is started, data processing, cache writing, and cache reading operations are performed according to the execution graph. The specific steps of these two operations are as follows:
[0061] The method of the present invention is a single-node caching technology, which means that the cache client and the cache server are located on the same physical machine. If the data transfer volume is too large, the performance of using gRPC alone may not reach the ideal level. Therefore, we have optimized the performance in this case through shared memory. The cache server pre-allocates 4GB of shared space. Shared memory allows multiple processes to share the same physical memory area without additional memory copying, enabling data to be directly visible and operable in the shared memory area, thereby reducing the latency and memory overhead during data transfer.
[0062] When data needs to be written to the cache, the cache client submits an asynchronous buffer request to the gRPC queue. After the gRPC thread returns the shared memory address, the cache client continuously writes data to the asynchronous buffer. Once the buffer is full, the client submits a cache write request to the gRPC queue. Subsequently, the gRPC thread submits the cache write request to the request queue of the cache server. This cache write request will wait in the request queue for the working thread of the server to write the data from the shared memory to the cache memory. It should be noted that during the process of writing the asynchronous buffer to the cache, the cache client does not block and wait for the cache server to complete the write and return a response, but starts using the next available asynchronous buffer. This minimizes the time the cache client waits for the cache server response. Finally, the cache result is returned to the client.
[0063] When cache data needs to be read, the cache client submits a data row request to the gRPC queue. The gRPC thread locates the position of the data row in memory, then allocates shared memory for the data row, and submits a data extraction request to the request queue of the cache server. The working thread of the server writes the data from the cache memory to the shared memory. Once the data is read into the shared memory, the gRPC working thread returns the shared memory location of the data row to the cache client. After the cache client copies the data row from the shared memory to the target location, it submits a request to release the shared memory to the gRPC queue, and the gRPC thread completes the release process.
[0064] For the cache writing and cache reading operation processes, the present invention also considers optimizing memory management. The specific optimization process is as follows:
[0065] When optimizing memory management, based on the NUMA framework, according to the ratio of the memory space of the cache server to the total system memory space, a cache memory pool perceived by the NUMA framework is configured. Each cache memory pool consists of memory blocks (Arenas). The number of memory blocks is set to the number of CPUs, and memory space is allocated to the memory blocks through specific NUMA nodes. In addition, we bind each thread in the gRPC thread group and the worker thread group for processing cache services to a specific NUMA node. This method has the following advantages:
[0066] (1) When the gRPC thread processes a request, it continuously allocates memory to create a CacheServerRequest object as a label for the gRPC request. By binding to a certain node, the gRPC thread can access the local NUMA memory, avoiding memory calls to other CPUs, thereby reducing access latency.
[0067] (2) Similarly, when the worker thread of the cache server processes a cache request, we first determine the NUMA node corresponding to the worker thread and select a memory block from that node for memory allocation instead of random allocation. This ensures that the worker thread always caches data in the memory of the local NUMA node, thereby reducing access latency.
[0068] The present invention also takes into account multi-task training during AI training. In the present invention, a single cache server can serve multiple cache clients, and the cache data required by the user can be uniquely identified by the session ID of the cache client and the CRC code of the cache graph structure in the data processing pipeline. Therefore, in order to share the same cache service, the pipelines do not have to be exactly the same. As long as the same session ID is specified, the present invention is also very suitable for distributed training.
[0069] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula that is closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.
[0070] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0071] The above has described in detail an embodiment of the present invention, but the above content is only a preferred embodiment of the present invention and cannot be considered as defining the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.
Claims
1. A multi-stage caching method for accelerating AI data processing, characterized in that: The following steps are involved: Build a data processing pipeline through the top-level interface, insert cache operations in the data processing pipeline in advance and mark the position; A cache client, a cache server and a shared space are set up, and the data processing pipeline communicates through the cache client and the cache server. The cache server is used to store cache data sets, and the cache data sets include the loaded original data sets and the data sets after preprocessing. The cache client and the cache server use the gRPC framework and the shared space to exchange information; Convert the constructed data processing pipeline into a graph structure and perform node verification on the graph structure. After verification, divide the cached dataset into mappable dataset and non-mappable dataset according to whether the cached dataset supports full mapping to memory. For different datasets, use corresponding methods to convert the graph structure, generate an execution graph and optimize repeated operators. Start the data processing pipeline and perform data processing according to the generated execution graph. When the data processing pipeline performs cache writing and cache reading of the cache data set, the cache client is called to send request information to the cache server. The cache server receives the request information and performs corresponding operations. After completing the operation, it returns the corresponding information to the cache client.
2. A multi-stage caching method for accelerating AI data processing according to claim 1, characterized in that: When the gRPC framework is used between the cache client and the cache server, variables rq and reply are set for each request sent by the cache client to the cache server, wherein rq is used to record information that needs to be sent to the cache server, and reply is used to obtain response information sent from the cache server.
3. The multi-stage caching method for accelerating AI data processing according to claim 1, characterized in that: The specific process of node verification for the graph structure is as follows: Call the cache verification function, starting from the starting node of the graph structure, traverse all nodes of the entire graph structure according to the data processing path, and filter out nodes for batch processing, connection, filtering, skipping, retrieval and compression operations respectively. When there is no pre-inserted cache operation in the subsequent data processing path of any node, the verification is judged to be passed, otherwise it is judged to be a verification error, and an error report is output after the traversal is completed to remind the user to modify the insertion position of the cache operation.
4. A multi-stage caching method for accelerating AI data processing according to claim 3, characterized in that: After the judgment and verification are passed, the specific process of converting the graph structure is as follows: When the cache data set is a mappable data set, a cache merge node is created in the graph structure and inserted into the map node as the parent node of the map node. The sampler is obtained from the data loading node, and a cache search node is created based on the sampler. The sampler is replaced with the cache search node to search for and return the cached data in the cache merge node. The Build() function of all nodes in the graph structure is called to build the corresponding operator of each node in turn, generate an execution graph, and mark the operator corresponding to the cache merge node as the cache merge operator, and the operator corresponding to the cache search node as the cache search operator. When the cached data set is an unmappable data set, obtain the sampler corresponding to the data loading node belonging to the leaf node in the graph structure, create a cache node based on the sampler, and use the cache node to replace the sampler. Insert the cache node into the map node as the parent node of the map node, call the Build() function of all nodes in the graph structure, build the operator corresponding to each node runtime in turn, generate an execution graph, and mark the operator corresponding to the cache node as a cache operator.
5. A multi-stage caching method for accelerating AI data processing according to claim 4, characterized in that: The specific process of repeated operator optimization on the graph structure is as follows: Determine the type of cache data set. When the cache data set is a mappable data set, obtain the position of the repeat operator in the operation queue. If the repeat operator is located before the cache search operator and the cache merge operator, set the repeat count of the cache search operator according to the repeat count set in the repeat operator. Otherwise, set the repeat count of the cache search operator according to the repeat count set in the cache merge operator. When the cached data set is an unmappable data set, the data that is repeatedly processed by the repeated operator is marked as a cache object, and the graph structure is traversed from top to bottom. During the traversal, the cache objects are added to the cache node stack one by one. When the repeated optimization mechanism traverses the graph structure from bottom to top, the corresponding cache objects are taken out from the cache node stack one by one, and the cache object repeat count is set to 1.
6. A multi-stage caching method for accelerating AI data processing according to claim 4, characterized in that: The specific operation process of the cache merge operator and cache search operator is as follows: When the data loading operator performs data loading operations, it provides the sampler to the cache search operator to obtain the ID of the data row. After the cache search operator obtains the ID, it starts N child threads in the main thread, where N is a positive integer. After collecting a preset number of IDs, the main thread assigns the ID to the child threads through the internal queue. The child threads send cache query requests to the cache server in batches. If the ID of the cache query request hits the cache server, the data stream corresponding to the ID is sent to the cache merge operator. Otherwise, the cache search operator sends the missed ID to the data loading operator. The data loading operator re-reads the parsed data stream and sends the data stream with placeholders to the cache merge operator. The cache merge operator continuously waits for and receives data streams, merges the data streams corresponding to the hit IDs and stores them in the cache server. For data streams containing placeholders, the placeholders are replaced with actual data. After the replacement is completed, the complete data stream is stored in the cache server.
7. The multi-stage caching method for accelerating AI data processing according to claim 4, characterized in that: The specific operation process of the cache operator is as follows: The cache operator starts M sub-threads in the main thread, where M is a positive integer. Each sub-thread remains in a blocked state before successfully storing the data in the cache server and receiving the return information. When the data storage tasks of all sub-threads are completed and the return information is obtained, the cache operator switches to the acquisition state and starts reading the data stored in the cache server through the cache operator.
8. The multi-stage caching method for accelerating AI data processing according to claim 2, characterized in that: The specific process of the data processing pipeline performing cache write and cache read operations on the cache data set is as follows: When writing to the cache, the data processing pipeline calls the cache client to submit an asynchronous buffer request to the gRPC queue. After receiving the request, the gRPC queue returns a shared memory address. Then the cache client starts to continuously write the cache data set to the shared memory. When the shared memory storage is full, the buffer client submits a cache write request to the gRPC queue. The gRPC thread submits the cache write request to the request queue of the cache server. The working thread of the cache server writes the cache data set from the shared memory to the memory of the cache server in turn according to the request queue. When performing cache reading, the data processing pipeline calls the cache client to submit a data row request to the gRPC queue. After receiving the request, the gRPC queue locates the data row in the memory of the cache server and allocates shared memory for the data row. At the same time, it submits a data extraction request to the request queue of the cache server. The working thread of the cache server writes the cache data set from the memory of the cache server to the shared memory. When the cache data set is successfully read into the shared memory, the gRPC queue returns the shared memory address corresponding to the data to the cache client. The cache client copies the cache data set from the shared memory to the target location in the data processing pipeline, and submits a request to release the shared memory to the gRPC queue. The gRPC thread completes the memory release.
9. A multi-stage caching method for accelerating AI data processing according to claim 8, characterized in that: The data processing pipeline also includes optimizing memory management during the cache write and cache read operations of the cache data set. The specific optimization process is as follows: When optimizing memory management, based on the NUMA framework, according to the ratio of the memory space of the cache server and the total memory space of the system, a NUMA framework-aware cache memory pool is configured. Each cache memory pool consists of memory blocks. The number of memory blocks is set to the number of CPUs, and memory space is allocated to the memory blocks through specific NUMA nodes.
Citation Information
Patent Citations
Streaming data-oriented machine learning model online service deployment method and system
CN115756875A
Communication protocol, and a method thereof for accelerating artificial intelligence processing tasks
US11570257B1