A method, apparatus, and storage medium for caching data in blocks
Through the method of chunked data cache, the problem of insufficient memory space caused by full data loading in the prior art is solved, and the cache space is more efficiently utilized without increasing the storage space, thereby reducing hardware costs.
Patent Information
- Application Number
- CN202011515896.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-12-21
AI Technical Summary
The existing big data caching mechanism is insufficient memory space due to full data loading, and the cache operation cannot be completed. The method of increasing storage space will lead to increased hardware costs and waste of resources.
Using the method of chunking cache data, by receiving the cached data request for the computing task, multiple data chunks of the target data are determined, and after the calculation unit executes and obtains the results, the data chunks are cached and cleared one by one until all chunks are cached.
Effectively utilize existing storage space, reduce hardware costs, reduce insufficient storage space, and improve the utilization rate of cached data.
Smart Images

Figure CN112685334B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular, to a method, apparatus, and storage medium for caching data in blocks. Background Art
[0002] As is well known, data caching is an effective way to improve data access speed and is widely used in various data processing systems. In recent years, with the increasing development and popularization of network communication and computer technology, the application of big data has become more and more extensive, which also puts higher requirements on the data caching space, especially for large data processing systems with multiple cache levels.
[0003] The existing big data caching mechanism usually adopts the method of loading all data. Often, due to the overly large amount of data and insufficient memory space, the caching operation cannot be completed. In addition, when using the method of loading all data, even if the amount of data is not large, when the memory space or disk space is occupied by some key processes in large quantities and becomes insufficient, the caching operation often fails to be completed frequently.
[0004] To solve the above problems, simply solving them by adding storage space will inevitably lead to an increase in the cost of hardware construction and maintenance. Moreover, for some systems that do not have scalability and cannot increase the storage space anymore, this method may mean reconstruction, resulting in a great waste of resources.
[0005] Therefore, how to improve the data caching method to make more full use of the caching space and reduce the situation of insufficient storage space without increasing the storage space is still a technical problem to be solved. Summary of the Invention
[0006] In view of the above problems, the inventor of the present invention creatively provides a method, apparatus, and storage medium for caching data in blocks.
[0007] According to the first aspect of the embodiments of the present invention, a method for caching data in blocks includes: receiving a caching data request for a computing task, where the computing task includes a plurality of concurrently executable computing units; determining target data required by the computing task and a plurality of data blocks included in the target data, where each of the plurality of data blocks is used for the calculation of at least one of the plurality of computing units; caching one of the plurality of data blocks, and after confirming that the computing unit corresponding to the corresponding data block has been executed and a calculation result has been obtained, caching the next data block until each of the plurality of data blocks has been cached.
[0008] According to an embodiment of the present invention, before determining the target data required for the computing task and the multiple data chunks included in the target data, the method further includes: configuring the number of data chunks included in the target data; dividing the target data into multiple independent data chunks according to the number of data chunks.
[0009] According to an embodiment of the present invention, before caching one of the multiple data chunks, the method further includes: determining whether all data chunks of the target data can be cached, and if not, proceeding with the subsequent operations.
[0010] According to an embodiment of the present invention, before caching one of the multiple data chunks, the method further includes: obtaining the identifiers of each of the multiple data chunks and sorting all the identifiers of the multiple data chunks to obtain an ordered queue; correspondingly, caching one of the multiple data chunks includes: taking out the identifier of a data chunk from the ordered queue and caching the data chunk corresponding to the corresponding identifier.
[0011] According to an embodiment of the present invention, confirming that the computing unit corresponding to the corresponding data chunk is executed and obtaining a calculation result includes: during the concurrent execution of the computing unit, obtaining the reference count of the corresponding data chunk, where the reference count is the number of times the corresponding data chunk is used by all computing units; when the reference count of the corresponding data chunk is 0, it is confirmed that the computing unit corresponding to the corresponding data chunk is executed and the calculation result is obtained.
[0012] According to an embodiment of the present invention, before obtaining the reference count of the corresponding data chunk, the method further includes: analyzing each computing unit of the computing task to obtain the number of times each computing unit uses the target data; accumulating the number of times each computing unit uses the corresponding data chunk to obtain the reference count of the target data; obtaining the identifier of each of the multiple data chunks included in the target data; and recording the reference count of the target data as the reference count of the corresponding data chunk through the identifier of each data chunk.
[0013] According to an embodiment of the present invention, during the concurrent execution of the computing unit, the method further includes: if the computing unit uses the cached corresponding data chunk, decrementing the reference count of the corresponding data chunk by 1.
[0014] According to an embodiment of the present invention, before caching the next data chunk, the method further includes: clearing the corresponding data chunk from the cache.
[0015] According to a second aspect of an embodiment of the present invention, a device for caching data in blocks, the device includes: a cached data request receiving module, which receives a cached data request of a computing task, where the computing task includes a plurality of concurrently executable computing units; a target data and data block determination module, configured to determine target data required by the computing task and a plurality of data blocks included in the target data, where each of the plurality of data blocks is used for the computation of at least one of the plurality of computing units; a data block caching module, configured to cache one of the plurality of data blocks, and after confirming that the computing unit corresponding to the corresponding data block has been executed and a computation result has been obtained, clear the current data block and cache the next data block until each of the plurality of data blocks has been cached.
[0016] According to a third aspect of an embodiment of the present invention, a computer storage medium is provided, the storage medium includes a set of computer-executable instructions, which are used to execute the method for caching data in blocks as described in any one of the above when the instructions are executed.
[0017] An embodiment of the present invention provides a method, a device, and a storage medium for caching data in blocks. The method includes: after receiving a cached data request of a computing task, first determining target data required by the computing task and a plurality of data blocks included in the target data; then, each time only caching one of the data blocks in the cache, and after confirming that the computing unit corresponding to the corresponding data block has been executed and a computation result has been obtained, caching the next data block until each of the plurality of data blocks has been cached. At this time, all computing units of the above computing task have also been executed, and the subsequent computing process can be continued to obtain the final result of the computing task. In the above process of caching data in blocks, each time only one of the plurality of data blocks is cached, which requires less storage space for caching data, can make the best use of the existing storage space, reduce the hardware cost, and reduce the situation where the storage space is insufficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become easily understandable. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, where:
[0019] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.
[0020] Figure 1 It is a schematic flowchart of the implementation of the method for caching data in blocks according to an embodiment of the present invention;
[0021] Figure 2 It is a schematic flowchart of a specific implementation of the application of the method for caching data in blocks according to an embodiment of the present invention;
[0022] Figure 3 This is a schematic diagram of the component structure of the block caching data device according to an embodiment of the present invention. Detailed implementation manners
[0023] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
[0024] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without conflict, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0025] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0026] Figure 1 shows the implementation process of the method for block caching data according to an embodiment of the present invention. Refer to Figure 1 , the method includes: operation 110, receiving a cache data request for a computing task, where the computing task includes a plurality of concurrently executable computing units; operation 120, determining the target data required by the computing task and a plurality of data blocks included in the target data, where each data block in the plurality of data blocks is used for the calculation of at least one computing unit in the plurality of computing units; operation 130, caching one data block in the plurality of data blocks, and after confirming that the computing unit corresponding to the corresponding data block has been executed and the calculation result has been obtained, caching the next data block until each data block in the plurality of data blocks has been cached.
[0027] It should be noted that the main purpose of data caching is to enable the computing unit to obtain the data required for computing more quickly. Therefore, the method of caching data in chunks in the embodiments of the present invention is usually combined with the scheduling and execution process of computing tasks and carried out in coordination. Usually, different threads can be used to separately perform task scheduling and execution and data caching operations, and another main control program is set up to promote them in coordination; or the main thread and subordinate thread method can be used. Implementers can choose any applicable implementation method according to needs.
[0028] In operation 110, the cached data request may be just a triggering process or a command to initialize the cache management tool. Through this request, the cache management program can make various resource preparations for the next step of caching data, such as obtaining the corresponding storage space, etc.
[0029] The computing unit mainly refers to the smallest executable computing unit. For example, a certain function, subtask, or a certain operation, etc.
[0030] In operation 120, the target data here is usually large, so it will be divided into independent data chunks according to needs. By independent, it mainly means that there is no coupling relationship between the data chunks, and each data chunk includes at least all the data required for one calculation of the corresponding computing unit. For example, a table stores 10,000 records, and one record is needed in each calculation, then each record is the smallest unit that can be chunked. If the predefined number of chunks is 10, it can be divided into 10 data chunks, and each data chunk can include a certain number (for example, thousands to hundreds) of records. The data in each data chunk does not repeat each other and together is exactly the original 10,000 records. The data chunks required for the computing task can be informed by the computing task; or a correspondence between the computing task and the data can be maintained in advance, and the data block identifier corresponding to the task can be obtained through the task identifier at runtime; and any other feasible ways.
[0031] In operation 130, it must be ensured that all computing units that need to use this data chunk have obtained the corresponding data. Otherwise, if a computing unit fails to obtain the data and cannot complete the calculation, it may cause the entire computing task to fail to continue the subsequent calculation. When it is confirmed that the computing units corresponding to the corresponding data chunks are executed and the calculation results are obtained, the chunk data to be used by each computing unit and the execution status of the computing unit can be detected one by one, but this method will increase the corresponding operations and occupy a part of the computing resources; or the reference count of the data block can be obtained in advance, and the caching and execution of the data can be promoted in coordination through the reference count during execution; and any other feasible ways.
[0032] According to an embodiment of the present invention, before determining the target data required for a computing task and the multiple data chunks included in the target data, the method further includes: configuring the number of data chunks included in the target data; and dividing the target data into multiple independent data chunks according to the number of data chunks.
[0033] In this embodiment, the implementer can configure the number of data chunks according to the size of the storage space and other related requirements, and divide the target data into multiple independent data chunks according to the number of data chunks. In this way, the size of the chunks can be flexibly adjusted according to the implementation conditions, making the utilization rate of the cache higher.
[0034] According to an embodiment of the present invention, before caching one of the multiple data chunks, the method further includes: determining whether all data chunks of the target data can be cached; if not, then continue with the subsequent operations.
[0035] If the storage space available for cache allocation and use is large enough to hold all the data chunks, then using the traditional caching method, that is, caching the entire target data, will reduce the operation of replacing data chunks and is more efficient. Therefore, in this embodiment, a judgment will be made first. In this way, the best method for caching data can be selected according to different operating conditions.
[0036] According to an embodiment of the present invention, before caching one of the multiple data chunks, the method further includes: obtaining the identifier of each data chunk in the multiple data chunks and sorting all the identifiers of the multiple data chunks to obtain an ordered queue; correspondingly, caching one of the multiple data chunks includes: taking out the identifier of a data chunk from the ordered queue and caching the data chunk corresponding to the corresponding identifier.
[0037] In this embodiment, by sorting the identifiers of the data chunks, the data chunks can be cached in an orderly manner to ensure that each data chunk is cached without omission. At the same time, through the sorting method, it can also be used as a means of cooperative operation with the computing task, that is, the same sorting can be adopted in another thread for scheduling and executing the computing task to read the cached data.
[0038] According to an embodiment of the present invention, confirming that the computing unit corresponding to the corresponding data chunk is executed and obtaining the calculation result includes: during the concurrent execution of the computing unit, obtaining the reference count of the corresponding data chunk, where the reference count is the number of times the corresponding data chunk is used by all computing units; when the reference count of the corresponding data chunk is 0, it is confirmed that the computing unit corresponding to the corresponding data chunk is executed and the calculation result is obtained.
[0039] In this embodiment, the reference count of data chunks is obtained in advance, and at runtime, the reference count is used to coordinate the calculation process and the data caching process.
[0040] According to an embodiment of the present invention, before obtaining the reference count of the corresponding data chunk, the method further includes: analyzing each computing unit of the computing task to obtain the number of times each computing unit uses the target data; accumulating the number of times each computing unit uses the corresponding data chunk to obtain the reference count of the target data; obtaining the identifier of each data chunk included in the target data; and recording the reference count of the target data as the reference count of the corresponding data chunk through the identifier of each data chunk.
[0041] In this embodiment, by analyzing the reference count of each computing unit for the target data and accumulating it, the total reference count of the target data is obtained, and this reference count is also the reference count of each data chunk included in the target data. This analysis is usually obtained through static analysis of the call relationship between the code and the data before runtime. Implementers can use existing code analysis tools to obtain the call relationship between the code and the data, and on this basis, count the reference count of each computing unit for the target data.
[0042] According to an embodiment of the present invention, during the concurrent execution of computing units, the method further includes: if a computing unit uses the cached corresponding data chunk, then decrementing the reference count of the corresponding data chunk by 1.
[0043] During the process of controlling or scheduling the concurrent execution of each computing unit, usually some processing logic can be added after the completion of the calculation process, which can include the operation of decrementing the reference count by 1.
[0044] According to an embodiment of the present invention, before caching the next data chunk, the method further includes: clearing the corresponding data chunk from the cache.
[0045] When the reference count of the data chunk is 0, or after it has been confirmed that the computing unit corresponding to the data chunk has been executed and the calculation result has been obtained, it can basically be determined that the current computing task no longer requires the data chunk in the cache. In this embodiment, clearing the corresponding data chunk can free up more storage space for the next data chunk, further improving the utilization rate of the cache.
[0046] Figure 2 The specific implementation process diagram of an application of the method for caching data in chunks according to an embodiment of the present invention is shown. This application uses the cache management tool in the Spark platform and combines the scheduling management of the computing task to achieve the chunk caching of the data chunks of the target data.
[0047] In Spark, an Elastic Distributed Dataset (RDD) can be defined, and the dataset can be divided into multiple data chunks. The number of data chunks can be predefined according to the number of data files, the default parallelism of Spark, or the calculation output. In addition, Spark also provides a data caching management tool (Spark Cache), which is very suitable for implementing the method of caching data in chunks provided by the embodiments of the present invention.
[0048] As Figure 2 shown, the specific steps of this process include:
[0049] Step 2010, analyze the data reference count from the Directed Acyclic Graph (DAG);
[0050] In Spark, the DAG is used to model the relationship of the RDD, which describes the dependency relationship of the RDD. Therefore, by analyzing the DAG, the reference count of each data chunk can be obtained.
[0051] The reference count can be stored in a temporary variable and read or updated by the master program through parameters. The operations on the temporary variable can be defined in the corresponding operations after the logical judgment of the master program scheduling concurrent tasks.
[0052] Step 2020, start the calculation, execute multiple concurrent calculations (Tasks) according to the DAG graph, and at the same time start another cache management thread to manage and operate the cached data;
[0053] Step 2030, after the cache management thread receives the cache request, determine all the data chunks corresponding to the calculation tasks.
[0054] Step 2040, sort the Ids of the data chunks to obtain an ordered queue;
[0055] Step 2050, take out the Id of a data chunk from the ordered queue (the first time is to take the value of the first element of the queue, and then take the next element value in the queue in turn), and cache the corresponding data chunk;
[0056] Step 2060, in the main thread responsible for task scheduling and executing multiple concurrent calculations, continue the calculation process according to the DAG, including reading the data in the cache;
[0057] Step 2070, detect whether the data in the cache has cached a data chunk (for the first time) or whether the data chunk has been updated. If so, continue to step 2080. If not, wait;
[0058] Step 2080: Use the data chunks in the cache to complete corresponding calculations. If any calculation uses a data chunk in the cache, decrement the reference count by 1.
[0059] For each concurrently executed calculation, first determine the ID of the next data chunk to be used and compare it with the ID of the data chunks cached in the cache: If they are the same, retrieve the data for calculation; if they are different, block and wait to be awakened after a new data chunk becomes available.
[0060] Step 2090: Meanwhile, in the cache management thread, continuously detect whether the reference count of the data chunk is 0. If so, proceed to Step 2100; if not, wait.
[0061] Step 2100: Clear the data chunk in the cache.
[0062] Step 2110: Determine whether there are still data chunks that have not been cached (whether it is already the last element in the ordered queue). If there are still data chunks (not the last element), obtain the next data chunk and return to Step 2050. If there are no more data chunks (already the last element), end the cache management thread.
[0063] Step 2120: Meanwhile, in the main thread, after using the data block in the cache to complete the corresponding calculation, detect whether there are still calculations that have not been completed. If so, return to Step 2060 to continue reading the data in the cache. If not, end the calculation task and return the calculation result.
[0064] It should be noted that the specific implementation process of the above Embodiment 1 is only for illustrative purposes and does not limit the implementation manner or application scenario of the embodiments of the present invention. Implementers can adopt any applicable implementation manner according to specific implementation conditions and apply it to any applicable application scenario.
[0065] Furthermore, an embodiment of the present invention also provides a device for caching data in chunks, as Figure 3 shown. The device 30 includes: a cache data request receiving module 301, which receives a cache data request for a calculation task, where the calculation task includes multiple concurrently executable calculation units; a target data and data chunk determination module 302, which is used to determine the target data required for the calculation task and the multiple data chunks included in the target data, where each data chunk in the multiple data chunks is used for the calculation of at least one calculation unit among the multiple calculation units; a data chunk caching module 303, which is used to cache one data chunk among the multiple data chunks, and after confirming that the calculation unit corresponding to the data chunk has been executed and the calculation result has been obtained, clear the current data chunk and cache the next data chunk until each data chunk among the multiple data chunks has been cached.
[0066] According to an implementation manner of Embodiment 1 of the present invention, the apparatus 30 further includes: a configuration module for the number of data chunks, configured to configure the number of data chunks included in the target data; a data chunk division module, configured to divide the target data into a plurality of mutually independent data chunks according to the number of data chunks.
[0067] According to an implementation manner of Embodiment 1 of the present invention, the apparatus 30 further includes a data cache load judgment module, configured to judge whether all data chunks of the target data can be cached. If not, the subsequent operations are continued.
[0068] According to an implementation manner of Embodiment 1 of the present invention, the apparatus 30 further includes a data chunk sorting module, configured to obtain the identifier of each data chunk in the plurality of data chunks and sort all the identifiers of the plurality of data chunks to obtain an ordered queue; correspondingly, the data chunk caching module is specifically configured to take out the identifier of a data chunk from the ordered queue and cache the data chunk corresponding to the corresponding identifier.
[0069] According to an implementation manner of Embodiment 1 of the present invention, the data chunk caching module 303 includes: a reference count obtaining sub-module, configured to obtain the reference count of the corresponding data chunk during the concurrent execution of the computing units, where the reference count is the number of times the corresponding data chunk is used by all computing units; a computing unit execution status judgment sub-module, configured to confirm that the computing unit corresponding to the corresponding data chunk has been executed and obtain a computing result when the reference count of the corresponding data chunk is 0.
[0070] According to an implementation manner of Embodiment 1 of the present invention, the apparatus 30 further includes a data chunk reference count statistics module, configured to analyze each computing unit of the computing task to obtain the number of times each computing unit uses the target data; accumulate the number of times each computing unit uses the corresponding data chunk to obtain the reference count of the target data; obtain the identifier of each data chunk included in the target data; and record the reference count of the target data as the reference count of the corresponding data chunk through the identifier of each data chunk.
[0071] According to an implementation manner of Embodiment 1 of the present invention, the apparatus 30 further includes a reference count update module, configured to decrement the reference count of the corresponding data chunk by 1 if the computing unit uses the cached corresponding data chunk.
[0072] According to an implementation manner of Embodiment 1 of the present invention, the apparatus 30 further includes a data clearing module, configured to clear the corresponding data chunk from the cache.
[0073] According to the third aspect of the embodiments of the present invention, there is provided a computer storage medium, where the storage medium includes a set of computer executable instructions, which are used to execute the method for caching data in chunks as described in any one of the above when the instructions are executed.
[0074] It should be noted here that the descriptions of the device embodiments for block caching data above and the descriptions of the computer storage medium embodiments above are similar to the descriptions of the foregoing method embodiments and have beneficial effects similar to those of the foregoing method embodiments. Therefore, they will not be elaborated. For the technical details not disclosed in the descriptions of the device embodiments for block caching data and the computer storage medium embodiments of the present invention, please refer to the descriptions of the foregoing method embodiments of the present invention for understanding. For the sake of saving space, they will not be elaborated here.
[0075] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.
[0076] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another device, or some features can be ignored or not executed. In addition, the couplings, direct couplings, or communication connections between the components shown or discussed may be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical or other forms.
[0077] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they may be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0078] In addition, each functional unit in the embodiments of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit; the above integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.
[0079] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments. The aforementioned storage medium includes various media that can store program codes, such as removable storage media, read-only memory (ROM), magnetic disks, or optical discs.
[0080] Alternatively, if the above integrated units of the present invention are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media that can store program codes, such as removable storage media, ROM, magnetic disks, or optical discs.
[0081] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for block-caching data, the method comprises: Receiving a cache data request for a computing task, where the computing task includes multiple concurrently executable computing units; Determining target data required by the computing task and multiple data blocks included in the target data, where each data block in the multiple data blocks is for the computation of at least one computing unit among the multiple computing units; Caching one data block among the multiple data blocks, and after confirming that the computing unit corresponding to the corresponding data block has been executed and obtained a computation result, caching the next data block until each data block among the multiple data blocks has been cached; The confirmation that the computing unit corresponding to the corresponding data block has been executed and obtained a computation result includes: during the concurrent execution of the computing units, obtaining the reference count of the corresponding data block, where the reference count is the number of times the corresponding data block is used by all computing units; if the computing unit uses the cached corresponding data block, then decrement the reference count of the corresponding data block by 1, and when the reference count of the corresponding data block is 0, it is confirmed that the computing unit corresponding to the corresponding data block has been executed and obtained a computation result; Before obtaining the reference count of the corresponding data block, the method further includes: analyzing each computing unit of the computing task to obtain the number of times each computing unit uses the target data; accumulating the number of times each computing unit uses the corresponding data block to obtain the reference count of the target data; obtaining the identifier of each data block included in the target data; and recording the reference count of the target data as the reference count of the corresponding data block through the identifier of each data block.
2. The method according to claim 1, before determining the target data required by the computing task and the multiple data blocks included in the target data, the method further comprises: Configuring the number of data blocks included in the target data; Dividing the target data into multiple mutually independent data blocks according to the number of data blocks.
3. The method according to claim 1, before caching one data block among the multiple data blocks, the method further comprises: Judging whether all data blocks of the target data can be cached, if not, then continuing with the subsequent operations.
4. The method according to claim 1, before caching one data block among the multiple data blocks, the method further comprises: Obtaining the identifier of each data block among the multiple data blocks and sorting all the identifiers of the multiple data blocks to obtain an ordered queue; Correspondingly, caching one data block among the multiple data blocks includes: Taking out the identifier of a data block from the ordered queue and caching the data block corresponding to the corresponding identifier.
5. The method according to claim 1, before caching the next data block, the method further comprises: Clearing the corresponding data block from the cache.
6. A device for block-caching data, the device comprises: Cache data request receiving module, which receives cache data requests for computing tasks, where the computing tasks include multiple concurrently executable computing units; Target data and data block determination module, which is used to determine the target data required by the computing task and the multiple data blocks included in the target data, where each of the multiple data blocks is used for the computation of at least one of the multiple computing units; Data block caching module, which is used to cache one of the multiple data blocks. After confirming that the computing unit corresponding to the corresponding data block has been executed and the computation result has been obtained, the current data block is cleared, and the next data block is cached until each of the multiple data blocks has been cached; The confirmation that the computing unit corresponding to the corresponding data block has been executed and the computation result has been obtained includes: during the concurrent execution of the computing unit, obtaining the reference count of the corresponding data block, where the reference count is the number of times the corresponding data block is used by all computing units; if the computing unit uses the cached corresponding data block, the reference count of the corresponding data block is decremented by 1, and when the reference count of the corresponding data block is 0, it is confirmed that the computing unit corresponding to the corresponding data block has been executed and the computation result has been obtained; Before obtaining the reference count of the corresponding data block, it further includes: analyzing each computing unit of the computing task to obtain the number of times each computing unit uses the target data; accumulating the number of times each computing unit uses the corresponding data block to obtain the reference count of the target data; obtaining the identifier of each data block included in the target data; and recording the reference count of the target data as the reference count of the corresponding data block through the identifier of each data block.
7. A computer storage medium, on which program instructions are stored, wherein, the program instructions are used to execute the method for caching data in blocks as described in any one of claims 1 to 5 when running.
Citation Information
Patent Citations
Data storage method and storage equipment
CN108829613A