Data stratification method and system for storage-computing integrated architecture for computing-intensive systems
By adopting the integrated storage and computing architecture data hierarchy method in computing-intensive systems, the problem of large power consumption of storage walls and memory access is solved, and the performance optimization of hybrid media storage systems and effective management of hot and cold data is achieved.
Patent Information
- Application Number
- CN202411358136.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-09-27
AI Technical Summary
When facing computing-intensive systems, the prior art is difficult to effectively solve the problem of large power consumption of storage walls and memory access, and the hardware implementation cost of hybrid memory management is high and has low flexibility.
The data hierarchy method of storage and computing integrated architecture for computing-intensive systems is adopted. By setting up a control module to redirect concurrent read and write tasks, the file pre-retrieval module is set to distinguish data attributes according to the file type, and data migration and storage are carried out based on these attributes.
The overall performance of hybrid media storage systems under computing-intensive tasks is optimized, the performance overhead brought by data migration is reduced, and the dynamic identification and management of hot and cold data is improved.
Smart Images

Figure CN119293010B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data stratification and hybrid storage systems, and in particular to a data stratification method and system for a storage-computing integrated architecture for computing-intensive systems. Background Art
[0002] Among the existing mainstream technical solutions, high-bandwidth data transmission technologies such as optical interconnection have provided limited relief to the intensive memory access problem in terms of transmission rate, but their functions and performance are still far from application requirements and have high procurement costs. Although near-storage computing has alleviated the storage wall problem, it has not fundamentally eliminated the storage wall. The volatile storage SRAM / DRAM technology in storage and computing integration is mature, but it affects the operating speed of the processor, and it is difficult to have a good compromise between performance and capacity, and there is a bottleneck problem. The characteristics of non-volatile storage are more suitable for the implementation of storage and computing integration, but in specific applications, the implementation of hybrid memory management through hardware will bring high costs and low flexibility.
[0003] Therefore, the prior art still has defects. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a data stratification method and system for a storage-computation integrated architecture for computing-intensive systems in view of the above-mentioned defects of the prior art. The technical solution adopted by the present invention is as follows:
[0005] In a first aspect, the present invention provides a data stratification method for a storage-computation integrated architecture for a computing-intensive system, wherein the method comprises:
[0006] A data stratification method for storage-computation integrated architecture of a computing-intensive system, characterized in that the method comprises:
[0007] Setting a control module, and redirecting part of the file data corresponding to the concurrent read and write tasks based on the control module;
[0008] Setting a file pre-retrieval module, based on which the file pre-retrieval module distinguishes data attributes of the file data according to the file type, wherein the data attributes include cold data and hot data;
[0009] The file data is migrated and stored based on the determined data attributes.
[0010] In one implementation, the method further comprises:
[0011] Set load thresholds;
[0012] Determine the load of the device, and when the load of the device exceeds the load threshold, add the excess load to the storage module of the next level;
[0013] When the cache hit rate of a device reaches the maximum performance, the data inflow of the device is reduced.
[0014] In one implementation, it is characterized in that redirecting part of the file data corresponding to the concurrent read and write tasks based on the control module includes:
[0015] A scheduler is set, and based on the scheduler, when the concurrency level of the read and write tasks reaches a threshold, part of the file data in the non-volatile random access memory is redirected to the solid state drive.
[0016] In one implementation, it is characterized in that the non-volatile random access memory includes a write operation cache and a read operation cache, the write operation cache is used to detect the task queue status, and the read operation cache is used to avoid log record reconstruction operations caused by file data redirection.
[0017] In one implementation, it is characterized in that a submodule including a file pre-reading program is provided in the scheduler, which is used to distinguish the hotness and coldness of data and make reasonable feedback to the scheduler according to whether the read-write cache area hits.
[0018] In one implementation, it is characterized in that the file pre-reading program is used to generate a file pre-reading information block, and the file pre-reading information block includes: file feature parameters, read and write operation hit parameters, file IO access number parameters, file address and file mapping information.
[0019] In one implementation, the method further includes:
[0020] The NOVA file system is used as a log-structured file system to provide consistency between metadata and file data when files are updated during copy-on-write, and to enable the main control unit to bypass the dynamic random access memory to directly access the non-volatile main memory.
[0021] In a second aspect, an embodiment of the present invention further provides a data tiering system with a storage-computing integrated architecture for a computing-intensive system, wherein the system comprises:
[0022] A control module, used to redirect part of the file data corresponding to the concurrent read and write tasks;
[0023] A file pre-retrieval module, used to distinguish data attributes of the file data according to the file type, wherein the data attributes include cold data and hot data;
[0024] The data migration and storage module is used to migrate and store the file data based on the determined data attributes.
[0025] In the third aspect, an embodiment of the present invention further provides a terminal, wherein the terminal includes a memory, a processor, and a storage-computing integrated architecture data stratification program for computing-intensive systems stored in the memory and executable on the processor, and when the processor executes the storage-computing integrated architecture data stratification program for computing-intensive systems, the steps of the storage-computing integrated architecture data stratification method for computing-intensive systems of any one of the above-mentioned schemes are implemented.
[0026] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein a data tiering program for a storage-computing integrated architecture for a computing-intensive system is stored on the computer-readable storage medium, and when the data tiering program for a storage-computing integrated architecture for a computing-intensive system is executed by a processor, the steps of the data tiering method for a storage-computing integrated architecture for a computing-intensive system described in any one of the above-mentioned schemes are implemented.
[0027] Beneficial effects: Compared with the prior art, the present invention provides a data tiering method for a storage-computation integrated architecture for computing-intensive systems. The present invention first sets a control module, and redirects part of the file data corresponding to concurrent read and write tasks based on the control module. Then, a file pre-retrieval module is set, and based on the file pre-retrieval module, the data attributes of the file data are distinguished according to the file type, and the data attributes include cold data and hot data. Finally, the file data is migrated and stored based on the determined data attributes. The present invention divides the data into cold and hot layers, and on the basis of the stratification of the capacity layer and the performance layer, optimizes the cache data scheduling in high-concurrency multi-threaded task scenarios, and is suitable for computing-intensive systems. It optimizes and improves the traditional data tiering architecture, can optimize the overall performance of the hybrid media storage system under computing-intensive tasks, can dynamically identify cold and hot data, and reduce the performance overhead caused by data migration. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is a schematic diagram of the principles of hot and cold data migration and storage.
[0029] Figure 2 A flowchart of a preferred embodiment of a data stratification method for a storage-computing integrated architecture for computing-intensive systems provided in an embodiment of the present invention.
[0030] Figure 3 The performance characteristics of various storage media as the number of threads increases.
[0031] Figure 4 A schematic diagram of a scheduler in a data tiering method for a storage-computing integrated architecture for computing-intensive systems provided in an embodiment of the present invention.
[0032] Figure 5A schematic diagram of the generation and function of file information blocks in a data stratification method for a storage-computing integrated architecture for computing-intensive systems provided in an embodiment of the present invention.
[0033] Figure 6 A schematic diagram of the architecture of a data stratification device with a storage-computing integrated architecture for computing-intensive systems provided in an embodiment of the present invention.
[0034] Figure 7 A functional block diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0035] In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0036] There are several solutions to the problems of storage wall and high power consumption of memory access caused by current intensive computing and intensive memory access:
[0037] Optical interconnect technology: High-speed data transmission and reduced power consumption are achieved through optical interconnect technology.
[0038] Hybrid computing storage: 2.5D / 3D stacking technology allows multiple chips to be stacked together, increasing storage density, expanding storage capacity, and increasing storage bandwidth by increasing parallel width or using serial transmission.
[0039] Near-storage computing: Computer systems commonly use a multi-level storage hierarchy and high-density on-chip storage to alleviate memory access latency and power consumption.
[0040] Computational storage. The storage algorithm is embedded in the memory particle, so that the storage unit has computing function. Data does not need a separate computing component to complete the calculation, but is stored and calculated in the storage unit, eliminating data access delay and power consumption. In the selection of storage media, there are currently phase change storage (PCM) based on non-volatile storage, resistive memory ReRAM / memristor, floating gate devices and flash memory FLASH and SRAM and DRAM based on volatile storage.
[0041] As the gap between processor and memory process improvement increases, the scissors fork of the memory wall under the von Neumann architecture continues to increase, and the problem of the memory access power consumption wall becomes increasingly prominent. The industry and academia have begun to shift their focus from computing to storage. Today, artificial intelligence is developing rapidly. The amount of data processed by the neural network streaming algorithm in the field of deep learning is very large, while the calculation in the artificial intelligence algorithm does not require high precision. In recent years, the academic and industrial circles have continued to invest, and non-volatile memory has developed rapidly, such as phase change memory PCM and resistive random access memory RRAM. Due to the natural integration of non-volatile memory for computing and storage, it is moving towards the integration of storage and computing based on non-volatile memory, and building a storage system based on non-volatile memory. However, due to current technical and cost issues, computational storage chips are challenging. Optane PMem based on PCM has been used to some extent, using hybrid storage media to meet the intensive memory access requirements of neural networks in artificial intelligence and reduce memory power consumption.
[0042] For storage system design, tiering and caching are two important concepts. Computer storage systems use the principle of locality to reduce the gap in data access performance among storage media with different characteristics (bandwidth, latency, capacity, power consumption, cost, etc.) in the form of caching. Tiering can improve the flexibility of storage needs, optimize data management, and reduce the total cost of ownership. There is some frequently accessed data, called hot data, and some infrequently accessed data, called cold data. The hot and cold data are distinguished by algorithms, and the hot data is placed in the cache. While calculating, data with high access frequency is migrated to media with high read and write performance according to statistical data attributes such as the heat of IO data, such as Figure 1 The high-performance data layer in the storage system is used, and the data with low access frequency is migrated to the low-performance media for storage, such as Figure 1 Since hot and cold data are dynamic, their attributes will change, resulting in overhead caused by frequent migration. Therefore, algorithm strategies are needed to dynamically identify hot and cold data and reduce their performance overhead.
[0043] This embodiment provides a data stratification method for a storage-computing integrated architecture for computing-intensive systems. The method can be applied to a terminal system, which can be a computer terminal system or other intelligent product terminal system. Figure 2 As shown in the figure, the data stratification method for storage-computing integrated architecture for computing-intensive systems of this embodiment includes the following steps:
[0044] Step S100: Setting a control module, and redirecting part of the file data corresponding to the concurrent read and write tasks based on the control module.
[0045] Step S200: Setting a file pre-retrieval module, based on which the file pre-retrieval module distinguishes data types of the file data according to file types, the data types including cold data and hot data;
[0046] Step S300: Migrate and store the file data based on the determined data attributes.
[0047] At present, hybrid storage systems are composed of the following classic storage media: hard disk drives (HDDs), solid-state drives (SSDs), dynamic random access memories (DRAMs), and non-volatile main memories (NVMMs) that have become increasingly popular in recent years. Optane NVRAM is a widely used non-volatile random access memory with read and write latency close to that of DRAM, but its write throughput is very slow and inversely proportional to the number of threads, which means negative scalability under multi-threading. For hybrid storage systems under compute-intensive tasks, the problem of slow write throughput will further affect concurrent read operations, and the read and write operations when new data is generated will reduce overall performance. That is, new decisions need to be made to optimize and improve system performance in compute-intensive task multi-threading scenarios. For example Figure 3 As shown in , existing storage media have similar performance characteristics for different concurrencies, that is, as the number of threads increases. In the scenario of computationally intensive tasks, the performance of two adjacent layers of the storage system is similar, and there is a performance loss in managing it using traditional caching and data tiering methods.
[0048] In order to optimize the overall performance of the hybrid media storage system under computing-intensive system tasks, the focus is on optimizing the performance loss of read and write operations under NVM and SSD media. This embodiment makes two improvements. The first is to set a control module, based on which part of the file data corresponding to the concurrent read and write tasks is redirected. The second is to set a file pre-retrieval module, based on which the file pre-retrieval module distinguishes the data attributes of the file data according to the file type, and the data attributes include cold data and hot data, so as to migrate and store the file data based on the determined data attributes. This embodiment can optimize the cache allocation of read and write operations of frequently accessed data and reduce the performance overhead caused by data migration.
[0049] Specifically, this embodiment first models the cache throughput, and the expected delay of a single request is T cache,1 =H·T hit + ( 1-H)·T miss , where H is the cache hit rate, T hit is the hit delay, T miss is the miss latency. Concurrent bandwidth when caching Among them, R lo and R hiThe processing speed of the capacity layer and the performance layer. The inverse of the average time of each request is used to calculate the bandwidth. The model data shows that the cache performance is limited by the bandwidth of the slow device and it is difficult to play the performance of the fast device. The traditional cache method is not an effective method in computing-intensive scenarios, that is, concurrency and high cache hit rate. Too many requests may be concentrated on the same device, resulting in too many concurrent tasks.
[0050] To this end, this embodiment also sets a load threshold, which is an upper limit parameter of the performance of the storage system device, including the maximum possible performance parameters provided by each device. This embodiment determines the load of the device, and when the load of the device exceeds the load threshold, the excess load is added to the SSD storage module of the next level. And when the cache hit rate of the device reaches the maximum performance, the data inflow of the device is reduced. In other words, when the device is fully loaded, the data inflow replacement between the two devices is controlled. After the performance of a device reaches the upper limit parameter, further increasing the load does not reduce its processing capacity. Through feedback from specific working conditions, for the hybrid storage system of Optane SSD and NVMM, the write performance is significantly smaller than the read performance difference.
[0051] Since the existing cache block is based on Optane NVRAM, it is not connected to the processing module through the PCIe bus like SSD, and the storage unit is accessed according to the memory address. Optane DIMM is connected to the memory bus. When the number of threads increases and Optane DIMM is accessed concurrently, it brings about the space competition problem of the write operation cache block, resulting in the accumulation of tasks in the queue. The write cache area is divided in NVM and redesigned on the I / O stack to improve system efficiency.
[0052] Based on this, Figure 4 As shown, this embodiment sets a scheduler, and based on the scheduler, when the concurrency level of the read and write tasks reaches a threshold, part of the file data in the non-volatile random access memory (NVRMM) is redirected to the solid state drive. The non-volatile random access memory includes a write operation cache area and a read operation cache area, the write operation cache area is used to detect the task queue status and cache hot data to reduce the overhead of frequent rewriting, and the read operation cache area is used to avoid the log record reconstruction operation caused by the redirection of file data.
[0053] Based on the cache concurrency bandwidth and the basic performance of the storage medium mentioned above, when the task concurrency level is high (generally when the number of threads is greater than 16), feedback is given to the scheduler, and the scheduler redirects part of the file data in NVRM to the SSD, with a data redirection split rate of 60%. Whenever the workload locality changes, the optimization process will be restarted, the scheduler will re-enter the initial state, and the current read and write request data will be collected and allocated.
[0054] In addition, the following issues should be considered in practical applications. The first is the mobile overhead of data redirection; the second is the consistency problem of multiple processes accessing the same file under high concurrency. Currently common file systems suitable for NVM / SSD hybrid storage include NOVA, Ext4 (Ext4-DAX), XFS, BPFS, etc. This embodiment uses the NOVA file system as a log-structured file system to provide consistency between metadata and file data when copy-on-write is used for file updates, and enables the main control unit to bypass the dynamic random access memory to directly access the non-volatile main memory, thereby avoiding the overhead caused by temporary data copying.
[0055] While calculating, this embodiment migrates the file data with high access frequency to the medium with high read / write performance based on the determined data attributes, and migrates the file data with low access frequency to the medium with low performance. In order to reduce the overhead problem caused by the frequent rewriting of hot data and the data migration, the scheduler includes a submodule of the file pre-reading program, which is used to distinguish the hotness of the data and make reasonable feedback to the scheduler based on whether the read / write cache area hits. Figure 5 As shown, the file pre-reading program generates a file pre-reading information block (Predict Block) through the file, including five main parameters: file feature parameter file_feature, read and write operation hit parameter cache_info, file IO access number parameter; I / O_history, file address file_location and file mapping information file_mapping. Among them, the first 16 bytes of metadata information are used as file feature parameter file_feature. The read and write operation hit parameter cache_info is a Boolean value. The file pre-reading program generates a file pre-reading information block through file data and transmits it to the scheduler. The scheduler schedules the LRU algorithm (Least recently used) through the file feature parameter and the IO access number parameter. When migrating and storing hot and cold data, the hot data is directly allocated to the NVRAM, and is no longer dynamically redirected according to the number of concurrent tasks. Some hot data and cold data that cause the NVRAM load to be too large and exceed the threshold are allocated to the SSD.
[0056] In summary, this embodiment optimizes the reduction of data movement and the performance and power consumption overhead caused by it in computing-intensive systems, and considers the read-write cache hits and data hotness and coldness issues in the storage system design, redirects part of the data, optimizes the cache strategy in multi-threaded scenarios, and designs a storage system that combines a hybrid storage medium of non-volatile storage and volatile storage. It also adopts a reasonable file system considering its application scenarios and storage characteristics, thereby ensuring the atomicity of operations and the consistency of file data.
[0057] Based on the above embodiments, the present invention also provides a data tiering system with a storage-computing integrated architecture for computing-intensive systems. Figure 6 As shown in, the system includes: a control module 10, the file pre-retrieval module 20 and a data migration and storage module 30. The control module 10 is used to redirect part of the file data corresponding to the concurrent read and write tasks. The file pre-retrieval module 20 is used to distinguish the data attributes of the file data according to the file type, and the data attributes include cold data and hot data. The data migration and storage module 30 is used to migrate and store the file data based on the determined data attributes.
[0058] In one implementation, the system further includes:
[0059] A load threshold setting module, and setting the load threshold;
[0060] A load transfer module, used to determine the load of the capacity device, and when the load of the capacity device exceeds the load threshold, add the excess load to the storage module of the next level;
[0061] The data inflow control module is used to reduce the data inflow of the capacity device when the cache hit rate of the capacity device reaches the maximum performance.
[0062] In one implementation, the control module 10 includes:
[0063] The scheduler is used to redirect part of the file data in the non-volatile random access memory to the solid-state drive when the concurrency level of the read and write tasks reaches a threshold.
[0064] In one implementation, the non-volatile random access memory includes a write operation cache and a read operation cache, the write operation cache is used to detect the task queue status, and the read operation cache is used to avoid log record reconstruction operations caused by file data redirection.
[0065] In one implementation, a submodule including a file pre-reading program is provided in the scheduler, which is used to distinguish the hotness and coldness of data and to provide reasonable feedback to the scheduler according to whether the read-write cache area hits.
[0066] In one implementation, the file pre-reading program is used to generate a file pre-reading information block, which includes: file feature parameters, read and write operation hit parameters, file IO access number parameters, file address and file mapping information.
[0067] In one implementation, the system further includes:
[0068] The file system setting module is used to adopt the NOVA file system as a log-structured file system to provide consistency between metadata and file data when files are updated during copy-on-write, and to enable the main control unit to bypass the dynamic random access memory to directly access the non-volatile main memory.
[0069] The working principles of each module in the storage-computing integrated architecture data tiering device for computing-intensive systems in this embodiment are the same as the principles of each step in the above method embodiment, and will not be repeated here.
[0070] Based on the above embodiment, the present invention further provides a terminal, the principle block diagram of the terminal can be as follows: Figure 7 The terminal may include one or more processors 100 ( Figure 7 Only one is shown), a memory 101 and a computer program 102 stored in the memory 101 and executable on one or more processors 100, for example, a data tiering program for a storage-computing integrated architecture for a computing-intensive system. When one or more processors 100 execute the computer program 102, each step in the embodiment of the method for data tiering for a storage-computing integrated architecture for a computing-intensive system can be implemented. Alternatively, when one or more processors 100 execute the computer program 102, the functions of each module / unit in the embodiment of the method for data tiering for a storage-computing integrated architecture for a computing-intensive system can be implemented, which is not limited here.
[0071] In one embodiment, the processor 100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0072] In one embodiment, the memory 101 may be an internal storage unit of an electronic device, such as a hard disk or memory of the electronic device. The memory 101 may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 101 may also include both an internal storage unit of the electronic device and an external storage device. The memory 101 is used to store computer programs and other programs and data required by the terminal. The memory 101 may also be used to temporarily store data that has been output or is to be output.
[0073] Those skilled in the art will understand that Figure 7 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the scheme of the present invention, and does not constitute a limitation on the terminal to which the scheme of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0074] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, operating database or other media used in the embodiments provided by the present invention may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double operational data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data stratification method for storage-computation integrated architecture for computing-intensive systems, characterized in that: The method comprises: Setting a control module, and redirecting part of the file data corresponding to the concurrent read and write tasks based on the control module; Setting a file pre-retrieval module, based on which the file pre-retrieval module distinguishes data attributes of the file data according to the file type, wherein the data attributes include cold data and hot data; Migrating and storing the file data based on the determined data attributes; The method further comprises: Setting a load threshold, where the load threshold is an upper limit parameter of the storage system device performance, including the maximum possible performance parameter provided by each device; Determine the load of the device, and when the load of the device exceeds the load threshold, add the excess load to the storage module of the next level; When the cache hit rate of the device reaches the maximum performance, reducing the amount of data flowing into the device; The redirecting of part of the file data corresponding to the concurrent read and write tasks based on the control module includes: Setting a scheduler, based on the scheduler, redirecting part of the file data in the non-volatile random access memory to the solid state drive when the concurrency level of the read and write tasks reaches a threshold; Whenever the workload locality changes, the optimization process is restarted, the scheduler re-enters the initial state, and the current read and write request data is collected and allocated; The non-volatile random access memory includes a write operation cache and a read operation cache, wherein the write operation cache is used to detect the task queue status, and the read operation cache is used to avoid the log record reconstruction operation caused by the file data redirection; the scheduler is provided with a submodule including a file pre-reading program, which is used to distinguish the hotness and coldness of the data and make reasonable feedback to the scheduler according to whether the read and write caches are hit; The file pre-reading program is used to generate a file pre-reading information block, and the file pre-reading information block includes: file feature parameters, read and write operation hit parameters, file IO access number parameters, file address and file mapping information; The file pre-reading program generates a file pre-reading information block through file data and transmits it to the scheduler. The scheduler schedules the LRU algorithm through file feature parameters and file IO access number parameters. When migrating and storing hot and cold data, the hot data is directly allocated to the non-volatile random access memory, and is no longer dynamically redirected according to the number of concurrent tasks. Some hot data and cold data that cause the non-volatile random access memory to be overloaded and exceed the threshold are allocated to the solid-state hard disk; The method further comprises: The NOVA file system is used as a log-structured file system to provide consistency between metadata and file data when files are updated during copy-on-write, and to enable the main control unit to bypass the dynamic random access memory to directly access the non-volatile main memory.
2. A data tiering system for storage and computing integrated architecture for computing-intensive systems, the system being used to implement the steps of the data tiering method for storage and computing integrated architecture for computing-intensive systems as described in claim 1, characterized in that: The system comprises: A control module, used to redirect part of the file data corresponding to the concurrent read and write tasks; A file pre-retrieval module, used to distinguish data attributes of the file data according to the file type, wherein the data attributes include cold data and hot data; The data migration and storage module is used to migrate and store the file data based on the determined data attributes.
3. A terminal, characterized in that: The terminal includes a memory, a processor, and a storage-computing integrated architecture data stratification program for computing-intensive systems, which is stored in the memory and can be run on the processor. When the processor executes the storage-computing integrated architecture data stratification program for computing-intensive systems, the steps of the storage-computing integrated architecture data stratification method for computing-intensive systems as described in claim 1 are implemented.
4. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a data stratification program for a storage-computing integrated architecture for a computing-intensive system. When the data stratification program for a storage-computing integrated architecture for a computing-intensive system is executed by a processor, the steps of the data stratification method for a storage-computing integrated architecture for a computing-intensive system as described in claim 1 are implemented.
Citation Information
Patent Citations
Consistent hash-based hierarchical mixed storage system and method
CN107844269A
Hybrid storage method and system for data layout and scheduling based on segment mapping
CN111078143A