Data Storage and Caching Optimization Method and System Based on Alluxio

By calculating file priorities on Alluxio and formulating transmission strategies, as well as intelligent prefetching and replacement strategies, the problem of low data access efficiency in multi-data center environments is solved, and the I/O performance and overall performance of the system are improved.

CN119200982BActive Publication Date: 2025-05-27SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411264707.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2025-05-27
Estimated Expiration
2044-09-10

AI Technical Summary

Technical Problem

In a multi-data center environment, due to geographical distribution, network latency and bandwidth limitations, data access across data centers is low, affecting the I/O performance of the system.

Method used

Using Alluxio-based data storage and cache optimization methods, by calculating the priority of files and formulating transmission policies based on priority, high-priority files can be obtained higher resource priority in asynchronous storage, and reducing the frequency of remote calls across data centers through intelligent prefetching and replacement strategies.

Benefits of technology

Improves data writing efficiency, reduces the risk of data loss, improves the I/O performance and overall performance of the system, and enhances the availability of Alluxio in a multi-data center environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119200982B_ABST
    Figure CN119200982B_ABST
Patent Text Reader

Abstract

The present invention discloses a data storage and caching optimization method and system based on Alluxio. The method includes: obtaining a file to be written, and calculating the priority of the file according to four attributes of the file to be written; the four attributes of the file to be written include: file access frequency, file capacity, file importance level, and file freshness; storing the file to be written into the Alluxio cache and the underlying storage according to the file transfer policy corresponding to the file priority; obtaining data to be read, and determining whether the data to be read is in the cache. If so, returning the data in the cache to the user; if not, returning the data in the underlying storage to the user; performing a prefetch operation on the data to be read and the data blocks associated with the data to be read according to the association rule. If the cache utilization rate exceeds a set threshold, a cache replacement operation is executed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data storage and cache optimization, and particularly to a data storage and cache optimization method and system based on Alluxio. Background Art

[0002] In the current architectures of big data and cloud computing, data centers are usually distributed in multiple geographical locations. However, the data access and storage efficiency in a multi-data center environment is often affected by geographical distribution, network latency, and bandwidth limitations between data centers. Especially in the cross-data center scenario, due to the long physical distance between different data centers, network latency becomes an important bottleneck affecting the system I / O performance. The synchronization and data transfer speed between data centers are much slower than those within the same data center. Such latency and bandwidth limitations significantly reduce the data access speed, thereby affecting the overall system performance. Summary of the Invention

[0003] To solve the deficiencies of the prior art, the present invention provides a data storage and cache optimization method and system based on Alluxio;

[0004] On the one hand, a data storage and cache optimization method based on Alluxio is provided, including:

[0005] Obtain a file to be written, and calculate the priority of the file according to four attributes of the file to be written; the four attributes of the file to be written include: file access frequency, file capacity, file importance level, and file freshness;

[0006] Store the file to be written into the Alluxio cache and the underlying storage according to the file transfer policy corresponding to the file priority;

[0007] Obtain data to be read, and determine whether the data to be read is in the cache. If so, return the data in the cache to the user; if not, return the data in the underlying storage to the user;

[0008] Perform a prefetch operation on the data to be read and the data blocks associated with the data to be read according to the association rule. If the cache usage rate exceeds a set threshold, perform a cache replacement operation.

[0009] On the other hand, a data storage and cache optimization system based on Alluxio is provided, including:

[0010] An obtaining module, which is configured to: obtain a file to be written, and calculate the priority of the file according to four attributes of the file to be written; the four attributes of the file to be written include: file access frequency, file capacity, file importance level, and file freshness;

[0011] A writing module, which is configured to: store the file to be written into the Alluxio cache and the underlying storage according to the file transfer policy corresponding to the file priority;

[0012] A reading module, which is configured to: obtain the data to be read, determine whether the data to be read is in the cache, and if so, return the data in the cache to the user; if not, return the data in the underlying storage to the user;

[0013] A prefetching module, which is configured to: perform a prefetching operation on the data to be read and the data blocks associated with the data to be read according to the association rule, and if the cache utilization rate exceeds the set threshold, perform a cache replacement operation.

[0014] The above technical solution has the following advantages or beneficial effects:

[0015] By introducing a file priority allocation mechanism, the present invention can dynamically adjust the data transmission and storage policies according to indicators such as the importance, access frequency, size, and freshness of the file, ensuring that high-priority files obtain higher resource priorities during the asynchronous storage process. The combined use of synchronous writing and asynchronous writing not only improves the data writing efficiency but also effectively reduces the risk of data loss caused by failures before asynchronously writing to the underlying storage system, thereby ensuring the high availability and persistence of the data.

[0016] The present invention adopts an intelligent prefetching and replacement strategy based on the correlation between data blocks, which can actively prefetch the data blocks associated with the currently accessed data during the data access process, reduce the remote call frequency in the cross-data center scenario, and reduce the data access latency. Especially when the system load is low, the intelligent prefetching mechanism can make full use of idle resources, improve the cache hit rate and the I / O performance of the system, and further enhance the availability of Alluxio in the multi-data center environment.

[0017] By designing a dynamic load balancing mechanism, the present invention can monitor the resource usage of each node in real time and dynamically adjust the task allocation and weights according to the load conditions, avoiding resource overload while maximizing the use of system resources. This optimization strategy ensures that the system can still operate efficiently in the high-concurrency scenario, reduces the performance bottleneck caused by load imbalance, and thus improves the overall performance and response speed of the system.

[0018] The present invention will comprehensively consider the file access frequency, size, importance, and freshness to implement the allocation of file priorities, formulate different transmission policies for files with different priorities, reduce the risk of data loss, and improve the overall performance of the system.

[0019] By considering the access frequency and timing correlation between data blocks, and combining the network conditions and system load in the cross-data center scenario, intelligent prefetching of associated data blocks is performed to reduce the frequency of remote calls and lower the data access latency. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments and descriptions thereof of the invention are used to explain the invention and do not constitute an improper limitation of the invention.

[0021] Figure 1 It is a flowchart of the method for the first embodiment;

[0022] Figure 2 It is a data transmission strategy diagram for the first embodiment;

[0023] Figure 3 It is a flowchart of dynamic load balancing for the first embodiment;

[0024] Figure 4 It is for parallel mining of association rules in the first embodiment;

[0025] Figure 5 It is a flowchart of intelligent cache replacement for the first embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs.

[0027] As an open-source distributed file system, Alluxio can provide an efficient data access layer to alleviate these problems. By caching hot data in the local data center and using an asynchronous write mechanism to reduce the frequency of cross-data center data transfer. It supports multiple ways to write data, which can be specified when writing data. Alluxio supports four ways to write data, namely, data is only retained in Alluxio (MUST_CACHE), data is written to both Alluxio and the underlying persistent storage simultaneously (CACHE_THROUGH), data is asynchronously written to the underlying layer (ASYNC_THROUGH), and data is only written to the underlying layer (THROUGH). Although the synchronous method can ensure data persistence, the speed of writing to the underlying storage is usually much slower than writing to local storage. Therefore, the synchronous write speed will be comparable to the speed of writing to the underlying storage. Asynchronous storage is a way to separate the data write operation from the data persistence operation, enabling data to be first written to a high-speed cache layer and then asynchronously written to a low-speed persistent layer in the background, thereby improving the data write performance and efficiency. This method can provide data writing at a speed close to that of MUST_CACHE and can complete data persistence.

[0028] In order to write data at the memory speed while persisting data, the asynchronous storage strategy provided by Alluxio is proposed, where data is synchronously written to the worker and asynchronously written to the underlying storage system. However, if a failure occurs before the data is asynchronously written to the underlying storage system, it may lead to data loss. And since data transfer and storage operations are performed in the background, they may compete with foreground applications for system resources (such as CPU, memory, and bandwidth), affecting the overall performance. Therefore, research an asynchronous storage optimization strategy to reduce the risk of data loss.

[0029] In the era of big data and cloud computing, the efficiency of data processing and storage has a crucial impact on the overall performance of the system. As an open-source distributed file system, Alluxio aims to provide an efficient data access layer between the computing framework and the storage system, thereby enhancing the performance of data-intensive applications. However, in cross-domain scenarios, network latency becomes an important bottleneck affecting the system I / O speed. Especially when the computing speed has increased significantly, this mismatch is more obvious. The bottleneck of I / O speed not only limits the efficiency of data processing but also directly affects the quality of external services. Using prefetching technology and replacement technology to persistently retain hot data in Alluxio can avoid frequent remote calls, reduce the number of accesses to the underlying storage, and reduce data access latency. Therefore, research a cache prefetching and replacement strategy based on Alluxio in cross-domain scenarios to improve the availability and performance of Alluxio.

[0030] Example 1

[0031] This embodiment provides a data storage and cache optimization method based on Alluxio;

[0032] As Figure 1 shown, the data storage and cache optimization method based on Alluxio includes:

[0033] S101: Obtain the file to be written, and calculate the priority of the file according to four attributes of the file to be written; the four attributes of the file to be written include: file access frequency, file capacity, file importance level, and file freshness;

[0034] S102: Store the file to be written in the Alluxio cache and the underlying storage according to the file transfer policy corresponding to the file priority;

[0035] S103: Obtain the data to be read, and determine whether the data to be read is in the cache. If so, return the data in the cache to the user; if not, return the data in the underlying storage to the user;

[0036] S104: Perform a prefetch operation on the data to be read and the data blocks associated with the data to be read according to the association rule. If the cache usage rate exceeds the set threshold, perform a cache replacement operation.

[0037] It should be understood that in S101: Obtain the file to be written, each data packet of the file contains detailed information to ensure the effective management and tracking of data transmission. Each file data packet includes: file ID, priority, source node and target node, data block information (data block size and storage location), transmission status, checksum, replica information, and timestamp.

[0038] The file ID is the unique identifier of the file, used to distinguish different files. The file priority is divided into three levels: high, medium, and low. The priority level determines different transmission policies for the file. High-priority files need to be transmitted immediately and ensure data persistence. Medium-priority files allow a certain delay, while low-priority files can be transmitted when resources are idle. The transmission progress records the current transmission status of the data packet, such as "waiting for transmission", "transmitting", "transmission completed", "transmission failed", etc. The transmission status is updated in real time to monitor the transmission progress and handle exceptions in a timely manner. A checksum is calculated for each data block before transmission and verified after transmission to ensure that the data has not been tampered with or damaged during transmission. For high-priority data, record the number and storage location of the data block replicas to quickly recover the data in case of node failure. The replica information ensures high availability and reliability of the data. The timestamp records timestamp information such as the creation time, transmission start time, and transmission completion time of each data packet.

[0039] Further, the file access frequency refers to:

[0040]

[0041] Among them, F i represents the access frequency of file i, n represents the number of time periods, and w j represents the weight of the j-th time period, and A j represents the number of accesses of file i within the j-th time period.

[0042] For example, the time is divided into days, and the number of accesses to the file is recorded every day. A weight w j is assigned to each time period. The weight of a more recent time period is larger, and the weight of a more distant time period is smaller.

[0043] It should be understood that the access frequency of a file is measured by comprehensively considering time factors, the number of accesses, and weights by monitoring and tracking the usage of each file in real time. The higher the access frequency, the higher the priority.

[0044] Further, the file capacity is obtained from the metadata of the file. Files with a larger capacity require more transmission and storage resources.

[0045] Further, the file importance level is determined by user selection. When uploading a file, the user can determine the importance of the file by selecting a level from 1 to 10. The larger the number, the more important the file and the higher the priority.

[0046] Further, the file freshness is determined according to the last modification time of the file. The file freshness is obtained by subtracting the last modification time from the current time. The smaller the obtained value, the fresher the data.

[0047] Further, calculating the priority of the file according to the four attributes of the file to be written includes:

[0048] Assume that for file i, the original metrics are: access frequency F i , file size S i , file importance I i and data freshness N i . Standardize each metric:

[0049]

[0050] Combining the standardized metrics, calculate the priority P i of each file according to the weight ratio as:

[0051] P i = aF i+bS i +cI i +dN i ,

[0052] Among them, a, b, c, and d are weights, and a + b + c + d = 1.

[0053] Furthermore, as Figure 2 shown, S102: Store the file to be written into the Alluxio cache and the underlying storage according to the file transfer policy corresponding to the file priority, including:

[0054] If the file priority is high priority, use the synchronous write mechanism and the dual-copy mechanism;

[0055] The synchronous write mechanism means that while writing the file to be written into the Alluxio cache, also write the file to be written into the underlying storage;

[0056] The dual-copy mechanism means that each data block is copied to another worker node;

[0057] After the data is written to the underlying storage, start data verification. The data verification is used to determine whether the written file is damaged. If the verification is successful, it means the file writing is completed; if the verification fails, it means the file needs to be rewritten.

[0058] It should be understood that the Alluxio cache refers to a distributed memory storage layer. Alluxio can cache hot data (i.e., frequently accessed data) in memory, making data reading and processing more efficient. Through the Alluxio cache, the application can quickly obtain data from memory without directly accessing the underlying storage, greatly reducing I / O latency and improving data processing performance.

[0059] It should be understood that the underlying storage refers to the persistent storage system or device that actually stores data. It is responsible for storing data in the long term and ensuring the security, persistence, and reliability of the data. The underlying storage usually includes various forms of storage systems, such as distributed file systems (such as HDFS), object storage (such as Amazon S3, Azure Blob Storage), local hard disks, or network storage devices. The underlying storage is more stable and persistent than the cache, and the data will not disappear even in the case of system restart or memory loss. Alluxio will finally write the file or data to the underlying storage after processing to ensure that the data can be stored in the long term and accessed later.

[0060] It should be understood that since the priority of a file reflects its importance and urgency, a synchronous write mechanism is considered for high-priority files. While writing to the Alluxio cache, it is immediately written to the underlying storage. To ensure high availability and reliability of the data, a dual-copy mechanism is maintained until the write is completed, that is, each data block is copied to another worker node. After the data is written to the underlying storage, the system will automatically trigger a data verification process. Checksums (such as MD5 or SHA-256) are used to verify the data blocks to ensure that no data corruption or tampering has occurred during transmission. Once the verification is successful, the system will generate a confirmation message and update the file status to "write completed". If the verification fails, a data retransmission mechanism is triggered to perform the write operation again until the data verification is successful.

[0061] Further, the step S102: storing the file to be written in the Alluxio cache and the underlying storage according to the file transfer policy corresponding to the file priority further includes:

[0062] If the file priority is medium, an asynchronous write mechanism and a load balancing mechanism are used;

[0063] The asynchronous write mechanism means that the file to be written is first written to the Alluxio cache, and then the file to be written is written to the underlying storage; during the writing process, the intermediate state of the saved data is periodically checked through a checkpoint mechanism.

[0064] It should be understood that the checkpoint mechanism is a mechanism to ensure that the data processing can be restored even in case of a failure. It periodically saves the current data processing state as a "checkpoint" during the processing of tasks, so that when a failure occurs, the system can restart the processing from the most recent checkpoint instead of starting from scratch, thereby reducing the data recovery time and resource consumption.

[0065] Further, the specific steps of periodically checking the intermediate state of the saved data through the checkpoint mechanism include:

[0066] (11) Set the checkpoint interval: Configure a set time interval or data volume (such as every minute or for every certain amount of data processed) as the frequency of creating checkpoints;

[0067] This ensures that checkpoints are not created too frequently, which would affect system performance, but are frequent enough to effectively save the intermediate state;

[0068] (12) Save the checkpoint: When the set time interval or data volume is reached, save the progress of the current file write operation as a checkpoint;

[0069] The progress described above includes: the offset of the data, the writing status, the metadata of the file, and the system environment information; the offset of the data is the current position where the file is written; the writing status is the completed data blocks or shards; the metadata of the file is the location information of the file in storage; the system environment information is the status of the current processing node, resource usage, etc.

[0070] (13) Store checkpoint information;

[0071] Through the storage checkpoint, the system can accurately locate the writing progress during fault recovery.

[0072] (14) Recovery when a fault occurs: If a fault occurs during the asynchronous writing process, restart the writing operation from the nearest checkpoint.

[0073] As Figure 3 shown, the load balancing mechanism includes:

[0074] (21) Based on the real-time load situation, adopt the weighted round-robin algorithm to allocate transmission tasks: The weighted round-robin algorithm assigns an initial weight to each worker node;

[0075] (22) During the polling process of the transmission task, the monitoring agent sub-node of the worker node itself collects the current load data of the current worker node at a set time interval and uploads the load data to the central monitoring server;

[0076] The central monitoring server determines whether the load of the current worker node exceeds the set first threshold. If it exceeds the set first threshold, the weight of the current worker node is reduced; determines whether the load of the current worker node is lower than the set second threshold, and then the weight of the current worker node is increased; if the load of the current worker node is lower than the first threshold and higher than the second threshold, the weight of the current worker node remains unchanged;

[0077] (23) Allocate tasks according to the weights.

[0078] It should be understood that a worker node refers to a machine or server responsible for storing data and processing file transmission. Each worker node is equipped with a monitoring agent sub-node that regularly collects the resource usage data of the node and sends this data to the central monitoring server to help achieve load balancing and resource optimization scheduling.

[0079] The monitoring agent sub-node refers to a lightweight monitoring component or service deployed on each worker node, which is responsible for real-time collecting and monitoring the resource usage and running status of the worker node. Its main function is to regularly collect various key resource metrics of the worker node and report this data to the central monitoring server to support the system's load balancing and resource optimization decisions.

[0080] The central monitoring server refers to the core component or server in a distributed system responsible for centralized management and monitoring of all worker nodes. It receives and analyzes the resource usage data uploaded by each worker node, evaluates the overall load of the system in real time, and makes intelligent scheduling decisions based on this data, such as load balancing, task allocation, etc.

[0081] Furthermore, the (21) weighted round-robin algorithm assigns an initial weight to each worker node, where the initial weight W i is calculated by the formula:

[0082] W i = αC i + βM i + γB i + δI i + εD i ;

[0083] where C i represents the CPU capacity of worker node i, M i represents the memory capacity of worker node i, B i represents the network bandwidth of node i, I i represents the input / output I / O performance of worker node i, D i represents the distance of worker node i from the central node. α, β, γ, δ, and ε are weight coefficients used to adjust the influence of each resource capacity and distance on the total weight.

[0084] Furthermore, during the round-robin process of transmitting tasks, the monitoring agent sub-node of the worker node itself regularly collects the current load data of the current worker node, including:

[0085] During the round-robin process of transmitting tasks, the current load data of the worker node is collected every minute. The current total load L of worker node i i is calculated by the formula:

[0086] L i = αL cpu,i + βL mem,i + γL bw,i + δL io,i

[0087] where L cpu,i represents the CPU load of worker node i, L mem,i represents the memory load of worker node i, L bw,i represents the network bandwidth load of worker node i, L io,i represents the input / output I / O load of worker node i. The weight W is adjusted according to the current load L i . i

[0088] Further, the (22) central monitoring server determines whether the load of the current working node exceeds a set first threshold. If it exceeds the set first threshold, the weight of the current working node is reduced; it determines whether the load of the current working node is lower than a set second threshold, and if so, the weight of the current working node is increased. Specifically, it includes:

[0089] If the load Li of working node i i exceeds the preset high load threshold T1, the weight of working node i is reduced to W' i ;

[0090] W i ' = W i - λ(Li i - T1), if Li i > T1

[0091] If the load Li of working node i i is lower than the set low threshold T2, the weight of working node i is increased to W′ i , and the adjusted weight W′ i is calculated by the formula:

[0092] W′ i = W i + η(T2 - Li i ), if Li i < T2

[0093] where λ and η are coefficients for adjusting the weight.

[0094] It should be understood that (23) tasks are assigned according to the weights, specifically including:

[0095] W i is the weight of the i-th node; in each polling cycle, all working nodes are traversed, and tasks are assigned in order of the weights. Nodes with larger weights are assigned more tasks in each round, while nodes with smaller weights are assigned fewer tasks.

[0096] Exemplarily, (23) tasks are assigned according to the weights, specifically including:

[0097] If there are three working nodes A, B, and C, and their weights are W A = 5, W B = 3, W C = 2, the tasks will be assigned according to the following pattern:

[0098] Polling 1: Working node A receives 5 tasks, B receives 3 tasks, and C receives 2 tasks.

[0099] Polling 2: If the load of worker node A becomes high and the weight may be adjusted to W A = 4, then during the next task polling, A will receive 4 tasks, and the task allocation for B and C will also be dynamically adjusted according to the load.

[0100] It should be understood that medium-priority files are written asynchronously. The data is first written to the Alluxio cache and then triggers an asynchronous write to the underlying storage. During the writing process, the system periodically saves the intermediate state of the data through a checkpoint mechanism. This mechanism ensures that in case of a failure, the writing process can be restarted from the most recent checkpoint without starting from scratch, reducing the data recovery time and system resource consumption.

[0101] To avoid resource contention, the system dynamically monitors the load conditions of each worker node and distributes the transmission tasks of medium-priority files to nodes with lower load. Based on the real-time load conditions, the system uses a weighted round-robin algorithm to allocate transmission tasks. The weighted round-robin algorithm assigns a weight to each node, and the weight is calculated based on the node's resource capabilities (such as CPU, memory, network bandwidth, and I / O) and the distance between nodes. First, an initial weight is set for each node, and the weight reflects the maximum resource processing capacity of the node. The initial weight of each node is calculated based on its maximum resource processing capacity.

[0102] This dynamic adjustment mechanism can ensure that the system remains balanced when the load changes, avoiding overloading of certain nodes and resulting in a decline in system performance. The system analyzes the load conditions of each node in real time. The monitoring agent sends the collected resource usage data to the central monitoring server, and the server analyzes these data through algorithms to identify nodes with lower current load. These nodes will preferentially receive new transmission tasks, thus achieving load balancing.

[0103] After the transmission task allocation, the system continuously monitors the execution of the tasks. If a failure is detected in a certain node, the system will immediately stop allocating new tasks to that node and reallocate the unfinished tasks to other nodes with lower load and closer distance. Ensure that the transmission tasks can be quickly restored in case of a failure, reducing the interruption time of data transmission.

[0104] Furthermore, the step S102: storing the file to be written in the Alluxio cache and the underlying storage according to the file transmission policy corresponding to the file priority further includes:

[0105] If the file priority is low priority, an asynchronous write mechanism and a data compression mechanism are used;

[0106] The asynchronous write mechanism means that the file to be written is first written to the Alluxio cache, and when the system resources are idle, the file to be written is then written to the underlying storage.

[0107] The data compression mechanism refers to:

[0108] When the network bandwidth utilization rate is higher than the set threshold, while the CPU utilization rate and the input / output interface utilization rate are both lower than the set threshold, select the LZ4 compression algorithm to compress the file to be written;

[0109] When the network bandwidth utilization rate is lower than the set threshold, while the CPU utilization rate and the input / output interface utilization rate are both higher than the set threshold, select the Gzip compression algorithm to compress the file to be written;

[0110] When the network bandwidth utilization rate, the CPU utilization rate, and the input / output interface utilization rate are all within the set threshold range, select the Zstandard compression algorithm to compress the file to be written.

[0111] It should be understood that the LZ4 compression algorithm is a lossless data compression algorithm, known for its high-speed compression and decompression speed. It can compress and decompress data in a very short time.

[0112] It should be understood that the Gzip compression algorithm is a commonly used lossless data compression algorithm. It can significantly compress the file size while ensuring data integrity. Gzip is often used to reduce the file size to save space and bandwidth in file storage and file transmission. When the network bandwidth utilization rate is low while the CPU and I / O utilization rates are high, the purpose of selecting the Gzip compression algorithm is to use its high compression ratio to reduce the file size, thereby reducing the pressure on the CPU and I / O and efficiently using network resources.

[0113] It should be understood that the Zstandard compression algorithm is a lossless data compression algorithm. It aims to provide the best balance among compression speed, decompression speed, and compression ratio. It can provide an extremely fast processing speed while maintaining an efficient compression ratio.

[0114] It should be understood that the data of low-priority files is first written to the cache, and asynchronous writing to the underlying storage is triggered when the system resources are idle. After the file is written to the Alluxio cache, the system continuously monitors the current resource usage situation. Considering the utilization rates of network bandwidth, CPU, and I / O comprehensively to decide whether to trigger asynchronous writing. When the utilization rates are all lower than their respective set thresholds, the asynchronous writing process is triggered to transfer the data of low-priority files from the cache to the underlying storage, avoiding the resource contention of low-priority tasks for high-priority tasks, thereby optimizing the overall performance.

[0115] To reduce the amount of data transmitted and bandwidth occupancy, the system judges the file size when transmitting low-priority files. For files with a capacity larger than the set threshold, it is selected to compress the file first and then transmit it. In the selection of compression algorithms, the type of compression algorithm to be used is dynamically judged according to the real-time resource utilization rate and task requirements.

[0116] When the system detects that the network bandwidth utilization rate is high while the CPU and I / O utilization rates are low, this usually indicates that the transmission bandwidth is the current main bottleneck. In this case, the system will select the LZ4 algorithm because LZ4 provides very fast compression and decompression speeds, which can quickly reduce the amount of data and relieve the bandwidth pressure. LZ4 is suitable for scenarios where a large amount of data needs to be transmitted and the transmission needs to be completed quickly.

[0117] When the system detects that the CPU and I / O utilization rates are high while the network bandwidth utilization rate is low, this indicates that the system resources are relatively sufficient and are suitable for more efficient compression. In this case, the system selects the Gzip algorithm. The compression ratio of Gzip is better than that of LZ4. Although the compression and decompression speeds are slightly slower, it can significantly reduce the amount of data transmitted and is applicable to scenarios where high compression ratio is required but the transmission delay requirement is not high.

[0118] When the system detects that the utilization rates of CPU, I / O, and network bandwidth are all at a medium level and a balance needs to be achieved between the compression ratio and the compression speed, the Zstandard (Zstd) algorithm will be selected. Zstd provides a relatively high compression ratio while having relatively fast compression and decompression speeds, and is applicable to situations where comprehensive performance is required.

[0119] Further, the step S103: obtaining the data to be read, judging whether the data to be read is in the cache, if so, returning the data in the cache to the user; if not, returning the data in the underlying storage to the user, specifically includes:

[0120] When Alluxio receives a data read request from a user, it first confirms whether the requested data is already in the Alluxio cache;

[0121] If the data is in the cache, Alluxio directly reads from the cache and returns it to the user;

[0122] If the data is not in the cache, Alluxio retrieves the data from the underlying storage (such as HDFS, S3, etc.) and returns it to the user;

[0123] At the same time, Alluxio caches the data obtained from the underlying storage locally and updates the metadata for the next access.

[0124] Further, as Figure 5As shown, S104: According to the association rules, perform a prefetch operation on the data to be read and the data blocks associated with the data to be read. If the cache usage rate exceeds the set threshold, perform a cache replacement operation, specifically including:

[0125] S104-1: Set the size and step size of the sliding window;

[0126] S104-2: Calculate the closed frequent item sets for the data within the sliding window;

[0127] S104-3: Construct a Trie tree to store the closed frequent item sets, merge the same prefixes, and obtain the association relationships between data blocks based on the Trie tree;

[0128] S104-4: Determine whether the data accessed by the user is in the cache. If it is, enter S104-5; if not, enter S104-7;

[0129] S104-5: Determine whether the data blocks associated with the current data block are in the cache. If they are, return to S104-1; if not, mark the data blocks associated with the current data block as "prefetch candidates" and enter S104-6; the associated data blocks are obtained based on the Trie tree. If data block A and data block B belong to the same closed frequent item set, it means that data block A and data block B are associated; otherwise, it means that data block A and data block B are not associated; the "prefetch candidate" refers to the candidate data blocks waiting to be selected;

[0130] S104-6: Determine whether the load is lower than the set threshold. If not, return to S104-1; if it is, store the data blocks marked as "prefetch candidates" in the cache and return to S104-1;

[0131] S104-7: Determine whether the cache usage rate exceeds the set threshold. If not, store the data block accessed by the user in the cache and enter S104-5; if it is, remove the data block based on the LRU policy and mark the associated data blocks; enter S104-8;

[0132] S104-8: Determine whether the time difference between the most recent access time of the associated data block and the current time exceeds the set time threshold. If it is, remove the current data block and return to S104-7; if not, return to S104-7.

[0133] Further, in S104-1, setting the size and step length of the sliding window includes: the size of the sliding window is the amount of data contained within the sliding window, and the step length of the sliding window is the amount of new data added each time the window slides; considering the data in the window as a queue, if the amount of data in the queue exceeds the set window size W, the earliest data added to the queue is removed to ensure that only the latest W pieces of data are in the queue; marking the removed data as "old data"; when new data arrives, adding the new data to the current window and marking it as "new data". The data previously marked as "new data" has the "new data" mark removed after a new round of data addition. In other words, only the data added in the latest round is "new data".

[0134] It should be understood that due to the large amount of data to be mined, full-scale data mining may be a heavy task, so a sliding window strategy is adopted. First, the size of the window and the sliding step length need to be set. The window size determines the amount of data contained in each window, and the sliding step length determines the amount of new data added each time the window slides.

[0135] To manage the data in the window, a queue needs to be maintained to store the data in the current window. Whenever new data arrives, it is added to the end of the queue, and at the same time, the size of the queue is checked. If the amount of data in the queue exceeds the set window size W, the earliest added data is removed from the front of the queue to ensure that there are always only the latest W pieces of data in the queue.

[0136] During the data update process, when new data arrives, it is first added to the current window and marked as "new data" to distinguish which data is newly added and which data already exists, thus achieving incremental update.

[0137] Regularly check the size of the queue. When the amount of data in the queue exceeds the window size W, the earliest data is removed from the queue, and these removed data are marked as "old data".

[0138] Further, in S104-2: calculating the closed frequent item sets for the data within the sliding window includes:

[0139] S104-21: Each time the window slides, only process the data marked as "new data", and calculate the closed frequent item sets of the "new data" within the window; merge the closed frequent item sets of the "new data" within the window with the closed frequent item sets of all the data within the window to update the closed frequent item set list;

[0140] S104-22: Divide the entire window into several subsets; assign the subsets to different worker nodes;

[0141] S104-23: Based on the list of closed frequent item sets, each worker node calculates the local closed frequent item sets and performs pruning within the local scope;

[0142] S104-24: After each worker node completes the mining of local closed frequent item sets, local merging is first performed on some nodes, and then the merged results are sent to the central monitoring server for final merging;

[0143] S104-25: Deduplicate the results of the closed frequent item sets generated by each node, merge the same closed frequent item sets, and recalculate the support degree of the merged closed frequent item sets to ensure the correctness of its support degree within the global scope.

[0144] S104-26: Perform final pruning on the merged global closed frequent item sets, removing the item sets that do not meet the support degree threshold, that is, the item sets that are locally closed frequent item sets but not globally closed frequent item sets.

[0145] Further, the closed frequent item set means that if the support degree of an item set (that is, the proportion of transactions containing this item set in the total transactions) meets the set threshold, then this item set is called a frequent item set.

[0146] The support degree of an item set is defined as:

[0147]

[0148] A transaction represents a set of data records processed in a sliding window. Each transaction usually consists of multiple data items. For example, in a data stream processing, each transaction may correspond to a purchase behavior, a set of user operation records, or a specific event. In frequent item set mining, the system calculates the support degree of item sets based on these transactions and analyzes which combinations of data items often appear together. The total number of transactions in the sliding window is the total number of data records currently covered by the window and serves as the basis for calculating frequent item sets.

[0149] Further, the item set refers to a set of data composed of multiple data item combinations in the transactions within the sliding window. It is usually used to describe the item combinations that appear together in a transaction database or dataset. An item set can be used to represent several elements that appear simultaneously in a transaction or record and is widely used in tasks such as association rule mining and frequent item set mining. A frequent item set is an item set that appears above a certain support degree threshold. It reflects the frequently occurring item combinations in the data.

[0150] The item set I refers to a combination of several data items that appear simultaneously in the transactions within the sliding window. It can be used to represent the data items that often appear together in a transaction database or dataset.

[0151] Illustrate the meaning of the item set I with an example:

[0152] Assume that there is the following transaction data in the current sliding window (each transaction is a set of events or data items that occur simultaneously):

[0153] Transaction 1: {apple, banana, pear};

[0154] Transaction 2: {apple, pear};

[0155] Transaction 3: {banana, grape};

[0156] Transaction 4: {apple, banana};

[0157] In this case, an itemset I can be any combination of items that appear simultaneously in a transaction. For example:

[0158] The itemset {apple, pear} means that apples and pears appear simultaneously in the same transaction (appear in Transaction 1 and Transaction 2).

[0159] The itemset {banana, grape} means that bananas and grapes appear simultaneously in the same transaction (appear in Transaction 3).

[0160] The itemset {apple, banana} means that apples and bananas appear simultaneously in the same transaction (appear in Transaction 1 and Transaction 4).

[0161] If the support of this itemset I (i.e., the frequency at which this itemset appears in the transactions within the sliding window) meets the set threshold, then it is a frequent itemset. For example, the itemset {apple, pear} appears in 2 transactions. If the threshold is 50%, then it is a frequent itemset.

[0162] Regarding the processing of "new data":

[0163] When the window slides each time, only the data marked as "new data" will be processed. For example, if Transaction 4 is the latest (i.e., newly slid into the window), then the current sliding window will only process the data items {apple, banana} in Transaction 4 to update the frequent itemset.

[0164] Furthermore, the closed frequent itemset is a special type of frequent itemset. It not only meets the conditions of a frequent itemset (i.e., its support is higher than the set threshold), but also has closure. The definition of closure is: If the superset of a certain frequent itemset (i.e., a larger itemset that contains this itemset) has the same support as it, then this itemset is no longer a closed frequent itemset, and the superset is regarded as the closed frequent itemset. Therefore, the closed frequent itemset represents those item sets that cannot be covered by a larger itemset.

[0165] The closed frequent itemset is not only a frequent itemset, but also there is no other larger itemset in the data that contains this itemset and has the same support.

[0166] It should be understood that when sliding the window each time, only those transactions marked as "new data" are processed to calculate their closed frequent itemsets. For the removal process of old data, when data is removed from the window, the system needs to update the current frequent itemsets and remove those affected itemsets, that is, those whose support no longer meets the threshold due to the removed data. Finally, the mining results of the new data are merged with the results of the closed frequent itemsets in the current window to update the list of closed frequent itemsets. This requires recalculating the support of all affected itemsets to ensure their support in the current window.

[0167] When mining closed frequent itemsets, the mining algorithm idea used is the parallelized CHARM algorithm, which divides the entire data window into several subsets, and each subset can be processed independently.

[0168] The CHARM algorithm is from: Zaki, Mohammed J., and Ching-Jui Hsiao. "CHARM: An efficient algorithm for closed itemset mining." Proceedings of the 2002 SIAM international conference on data mining. Society for Industrial and Applied Mathematics, 2002.

[0169] The partitioning strategy adopts the method of transaction partitioning, dividing the data window equally by quantity to ensure the balance of the data volume of each subset and avoid over-concentration of calculations in some subsets.

[0170] In the parallel computing stage, the distributed computing framework Apache Hadoop is used to allocate tasks to different computing nodes, and a hash table is used to record local frequent itemsets and their supports.

[0171] Each node runs the CHARM algorithm independently, performs local closed frequent itemset mining on its own subset, and performs pruning within the local scope.

[0172] After each computing node completes local closed frequent itemset mining, local merging is first performed on some nodes, and then the merged results are sent to the central node for final merging.

[0173] Deduplication is performed on the results of the closed frequent itemsets generated by each node, the same closed frequent itemsets are merged, and the support of the merged closed frequent itemsets is recalculated to ensure its correct support in the global scope.

[0174] Perform final pruning on the merged global closed frequent item sets, removing item sets that do not meet the support threshold, that is, item sets that are locally closed frequent item sets but not globally closed frequent item sets.

[0175] Further, S104-21: Merge the closed frequent item sets of the "new data" within the window with the closed frequent item sets of all data within the window, specifically including:

[0176] In the sliding window strategy, when "new data" enters the window and its closed frequent item sets are calculated, merge the closed frequent item sets of the "new data" with the closed frequent item sets of all data within the current window.

[0177] First, for each closed frequent item set of the "new data", check whether it already exists in the frequent item sets of the current window. If it exists, update the support of this item set; if it does not exist, add it as a new item set.

[0178] Next, recalculate the support of all item sets within the entire window, and remove item sets with support lower than the set threshold from the frequent item sets.

[0179] Finally, during the merging process, if a superset of an item set has the same support, then this item set is removed.

[0180] This merging process can dynamically update the frequent item sets, reflect the data increment change, and at the same time maintain the integrity and correctness of the frequent item sets.

[0181] A superset is a larger set formed by adding other data items based on a certain item set. Equivalently, a superset is a larger set formed by adding other data items based on a certain item set.

[0182] Illustrate with an example:

[0183] Suppose we have an item set A = {apple, banana}, then the item set {apple, banana, pear} is a superset of A because it contains all the data items of A and also has an additional data item "pear".

[0184] Item set A = {apple, banana}

[0185] Item set B = {apple, banana, pear}

[0186] Item set B is a superset of item set A because B contains all the elements of A and has more elements.

[0187] In the definition of closed frequent item sets, if a superset of a frequent item set has the same support as it (that is, they appear in the same number of transactions), then this frequent item set is not a closed frequent item set, and the superset is the closed frequent item set.

[0188] Example: Suppose the itemset {apple, banana} appears in 3 transactions, and its superset {apple, banana, pear} also appears in the same 3 transactions. Then {apple, banana} is not a closed frequent itemset, and only {apple, banana, pear} is a closed frequent itemset because it contains more items and there is no larger itemset that appears in the same transactions.

[0189] Furthermore, the S104-21: Update the list of closed frequent itemsets, including:

[0190] The list of closed frequent itemsets is a structure that stores all the closed frequent itemsets and their corresponding support degrees within the current window. It is used to track the frequent patterns in the dataset and maintain its accuracy through an incremental update mechanism. In the sliding window, when "new data" enters, the new closed frequent itemsets need to be merged and updated with the current list of closed frequent itemsets.

[0191] List of closed frequent itemsets = [{itemset: {A, B}, support degree: 0.8}, {itemset: {A, C}, support degree: 0.3}].

[0192] Furthermore, the S104-22: Divide the entire window into several subsets, including:

[0193] Dividing the entire window into several subsets means dividing the data within the window.

[0194] In the process of data mining, to improve the computational efficiency and utilize the advantages of parallel processing, a large dataset (such as all the data within the sliding window) is usually split into multiple subsets, and these subsets are assigned to different computing nodes for parallel processing.

[0195] Furthermore, the S104-22: Assign the subsets to different computing nodes, including:

[0196] In the parallel computing stage, the distributed computing framework Apache Hadoop is used to assign tasks to different computing nodes.

[0197] Furthermore, the S104-23: Based on the list of closed frequent itemsets, each computing node calculates the local closed frequent itemsets, including:

[0198] The mining process of local closed frequent itemsets is similar to that of global closed frequent itemsets, but it only acts on the data of each subset. Specifically, a local closed frequent itemset refers to a frequent itemset calculated within a certain data subset and satisfies the definition of a closed frequent itemset.

[0199] Furthermore, as Figure 4 shown, the S104-23: And perform pruning processing within the local scope, including:

[0200] In local calculation, if the supersets of a certain itemset have the same support, then this itemset is no longer a closed frequent itemset and needs to be removed. In addition, if the support of a certain itemset is lower than the set minimum support threshold, this itemset will also be pruned to avoid it entering the subsequent calculation stage.

[0201] After local pruning is completed, the nodes will merge the calculation results and further perform pruning operations globally. Global pruning will remove those item sets that are closed frequent item sets within the local scope but no longer meet the conditions of closed frequent item sets due to the decrease in support globally. The pruning operation reduces the computational complexity by removing redundant item sets and improves the overall efficiency of the system.

[0202] A superset refers to a larger itemset formed by adding one or more data items to the original itemset. In other words, the superset of itemset A is a set that contains all the data items of itemset A and at least one other data item.

[0203] Furthermore, for S104-24: After each computing node completes the mining of local closed frequent item sets, local merging is first performed on some nodes, specifically including:

[0204] When each computing node completes the calculation of local closed frequent item sets, first, the local closed frequent item sets generated by each computing node are merged. The local merging process refers to: initially integrating the local closed frequent item sets generated among the same group of nodes to reduce redundant and duplicate item sets.

[0205] During merging, compare the local frequent item sets of each node, find the same item sets, and add their supports.

[0206] If it is found during the local merging process that the supports of some item sets do not reach the set minimum threshold, then directly perform pruning within the local scope and remove these item sets.

[0207] By adopting this method of first performing local merging on some nodes, the amount of frequent item set data transmitted to the central node can be effectively reduced, the computational and transmission overheads can be lowered, and thus the efficiency of the entire distributed system can be improved.

[0208] Furthermore, for S104-24: Then, the merged result is sent to the central monitoring server for final merging, specifically including:

[0209] After each computing node completes the mining of local closed frequent item sets and performs preliminary local merging, the merged result is sent to the central monitoring server through the communication mechanism built into the distributed computing framework Hadoop.

[0210] Each node packs the merged local frequent item sets into a result set and sends the result set to the central monitoring server via network transmission.

[0211] After receiving the local results from all computing nodes, the central monitoring server is responsible for merging and recalculating the support to ensure the correct support of the final closed frequent item sets globally.

[0212] Meanwhile, during the merging process, the central node also performs deduplication and pruning to remove redundant or item sets that do not meet the support threshold, thus generating a complete list of global closed frequent item sets.

[0213] Furthermore, the step S104-25: deduplicate the closed frequent item set results generated by each node, including:

[0214] After each node generates local closed frequent item sets, the central node first performs deduplication on these results.

[0215] This process is achieved by comparing the frequent item sets uploaded by different nodes, identifying the same item sets, and merging them.

[0216] For example, if the content of the item sets generated by different nodes is the same, they are considered duplicate item sets.

[0217] During the merging, the central node accumulates the support of these duplicate item sets to ensure that the support reflects the true frequency of the item set globally. Through this deduplication and merging operation, the system can avoid redundant storage and processing and ensure the correctness and accuracy of the global closed frequent item set list.

[0218] Furthermore, the step S104-25: merge the same closed frequent item sets, including:

[0219] Merging the same closed frequent item sets means that in the results of local closed frequent item sets generated by each computing node, the central node integrates the item sets with the same content.

[0220] If multiple nodes independently calculate the same closed frequent item set, these item sets are duplicates globally, so they need to be merged.

[0221] During the merging process, the central node adds up the support of the same item sets to reflect the frequency of the item set in the entire dataset.

[0222] This process ensures that each item set in the global closed frequent item set list appears only once and its support accurately represents the frequency of the item set in all data.

[0223] Further, S104-26: Perform final pruning on the merged global closed frequent item sets to remove item sets that do not meet the support threshold, that is, item sets that are locally closed frequent item sets but not globally closed frequent item sets, including:

[0224] The process of performing final pruning on the merged global closed frequent item sets is to ensure that only item sets that meet the global support threshold are retained.

[0225] Although some item sets meet the conditions of closed frequent item sets in local nodes, globally, due to the possible decrease in support, they no longer meet the criteria of frequent item sets.

[0226] Therefore, the system will remove those item sets with support lower than the set threshold by comparing the global support of each item set.

[0227] In addition, if a superset of an item set has the same support, then the item set will also be pruned because it is no longer a closed item set globally.

[0228] Through this pruning operation, it is ensured that the finally retained item sets are both frequent and closed globally.

[0229] An item set that does not meet the support threshold refers to an item set whose support is lower than the pre-set minimum support threshold in the global data range. Support measures the frequency of an item set appearing in the dataset. If the support of an item set does not meet the threshold requirement, it is considered not frequent enough globally and cannot be regarded as a frequent item set.

[0230] Further, S104-3: Build a Trie tree to store closed frequent item sets and merge the same prefixes, including:

[0231] S104-31: Use building a Trie tree to store closed frequent item sets; each node represents an item, and each path represents a closed frequent item set;

[0232] S104-32: During the storage process, use the child nodes with the same prefix for merging to reduce duplicate storage; when inserting a frequent item set, check if there is a common prefix. For example, for the frequent item sets "ABC" and "ABD", the prefix "AB" is common;

[0233] S104-33: If the prefix node does not exist, create a prefix node. If the "AB" node does not exist, create node "A", and then create node "B" under node "A";

[0234] S104-34: Based on the prefix nodes, create branches for different child nodes. For "ABC", create node "C" under node "B"; for "ABD", create node "D" under node "B".

[0235] S104-35: At each termination node (i.e., leaf node), store the support degree of this frequent item set.

[0236] It should be understood that the finally constructed Trie tree contains the organizational structures of all closed frequent item sets: each node represents an item, the path represents a frequent item set, the item sets with a common prefix share the path, and the leaf node stores the support degree of the corresponding frequent item set.

[0237] Furthermore, in the process of the said S104-32: storing, use the child nodes with the same prefix for merging, including:

[0238] In the process of storing frequent item sets, merge the child nodes with the same prefix by constructing a Trie tree: when inserting a new frequent item set, first start from the root node and layer by layer check whether there is a prefix node of this item set in the tree. If it exists, continue to insert the subsequent items along this path; if it does not exist, create a new node. In this way, if different frequent item sets share the same prefix (such as "ABC" and "ABD" sharing the prefix "AB"), their common prefix paths will only be stored once, thus reducing the storage amount of duplicate nodes and improving the storage efficiency.

[0239] Furthermore, in the said S104-34: based on the prefix nodes, create branches for different child nodes. For "ABC", create node "C" under node "B"; for "ABD", create node "D" under node "B"; including:

[0240] In the process of storing frequent item sets, merge the child nodes with the same prefix by constructing a Trie tree: when inserting a new frequent item set, first start from the root node and layer by layer check whether there is a prefix node of this item set in the tree. If it exists, continue to insert the subsequent items along this path; if it does not exist, create a new node. In this way, if different frequent item sets share the same prefix (such as "ABC" and "ABD" sharing the prefix "AB"), their common prefix paths will only be stored once, thus reducing the storage amount of duplicate nodes and improving the storage efficiency. Among them, ABC and ABD have a common prefix AB, and the subsequent C and D are different branches.

[0241] In this way, all item sets with the same prefix share the same prefix path. This merging method not only reduces the storage quantity of nodes, but also improves the query efficiency and optimizes the storage and search performance.

[0242] An appropriate organizational storage rule set can find association rules more quickly and effectively, thus accelerating the search and improving efficiency.

[0243] Further, in step S104-7, based on the LRU policy, removing data blocks includes:

[0244] Alluxio itself provides the LRU policy. When the utilization rate of the cache exceeds a set threshold (e.g., 80%), the system will activate the LRU policy to find and remove the data blocks that have not been accessed for the longest time. Considering the correlation between data blocks, the associated data blocks of the removed data blocks will also be removed.

[0245] The LRU (Least Recently Used) policy is a cache replacement algorithm mainly used to manage data in the cache. When the cache space is limited and space needs to be freed up for new data, the LRU policy selects the cache item that has not been used for the longest time for elimination. This algorithm assumes that the recently used data is likely to be used again, while the data that has not been used for a long time is less likely to be used.

[0246] Further, in step S104-7, marking the associated data blocks includes:

[0247] Each data block can have a status field in the metadata indicating its current marked status (such as "prefetch candidate" or "to be removed"). This metadata is stored in the cache together with the data block. Once the status changes, the change in the mark can be reflected by updating the metadata.

[0248] In the cross-domain scenario where the Alluxio cache is combined with the underlying storage, network latency may cause the I / O speed to be unable to keep up with the computing speed, thus affecting the performance of external services. By using prefetching and replacement technologies, hot data can be persistently retained in the Alluxio cache, reducing the number of remote calls and accesses to the underlying storage, and reducing data access latency and communication frequency under high load. Therefore, the present invention will adopt a cache policy based on the correlation between data blocks, use the parallelized CHARM algorithm to obtain the association rules between data blocks, and adjust the cache content according to the obtained association rules to implement cache prefetching and replacement operations, thereby improving the availability of Alluxio.

[0249] The current prefetch policy of Alluxio is that when the data requested by the user is not in the Alluxio cache, Alluxio will obtain the required data from the underlying storage system (such as HDFS, S3, etc.) and store it in the Alluxio cache.

[0250] Based on the current Alluxio prefetch policy, considering the correlation of data blocks and system load, design a more intelligent prefetch policy.

[0251] When a user accesses data block a, if data block a is not in the cache, Alluxio fetches the data block a from the underlying storage system, stores it in the cache, and locates several data blocks associated with a to check whether these data blocks are in the cache. If these associated data blocks are not in the cache, mark them as prefetch candidates.

[0252] Idle-time prefetching is triggered when the system load is lower than the set threshold. Select the data blocks marked as prefetch candidates from the prefetch candidate list, batch-fetch them from the underlying storage system, and store them in the cache.

[0253] When the cache utilization rate exceeds the set upper limit (e.g., 50%), based on the LRU policy, remove the least recently used data block, and mark the data blocks associated with the replaced data block a as candidate removed data blocks;

[0254] If the time since the last access of the associated data block exceeds the set time threshold from the current time, it means that the data block has also been rarely accessed recently. Consider removing them together to free up more cache space. If the time since the last access of the associated data block is less than the set time threshold from the current time, it means that the data block has been accessed recently, and keep them in the cache.

[0255] Embodiment 2

[0256] This embodiment provides a data storage and cache optimization system based on Alluxio, including:

[0257] An acquisition module, which is configured to: acquire a file to be written, and calculate the priority of the file according to four attributes of the file to be written; the four attributes of the file to be written include: file access frequency, file capacity, file importance level, and file freshness;

[0258] A writing module, which is configured to: store the file to be written in the Alluxio cache and the underlying storage according to the file transfer policy corresponding to the file priority;

[0259] A reading module, which is configured to: acquire the data to be read, determine whether the data to be read is in the cache, and if so, return the data in the cache to the user; if not, return the data in the underlying storage to the user;

[0260] A prefetching module, which is configured to: perform a prefetch operation on the data to be read and the data blocks associated with the data to be read according to the association rule, and perform a cache replacement operation if the cache utilization rate exceeds the set threshold.

[0261] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. The data storage and cache optimization method based on Alluxio is characterized by: include: Obtain a file to be written, and calculate the priority of the file according to four attributes of the file to be written; The four attributes of the file to be written include: file access frequency, file capacity, file importance level and file freshness; Store the files to be written into the Alluxio cache and underlying storage according to the file transfer policy corresponding to the file priority. Get the data to be read and determine whether the data to be read is in the cache. If yes, return the data in the cache to the user; if not, return the data in the underlying storage to the user; According to the association rules, pre-fetch operations are performed on the data to be read and the data blocks associated with the data to be read. If the cache usage rate exceeds the set threshold, a cache replacement operation is performed; The calculating the priority of the file according to the four attributes of the file to be written includes: Assume that for the file For , file size , Importance of Documents and the freshness of the data , standardize each indicator; combine the standardized indicators and calculate the priority of each file according to the weight ratio for: ,in, , , and is the weight, ; According to the file transfer policy corresponding to the file priority, the files to be written are stored in the Alluxio cache and underlying storage, including: If the file priority is high, the synchronous write mechanism and the dual copy mechanism are used; The synchronous writing mechanism means writing the file to be written into the Alluxio cache and the underlying storage at the same time; the dual copy mechanism means copying each data block to another working node; After the data is written into the underlying storage, data verification begins. The data verification is used to determine whether the written file is damaged. If the verification succeeds, it means that the file is written. If the verification fails, it means that the file needs to be rewritten. According to the file transfer policy corresponding to the file priority, the file to be written is stored in the Alluxio cache and the underlying storage. If the file priority is medium, the asynchronous writing mechanism and load balancing mechanism are used. The asynchronous writing mechanism means that the file to be written is first written to the Alluxio cache, and then written to the underlying storage; during the writing process, the intermediate state of the saved data is regularly checked through the checkpoint mechanism; According to the file transfer policy corresponding to the file priority, the files to be written are stored in the Alluxio cache and the underlying storage, which also includes: If the file priority is low, the asynchronous writing mechanism and data compression mechanism are used; The asynchronous writing mechanism means that the file to be written is first written to the Alluxio cache, and then written to the underlying storage when the system resources are idle; The load balancing mechanism includes: based on the real-time load situation, using a weighted polling algorithm to allocate transmission tasks: the weighted polling algorithm allocates an initial weight to each working node; During the polling process of the transmission task, the monitoring agent subnode of the working node collects the current load data of the current working node at the set time interval and uploads the load data to the central monitoring server; The central monitoring server determines whether the load of the current working node exceeds the set first threshold. If it exceeds the set first threshold, the weight of the current working node is reduced; determines whether the load of the current working node is lower than the set second threshold, and increases the weight of the current working node; if the load of the current working node is lower than the first threshold and higher than the second threshold, the weight of the current working node is kept unchanged; Assign tasks based on weights; The weighted round-robin algorithm assigns an initial weight to each working node, where the initial weight The calculation formula is: ; in, Represents a working node CPU power, Represents a working node Memory capacity, Represents a working node The network bandwidth Represents a working node I / O performance, Represents a working node The distance from the central node; , , , and is the weight coefficient, which is used to adjust the impact of each resource capacity and distance on the total weight; During the polling process of the transmission task, the monitoring agent subnode of the working node itself regularly collects the current load data of the current working node, including: During the polling process of the transmission task, the current load data of the working node is collected every minute. The current total load The calculation formula is: in, Represents a working node The CPU load, Represents a working node The memory load, Represents a working node The network bandwidth load, Represents a working node Input and output I / O load, according to the current load Adjustment Rights ; The central monitoring server determines whether the load of the current working node exceeds a set first threshold, and if so, reduces the weight of the current working node; determines whether the load of the current working node is lower than a set second threshold, and increases the weight of the current working node, specifically including: If the worker node Load Exceeds the preset high load threshold , then reduce the number of working nodes The weight of ; If the worker node Load Below the set low threshold , then add working nodes The weight of , the adjusted weight The calculation formula is: in, and is the coefficient for adjusting the weight.

2. The data storage and cache optimization method based on Alluxio as claimed in claim 1, characterized in that: The data compression mechanism refers to: when the network bandwidth utilization is higher than a set threshold, and the CPU utilization and the input and output interface utilization are lower than the set threshold, the LZ4 compression algorithm is selected to compress the file to be written; When the network bandwidth utilization is lower than the set threshold, and the CPU utilization and input / output interface utilization are higher than the set threshold, the Gzip compression algorithm is selected to compress the file to be written; When the network bandwidth utilization, CPU utilization, and input / output interface utilization are all within the set threshold range, the Zstandard compression algorithm is selected to compress the files to be written.

3. The data storage and cache optimization method based on Alluxio as claimed in claim 1, characterized in that: According to the association rules, the data to be read and the data blocks associated with the data to be read are pre-fetched. If the cache usage exceeds the set threshold, a cache replacement operation is performed, which specifically includes: (1): Set the size of the sliding window and the step size of the sliding window; (2): Calculate the closed frequent itemsets for the data in the sliding window; (3): Build a Trie tree to store closed frequent itemsets, merge the same prefixes, and obtain the association relationship between data blocks based on the Trie tree; (4): Determine whether the data accessed by the user is in the cache. If so, proceed to (5); if not, proceed to (7); (5): Determine whether the data block associated with the current data block is in the cache. If so, return to (1); if not, mark the data block associated with the current data block as a "pre-fetch candidate" and enter (6); the associated data blocks are obtained based on the Trie tree. If data block A and data block B belong to the same closed frequent item set, it means that data block A and data block B are associated. Otherwise, it means that data block A and data block B are not associated. The pre-fetch candidate refers to the candidate data block waiting to be selected; (6): Determine whether the load is lower than the set threshold. If not, return to (1); if yes, store the data block marked as "prefetch candidate" in the cache and return to (1); (7): Determine whether the cache usage rate exceeds the set threshold. If not, store the data block accessed by the user in the cache and enter (5); if yes, remove the data block based on the LRU strategy and mark the associated data blocks; enter (8); (8): Determine whether the difference between the most recent access time of the associated data block and the current time exceeds the set time threshold. If so, remove the current data block and return to (7); if not, return to (7).

4. The data storage and cache optimization method based on Alluxio as claimed in claim 3 is characterized in that: The size of the sliding window and the step size of the sliding window are set, including: the size of the sliding window is the amount of data contained in the sliding window, and the step size of the sliding window is the amount of new data added each time the window slides; the data in the window is regarded as a queue, if the amount of data in the queue exceeds the set window size W, the earliest data added to the queue is removed to ensure that there are only the latest W data in the queue; the removed data is marked as "old data"; when new data arrives, the new data is added to the current window and marked as "new data"; the data previously marked as "new data" has its "new data" mark removed after a new round of data is added.

5. The data storage and cache optimization method based on Alluxio as claimed in claim 3, characterized in that: For the data in the sliding window, calculate the closed frequent itemsets, including: Each time the window slides, only the data marked as "new data" is processed, and the closed frequent itemsets of the "new data" in the window are calculated; the closed frequent itemsets of the "new data" in the window are merged with the closed frequent itemsets of all the data in the window, and the closed frequent itemset list is updated; Divide the entire window into several subsets; assign the subsets to different working nodes; Based on the closed frequent itemset list, each working node calculates the local closed frequent itemset and performs pruning in the local range; After each working node completes the local closed frequent itemset mining, it first performs local merging on some nodes, and then sends the merged results to the central monitoring server for final merging; Remove duplicate closed frequent item sets generated by each node, merge identical closed frequent item sets, and recalculate the support of the merged closed frequent item sets to ensure that their support is correct globally. The final pruning is performed on the merged global closed frequent itemsets to remove itemsets that do not meet the support threshold.

6. The data storage and cache optimization system based on Alluxio is characterized by: include: An acquisition module is configured to: acquire a file to be written, and calculate a priority of the file according to four attributes of the file to be written; The four attributes of the file to be written include: file access frequency, file capacity, file importance level and file freshness; The write module is configured to store the files to be written into the Alluxio cache and underlying storage according to the file transfer policy corresponding to the file priority; A reading module is configured to: obtain the data to be read, determine whether the data to be read is in the cache, and if so, return the data in the cache to the user; if not, return the data in the underlying storage to the user; A pre-fetch module is configured to: perform a pre-fetch operation on the data to be read and the data blocks associated with the data to be read according to the association rule, and perform a cache replacement operation if the cache usage rate exceeds a set threshold; The calculating the priority of the file according to the four attributes of the file to be written includes: Assume that for the file For , file size , Importance of Documents and the freshness of the data , standardize each indicator; combine the standardized indicators and calculate the priority of each file according to the weight ratio for: ,in, , , and is the weight, ; According to the file transfer policy corresponding to the file priority, the files to be written are stored in the Alluxio cache and underlying storage, including: If the file priority is high, the synchronous write mechanism and the dual copy mechanism are used; The synchronous writing mechanism means writing the file to be written into the Alluxio cache and the underlying storage at the same time; the dual copy mechanism means copying each data block to another working node; After the data is written into the underlying storage, data verification begins. The data verification is used to determine whether the written file is damaged. If the verification succeeds, it means that the file is written. If the verification fails, it means that the file needs to be rewritten. According to the file transfer policy corresponding to the file priority, the file to be written is stored in the Alluxio cache and the underlying storage. If the file priority is medium, the asynchronous writing mechanism and load balancing mechanism are used. The asynchronous writing mechanism means that the file to be written is first written to the Alluxio cache, and then written to the underlying storage; during the writing process, the intermediate state of the saved data is regularly checked through the checkpoint mechanism; According to the file transfer policy corresponding to the file priority, the files to be written are stored in the Alluxio cache and the underlying storage, which also includes: If the file priority is low, the asynchronous writing mechanism and data compression mechanism are used; The asynchronous writing mechanism means that the file to be written is first written to the Alluxio cache, and then written to the underlying storage when the system resources are idle; The load balancing mechanism includes: based on the real-time load situation, using a weighted polling algorithm to allocate transmission tasks: the weighted polling algorithm allocates an initial weight to each working node; During the polling process of the transmission task, the monitoring agent subnode of the working node collects the current load data of the current working node at the set time interval and uploads the load data to the central monitoring server; The central monitoring server determines whether the load of the current working node exceeds the set first threshold. If it exceeds the set first threshold, the weight of the current working node is reduced; determines whether the load of the current working node is lower than the set second threshold, and increases the weight of the current working node; if the load of the current working node is lower than the first threshold and higher than the second threshold, the weight of the current working node is kept unchanged; Assign tasks based on weights; The weighted round-robin algorithm assigns an initial weight to each working node, where the initial weight The calculation formula is: ; in, Represents a working node CPU power, Represents a working node Memory capacity, Represents a working node The network bandwidth Represents a working node I / O performance, Represents a working node The distance from the central node; , , , and is the weight coefficient, which is used to adjust the impact of each resource capacity and distance on the total weight; During the polling process of the transmission task, the monitoring agent subnode of the working node itself regularly collects the current load data of the current working node, including: During the polling process of the transmission task, the current load data of the working node is collected every minute. The current total load The calculation formula is: in, Represents a working node The CPU load, Represents a working node The memory load, Represents a working node The network bandwidth load, Represents a working node Input and output I / O load, according to the current load Adjustment Rights ; The central monitoring server determines whether the load of the current working node exceeds a set first threshold, and if so, reduces the weight of the current working node; determines whether the load of the current working node is lower than a set second threshold, and increases the weight of the current working node, specifically including: If the worker node Load Exceeds the preset high load threshold , then reduce the number of working nodes The weight of ; If the worker node Load Below the set low threshold , then add working nodes The weight of , the adjusted weight The calculation formula is: in, and is the coefficient for adjusting the weight.