Mass data storage optimization method and device and storage medium
By analyzing the access frequency and time of massive data, dividing it into different categories and adopting appropriate compression and deduplication processing, combining multi-level caching and real-time resource provisioning, the problem that traditional memory management methods cannot efficiently store and quickly access massive data, significantly improving system performance and storage efficiency.
Patent Information
- Application Number
- CN202510211371.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional memory management methods cannot effectively solve the needs of efficient storage and fast access of massive data, especially in frequently accessed data processing, which has problems such as low cache hit rate and data access latency.
By analyzing the access frequency and access time of the data, the data is divided into hot data, temperature data and cold data, and the appropriate compression algorithm is used for processing. Before storing data, use the data fingerprint for deduplication and monitor memory usage and compression algorithm performance in real time to dynamically provision system resources. At the same time, set up multi-level cache and cache it to the corresponding cache layer according to the data type.
It significantly improves storage space utilization and overall system performance, reduces system response time, improves user experience, and ensures the efficiency of the compression and decompression process.
Smart Images

Figure CN120216504A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, and storage medium for optimizing the storage of massive data. Background Art
[0002] With the development of big data technology, the rapid growth of data volume has brought huge challenges to storage, transmission, and processing. Traditional memory management methods can no longer meet the requirements of efficient storage and fast access to massive data. Especially when dealing with frequently accessed data, how to optimize the memory storage space and improve the system response speed has become an urgent problem to be solved.
[0003] Existing memory optimization technologies mainly rely on traditional cache acceleration strategies. However, when facing a large amount of data, these methods often face problems such as low cache hit rate and high data access latency. In addition, existing compression and deduplication technologies often have insufficient optimization in large-scale data applications. For example, the compression algorithm cannot be flexibly adjusted, and the performance of the deduplication process is relatively low. Summary of the Invention
[0004] This application provides a method, device, and storage medium for optimizing the storage of massive data, which realizes efficient storage and fast access to massive data, and significantly improves the storage space utilization rate and the overall system performance.
[0005] Analyze the target data according to the access frequency and / or access time of the data, and divide the target data into hot data, warm data, and cold data; For the hot data, warm data, and cold data, use corresponding compression algorithms for compression processing respectively; Before storing the target data, deduplicate the target data according to the fingerprint of the target data; Real-time monitor the usage of memory, the access frequency of the target data, and the performance of the compression algorithm, and allocate system resources according to the monitoring results; Set up a multi-level cache, and cache different types of target data into the multi-level cache according to the types into which the target data is divided.
[0006] On the other hand, this application provides a device for optimizing the storage of massive data, and the device includes: A classification module, configured to analyze the target data according to the access frequency and / or access time of the data, and divide the target data into hot data, warm data, and cold data; A compression module, configured to use corresponding compression algorithms for compression processing respectively for the hot data, warm data, and cold data; A deduplication module, configured to deduplicate the target data according to the fingerprint of the target data before storing the target data; A deployment module, configured to monitor the usage of memory, the access frequency of the target data, and the performance of the compression algorithm in real time, and deploy system resources according to the monitoring results; A cache module, configured to set up multiple levels of caches, and cache different types of target data into the multiple levels of caches according to the types into which the target data is divided.
[0007] In a third aspect, the present application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the technical solution of the above-mentioned massive data storage optimization method are implemented.
[0008] In a fourth aspect, the present application provides a storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the technical solution of the above-mentioned massive data storage optimization method are implemented.
[0009] As can be seen from the technical solutions provided by the present application above, on the one hand, for hot data, warm data, and cold data, corresponding compression algorithms are respectively used for compression processing, which can ensure the efficiency of the compression and decompression processes and avoid negative impacts on system performance; on the other hand, before storing the target data, the target data is deduplicated according to the fingerprint of the target data, which can improve the utilization rate of the storage space and reduce unnecessary data redundancy; in the third aspect, multiple levels of caches are set up, and different types of target data are cached into the multiple levels of caches according to the types into which the target data is divided, so that the frequently accessed data can be stored in the high-speed cache, greatly improving the data access speed, reducing the system response time, and enhancing the user experience. Description of the Drawings
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 is a flowchart of the massive data storage optimization method provided by the embodiment of the present application; Figure 2 is a structural schematic diagram of the massive data storage optimization device provided by the embodiment of the present application; Figure 3It is a schematic diagram of the structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0014] In this specification, adjectives such as first and second can only be used to distinguish one element or action from another element or action, and do not necessarily require or imply any actual such relationship or order. Where circumstances permit, reference to an element or component or step (etc.) should not be construed as being limited to only one of the element, component, or step, but may be one or more of the element, component, or step, etc.
[0015] In this specification, for ease of description, the dimensions of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0016] With the development of big data technology, the rapid growth of data volume has brought huge challenges to storage, transmission, and processing. Traditional memory management methods can no longer meet the requirements of efficient storage and fast access to massive data. Especially when dealing with frequently accessed data, how to optimize the memory storage space and improve the system response speed has become an urgent problem to be solved. Existing memory optimization technologies mainly rely on traditional cache acceleration strategies, but these methods often face problems such as low cache hit rate and high data access latency when dealing with large-scale data volumes. In addition, existing compression and deduplication technologies often have insufficient optimization problems in large-scale data applications. For example, the compression algorithm cannot be flexibly adjusted, and the performance of the deduplication process is low.
[0017] In view of the above problems in the prior art, the present application proposes a method for optimizing the storage of massive data. The flowchart is as shown in the attached Figure 1 figure, and mainly includes steps S101 to S105, which are described in detail as follows: Step S101: Analyze the target data according to the access frequency and / or access time of the data, and divide the target data into hot data, warm data, and cold data.
[0018] In a large-scale data processing system, the access frequency of data often has significant temporal characteristics. Some data is frequently accessed within a short period, while some data is rarely accessed. Analyzing target data based on its access frequency and / or access time, and classifying the target data into hot data, warm data, and cold data. Subsequently, adopting different processing methods according to the hotness and coldness of the data can save resources and improve the performance of the system. As an embodiment of this application, analyzing target data based on its access frequency and / or access time and classifying the target data into hot data, warm data, and cold data can be as follows: using a time window algorithm to count the access frequency of target data; determining target data with an access count greater than a first access threshold within the time window as hot data, determining target data with an access count less than a second access threshold as cold data, and determining target data with an access count between the first access threshold and the second access threshold as warm data, where the first access threshold is greater than the second access threshold.
[0019] Although the time window algorithm involved in the above embodiments can be used to statistically analyze data characteristics within a certain time period, the size of the time window directly affects the accuracy of data cold and hot classification. For example, if the time window is too small, long-term access patterns may not be captured, resulting in frequent adjustments. Conversely, if the time window is too large, changes in data access cannot be reflected in a timely manner, leading to classification lag. From another perspective, data access patterns may change over time, such as burst traffic or periodic access. A fixed time window cannot reflect these changes in a timely manner, resulting in inaccurate classification of data cold and hot. A smaller time window requires more frequent calculations, increasing the system burden, while a larger time window may reduce the response speed. Dynamic adjustment can find a balance between the two. In addition, through dynamic adjustment, the system can automatically optimize resource allocation according to the current load. For example, when the access is stable, a larger time window can be used to reduce calculations, and when there are fluctuations, a smaller time window can be used to improve the accuracy of data cold and hot classification. Finally, different application scenarios may have different data access characteristics. Dynamically adjusting the size of the time window makes the method more general and adaptable. Therefore, it is necessary to dynamically adjust the time window of the above embodiments to balance classification accuracy and computational overhead in different situations. That is, the above embodiments may further include: predicting the access frequency of target data; dynamically adjusting the size of the time window in the time window algorithm according to the predicted access frequency of the target data. Specifically, the implementation of predicting the access frequency of the target data may be achieved by establishing a data access frequency trend prediction model. This model is trained based on the historical data access frequency sequence using a recurrent neural network (RNN) or a long short-term memory network (LSTM) in deep learning to predict the data access frequency trend in the future for a period of time. As for dynamically adjusting the size of the time window in the time window algorithm according to the predicted access frequency of the target data, it may specifically be: making a comprehensive judgment based on the prediction result and the currently monitored data access frequency fluctuation situation. When it is predicted that the access frequency of the target data will increase significantly and the current fluctuation is large, an exponential reduction of the time window is adopted to quickly capture the change in data heat. When it is predicted that the access frequency of the target data is stable and the current fluctuation is small, a logarithmic increase of the time window is adopted to reduce the computational overhead while ensuring the accuracy of data heat judgment. By dynamically adjusting the size of the time window, the system can effectively improve the cache hit rate and optimize the use of memory storage space.
[0020] As an embodiment of the present application, dynamically adjusting the size of the time window in the time window algorithm according to the predicted access frequency of the target data can also be achieved through steps S1011 to S1014, which are described in detail as follows: Step S1011: Initialize the quantization model for time window prediction.
[0021] Build a quantum optimization model and model the problem of adjusting the time window as a quantum optimization problem. For example, the goal is to maximize the capture accuracy of data heat change while minimizing the computational overhead.
[0022] Step S1012: Use a quantum algorithm to dynamically adjust the size of the time window in the time window algorithm.
[0023] Specifically, the quantum annealing algorithm (Quantum Annealing) or the quantum approximate optimization algorithm (Quantum Approximate Optimization Algorithm, QAOA) can be used to dynamically adjust the size of the time window. Quantum annealing can quickly find the global optimal solution, thus avoiding falling into the local optimal solution.
[0024] Step S1013: Update the quantum state.
[0025] As the data access pattern changes, the quantum optimization model updates the size of the time window in real time according to the access frequency and heat fluctuation of the data to ensure the optimal window adjustment strategy.
[0026] Step S1014: Optimize the adjustment strategy of the time window.
[0027] The quantum optimization model continuously performs feedback and iterative updates. As the system runs, the optimization effect becomes more and more accurate and adapts to different data access patterns.
[0028] Compared with the traditional time window adjustment scheme, the above dynamic time window adjustment method based on quantum optimization can find the global optimal solution more quickly. Especially when there are a large number of data access changes, it can provide an accurate time window adjustment strategy. Therefore, it is a time window adjustment method with strong adaptability and can handle complex and unpredictable access patterns.
[0029] As another embodiment of this application, dynamically adjusting the size of the time window in the time window algorithm according to the predicted access frequency of the target data can also be achieved through steps S’1011 to S’1013, and the detailed description is as follows: Step S’1011: Build an environment for reinforcement learning.
[0030] In the embodiment of this application, building an environment for reinforcement learning specifically includes defining the environment and space of reinforcement learning, including the state space and the action space. Among them, the state space can be the current data access frequency, access pattern, the size of the time window, and so on. The action space is the possible actions for adjusting the size of the time window (such as shrinking or expanding the time window, etc.).
[0031] Step S’1012: Design the reward function for the action.
[0032] Specifically, a reward function can be set to quantify the effect of the time window adjustment action. For example, the reward function can be defined based on the balance between the accuracy of the captured data heat change and the computational cost. When the time window can effectively capture the heat change and minimize the computational cost, the reward value is higher.
[0033] Step S’1013: Optimize the adjustment strategy of the time window.
[0034] Through multiple rounds of interaction with the environment, the system uses reinforcement learning algorithms (such as Q-learning, Deep Q-Network, etc.) to learn how to adjust the time window size, and the feedback after each adjustment will be used to optimize the strategy. Over time, the system will continuously adjust its strategy, optimize the selection of the time window, and make it adapt to the changing data access patterns.
[0035] Compared with the traditional time window adjustment scheme, the above-mentioned dynamic time window adjustment method based on reinforcement learning can adapt to the dynamically changing environment, has strong adaptability, and automatically optimizes the time window adjustment strategy through continuous trial and error and feedback without manual intervention. Therefore, it can combine with deep reinforcement learning to handle complex non-linear problems.
[0036] Step S102: For hot data, warm data, and cold data, use corresponding compression algorithms for compression processing.
[0037] As an embodiment of the present application, for hot data, warm data, and cold data, the corresponding compression algorithms can be used for compression processing respectively as follows: use a fast compression algorithm to compress hot data; use an algorithm that balances compression ratio and compression speed to compress warm data; and use a high compression ratio algorithm to compress cold data. Among them, using a fast compression algorithm to compress hot data can be to use the LZ4 algorithm to compress hot data. Specifically, when using the LZ4 algorithm to compress hot data, an adaptive parameter adjustment mechanism based on a genetic algorithm can be set, that is: define a set of parameter sets related to the performance of the LZ4 algorithm as gene encoding, and use the comprehensive evaluation index of compression ratio and compression speed as the fitness function; during the compression process, according to the specific characteristics of hot data, such as data type, data length, etc., randomly generate an initial population, and through the selection, crossover, and mutation operations of the genetic algorithm, iteratively optimize the parameter set, so that the parameters of the LZ4 algorithm can be dynamically adjusted under hot data with different characteristics to achieve the optimal compression efficiency. In summary, using the LZ4 algorithm to compress hot data mainly includes key links such as initializing the population, designing the fitness function, genetic operations, and the optimization process. Among them, initializing the population means defining the encoding method of compression parameters, such as compression block size, dictionary size, etc., converting these parameters into the form of a genome, and initializing a population, where each individual represents a combination of compression parameters; the fitness function evaluates the pros and cons of each individual based on the size of the compressed file and the compression / decompression speed. Specifically, the trade-off between the compression ratio and the compression and decompression time determines the fitness value; genetic operations generate the next generation of individuals through selection, crossover, and mutation operations. The selection operation is based on the fitness value, and the crossover and mutation operations are to explore better combinations of compression parameters. As for the optimization process, it means that after multiple generations of iteration, the genetic algorithm can find a set of optimal compression parameters, thereby optimizing the LZ4 algorithm to achieve the best compression effect.
[0038] The temperature data is compressed using an algorithm that balances the compression ratio and the compression speed. Specifically, it can be a method that combines an improved Huffman compression algorithm with run-length encoding (RLE). That is, first, run-length encoding is performed on the temperature data to reduce consecutive repeating elements in the data. Then, for the data after run-length encoding, a Huffman tree is dynamically constructed based on the probability distribution of the data to increase the compression ratio. At the same time, the generation process of Huffman encoding is optimized to reduce the time overhead of encoding and decoding, thereby balancing the compression ratio and the compression speed. As for compressing the cold data using a high-compression ratio algorithm, specifically, a compression method based on wavelet transform and arithmetic coding can be used. That is, first, wavelet transform is performed on the cold data to decompose the data into different frequency sub-bands, highlighting the main features of the data and reducing data redundancy. Then, the transformed coefficients are quantized, and arithmetic coding is used to encode the quantized coefficients. In this way, a high compression ratio for the cold data is achieved.
[0039] Step S103: Before storing the target data, deduplicate the target data according to the fingerprint of the target data.
[0040] For massive data, if there are duplicate parts, storing the duplicate data will greatly occupy storage space. To avoid this situation, in the embodiments of the present application, before storing the target data, the target data is deduplicated according to the fingerprint of the target data. Specifically, before storing the target data, deduplicating the target data according to the fingerprint of the target data can be achieved through steps S1031 to S1034, which are described in detail as follows: Step S1031: Construct a data fingerprint library based on a Bloom filter.
[0041] In the embodiments of the present application, constructing a data fingerprint library based on a Bloom filter can specifically be a structure that combines a Bloom filter and a counting Bloom filter. That is, first, the Bloom filter is used to quickly determine whether a data fingerprint may exist in the fingerprint library. When the Bloom filter determines that a data fingerprint may exist, further verification is performed through the counting Bloom filter. The counting Bloom filter can not only record whether a fingerprint exists but also record the number of times the fingerprint appears. For data with a repetition count reaching a certain threshold, a more efficient storage method is adopted, such as storing it in a dedicated duplicate data storage area. And during the subsequent deduplication process, the processed duplicate data is skipped according to the records of the counting Bloom filter, thereby greatly reducing the computational overhead in the data deduplication process.
[0042] Step S1032: Calculate the fingerprint of the target data based on a hash algorithm.
[0043] Specifically, calculating the fingerprint of the target data based on the hash algorithm can be achieved by introducing the context information of the data on the basis of the traditional hash algorithm; taking the byte information of a certain length before and after the data as additional input and participating in the hash calculation process together with the original data; dynamically adjusting the initial seed value of the hash algorithm and calculating an adaptive seed value according to the size and type of the data. This solution reduces the hash collision rate at the cost of increasing the computational complexity and accuracy of the hash algorithm, and can calculate the fingerprint of the target data more quickly and accurately.
[0044] Step S1033: Match the fingerprint of the target data with the fingerprints in the data fingerprint library.
[0045] Step S1034: If, after step S1033, the fingerprint of the target data matches the fingerprints in the data fingerprint library successfully, determine that the target data is duplicate data.
[0046] Step S104: Monitor the memory usage, the access frequency of the target data, and the performance of the compression algorithm in real time and allocate system resources according to the monitoring results.
[0047] Specifically, step S104 can be implemented as follows: By establishing a performance model to predict the performance of the system under different resource allocation strategies, this performance model comprehensively considers multiple factors such as memory usage rate, data access latency, compression and decompression time, etc., and selects the optimal resource allocation strategy according to the prediction results. It should be noted that the performance model in the above embodiments is trained using the reinforcement learning algorithm, that is, the state space is defined as the set of various monitoring parameters of the system, including memory usage rate, data access frequency, compression algorithm performance, etc.; the action space is defined as various possible resource allocation strategies, such as adjusting the compression algorithm, allocating memory space, etc.; through the continuous interaction between the agent and the system environment, according to the reward signal feedback by the system (calculated based on comprehensive indicators such as reduced memory usage rate, shortened data access latency, improved compression and decompression efficiency, etc.), continuously optimize the performance model, so that the agent can learn the optimal resource allocation strategy to achieve more intelligent and efficient resource allocation. By monitoring and dynamically allocating resources in real time, it can ensure that the system can operate efficiently under different load conditions, avoid resource waste and performance bottlenecks, and improve the overall performance and reliability of the system.
[0048] From steps S1030 and S104 of the above embodiments, on the one hand, by combining the hash algorithm and the Bloom filter, the traditional data deduplication method can be improved, effectively reducing the storage space of duplicate data; on the other hand, since the redundant data is reduced, the time and resources required for the system to perform data retrieval and processing are also correspondingly reduced, further improving the data access speed.
[0049] Step S105: Set up a multi-level cache, and cache different types of target data into the multi-level cache according to the types into which the target data is classified.
[0050] As an embodiment of the present application, setting up a multi-level cache and caching different types of target data into the multi-level cache according to the types into which the target data is classified can be implemented through steps S1051 to S1053, and the detailed description is as follows: Step S1051: Set up a primary cache and a secondary cache.
[0051] Considering that the target data is only classified into three categories, namely hot data, cold data, and warm data, etc., the embodiment of the present application can set up a primary cache and a secondary cache.
[0052] Step S1052: Cache the hot data into the primary cache, and cache the warm data and / or cold data into the secondary cache.
[0053] In the embodiment of the present application, the primary cache can adopt a high-speed SRAM memory to store the hottest data, that is, the hot data in the target data, and the secondary cache adopts a DRAM memory with a larger capacity to store relatively hotter data, that is, warm data and / or cold data.
[0054] Step S1053: Dynamically adjust the data in the primary cache and / or secondary cache through a cache replacement algorithm.
[0055] Specifically, dynamically adjusting the data in the primary cache and / or secondary cache through a cache replacement algorithm can be: dynamically adjusting the heat factor according to the access frequency and access time interval of the data; when making a cache replacement decision, comprehensively evaluate and adjust the data in the primary cache and / or secondary cache in combination with the heat factor. Specifically, when making a cache replacement decision, not only consider the recent usage time of the data, but also comprehensively evaluate it in combination with the heat factor. For data with a higher heat factor, even if its recent usage time is earlier, it is preferentially retained in the cache, so as to improve the cache hit rate and further optimize the fast response ability to frequently accessed data.
[0056] From the above attachment Figure 1As can be seen from the exemplary method for optimizing mass data storage, on the one hand, for hot data, warm data, and cold data, corresponding compression algorithms are respectively used for compression processing, which can ensure the efficiency of the compression and decompression processes and avoid negative impacts on system performance; on the other hand, before storing the target data, duplicate elimination is performed on the target data according to the fingerprint of the target data, which can improve the utilization rate of storage space and reduce unnecessary data redundancy; thirdly, a multi-level cache is set, and different types of target data are cached in the multi-level cache according to the types into which the target data is divided, so that frequently accessed data can be stored in the high-speed cache, greatly improving the speed of data access, reducing the system response time, and enhancing the user experience.
[0057] Please refer to the appendix Figure 2 , which is a mass data storage optimization device provided by an embodiment of the present application. The device may include a classification module 201, a compression module 202, a duplicate elimination module 203, a deployment module 204, and a cache module 205, which are described in detail as follows: The classification module 201 is used to analyze the target data according to the access frequency and / or access time of the data, and divide the target data into hot data, warm data, and cold data; The compression module 202 is used to perform compression processing on hot data, warm data, and cold data respectively by using corresponding compression algorithms; The duplicate elimination module 203 is used to perform duplicate elimination on the target data according to the fingerprint of the target data before storing the target data; The deployment module 204 is used to monitor the usage of memory, the access frequency of the target data, and the performance of the compression algorithm in real time, and deploy system resources according to the monitoring results; The cache module 205 is used to set a multi-level cache, and cache different types of target data in the multi-level cache according to the types into which the target data is divided.
[0058] From the above appendix Figure 2 As can be seen from the exemplary mass data storage optimization device, on the one hand, for hot data, warm data, and cold data, corresponding compression algorithms are respectively used for compression processing, which can ensure the efficiency of the compression and decompression processes and avoid negative impacts on system performance; on the other hand, before storing the target data, duplicate elimination is performed on the target data according to the fingerprint of the target data, which can improve the utilization rate of storage space and reduce unnecessary data redundancy; thirdly, a multi-level cache is set, and different types of target data are cached in the multi-level cache according to the types into which the target data is divided, so that frequently accessed data can be stored in the high-speed cache, greatly improving the speed of data access, reducing the system response time, and enhancing the user experience.
[0059] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. AsFigure 3 As shown, the electronic device 3 in this embodiment mainly includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program for optimizing the storage of massive data. When the processor 30 executes the computer program 32, it implements the steps in the embodiment of the above-mentioned massive data storage optimization method, such as Figure 1 the steps S101 to S105 shown. Alternatively, when the processor 30 executes the computer program 32, it implements the functions of each module / unit in the above-mentioned device embodiments, such as Figure 2 the functions of the classification module 201, compression module 202, deduplication module 203, allocation module 204, and cache module 205 shown.
[0060] Exemplarily, the computer program 32 for optimizing the storage of massive data mainly includes: analyzing target data according to the access frequency and / or access time of the data, and dividing the target data into hot data, warm data, and cold data; respectively performing compression processing on the hot data, warm data, and cold data using corresponding compression algorithms; before storing the target data, deduplicating the target data according to the fingerprint of the target data; monitoring the usage of the memory, the access frequency of the target data, and the performance of the compression algorithm in real time, and allocating system resources according to the monitoring results; setting up multiple levels of caches, and caching different types of target data to the multiple levels of caches according to the types into which the target data is divided. The computer program 32 can be divided into one or more modules / units. One or more modules / units are stored in the memory 31 and executed by the processor 30 to complete this application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3. For example, the computer program 32 can be divided into the functions of the classification module 201, compression module 202, deduplication module 203, allocation module 204, and cache module 205 (modules in the virtual device). The specific functions of each module are as follows: The classification module 201 is used to analyze target data according to the access frequency and / or access time of the data, and divide the target data into hot data, warm data, and cold data; the compression module 202 is used to respectively perform compression processing on the hot data, warm data, and cold data using corresponding compression algorithms; the deduplication module 203 is used to deduplicate the target data according to the fingerprint of the target data before storing the target data; the allocation module 204 is used to monitor the usage of the memory, the access frequency of the target data, and the performance of the compression algorithm in real time, and allocate system resources according to the monitoring results; the cache module 205 is used to set up multiple levels of caches, and cache different types of target data to the multiple levels of caches according to the types into which the target data is divided.
[0061] The electronic device 3 may include but is not limited to a processor 30 and a memory 31. Those skilled in the art can understand that Figure 3 merely examples of the electronic device 3, which do not constitute a limitation on the electronic device 3, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.
[0062] The so-called processor 30 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0063] The memory 31 may be an internal storage unit of the electronic device 3, such as the hard disk or memory of the electronic device 3. The memory 31 may also be an external storage device of the electronic device 3, such as a plug-in hard disk equipped on the electronic device 3, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 31 may also include both an internal storage unit and an external storage device of the electronic device 3. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 may also be used to temporarily store data that has been output or is to be output.
[0064] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above-mentioned device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0065] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0066] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0067] In the embodiments provided in this application, it should be understood that the disclosed device / equipment and method can be implemented in other ways. For example, the device / equipment embodiments described above are only illustrative. For example, the division of modules or units is only a logical functional division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the device or unit can be in electrical, mechanical or other forms.
[0068] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0069] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0070] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program for the massive data storage optimization method can be stored in a storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments, that is, analyze the target data according to the access frequency and / or access time of the data, and divide the target data into hot data, warm data, and cold data; for hot data, warm data, and cold data, respectively, perform compression processing using corresponding compression algorithms; before storing the target data, perform deduplication on the target data according to the fingerprint of the target data; monitor the usage of the memory, the access frequency of the target data, and the performance of the compression algorithm in real time and allocate system resources according to the monitoring results; set up multiple levels of caches, and cache different types of target data into the multiple levels of caches according to the types into which the target data is divided. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The storage medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the storage medium does not include electrical carrier signals and telecommunication signals.
[0071] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application. The specific implementation manners described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above is only the specific implementation manner of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application should all be included in the protection scope of the present invention.
Claims
1. A method for optimizing mass data storage, characterized in that: The method comprises: Analyze the target data according to the access frequency and / or access time of the data, and divide the target data into hot data, warm data, and cold data; For the hot data, warm data and cold data, corresponding compression algorithms are used to perform compression processing respectively; Before storing the target data, deduplicating the target data according to the fingerprint of the target data; Monitor the usage of memory, the access frequency to the target data and the performance of the compression algorithm in real time and allocate system resources according to the monitoring results; A multi-level cache is set, and different types of target data are cached in the multi-level cache according to the types into which the target data are divided.
2. The method for optimizing mass data storage according to claim 1, characterized in that: The analyzing the target data according to the access frequency and / or access time of the data and dividing the target data into hot data, warm data and cold data includes: Using a time window algorithm to count the access frequency of the target data; Target data whose access times within a time window are greater than a first access threshold are determined as the hot data, target data whose access times are less than a second access threshold are determined as the cold data, and target data whose access times are between the first access threshold and the second access threshold are determined as the warm data, and the first access threshold is greater than the second access threshold.
3. The method for optimizing mass data storage according to claim 2, characterized in that: The method further comprises: predicting the access frequency of the target data; The time window size in the time window algorithm is dynamically adjusted according to the predicted access frequency of the target data.
4. The method for optimizing mass data storage according to claim 1, characterized in that: The hot data, warm data and cold data are compressed using corresponding compression algorithms respectively, including: Compressing the hot data using a fast compression algorithm; compressing the warm data using an algorithm that balances compression ratio and compression speed; and The cold data is compressed using a high compression ratio algorithm.
5. The method for optimizing mass data storage according to claim 1, characterized in that: Before storing the target data, deduplication of the target data is performed according to the fingerprint of the target data, including: Build a data fingerprint library based on Bloom filter; Calculating the fingerprint of the target data based on a hash algorithm; matching the fingerprint of the target data with the fingerprint in the data fingerprint library; If the match is successful, it is determined that the target data is duplicate data.
6. The method for optimizing mass data storage according to claim 1, characterized in that: The step of setting a multi-level cache and caching different types of target data into the multi-level cache according to the types into which the target data are divided includes: Set up the first-level cache and the second-level cache; Cache the hot data in the first-level cache, and cache the warm data and / or cold data in the second-level cache; The data in the first-level cache and / or the second-level cache is dynamically adjusted through a cache replacement algorithm.
7. The method for optimizing mass data storage according to claim 6, characterized in that: The dynamically adjusting the data in the first-level cache and / or the second-level cache by using a cache replacement algorithm includes: Dynamically adjust the heat factor based on the access frequency and access time interval of the data; When making a cache replacement decision, the data in the first-level cache and / or the second-level cache is adjusted after a comprehensive evaluation in combination with the heat factor.
8. A mass data storage optimization device, characterized in that: The device comprises: A classification module, used to analyze target data according to access frequency and / or access time of the data, and classify the target data into hot data, warm data and cold data; A compression module, used for compressing the hot data, warm data and cold data using corresponding compression algorithms respectively; a deduplication module, used for deduplicating the target data according to the fingerprint of the target data before storing the target data; A deployment module, used to monitor the usage of memory, the access frequency to the target data and the performance of the compression algorithm in real time and deploy system resources according to the monitoring results; The cache module is used to set a multi-level cache and cache different types of target data into the multi-level cache according to the types into which the target data are divided.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Storage system optimization method and device, electronic equipment, medium and product
CN120909532A
Multi-level caching method based on distributed storage and related equipment
CN121009121A
Data caching and updating method, equipment and medium
CN121051122A
Data processing method and device, computer equipment and storage medium
CN121188024A
Block chain storage optimization method, system and equipment based on hierarchical compression and dynamic fragmentation and medium
CN121350146A