Streaming data storage method combined with adaptive compression
By combining first-order and second-order analysis tasks with B+ tree index and hash index, the storage priority and position of data segments are dynamically adjusted, solving the problem of storage space fragmentation in the adaptive compression method, and achieving efficient data storage and access.
Patent Information
- Application Number
- CN202510454983.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Adaptive compression methods lead to fragmentation of storage space and waste of resources during storage processes, and it is impossible to effectively manage the distribution of data segments and optimize the storage structure.
The characteristics of the data segment are evaluated through first-order and second-order analysis tasks, combined with B+ tree index and hash index, dynamically adjust the storage priority and position of the data segment, perform fragmented management, optimize the compression algorithm selection of the data segment, and realize adaptive mapping operations.
Improve storage efficiency, reduce storage costs, optimize data access speed, avoid waste of storage space, and ensure flexibility and efficiency of data stream processing.
Smart Images

Figure CN120295582A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage, and more specifically, to a streaming data storage method combined with adaptive compression. Background Art
[0002] Adaptive compression is a method of dynamically adjusting compression algorithms and strategies. It can select the most suitable compression method in real time according to the characteristics of data and storage requirements. By analyzing features such as data redundancy, entropy value, and access frequency, it can automatically adjust the compression ratio and algorithm to optimize data storage efficiency and access speed. Adaptive compression can automatically adapt to data changes in different operating environments, improve storage efficiency, and reduce storage costs.
[0003] During the adaptive compression process, segmented storage is usually performed, which relies on local decompression to improve access efficiency. However, local decompression will cause the problem of storage space fragmentation. When different data segments use different compression algorithms, or some parts of the data stream are not effectively allocated for compression, it will lead to uneven distribution of storage space, resulting in waste of storage space. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present invention provide a streaming data storage method combined with adaptive compression, which can solve the problems mentioned in the above background art through fragmentation evaluation and management, and processing of eigenvectors of data.
[0005] To achieve the above object, the present invention provides the following technical solution: A streaming data storage method combined with adaptive compression, including:
[0006] Receiving real-time data and performing a first-order analysis task, where the first-order analysis task determines the segmentation method and compression algorithm for each data segment based on the eigenvector of the data;
[0007] Continuously inputting the data stream and performing a second-order analysis task, analyzing the data stream in a real-time processing pipeline, and evaluating the distribution characteristics of the currently stored data segments through a B+ tree index; if the evaluation of the second-order analysis task meets the expectation, the reconstruction of the data segment is completed, otherwise, return to the first-order analysis task for re-evaluation;
[0008] Evaluating the delayed access degree of the reconstructed data segment. If the delayed access degree exceeds the expectation, adjust the storage priority within a preset sliding time window based on the access frequency and timeliness of the data segment through hash indexing and queue management to obtain time-sensitive data segments;
[0009] Performing a fragmentation evaluation on the time-sensitive data segment. If the fragmentation evaluation shows the existence of storage fragments, perform a fragmentation management task, otherwise, store the time-sensitive data segment according to the original allocation and continue the processing;
[0010] According to the results of fragmented evaluation, use a hash table and a B+ tree index to divide it into an optimized data segment and a merged data segment, and the data segments that are not divided are still retained as time-sensitive data segments;
[0011] Assign different weights to the optimized data segment, the merged data segment, and the time-sensitive data segment in the streaming data priority calculation model, and perform the adaptive mapping operation of the streaming data segment.
[0012] In a preferred embodiment, the compression potential of the data segment is defined by the feature vector of the data. The feature vector of the data includes redundancy, entropy value, and access frequency in the first-order analysis task; based on the compression potential, the data segment is defined as the first data segment, the second data segment, and the third data segment;
[0013] The selection of the first data segment includes: if the redundancy of the data segment is greater than the preset redundancy threshold, then use a type of compression ratio algorithm to perform compression. The type of compression ratio algorithm includes LZ77 or Huffman coding;
[0014] The selection of the second data segment includes: if the entropy value of the data segment is greater than the preset entropy threshold and the storage requirement of the data segment exceeds the preset target storage requirement, then use a type of compression ratio algorithm to perform compression; the type of compression ratio algorithm includes Zlib; the storage requirement of the data segment includes the number of bytes of the data and the complexity of the data;
[0015] The selection of the third data segment includes: if the data segment is not assigned to the first data segment or the second data segment within the preset time window, or the access frequency of the data segment is greater than the preset access frequency, then select not to compress or use a type of compression ratio algorithm to perform compression; the type of compression ratio algorithm includes Run-Length Encoding.
[0016] In a preferred embodiment, establish a distribution feature model based on the data segment distribution characteristics in the second-order analysis task. The distribution feature model is constructed based on data redundancy, data access frequency, data segment storage density, and data segment location information as input variables; calculate the distribution feature value based on the distribution feature model. If the distribution feature value is greater than the preset distribution feature threshold, it is determined that the evaluation of the second-order analysis task meets the expectations;
[0017] The goal of data redundancy in the distribution feature model is to quantify the repeatability of information in the data segment. By calculating the frequencies of each field in the data segment and comparing them with the frequency peak and average frequency, the redundancy degree of the data is measured;
[0018] The goal of data access frequency in the distribution feature model is to measure the index of the access activity of the data during the storage process. Based on reflecting the hot spot attribute of the data segment, evaluate it through time series data pairs to predict whether the data segment is a frequently accessed hot data;
[0019] The goal of the data segment storage density in the distribution feature model is to measure the space utilization of data in the storage area; by comparing the storage occupancy of each field within the data segment with the storage space of the entire data segment, the utilization efficiency of the storage space is measured to determine whether there is waste of storage resources in the data segment.
[0020] The goal of the data segment location information in the distribution feature model is to evaluate the location of the data segment in the storage system and optimize the storage layout; it evaluates the efficiency of the storage location by calculating the distance between each data segment and other data segments in the system.
[0021] In a preferred embodiment, when obtaining the time-sensitive data segment, each data segment is first evaluated to calculate its access frequency and timeliness, so that the required data segments enter the cache preferentially.
[0022] In evaluating the delayed access degree, each data segment is evaluated to calculate its access frequency and timeliness, so that the required data segments enter the cache preferentially, and then the priority of each data segment is calculated according to the product of [].
[0023] A unique storage location identifier is assigned to each data segment through a hash index, and the storage location of the data segment is established based on the storage location identifier.
[0024] According to the priority obtained from the access frequency and timeliness, the storage priority of the data segment is dynamically adjusted, and a queue is constructed according to the change of the real-time data stream, and the position of the data segment in the queue is adjusted in real time.
[0025] Within the sliding time window, the access to each data segment is traced, and a new priority of the data segment is calculated and generated based on the initial priority and the timeliness decay factor, and the timeliness decay factor weakens the priority of the data segment over time.
[0026] Finally, the time-sensitive data segment is obtained.
[0027] In a preferred embodiment, fragmentation evaluation is performed based on the fragmentation degree and storage efficiency of the time-sensitive data segment; a fragmentation evaluation model is constructed through fragmentation evaluation; the time-sensitive data segment is split into multiple storage blocks, and each storage block is assigned a storage state, and the storage state includes free, occupied or fragmented.
[0028] The fragmentation degree is used to measure the effective utilization of the storage space, and it represents the ratio of the unused space to the used space in the storage block based on the ratio between the free space and the actual storage space of the data segment.
[0029] The storage efficiency is an index used to measure the occupancy rate of the storage space by the storage system.
[0030] Taking the fragmentation degree and storage efficiency of the timeliness data segment as input variables, each storage block is calculated. Based on the fragmentation evaluation model, the fragmentation score of the entire timeliness data segment is calculated. The fragmentation score is used to reflect the fragmentation degree generated during the storage process of the timeliness data segment. Based on the fragmentation score, the timeliness data segment is divided into an optimized data segment and a merged data segment. If the fragmentation score is less than the lower limit of the preset fragmentation score threshold, it is divided into an optimized data segment. If the fragmentation score is greater than the upper limit of the preset fragmentation score threshold, it is divided into a merged data segment. For the data segments that are not divided out, they remain as timeliness data segments.
[0031] In a preferred embodiment, the goal of the fragmentation management task is to reduce fragmentation and improve the utilization efficiency of the storage space by merging and reorganizing storage blocks. Based on this, a reorganization cost function is constructed to quantify the resource consumption and effect of the fragmentation management task through the reorganization cost function.
[0032] In the execution of the fragmentation management task, by calculating the size and weight factor of each storage block, and combining the fragmentation score and storage compactness for calculation, the reorganization cost of each storage block is obtained. Then, based on the density and fragmentation degree of the storage block, a merging strategy is selected. According to the calculated reorganization cost, it is judged whether the storage block needs to be reconstructed. If the reorganization cost of reconstructing the storage block is lower than the preset reorganization cost threshold, the reconstruction is executed; otherwise, the storage is allocated as original and the process continues.
[0033] In a preferred embodiment, it further includes a streaming data segment adaptive mapping operation:
[0034] Calculate the input variables of each data segment: Perform a first-order analysis task according to the characteristics of the streaming data segment. In the first-order analysis task, the validity characteristics of each data segment are supplemented and collected. The validity characteristics include access frequency, compression rate, and storage space. Taking the validity characteristics as input variables, a weighted calculation is performed to obtain the final priority of each data segment.
[0035] Calculate the priority of the data segment: Calculate the priorities of the optimized data segment, the merged data segment, and the timeliness data segment. Assign an independent priority value to each data segment, and based on the independent priority value, influence the subsequent storage and access scheduling decisions.
[0036] Priority scheduling: Comprehensively evaluate the priorities of the optimized data segment, the merged data segment, and the timeliness data segment according to the global weight coefficient to obtain the comprehensive priority of each data segment. The comprehensive priority determines the mapping order of the data segment in the streaming data storage, that is, the priority of the data segment being stored and accessed.
[0037] Perform adaptive mapping operation on data segments: According to the comprehensive priority of each data segment, perform an adaptive mapping operation to rearrange the storage location of the data segment in the streaming storage system or schedule the access priority.
[0038] Technical effects and advantages of the present invention:
[0039] 1. By dynamically selecting an appropriate compression algorithm based on the feature vector of the data, an optimal ratio between the compression ratio of the data and the storage space is formed, improving the storage efficiency and reducing the storage cost;
[0040] 2. Through real-time data analysis and reconstruction, based on the B+ tree index to evaluate the distribution characteristics of data segments, dynamically adjust the storage structure of data segments to ensure flexibility in data stream processing;
[0041] 3. By combining the access frequency, timeliness, and fragmentation of the data, and adjusting the priority of data segments based on hash index and queue management, the access speed and storage performance of the data are optimized to ensure that time-sensitive data segments are processed first;
[0042] 4. Through fragmentation evaluation and management, it is conducive to identifying storage fragments and optimizing time-sensitive data segments, avoiding waste of storage space, and thus improving the utilization rate of storage resources;
[0043] 5. Based on the comprehensive priority of each data segment, perform an adaptive mapping operation to dynamically adjust the storage location and access priority of the data segment, so as to ensure relatively efficient access operations in streaming data storage. Description of the Drawings
[0044] Figure 1 is the flow of the present invention Figure 1 .
[0045] Figure 2 is the flow of the present invention Figure 2 . Detailed Embodiments
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0047] Referring to the attached specification Figure 1-2 , a method for streaming data storage combined with adaptive compression according to an embodiment of the present invention includes:
[0048] Receive real-time data and perform first-order analysis tasks. The first-order analysis tasks determine the segmentation method and compression algorithm for each data segment based on the eigenvectors of the data;
[0049] Continuously input the data stream and perform second-order analysis tasks. Analyze the data stream in the real-time processing pipeline, and evaluate the distribution characteristics of the currently stored data segments through a B+ tree index; if the evaluation of the second-order analysis tasks meets the expectations, complete the reconstruction of the data segments, otherwise return to the first-order analysis tasks for re-evaluation;
[0050] Evaluate the latency access degree of the reconstructed data segments. If the latency access degree exceeds the expectations, based on the access frequency and timeliness of the data segments, adjust the storage priority within a preset sliding time window through a hash index and queue management to obtain time-sensitive data segments;
[0051] Perform fragmentation evaluation on the time-sensitive data segments. If the fragmentation evaluation shows the existence of storage fragmentation, perform fragmentation management tasks, otherwise the time-sensitive data segments are stored according to the original allocation and continue to be processed;
[0052] According to the results of the fragmentation evaluation, use a hash table and a B+ tree index to divide them into optimized data segments and merged data segments, and the data segments that are not divided remain as time-sensitive data segments;
[0053] Assign different weights to the optimized data segments, merged data segments, and time-sensitive data segments in the streaming data priority calculation model, and perform the adaptive mapping operation of the streaming data segments.
[0054] Define the compression potential of the data segments by the eigenvectors of the data. The eigenvectors of the data include redundancy, entropy value, and access frequency in the first-order analysis tasks; based on the compression potential, the data segments are defined as the first data segments, the second data segments, and the third data segments;
[0055] The selection of the first data segments includes: if the redundancy of the data segment is greater than the preset redundancy threshold, use a type of compression ratio algorithm to perform compression. The type of compression ratio algorithm includes LZ77 or Huffman coding; the type of compression ratio algorithm refers to a high compression ratio algorithm;
[0056] The selection of the second data segments includes: if the entropy value of the data segment is greater than the preset entropy threshold and the storage requirement of the data segment exceeds the preset target storage requirement, use a type of compression ratio algorithm to perform compression; the type of compression ratio algorithm includes Zlib, and the type of compression ratio algorithm is also a medium compression ratio algorithm; the storage requirement of the data segment includes the number of bytes of the data and the complexity of the data. The complexity of the data includes but is not limited to the number of data fields and the attribute types;
[0057] The selection of the third data segment includes: if the data segment is not allocated to the first data segment or the second data segment within the preset time window, or the access frequency of the data segment is greater than the preset access frequency, then select not to compress or execute compression using three types of compression ratio algorithms; the three types of compression ratio algorithms include Run-Length Encoding, and the three types of compression ratio algorithms are also low compression ratio algorithms.
[0058] Establish a distribution feature model based on the data segment distribution characteristics in the second-order analysis task. The distribution feature model is based on the data redundancy R d , the data access frequency A f , the data segment storage density D s , the data segment position information P l and is constructed with these as input variables; calculate the distribution feature value based on the distribution feature model. If the distribution feature value is greater than the preset distribution feature threshold, it is determined that the evaluation of the second-order analysis task meets the expectations.
[0059] The data redundancy R d The goal of the distribution feature model is to quantify the repeatability of information in the data segment. By calculating the frequencies of each field in the data segment and comparing them with the frequency peak and average frequency, the redundancy degree of the data is measured. Among them, the data segment with high redundancy represents a high compression potential because it contains a large amount of repetitive information data.
[0060] The data access frequency A f The goal of the distribution feature model is an indicator to measure the access activity of data during storage. The data segment with a higher access frequency should be optimized and stored efficiently first. Based on A f it reflects the hot spot attribute of this data segment. Through time series data for A f to evaluate and predict whether the data segment is hot data that is frequently accessed.
[0061] The data segment storage density D s The goal of the distribution feature model is to measure the space utilization of data in the storage area. Among them, the data segment with a higher storage density means that this data segment occupies less space when stored, so it is more efficient when stored; by comparing the storage occupancy of each field in the data segment with the storage space of the entire data segment, the utilization efficiency of the storage space is measured, and it is judged whether there is waste of storage resources in the data segment.
[0062] The data segment position information P l The goal in the distribution feature model is to evaluate the position of the data segment in the storage system and optimize the storage layout; it evaluates the efficiency of the storage position by calculating the distance between each data segment and other data segments in the system; among them, the closer the storage position is to the data segment with high access frequency, the higher the data access efficiency. Based on this, optimize the storage strategy of the data and reduce unnecessary access latency.
[0063] Based on the data redundancy R d 、the data access frequency A f 、the storage density D of the data segment s 、the position information P of the data segment l When constructing the distribution feature model, it can be expressed as:
[0064]
[0065] where f i is the occurrence frequency of the i-th field in the data segment, and f i is used to reflect the degree of repetition of each field in the data segment. The higher f i is, the stronger the redundancy of the field; f max is the maximum occurrence frequency of the fields in the data segment, which is used to standardize the occurrence frequency of the fields. The purpose of doing this is to make the redundancy independent of the size and content of the data segment; f avg is the average frequency of the fields in the data segment, representing the medium level of the frequency distribution of different fields in the data segment; n is the number of fields in the data segment; where By calculating the logarithmic difference between the field frequency and the average frequency, the redundancy of high-frequency fields is further emphasized; a large deviation indicates that the field frequency is much higher than the average level, showing that the field has strong redundancy and is suitable for compression;
[0066] where a t is the number of accesses to the data segment at time point t, representing the access frequency of the data segment within a specific time window. The more accesses there are, the more frequently the data segment is accessed; A max is the maximum value of the number of accesses to the data segment within the time window; A avg is the average value of the number of accesses to the data segment within the time window, reflecting the overall level of data access and used to calculate the average value of the number of accesses at all time points; T represents the length of the evaluation time period; where By calculating the logarithmic difference between the access frequency and the average frequency, the fluctuation degree of the access frequency is measured. A large deviation indicates that the access frequency is much higher than the average level, suggesting that the data segment is hot data;
[0067] where s i is the storage space occupied by the i-th field in the data segment, which is used to measure the size of each field in storage. The larger the storage space occupied by the field, the lower the storage density of the field; S total is the total storage space of the data segment; a higher storage density D of the data segment s means that the storage space is utilized more fully and the data storage efficiency is higher;
[0068] where di The storage distance at the i-th position of the data segment in the storage area, representing the physical position of the data segment in the storage area. A smaller d i value indicates that the data segment is located at a closer storage position and has a faster access speed; D total is the total storage distance of all data segment positions in the storage system, representing the comprehensive layout of all data segment storage positions; D avg is the average storage distance of the data segment positions in the storage system, D avg which is used to reflect the average layout level of the storage area; a smaller d i means that the data segment is stored at a position with a higher priority, which also means it can respond to access requests more quickly;
[0069] Based on the above formula, the calculated data redundancy R d , data access frequency A f , data segment storage density D s , data segment position information P l values are obtained and weighted and combined to finally obtain the distribution feature model M d ;
[0070] M d = w1R d + w2A f + w3D s + w4P l
[0071] where w1, w2, w3, and w4 respectively correspond to the weights of the data redundancy R d , data access frequency A f , data segment storage density D s , data segment position information P l , and the weights can be adjusted according to the actual application scenario.
[0072] When obtaining time-sensitive data segments, first evaluate each data segment D i , calculate its access frequency Freq(D i ) and timeliness St(D i ), and make the required data segments enter the cache first;
[0073] In evaluating the degree of delayed access, evaluate each data segment D i , calculate its access frequency Freq(D i ) and timeliness St(D i ), make the required data segments enter the cache first, and then calculate the priority P(D i ) of each data segment according to the product of Freq(D i ) and St(D i ); P(Di ) = Freq(D i ) × St(D i );
[0074] Assign a unique storage location identifier to each data segment D through the hash index H(D i ) and establish the storage location of the data segment based on the storage location identifier; H(D i ) is expressed as: H(D i ) = hash(ID(D i )); i
[0075] Dynamically adjust the storage priority of the data segment D according to P(D i ) obtained from the access frequency Freq(D i ) and the timeliness St(D i ), and construct a queue according to the change of the real-time data stream, and adjust the position of the data segment D i in the queue in real time; i
[0076] Within the sliding time window, track the access to each data segment D i , calculate and generate the new priority New Priority(D i ) of the data segment D based on the initial priority Queue(D i ) and the timeliness decay factor α. The timeliness decay factor α weakens the priority of the data segment D i over time; assume T i is the current time, T now (D last ) is the time of the last access to the data segment D i . Based on T i - T now (D last ) represents the time difference from the last access to the current time. The expression of New Priority(D i ) is: i
[0077]
[0078] Finally, obtain the time-sensitive data segment;
[0079] The access frequency Freq(D i ) refers to the number of times a certain data segment D i is accessed within a given time window, and is used to evaluate the activity of the data segment D i in the streaming data. Record the number of accesses to each data segment D i within the sliding time window, and set the time window as Tw , then the access frequency Freq(D i ) is expressed as:
[0080]
[0081] where access(D i , t) represents the number of times the data segment D i is accessed at time t;
[0082] The timeliness St(D i ) is used to measure the degree to which the data segment D i remains valid or useful in the current streaming data system; the timeliness is directly related to the expiration degree of the data, the access pattern, and the storage requirements; in the function expression of St(D i ), by considering the T i of the data segment D last (D i ) and when T now , a threshold θ is set to define its validity period; the timeliness St(D i ) is expressed as:
[0083]
[0084] where α is the timeliness decay factor, T last (D i ) is the last access time of the data segment D i , T now is the current time, indicating the expiration degree of the data segment, and a smaller α will decay more slowly.
[0085] Perform fragmentation evaluation based on the fragmentation degree and storage efficiency of the timeliness data segment; construct a fragmentation evaluation model through fragmentation evaluation; split the timeliness data segment into multiple storage blocks B1, B2, …, B n , and each storage block B i is given a storage state, and the storage state includes free, occupied, or fragmented; the size of the storage block |B i | is determined by the storage system according to the capacity of the storage unit and the data compression algorithm;
[0086] Based on the fragmentation degree F, it is used to measure the effective utilization of the storage space. It is based on the ratio between the free space and the actual storage space of the data segment, and represents the ratio of the unused space to the used space in the storage block B i ; among them, the larger the value of the fragmentation degree F, the more serious the space waste of data storage, and the lower the storage efficiency of the system. The formula for the fragmentation degree F is expressed as:
[0087]
[0088] Among them, S unused,i represents the unused storage space in the i-th storage block B i ; S total,i represents the total storage space of the i-th storage block B i ; R i represents the density factor of the storage block B i , and the density factor is used to measure the tightness of the storage block usage;
[0089] The storage efficiency is an index used to measure the occupancy rate of the storage space by the storage system. The storage efficiency is represented by the storage compactness ρ. The higher the storage compactness ρ, the higher the utilization rate of the storage space. The storage efficiency is expressed as:
[0090]
[0091] Among them, S used,i represents the used storage space in the storage block B i ; E i represents the weighting factor used to consider the storage timeliness and access frequency of the data segment; the larger the value of the storage compactness ρ, the more efficient the storage;
[0092] Taking the fragmentation degree and storage efficiency of the time-sensitive data segment as input variables, calculate each storage block B i , and calculate the fragmentation score F score of the entire time-sensitive data segment based on the fragmentation evaluation model. The fragmentation score F score is used to reflect the fragmentation degree generated during the storage process of the time-sensitive data segment. Based on the fragmentation score F score , divide the time-sensitive data segment into an optimized data segment and a merged data segment; if the fragmentation score is less than the lower limit of the preset fragmentation score threshold, divide it into an optimized data segment, because the optimized data segment has already efficiently utilized the storage space and is suitable for further optimizing the storage layout; if the fragmentation score is greater than the upper limit of the preset fragmentation score threshold, divide it into a merged data segment, and the merged data segment needs to reduce fragmentation and improve storage efficiency through a merge operation; for the data segment that has not been divided, it remains a time-sensitive data segment and continues to maintain high-priority storage and access; the fragmentation score F score is expressed as:
[0093]
[0094] Among them, F i represents the fragmentation degree of the i-th storage block; α i is the weight factor of the fragmented block; ρ j is the storage compactness of the j-th data segment; β j is the weight factor of the storage compactness; c is the number of storage blocks; v is the number of data segments.
[0095] The goal of the fragmented management task is to reduce fragmentation and improve the utilization efficiency of storage space by merging and reorganizing storage blocks. Based on this, a reorganization cost function is constructed to quantify the resource consumption and effect of the fragmented management task through the reorganization cost function.
[0096] In the execution of the fragmented management task, by calculating the size i of each storage block B and the weight factor W i , and combining the fragmentation score F score and the storage compactness ρ j for calculation, the reorganization cost C reorg of each storage block is obtained; then, based on the density and fragmentation degree of the storage blocks, a merging strategy is selected, and the storage blocks with lower density and higher fragmentation degree are preferentially merged to reduce fragmentation and increase storage density. According to the calculated reorganization cost C reorg , it is judged whether the storage blocks need to be reconstructed. If the reorganization cost C reorg of reconstructing the storage blocks is lower than the preset reorganization cost threshold, then the reconstruction is executed; otherwise, the storage is allocated as originally and the processing continues.
[0097]
[0098] Among them, the reorganization cost C reorg is used to represent the computing resources required for executing the fragmented management task. is the size of the storage block B i , representing the physical space of the storage block; W i is the weight factor of each storage block B i ; F score is the fragmentation score, indicating the severity of fragmentation. If the fragmentation score increases, the need for merging and reconstruction also increases, so its influence will be increased in the reorganization cost; λ is the weight coefficient, and λ is used to adjust the influence of the fragmentation score on the reorganization cost. The larger λ is, the more serious the fragmentation problem is, and the higher the reorganization cost of the system; ζ is the influence coefficient of the storage compactness on the reorganization cost; the storage compactness ρ j is used to measure the occupancy efficiency of data in storage. When ρ j is relatively high, it means that the storage space is better utilized and the reorganization cost is relatively low.
[0099] It also includes the adaptive mapping operation of the streaming data segment:
[0100] Calculate the input variables for each data segment: Perform a first-order analysis task based on the characteristics of the streaming data segment. Supplement and collect the validity characteristics of each data segment in the first-order analysis task. The validity characteristics include access frequency, compression ratio, and storage space. Use the validity characteristics as input variables and perform weighted calculations to obtain the final priority of each data segment;
[0101] Calculate the priority of the data segment: Calculate the priorities of the optimized data segment, merged data segment, and timeliness data segment. Assign an independent priority value to each data segment, and based on the independent priority value, influence subsequent storage and access scheduling decisions;
[0102] Priority scheduling: Comprehensively evaluate the priorities of the optimized data segment, merged data segment, and timeliness data segment according to the global weight coefficient to obtain the comprehensive priority of each data segment. The comprehensive priority determines the mapping order of the data segment in the streaming data storage, that is, the priority of the data segment to be stored and accessed;
[0103] Perform the adaptive mapping operation of the data segment: According to the comprehensive priority of each data segment, perform the adaptive mapping operation to rearrange the storage location or schedule the access priority of the data segment in the streaming storage system, so as to ensure the efficiency and real-time performance of data access;
[0104] It should be further explained for the above embodiments:
[0105] The priority P of the optimized data segment opt The optimized data segment in is the data segment after fragmentation management and reconstruction, which has low latency access and high access frequency; Calculate its priority according to the validity of access frequency, compression ratio, and storage space;
[0106] P opt = α·F opt + β·C opt - γ·S opt
[0107] Where F opt is the access frequency of the optimized data segment; C opt is the compression ratio of the optimized data segment. The higher the compression ratio, the larger the value, which means more advantages in the embodiment; S opt is the storage space usage of the optimized data segment; α, β, γ are weight coefficients, which respectively represent the influence of access frequency, compression ratio, and storage space on the priority;
[0108] The priority P of the merged data segment merge The merged data segment in is the data segment formed by merging multiple smaller data segments. Its priority is jointly determined by multiple factors. Calculate its priority through the historical access frequency, latency access degree, and time span of the merged data segment;
[0109] P merge = δ·H merge + λ·D merge - θ·T merge
[0110] Where H merge is the historical access frequency of the merged data segment; D merge is the latency access degree of the merged data segment; T merge is the time span of the merged data segment; δ, λ, θ are weight coefficients representing the impacts of historical access frequency, latency access degree, and time span on the priority of the merged data segment respectively;
[0111] The priority P timeliness of the time-sensitive data segment in the time-sensitive data segment refers to a data segment with time-sensitive requirements, usually requiring a high access priority during storage. The priority P timeliness of the time-sensitive data segment combines the time-sensitive requirements of storage and the access frequency of the real-time data stream during calculation;
[0112] P timeliness = μ·F timeliness + ν·A timeliness - ξ·S timeliness
[0113] Where F timeliness is the real-time access frequency of the time-sensitive data segment; A timeliness is the time-sensitive requirement of the time-sensitive data segment; S timeliness is the storage resource occupancy of the time-sensitive data segment; μ, ν, ξ are weight coefficients, which represent the impacts of access frequency, time-sensitive requirement, and storage resource occupancy on the priority of the time-sensitive data segment;
[0114] The generation goal of the priority calculation model for the streaming data segment is to calculate the priority of each data segment according to its characteristics, so as to determine how to adaptively map the storage and access order of the data segments;
[0115] P total = α1·P opt + α2·P merge + α3·P timeliness
[0116] Where P total is the comprehensive priority of each data segment; P opt is the priority of the optimized data segment; P merge is the priority of the merged data segment; P timeliness is the priority of the time-sensitive data segment; α1, α2, α3 are global weight coefficients, which are used to adjust the impacts of the optimized data segment, the merged data segment, and the time-sensitive data segment in the final priority calculation.
[0117] The foregoing are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A streaming data storage method combined with adaptive compression, comprising: Receiving real-time data and performing a first-order analysis task, where the first-order analysis task determines the segmentation method and compression algorithm for each data segment based on the feature vector of the data; It is characterized in that: Continuously inputting the data stream and performing a second-order analysis task, analyzing the data stream in a real-time processing pipeline, and evaluating the distribution characteristics of the currently stored data segments through a B+ tree index; if the evaluation of the second-order analysis task meets the expectation, the reconstruction of the data segments is completed, otherwise, return to the first-order analysis task for re-evaluation; Evaluating the delayed access degree of the reconstructed data segments. If the delayed access degree exceeds the expectation, based on the access frequency and timeliness of the data segments, adjust the storage priority within a preset sliding time window through hash indexing and queue management to obtain time-sensitive data segments; Performing fragmentation evaluation on the time-sensitive data segments. If the fragmentation evaluation shows the existence of storage fragmentation, perform a fragmentation management task, otherwise, the time-sensitive data segments are stored according to the original allocation and continue to be processed; According to the results of the fragmentation evaluation, use a hash table and a B+ tree index to divide them into optimized data segments and merged data segments, and the data segments that are not divided remain as time-sensitive data segments; Assign different weights to the optimized data segments, merged data segments, and time-sensitive data segments in the streaming data priority calculation model, and perform an adaptive mapping operation on the streaming data segments.
2. The streaming data storage method combined with adaptive compression according to claim 1, characterized in that: Define the compression potential of the data segments by the feature vector of the data. The feature vector of the data includes redundancy, entropy value, and access frequency in the first-order analysis task; define the data segments as the first data segments, second data segments, and third data segments based on the compression potential; The selection of the first data segments includes: if the redundancy of the data segment is greater than the preset redundancy threshold, use a type of compression ratio algorithm to perform compression, and the type of compression ratio algorithm includes LZ77 or Huffman coding; The selection of the second data segments includes: if the entropy value of the data segment is greater than the preset entropy threshold and the storage requirement of the data segment exceeds the preset target storage requirement, use a type of compression ratio algorithm to perform compression; the type of compression ratio algorithm includes Zlib; the storage requirement of the data segment includes the number of bytes of the data and the complexity of the data; The selection of the third data segments includes: if the data segment is not assigned to the first data segment or the second data segment within the preset time window, or the access frequency of the data segment is greater than the preset access frequency, select not to compress or use a type of compression ratio algorithm to perform compression; the type of compression ratio algorithm includes Run-Length Encoding.
3. The streaming data storage method combined with adaptive compression according to claim 2, characterized in that: Establish a distribution feature model based on the data segment distribution characteristics in the second-order analysis task. The distribution feature model is constructed based on data redundancy, data access frequency, data segment storage density, and data segment location information as input variables; calculate the distribution feature value based on the distribution feature model. If the distribution feature value is greater than the preset distribution feature threshold, it is determined that the evaluation of the second-order analysis task meets the expectation; The goal of data redundancy in the distribution feature model is to quantify the repetition of information in a data segment. By calculating the frequencies of each field in the data segment and comparing them with the frequency peak and average frequency, the redundancy degree of the data is measured. The goal of data access frequency in the distribution feature model is an indicator to measure the access activity of data during storage. Based on reflecting the hot attributes of the data segment, it is evaluated through time series data pairs to predict whether the data segment is hot data with frequent access. The goal of data segment storage density in the distribution feature model is to measure the space utilization of data in the storage area. By comparing the storage occupancy of each field within the data segment with the storage space of the entire data segment, the utilization efficiency of the storage space is measured to determine whether there is waste of storage resources in the data segment. The goal of data segment location information in the distribution feature model is to evaluate the location of the data segment in the storage system and optimize the storage layout. It evaluates the efficiency of the storage location by calculating the distance between each data segment and other data segments in the system.
4. A streaming data storage method combined with adaptive compression according to claim 3, characterized in that: When obtaining time-sensitive data segments, first evaluate each data segment, calculate its access frequency and timeliness, and make the required data segments enter the cache first. In evaluating the delayed access degree, evaluate each data segment, calculate its access frequency and timeliness, make the required data segments enter the cache first, and then calculate the priority of each data segment according to the product of [parameters not provided in the original]. Assign a unique storage location identifier to each data segment through a hash index, and establish the storage location of the data segment based on the storage location identifier. Dynamically adjust the storage priority of the data segment according to the priority obtained from the access frequency and timeliness, and construct a queue according to the changes in the real-time data stream, and adjust the position of the data segment in the queue in real time. Within the sliding time window, track the access of each data segment, calculate and generate the new priority of the data segment based on the initial priority and the timeliness decay factor, and the timeliness decay factor weakens the priority of the data segment over time. Finally, obtain the time-sensitive data segments.
5. A streaming data storage method combined with adaptive compression according to claim 4, characterized in that: Perform fragmentation evaluation based on the fragmentation degree and storage efficiency of the time-sensitive data segment; construct a fragmentation evaluation model through fragmentation evaluation; split the time-sensitive data segment into multiple storage blocks, and assign a storage state to each storage block, and the storage state includes free, occupied or fragmented. The fragmentation degree is used to measure the effective utilization of the storage space, and it is based on the ratio between the free space and the actual storage space of the data segment to represent the ratio of the unused space to the used space in the storage block. The storage efficiency is an indicator used to measure the occupancy rate of the storage space by the storage system. Each storage block is calculated with the fragmentation degree and storage efficiency of the timeliness data segment as input variables. Based on the fragmentation evaluation model, the fragmentation score of the entire timeliness data segment is calculated. The fragmentation score is used to reflect the fragmentation degree generated during the storage process of the timeliness data segment. Based on the fragmentation score, the timeliness data segment is divided into an optimized data segment and a merged data segment. If the fragmentation score is less than the lower limit of the preset fragmentation score threshold, it is divided into the optimized data segment. If the fragmentation score is greater than the upper limit of the preset fragmentation score threshold, it is divided into the merged data segment. For the data segments that have not been divided, they remain as timeliness data segments.
6. A streaming data storage method combined with adaptive compression according to claim 5, characterized in that: The goal of the fragmentation management task is to reduce fragmentation and improve the utilization efficiency of the storage space by merging and reorganizing storage blocks. Based on this, a reorganization cost function is constructed, and the resource consumption and effect of the fragmentation management task are quantified through the reorganization cost function. During the execution of the fragmentation management task, by calculating the size and weight factor of each storage block, and combining the fragmentation score and storage compactness for calculation, the reorganization cost of each storage block is obtained. Then, based on the density and fragmentation degree of the storage block, a merging strategy is selected. According to the calculated reorganization cost, it is judged whether the storage block needs to be reconstructed. If the reorganization cost of reconstructing the storage block is lower than the preset reorganization cost threshold, the reconstruction is executed; otherwise, it is stored according to the original allocation and the processing continues.
7. A streaming data storage method combined with adaptive compression according to claim 6, characterized in that: It further includes an adaptive mapping operation for the streaming data segment: Calculating the input variables for each data segment: performing a first-order analysis task according to the characteristics of the streaming data segment. In the first-order analysis task, the validity characteristics of each data segment are supplemented and collected. The validity characteristics include access frequency, compression ratio, and storage space. Taking the validity characteristics as input variables, a weighted calculation is performed to obtain the final priority of each data segment. Calculating the priority of the data segment: calculating the priorities of the optimized data segment, the merged data segment, and the timeliness data segment, and assigning an independent priority value to each data segment. Based on the independent priority value, the subsequent storage and access scheduling decisions are affected. Priority scheduling: comprehensively evaluating the priorities of the optimized data segment, the merged data segment, and the timeliness data segment according to the global weight coefficient to obtain the comprehensive priority of each data segment. The comprehensive priority determines the mapping order of the data segment in the streaming data storage, that is, the priority of the data segment to be stored and accessed. Performing the adaptive mapping operation for the data segment: according to the comprehensive priority of each data segment, performing the adaptive mapping operation to rearrange the storage location of the data segment in the streaming storage system or schedule the access priority.
Citation Information
Patent Citations
Method and system for self-adaptation data compression and decompression and storage device
CN103516369A
Joining tables in a mapreduce procedure
CN103620601A
Processing method and device for storing fragmented data, storage medium and chip
CN115729856A
Big data analysis method based on adaptive data management and privacy protection
CN119088832A
Efficient memory allocation and recovery method for Android device
CN119645634A