Intelligent hierarchical storage method for unstructured data

By building an intelligent hierarchical storage method for unstructured data, combining data attenuation function and machine learning model, dynamically optimizing the storage strategy of unstructured data, the problem of failure to consider multi-dimensional features in traditional storage methods is solved, and efficient and intelligent data management is achieved.

CN120277171AActive Publication Date: 2025-07-08DETSERWEI TECH CO LTD

Patent Information

Application Number
CN202510363689.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-08
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

In the unstructured data storage management, the prior art fails to effectively consider multi-dimensional characteristics such as data timeliness attenuation, content sensitivity and business value, resulting in disconnection between storage decisions and actual data value, affecting data compliance and access efficiency.

Method used

Using an intelligent hierarchical storage method, we use the initial value decay function of unstructured data, combine historical access records and data sensitivity, dynamically adjust storage hierarchical decisions, and use machine learning models to optimize access probability prediction and value correction, and dynamically optimize data storage strategies.

Benefits of technology

It realizes multi-dimensional data value evaluation, reduces error rate, reduces manual labeling costs, supports policy redundancy and cross-verification, and improves the intelligence and efficiency of data management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277171A_ABST
    Figure CN120277171A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent hierarchical storage method for unstructured data, and relates to the technical field of unstructured data processing, and the method specifically comprises the steps: determining an initial value through data feature analysis; establishing an attenuation model to output a dynamic value; generating a preliminary storage hierarchy according to the value attenuation degree; using the access record to train a prediction model, and evaluating a future access probability; adjusting a data value weight according to the access probability; and determining an optimal storage hierarchy based on the adjusted comprehensive value index. According to the method, dynamic matching of storage resources and access popularity is realized through a dual value evaluation mechanism; various reference factors are fused, real-time dynamic quantification of data values is realized, and limitation of static evaluation is broken through; through collaborative optimization of a value attenuation model and machine learning prediction, the core pain points of storage strategy lag and low resource utilization rate in a traditional method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unstructured data processing, and specifically to an intelligent hierarchical storage method for unstructured data. Background Art

[0002] With the rapid development of big data and artificial intelligence technologies, the scale of unstructured data (such as text, images, audio, video, etc.) has shown an explosive growth. According to IDC statistics, unstructured data already accounts for more than 80% of the total global data volume, and its storage management has become a core challenge for enterprise IT infrastructure. Traditional hierarchical storage technologies mainly divide data into levels such as hot, warm, and cold based on static rules such as data access frequency and storage cost, but there are the following limitations: Single dimension of value evaluation: Existing methods mostly rely on manually preset strategies (such as access frequency statistics based on LRU), without considering multi-dimensional features such as data timeliness decay, content sensitivity, and business value, resulting in a disconnection between storage decisions and the actual value of data. For example, medical imaging data needs to be retained for a long time due to compliance requirements, but its access frequency may be extremely low, and traditional cold storage strategies cannot meet compliance and rapid retrieval requirements. Summary of the Invention

[0003] Aiming at the deficiencies of the existing technology, the present invention provides an intelligent hierarchical storage method for unstructured data.

[0004] In order to achieve the above object, the technical solution of the present invention is as follows:

[0005] An intelligent hierarchical storage method for unstructured data, comprising the following steps:

[0006] Obtain unstructured data, and determine the initial value of the unstructured data based on the unstructured data;

[0007] Construct a decay function based on the initial value of the unstructured data, and obtain the decayed decay value;

[0008] Determine a storage level factor one according to the decayed decay value, and generate a storage level decision one according to the storage level factor one;

[0009] Obtain the historical access record of the unstructured data, and determine the predicted value of the access probability of the unstructured data according to the historical access record;

[0010] Correct the decayed value of the unstructured data according to the predicted value of the access probability, and output the corrected decayed value;

[0011] Determine a storage level factor two according to the corrected decayed value, and generate a storage level decision two.

[0012] The present invention also discloses an intelligent hierarchical storage system for unstructured data, which is characterized in that it executes the above-mentioned intelligent hierarchical storage method for unstructured data, including

[0013] A data acquisition and preprocessing module, which is used to obtain unstructured data and extract the data generation timestamp and historical access records of the unstructured data based on the unstructured data;

[0014] An initial value evaluation module, which is used to determine the timeliness factor, access value factor, and data sensitivity, and determine the initial value based on the determined timeliness factor, access value factor, and data sensitivity; wherein the timeliness factor and access value factor are respectively determined based on the data generation timestamp and historical access records;

[0015] A value decay and correction module, which is used to construct a decay function based on the initial value of the unstructured data, and obtain the corrected decay value according to the predicted access probability value;

[0016] A storage decision and execution module, which is used to determine the storage level factor according to the decayed decay value and generate a storage level decision one according to the storage level factor one, determine the storage level factor two according to the corrected decay value and generate a storage level decision two according to the storage level factor two, and generate a storage level decision three by judging the migration condition based on the hysteresis interval and the fluctuation compensation value;

[0017] A dynamic parameter optimization module, which is used to update the sensitivity weight, update the characteristic lifetime parameter and the decay curve shape parameter based on the storage cost and access delay, and dynamically adjust the upper threshold and the lower threshold of the hysteresis interval according to the value fluctuation standard deviation.

[0018] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0019] 1. Break through the traditional single-frequency statistics, incorporate timeliness decay, access variance normalization, and multi-modal sensitivity into a unified quantization framework, effectively reducing the error rate; automatically identify the data that needs to be retained for a long time or stored with high security through sensitivity analysis, reducing the manual annotation cost;

[0020] 2. Generate different storage decisions based on the decay value and the corrected value respectively, support policy redundancy and cross-verification, decision one reflects the natural decay state of the data, and decision two reflects the external access intervention, which is convenient for the administrator to conduct root cause analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The disclosure of the present invention will be described with reference to the accompanying drawings. It should be understood that the drawings are only for illustrative purposes and are not intended to limit the scope of protection of the present invention. In the drawings, the same reference numerals are used to refer to the same components. Among them:

[0022] Figure 1 is the process flow diagram of the present invention;

[0023] Figure 2 is the initial value data flow diagram of the present invention;

[0024] Figure 3 is the data flow diagram of the corrected decay value of the present invention;

[0025] Figure 4 is the data sensitivity update flow diagram of the present invention;

[0026] Figure 5 is the migration judgment flow diagram of the present invention. Detailed implementation manners

[0027] It is easy to understand that according to the technical solution of the present invention, without changing the essence of the present invention, those of ordinary skill in the art can propose various structural ways and implementation ways that can be mutually replaced. Therefore, the following detailed implementation manners and the accompanying drawings are only illustrative descriptions of the technical solution of the present invention, and should not be regarded as all of the present invention or as a limitation or restriction on the technical solution of the present invention.

[0028] Application overview:

[0029] As mentioned above, unstructured data (such as images, videos, logs, documents, etc.) has become the core assets of enterprises with the popularization of technologies such as cloud computing and the Internet of Things. Due to its huge volume, there are great difficulties in how to store the data.

[0030] The following are two relatively common processing solutions:

[0031] Rely on manually preset strategies (such as the LRU algorithm to count the recent access times) or simple rules (such as the file creation time), without comprehensively considering multiple features such as data timeliness decay, content sensitivity, and business relevance.

[0032] Adopt a fixed archiving period (such as automatic downgrading after 30 days), but the data value decay law varies significantly due to type and scenario differences.

[0033] However, the above solutions all have their own disadvantages.

[0034] For the first one, medical imaging data needs to be stored for a long time due to compliance requirements, but its access frequency may be extremely low. Although traditional cold storage strategies reduce costs, they lead to a soaring delay in emergency retrieval, affecting the diagnosis and treatment efficiency.

[0035] For the second one, the popularity of news videos usually shows a "burst period - slow decay" curve, and the fixed life cycle model will downgrade it prematurely, missing the long-tail access benefits.

[0036] In view of the above defects in the prior art, the basic concept of this application is to combine rules with a machine learning model, and through a value decay model and a self-learning compensation mechanism, dynamically optimize the data storage strategy. Specifically, first, calculate the value of data decaying over time through the value decay model, and then use the self-learning compensation mechanism to dynamically adjust the decay curve by combining access probability prediction and value correction, so as to maximize the utilization value of data while ensuring optimized storage costs. Finally, the system dynamically selects the most suitable storage strategy according to the corrected value, realizing efficient and intelligent data management.

[0037] After introducing the basic concept of the present invention, the embodiments of the present invention will be specifically introduced below with reference to the accompanying drawings.

[0038] Embodiment 1:

[0039] As Figure 1 shown, an intelligent hierarchical storage method for unstructured data includes the following steps:

[0040] Obtain unstructured data and determine the initial value of the unstructured data based on the unstructured data.

[0041] As Figure 2 shown, determining the initial value of unstructured data based on unstructured data includes:

[0042] Obtain the data generation timestamp t create of the unstructured data, the historical access record Access i of the i-th access event, and the data sensitivity S;

[0043] Calculate the timeliness factor τ according to the data generation timestamp t create :

[0044] where t now is the current time, and λ is the industry decay coefficient.

[0045] Calculate the access value factor ν according to the historical access record Access i :

[0046] where, u i is the time decay weight of the i-th access event, ∈ is a constant to prevent division by zero, and Var(Access) is the variance of the historical access record Access.

[0047] u i is usually inversely proportional to the time distance of the access event, that is, the closer the access event is to the current time, the greater the weight; the farther the access event is, the smaller the weight. A commonly used method for calculating the time decay weight is the exponential decay model:

[0048] Among them, t now is the current time, t i is the time of the i-th access event, and λ is the decay coefficient, which controls the decay speed of the weight over time.

[0049] Suppose the historical access records are as follows:

[0050]

[0051] Calculate the access value factor ν:

[0052]

[0053] Calculate the initial value V0 according to the timeliness factor τ, the access value factor ν, and the data sensitivity S:

[0054] V0 = α·τ + β·ν + γ·S, where α, β, and γ are the weights for calculating the initial value V0 from the timeliness factor τ, the access value factor ν, and the data sensitivity S in sequence.

[0055] The calculation process of the data sensitivity S is as follows:

[0056] Use a natural language processing model to detect sensitive keywords (such as "ID number", "bank card number") in unstructured data and calculate the text sensitivity S text ;

[0057] Example:

[0058] Use an object detection model to identify sensitive regions (such as faces, license plates) in unstructured data and calculate the image sensitivity S image ;

[0059] Example:

[0060] Use a voiceprint recognition model to detect sensitive voice content (such as personal identity information) and calculate the audio sensitivity S audio ;

[0061] Example:

[0062] Assign compliance labels to the data (such as GDPR = 0.9, HIPAA = 0.95), and assign compliance sensitivity values S according to regulatory requirements reg ;

[0063] Example:

[0064] Assign a score according to the impact degree of the data on the business (e.g., transaction log = 0.9, test data = 0.2), and calculate the business value sensitivity S biz ;

[0065] Example:

[0066] Calculate the data sensitivity S:

[0067] S = w1·S text + w2·S image + w3·S audio + w4·S reg + w5·S biz ,

[0068] where w1, w2, w3, w4, w5 are the weights of the text sensitivity S text , image sensitivity S image , audio sensitivity S audio , compliance sensitivity value S reg , business value sensitivity S biz respectively, and ∑w i = 1, w i is the dynamic weight of the i-th factor.

[0069] Example:

[0070] Financial transaction log data:

[0071] Data content: The keyword "bank card number" is detected, S text = 0.8.

[0072] Regulatory compliance: Protected by GDPR, S reg = 0.9.

[0073] Business value: The transaction log is crucial to the business, S biz = 0.9.

[0074] Comprehensive sensitivity: S = 0.4·0.8 + 0.3·0.9 + 0.3·0.9 = 0.86.

[0075] Test data:

[0076] Data content: No sensitive keywords, S text = 0.1.

[0077] Regulatory compliance: No special regulatory requirements, S reg = 0.5.

[0078] Business value: The test data has a low impact on the business, S biz = 0.2.

[0079] Comprehensive sensitivity: S = 0.4·0.1 + 0.3·0.5 + 0.3·0.2 = 0.25.

[0080] Construct a decay function for the initial value based on unstructured data and obtain the decayed value after decay.

[0081] The decayed value after decay is:

[0082]

[0083] where V0 is the initial value;

[0084] η0 is the characteristic lifetime parameter, representing the time required for the data value to decay to (about 36.8%) of the initial value, controlling the overall shape of the decay curve. The larger η is, the slower the decay speed.

[0085] Example:

[0086] η = 72 hours: It means that the data value decays to 36.8% of the initial value after 72 hours.

[0087] β0 is the decay curve shape parameter, controlling the shape of the decay curve.

[0088] β < 1: The decay speed gradually slows down (fast decay in the early stage, slow decay in the later stage);

[0089] β = 1: The decay speed is constant (exponential decay);

[0090] β > 1: The decay speed gradually accelerates (slow decay in the early stage, fast decay in the later stage).

[0091] Example:

[0092] β = 1.2: It means that the data value decay speed gradually accelerates.

[0093] V(t) is the value of the data at time t, reflecting the current importance or sensitivity of the data, and is used to determine the storage level and migration strategy of the data.

[0094] Determine the storage level factor one according to the decayed value after decay, and generate the storage level decision one according to the storage level factor one;

[0095] Divide the storage levels according to the value threshold, as shown in the following table:

[0096]

[0097] As can be seen from the above table, the decay function V(t) is used as the input for storage level decision. According to the decayed value, the data is divided into three levels: hot, warm, and cold for storage.

[0098] Such asFigure 3 As shown, obtain the historical access records of unstructured data, and determine the predicted value of the access probability of the unstructured data according to the historical access records;

[0099] The implementation steps for determining the predicted value of the access probability of unstructured data according to the historical access records are as follows:

[0100] Input the historical access record Access i , and normalize the historical access record Access i to the interval [0, 1], and divide the time series data into sliding windows X of a fixed length as the input features of the model;

[0101] Use the trained LSTM-TCN hybrid model to process the sliding window X to obtain the predicted value P of the access probability access , P access = LSTM-TCN(Access i );

[0102] In the LSTM-TCN hybrid model, the loss function is:

[0103] where Y i is the true value, is the predicted value;

[0104] Use the Adam optimizer, divide the data set into a training set and a validation set, and iteratively train the model until the loss function converges or reaches the maximum number of training epochs.

[0105] Revise the decay value of the unstructured data according to the predicted value of the access probability, and output the revised decay value;

[0106] The implementation process is as follows:

[0107] Obtain the data value V(t) at the current time point t and the predicted value P of the access probability access ;

[0108] Perform value correction calculation to obtain the revised decay value V adjusted (t), and the correction formula is:

[0109] V adjusted (t) = V(t) · [1 + α · (P access - 0.5)], where α is the industry correction coefficient used to control the amplitude of the correction, and 0.5 is a reference value used to judge the high or low access probability.

[0110] If P access > 0.5, it means the access probability is relatively high;

[0111] At this time (Paccess - 0.5) is a positive number, and the correction factor α·(P access - 0.5) is also a positive number.

[0112] [1 + α·(P access - 0.5)] in the formula is greater than 1, so V adjusted (t) > V(t), that is, the value is corrected upward.

[0113] If P access < 0.5, it means the access probability is relatively low;

[0114] At this time, (P access - 0.5) is a negative number, and the correction factor α·(P access - 0.5) is also a negative number.

[0115] [1 + α·(P access - 0.5)] in the formula is less than 1, so V adjusted (t) < V(t), that is, the value is corrected downward.

[0116] If P access = 0.5, it means the access probability is at the baseline level;

[0117] At this time, (P access - 0.5) = 0, and the correction factor is 0.

[0118] [1 + α·(P access - 0.5)] = 1, so V adjusted (t) = V(t), that is, the value remains unchanged.

[0119] Example:

[0120] Use the LSTM - TCN model to predict the access probability for the next 7 days: P access = LSTM - TCN(Access i ) = 0.6;

[0121] Calculate the access probability deviation: P access - 0.5 = 0.6 - 0.5 = 0.1;

[0122] Calculate the correction factor: α·(P access - 0.5) = 0.2·0.1 = 0.02;

[0123] Dynamic correction attenuation: V adjusted (t) = 1000·[1 + 0.02] = 1000·1.02 = 1020.

[0124] Determine the storage level factor two based on the corrected attenuation value, and generate the storage level decision two.

[0125] Example 2: On the basis of Example 1, this example adds a process for adjusting data sensitivity, as follows Figure 4 shown;

[0126] After calculating the data sensitivity S, use the reinforcement learning model to calculate the feedback score Q of each factor i , and dynamically adjust the weights:

[0127] where exp(Q i ) is the exponential function value of the feedback score of the i-th factor, is the normalization factor to ensure that the sum of all weights is 1.

[0128] According to the dynamically adjusted weights, update the data sensitivity S again.

[0129] Example:

[0130] Suppose the text sensitivity S text , image sensitivity S image , audio sensitivity S audio , compliance sensitivity value S reg , business value sensitivity S biz are 0.8, 0.6, 0.9, 0.5, 1.0 in sequence, and the initial feedback scores of these five parameters are 3, 2, 4, 1, 5 in sequence.

[0131] Calculate exp(Q i ) = [20.09, 7.39, 54.60, 2.72, 148.41];

[0132] Calculate the normalization factor

[0133] Calculate the weight w i = [0.086, 0.032, 0.234, 0.012, 0.636];

[0134] Use the reinforcement learning model to update Q i (Suppose it is 4, 3, 5, 2, 6 after update);

[0135] Recalculate the weight w i = [0.119, 0.048, 0.324, 0.016, 0.493];

[0136] Recalculate the sensitivity S = 0.119·0.8 + 0.048·0.6 + 0.324·0.9 + 0.016·0.5 + 0.493·1.0 = 0.879.

[0137] Dynamically optimize the calculation of data sensitivity through a reinforcement learning model, thereby improving the accuracy and adaptability of data processing.

[0138] Specifically: Use a reinforcement learning model to calculate the feedback score Q of each factor i , and amplify the score difference through the exponential function exp(Q i ). Dynamically adjust the weight wi of each factor according to the feedback score to ensure that the weight can reflect the actual importance of each factor. Recalculate the data sensitivity according to the adjusted weight to make it more accurate and reliable. By dynamically adjusting the weight, the data sensitivity can adapt to changes in data characteristics and the environment, improving the effect of data processing and analysis. Make the calculation process of data sensitivity have self-learning and self-adaptive capabilities, reduce manual intervention, and improve the intelligent level of the system.

[0139] This step dynamically adjusts the weight through a reinforcement learning model, optimizes the calculation of data sensitivity, ensures that data processing is more accurate, flexible and intelligent, and provides a more reliable basis for subsequent storage strategy optimization and data analysis.

[0140] Example 3, on the basis of Example 1, this example adds the following process:

[0141] Based on the storage cost Cost and access latency Latency, update the parameter feature lifetime parameter η0 and decay curve shape parameter β0 using the Bayesian optimization framework.

[0142] Bayesian optimization is a global optimization method based on a probability model, suitable for scenarios where the calculation cost of the objective function is high or the derivative cannot be directly obtained.

[0143] Example:

[0144] Initial parameters: η0 = 72, β0 = 1.2;

[0145] Input: Storage cost Cost = $1000, access latency Latency = 50 milliseconds;

[0146] Decay function:

[0147] Process: Use Bayesian optimization to update the parameters:

[0148] η0 = 90, β0 = 1.1;

[0149] Input: Optimized decay function:

[0150] By using the Bayesian optimization framework to dynamically update the feature lifetime parameter η0 and decay curve shape parameter β0, it can effectively balance the storage cost and access latency and optimize the data storage strategy. This method has the following advantages:

[0151] Global optimization: Bayesian optimization can find the global optimal solution and avoid falling into local optima.

[0152] High efficiency: Through the surrogate model and acquisition function, the number of calculations of the objective function is reduced.

[0153] Self-adaptability: It can dynamically adjust parameters and decay curves according to data characteristics and environmental changes.

[0154] Example 4, on the basis of Example 1, this example adds a process of migration judgment;

[0155] Specifically as Figure 5 shown, the specific implementation is as follows:

[0156] Determining the storage level factor according to the decayed decay value further includes:

[0157] Define the upper threshold θ high and the lower threshold θ low ;

[0158] Calculate the standard deviation of value fluctuation σ v :

[0159] Calculate the fluctuation compensation value Compensation, Compensation = εσ v , where ε is the compensation constant;

[0160] Judge the migration condition:

[0161]

[0162] Warming migration: Move data from the cold layer → warm layer or warm layer → hot layer;

[0163] Cooling migration: Move data from the hot layer → warm layer or warm layer → cold layer;

[0164] Generate the storage level decision three.

[0165] Example: The current data value V(t) is 0.81, the upper threshold θ high is 0.8, the lower threshold θ low is 0.4, the compensation constant ε is 0.1, and the standard deviation of value fluctuation σ v is 0.11.

[0166] Calculate the fluctuation compensation value Compensation = 0.011.

[0167] Warming migration threshold: 0.8 + 0.011 = 0.811;

[0168] Cooling migration threshold: 0.4 - 0.011 = 0.389;

[0169] The current value V(t) does not exceed the warming migration threshold, and migration is not triggered.

[0170] By introducing a hysteresis interval, frequent migrations caused by short-term fluctuations in data value are avoided, the system load is reduced, invalid migration operations are reduced, and storage resource allocation is optimized. The hysteresis interval is dynamically adjusted to adapt to the fluctuation characteristics of different data values.

[0171] Standard deviation of value fluctuation σ v is obtained based on the sequence V of data values V(t) within the acquisition time window T. The data value sequence V = [V(t1), V(t2),..., V(t n )], where t1, t2,..., t n are the sampling time points within the time window respectively, and n is the number of samplings; the standard deviation of value fluctuation σ v has the following calculation formula:

[0172] In the formula, μ v is the mean value of the data value sequence V.

[0173] The standard deviation of value fluctuation σ v is used to measure the degree of fluctuation of the data value V(t) and reflects the stability of the data value. It is a statistic of the change in data value over a period of time and is used to optimize the storage migration strategy to avoid frequent migrations.

[0174] Example: Within the time window T = 24 hours, the data value V(t) is collected once per hour:

[0175] V = [0.9, 0.85, 0.88, 0.82, 0.78, 0.75, 0.72, 0.70, 0.68, 0.65, 0.63, 0.60,...];

[0176] Calculate the mean value:

[0177] Calculate the variance:

[0178] Calculate the standard deviation:

[0179] By introducing σ v , frequent migrations caused by short-term fluctuations in data value are avoided, the system load is reduced; the hysteresis interval is dynamically adjusted to adapt to the fluctuation characteristics of different data values; invalid migration operations are reduced, and storage resource allocation is optimized.

[0180] Embodiment 5. The present invention provides an intelligent hierarchical storage system for unstructured data, which executes the above-mentioned intelligent hierarchical storage method for unstructured data, including

[0181] A data collection and preprocessing module, which is used to obtain unstructured data and extract the data generation timestamp and historical access records of the unstructured data based on the unstructured data; obtain unstructured data through data collection tools (such as crawlers, APIs, etc.). Use a timestamp extraction tool and an access log analysis tool to generate timestamps and access records.

[0182] An initial value evaluation module, which is used to determine the timeliness factor, access value factor, and data sensitivity, and determine the initial value based on the determined timeliness factor, access value factor, and data sensitivity; among them, the timeliness factor and access value factor are determined based on the data generation timestamp and historical access records respectively;

[0183] A value decay and correction module, which is used to construct a decay function based on the initial value of the unstructured data and calculate the decay value of the data. Obtain the corrected decay value according to the predicted access probability value and dynamically correct the decay value.

[0184] A storage decision and execution module, which is used to determine the storage level factor one according to the decayed decay value and generate the storage level decision one according to the storage level factor one, determine the storage level factor two according to the corrected decay value and generate the storage level decision two according to the storage level factor two, and judge the migration condition based on the hysteresis interval and the fluctuation compensation value to generate the storage level decision three;

[0185] A dynamic parameter optimization module, which is used to update the sensitivity weight, update the characteristic lifetime parameter and the decay curve shape parameter based on the storage cost and access delay, and dynamically adjust the upper threshold and the lower threshold of the hysteresis interval according to the value fluctuation standard deviation.

[0186] The technical scope of the present invention is not limited to the content described above. Those skilled in the art can make various deformations and modifications to the above embodiments without departing from the technical idea of the present invention, and these deformations and modifications should all fall within the protection scope of the present invention.

Claims

1. An intelligent hierarchical storage method for unstructured data, characterized in that: The steps are as follows: Obtain unstructured data and determine the initial value of the unstructured data based on the unstructured data; Construct an attenuation function based on the initial value of the unstructured data and obtain the attenuated value after attenuation; Determine the storage level factor one according to the attenuated value after attenuation, and generate a storage level decision one according to the storage level factor one; Obtain the historical access record of the unstructured data and determine the predicted value of the access probability of the unstructured data according to the historical access record; Correct the attenuated value of the unstructured data according to the predicted value of the access probability, and output the corrected attenuated value; Determine the storage level factor two according to the corrected attenuated value and generate a storage level decision two.

2. The intelligent hierarchical storage method for unstructured data according to claim 1, wherein: The determining the initial value of the unstructured data based on the unstructured data includes: Obtain the data generation timestamp t of unstructured data based on unstructured data create and the historical access record Access of the i-th access event i and the data sensitivity S; Generate a timestamp t based on the data create Calculate the timeliness factor τ: where t now is the current time and λ is the industry decay coefficient; According to the historical access record Access i Calculate the access value factor ν: where u i is the time decay weight of the i-th access event, ∈ is a constant to prevent division by zero, and Var(Access) is the variance of the historical access record Access; Calculate the initial value V0 according to the timeliness factor τ, the access value factor ν, and the data sensitivity S: V0 = α·τ + β·ν + γ·S, where α, β, and γ are the weights for calculating the initial value V0 of the timeliness factor τ, the access value factor ν, and the data sensitivity S in sequence.

3. The intelligent hierarchical storage method for unstructured data according to claim 2, characterized in that: The calculation process of the data sensitivity S is as follows: Use a natural language processing model to detect sensitive keywords in unstructured data and calculate the text sensitivity S text ; Identify sensitive regions in unstructured data using a target detection model and calculate the image sensitivity S image ; Use a voiceprint recognition model to detect sensitive speech content and calculate the audio sensitivity S audio ; Attach compliance labels to the data and assign a compliance sensitivity value S according to regulatory requirements reg ; Assign a score according to the impact degree of the data on the business, and calculate the business value sensitivity S biz ; Calculate the data sensitivity S: S = w1·S text + w2·S image + w3·S audio + w4·S reg + w5·S biz , where w1, w2, w3, w4, and w5 are the weights of the text sensitivity S text , the image sensitivity S image , the audio sensitivity S audio , the compliance sensitivity value S reg , the business value sensitivity S biz , and ∑w i = 1, where w i is the dynamic weight of the i-th factor.

4. The intelligent hierarchical storage method for unstructured data according to claim 4, characterized in that: Use a reinforcement learning model to calculate the feedback score Q of each factor i , and dynamically adjust the weights: where exp(Q i ) is the exponential function value of the feedback score of the i-th factor; Update the data sensitivity S again according to the dynamically adjusted weight.

5. The intelligent hierarchical storage method for unstructured data according to claim 1, wherein: The attenuated value after attenuation is: where V0 is the initial value, η0 is the characteristic life parameter, β0 and is the decay curve shape parameter; Based on the storage cost Cost and the access latency Latency, update the parameter characteristic life parameter η0 and the attenuation curve shape parameter β0 using the Bayesian optimization framework.

6. The intelligent hierarchical storage method for unstructured data according to claim 1, characterized in that: The determining the storage level factor according to the attenuated value after attenuation further includes: Define the upper threshold θ of the hysteresis interval high , the lower threshold θ low ; Calculate the standard deviation of value fluctuations σ v : Calculate the fluctuation compensation value Compensation, Compensation = εσ v , where ε is the compensation constant; Judge the migration condition: Generate a storage level decision three.

7. The intelligent hierarchical storage method for unstructured data according to claim 1, characterized in that: The standard deviation of value fluctuations σ v is obtained based on the data value sequence V within the acquisition time window T, and the data value sequence V = [V(t1), V(t2),..., V(t n )], where t1, t2,..., t n are respectively the sampling time points within the time window, and n is the number of samplings; the formula for calculating the standard deviation of value fluctuations σ v is as follows: where μ v is the mean value of the data value sequence V.

8. The intelligent hierarchical storage method for unstructured data according to claim 1, wherein: The implementation steps of determining the predicted value of the access probability of the unstructured data according to the historical access record are as follows: Input historical access record Access i , and normalize the historical access record Access i to the interval [0, 1], and split the time series data into sliding windows X of a fixed length as the input features of the model; Using the trained LSTM-TCN hybrid model, process the sliding window X to obtain the access probability prediction value P access , P access = LSTM-TCN(Access i ); In the LSTM-TCN hybrid model, the loss function is: where Y i is the true value, and is the predicted value; Use the Adam optimizer, divide the data set into a training set and a validation set, and iteratively train the model until the loss function converges or reaches the maximum number of training epochs.

9. The intelligent hierarchical storage method for unstructured data according to claim 1, characterized in that: The implementation process of correcting the attenuated value of the unstructured data according to the predicted value of the access probability is as follows: Obtain the data value V(t) at the current time point t and the predicted access probability value P access ; Perform value correction calculation to obtain the corrected decay value V adjusted (t), and the correction formula is: V adjusted (t) = V(t)·[1 + α·(P access - 0.5)], where α is the industry correction coefficient.

10. An intelligent hierarchical storage system for unstructured data, characterized in that: Execute the intelligent hierarchical storage method for unstructured data as described in any one of claims 1 to 9, including A data acquisition and preprocessing module, which is used to obtain unstructured data and extract the data generation timestamp and historical access record of the unstructured data based on the unstructured data; An initial value evaluation module, which is used to determine the timeliness factor, the access value factor, and the data sensitivity, and determine the initial value based on the determined timeliness factor, access value factor, and data sensitivity; Wherein the timeliness factor and the access value factor are respectively determined based on the data generation timestamp and the historical access record; A value attenuation and correction module, which is used to construct an attenuation function based on the initial value of the unstructured data, and obtain the corrected attenuated value according to the predicted value of the access probability; A storage decision and execution module, which is used to determine the storage level factor one according to the attenuated value after attenuation and generate a storage level decision one according to the storage level factor one, determine the storage level factor two according to the corrected attenuated value and generate a storage level decision two according to the storage level factor two, and judge the migration condition based on the hysteresis interval and the fluctuation compensation value to generate a storage level decision three; A dynamic parameter optimization module, which is used to update the sensitivity weights, update the feature lifetime parameters and decay curve shape parameters based on the storage cost and access latency, and dynamically adjust the upper threshold and lower threshold of the hysteresis interval according to the standard deviation of value fluctuations.

Citation Information

Patent Citations

  • Data access method capable of facing file transfer protocol (ftp) service

    CN103152377A

  • Intelligent storage automatic grading method and device, storage medium and electronic equipment

    CN112306406A

  • File storage management method and storage service equipment

    CN117708044A

  • Dynamic data grading method

    CN117807535A

  • Multi-level data node storage indexing method and system

    CN118626685A

Cited By

  • Data quality evaluation method and system based on rule engine and machine learning

    CN121030265A

  • Unstructured storage hierarchical strategy optimization method based on machine learning

    CN121116940A

  • A machine learning-based unstructured storage tiering policy optimization method

    CN121116940B

  • Intelligent grading and excess subscription management system and method for GPU video memory

    CN121722577A