Data management method based on distributed storage system
By constructing a data management method based on behavior perception, the problem of policy mismatch in distributed storage systems is solved, dynamic perception and differentiated control of data access behavior is realized, and the accuracy of data management and system stability are improved.
Patent Information
- Application Number
- CN202510741289.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing distributed storage systems lack dynamic perception and classification management capabilities for data access behavior, resulting in policy mismatch, and problems such as redundant replica distribution imbalance, hot and cold resource configuration misalignment, and increased scheduling path delay.
By collecting and standardizing data access logs, a visit behavior sequence is constructed, access mutation rate, periodicity and span indicators are extracted to generate fusion behavior feature vectors, behavior pattern clustering and label division, strategy structures are generated, and scheduling calculations are performed in combination with node resource status, replica creation, compression configuration and hierarchical deployment are realized, and behavior path response processes are analyzed to generate scheduling management signals.
It realizes dynamic perception and differentiated control of data access behavior in distributed storage systems, improves the accuracy of data management and the stability of system operation, and reduces the problem of mismatch in policy execution.
Smart Images

Figure CN120255824A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management. More specifically, the present invention relates to a data management method based on a distributed storage system. Background Art
[0002] As a storage architecture for massive data, high scalability, and high elastic fault tolerance requirements, the distributed storage system is widely used in cloud platforms, Internet services, edge computing, and data-intensive application scenarios. The distributed storage system divides data into multiple segments and stores them distributively on different storage nodes to achieve data redundancy backup, concurrent read and write access, and replacement of faulty nodes. This system usually has basic functions such as a replica mechanism, heat awareness, multi-replica scheduling, and consistency control, and supports decoupling of storage services from the upper-layer business system.
[0003] Deficiencies of the prior art: In a distributed storage system, strategies such as data replica deployment, compression control, and storage heat stratification generally rely on preset rules or static templates, lacking the ability to dynamically perceive and classify manage data access behaviors. When behavioral characteristics of data objects change during operation, such as access pattern drift, sudden high-frequency reads, phased usage, or collaborative access, the storage system cannot effectively identify and respond with strategies, resulting in continuous mismatch of the original strategies, problems such as unbalanced distribution of redundant replicas, misalignment of cold and hot resource configurations, and increased scheduling path latency. Especially in a deployment environment with multiple nodes and heterogeneous resources, the lack of dynamic coupling between strategies and resource states further limits the accuracy, adaptability, and resource optimization ability of data management strategies. Summary of the Invention
[0004] To overcome the above-mentioned defects of the prior art, there are the following solutions to solve the problem of mismatch of storage management behavior strategies in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions: A data management method based on a distributed storage system, including the following steps: Collect and standardize data access logs in the distributed storage system, construct an access behavior sequence sorted by time, and extract access mutation rate, access periodicity, and access span metrics to generate a fusion-type behavior feature vector; Perform behavior pattern clustering and label division based on the fusion-type behavior feature vector, construct a mapping relationship corresponding to behavior categories and policy intentions, and generate a policy structure body to perform resource configuration control; Perform scheduling calculations on the node resource status according to the policy structure body, construct a cost function to select the optimal node, complete replica creation, compression configuration, and hierarchical deployment, establish a consistency path and record the scheduling results, and generate replica running status information for policy evolution; Analyze the behavioral path response process after the strategy deployment, obtain the management scheduling information in the behavioral path compression and execution displacement trend and analyze it, and perform scheduling management according to different signals generated by the analysis results.
[0006] In a preferred embodiment, collect and standardize the data access logs in the distributed storage system, and construct an access behavior sequence sorted by time. The specific steps include: Collect all access events of data objects in the data access logs. Each access event record includes the access time, the initiating node, the operation type, and the data volume. Perform standardization operations on the collected data, including unifying the timestamp accuracy to the second-level UTC format, normalizing the operation type to a unified identifier, filtering out non-business accesses, and merging duplicate access requests within a set time interval. Form an event sequence sorted by time from the access logs of the standardized data objects, and convert the preprocessed event sequence into a structured access behavior sequence.
[0007] In a preferred embodiment, extract access mutation rate, access periodicity, and access span metrics to generate a fusion-type behavior feature vector. The specific steps include: Take each valid access event in the event sequence as a behavior node of the data object and form a continuous behavior trajectory. Extract the access rate fluctuation feature, calculate the difference in access volume and the difference in time interval between adjacent accesses, judge whether there is a rate mutation in the access behavior, and if there is a rate mutation, obtain and use it as the access mutation rate metric. Perform periodic feature extraction, perform frequency domain analysis on the access time series, and observe the activities in the time period as the access concentration metric. Perform access concentration and span extraction. Use the time interval between the first and last accesses as the behavior span, and use the ratio of the front and back of the access concentration in the behavior span as the access concentration. Summarize the access concentration and span as the access span metric. Use the frequency domain analysis method to detect the periodic trend of the access behavior sequence, calculate the behavior fluctuation amplitude corresponding to each frequency, take the periodic component with the largest behavior fluctuation amplitude as the dominant period, and use a continuous time series formed by the change trend of the number of accesses per unit time as the access density change trajectory. Perform analysis on the context label information bound to the data object. The dimensions of the context label information include the data source module, the tenant or project code to which it belongs, the data function role label, whether it is collaborative access data, and the task priority. Perform context vector representation according to the context embedding function, and project the context label into adjacent positions in the Euclidean space according to the semantic structure. After completing the access behavior modeling, a fusion behavior feature vector is generated.
[0008] In a preferred embodiment, based on the fusion behavior feature vector, behavior pattern clustering and label partitioning are performed, a mapping relationship corresponding to the behavior category and policy intention is constructed, and a policy structure is generated to perform resource allocation control. The specific steps include: Use the logarithmic function to regularize the values of different feature dimensions, and use the clustering method optimized based on density gradient to divide the feature space; Define behavior categories corresponding to different access patterns and business semantics. The behavior categories specifically include: periodic access type, sudden hot spot type, long-term cold storage type, collaborative sharing type, and unstable type; Assign a unique behavior label to each data object; Map the behavior label to a set of policy parameters. The policy structure includes fields: the number of replicas, the storage heat level identifier, whether to enable compression, the replica deployment area, and the consistency level requirement; According to the differences in behavior labels, perform different rule mappings; For the periodic access type, set the periodic compression and delayed replica refresh policies; For the sudden hot spot type, set the main replica to reside and enable the cache layer snapshot; For the long-term cold storage type, perform compression, replica redundancy, and archival storage processing; For the collaborative sharing type, set the replica distribution and enable request scheduling; For the unstable type, trigger periodic behavior recognition.
[0009] In a preferred embodiment, according to the policy structure, the node resource status is scheduled and calculated, a cost function is constructed for optimal node selection, replica creation, compression configuration, and hierarchical deployment are completed, a consistency path is established and the scheduling result is recorded, and replica running status information is generated for policy evolution, including the following steps: Construct a set of dynamic state parameters for each storage node to form a resource status vector. The resource status vector includes CPU occupancy rate, disk space usage rate, network bandwidth occupancy ratio, real-time packet loss rate, and node error history; On the basis of meeting the policy constraints, perform replica deployment, construct a cost function, which measures the cost of each storage node as a candidate replica deployment point, and select storage nodes based on the cost function; Sort the candidate storage nodes based on the cost function result, and select the node combination with the smallest total cost from them for replica deployment; Construct an ascending sorting candidate set according to the cost function, select nodes according to the ascending sorting candidate set and form a target deployment set; Send deployment instructions to each node in the target deployment set to perform deployment operations: The deployment operation refers to the physical write initialization of the data replica. If the compression strategy is 1, the compressor is enabled to perform structural compression on the data replica, and according to the storage heat level identifier, the replica is stored in the corresponding level storage medium. Then, according to the consistency level requirements, the data synchronization path and confirmation mechanism between the master and slave replicas are established; After the deployment is completed, a scheduling task execution log is generated, binding the task number and behavior label.
[0010] In a preferred embodiment, analyzing the behavior path response process after the strategy deployment, obtaining and analyzing the management scheduling information in the behavior path compression and execution displacement trend, includes the following steps: Acquire management scheduling information generated during the behavior path compression and execution displacement trend analysis, the management scheduling information including the behavior path compression index and the strategy guidance offset displacement index; The behavior path compression index indicates whether the actual access path of the data object tends to be stable and convergent after the policy deployment; The policy-guided offset displacement index indicates the degree of topological offset between the target deployment area set by the policy structure and the actual replica scheduling and access behavior; The obtained behavior path compression index and strategy-guided offset displacement index are combined to generate a data management coefficient; The behavior path compression index is positively correlated with the data management coefficient, and the strategy-guided offset displacement index is negatively correlated with the data management coefficient.
[0011] In a preferred embodiment, scheduling management is performed according to different signals generated by the analysis results, including the following steps: Compare the generated data management coefficient with the set data management threshold; If the data management coefficient is greater than or equal to the data management threshold, a stable execution signal is generated, and the current strategy structure and corresponding scheduling path are maintained, and the original scheduling plan continues to be executed; If the data management coefficient is less than the data management threshold, an evolution adjustment signal is generated, the scheduling plan is stopped, the behavior re-identification mechanism is triggered, the behavior labels are reallocated and the new strategy structure is mapped.
[0012] The technical effects and advantages of the data management method based on the distributed storage system of the present invention are as follows: The present invention realizes dynamic perception and differential control of data access behavior in a distributed storage system by constructing a data management mechanism based on behavior perception and policy linkage. By collecting and standardizing access logs, extracting access mutation rate, periodicity, and access span metrics, generating a behavior feature vector that integrates context semantics, and identifying behavior patterns based on clustering methods, a policy structure body with replica quantity, compression strategy, heat level, and consistency requirements is mapped and generated. Combining the node resource status to construct a scheduling cost function, the optimal deployment of replicas, hierarchical storage, and consistency configuration are realized; After the policy is executed, analyze the compressibility and offset trend of the behavior path, extract management scheduling features, construct a data management coefficient, and generate a stable execution or policy adjustment signal, so as to realize the policy closed-loop control and resource dynamic scheduling driven by behavior, reduce the problem of policy execution mismatch caused by the lack of behavior perception in the distributed storage system, and significantly improve the accuracy of data management and the stability of system operation in the distributed system. Brief Description of the Drawings
[0013] Figure 1 It is a schematic flow diagram of the data management method based on a distributed storage system of the present invention. Detailed Embodiment
[0014] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0015] To achieve the above object, Figure 1 A schematic structural diagram of the data management method based on a distributed storage system of the present invention is given, which specifically includes the following steps; Collect and standardize the data access logs in the distributed storage system, construct an access behavior sequence sorted by time, and extract access mutation rate, access periodicity, and access span metrics to generate a fusion behavior feature vector; Based on the fusion behavior feature vector, perform behavior pattern clustering and label division, construct a mapping relationship corresponding to behavior categories and policy intentions, generate a policy structure body, and perform resource configuration control; According to the policy structure body, perform scheduling calculations on the node resource status, construct a cost function to select the optimal node, complete replica creation, compression configuration, and hierarchical deployment, establish a consistency path, record the scheduling results, generate replica running status information, and perform policy evolution; Analyze the behavioral path response process after the analysis strategy is deployed, obtain the management scheduling information in the behavioral path compression and execution displacement trend, and analyze it. Perform scheduling management according to different signals generated by the analysis results.
[0016] In a distributed storage system, the access behaviors of different data objects show significant heterogeneity: some data has a high degree of periodicity, some data has short-term high-frequency access in specific scenarios, and some data may be in a low-frequency and passive reading state for a long time.
[0017] Step 1: Model and extract features of data access behaviors, that is, model the access behaviors of each data object and extract behavior vectors that can reflect its timing features, operation features, and context features. The specific steps are as follows: Collect and standardize data access logs. In the data path integrated in the distributed storage platform, collect all access events of the data objects therein. Each access event record includes the access time, the initiating node, the operation type (such as reading, writing, deleting, etc.), and the data volume. Perform standardization operations on the collected data. The standardization operations include unifying the timestamp accuracy to the second-level UTC format, normalizing the operation type to a unified identifier, filtering out non-business accesses, including health checks and heartbeat background task accesses, and merging duplicate access requests within a set time interval (for example, the set time interval is 0.1 seconds). Finally, the access logs of each data object form an event sequence sorted by time, in the format: (access time, operation type, access volume, source identifier, task context). This sequence serves as the original input for subsequent behavior analysis.
[0018] Construct the access behavior sequence and extract dynamic features, identify the access patterns of data objects in the time dimension, and convert the preprocessed event sequence into a structured access behavior sequence. During the construction process, define each valid access event in the event sequence as a behavior node of the data object to form a continuous behavior trajectory. Extract a set of dynamic access features to determine its access rate, fluctuation degree, and mutation characteristics. Model from the following three dimensions: Access rate fluctuation feature: Statistically analyze the difference in access volume and the difference in time interval between adjacent accesses to determine whether there is an obvious rate mutation in the access behavior. If writing and reading frequently alternate in a certain period of access, or the data volume suddenly increases or decreases, then its behavior pattern is closer to high-fluctuation data, such as AI training task input files or log aggregation systems. Use the above as the designed access mutation rate index to represent the proportion of mutation points in the entire sequence. Periodic feature extraction. Periodicity is a core metric for measuring behavioral regularity. Perform frequency-domain analysis on the access time series to observe its activity intensity over specific time periods (such as minutes, hours, days). If the behavior repeats within 24 hours, it indicates that the data is related to daily scheduled tasks. If there are peaks in high-frequency accesses per hour, it may be related to database caching or periodic synchronization. Use the above content as the access periodicity metric. By scanning multiple candidate periods (such as 10 minutes, 1 hour, 1 day, etc.), extract its dominant period and determine the strength of periodicity. Access concentration and span extraction. Define the behavior span as the time interval between the first and last accesses. Use the ratio of access concentration within the behavior span as the access concentration. This is used to distinguish between stable access and phased access data. For example, if a certain data is accessed 100 times concentratedly in the last hour and there is no access record in the previous 24 hours, it is regarded as phased access type. Use the above content as the access span metric.
[0019] Some data objects have obvious periodic behavior characteristics, such as regular backup tasks, periodic indicator writing, and daily log archiving. Perform periodic feature recognition and behavior density modeling on the data, and use the frequency-domain analysis method to detect the periodic trend of the access behavior sequence. By constructing a frequency scanning grid at multiple time scales (such as ten minutes, one hour, one day), calculate the behavior fluctuation amplitude corresponding to each frequency respectively. Take the period component with the largest behavior fluctuation amplitude as the dominant period. In the recognition result, if there is a continuous and stable amplitude peak at a specific period, it can be considered that the data has periodic access characteristics. This feature is an important criterion for distinguishing periodic and non-periodic data. At the same time, according to the change trend of the access quantity per unit time, calculate the access density change trajectory. The access density change trajectory is a continuous time series composed of the change trends of the access quantity in consecutive time periods per unit time, which helps to identify short-term burst accesses or distributed behavior transfers. Combining the mutation rate and access density fluctuations, it can be judged whether there is a behavior phase transition, such as the access evolution from cold to hot or from hot to cold. In a distributed storage system, although the behavior patterns (such as mutation rate, periodicity, access density, etc.) extracted from the access time series can reflect certain usage trends, they still lack auxiliary information on the usage context. In actual business, the access behavior of data objects is often jointly affected by non-time series factors such as its business role, tenant scenario, generation source, and task binding relationship. For example, even if two data objects have similar access frequencies and periodic structures, one may be a log archive file and the other may be a model parameter snapshot, and they should be completely different in terms of replica deployment, consistency requirements, and scheduling strategies. Therefore, relying solely on behavior characteristics may cause policy misjudgment or classification ambiguity. The context label information bound to the data object includes the following dimensions: Data source module (such as log center, AI task output, archiving system); tenant or project code (used for resource isolation and access privilege control); data function role label (such as input source, task intermediate results, final output); whether the data is accessed collaboratively (such as whether it is accessed by multiple nodes); task priority (Such as whether to bind real-time tasks or asynchronous tasks).
[0020] Introduce context embedding function: ,in, is the j-th context label, j=1, 2,…, k, k is the number of label categories; is the context encoding function (which can use distributed embedding table, sparse hash map or low-rank coding); is the context vector representation, Represents the real number R dimensional vector space, generally a dense vector, for example, a vector in a three-dimensional space belongs to ; Project context labels with similar semantic structures to adjacent positions in the Euclidean space to maintain the consistency of their task semantics or resource preferences; After completing the access behavior modeling part (such as mutation rate, cycle, access concentration, etc.), let the behavior vector be , represents the behavior vector of the ith data object, that is, the feature representation extracted by modeling its access behavior sequence. Represents the real number R dimensional vector space, the final fusion behavior feature vector is constructed as follows: ,in, Represents a vector concatenation operation; is the full-dimensional behavior representation that is ultimately used for behavior pattern recognition. Represents the real number R dimensional vector space.
[0021] Step 1 has completed the modeling of the access behavior of each data object and output structured behavior feature vectors. These vectors carry the behavior information of the data object, but the feature vectors themselves cannot directly guide the system to adopt what kind of resource allocation and policy control. Therefore, the behavior feature vectors need to be parsed into clear behavior category labels, and these labels are further bound to specific storage policies to perform differentiated responses; Step 2: Behavior pattern recognition and policy intent mapping. The specific steps are as follows: Introduce a non - linear mapping function (such as logarithmic function, Sigmoid function) to normalize the values of different feature dimensions. Taking the logarithmic mapping as an example, its form is: , where is a tuning parameter used to control the compression degree, usually selected between 0.1 and 1.0; x is the original feature value, that is, a certain original numerical value in the behavior feature vector; is the normalized feature value after non - linear mapping, and log represents the logarithmic function. This operation can compress large values into a comparable range and enhance the discriminability of low - value features. It is especially suitable for features with low amplitude and high information density such as access mutation rate. After completion, all features are normalized to a unified relative magnitude, facilitating the clustering algorithm to reasonably determine the boundaries in the multi - dimensional space; Perform behavior pattern clustering and label assignment. Use a clustering method optimized based on density gradient (such as spectral clustering based on bimodal point convergence or minimum redundancy clustering) to divide the feature space, avoiding excessive assumptions about the Euclidean space distribution. Define the center of each class as , is the dimensional vector space of real numbers R, is the center of the overall behavior feature distribution of the clustered data object J. For any assign a behavior class label: , where is a non - Euclidean metric function, such as a projection distance function or a negative entropy kernel function under the Mahalanobis distance structure, used to avoid misjudgment of non - convex distributions by linear distances, represents the input vector of the i - th data object after feature normalization or mapping processing; Pre - define several typical behavior categories, corresponding to different access patterns and business semantics, specifically including: periodic access type , with fixed - frequency timed writing or querying behaviors; sudden hot - spot type , with high - frequency writing or reading in a short time; long - term cold storage type , with no active access for a long time or only low - frequency reading; collaborative sharing type , with frequent access by multiple nodes or users; non - stable type , with frequent jumps in behavior patterns and access noise.
[0022] Each data object is finally assigned a unique behavior label , and this label is used as the behavior recognition result for policy reasoning.
[0023] The behavior label itself is just a semantic classification result and does not have the ability to directly control the operations of the storage system. Therefore, map the behavior label to a set of policy parameters; Design a policy structure , the policy structure includes the following fields: the number of replicas (determining the redundancy level), the storage heat level identifier (determining the data placement layer, such as hot storage, archive, cold chain), whether compression is enabled (controlling whether to compress and the compression depth), the replica deployment area (such as deploying to edge nodes or the backbone core), and the consistency level requirement (controlling the replica synchronization intensity); According to different behavior labels, the mapping rules are as follows: for periodic access types, set regular compression and delayed replica refresh policies; for sudden hotspot types, set the main replica to reside in the core node and enable fast snapshots at the cache level; for cold data types, enable high compression, low replica redundancy, and archive storage; for collaborative access types, set multi-point distribution of replicas and enable request-based near scheduling; for unstable types, trigger short-cycle behavior re-identification and rapid policy iteration channels.
[0024] Write the policy structure into the policy configuration center and establish an index for each data object for the subsequent policy execution scheduling module to read and perform operations such as replica creation, compression management, hot and cold tiering, and data migration accordingly. If the system has an online policy adjustment mechanism, the historical sequences of behavior labels and behavior characteristics will be continuously monitored to trigger the policy evolution mechanism.
[0025] Each data object is assigned a behavior label, and based on this, a policy structure is generated. This policy structure includes key fields such as the number of replicas, the replica deployment area, the storage heat level, the compression requirement, and the consistency level. The policy structure still belongs to a static configuration template and cannot be directly executed. Due to the dynamic changes in the resource status of each node, frequent fluctuations in the network status, and significant differences in disaster recovery costs in a distributed system, the implementation effects of policies vary greatly at different times and in different spaces; In a distributed storage system, there are significant heterogeneities and load fluctuations among nodes, and the execution costs of the same policy vary significantly at different time points and on different nodes. For example, if a node with disk I / O approaching saturation continues to receive replica write requests, it will further degrade performance and may cause task backlogs or service drifts; Step 3, perform policy linkage execution and resource-aware scheduling, that is, select the optimal combination of storage nodes, path scheduling, and storage layout methods on the premise of meeting the policy intent. The specific steps are as follows: Construct a set of dynamic state parameters for each storage node to form a resource status vector. This vector should at least include the following content: CPU occupancy rate, reflecting the usage of computing resources; disk space utilization rate, reflecting the storage pressure; current I / O load (such as the length of the write / read request queue); network bandwidth occupancy ratio and real-time packet loss rate; node error history (such as the number of abnormal responses within 30 minutes); task completion reliability assessment (based on the latency and failure of historical task executions).
[0026] Perform policy constraint interpretation and construction of the scheduling objective function, map the policy to the optimal node combination in the current resource state, that is, on the basis of meeting the policy constraints, complete replica deployment at the lowest cost, measure the cost of each storage node as a candidate replica deployment point, and select storage nodes based on this cost function; Construct the cost function as follows: , where is the CPU occupancy rate of the node; is the disk usage rate of the node; is the current bandwidth throughput of the node; is a constant to prevent division by zero; and and are the weights of the CPU occupancy rate of the scheduling node, the weight of the node disk usage rate, and the weight of the current bandwidth throughput of the node respectively, and can be defined differently according to actual needs; is the execution cost of deploying the replica of the i-th data object to the j-th storage node; It should be noted that the task completion reliability evaluation is used to measure the execution success rate and execution quality of a storage node for the system-assigned tasks (such as replica creation, synchronous writing, compression operation, heat migration, etc.) in the past scheduling cycle, and a segmented quantitative score is given according to the completion quality of each task as the task completion reliability evaluation; the design idea of the cost function is that the more tense the resource usage, the weaker the network throughput capacity, and the lower the historical task success rate, the greater the scheduling cost of the node, and the less suitable it is as the target replica deployment point.
[0027] Based on the results of the cost function, sort the candidate storage nodes, and select the node combination with the smallest total cost from them for replica deployment; Construct an ascending-order sorting candidate set according to the cost function, sort according to the size of the cost function, and the smaller the cost function, the higher the sorting; select the first N nodes of the ascending-order sorting candidate set to form the target deployment set; Send deployment instructions to each node in the target deployment set and perform the following deployment operations: The deployment operation refers to the physical write initialization of the data replica. If the compression policy is 1, enable the compressor, perform structural compression on the replica, and store the replica in the corresponding hierarchical storage medium (such as SSD cache, HDD archive) according to the storage heat level identifier, and then establish a data synchronization path and confirmation mechanism between the master and slave replicas according to the consistency level requirements.
[0028] After the deployment is completed, generate a scheduling task execution log and bind the task number and behavior label; Write the behavior tags, deployment replica list, and final policy structure into the policy status table, record the actual deployment success timestamp, path, and resource usage, point the current state pointer of the data object to the policy replica mapping, initialize the monitoring tasks for each replica, and record the synchronization status, access latency, and error count for generating the policy execution feedback vector.
[0029] Step 4: Conduct self-evolution analysis of the behavior feedback closed-loop policy. The specific steps are as follows: Analyze the behavior path response process after policy deployment, and obtain the management scheduling information in the behavior path compression and execution displacement trends. The management scheduling information includes the behavior path compression index and the policy guidance offset displacement index. The behavior path compression index is used to measure the degree to which the actual access path of the data object tends to be stable and convergent after policy deployment. The behavior path compression index is based on the ratio between the cumulative total number of hops of the access path and the number of path uniquenesses. By modeling the jumpiness and distribution aggregation degree of the access path, it reflects whether the system has formed a behavior path with structural consistency after policy execution, measures the distribution aggregation and execution trajectory compression of the access path, and thus reflects the convergence state of the behavior path under policy guidance. If the access path has a high concentration, strong repeatability, and a stable hop count distribution, it indicates that the scheduling path has been gradually compressed and a stable access trajectory has been formed. The behavior path compression index is close to 1, indicating that the policy scheduling has a good path aggregation effect. On the contrary, if the access path distribution is divergent, changes frequently, and the hop counts are inconsistent, the value of the behavior path compression index approaches 0, indicating that there is a behavior path oscillation in the system and the guiding effect of the policy on the behavior weakens. The role of this index is to assist in judging whether the policy has achieved structural convergence at the path level and is a comprehensive evaluation index for the path consistency of the policy execution result and the behavior response stability. The acquisition logic of the behavior path compression index is as follows: After the policy structure is executed, within the specified observation period collect all access request records associated with the data object. represents the start time point of the observation period. represents the duration (time length) of the observation period, that is, the time window length observed backward from Express the node jump structure of each access as an ordered path vector to form a path set: , where u is the number of accesses within the observation period. For the path set statistically sum up the total number of hops of all paths: , where represents the number of hops in path . Identify the number of unique path structures , the expression is: , where , represents the set of unique paths after deduplication from , based on the judgment of node sequence congruence; calculate the behavior path compression index, and the calculation expression is: .
[0030] The policy-guided offset displacement index is used to measure the offset degree in the topological structure between the target deployment area set by the policy structure and the actual replica scheduling and access behavior; The policy-guided offset displacement index calculates the minimum structural displacement distance of the actual replica access or scheduling landing point relative to the set of policy target nodes, and constructs a logarithmic proportional measure of the path displacement, reflecting whether there is an obvious spatial deflection in the physical node level during the policy execution. If the actual access still revolves around the policy expected area and the system scheduling maintains structural concentration, the value of the policy-guided offset displacement index is relatively low; if the actual operation nodes after the policy deployment gradually move away from the initial target area and cause scheduling drift, the policy-guided offset displacement index rises rapidly, indicating that the policy control force gradually weakens in the geographical or topological dimension. This policy-guided offset displacement index can be used to identify the structural offset problem in the policy execution, guide whether to perform policy redirection, node migration or consistency level adjustment, and is an important spatial dimension index for the effectiveness of policy control; The acquisition logic of the policy-guided offset displacement index is as follows: Extract the set of policy target replica nodes when the data object is scheduled from the system deployment log: , where is the jth storage node in the distributed system. During the set policy execution period, collect the actual nodes involved in the scheduling behavior, access behavior and replica management behavior generated by the data object to form the set of actual operation landing points , where N is the complete set of callable operation response actions, is the kth storage node in the distributed system; Calculate the structural distance between two nodes , in the system network layer to determine the system-guided offset: , calculate the policy-guided offset displacement index, and the calculation expression is: .
[0031] It should be noted that the policy structure contains a deployment location constraint field, which indicates the set of node areas where the policy expects replicas to be deployed, such as specifying "edge node set", "cross-availability zone distribution", "core backbone node", etc. The policy target replica node set is the expected target operation area when the policy is deployed; the actual operation landing point set reflects the actual operation landing point of the system under the guidance of the current policy; the structural distance is determined by the number of network layer hops (such as the three-layer switching distance), which indicates the total minimum aggregate displacement in the topology between the actual operation behavior of the system and the policy target area.
[0032] The obtained behavior path compression index and strategy-guided offset displacement index are combined to generate the data management coefficient. The expression is: , where , are the preset proportional coefficients of the behavior path compression index and the strategy-guided offset displacement index, respectively, and , Both are greater than 0.
[0033] The specific method of jointly generating the data management coefficient may involve multiple algorithms and models, which depends on the actual situation and application requirements. The present embodiment may use a weighted summation method to combine the behavior path compression index and the strategy-guided offset displacement index to generate a comprehensive data management coefficient. This data management coefficient can be used as an input for distinguishing nodes in a distributed storage path and used to determine the final storage path situation.
[0034] It should be noted that the size of the preset proportional coefficient is a specific value obtained by quantifying each parameter. In order to facilitate subsequent comparison, the size of the coefficient depends on the amount of sample data and the preset proportional coefficient initially set by technical personnel in this field for each group of sample data. It is not unique, as long as it does not affect the proportional relationship between the parameter and the quantized value. For example, the behavior path compression index is proportional to the data management coefficient. The behavior path compression index and the strategy-guided offset displacement index are normalized to have the same dimension and range. This can be achieved by subtracting the mean from the original data and dividing it by the standard deviation, or mapping the data to the range of [0, 1].
[0035] The larger the behavior path compression index and the smaller the policy guidance deviation displacement index, the larger the jointly generated data management coefficient. This indicates that under the current policy control, the access path of the data object shows a high degree of structural aggregation, and at the same time, the scheduling behavior is highly concentrated in the target area set by the policy. The distributed storage system has achieved good scheduling convergence and policy response consistency in both the behavior path layer and the spatial structure layer. This state reflects a strong correlation between the behavior modeling result and the policy structure body, indicating that the policy deployment is accurate and the execution is stable. The storage system has good resource scheduling controllability and behavior management capabilities, which is an ideal embodiment of the execution effect of the behavior-driven policy; The smaller the behavior path compression index and the larger the policy guidance deviation displacement index, the smaller the jointly generated data management coefficient. This indicates that under the current policy structure control, there are still significant jumps or structural divergences in the access path of the data object. At the same time, the actual scheduling nodes deviate significantly from the policy target deployment area, and the system scheduling behavior fails to effectively converge to the expected structure. Such a state indicates that the current behavior characteristics may have evolved, or the policy deployment plan cannot adapt to the changes in system resources, resulting in out-of-control behavior responses, scattered paths, and a decline in the policy execution effect. At this time, the behavior re-identification and policy structure reconstruction mechanism should be initiated to re-establish a stable and efficient data management chain; Compare the generated data management coefficient with the pre-set data management threshold to generate a stable execution signal and an evolution adjustment signal; If the data management coefficient is greater than or equal to the data management threshold, and a stable execution signal is generated at this time, it indicates that the current access path of the data object has formed a stable aggregation trend under the policy control, and the actual scheduling behavior highly fits the policy deployment area. The policy control effect is good, the resource allocation is reasonable, and the system does not need to adjust the policy. Maintain the current policy structure body and the corresponding scheduling path, and the system continues to execute the original scheduling plan; If the data management coefficient is less than the data management threshold and an evolution adjustment signal is generated, it indicates that there is an obvious divergence trend in the current behavior path, or there has been an offset in the actual replica scheduling area where the policy guidance fails. The system scheduling efficiency and policy adaptability decline, triggering the behavior re-identification mechanism, stopping the scheduling plan, recalculating the behavior feature vector, reassigning the behavior label and mapping a new policy structure body, and at the same time updating the replica deployment and resource path, entering the policy self-evolution process.
[0036] Through the signal linkage mechanism, the system can achieve closed-loop determination of the policy execution status and differential response control based on the comparison results of thresholds. Without relying on manual rules, the system can dynamically determine whether the current policy is still effective and adaptable, thereby triggering policy reconstruction, resource rescheduling, or behavior re-identification in a timely manner, improving the accuracy and interpretability of data management decisions, enhancing the intelligent control ability of the distributed storage system in multiple scenarios and multiple behavior patterns, and contributing to the realization of more efficient and reliable data storage policy optimization goals.
[0037] It should be noted that the threshold information related to this embodiment is pre-set by professionals and will not be explained in detail here. In this embodiment, some parameter English letters are the same, but different meanings are explained when used, and they will not be explained one by one here.
[0038] The present invention realizes dynamic perception and differential control of data access behavior in a distributed storage system by constructing a data management mechanism based on behavior perception and policy linkage. By collecting and standardizing access logs, extracting access mutation rate, periodicity, and access span indicators, generating a behavior feature vector that integrates context semantics, and identifying behavior patterns based on clustering methods, mapping to generate a policy structure body with replica quantity, compression policy, heat level, and consistency requirements, and constructing a scheduling cost function in combination with node resource status to achieve optimal deployment of replicas, hierarchical storage, and consistency configuration. After the policy is executed, analyze the compressibility and offset trend of the behavior path, extract management scheduling features, construct a data management coefficient, and generate a stable execution or policy adjustment signal, thereby realizing policy closed-loop control and resource dynamic scheduling driven by behavior, reducing the problem of policy execution mismatch caused by the lack of behavior perception in the distributed storage system, and significantly improving the accuracy of data management and the stability of system operation in the distributed system.
[0039] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0040] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.
[0041] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0042] In addition, in each embodiment of this application, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0043] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.
[0044] Finally: The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A data management method based on a distributed storage system, characterized in that: It includes the following steps: Collect and standardize the data access logs in the distributed storage system, construct an access behavior sequence sorted by time, and extract access mutation rate, access periodicity, and access span metrics to generate a fusion behavior feature vector; Based on the fusion behavior feature vector, perform behavior pattern clustering and label division, construct the mapping relationship between corresponding behavior categories and policy intentions, generate a policy structure body, and perform resource configuration control; According to the policy structure body, perform scheduling calculation on the node resource status, construct a cost function to select the optimal node, complete replica creation, compression configuration, and hierarchical deployment, establish a consistency path and record the scheduling result, generate replica running status information for policy evolution; Analyze the behavior path response process after policy deployment, obtain the management scheduling information in the behavior path compression and execution displacement trend and analyze it, and perform scheduling management according to different signals generated by the analysis result.
2. The data management method based on a distributed storage system according to claim 1, characterized in that: Collect and standardize the data access logs in the distributed storage system, construct an access behavior sequence sorted by time. The specific steps include: Collect all access events of data objects in the data access logs. Each access event record includes access time, initiating node, operation type, and data volume; Perform standardization operations on the collected data, including unifying the timestamp accuracy to the second-level UTC format, normalizing the operation type to a unified identifier, filtering non-business accesses, and merging duplicate access requests within a set time interval; Form an event sequence sorted by time from the access logs of the standardized data objects, and convert the preprocessed event sequence into a structured access behavior sequence.
3. The data management method based on a distributed storage system according to claim 2, wherein: Extract access mutation rate, access periodicity, and access span metrics to generate a fusion behavior feature vector. The specific steps include: Take each valid access event in the event sequence as a behavior node of the data object and form a continuous behavior trajectory; Extract the access rate fluctuation feature, calculate the difference in access volume and the difference in time interval between adjacent accesses, judge whether there is a rate mutation in the access behavior. If there is a rate mutation, obtain and use it as the access mutation rate metric; Perform periodic feature extraction, perform frequency domain analysis on the access time series, and observe the activities in the time period as the access concentration metric; Perform access concentration and span extraction. Use the time interval between the first and last accesses as the behavior span, use the ratio of the front and back of the access concentration in the behavior span as the access concentration, and summarize the access concentration and span as the access span metric; Use the frequency domain analysis method to detect the periodic trend of the access behavior sequence, calculate the behavior fluctuation amplitude corresponding to each frequency, take the periodic component with the largest behavior fluctuation amplitude as the dominant period, and use a continuous time series composed of the change trend of the number of accesses per unit time as the access density change trajectory; Perform analysis on the context label information bound to the data object. The dimensions of the context label information include data source module, tenant or project code, data function role label, whether it is collaborative access data, and task priority. Perform context vector representation according to the context embedding function, and project the context tags to adjacent positions in the Euclidean space according to the semantic structure; After completing the access behavior modeling, generate a fused behavior feature vector.
4. The data management method based on a distributed storage system according to claim 3, wherein: Based on the fused behavior feature vector, perform behavior pattern clustering and label partitioning, construct the mapping relationship between the corresponding behavior categories and policy intentions, and generate a policy structure body for resource allocation control. The specific steps include: Use the logarithmic function to regularize the numerical values of different feature dimensions, and use the clustering method optimized based on density gradient to divide the feature space; Define behavior categories corresponding to different access patterns and business semantics. The behavior categories specifically include: periodic access type, sudden hot spot type, long-term cold storage type, collaborative sharing type, and unstable type; Assign a unique behavior label to each data object; Map the behavior label to a set of policy parameters. The policy structure body includes fields: number of replicas, storage heat level identifier, whether to enable compression, replica deployment area, and consistency level requirement; Perform different rule mappings according to different behavior labels; For the periodic access type, set the periodic compression and delayed replica refresh policies; For the sudden hot spot type, set the main replica residency and enable the cache layer snapshot; For the long-term cold storage type, perform compression, replica redundancy, and archival storage processing; For the collaborative sharing type, set the replica distribution and enable request scheduling; For the unstable type, trigger periodic behavior recognition.
5. The data management method based on a distributed storage system according to claim 4, characterized in that: Perform scheduling calculations on the node resource status according to the policy structure body, construct a cost function for optimal node selection, complete replica creation, compression configuration, and hierarchical deployment, establish a consistency path and record the scheduling results, and generate replica running status information for policy evolution, including the following steps: Construct a set of dynamic state parameters for each storage node to form a resource status vector. The resource status vector includes CPU occupancy rate, disk space usage rate, network bandwidth occupancy ratio, real-time packet loss rate, and node error history; On the basis of meeting the policy constraints, perform replica deployment, construct a cost function, which measures the cost of each storage node as a candidate replica deployment point, and select storage nodes based on the cost function; Sort the candidate storage nodes based on the cost function results, and select the node combination with the minimum total cost from them for replica deployment; Construct an ascending-order sorted candidate set according to the cost function, select nodes according to the ascending-order sorted candidate set and form a target deployment set; Send deployment instructions to each node in the target deployment set to execute the deployment operation: The deployment operation refers to the physical writing initialization of the data replica. If the compression policy is 1, enable the compressor, perform structural compression on the data replica, and according to the storage heat level identifier, store the replica in the corresponding level storage medium, and then establish a data synchronization path and confirmation mechanism between the master and slave replicas according to the consistency level requirement; After the deployment is completed, generate a scheduling task execution log and bind the task number and the behavior label.
6. The data management method based on a distributed storage system according to claim 5, characterized in that: Analyze the behavior path response process after the policy deployment, obtain the management scheduling information in the behavior path compression and execution displacement trend and analyze it, including the following steps: Acquire management scheduling information generated during the behavior path compression and execution displacement trend analysis process, the management scheduling information including the behavior path compression index and the strategy guidance offset displacement index; The behavior path compression index indicates whether the actual access path of the data object tends to be stable and convergent after the policy deployment; The policy-guided offset displacement index indicates the degree of topological offset between the target deployment area set by the policy structure and the actual replica scheduling and access behavior; The obtained behavior path compression index and strategy-guided offset displacement index are combined to generate a data management coefficient; The behavior path compression index is positively correlated with the data management coefficient, and the strategy-guided offset displacement index is negatively correlated with the data management coefficient.
7. The data management method based on a distributed storage system according to claim 6, wherein: Scheduling management is performed based on different signals generated by the analysis results, including the following steps: Compare the generated data management coefficient with the set data management threshold; If the data management coefficient is greater than or equal to the data management threshold, a stable execution signal is generated, and the current strategy structure and corresponding scheduling path are maintained, and the original scheduling plan continues to be executed; If the data management coefficient is less than the data management threshold, an evolution adjustment signal is generated, the scheduling plan is stopped, the behavior re-identification mechanism is triggered, the behavior labels are reallocated and the new strategy structure is mapped.
Citation Information
Patent Citations
Heterogeneous data resource management method oriented to computing power network
CN116594771A
SaaS-based cloud platform data storage method
CN118075293A
Distributed financial data management method and system under cloud infrastructure
CN118245887A
Customer portrait key data mining method and system based on space-time big data
CN118797542A
Distributed multi-tenant data security isolation system and method
CN119402233A
Cited By
Big data resource processing method based on cloud database service
CN120469843A
A big data resource processing method based on cloud database service
CN120469843B
Intelligent data storage method based on cloud platform
CN120872250A
A cloud-based intelligent data storage method
CN120872250B
Data center data management method based on artificial intelligence
CN121255476A