Distributed metadata management method and storage system based on cloud computing

By building a distributed metadata management method for cloud computing, the access frequency is collected, the popularity and resource distribution model is constructed, the migration cost is evaluated, and the optimal scheduling strategy is generated, which solves the problem of lack of global optimization and uncontrollable migration of replica scheduling, and efficient resource utilization and stability improvement are achieved.

CN120335723AInactive Publication Date: 2025-07-18王建福
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510451897.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, replica scheduling lacks global optimization capabilities, cannot dynamically adapt to access changes, and the cost of replica migration is uncontrollable, resulting in imbalance in replica hot matching, misalignment of deployment location and access path, frequent and uncontrollable migration.

Method used

A distributed metadata management method based on cloud computing is built, and the optimal scheduling strategy is generated by collecting access frequency, building access popularity distribution, evaluating cache resources and migration costs, and dynamic adjustment and optimization of cache strategy is achieved by using iterative solution.

Benefits of technology

It improves the replica scheduling accuracy and system stability, achieves high resource utilization and low data access latency, avoids the problems of hot and cold misalignment of replicas and high redundancy, and scheduling decisions can be quantified and optimized and interpreted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335723A_ABST
    Figure CN120335723A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cloud computing and distributed storage, and discloses a distributed metadata management method and storage system based on cloud computing, and the method comprises the following steps: S1, collecting the access frequency of a plurality of metadata objects in a set time period; s2, performing weighting processing on the access frequency, and constructing access heat distribution under multiple time scales; s3, obtaining the available cache capacity of a plurality of cache nodes, and constructing cache resource weight distribution; s4, assessing the access cost and migration cost between each metadata object and each cache node; and S5, generating a scheduling strategy between the metadata object and the cache node based on a preset optimization target, wherein the optimization target is access cost minimization. According to the method, a heat, resource and cost joint modeling mechanism is constructed, and optimal scheduling and feedback adaptive control are combined, so that accurate decision, dynamic adjustment and migration cost controllability of copy deployment are realized, and the performance and stability of distributed metadata management are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing and distributed storage, and specifically to a distributed metadata management method and storage system based on cloud computing. Background Art

[0002] In existing cloud computing and distributed systems, to improve the availability and access performance of metadata, metadata replicas are usually deployed among multiple cache nodes. Traditional replica scheduling methods are mostly based on strategies such as fixed replica numbers, average distribution, on-demand pulling, or static popularity configuration. Some systems also introduce a logic for adjusting the number of replicas based on access frequency to alleviate access hotspots. However, most of these strategies rely on preset rules and lack a comprehensive evaluation of system status, transmission cost, and node capabilities. Especially in scenarios of high concurrency, resource heterogeneity, or cross-regional deployment, the efficiency and flexibility of replica scheduling are easily restricted.

[0003] In this type of existing technology, the replica deployment process often adopts a one-time mapping method: generating strategies based on access statistical information, then deploying replicas, and no longer actively correcting; the access popularity model usually only depends on the number of accesses and is not jointly modeled with the cache resource status or scheduling cost; the scheduling process does not form an optimization objective function, and there is even less a policy evolution mechanism driven by a feedback link. Although this design method is simple to deploy, it is prone to problems such as unbalanced matching of replica popularity, misalignment between deployment locations and access paths, frequent and uncontrollable migrations, etc. In contrast, the present invention not only improves the accuracy of replica scheduling and system stability by constructing a unified scheduling model of "popularity-resource-cost", introducing a scheduling feedback mechanism, and evolution path control, but also effectively solves key defects in the prior art such as static deployment decisions, non-convergent scheduling, and uncontrollable migration costs. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a distributed metadata management method and storage system based on cloud computing, which solves the problems in the prior art that replica scheduling lacks global optimization ability, cannot dynamically adapt to access changes, and the replica migration cost is uncontrollable.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A distributed metadata management method and storage system based on cloud computing, including the following steps: S1: Collect the access frequencies of multiple metadata objects within a set time period; S2: Perform weighted processing on the access frequencies to construct an access popularity distribution under multiple time scales; S3: Obtain the available cache capacities of multiple cache nodes and construct a cache resource weight distribution; S4: Evaluate the access cost and migration cost between each metadata object and each cache node; S5: Generate a scheduling policy between the metadata object and the cache nodes based on a preset optimization goal, where the optimization goal is to minimize the access cost; S6: Construct a resource game model based on the non - cooperative behavior among cache nodes, and use an iterative solution method to obtain an optimal caching policy that satisfies the equilibrium condition; S7: Deploy metadata replicas on each cache node according to the optimal caching policy, and perform cache migration operations; S8: Repeat the above steps at a set time period to achieve dynamic adjustment and continuous optimization of the caching policy.

[0006] Preferably, the multi - time - scale access heat distribution is obtained by performing exponential decay weighted processing on the access frequency, and the access frequency weight is higher for time closer to the current time.

[0007] Preferably, the scheduling policy generation step is solved using a regularization optimization method with a smoothing term to improve computational stability.

[0008] Preferably, the cache replica deployment includes preferentially storing in low - latency nodes, secondarily storing in local SSDs, and then degrading to remote storage nodes.

[0009] Preferably, the cache policy update period is dynamically adjusted according to the system load.

[0010] A distributed metadata management and storage system based on cloud computing, comprising: An access collection module for obtaining the access frequency of metadata objects in real - time; A heat modeling module for constructing an access heat distribution based on a multi - scale time window; A cache resource modeling module for generating resource weights according to the capacity of cache nodes; A cost evaluation module for calculating the access cost and migration cost between an object and a cache node; A policy generation module for generating a resource scheduling policy that minimizes the access cost; An equilibrium solving module for constructing a multi - node caching game model and solving for the caching policy in the equilibrium state through iteration; A policy execution module for deploying metadata replicas according to the caching policy and completing migration operations; An update module for periodically triggering access heat reconstruction and policy update to support dynamic optimization.

[0011] Preferably, the equilibrium solving module includes a game modeler for constructing a caching game graph and a solving engine for iteratively calculating the Nash equilibrium.

[0012] Preferably, the policy generation module uses a neural network to assist in evaluating the impact of the scheduling scheme on the access cost and supports self-learning optimization of the policy.

[0013] Preferably, the cost evaluation module updates the delay information and network load of each node in real time for feedback adjustment of the cache deployment policy.

[0014] Preferably, all modules of the system run on a distributed cloud platform cluster in the form of containerized microservices, supporting dynamic scaling at the module level.

[0015] The present invention provides a distributed metadata management method and storage system based on cloud computing, having the following beneficial effects: 1. The present invention adopts a joint strategy of access heat modeling, cache resource modeling, and cost trade-off optimization to globally model the replica scheduling, achieving the technical effects of high resource utilization rate and low data access latency. Compared with the existing solutions that only make static allocations based on the number of replicas or access frequencies, it solves the structural problems of rigid policies, lagging responses, and inability to consider migration costs.

[0016] 2. The present invention adopts a periodic feedback mechanism, enabling the system to self-perceive deployment errors and access offsets, and automatically recycle old states and adjust the scheduling scheme. This is not complex to implement, but it brings significant benefits. Compared with the design in traditional solutions that no longer pay attention to the actual hit situation after deployment, it avoids the situation of misalignment between hot and cold replicas and high redundancy.

[0017] 3. By constructing a scheduling mapping matrix and cooperating with a weighted cost model, the present invention forms a unified objective function solving mechanism, achieving the technical effects of quantifiable scheduling decisions and interpretable optimization. Compared with the existing methods that rely on manual rules and cannot form a stable convergence path, it breaks through the limitations of large randomness in replica distribution and high operation and maintenance difficulty.

[0018] 4. The present invention introduces evolutionary mechanisms such as scheduling round control, event-triggered scheduling, and scheduling freezing, making the entire replica management process not a one-time calculation but a process that can grow, be adjusted, and have memory. This cannot be achieved in the previous methods where deployment is terminated immediately, solving the core shortcoming of being unable to dynamically adapt to access changes. Brief Description of the Drawings

[0019] Figure 1 It is a flowchart of the method steps of the present invention. Detailed Embodiments

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0021] Please refer to the attached Figure 1 , the embodiment of the present invention provides a distributed metadata management method based on cloud computing, including the following steps: S1: Collect the access frequencies of multiple metadata objects within a set time period; In the distributed metadata management method based on cloud computing of the present invention, the entire technical solution takes the metadata access requests on the cloud platform as the core input for scheduling optimization.

[0022] Generally, to ensure the rationality of the subsequent cache scheduling strategy, it is necessary to perform accurate and dynamic data collection on the actual access behavior. Step S1 is the basic input step in the whole process. Its goal is to obtain the access frequency information of users for metadata objects and perform time serialization processing to provide raw data support for constructing the access heat model.

[0023] As a prerequisite, this step is usually implemented through a lightweight proxy module deployed on edge nodes or virtual machines in a distributed cloud storage platform. This module can monitor the event level of metadata layer access by listening to system calls, API interface requests, or file handle behaviors.

[0024] In this embodiment, the access collection module is deployed in container nodes under the Kubernetes environment. Combining the monitoring interfaces provided by the cloud platform (such as K8sSidecar, DaemonSet), it can non-invasively capture metadata operations on the distributed file system, including but not limited to: Creation of files or directories (Create), metadata reading (GetAttr, Stat), metadata modification (SetAttr, Chmod), object deletion (Unlink, Rmdir), path access behavior (Open, Lookup).

[0025] Generally, the system uniformly abstracts the above access events into "metadata access requests". Each request is recorded as a triple with an event timestamp, an object identifier (ObjectID), and the source IP. All requests are written into the Kafka message stream by the collector to ensure asynchronous high-throughput processing capabilities.

[0026] In a possible implementation, the access behavior statistics are carried out in the way of "sliding time window". That is, at any time point t, the system looks back at the access records in the recent period (set as seconds) to calculate the access frequency of each metadata object.

[0027] Specifically, for the metadata object set , the weighted access frequency of the object within the time t is defined as: ; where: represents the weighted access frequency of the object within the current time window; is the time window length (for example, 300 seconds); represents the original access count of the object at the time point ; is the time weight function, which is used to express the time decay weight; In some embodiments, can be selected in the form of exponential decay, such as: ; where is the exponential decay coefficient, which controls the weight influence degree of the recent access behavior.

[0028] As an option, to enhance the system's ability to recognize multiple access patterns, the system can maintain multiple different value sliding windows in parallel. For example, the access frequencies at the scales of 1 minute, 5 minutes, and 30 minutes are retained simultaneously to form a multi-scale access frequency vector.

[0029] In some implementations, to reduce the computational and storage overheads, the access records of each object do not have to be accurate to every second, but a time bucket aggregation mechanism is adopted. For example, the time line is divided into several non-overlapping time segments, and the access frequency and time decay value are aggregated once for each time slice. The aggregation result can be used to approximately restore the continuous time decay weight model.

[0030] In addition, for engineering practicality considerations, in some embodiments, a circular buffer or a counting hash table is used to maintain the access frequency data, minimizing the space overhead to the greatest extent and supporting the rapid clearing of expired records.

[0031] To ensure the cross-node consistency of the access behavior data, the data synchronized between multiple collection nodes will be marked with a unified time stamp. The system can rely on the service time provided by the cloud platform (such as the Kubernetes NTP service) for unified time benchmark alignment to ensure that the event alignment accuracy is better than the millisecond level.

[0032] In some cases, to prevent malicious access behaviors or transient bursts from affecting scheduling judgments, the access frequency data is also preprocessed through methods such as percentile peak clipping, data smoothing, or moving average filtering.

[0033] In summary, step S1 is a key input means for obtaining the dynamic state of the system in the technical solution of the present invention, determining the initial boundary conditions of the subsequent scheduling model. Its implementation method has stability and low interference, can operate in a high-concurrency environment for a long time, and supports high-frequency real-time access statistics requirements.

[0034] S2: Perform weighted processing on the access frequency to construct the access heat distribution under multiple time scales; After completing the continuous collection of metadata access requests in step S1, the system uses these original access frequency data as input and enters the subsequent modeling stage. Step S2 is mainly responsible for constructing an access heat distribution model under multiple time scales based on the time-series data. This heat model, as one of the inputs for scheduling optimization, determines the evaluation method of the "activity level" of each metadata object in different time periods. The abstraction and transformation of access characteristics in this step directly affect the effectiveness and convergence stability of the scheduling results. Based on the dynamic evolution form of heat probability, it is beneficial to depict the local aggregation and time decay characteristics of metadata access. Its modeling strategy needs to balance real-time performance, expressive ability, and resource constraints.

[0035] In this embodiment, the system probabilistically processes the access behavior of each metadata object based on the access frequency sequence obtained in step S1. of the access behavior.

[0036] Generally, to avoid the excessive lag effect of historical access records on the current state, a time decay weight is introduced in the calculation of access heat. This weight depends on the access frequency sequence of the object within the time window and is calculated through a continuous time function.

[0037] Specifically, let represent the current time, be the selected sliding window length, in seconds, minutes, or other time units. The object at time The original access frequency within is , which has been statistically obtained in S1.

[0038] In a possible implementation, the weighted access frequency of the object is defined as: ; where, represents the decay frequency value of the object within the current window; is a time decay function that satisfies and .

[0039] As an option, an exponential function can be adopted: ; wherein is a time decay control parameter. The larger the value, the more the system tends to focus on recent access records.

[0040] In some embodiments, to control the computational complexity, time s can be defined at discrete time steps, for example, discretely divided in units of seconds, 5 seconds, or 10 seconds.

[0041] The above can form a set of probability densities after normalization: ; wherein represents the unit probability heat value of object and satisfies .

[0042] As a normalized form, this probability distribution reflects the relative activity levels among all objects at the current moment.

[0043] In some embodiments, the system can concurrently maintain heat distributions at multiple different time scales. Let represent the window lengths of multiple time scales, and calculate respectively: ; This structure allows the system to simultaneously perceive short-term burst hotspots and long-term high-frequency accessed objects, thereby supporting the scheduling policy to react within different time ranges.

[0044] In one possible implementation, the multi-scale heat can further form a vector form: ; wherein is the heat at one time scale corresponding to each dimension.

[0045] Generally, this heat vector is input to the subsequent scheduling module for jointly considering the cache allocation probability and the replica deployment strategy.

[0046] In some other scenarios, the system can also perform dimensionality reduction on the original heat vector. For example, extract the main evolution direction through principal component analysis (PCA), or obtain a unified heat score through a weighted average form.

[0047] In addition, the system supports dynamically adjusting the value of τ during the actual deployment process, and performs adaptive parameter adjustment according to operating conditions such as the current access volume, system load, and time period.

[0048] During specific engineering deployment, to improve the real-time performance of the popularity model update, the system can adopt a sliding window cache structure or an incremental update method, without repeatedly scanning the full access log, thereby improving the system response efficiency.

[0049] To ensure the robustness of the probability model in different access density environments, during the popularity normalization process, a minimum truncation, noise suppression, or outlier filtering mechanism can also be introduced.

[0050] Through the above construction method, this step successfully transforms the original discrete access events into a continuous, differentiable, and adjustable popularity probability expression form, providing a mathematical input basis for subsequent scheduling solutions.

[0051] S3: Obtain the available cache capacities of multiple cache nodes and construct a cache resource weight distribution; After completing the acquisition of metadata access behavior (step S1) and the construction of the access popularity probability distribution (step S2), the system needs to further establish a description of the availability of cache resources to support the subsequent solution of the resource scheduling strategy. The task of step S3 is to extract the available cache resource information from the underlying cache nodes and transform it into a structured numerical weight form that can be processed by the optimization model. The cache resource modeling and the access popularity modeling are generated in parallel, and both are used as the input end components for constructing the resource scheduling mapping.

[0052] In this embodiment, the cache node refers to a set of logical nodes deployed in the cloud computing platform that can be used to store metadata copies. This set is denoted as , where is the total number of nodes. Each cache node usually corresponds to a physical container, a virtual computing node, an edge gateway device, or other addressable cache resource entities.

[0053] Generally, each cache node has an independent cache capacity limit. This capacity is denoted as , indicating the upper limit of the allocable cache resources owned by node . The unit can be GB, MB, or the number of blocks, depending on the deployment architecture of the cache system. For convenient unified processing, the system normalizes all cache capacities to obtain a resource distribution weight vector: ; Among them, represents the resource probability weight allocated to node , satisfying: ; The normalized vector ν = [v1, v2, …, vk] is a probabilistic modeling representation of the cache resources in the system, which semantically aligns with the access popularity distribution μ generated in step S2.

[0054] In a possible implementation, the system also dynamically monitors the node resource status, including CPU occupancy, memory usage, current cache write pressure, etc. Based on the monitoring results, the system can make real-time adjustments. For example, when a network anomaly occurs at a certain node or the IO waiting time exceeds the set threshold, the system will discount its cache capacity:

[0055] where represents the dynamic weight adjustment factor, which is automatically calculated by the system through a fault detection mechanism or a load prediction model. The updated will replace the original capacity to participate in the normalization process to ensure that the resource modeling fits the system operation status.

[0056] Specifically, the cache resource modeling not only includes the static expression of the storage capacity, but also should reflect the sensitivity to the distributed network structure. In some deployment scenarios, factors such as the physical distance, network latency, and transmission rate between the node and the request source will indirectly affect the actual cache availability. Therefore, the system can introduce a geographical awareness factor , which represents the node proximity to the main access path.

[0057] As an option, the final cache node weight can be calculated by fusing the capacity, dynamic state, and topology factors: ; In some embodiments, can be automatically generated through network ranging, access route analysis, or the available zone weights returned by the cloud platform API, supporting policy balancing in cross-available zone or cross-region deployment environments.

[0058] In addition, to improve the frequency and efficiency of modeling updates, the system can control the resource status sampling period through an asynchronous task or a timer mechanism, and the sampling results are sent to the modeling engine through a message middleware. The modeling engine maintains a node status cache table and loads the latest status data to participate in the normalization calculation before each round of scheduling optimization.

[0059] Some implementations also support a threshold trigger mechanism, that is, when the change amplitude of a certain node status exceeds a certain proportion, the resource modeling result is immediately refreshed to improve the scheduling accuracy.

[0060] In summary, this step constructs a structured expression form of cache resource distribution with multi-dimensional inputs, which serves as the probability input at the resource end corresponding to the access popularity, laying a numerical foundation for the transmission mapping calculation of subsequent scheduling strategies. It also has high adaptability in cross-node, cross-region, or heterogeneous cache environments. The model update mechanism is flexible and capable of responding to performance fluctuations and network bursts.

[0061] S4: Evaluate the access cost and migration cost between each metadata object and each cache node; After completing the access popularity modeling (step S2) and cache resource modeling (step S3), the system still lacks a constraint index for quantifying the resource allocation cost. The popularity distribution reflects the access intention, the resource modeling expresses the ability boundary, and step S4 fills the evaluation scale between the two. This step constructs a transmission cost model to provide a computable objective function basis for scheduling optimization. The combination of access cost and migration cost determines the deployment tendency of metadata replicas on different cache nodes.

[0062] In this embodiment, the system constructs a cost quantification index for each pair of metadata objects and cache nodes to form a cost matrix .

[0063] Generally, this cost includes two parts: one is the access cost, and the other is the migration cost. Since they cannot be directly added numerically, standardization processing and weighted combination are required.

[0064] Specifically, the access cost represents the average delay generated by transmitting the object from node to the requesting user, usually measured in milliseconds or seconds. This value can be monitored in real time by the network RTT measurement module and historical samples can be recorded.

[0065] As an option, the system can use the exponential moving average method to smooth the delay to reduce the impact of instantaneous jitter: ; where ρ is the smoothing factor, is the currently measured delay value.

[0066] In a possible implementation, the system can also correct the delay value by combining the packet loss rate or path fluctuation factor to improve the robustness of the evaluation.

[0067] The migration cost represents moving the metadata object from its current location to the cache node Required bandwidth cost. This cost can be expressed as the proportion of bandwidth consumed by data transmission per unit time.

[0068] Generally, this cost depends on the object size , link bandwidth and network congestion status. The calculation method can be expressed as: ; where is the data size of the object ; is the link bandwidth from the original position of the object to the node ; is the network congestion adjustment factor, reflecting network stability and queuing time.

[0069] In some embodiments, can be obtained by estimating the length of the transmission queue in real time.

[0070] In actual use, to unify the two cost dimensions, the system combines them with weights to obtain the total cost: ; where represents the total deployment cost of the metadata object at the node ; is the access cost weight; is the migration cost weight; + 0, not a fixed ratio, and the system can adjust dynamically according to the load.

[0071] As a strategy, the system automatically increases the value during network congestion periods to reduce the frequency of replica migration. When the access pressure is low, it increases the weight and gives priority to optimizing response latency.

[0072] In some embodiments, the cost matrix C is cached in memory and triggered for update every fixed period or when the network state changes drastically. If the system detects abnormal packet loss rates or link failures in some nodes, it will set the corresponding to infinity to prevent the scheduler from assigning tasks to unreachable nodes.

[0073] In addition, for performance considerations under high concurrency, the system can pre-compute some static terms such as to reduce the computational load in each round of optimization iteration.

[0074] Finally, the cost matrix C will be passed as input to the scheduling policy generation module and participate in the calculation and solution of the objective function together with the access heat distribution μ in step S2 and the resource distribution ν in step S3.

[0075] This step establishes a unified standard for measuring the "cost" of replica allocation and is a key component in constructing the scheduling optimization problem. It has a flexible structure, high scalability, and can adapt to network topologies and operating conditions in various deployment environments.

[0076] S5: Generate a scheduling policy between metadata objects and cache nodes based on a preset optimization goal, where the optimization goal is to minimize the access cost; After consecutive modeling processes in steps S2 to S4, the system has respectively obtained the access heat distribution, cache resource distribution, and deployment cost matrix. These three sets of data constitute the core input of the scheduling solution problem. The role of step S5 is to formalize the scheduling task as an optimization problem and solve it to finally obtain the deployment mapping of metadata replicas. This step follows the previous model but does not repeat the expression. Instead, it introduces a mathematical optimization framework to complete the mapping decision between resources and access intentions.

[0077] In this embodiment, the system sets the mapping decision variable as to represent whether the metadata object is deployed on the cache node . If , it means the replica is deployed on this node; otherwise, no replica configuration is performed.

[0078] Generally, this set of variables constitutes a mapping matrix , where is the number of metadata objects, and is the number of cache nodes. The system goal is to solve the optimal to minimize the overall deployment cost.

[0079] Specifically, the scheduling objective function can be expressed as: ; where is the probability weight of the object in the access heat distribution; is the deployment cost matrix obtained from step S4; is the scheduling variable representing the deployment choice.

[0080] Under the above goal, the system needs to introduce a set of constraints to ensure the rationality of the deployment plan and the controllability of the system boundary.

[0081] As an option, the capacity constraint is used to limit the replica load of each node not to exceed its cache capacity. The constraint form is as follows: ; wherein, represents the data size of the object ; represents the available capacity of the cache node (see step S3).

[0082] In some embodiments, to prevent replica islands, the minimum replica redundancy of each object may also be set, specifically: ; wherein, represents the minimum number of replicas required for the object , usually obtained from system policy configuration.

[0083] In a possible implementation, the system introduces a boolean tensor , records the historical deployment status, and uses it as a reference item for variable X, and adds a historical migration penalty term in the scheduling optimization: ; wherein, is the migration penalty factor; represents whether the deployment of the object replica has changed.

[0084] This term is used to constrain the replica change frequency and reduce the instability of the system.

[0085] In some implementations, the scheduler uses the integer linear programming (ILP) method to solve the above optimization problem. If the problem scale is large, the system can switch to a heuristic algorithm, such as the greedy deployment method, the distributed random sampling strategy, or the approximation algorithm based on convex relaxation.

[0086] As an alternative strategy, the system can also perform calculations based on the distributed graph optimization framework, specifically including classifying nodes for scheduling variables to improve the convergence efficiency in a parallel computing manner.

[0087] In high-concurrency scenarios, the scheduler supports the incremental optimization mechanism. That is, based on the historical deployment solution , only optimize the set of objects whose access heat has changed recently to reduce the overall solution pressure.

[0088] The finally output is the current optimal replica scheduling mapping relationship. This structure will be used in the actual replica deployment system, and the background replica management module will complete replica synchronization and cache configuration.

[0089] Through this step, the system realizes the overall scheduling strategy inference starting from probability heat and resource constraints. This process is abstracted as solving a constrained optimization problem, with a flexible boundary control mechanism and dynamic update ability. Its mathematical form can be directly migrated to other heterogeneous cache systems, with good generality and engineering feasibility.

[0090] S6: Construct a resource game model based on the non-cooperative behavior among cache nodes, and use an iterative solution method to obtain the optimal cache strategy that satisfies the equilibrium condition; After completing the solution of the optimal replica scheduling mapping (step S5), the system still needs to implement the theoretical mapping result to achieve the actual update of cache content and replica consistency management. Step S6 exactly undertakes this task. Its core lies in converting the scheduling matrix into actual replica operation instructions in the cache system and implementing scheduling through the network layer. The previously solved scheduling variable matrix has direct execution semantics in this step.

[0091] In this embodiment, the replica deployment module first parses the optimal scheduling solution ={ }, and makes a difference comparison with the current replica state matrix ={ } to obtain the operation set that needs to be executed: ; Among them, if , it means that a new replica needs to be added, and the object is sent to the node ; if , it means that a replica needs to be deleted, and the object is removed from the node ; if , it means that the deployment state of this object remains unchanged.

[0092] Generally, the replica deployment operation is divided into two stages: transmission and local disk writing. The system uses an asynchronous controller to issue scheduling instructions, and the distributed file system client completes replica transmission at the target node.

[0093] Specifically, the replica transmission path selects a node from the current replica source node set as the source node and initiates data migration to the target node . The system selects the optimal source node according to the following criteria: ; Among them, represents the link transmission cost between node and node , which can be measured based on the bandwidth-delay ratio or congestion window size.

[0094] As an option, the transmission scheduling module supports a multi-source concurrent replication strategy. If the object supports the sharding mechanism, the replicas can be partitioned according to the shard set and transmitted to the target node concurrently from multiple nodes . This method can improve the transmission efficiency of large objects.

[0095] In a possible implementation, to prevent the system resources from being occupied by large-scale migration operations, the system sets an upper limit on the maximum number of replicated copies per round . When , only the most urgent or cost-optimal subset is selected for execution, and the remaining tasks are deferred to the next scheduling cycle.

[0096] The replica disk write adopts the Copy-on-Write mechanism. After the data is received in the local buffer, it is asynchronously written to the persistent storage device. After the write is completed, the local metadata directory of the node is updated, and the central scheduler is notified to update the T matrix.

[0097] In some embodiments, the system maintains a version number identifier during the replica change process and performs consistency verification between replicas. This mechanism is applicable to cache architectures that support write-many read-many scenarios to ensure the validity of replicas.

[0098] As a synchronization strategy, the system introduces a lightweight consistency protocol, such as the Last-Write-Wins policy based on timestamps. Each node records the most recent write time of the object , and only performs an overwrite write operation when .

[0099] In some deployment modes, the replica deployment operation is forwarded and completed through an edge controller or a cloud proxy. This mechanism can be applied to heterogeneous environments such as device-level caching, edge-side storage, or cross-cloud deployment. The system provides a unified interface encapsulation to shield the underlying execution differences.

[0100] After the deployment is completed, the scheduler performs a global consistency verification once, and the verification targets are: ; If there are deviations, the compensation operation is automatically triggered, and the above-mentioned transmission and disk write processes are repeated until convergence to the target state. The verification process can be completed by comparing hash values, object version numbers, or file metadata.

[0101] Finally, the cache system completes the actual update of the replica deployment plan, and the mapping optimization results take effect at the data layer, closing the resource scheduling loop. Step S6 has the ability to interact with the resource allocation and solution layer, and can iteratively enhance the deployment stability and replica reachability in multiple rounds of scheduling.

[0102] S7: Deploy metadata replicas on each cache node according to the optimal cache policy and perform cache migration operations; After the replica deployment is completed (step S6), the scheduling loop of the system has not been formed. The deployment results need to be transmitted back to the upper-layer scheduling logic to achieve adaptive correction and dynamic feedback of future scheduling policies. Step S7 is the key to constructing this feedback chain. It focuses on the sampling of deployment results, status updates, and the backfilling of the scheduler status, enabling the entire replica management mechanism to have the characteristics of evolutionary ability and historical memory.

[0103] In this embodiment, the system introduces a status feedback module to regularly collect deployment status snapshots from each cache node. The snapshot data includes but is not limited to the following: The current replica existence matrix , and compare it with the scheduling solution ; Object availability identifier, indicating whether the replica is readable locally on the node; Node load information, such as the current I / O queue depth, CPU utilization, etc.; Replica access logs, including the local hit count of each object on each node .

[0104] Generally, the system reports the above information to the scheduling center through the message bus or the embedded heartbeat protocol. The feedback period can be configured, and the default is aligned with the scheduling period to maintain the timing consistency between the scheduling status and the actual system.

[0105] Specifically, the feedback data is used to update the following model components: First, update the actual replica status matrix , for differential deployment decisions in subsequent rounds of scheduling; second, according to the access logs , correct the access heat distribution . The following sliding update strategy is adopted: ; Among them, is the update step size, represents the number of accesses of object on node .

[0106] As an option, the system supports a heat adjustment mechanism based on local hit rate. If an object is frequently hit on multiple nodes, the system can automatically increase its redundancy threshold , and retain more copies in the next round of deployment.

[0107] In a possible implementation, the system also incorporates the migration frequency into the feedback loop. For each object, record its number of replica migrations within the time window . . If exceeds the preset threshold , the system marks it as a "migration-sensitive object" and imposes a punitive treatment on its deployment in subsequent scheduling.

[0108] The form of punishment can be adjusted as follows: ; where γ is the punishment factor, is the updated cost item. This strategy can effectively suppress the system instability caused by high-frequency migrations.

[0109] In some embodiments, the feedback module also incorporates an asynchronous exception capture mechanism. When abnormal situations such as deployment failure, target node unreachability, or replica verification inconsistency are detected, the system immediately marks the corresponding entry as a "failed replica" and triggers an emergency correction process through the scheduler.

[0110] As a supplementary mechanism, the system sets an object active status indicator , calculates the object activity based on the access frequency per unit time, and maintains a hierarchical priority queue for replica updates according to : ; where represents the statistical window length. High-activity objects will be given priority in terms of scheduling accuracy and replica consistency during the scheduling process.

[0111] In some deployment architectures, the feedback mechanism works in coordination with the replica consistency protocol. After deploying or deleting a replica, the system waits for a consistency confirmation signal. If the confirmation fails, it automatically marks the feedback status as "not effective" and delays the update of the scheduling matrix.

[0112] Finally, step S7 completes the recovery of multi-dimensional data such as replica deployment status, access behavior, and system load, as well as model update, providing data support for the next round of scheduling. The feedback mechanism runs through the entire replica life cycle, forming an internal loop path for adaptive replica optimization, enhancing the system's controllability and policy closed-loop ability.

[0113] S8: Repeat the above steps at a set time interval to achieve dynamic adjustment and continuous optimization of the caching policy; After the system completes the replica status feedback and heat correction (step S7), without a periodic restart mechanism or human intervention, the scheduling process will tend to be static and the policy evolution ability will be limited. Step S8 aims to break this local steady state and achieve continuous evolution of the policy and self-update of the structure through scheduling round management, scheduling trigger design, and model warm start control. This step serves as both the end point of the process and the starting node of the next round of scheduling, constructing a cyclic decision-making path.

[0114] In this embodiment, the system introduces a scheduling period controller to define a global scheduling period T, which is used to drive the main loop of the scheduling execution. Every T unit of time, the scheduler restarts and enters from step S1 again.

[0115] Generally, the scheduling period T depends on dynamic parameters such as the system access intensity, object change frequency, and resource change rate. In some embodiments, the scheduling period can be dynamically adjusted by a prediction model. For example, an exponential regression model is used to predict future access fluctuations: ; where represents the current access intensity, represents the predicted value for the next period, and is the smoothing parameter.

[0116] As an option, the system supports an asynchronous trigger mechanism. When specific conditions are met (such as a decrease in cache hit rate, a sudden change in load index, or an abnormal replica loss rate), the current period can be terminated early and a new round of scheduling can be forced to enter.

[0117] In a possible implementation, an event-driven model is introduced. Define a trigger event set , when any is detected as the "activated" state, the entire process from step S1 to S8 is immediately triggered.

[0118] In some embodiments, the scheduling state is stored as a snapshot , including the current heat distribution , resource distribution , cost matrix , replica mapping and feedback status information. Before each restart of the scheduling, the system can choose to warm start from the previous snapshot or cold start from the initial state again.

[0119] Specifically, if warm start is selected, the scheduling solver uses as the initial solution input and executes an incremental optimization algorithm, such as greedy insertion, local perturbation, or gradient correction strategy based on simulated annealing. This mechanism can reduce the solution complexity and avoid repeated calculations.

[0120] In certain deployment scenarios, a round numbering mechanism is supported. Define the scheduling round number , and each time a complete process is executed , the scheduling result is written into the meta-scheduling record set , supporting historical backtracking and trend analysis.

[0121] As an extended module for policy evolution, the system also supports the "cooling period" design. That is, introduce the shortest interval between two consecutive deployments , to prevent replica oscillations caused by excessive scheduling frequency. If the current system state change is not sufficient to break through , then skip the scheduling cycle and only perform data collection and status observation.

[0122] In some implementation details, the system also designs a scheduling stability index , indicating the degree of difference between two adjacent rounds of deployment: ; If continues to decline and is lower than the stability threshold δ, the system enters the scheduling freeze mode and only maintains basic replica monitoring and feedback sampling.

[0123] Finally, step S8 constructs the boundary and restart mechanism of the scheduling life cycle. The scheduling main loop can continue to run and evolve dynamically. The system seeks a balance between stability and perturbation, realizes elastic switching of the scheduling path between cycles and events, and ensures the robustness and adaptability of replica management during long-term operation.

[0124] A distributed metadata management storage system based on cloud computing includes: An access collection module for obtaining the access frequency of metadata objects in real time; This module is responsible for real-time monitoring of access requests for metadata objects and collecting the following key data: Access frequency : The number of accesses to an object within a unit of time ; Access source : Record the access contributions of different user groups or business applications; Access time distribution : Build a time series prediction model to discover periodic access patterns; Access operation type (read / write), used to distinguish metadata consistency requirements A heat modeling module for constructing an access heat distribution based on a multi-scale time window; This module is used to quantify the access heat of metadata objects and dynamically adjust the replica distribution in combination with time decay modeling. The heat modeling uses an exponential decay model, and the formula is as follows: ; Among them: is the object at time of the comprehensive access popularity; represents the access frequency in the past time period; is the weight of different time windows, ensuring both short-term sudden increase and long-term popularity; is the decay factor, controlling the influence weight of historical access.

[0125] This module can also optimize the popularity modeling by combining machine learning prediction methods (such as LSTM, ARIMA) to make the scheduling decision more accurate.

[0126] Cache resource modeling module, used to generate resource weights according to the capacity of cache nodes; This module is used to describe the available resource status of cache nodes, including storage capacity, computing power, current load, etc. The cache node resource weight can be expressed as: ; Among them: is the node available storage capacity; is the computing resource (such as CPU, memory); is the current load, taking the reciprocal means that the higher the load, the lower the weight; , , , are weight parameters, used to balance the influence of different resources.

[0127] This module ensures that replicas are preferentially deployed on cache nodes with sufficient resources and low load, avoiding unbalanced replica distribution.

[0128] Cost evaluation module, used to calculate the access cost and migration cost between the object and the cache node; This module is mainly used to evaluate the cost of metadata access and the overhead brought by replica migration. The access cost mainly includes network transmission delay, storage read overhead, and the impact of replica hit rate, etc. The migration cost involves data transmission bandwidth consumption, load fluctuation risk, and resource occupancy of the target node.

[0129] During the replica scheduling process, the system needs to balance access efficiency and migration cost. Excessive and frequent replica migrations may lead to waste of system resources, while unreasonable replica storage locations may increase access latency. Therefore, the core task of this module is to find the optimal replica scheduling strategy through comprehensive analysis of access cost and migration cost, maximize access efficiency, and avoid unnecessary resource consumption.

[0130] A policy generation module, which is used to generate a resource scheduling policy that minimizes the access cost; Based on access popularity, cache resource status, and cost evaluation results, this module automatically generates a replica scheduling plan. The core goal is to optimize replica distribution to minimize access overhead and maximize system performance.

[0131] The policy generation module usually adopts heuristic algorithms or optimization algorithms, taking multiple factors into comprehensive consideration: The number of replicas of hot data: For data with high-frequency access, the system will increase the number of replicas to reduce access pressure; Replica storage location: Select the node closest to the user with the most abundant resources to ensure the lowest access latency; Migration strategy: If the resources of a certain node are tense or the replica hit rate drops, the system will actively adjust the replica location to adapt to changes in the access pattern.

[0132] In addition, this module supports the adaptive optimization of policies. When the access pattern changes significantly, the system can dynamically adjust policy parameters, such as increasing the number of replicas of certain key data, or reducing the storage redundancy of cold data, to improve overall performance.

[0133] An equilibrium solving module, which is used to construct a multi-node cache game model and solve the cache policy in the equilibrium state through an iterative method; The goal of this module is to optimize the replica scheduling process, balance the load among cache nodes, and ensure optimal access performance. Since cache resources are limited, replica scheduling needs to consider the competition relationship among multiple nodes.

[0134] In this module, the system will simulate the behavior of cache nodes and find the best replica storage plan through an iterative optimization method. Specifically, this module will: Evaluate the current load of each cache node to determine whether there is a situation of tight storage resources; Calculate the resource competition among nodes to ensure that the replica distribution will not be too concentrated on certain nodes; Optimize the replica distribution to minimize the access latency of the overall system and avoid overloading of individual nodes.

[0135] In addition, this module can also gradually optimize the replica scheduling scheme through an adaptive adjustment mechanism, enabling it to continuously adjust as the access pattern changes and ultimately achieve a stable load balancing state.

[0136] A policy execution module for deploying metadata replicas according to a caching policy and completing migration operations; This module is responsible for executing the caching policy, performing replica deployment and migration, including: Dynamic replica deployment: storing replicas on target cache nodes according to the optimal policy; Migration operation: ensuring data consistency and using parallel migration to reduce replica update latency; Caching invalidation handling: when the access popularity of replicas decreases or cache resources are strained, executing replica recycling or downgraded storage policies An update module for periodically triggering the reconstruction of access popularity and policy updates to support dynamic optimization; This module is used to periodically trigger the reconstruction of access popularity and policy updates to ensure that replica scheduling adapts to changes in the access pattern. The update policies include: Timed update: recalculating the replica deployment policy every time interval; Event trigger: immediately adjusting the replica distribution when the access pattern mutates or the cache node load exceeds the limit; Learning optimization: combining reinforcement learning methods to gradually converge the replica scheduling policy to the global optimum.

[0137] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A distributed metadata management method based on cloud computing, characterized in that It includes the following steps: S1: Collect the access frequencies of multiple metadata objects within a set time period; S2: Perform weighted processing on the access frequencies to construct an access heat distribution under multiple time scales; S3: Obtain the available cache capacities of multiple cache nodes and construct a cache resource weight distribution; S4: Evaluate the access cost and migration cost between each metadata object and each cache node; S5: Generate a scheduling policy between the metadata object and the cache node based on a preset optimization goal, and the optimization goal is to minimize the access cost; S6: Construct a resource game model based on the non-cooperative behavior between each cache node, and use an iterative solution method to obtain an optimal cache policy that meets the equilibrium condition; S7: Deploy metadata replicas on each cache node according to the optimal cache policy and perform cache migration operations; S8: Repeat the above steps at a set time period to achieve dynamic adjustment and continuous optimization of the cache policy.

2. The distributed metadata management method based on cloud computing according to claim 1, wherein The multi-time scale access heat distribution is obtained by performing exponential decay weighted processing on the access frequencies, and the access frequency weight closer to the current time is higher.

3. The distributed metadata management method based on cloud computing according to claim 1, characterized in that, The scheduling policy generation step is solved by using a regularization optimization method with a smoothing term to improve computational stability.

4. The distributed metadata management method based on cloud computing according to claim 1, wherein The cache replica deployment includes preferentially storing in low-latency nodes, secondarily storing in local SSDs, and then degrading to remote storage nodes.

5. The distributed metadata management method based on cloud computing according to claim 1, characterized in that, The cache policy update period is dynamically adjusted according to the system load.

6. A distributed metadata management storage system based on cloud computing, for use in the distributed metadata management method based on cloud computing according to claims 1-5, characterized in that, It includes: An access collection module for obtaining the access frequencies of metadata objects in real time; A heat modeling module for constructing an access heat distribution based on a multi-scale time window; A cache resource modeling module for generating resource weights according to the capacities of cache nodes; A cost evaluation module for calculating the access cost and migration cost between an object and a cache node; A policy generation module for generating a resource scheduling policy that minimizes the access cost; An equilibrium solving module for constructing a multi-node cache game model and solving for the cache policy in the equilibrium state through iteration; A policy execution module for deploying metadata replicas according to the cache policy and completing migration operations; An update module for periodically triggering access heat reconstruction and policy update to support dynamic optimization.

7. The distributed metadata management storage system based on cloud computing according to claim 6, characterized in that The equilibrium solving module includes a game modeler for constructing a cache game graph and a solving engine for iteratively calculating the Nash equilibrium.

8. The distributed metadata management storage system based on cloud computing according to claim 6, wherein The policy generation module uses a neural network to assist in evaluating the impact of the scheduling scheme on the access cost and supports self-learning optimization of the policy.

9. The distributed metadata management storage system based on cloud computing according to claim 6, wherein The cost evaluation module updates the latency information and network load of each node in real time for feedback adjustment of the cache deployment policy.

10. The distributed metadata management storage system based on cloud computing according to claim 6, wherein All modules of the system run on a distributed cloud platform cluster in a containerized microservice manner, supporting dynamic scaling at the module level.

Citation Information

Cited By

  • Intelligent hierarchical caching method and system based on hyper-converged architecture

    CN120705077A

  • Intelligent data storage method based on cloud platform

    CN120872250A

  • Memory access method and device, storage medium and computer equipment

    CN120929407A

  • Cache optimization method and device based on distributed storage

    CN121412185A

  • Enterprise-level image-text document management method, medium and system based on dynamic authority control

    CN121434113A