Cloud-edge collaborative data center hierarchical archival storage optimization method and system

By constructing a four-dimensional feature matrix and a spatiotemporal graph structure, data routing is dynamically optimized, solving the problems of disconnection risk and low resource utilization caused by network fluctuations in the cloud-edge collaborative architecture, and realizing low-latency hierarchical archiving of high-frequency data and cloud-based cold data settling.

CN120631277BActive Publication Date: 2025-11-07LINYI UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511127222.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-07
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Existing technologies fail to effectively address the risk of disconnection caused by dynamic network fluctuations in data center hierarchical archiving scenarios with cloud-edge collaborative architecture. This increases the archiving interruption rate, reduces storage resource utilization, and may cause high-frequency data to be mistakenly dumped to remote cloud nodes, leading to cross-domain access delays.

Method used

By constructing a four-dimensional feature matrix that integrates time, space, load, and network status, the network jitter rate is dynamically mapped as a constraint on network outage risk, and edge storage redundancy is used as a constraint on storage cost. Combined with the spatiotemporal graph structure, access hotspot areas are predicted, and data routing is optimized to achieve near-edge hierarchical archiving of high-frequency data.

Benefits of technology

Significantly reduces archiving interruption rate, improves storage resource utilization, optimizes cross-domain access latency, and enables low-latency hierarchical archiving of high-frequency data and cloud-based cold data settling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631277B_ABST
    Figure CN120631277B_ABST
Patent Text Reader

Abstract

The application provides a cloud-edge collaborative data center hierarchical archiving storage optimization method and system. In the application, a four-dimensional joint feature matrix is constructed based on network jitter rate, storage redundancy and access heat. The network jitter rate is mapped as a network outage risk constraint, and the storage redundancy is mapped as a storage cost constraint. A joint optimization objective is solved, and a storage distribution parameter is generated by locating the risk area. A cross-domain heat component is generated using the space-time dimension, and a dynamic heat is transferred by combining the parameters to construct a space-time graph. Future hotspots are predicted and high-frequency data sets are extracted. A routing mapping space is constructed by combining the heat properties of the data set and the node distance properties. High-frequency data is stored in a hierarchical archive in a distributed storage and indexed. By constructing a multi-dimensional feature matrix, jointly optimizing risk and cost, predicting access hotspots and fusing heat distance properties, the application realizes efficient, low-cost and network fluctuation-resistant hierarchical archiving storage optimization of high-frequency access data in a cloud-edge collaborative environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of cloud edge storage optimization, and particularly relates to a cloud edge collaborative data center hierarchical archiving storage optimization method and system. BACKGROUND

[0002] In the hierarchical archiving scenario of the data center in the cloud edge collaborative architecture, frequent network dynamic fluctuations on the edge side are prone to cause disconnection risks, resulting in a significant increase in cross-domain data access delay or even archiving operation failure. This requires a solution to have the ability to perceive and avoid network degradation areas in real time, to achieve a fine balance between limited edge storage resources and cloud resource costs, and to accurately capture and predict the migration law of data access hotspots in the time and space dimensions, and finally to achieve the optimization storage goal of archiving high-frequency data near the edge and sinking cold data to the cloud.

[0003] A current targeted solution is to use a spatiotemporal hotness prediction model based on LSTM for optimization. This solution extracts time and space features by analyzing historical access logs, constructs a time series prediction model of access hotness, uses the gating mechanism of LSTM to learn long-term dependencies to predict the distribution of future hotspot data, and according to a pre-set static storage cost threshold, caches the predicted high-frequency data to the edge node, and at the same time, combines a fixed period strategy to archive cold data to the cloud.

[0004] However, this existing solution has significant defects: its prediction process only relies on historical access patterns and fails to integrate key state parameters such as network jitter rate, resulting in a significant increase in archiving interruption rate when there is a sudden network degradation; the static storage cost threshold used by the solution cannot dynamically respond to actual changes in edge node load, resulting in low storage resource utilization; in addition, the solution lacks a collaborative optimization mechanism for the distance attribute between the cloud and the edge node and the access hotness, causing high-frequency data to be mistakenly sunk to a remote cloud node, significantly increasing cross-domain access delay. These deficiencies limit the effectiveness and efficiency of the solution in a dynamic network environment. SUMMARY

[0005] The present application provides a cloud edge collaborative data center hierarchical archiving storage optimization method and system to solve the problem of high archiving interruption rate, low storage utilization, and rapid increase in access delay caused by lack of network state perception, rigid cost control, and insufficient cross-domain collaboration in the prior art under dynamic network fluctuations.

[0006] In a first aspect, the present application provides a cloud edge collaborative data center hierarchical archiving storage optimization method, comprising:

[0007] According to the network jitter rate collected by the cloud edge network fluctuation scene, the edge storage redundancy and the cross-domain access heat, a joint feature matrix is constructed, which contains four dimensions of time dimension, space dimension, load dimension and network state dimension;

[0008] Based on the joint feature matrix, the network jitter rate of the network state dimension is mapped as a network outage risk constraint condition, and the edge storage redundancy of the load dimension is mapped as a storage cost constraint condition. A joint optimization objective is constructed according to the network outage risk constraint condition and the storage cost constraint condition, and an optimal solution of the joint optimization objective is solved. A network fluctuation parameter of the optimal solution is used to locate a network outage risk area, so as to generate a storage distribution parameter;

[0009] A cross-domain heat component is generated based on the time dimension and the space dimension of the joint feature matrix. A space-time graph structure is constructed by combining the storage distribution parameter. Dynamic access heat is transmitted through the space-time graph structure, so as to predict a future access hotspot area and extract a high-frequency access data set from the future access hotspot area;

[0010] A heat attribute is extracted from the high-frequency access data set. A distance attribute of a data center cloud storage node and an edge node is synchronously acquired. A data routing mapping space is constructed by fusing the heat attribute and the distance attribute. The high-frequency access data set is stored in a hierarchical manner and a distributed index is established in a distributed storage architecture of a data center based on the data routing mapping space.

[0011] Optionally, the joint feature matrix is used to map the network jitter rate of the network state dimension as a network outage risk constraint condition, and the edge storage redundancy of the load dimension as a storage cost constraint condition. A joint optimization objective is constructed according to the network outage risk constraint condition and the storage cost constraint condition, and an optimal solution of the joint optimization objective is solved. A network fluctuation parameter of the optimal solution is used to locate a network outage risk area, so as to generate a storage distribution parameter, which includes:

[0012] The network jitter rate is extracted from the network state dimension of the joint feature matrix, and a network outage risk threshold is set. The network jitter rate exceeding the network outage risk threshold is mapped as a network outage risk constraint condition;

[0013] The edge storage redundancy is extracted from the load dimension of the joint feature matrix, and a redundancy cost threshold is set. The edge storage redundancy lower than the redundancy cost threshold is mapped as a storage cost constraint condition;

[0014] A joint optimization objective containing the network outage risk constraint condition and the storage cost constraint condition is established;

[0015] adjusting network jitter parameters through an iterative optimization method to minimize the numerical value of the joint optimization target to obtain an optimal solution of the joint optimization target;

[0016] locating a disconnection risk region according to the position of the region in the optimal solution where the network jitter rate is continuously higher than the network disconnection risk threshold, and calculating storage distribution parameters including the number of risk region data backup copies and the compression ratio of non-risk region data backup copies according to the disconnection risk region.

[0017] Optionally, the cross-domain heat component is generated based on the time dimension and the space dimension of the joint feature matrix, a spatio-temporal graph structure is constructed by combining the storage distribution parameters, and dynamic access heat is transmitted through the spatio-temporal graph structure to predict a future access hotspot region and extract a high-frequency access data set from the future access hotspot region, including:

[0018] The time dimension feature value and the space dimension feature value are extracted from the joint feature matrix, and the time dimension feature value and the space dimension feature value at the same time point are multiplied to generate a cross-domain heat component;

[0019] The spatio-temporal graph structure is constructed based on the storage distribution parameters, the physical space grid is used as a graph node in the spatio-temporal graph structure, the cross-domain heat component is used as a node attribute value, and the reciprocal of the physical distance between grids is used as an edge weight value;

[0020] The transmission operation of the dynamic access heat is performed through the spatio-temporal graph structure, and the current node access heat value is fused with the weighted result of the access heat average value of adjacent nodes to update the access heat value of the current node;

[0021] The transmission operation of the dynamic access heat is repeatedly performed until the update amplitude of the access heat value of all nodes is less than a set convergence condition, and the predicted access heat of each grid in a future time window is output;

[0022] A number of grid regions with the highest predicted access heat are selected as future access hotspot regions, and a data set with an access frequency greater than a set threshold is extracted from the storage distribution corresponding to the future access hotspot regions as a high-frequency access data set.

[0023] Optionally, the heat attribute is extracted from the high-frequency access data set, the distance attribute of the data center cloud storage node and the edge node is synchronously obtained, and the data routing mapping space is constructed by fusing the heat attribute and the distance attribute, including:

[0024] The access frequency of each data block in a unit of time is extracted from the high-frequency access data set as a heat attribute, and the straight-line distance from the data center cloud storage node to each edge node is obtained as a distance attribute;

[0025] A three-dimensional coordinate space is established, which is composed of three coordinate axes, wherein the first coordinate axis corresponds to the hotness attribute, the second coordinate axis corresponds to the distance attribute, and the third coordinate axis corresponds to the routing priority weight;

[0026] An attribute fusion operation is performed on each data block, the hotness attribute and the distance attribute are weighted and calculated by a preset hotness coefficient and a distance coefficient to generate a routing priority weight value, and the routing priority weight value is mapped to the third coordinate axis of the three-dimensional coordinate space;

[0027] The three-dimensional coordinate point set of all data blocks is taken as a whole structure to form the data routing mapping space.

[0028] Optionally, the disconnected risk area is located according to a region position in which the network jitter rate in the optimal solution is continuously higher than the network disconnection risk threshold, and storage distribution parameters including a risk area data backup quantity and a non-risk area data copy compression ratio are calculated and generated according to the disconnected risk area, including:

[0029] An initial risk point set is formed by extracting monitoring points in which the network jitter rate continuously exceeds a preset time window from the optimal solution;

[0030] Geographically adjacent monitoring points in the initial risk point set are spatially clustered to generate a disconnected risk area with a continuous boundary;

[0031] The difference between the average network jitter rate of each disconnected risk area and the network disconnection risk threshold is calculated as an average jitter overrun value;

[0032] The data backup quantity of the risk area is set based on the average jitter overrun value, and a data copy compression ratio is set for the non-risk area;

[0033] The storage distribution parameters are generated according to the data backup quantity and the data copy compression ratio.

[0034] Optionally, the high-frequency access data set is stored in a hierarchical archive and a distributed index is established in a distributed storage architecture scheduled by the data routing mapping space in a data center, including:

[0035] Three storage levels are divided in the data routing mapping space according to the routing priority weight value, wherein the first storage level contains data blocks with routing priority weight values not less than a first preset threshold, the second storage level contains data blocks with routing priority weight values less than the first preset threshold but not less than a second preset threshold, and the third storage level contains data blocks with routing priority weight values less than the second preset threshold;

[0036] performing hierarchical storage operations in a distributed storage architecture scheduled by a data center, storing data blocks in the first storage hierarchy to edge storage areas of edge nodes, storing data blocks in the second storage hierarchy to an edge storage area of a single edge node, and uploading data blocks in the third storage hierarchy to a cloud storage node;

[0037] creating a distributed index table that records the unique identifier of each data block, the storage hierarchy to which it belongs, and the specific storage location coordinates;

[0038] synchronously storing a complete copy of the distributed index table to all edge nodes and cloud storage nodes;

[0039] When the hotness attribute value of a data block in the high-frequency access data set changes, the routing priority weight value of the data block in the data routing mapping space is recalculated and the distributed index table is updated.

[0040] Optionally, the network jitter rate, edge storage redundancy, and cross-domain access heat collected according to the cloud-edge network fluctuation scenario are used to construct a joint feature matrix, which contains four dimensions of time, space, load, and network state, including:

[0041] Collecting the network jitter rate, edge storage redundancy, and cross-domain access heat under the cloud-edge network fluctuation scenario;

[0042] Dividing a continuous time period into equal-length segments, and in each time segment, the cumulative time distribution of the cross-domain access heat is counted to generate time-dimension feature values, while the physical space is divided into grid regions, and in each grid region, the geographic distribution density of the cross-domain access heat is calculated to generate space-dimension feature values, and the network jitter rate is directly used as a network state-dimension feature value, and the edge storage redundancy is directly used as a load-dimension feature value;

[0043] Combining the time-dimension feature values, space-dimension feature values, load-dimension feature values, and network state-dimension feature values of the same time point in time sequence into a four-dimensional feature vector;

[0044] Arranging all the four-dimensional feature vectors of all time points in time sequence to form a joint feature matrix with rows and columns aligned, where the rows correspond to time points and the columns correspond to four-dimension feature values.

[0045] In a second aspect, the present application provides a cloud-edge collaborative data center hierarchical archiving storage optimization system, comprising:

[0046] The collection module is configured to collect network jitter, edge storage redundancy and cross-domain access heat according to a cloud-edge network fluctuation scenario, and construct a joint feature matrix, wherein the joint feature matrix comprises four dimensions of time dimension, space dimension, load dimension and network state dimension.

[0047] The generation module is configured to map the network jitter of the network state dimension to a network disconnection risk constraint condition, map the edge storage redundancy of the load dimension to a storage cost constraint condition based on the joint feature matrix, construct a joint optimization target according to the network disconnection risk constraint condition and the storage cost constraint condition, and solve an optimal solution of the joint optimization target, and locate a network disconnection risk area according to a network fluctuation parameter of the optimal solution to generate a storage distribution parameter.

[0048] The extraction module is configured to generate a cross-domain heat component based on the time dimension and the space dimension of the joint feature matrix, construct a space-time graph structure in combination with the storage distribution parameter, transfer dynamic access heat through the space-time graph structure, predict a future access hotspot area, and extract a high-frequency access data set from the future access hotspot area.

[0049] The storage module is configured to extract a heat attribute from the high-frequency access data set, synchronously acquire a distance attribute of a cloud storage node and an edge node of a data center, fuse the heat attribute and the distance attribute to construct a data routing mapping space, and perform hierarchical archival storage on the high-frequency access data set in a distributed storage architecture of the data center based on the data routing mapping space and establish a distributed index.

[0050] In a third aspect, the present application provides a computing device, comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the cloud-edge collaborative data center hierarchical archival storage optimization method of the first aspect.

[0051] In a fourth aspect, the present application provides a computer storage medium, which stores a computer program, and the computer program is executed by a computer to implement the cloud-edge collaborative data center hierarchical archival storage optimization method of the first aspect.

[0052] In the examples of the present application, according to the network jitter rate collected by the cloud-edge network fluctuation scene, the edge storage redundancy and the cross-domain access heat, a joint feature matrix is constructed, which contains four dimensions of time dimension, space dimension, load dimension and network state dimension; based on the joint feature matrix, the network jitter rate of the network state dimension is mapped as a network outage risk constraint condition, the edge storage redundancy of the load dimension is mapped as a storage cost constraint condition, a joint optimization objective is constructed according to the network outage risk constraint condition and the storage cost constraint condition, and the optimal solution of the joint optimization objective is solved, the network fluctuation parameter of the optimal solution is used to locate the network outage risk area, so as to generate a storage distribution parameter; based on the time dimension and the space dimension of the joint feature matrix, a cross-domain heat component is generated, a space-time graph structure is constructed by combining the storage distribution parameter, dynamic access heat is transmitted through the space-time graph structure, future access hotspot areas are predicted, and a high-frequency access data set is extracted from the future access hotspot areas; the heat attribute is extracted from the high-frequency access data set, the distance attribute of the data center cloud storage node and the edge node is synchronously acquired, the heat attribute and the distance attribute are fused to construct a data routing mapping space, and the high-frequency access data set is stored and indexed in a distributed storage architecture of the data center based on the data routing mapping space.

[0053] The technical scheme of the present application has the following beneficial effects:

[0054] The present application constructs a joint matrix that fuses four-dimensional features of time, space, load and network state, dynamically maps the network jitter rate as a network outage risk constraint and the storage redundancy as a cost constraint, generates a storage distribution parameter resistant to network fluctuations through joint optimization, further constructs a space-time graph structure using a heat component of the space-time dimension to transmit dynamic access heat, accurately predicts future access hotspot areas and extracts a high-frequency data set, finally fuses the heat attribute of the data itself and the distance attribute of the cloud-edge node to construct a routing mapping space, realizes near-edge low-delay hierarchical archiving of high-frequency data and cloud cold data settlement in a distributed storage architecture, synchronously establishes a global distributed index, significantly reduces the archiving interruption rate, improves the storage resource utilization rate and optimizes the cross-domain access delay.

[0055] Further, the network jitter rate is extracted from the network state dimension of the joint feature matrix, and a network outage risk threshold is set, and the jitter rate exceeding the threshold is mapped as a network outage risk constraint; the edge storage redundancy is extracted from the load dimension, and a redundancy cost threshold is set, and the redundancy below the threshold is mapped as a storage cost constraint; a joint optimization objective containing both types of constraints is established; the network fluctuation parameters are adjusted through iterative optimization to minimize the objective value, and the optimal solution is obtained; the network jitter rate region continuously exceeding the threshold in the optimal solution is located as the risk area, and finally the storage distribution parameters containing the risk area backup quantity and the non-risk area copy compression ratio are generated. The method dynamically converts the network jitter rate and the storage redundancy into the network outage risk constraint and the storage cost constraint, and accurately identifies the risk area through joint optimization, and generates adaptive storage distribution parameters accordingly: increases data backup in the identified high-risk area to enhance system fault tolerance, and compresses the data copy ratio in the low-risk area to optimize storage resource occupation, thereby effectively coordinating the control of network fluctuation environment and reducing the storage cost.

[0056] These aspects or other aspects of the present application will be more apparent in the following description of the embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0058] Figure 1 A flow chart of a cloud-edge collaborative data center hierarchical archiving storage optimization method provided by the present application is shown;

[0059] Figure 2 A scene diagram of a cloud-edge collaborative data center hierarchical archiving storage optimization method provided by the present application is shown;

[0060] Figure 3 A structural schematic diagram of a cloud-edge collaborative data center hierarchical archiving storage optimization system provided by the present application is shown;

[0061] Figure 4 A structural schematic diagram of a computing device provided by the present application is shown. DETAILED DESCRIPTION

[0062] In order to enable personnel in the technical field to better understand the present application scheme, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.

[0063] In some of the flowcharts described in the description and claims of the present application and in the above-described figures, a plurality of operations are included which occur in a particular order, but it should be clearly understood that the operations can be performed in an order other than that in which they appear or in parallel, and the serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, the flowcharts can include more or fewer operations, and the operations can be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this paper are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do "first" and "second" represent different types.

[0064] Research shows that the existing cloud-edge collaborative archiving scheme based on LSTM ignores real-time network fluctuations, uses static storage thresholds, and lacks a distance-heat collaborative mechanism, which still distributes high-frequency data in network degradation areas, resulting in high archiving interruption rate, low storage resource utilization, and mis-settlement of high-frequency data to remote cloud nodes, causing a surge in cross-domain access delay, making it difficult to meet the comprehensive needs of reliability, economy and timeliness in dynamic network environments.

[0065] To solve the above problems, the present application provides a cloud-edge collaborative data center hierarchical archiving storage optimization method, which maps the real-time network jitter rate dynamically to the network interruption risk constraint and the storage redundancy to the cost constraint by constructing a joint matrix integrating time, space, load and network state four-dimensional features, and solves the anti-fluctuation storage distribution parameters through joint optimization. Combined with the time and space dimensions, a heat transfer graph is constructed to predict access hotspots. Finally, the data heat attribute and node distance attribute are fused to form a routing mapping space, driving the near-edge hierarchical archiving of high-frequency data. This method not only actively avoids disconnected areas through real-time network perception, but also optimizes storage resource allocation through dynamic cost constraints, and further eliminates data mis-settlement through distance-heat coupling, fundamentally solving the archiving failure, resource waste and high delay problems of traditional schemes under network fluctuations.

[0066] The technical solutions in the embodiments of the present application will be described clearly and completely in the following combined with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0067] Figure 1 A flowchart of a cloud-edge collaborative data center hierarchical archiving storage optimization method is provided for the embodiments of the present application, as shown in Figure 1 The method comprises:

[0068] 101. Construct a joint feature matrix according to the network jitter rate, edge storage redundancy and cross-domain access heat collected under the cloud-edge network fluctuation scenario, the joint feature matrix containing four dimensions of time dimension, space dimension, load dimension and network state dimension;

[0069] Optionally, step 101 can specifically include the following steps:

[0070] 1011. Collect the network jitter rate, edge storage redundancy and cross-domain access heat under the cloud-edge network fluctuation scenario;

[0071] 1012. Divide the continuous time period into equal length segments, and in each time segment, count the cumulative time distribution of the cross-domain access heat to generate time dimension feature values, at the same time, divide the physical space into grid areas, and in each grid area, calculate the geographic distribution density of the cross-domain access heat to generate space dimension feature values, and directly take the network jitter rate as the network state dimension feature value, and directly take the edge storage redundancy as the load dimension feature value;

[0072] 1013. Combine the time dimension feature values, space dimension feature values, load dimension feature values and network state dimension feature values of the same time point in time sequence into four-dimensional feature vectors;

[0073] 1014. Arrange the four-dimensional feature vectors of all time points in time sequence to form a joint feature matrix with rows and columns aligned, the rows of the joint feature matrix corresponding to time points, and the columns corresponding to four dimension feature values.

[0074] In the above scheme, the network jitter rate: refers to the dynamic fluctuation index reflecting the stability of cloud-edge network transmission, contains the discrete degree and burst fluctuation characteristics of data transmission delay, and can be used for quantitative evaluation of network disconnection risk level. Edge storage redundancy: refers to the dynamic parameter representing the utilization state of edge node storage resources, contains the proportional relationship between available storage capacity and total capacity, and can be used to determine the storage cost optimization space. Cross-domain access heat: refers to the spatio-temporal distribution characteristics of geographically dispersed user access data sets, including access frequency per unit time and density changes in physical area, which can be used to identify high-frequency access hotspots. Time dimension feature value: refers to the periodicity quantization result extracted from continuous access behavior, which can be used to capture the time period distribution characteristics of access demand. Spatial dimension feature value: refers to the grid density index mapping the geographical concentration of access behavior, which can be used to locate spatial access hotspots. Load dimension feature value: refers to the native parameter directly representing the storage pressure of edge node, which can be used to drive the dynamic adjustment of storage strategy. Network state dimension feature value: refers to the original observation value directly reflecting the quality of network transmission channel, which can be used to build a network risk warning mechanism. Joint feature matrix: refers to the time series structured data set integrating multi-dimensional dynamic features, containing four types of feature vectors of time, space, load and network state aligned by time sequence, which can be used to support global optimization decision of cloud-edge collaborative storage.

[0075] In the embodiment of the application, first step 1011 synchronously collects raw data through three types of special interfaces, uses a network probe to send data packets to the edge node every 5 seconds, calculates the network jitter rate according to the fluctuation amplitude of the continuous packet delay, for example, the measured delay value sequence in 5 seconds is [80ms, 100ms, 205ms, 90ms, 150ms], then the network jitter rate = average absolute difference = 125ms; the capacity query API of the edge storage system is called to obtain the storage pool usage state, and the idle space ratio is used as the edge storage redundancy, for example, the total capacity of 10TB has 3.5TB of idle space, then the edge storage redundancy = 0.35; at the same time, the access log stream is parsed, the request times per minute are aggregated according to the geographical area to generate the cross-domain access heat, for example, region C has accumulated 200 accesses in the past 1 minute. The three types of data of network jitter rate, edge storage redundancy and cross-domain access heat constitute the basic raw materials for subsequent processing.

[0076] Secondly, by step 1012, based on the collected network jitter rate, edge storage redundancy and cross-domain access heat, the full-day time axis is divided into 96 15-minute segments, the total number of accesses in all regions in this period is accumulated as the time dimension characteristic value, for example, 8:00-8:15 is the 32nd segment, the total number of accesses in all regions in this period = 200 times in region C + 50 times in region D = 250 times; the physical service area is divided into a 1km² grid, and the spatial dimension characteristic value is calculated by dividing the number of accesses in the grid by the grid area, for example, region C is divided into 9 grids, the center grid C5 of region C has an area of 0.8km² and an access amount of 120 times, then the spatial dimension characteristic value = density value = 150 times / km²; the latest collected network jitter rate is directly used as the network state dimension characteristic value, and the edge storage redundancy is taken as the load dimension characteristic value.

[0077] Next, by step 1013, the time dimension characteristic value, the spatial dimension characteristic value, the load dimension characteristic value and the network state dimension characteristic value processed at the same time are combined into a machine-processable vector unit, a fixed-structure data container [time value, space value, load value, network state value] is created, the processing result at the current time is injected, and an accurate time stamp is attached to the vector. For example, at 8:00:05, the time value is 250, the space value is 150, the load value is 0.35, and the network state value is 125, then the corresponding vector is [250, 150, 0.35, 125], and the Unix timestamp is 1620000005.

[0078] Finally, by step 1014, all vectors are integrated in time sequence, the vectors with time stamps are stored in a time series database, and data in continuous periods is extracted at fixed intervals; a matrix with strictly aligned rows and columns is constructed, each row corresponds to a complete vector at a time, and each column fixedly represents a specific dimension. For example, 12 vectors in the 8:00-8:15 period are integrated to form a joint feature matrix of 12 rows x 4 columns, and the conversion from raw data to structured analysis base is completed.

[0079] In practical applications, in the collaborative scenario of edge node X and cloud data center Y, the system first performs data collection, measures the network delay fluctuation of node X every 5 seconds through ICMP network probe, measures the jitter rate at 8:00:00 to be 125 ms, at the same time, calls the storage API of node X to obtain the capacity state, calculates the redundancy to be 0.35, and parses the access log to count the total access amount of geographic region Z in the last minute to be 200 times. Then, feature processing is performed, the whole day is divided into 15-minute periods, and the time dimension value of the total region access amount in the 8:00-8:15 period is accumulated to be 250 times; region Z is divided into a 1km² grid, and the spatial dimension value of the 120 accesses in the center grid area of 0.8km² is calculated to be 150 times / km²; the aforementioned jitter rate 125 ms is directly used as the network state value, and the redundancy 0.35 is directly used as the load value. Subsequently, vector packaging is performed, and at 8:00:05, the four processed values are combined into a feature vector [250, 150, 0.35, 125] and marked with a timestamp. Finally, a matrix is constructed, and 12 vectors generated every 5 seconds in the 8:00-8:15 period are extracted to arrange a joint feature matrix of 12 rows x 4 columns in time sequence, wherein each row corresponds to a sampling time, and the four columns carry time value, spatial value, load value and network state value in turn.

[0080] The overall scheme of the above 101, by systematizing the collection of network jitter rate, edge storage redundancy and cross-domain access heat, and structuring and integrating the three types of key data into a joint matrix containing four-dimensional features of time, space, load and network state, a comprehensive data base for dynamically perceiving the cloud-edge environment is constructed, which not only converts the original network fluctuation, storage pressure and access behavior into standardized features that can be processed by machines, but also retains the dynamic correlation between dimensions through time sequence alignment, providing a unified multi-dimensional analysis input for subsequent risk avoidance, cost optimization and hotspot prediction, like installing a panoramic instrument panel on the system to monitor the network health, storage load and access heat.

[0081] 102、Based on the joint feature matrix, the network jitter rate of the network state dimension is mapped as a network outage risk constraint condition, the edge storage redundancy of the load dimension is mapped as a storage cost constraint condition, a joint optimization objective is constructed according to the network outage risk constraint condition and the storage cost constraint condition, and the optimal solution of the joint optimization objective is solved, and a network fluctuation parameter of the optimal solution is used to locate a network outage risk area to generate a storage distribution parameter;

[0082] Optionally, step 102 can specifically include the following steps:

[0083] 1021、Extract the network jitter rate from the network state dimension of the joint feature matrix and set a network outage risk threshold, and map the network jitter rate exceeding the network outage risk threshold as a network outage risk constraint condition;

[0084] 1022、extracting edge storage redundancy from the load dimension of the joint feature matrix and setting a redundancy cost threshold, mapping edge storage redundancy below the redundancy cost threshold as a storage cost constraint;

[0085] 1023、establishing a joint optimization objective containing both the network disconnection risk constraint and the storage cost constraint;

[0086] 1024、adjusting network fluctuation parameters through an iterative optimization method to minimize the value of the joint optimization objective to obtain an optimal solution of the joint optimization objective;

[0087] 1025、locating a disconnection risk region according to the location of the region where the network jitter rate continuously exceeds the network disconnection risk threshold in the optimal solution, and calculating a storage distribution parameter containing the number of data backup copies in the risk region and the compression ratio of data copies in the non-risk region according to the disconnection risk region.

[0088] In the step 1025, the following process can be specifically included: extracting monitoring points where the network jitter rate continuously exceeds a preset time window from the optimal solution to form an initial risk point set; spatially clustering geographically adjacent monitoring points in the initial risk point set to generate a disconnection risk region with a continuous boundary; calculating the difference between the average network jitter rate of each disconnection risk region and the network disconnection risk threshold as an average jitter overrun value; setting the number of data backup copies in the risk region based on the average jitter overrun value, and setting the data copy compression ratio in the non-risk region; and generating a storage distribution parameter according to the number of data backup copies and the data copy compression ratio.

[0089] In the above scheme, the network disconnection risk threshold refers to a jitter rate critical value triggering a network disconnection warning, contains monitored network fluctuation peak data, includes a continuous overrun duration determination condition, and can be used to identify a high-risk transmission path. The redundancy cost threshold refers to a warning line of edge storage resource occupation cost, contains a proportional relationship between idle capacity and total capacity, includes a dynamic load pressure coefficient, and can be used to control the economic allocation of storage copies. The joint optimization objective refers to a mathematical expression combining network disconnection risk and storage cost, and can be used to balance system reliability and resource overhead. The network fluctuation parameter refers to an adjustable variable affecting data transmission stability, and can be used to optimize network status. The disconnection risk region refers to a physical space range where network disconnection risk continuously exists, and can be used to guide data disaster recovery deployment. The storage distribution parameter refers to a set of operation instructions for controlling data copy distribution, and can be used to realize hierarchical storage optimization.

[0090] In the embodiments of the present application, first, all network jitter rates are extracted from the "network state value" column of the joint feature matrix by step 1021, and a preset network outage risk threshold is set. Network jitter rates exceeding the network outage risk threshold are converted into network outage risk constraint conditions, and risk constraint values are calculated for each monitoring point, with the calculation formula being as follows: risk constraint value = max (network jitter rate - threshold). For example, if the monitored network jitter rate is 170 ms, which is greater than the network outage risk threshold 150 ms, the risk constraint value is 20; and if the monitored network jitter rate is 155 ms, which is greater than the threshold 150 ms, the risk constraint value is 5. Finally, a network outage risk constraint set {monitoring point A: 20, monitoring point B: 5} is generated, and each entry represents a risk degree quantitative value of a specific location.

[0091] Secondly, edge storage redundancy is extracted from the load dimension of the joint feature matrix by step 1022, and a redundancy cost threshold is set. Redundancies lower than the redundancy cost threshold are converted into storage cost constraint conditions, and cost constraint values are calculated for each node, with the calculation formula being as follows: cost constraint value = max (redundancy cost threshold - edge storage redundancy). For example, if the monitored edge storage redundancy is 0.18 and the redundancy cost threshold is 0.25, the corresponding cost constraint value is 0.25-0.18 = 0.07; and if the monitored edge storage redundancy is 0.22, the corresponding cost constraint value is 0.25-0.22 = 0.03. The storage cost constraint set {node X: 0.07, node Y: 0.03} is output, and each entry reflects a storage resource gap of a specific node.

[0092] Then, the network outage risk constraint set and the storage cost constraint set are input into an optimization engine by step 1023, and a joint optimization objective function containing double constraints is established: wherein is a risk weight, is a cost weight. Taking the constraint set as an example, the risk constraint sum is 20+5 = 25, and the cost constraint sum is 0.07+0.03 = 0.1, so the initial objective value F(x) is 0.7x25+0.3x0.1 = 17.53.

[0093] Then, the system performs iterative optimization by using a genetic algorithm by step 1024: three groups of network parameters are randomly initialized, and the joint objective function value under each group of parameters is calculated; then the current optimal scheme is selected for gene mutation, and the function value of the new scheme is re-evaluated; through multiple rounds of iteration, the optimal scheme is retained and mutated in each round until the function value tends to be stable, and finally the optimal solution that minimizes the sum of the network outage risk constraint value and the storage cost constraint value is output, so as to realize the dynamic balance between network stability and storage economy. For example, the path scheme and the corresponding F(x) value are P1: 18.2, P2: 17.53, and P3: 19.1, the optimal scheme P2 is selected for parameter mutation, and a new scheme P2' and calculate F(x)=16.8, stop iteration when F(x) changes <0.01 for 3 rounds, output the optimal solution: scheme P2 ' , corresponding risk constraint sum 23, cost constraint sum 0.08.

[0094] Finally, the network jitter rate of the monitoring points exceeding the preset threshold for three times is filtered from the optimal solution through step 1025 to form an initial risk point set; then, a spatial clustering algorithm is used to aggregate the risk points that are geographically adjacent into physically continuous disconnected risk areas, such as points A, B, and C aggregated into area Z; subsequently, the difference between the average network jitter rate of each risk area and the disconnected network risk threshold is calculated as the average jitter overrun value; the number of data backups of the risk area is dynamically set based on the average jitter overrun value, and the calculation formula is as follows: data backup number = basic value 2 + average constraint value x adjustment coefficient 0.5, such as the data backup number of the physical risk area Z being 7; at the same time, the data copy compression ratio of the non-risk area is uniformly set, such as retaining 50% copies; finally, the storage distribution parameters containing the number of data backups of the risk area and the copy compression ratio of the non-risk area are integrated and output, such as {"risk area Z": 7, "non-risk area": 0.5}, to complete the closed-loop generation from risk identification to storage strategy.

[0095] In actual application, in the optimization scenario of the edge node X, the system first performs risk constraint mapping, extracts the network jitter rate sequence [125ms, 170ms, 155ms, 140ms] from the joint feature matrix, sets the disconnected network risk threshold to 150ms, and generates the disconnected network risk constraint set {monitoring point A: 20 (170-150), monitoring point B: 5 (155-150)}. Then, the cost constraint mapping is performed, the storage redundancy [0.35, 0.18, 0.22] is extracted, the cost threshold is set to 0.25, and the storage cost constraint set {node P: 0.07 (0.25-0.18), node Q: 0.03 (0.25-0.22)} is generated. Subsequently, the joint optimization target is constructed, the function is established, and the optimal solution F(x)=0.7x(20+5)+0.3x(0.07+0.03)=17.53 is obtained. Then, the genetic algorithm is iteratively optimized, F(x)=17.53 for the initial path B, and the new path B' is generated through parameter adjustment to reduce F(x) to 16.8. Finally, the storage parameters are generated, the monitoring point A with a jitter rate sequence [152, 158, 161] exceeding the threshold for three times is extracted, and the risk area Z is clustered, with an average jitter rate of 160ms, the average jitter overrun value = 160-150 = 10ms, and the number of backups set to = 2+10x0.5 = 7; the copy of the non-risk area is compressed to 50%, and the output parameters {"risk area Z backup number": 7, "non-risk area copy ratio": 0.5} are obtained.

[0096] The overall solution of 102 converts network fluctuation and storage load into quantifiable constraints, generates adaptive instructions through intelligent optimization, increases data backup in high-risk areas identified to improve reliability, and reduces data replicas in stable areas to save storage space, achieving dynamic balance between risk avoidance and cost control in a network fluctuation environment.

[0097] 103. Generate a cross-domain heat component based on the time dimension and the space dimension of the joint feature matrix, construct a spatio-temporal graph structure combining the storage distribution parameters, and transfer dynamic access heat through the spatio-temporal graph structure to predict future access hotspot areas and extract a high-frequency access dataset from the future access hotspot areas;

[0098] Optionally, step 103 can specifically include the following steps:

[0099] 1031. Extract time dimension feature values and space dimension feature values from the joint feature matrix, multiply the time dimension feature values and the space dimension feature values at the same time point to generate a cross-domain heat component;

[0100] 1032. Construct a spatio-temporal graph structure based on the storage distribution parameters, wherein the physical space grid is used as a graph node in the spatio-temporal graph structure, the cross-domain heat component is used as a node attribute value, and the reciprocal of the physical distance between grids is used as an edge weight value;

[0101] 1033. Perform a transfer operation of dynamic access heat through the spatio-temporal graph structure, fuse the weighted result of the current node access heat value and the adjacent node access heat mean value to update the own access heat value;

[0102] 1034. Repeat the transfer operation of dynamic access heat until the update amplitude of the access heat value of all nodes is less than the set convergence condition, and output the predicted access heat of each grid in the future time window;

[0103] 1035. Select a number of grid areas with the highest predicted access heat as future access hotspot areas, and extract a dataset with an access frequency greater than a set threshold from the storage distribution corresponding to the future access hotspot areas as a high-frequency access dataset.

[0104] In the above scheme, the cross-domain heat component refers to a comprehensive index reflecting the intensity of data access spatiotemporal distribution, which can be used to quantify the data demand intensity of a specific region at a specific time. The spatiotemporal graph structure refers to a network model describing the access heat transfer path, which can be used to simulate the dynamic diffusion process of access heat in geographical space. The dynamic access heat refers to the access demand prediction value migrating between grids over time, which can be used to iteratively update the regional access intensity in the future time window. The future access hotspot region refers to the predicted high data demand geographical range, which can be used to locate the storage location of high-frequency data that needs to be deployed preferentially. The high-frequency access dataset refers to the frequently accessed data subset in the hotspot region, which can be used to drive the resource tilt allocation of hierarchical archiving storage.

[0105] In the embodiment of the present application, first, the time dimension feature value of a specific time is read from the joint feature matrix through step 1031, that is, the cumulative access total in the time period and the space dimension feature value, that is, the access density of the geographical grid, are multiplied to generate the cross-domain heat component. Taking grid B3 as an example: the time dimension feature value is read as 300 times and the space dimension feature value is 120 times / km2, which are multiplied to obtain the cross-domain heat component 36,000. This value integrates the time cumulative effect and the space density and becomes the core index for quantifying regional data demand. This component value will be used as the initial attribute value of each node in the subsequent graph structure.

[0106] Secondly, through step 1032, the geographical grid division defined based on the storage distribution parameters is performed, such as the parameter in which the risk area Z contains grids B1, B2 and B3, and then the system constructs the spatiotemporal graph structure: each physical grid is converted into a graph node, and the cross-domain heat component calculated in 1031 is assigned to the corresponding node, such as creating node B3 and assigning node attribute value 36,000. Then, the spatial relationship between grids is calculated, and the reciprocal of the distance between the center points of two grids is taken as the weight value of the connected edge, for example, grid B2 is 1.5 kilometers away from grid B3, so the edge weight is set to . Finally, a topological network with weighted connections is formed, in which node B3 carries the initial heat 36,000 and is connected to node B2 through an edge with a weight of 0.67.

[0107] Next, the iterative update mechanism is started to simulate heat transfer. For each node, the weighted average value of the heat of its adjacent nodes is calculated, and the weight is the value of the connected edge. Then, the current value and the adjacent influence are fused according to the retention coefficient to update the access heat value of itself, and the calculation formula is as follows: new heat = 0.6*current value + 0.4*adjacent average value. Taking node B3 as an example: the heat of its only adjacent node B2 is 30,000, the connection weight is 0.67, and the weighted average calculation result is . Substituting B3 data, we get 0.6*36,000+0.4*30,000=33,600. This process reduces the hotness value of node B3 from 36,000 to 33,600, reflecting the diffusion effect of access demand from high-density areas to adjacent areas.

[0108] Then, the hotness transfer operation is repeated by step 1034: taking the updated B3 hotness value 33,600 as the new round of input, calculating the adjacent mean value again (still 30,000) and merging to get a new value 31,440 (0.6*33,600+0.4*30,000). After multiple iterations (5 rounds in the example), when the change in hotness is less than a preset threshold (such as 5%) for three consecutive times, it is determined to converge, and the stable hotness value 31,000 of node B3 is finally output as the predicted access intensity in the future period. This value reflects the initial hotness, spatial correlation and stable demand after multiple rounds of diffusion.

[0109] Finally, based on the converged graph structure by step 1035, the system first sorts all grids by predicted hotness: B3 (31,000) > B2 (28,000) > B1 (25,000), and selects the top two grids [B3, B2] as the future access hotspot areas. Then scan the storage data records in these areas: in grid B3, it is found that the historical access volume of file F1 is 120 times, which exceeds the set threshold of 80 times; while in grid B2, the access volume of file F2 is 75 times, which is lower than the threshold. Finally, the high-frequency access data set {F1} is extracted, and the accurate positioning from area prediction to specific data object is completed, providing target carriers for hierarchical storage.

[0110] In actual application, in the data prediction scenario of area M, the system first calculates the cross-domain hotness component, obtains the time dimension value 400 times in the 9:00-9:15 period and the space dimension value 150 times / km² of grid C2 from the joint feature matrix, and multiplies them to get the component value 400*150=60,000. Then construct the spatio-temporal graph structure, create nodes C1, C2 and C3 based on the grid division defined by the storage parameters, inject the component value 60,000 into the grid C2 node; calculate the distance between C2 and C3 as 2 km, and set the edge weight . Then perform hotness transfer, take the current value of C2 as 60,000, the adjacent C3 value as 45,000, calculate the adjacent weighted mean value ; update C2 hotness according to formula 0.6*60,000+0.4*45,000=54,000. After multiple rounds of iteration convergence, the first round 54,000→the second round 51,600→the third round 49,440→the fourth round 48,176, output C2 prediction value 48,000. Finally, extract high-frequency data, sort by predicted hotness to determine C2 as a hot area, scan C2 storage data with access frequency > 100 times, extract F3 with file access frequency 150 times and F4 with access frequency 120 times, and generate high-frequency data set {F3, F4}.

[0111] The overall scheme of the above 103 generates a cross-domain hotness component by fusing time and space dimension features, constructs a space-time graph network based on physical distance, dynamically simulates the inter-regional transfer of access hotness by means of iterative diffusion mechanism, accurately predicts future access hot spot areas, and extracts high-frequency access data set from them, realizes the connection from macro regional demand prediction to micro data object positioning, provides prospective data scheduling basis for subsequent hierarchical storage, and significantly improves the accuracy and timeliness of hot data near edge deployment.

[0112] 104, extract hotness attributes from the high-frequency access data set, synchronously acquire distance attributes of data center cloud storage nodes and edge nodes, fuse the hotness attributes and distance attributes to construct a data routing mapping space, and perform hierarchical archival storage and establish distributed index for the high-frequency access data set in the distributed storage architecture scheduled in the data center based on the data routing mapping space.

[0113] Optionally, step 104 can specifically include the following steps:

[0114] 1041, extract the access frequency of each data block in unit time as the hotness attribute from the high-frequency access data set, and acquire the straight-line distance from the data center cloud storage node to each edge node as the distance attribute;

[0115] 1042, establish a three-dimensional coordinate space composed of three coordinate axes, the first coordinate axis in the three-dimensional coordinate space corresponds to the hotness attribute, the second coordinate axis corresponds to the distance attribute, and the third coordinate axis corresponds to the routing priority weight;

[0116] 1043, perform attribute fusion operation on each data block, generate routing priority weight value by weighting calculation of the hotness attribute and distance attribute through preset hotness coefficient and distance coefficient, and map the routing priority weight value to the third coordinate axis of the three-dimensional coordinate space;

[0117] 1044, take the three-dimensional coordinate point set of all data blocks as a whole structure to constitute the data routing mapping space.

[0118] 1045、In the data routing mapping space, three storage levels are divided according to the routing priority weight values, the first storage level in the storage levels contains data blocks with routing priority weight values not less than a first preset threshold, the second storage level contains data blocks with routing priority weight values less than the first preset threshold but not less than a second preset threshold, and the third storage level contains data blocks with routing priority weight values less than the second preset threshold;

[0119] 1046、In the distributed storage architecture scheduled by the data center, hierarchical storage operations are performed, data blocks in the first storage level are stored in the edge storage area of the edge node, data blocks in the second storage level are stored in the edge storage area of a single edge node, and data blocks in the third storage level are uploaded to the cloud storage node;

[0120] 1047、A distributed index table is created, which records the unique identifier of each data block, the storage level to which it belongs, and the specific storage location coordinates;

[0121] 1048、A complete copy of the distributed index table is stored in all edge nodes and cloud storage nodes;

[0122] 1049、When the hotness attribute value of a data block in the high-frequency access data set changes, the routing priority weight value of the data block in the data routing mapping space is recalculated and the distributed index table is updated.

[0123] In the above scheme, the hotness attribute refers to a quantitative indicator reflecting the intensity of data access demand, including the cumulative number of times a data block is requested within a unit time, which can be used to measure the popularity of a data object. The distance attribute refers to a spatial parameter representing the physical separation between storage nodes, including the straight-line distance measurement (unit: kilometers) from the cloud data center to the edge node, which can be used to evaluate the potential delay cost of data transmission. The routing priority weight refers to a comprehensive score value driving storage decisions, including the linear superposition of the weighted calculation result of the hotness attribute and the inverse transformation result of the distance attribute, which can be used to determine the optimal storage level of the data block in the cloud-edge architecture. The data routing mapping space refers to a three-dimensional coordinate system that visualizes the storage strategy, including a spatial topology structure composed of the hotness attribute axis, the distance attribute axis, and the routing priority weight axis, which can be used to globally map the distribution location and storage strategy of all data blocks. The distributed index table refers to a dynamic directory recording data location information, including a key-value pair set of data block unique identifiers, storage level labels, and specific physical location coordinates, which can be used to support fast data positioning and retrieval across nodes.

[0124] In the embodiments of the present application, firstly, the hotness attribute of each data block is extracted from the high-frequency access data set through step 1041, the access frequency in the recent fixed period is counted by querying the log, and the quantized access intensity value is obtained as the hotness attribute; the spherical distance from the data center cloud node to each edge node is calculated by synchronously calling the geographic positioning service to obtain the accurate physical transmission cost parameter as the distance attribute. The hotness attribute and the distance attribute constitute the original input of the subsequent storage decision. For example, file F1 is accessed 120 times in the past 1 hour as the hotness attribute, and edge station E1 is 150 kilometers away from the cloud as the distance attribute.

[0125] Subsequently, a three-dimensional Cartesian coordinate system is established through steps 1042 and 1043, wherein the X axis bears the hotness attribute, the Y axis bears the distance attribute, and the Z axis bears the routing priority weight. Weighted fusion calculation is performed on each data block, and the formula is used to generate the priority value, wherein =0.7 is the hotness coefficient, =0.3 is the distance coefficient, is the hotness value, is the distance value, and the formula ensures that high access frequency and nodes with short distances obtain higher weights. The calculation result is mapped to the Z axis to form a three-dimensional coordinate point, realizing the conversion of abstract attributes to spatial positions. For example, the hotness value of file F1 is 120, and the distance value is 150, so the of F1 corresponds to the coordinate (120, 150, 84.002).

[0126] Secondly, the three-dimensional coordinate points of all data blocks such as F1 (120, 150, 84.002) and F2 (80, 300, 56.001) are integrated into a unified space model through step 1044 to form a data routing mapping space. The data routing mapping space intuitively presents the data characteristics through point cloud distribution, wherein the upper right quadrant (high X value + low Y value) aggregates high-priority data, and the lower left quadrant (low X value + high Y value) aggregates low-priority data, providing a visual basis for hierarchical decision-making. For example, the three coordinate points F1 (120, 150, 84.002) and F2 (80, 300, 56.001) in the data routing mapping space are located in the high-priority area and the low-priority area, respectively.

[0127] Next, by step 1045, a double-threshold dynamic partition is set in the routing space: when the priority value Z≥80, it belongs to the first storage level (high-value data that needs multi-node backup), when the priority value Z≥60 and Z<80, it belongs to the second level (medium-value data stored in a single node), and when the priority value Z<60, it belongs to the third level (low-value data settled in the cloud). This grading process is automatically completed without human intervention. For example, F1 corresponds to Z=84.002≥80, and it belongs to the first level. By step 1046, the storage operation is performed according to the level label, the first level data is copied to multiple edge nodes to improve access fault tolerance, the second level data is stored in a single edge node, and the third level data is uploaded to the cloud central storage. The synchronization triggers the data migration process to ensure that the physical location is consistent with the routing strategy, such as F1 being copied and stored to edge stations E1 and E2.

[0128] Then, by steps 1047 and 1048, a distributed index table is created to record the unique identifier, storage level label, and specific physical location coordinates of each data block. The index table is pushed to all edge nodes and cloud nodes through the message queue to realize real-time sharing of network location information and support millisecond-level data positioning of any node. For example, F1 corresponds to the index table record {"F1":"first level, location [E1, E2]"} and is synchronized to all network nodes.

[0129] Finally, by step 1049, when the data access mode changes, the recalculation process is automatically triggered: the hotness attribute value is updated, the priority formula operation is re-executed, and if the new priority value Z value causes level change, the storage location is adjusted and the index is updated, otherwise only the hotness field is refreshed. This mechanism ensures that the storage strategy always matches the real-time access demand. For example, the access volume of F1 increases to 150 times / hour, the hotness value is 150, the distance value remains 150, the priority value Z is recalculated, and the priority value Z is recalculated as 105.002, the level remains unchanged and only the hotness value of the index table is updated.

[0130] In practical applications, for video data block V2, the system first collects its hotness attribute as 150 times / minute based on the recent 10-minute access log statistics, and measures the physical distance from the cloud to edge node E3 as 250 kilometers by a laser range finder; based on the data corresponding to the above-mentioned hotness attribute and distance attribute, the routing priority weight is calculated in the three-dimensional coordinate space: , the coordinate point (150, 250, 105.0012) is generated, and the coordinate (90, 400, 63.00075) is obtained by processing the audio data block A1; all coordinates are integrated to build a routing mapping space; the levels are divided according to the double thresholds: 105.0012 of V2 is greater than or equal to 80, and 63.00075 of A1 is greater than or equal to 60 but less than 80, belonging to the second level; hierarchical storage is performed: V2 is copied to edge nodes E3 and E4, and A1 is stored to a single edge node E5; an index table record {"V2": "first level, location [E3, E4]", "A1": "second level, location [E5]"} is created; the index is synchronized to all 12 nodes in the network through the Pub or Sub message queue; when the access volume of V2 suddenly increases to 220 times per minute, the calculation is re-executed , the Z value is still greater than or equal to 80, but the heat amplitude is detected to be greater than the threshold, triggering storage optimization: a new edge node E6 is added to store a copy and the index is updated to [E3, E4, E6], realizing dynamic load balancing.

[0131] The overall scheme of the above 104, by fusing the heat attribute of high-frequency data and the distance attribute of cloud-edge nodes to build a three-dimensional routing mapping space, realizing intelligent hierarchical storage of data based on dynamic priority weight: high-heat near-edge data is copied to multiple nodes to guarantee low-latency access, medium-frequency data is deployed to a single edge node to balance efficiency and cost, and low-frequency remote data is settled in the cloud to release storage pressure, combined with the distributed index table synchronized in the whole network and the real-time strategy updating mechanism, a cloud-edge collaborative storage system is built, which takes into account access performance, resource utilization and dynamic adaptability, like an intelligent scheduling hub that can perceive data value and location cost for data centers.

[0132] The following is a complete example for steps 101-104, as shown in Figure 2 In the edge node P and cloud data center collaborative scenario, the system collects network jitter rate, edge storage redundancy, and cross-domain access heat of region Q every 5 minutes. The whole day is divided into 15-minute periods, and the total access volume of 8:00-8:15 period is 220 times as the time dimension value; region Q is divided into 1km² grid, and the access density of center grid R5 is 90 times / km² as the space dimension value; the network jitter rate is 125ms as the network state value, and the storage redundancy is 35% as the load value. At 8:00, the feature vector [220, 90, 0.35, 125] is generated, and 12 vectors of 8:00-8:15 period are integrated to form a 12x4 joint feature matrix, realizing environment state panoramic monitoring.

[0133] The network jitter rate sequence [125, 170, 155] ms is extracted from the matrix, the risk threshold 150 ms is set, and the risk constraint set {170:20, 155:5} is generated; the redundancy sequence [0.35, 0.18, 0.22] is extracted, the cost threshold 0.25 is set, and the cost constraint set {0.18:0.07, 0.22:0.03} is generated. The joint optimization target is established , the optimal solution is obtained by genetic algorithm iteration, path S, F(x)=16.8. The continuous over-threshold monitoring point is located: the grid R5 jitter rate is 152, 158 and 161 respectively for 3 times, all greater than the set risk threshold 150 ms, and the space clustering is the risk area T, the average over-limit value is calculated (20+18+16) / 3=18, the risk area backup number is set to 2+18×0.5=11, and the non-risk area copy is compressed to 50%, and the output parameter is {"risk area T":11,"non-risk area":0.5}.

[0134] The 8:00 time value 220 times and the grid R5 space value 90 times / km² are extracted from the matrix, and the cross-domain heat component is calculated as 220×90=19,800. Based on the storage parameters, the grid is taken as the node, the reciprocal of the distance between grids is taken as the edge weight, and the space-time graph is constructed. Then the heat transfer is performed: R5 new heat=0.6×19,800+0.4×adjacent weighted average 15,000=16,680, which converges to 18,500 after 5 rounds of iteration. R5 is determined as the future hotspot according to the prediction heat sorting (18,500>R6's 16,200), and the data with access quantity>100 times in this area is extracted: video file V1 corresponds to 150 times, and log file L1 corresponds to 120 times, and the high-frequency data set {V1, L1} is generated.

[0135] The V1 heat 150 times / minute is extracted, and the cloud-to-edge node P distance is 180 km; the routing priority , which corresponds to the coordinate (150, 180, 105.0017); according to Z≥80, the storage level is divided into the first level; V1 is stored to edge nodes P and Q; a distributed index {"V1":"first level, location [P, Q]", "L1":"second level, location [R]"} is created and synchronized in the whole network. When the V1 access quantity suddenly increases to 210 times / minute, the storage expansion is recalculated , triggering the storage expansion: adding edge node S to store the copy, and updating the index to {"V1":"first level, location [P, Q, S]"}.

[0136] Figure 3 An embodiment of the present application provides a structure diagram of a cloud-edge collaborative data center hierarchical archiving storage optimization system, as shown in Figure 3 , the system comprises:

[0137] The collection module 31 is used for collecting a joint feature matrix according to a network jitter rate collected in a cloud edge network fluctuation scene, an edge storage redundancy and a cross-domain access heat, and the joint feature matrix comprises four dimensions of a time dimension, a space dimension, a load dimension and a network state dimension.

[0138] The generation module 32 is used for mapping the network jitter rate in the network state dimension to a network disconnection risk constraint condition, mapping the edge storage redundancy in the load dimension to a storage cost constraint condition, constructing a joint optimization target according to the network disconnection risk constraint condition and the storage cost constraint condition and solving an optimal solution of the joint optimization target, and positioning a network disconnection risk area according to a network fluctuation parameter of the optimal solution to generate a storage distribution parameter.

[0139] The extraction module 33 is used for generating a cross-domain heat component based on the time dimension and the space dimension of the joint feature matrix, constructing a space-time graph structure in combination with the storage distribution parameter, transferring dynamic access heat through the space-time graph structure to predict a future access hotspot area and extract a high-frequency access data set from the future access hotspot area.

[0140] The storage module 34 is used for extracting a heat attribute from the high-frequency access data set, synchronously acquiring a distance attribute of a cloud storage node and an edge node of a data center, fusing the heat attribute and the distance attribute to construct a data routing mapping space, and performing hierarchical archive storage on the high-frequency access data set in a distributed storage architecture scheduled by the data center and establishing a distributed index based on the data routing mapping space.

[0141] Figure 3 The cloud edge collaborative data center hierarchical archive storage optimization system can perform Figure 1 The cloud edge collaborative data center hierarchical archive storage optimization method of the embodiment has the implementation principle and technical effects which will not be repeated. The specific operation manner of each module and unit of the cloud edge collaborative data center hierarchical archive storage optimization system in the above embodiment has been described in detail in the embodiment related to the method, and will not be described in detail here.

[0142] In one possible design, Figure 3 The cloud edge collaborative data center hierarchical archive storage optimization system of the embodiment can be implemented as a computing device, such as a server. Figure 4 As shown, the computing device can include a storage component 41 and a processing component 42.

[0143] The storage component 41 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 42.

[0144] The processing component 42 is configured to implement various processes and methods described in the present disclosure. For example, the processing component 42 can implement the processes described with respect to the various figures provided herein. Figure 1 The embodiment of the present disclosure provides a cloud-edge collaborative data center hierarchical archival storage optimization method.

[0145] The processing component 42 can include one or more processors to execute computer program instructions to complete all or part of the steps of the methods described above. Of course, the processing component can also be implemented as one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors or other electronic components, for performing the methods described above.

[0146] The storage component 41 is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0147] Of course, the computing device can also include other components, such as an input / output interface, a display component, a communication component, etc.

[0148] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.

[0149] The communication component is configured to facilitate wired or wireless communication between the computing device and other devices, etc.

[0150] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform, and the computing device can be a cloud server, and the processing component, the storage component, etc. can be a basic server resource rented or purchased from the cloud computing platform.

[0151] The embodiment of the present disclosure also provides a computer storage medium storing a computer program, and the computer program is executed by a computer to implement the above Figure 1 The embodiment of the present disclosure provides a cloud-edge collaborative data center hierarchical archival storage optimization method.

[0152] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0153] The apparatus embodiments described above are merely illustrative, wherein the units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0154] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, and the computer software products can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and include a plurality of instructions to make a computer device (which can be a personal computer, a server, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0155] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A cloud-edge collaborative data center hierarchical archival storage optimization method, characterized in that, The method comprises the steps of: According to the network jitter rate collected by the cloud edge network fluctuation scene, the edge storage redundancy and the cross-domain access heat, a joint feature matrix is constructed, which contains four dimensions of time dimension, space dimension, load dimension and network state dimension; Based on the joint feature matrix, the network jitter rate of the network state dimension is mapped to the network risk constraint condition, the edge storage redundancy of the load dimension is mapped to the storage cost constraint condition, the joint optimization objective is constructed according to the network risk constraint condition and the storage cost constraint condition, and the optimal solution of the joint optimization objective is solved, and the network fluctuation parameter of the optimal solution is used to locate the disconnection risk area to generate the storage distribution parameter; Based on the time dimension and space dimension of the joint feature matrix, a cross-domain heat component is generated, a space-time graph structure is constructed by combining the storage distribution parameter, and dynamic access heat is transmitted through the space-time graph structure to predict future access hotspot areas and extract high-frequency access data sets from the future access hotspot areas; The heat attribute is extracted from the high-frequency access data set, the distance attribute of the data center cloud storage node and the edge node is synchronously acquired, the heat attribute and the distance attribute are fused to construct a data routing mapping space, and the high-frequency access data set is stored and distributed in the distributed storage architecture of the data center based on the data routing mapping space. The method comprises the steps of: Extracting the network jitter rate from the network state dimension of the joint feature matrix and setting a network risk threshold, and mapping the network jitter rate exceeding the network risk threshold to the network risk constraint condition; Extracting the edge storage redundancy from the load dimension of the joint feature matrix and setting a redundancy cost threshold, and mapping the edge storage redundancy lower than the redundancy cost threshold to the storage cost constraint condition; Establishing a joint optimization objective containing the network risk constraint condition and the storage cost constraint condition; Adjusting the network fluctuation parameter by an iterative optimization method to minimize the value of the joint optimization objective to obtain the optimal solution of the joint optimization objective; According to the position of the area where the network jitter rate is continuously higher than the network risk threshold in the optimal solution, the disconnection risk area is located, and the storage distribution parameter containing the risk area data backup number and the non-risk area data copy compression ratio is calculated and generated according to the disconnection risk area.

2. The method of claim 1, wherein, The method comprises the steps of: Based on the time dimension and space dimension of the joint feature matrix, a cross-domain heat component is generated, a space-time graph structure is constructed by combining the storage distribution parameter, and dynamic access heat is transmitted through the space-time graph structure to predict future access hotspot areas and extract high-frequency access data sets from the future access hotspot areas; extracting time-dimension feature values and space-dimension feature values from the joint feature matrix, multiplying the time-dimension feature values and the space-dimension feature values at the same time point to generate a cross-domain heat component; constructing a space-time graph structure based on the storage distribution parameters, taking a physical space grid as a graph node in the space-time graph structure, taking the cross-domain heat component as a node attribute value, and taking an inverse of a physical distance between grids as an edge weight value; performing a dynamic access heat transfer operation through the space-time graph structure, fusing a weighted result of a current node access heat value and a neighboring node access heat mean value to update the current node access heat value; repeating the dynamic access heat transfer operation until an access heat value update amplitude of all nodes is less than a set convergence condition, and outputting predicted access heats of each grid in a future time window; selecting a number of grid regions with the highest predicted access heats as future access hotspot regions, and extracting a data set with an access frequency greater than a set threshold from a storage distribution corresponding to the future access hotspot regions as a high-frequency access data set.

3. The method of claim 1, wherein, extracting, from the high-frequency access data set, an access frequency of each data block in a unit time as a heat attribute, and simultaneously obtaining a distance attribute of a data center cloud storage node and an edge node, and fusing the heat attribute and the distance attribute to construct a data routing mapping space, including: extracting, from the high-frequency access data set, an access frequency of each data block in a unit time as a heat attribute, and simultaneously obtaining a distance attribute of a data center cloud storage node and an edge node, and fusing the heat attribute and the distance attribute to construct a data routing mapping space, including: establishing a three-dimensional coordinate space composed of three coordinate axes, in which a first coordinate axis corresponds to the heat attribute, a second coordinate axis corresponds to the distance attribute, and a third coordinate axis corresponds to a routing priority weight; performing an attribute fusion operation on each data block, performing weighted calculation on the heat attribute and the distance attribute through a preset heat coefficient and a distance coefficient to generate a routing priority weight value, and mapping the routing priority weight value to the third coordinate axis of the three-dimensional coordinate space; taking a three-dimensional coordinate point set of all data blocks as a whole structure to constitute the data routing mapping space.

4. The method of claim 1, wherein, locating a disconnection risk region according to a region position in which a network jitter rate in the optimal solution is continuously higher than the network disconnection risk threshold, and calculating and generating storage distribution parameters including a number of data backup copies of the risk region and a data copy compression ratio of a non-risk region, including: extracting, from the optimal solution, a monitoring point in which the network jitter rate continuously exceeds a preset time window to form an initial risk point set; spatially clustering geographically adjacent monitoring points in the initial risk point set to generate a disconnection risk region with a continuous boundary; calculating a difference between an average network jitter rate of each disconnection risk region and the network disconnection risk threshold as an average jitter overrun value; setting a number of data backup copies of the risk region based on the average jitter overrun value, and setting a data copy compression ratio of the non-risk region; generating storage distribution parameters according to the number of data backup copies and the data copy compression ratio.

5. The method of claim 1, wherein, The hierarchical archiving storage and distributed indexing of the high-frequency access dataset in the data center scheduled distributed storage architecture based on the data routing mapping space comprises: Three storage levels are divided in the data routing mapping space according to routing priority weight values, the first storage level in the storage levels contains data blocks with routing priority weight values not less than a first preset threshold, the second storage level contains data blocks with routing priority weight values less than the first preset threshold but not less than a second preset threshold, and the third storage level contains data blocks with routing priority weight values less than the second preset threshold; Hierarchical storage operations are performed in the data center scheduled distributed storage architecture, the data blocks in the first storage level are stored in the edge storage area of an edge node, the data blocks in the second storage level are stored in the edge storage area of a single edge node, and the data blocks in the third storage level are uploaded to a cloud storage node; A distributed index table is created, which records the unique identifier of each data block, the storage level to which the data block belongs, and the specific storage location coordinates; A complete copy of the distributed index table is synchronously stored in all edge nodes and cloud storage nodes; When the heat attribute value of a data block in the high-frequency access dataset changes, the routing priority weight value of the data block in the data routing mapping space is recalculated and the distributed index table is updated.

6. The method of claim 1, wherein, The network jitter rate, edge storage redundancy and cross-domain access heat collected according to the cloud edge network fluctuation scenario are used to construct a joint feature matrix, and the joint feature matrix contains four dimensions of time dimension, space dimension, load dimension and network state dimension, comprising: The network jitter rate, edge storage redundancy and cross-domain access heat under the cloud edge network fluctuation scenario are collected; A continuous time period is divided into equal length segments, the cumulative time distribution of the cross-domain access heat is counted in each time segment to generate time dimension feature values, at the same time, the physical space is divided into grid regions, the geographic distribution density of the cross-domain access heat is calculated in each grid region to generate space dimension feature values, and the network jitter rate is directly used as the network state dimension feature value, and the edge storage redundancy is directly used as the load dimension feature value; The time dimension feature values, space dimension feature values, load dimension feature values and network state dimension feature values at the same time are combined into four-dimensional feature vectors in time sequence; All the four-dimensional feature vectors at all times are arranged in time sequence to form a joint feature matrix with rows and columns aligned, and the rows of the joint feature matrix correspond to time points, and the columns correspond to four-dimensional feature values.

7. A cloud-edge collaborative data center hierarchical archival storage optimization system, characterized in that, Comprise: The acquisition module is used for collecting the network jitter rate, edge storage redundancy and cross-domain access heat according to the cloud edge network fluctuation scenario, and constructing a joint feature matrix, and the joint feature matrix contains four dimensions of time dimension, space dimension, load dimension and network state dimension. The generating module is configured to map, based on the joint feature matrix, the network jitter rate of the network state dimension to a network disconnection risk constraint condition, map the edge storage redundancy of the load dimension to a storage cost constraint condition, construct a joint optimization objective according to the network disconnection risk constraint condition and the storage cost constraint condition, and solve an optimal solution of the joint optimization objective, and locate a disconnection risk region according to a network fluctuation parameter of the optimal solution to generate a storage distribution parameter. The extracting module is configured to generate a cross-domain heat component based on the time dimension and the space dimension of the joint feature matrix, construct a spatio-temporal graph structure in combination with the storage distribution parameter, transfer dynamic access heat through the spatio-temporal graph structure, predict a future access hotspot region, and extract a high-frequency access data set from the future access hotspot region. The storing module is configured to extract a heat attribute from the high-frequency access data set, synchronously acquire a distance attribute of a data center cloud storage node and an edge node, fuse the heat attribute and the distance attribute to construct a data routing mapping space, and perform hierarchical archival storage on the high-frequency access data set and establish a distributed index in a distributed storage architecture of the data center based on the data routing mapping space. The mapping, based on the joint feature matrix, of the network jitter rate of the network state dimension to a network disconnection risk constraint condition, the mapping of the edge storage redundancy of the load dimension to a storage cost constraint condition, the construction of a joint optimization objective according to the network disconnection risk constraint condition and the storage cost constraint condition, the solving of an optimal solution of the joint optimization objective, and the locating of a disconnection risk region according to a network fluctuation parameter of the optimal solution to generate a storage distribution parameter, include: extracting a network jitter rate from the network state dimension of the joint feature matrix and setting a network disconnection risk threshold, and mapping a network jitter rate higher than the network disconnection risk threshold to a network disconnection risk constraint condition; extracting an edge storage redundancy from the load dimension of the joint feature matrix and setting a redundancy cost threshold, and mapping an edge storage redundancy lower than the redundancy cost threshold to a storage cost constraint condition; establishing a joint optimization objective containing the network disconnection risk constraint condition and the storage cost constraint condition; adjusting a network fluctuation parameter by an iterative optimization method to minimize a value of the joint optimization objective, so as to obtain an optimal solution of the joint optimization objective; locating a disconnection risk region according to a region position where the network jitter rate is continuously higher than the network disconnection risk threshold in the optimal solution, and calculating and generating a storage distribution parameter containing a risk region data backup number and a non-risk region data copy compression ratio according to the disconnection risk region.

8. A computing device, comprising: The storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the cloud-edge collaborative data center hierarchical archival storage optimization method of any one of claims 1 to 6.

9. A computer storage medium, characterized in that, The computer program is stored in the computer and is executed by the computer to implement the cloud-edge collaborative data center hierarchical archival storage optimization method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Optical fiber data storage management system and method based on big data

    CN120085812A

  • Meteorological metadata storage method and system based on machine learning

    CN120104579A