Data storage path dynamic optimization method based on edge calculation
By analyzing the storage node feature data and data feature information in the edge gateway, generating data migration strategies, and dynamically optimizing the data storage path, the problem that static configuration in the edge computing environment cannot adapt to dynamic access features is solved, and efficient and reliable data access and storage resource utilization are achieved.
Patent Information
- Application Number
- CN202510428037.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In an edge computing environment, traditional static storage path configuration cannot adapt to dynamically changing data access characteristics, resulting in unbalanced load of storage nodes and affecting data access efficiency.
By receiving node feature data of storage nodes in the edge gateway, performing multi-dimensional performance analysis, identifying nodes to be optimized, combining data feature information for hierarchical analysis, accurately locate the data to be optimized, and generating data migration strategies to dynamically optimize the data storage path.
It realizes intelligent and adaptive optimization of data storage paths, reduces cloud load, improves real-time and reliability of data access, and optimizes the utilization rate of storage resources.
Smart Images

Figure CN120104066A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data storage management technology, and more specifically, to a method for dynamically optimizing a data storage path based on edge computing. Background Art
[0002] With the rapid development of big data and cloud computing, the demand for data storage and access is growing rapidly. Traditional centralized storage architectures can no longer meet the requirements of large-scale data in terms of real-time, reliability and access efficiency. In order to alleviate the computing pressure in the cloud, edge computing has emerged as an emerging technology, which sinks computing and storage resources to the edge of the network to achieve local storage and processing of data, thereby improving access efficiency and reducing latency. However, in the edge environment, the static configuration of storage paths can no longer adapt to the complex and changeable data access characteristics. Due to the dynamic changes in factors such as data access frequency, storage node load and network bandwidth, fixed data storage paths are prone to: high-access data is distributed in low-performance or high-load nodes, affecting access efficiency; the load between storage nodes is unbalanced, some nodes have idle resources, and some nodes are under too much pressure. Therefore, an intelligent dynamic optimization method for data storage paths is urgently needed.
[0003] The patent with publication number CN118349359A discloses an edge computing method based on big data; it includes: allocating preprocessed data to a local or distributed storage system, and optimizing storage according to the importance and access frequency of the data; using the ant colony algorithm to dynamically optimize the location and movement path of edge nodes, which not only improves resource utilization efficiency but also reduces data transmission delay; by establishing a comprehensive objective function including resource utilization, data transmission delay and system performance, and setting constraints, training the edge computing model, the system can be self-optimized and adapt to the ever-changing data flow requirements, thereby improving data processing efficiency and response speed in the edge computing environment.
[0004] However, although the above-mentioned technology realizes the dynamic optimization of data storage paths, it focuses on the location optimization of edge nodes and the data storage distribution strategy, and mainly adjusts the deployment of edge computing nodes and data flow paths through the ant colony algorithm, but does not involve the dynamic migration mechanism of data between storage nodes; therefore, when the storage node load is uneven or the data access frequency changes, the above-mentioned technology cannot adaptively adjust the data storage path, and it is difficult to meet the data storage optimization needs in complex edge computing environments.
[0005] In view of this, the present invention proposes a data storage path dynamic optimization method based on edge computing to solve the above problems. Summary of the invention
[0006] In order to overcome the above-mentioned defects of the prior art and achieve the above-mentioned purpose, the present invention provides the following technical solution: a data storage path dynamic optimization method based on edge computing, applied to an edge gateway, comprising: S1: Receive node feature data sent by the storage node; S2: Perform multi-dimensional performance analysis on node feature data and use anomaly detection algorithms to identify nodes to be optimized in storage nodes; S3: receiving data feature information sent by the node to be optimized; S4: Perform hierarchical analysis on data feature information and combine hotspot data identification technology to accurately locate the data to be optimized in the node to be optimized; S5: Analyze the node feature data, generate a data migration strategy, and determine whether the migration node in the data migration strategy is an adaptation node. If the migration node is an adaptation node, dynamically optimize the storage path of the data to be optimized according to the data migration strategy; S6: If the migration node is not an adapter node, the global feature data sent by the cloud platform is received, the global feature data is comprehensively analyzed, the data migration strategy is optimized, and the storage path of the data to be optimized is dynamically optimized according to the optimized data migration strategy.
[0007] Furthermore, the node characteristic data includes storage capacity status, read / write performance status, load status and network resource status; wherein the storage capacity status includes total storage capacity and used capacity, the read / write performance status includes I / O read / write speed and IOPS, the load status includes CPU usage and memory usage, and the network resource status includes network bandwidth and network delay; The method for identifying a node to be optimized among storage nodes includes: Divide the used capacity in each set of node feature data by the corresponding total storage capacity to obtain the capacity utilization rate corresponding to each set of node feature data; replace the storage capacity state of each set of node feature data with the corresponding capacity utilization rate, and mark the replaced node feature data as replacement feature data; mark each set of replacement feature data obtained in real time as real-time feature data, and obtain historical feature data, which is the replacement feature data of normal nodes obtained at historical moments, and normal nodes are storage nodes that are not marked as nodes to be optimized; each data in each set of real-time feature data is respectively taken as a set of data sets with the same type of data in the historical feature data, and the data sets correspond to the data in the real-time feature data one by one; calculate the abnormal coefficient and coefficient threshold corresponding to each set of data sets in turn, and compare the abnormal coefficient corresponding to each set of data sets with the corresponding coefficient threshold; mark the data sets with abnormal coefficients greater than or equal to the coefficient threshold as abnormal sets, and do not mark the data sets with abnormal coefficients less than the coefficient threshold; mark the data in the real-time feature data corresponding to each set of abnormal sets as abnormal data, mark the real-time feature data corresponding to the abnormal data as abnormal feature data, and take the storage nodes corresponding to the abnormal feature data as nodes to be optimized.
[0008] Furthermore, the method for calculating the anomaly coefficient corresponding to the data set includes: Calculate the Euclidean distance between each two data in the data set in turn and mark them as data distances; sort the data distances corresponding to each data from small to large to generate a distance sorting table corresponding to each data; sort the data at the front in each distance sorting table The data distance of each is taken as the adjacent distance of the corresponding data. is an integer greater than 1; add the adjacent distances of each data in turn and then divide by , obtain the average distance corresponding to each data; mark the data corresponding to the real-time feature data in the data set as real-time data, and use the average distance of the real-time data as the abnormal coefficient of the data set; Methods for calculating coefficient thresholds corresponding to a data set include Count the number of data in the data set and mark it as the number of data; add the average distances corresponding to the data set in sequence, and then divide it by the number of data to obtain the overall mean corresponding to the data set; subtract the overall mean from the average distance of each data, and then square it to obtain the square of the difference of each data; add the square of the difference of each data in sequence and divide it by the number of data, and then take the square root to obtain the overall dispersion; multiply the overall dispersion by , plus the overall mean, as the coefficient threshold of the data set, .
[0009] Furthermore, the data characteristic information includes write frequency, read frequency, data size and access time, and the access time is the time when the data was last accessed; The method for locating the data to be optimized in the node to be optimized includes: Obtain the current time, mark all data in the node to be optimized as data to be analyzed, subtract the current time from the access time of each data to be analyzed, and take the absolute value to obtain the time difference of each data to be analyzed; preset a weight set, the weight set includes the write frequency, the read frequency and the weight coefficient corresponding to the time difference; standardize the write frequency, the read frequency and the time difference of each data to be analyzed, respectively, to obtain standardized data, the standardized data includes the standard write frequency, the standard read frequency and the standard time difference; multiply the reciprocal of the standard write frequency, the standard read frequency and the standard time difference of each data to be analyzed by the corresponding weight coefficient, and add them in sequence to obtain the heat value of each data to be analyzed; Each data to be analyzed is hierarchically classified to obtain the hierarchical type of each data to be analyzed; according to the hierarchical type of each data to be analyzed, the type label of each data to be analyzed is obtained, the type label is a digital label corresponding to the hierarchical type, and the hierarchical labels corresponding to different hierarchical types are all different; the type label of each data to be analyzed and the node feature data of the corresponding node to be optimized are taken as a group of analysis data, and the analysis data corresponds to the data to be analyzed one by one; each group of analysis data is input into the trained migration analysis model respectively, and the corresponding judgment label is predicted, the judgment label is the digital label corresponding to the migration judgment result, and the migration judgment result includes migration and retention, and migration and retention correspond to different judgment labels; according to the judgment label, the migration judgment result corresponding to each group of analysis data is obtained, if the migration judgment result is migration, then the data to be analyzed corresponding to the corresponding analysis data is marked as data to be optimized, if the migration judgment result is retention, then the data to be analyzed corresponding to the corresponding analysis data is not marked.
[0010] Furthermore, the method for obtaining the hierarchical type of each data to be analyzed includes: The write frequency and the read frequency of each data to be analyzed are added to obtain the access frequency of each data to be analyzed; a classification standard set is preset, and the classification standard set includes three classification sets, which correspond to the access frequency, heat value and data size respectively; wherein each classification set includes a first threshold and a second threshold, and the first threshold is greater than the second threshold; based on the preset classification rules, the access frequency, heat value and data size of each data to be analyzed are compared with the corresponding classification sets respectively to obtain the frequency type, heat type and size type corresponding to each data to be analyzed; the frequency type, heat type and size type of each data to be analyzed are combined to obtain the hierarchical type of each data to be analyzed.
[0011] Furthermore, the method for generating a data migration strategy includes: Build Group migration node set, each group migration node set includes Normal nodes; among them, is an integer greater than 1, is the number of data to be optimized; set an increasing numeric label for each set of migration nodes and mark it as a migration label. The range of the migration label is ; Randomly generated within the range of migration labels Candidate solutions, There is a one-to-one correspondence between candidate solutions and transition labels. ; Define the iterative process: for each candidate solution, two intervals are randomly generated The random coefficients between the candidate solutions are compared with the preset judgment coefficients to determine the update behavior of each candidate solution. The range of the judgment coefficients is ; Update each candidate solution according to the update behavior and obtain The best solution among the candidate solutions; The iterative process is executed cyclically. When the number of executions of the iterative process is greater than or equal to a preset iteration threshold, the iterative process is stopped and the best solution is obtained. The migration node set corresponding to the best solution is marked as the best set, and the best set is used as the data migration strategy.
[0012] Furthermore, construct Methods for group migration of node sets include: Randomly select from all normal nodes Normal nodes, according to Normal nodes build a set of migration nodes, Group migration node set, each group migration node set There is a one-to-one correspondence between normal nodes and the data to be optimized; Methods for determining update behavior include: Mark one of the random coefficients as the first coefficient and the other random coefficient as the second coefficient; compare the first coefficient and the second coefficient with the judgment coefficient respectively; if the first coefficient is less than the judgment coefficient, determine that the update behavior is an unknown update; if the first coefficient is greater than or equal to the judgment coefficient, and the second coefficient is less than the judgment coefficient, determine that the update behavior is a current update; if both the first coefficient and the second coefficient are greater than or equal to the judgment coefficient, determine that the update behavior is a known update; Methods for updating candidate solutions based on update behaviors include: When the update behavior is unknown update, from the interval A random unknown coefficient is generated in Subtract one, multiply by the unknown coefficient, and add one to obtain the unknown update value; update the candidate solution according to the unknown update value; When the update behavior is current update, from the interval A current coefficient is randomly generated, and the current coefficient is multiplied by the random solution to obtain a random update value, and the random solution is a random candidate solution; the current coefficient is subtracted by one, and then multiplied by the candidate solution to obtain a candidate update value; the random update value is added to the candidate update value to obtain the current update value; the candidate solution is updated according to the current update value; When the update behavior is a known update, from the interval A known coefficient is randomly generated, and the known coefficient is subtracted from one and multiplied by the candidate solution to obtain a candidate update value; the known coefficient is multiplied by the previous solution to obtain a previous update value, and the previous solution is the candidate solution in the previous iteration; the known coefficient is multiplied by the random solution to obtain a followed update value; the candidate update value, the previous update value and the followed update value are added in sequence to obtain a known update value; the candidate solution is updated according to the known update value.
[0013] Further, obtain Methods for finding the best solution among candidate solutions include: The migration node sets corresponding to the migration labels of each candidate solution are marked as the current set, and the normal nodes in each current set are obtained and marked as the current nodes; the access frequency, heat value and data size of each data to be optimized are taken as a set of parameters to be optimized, and the parameters to be optimized correspond to the data to be optimized one by one; the node feature data of the current node corresponding to each candidate solution, the node feature data of all nodes to be optimized and the parameters to be optimized of all data to be optimized are taken as evaluation data, and the evaluation data correspond to the candidate solutions one by one; each set of evaluation data is input into the trained benefit evaluation model respectively to evaluate the corresponding migration benefit ratio; the training process of the benefit evaluation model is consistent with the training process of the migration analysis model, and both are deep neural network models; the migration benefit ratios corresponding to each candidate solution are compared respectively, and the candidate solution with the largest migration benefit ratio is marked as the best solution.
[0014] Furthermore, the method for determining whether the migration node in the data migration strategy is an adaptation node includes: Obtain the storage nodes in the data migration strategy and mark them as migration nodes; obtain the data to be optimized corresponding to each migration node, and use the node feature data of each migration node and the parameters to be optimized corresponding to the data to be optimized as a set of prediction data, and the prediction data corresponds to the migration node one-to-one; input each set of prediction data into the trained feature prediction model to predict the corresponding migration feature data; the migration feature data is the node feature data corresponding to the migration node after the data to be optimized is migrated to the corresponding migration node; the training process of the feature prediction model is consistent with the training process of the migration analysis model, and both are deep neural network models; use the type label corresponding to each data to be optimized and the migration feature data of the corresponding migration node as a set of judgment data, and the judgment data corresponds to the data to be optimized one-to-one; input each set of judgment data into the trained migration analysis model to predict the corresponding judgment label and mark it as the prediction label; obtain the migration judgment result corresponding to each migration node according to the prediction label, if the migration judgment result is all retained, the corresponding migration node is marked as an adaptation node, if there is migration in the migration judgment result, the corresponding migration node is not marked.
[0015] Furthermore, the global feature data includes node feature data of all storage nodes; based on the global feature data, the migration node set is regenerated, and the data migration strategy is regenerated based on the regenerated migration node set; based on the regenerated data migration strategy, the data migration strategy generated in S5 is optimized.
[0016] The technical effects and advantages of the data storage path dynamic optimization method based on edge computing of the present invention are as follows: By analyzing the node feature data of storage nodes in real time, the anomaly detection algorithm is used to intelligently identify the storage nodes that need to be optimized; in addition, the data feature information of the data in the nodes to be optimized is analyzed to accurately locate the data to be migrated; at the same time, based on the evaluation of migration benefits, data migration strategies are independently generated, and the migration of data between storage nodes is adaptively adjusted to achieve dynamic optimization of data storage paths; in addition, when the edge gateway cannot meet the optimization requirements, the data migration strategy can be further optimized with the help of the global perspective of the cloud, and the synergy between the edge and the cloud can be used to improve data storage efficiency and access performance; giving full play to the advantages of edge computing to achieve intelligent and adaptive optimization of data storage paths, which helps to reduce cloud load, improve the real-time and reliability of data access, and optimize the utilization of storage resources at the same time, solving the problem that traditional static configuration cannot adapt to dynamic access characteristics, thereby significantly improving data management capabilities and overall performance in edge computing environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a flow chart of a method for dynamically optimizing data storage paths based on edge computing according to Example 1 of the present invention. DETAILED DESCRIPTION
[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] Example 1 See also Figure 1 As shown, the data storage path dynamic optimization method based on edge computing described in this embodiment includes: S1: The storage node collects node feature data and sends it to the edge gateway.
[0020] Node characteristic data includes storage capacity status, read / write performance status, load status and network resource status; among them, storage capacity status includes total storage capacity and used capacity, read / write performance status includes I / O read / write speed and IOPS, load status includes CPU usage and memory usage, network resource status includes network bandwidth and network delay; total storage capacity is the maximum storage space of the storage node, and used capacity is the occupied storage space on the storage node; I / O read / write speed is the speed at which the storage node reads or writes data per unit time, and IOPS is the number of input / output operations that the storage node can perform per second; CPU usage is the proportion of CPU resources occupied by the storage node, and memory usage is the ratio of the used memory of the storage node to the total available memory; network bandwidth is the maximum data transmission rate between the storage node and the edge gateway, and network delay is the time required for data to be transmitted from the storage node to the edge gateway.
[0021] The storage capacity status is used to determine the usage of storage resources in the storage node, which helps to identify storage nodes with insufficient or redundant storage capacity and decide whether data needs to be migrated to other nodes. When the used capacity of a storage node approaches saturation, some of the data needs to be migrated to other storage nodes to prevent a single storage node from overflowing and to ensure data writing and reading efficiency.
[0022] The read and write performance status is used to evaluate the read and write performance and carrying capacity of the storage node, which helps to identify the storage performance bottleneck of the storage node and determine whether it is necessary to migrate high-frequency access data to a higher-performance storage node. When the I / O read and write speed continues to decrease, it indicates that the storage node has a performance bottleneck or the media is aging. It is necessary to migrate high-frequency access data to a storage node with a higher I / O read and write speed to improve access efficiency. When the IOPS is close to the upper limit, it indicates that the storage performance of the storage node is saturated, resulting in access delays. It is necessary to migrate low-frequency access data to a node with lower IOPS to optimize resource utilization efficiency.
[0023] The load status is used to determine the computing load of the storage node, which helps to identify storage nodes with overloaded or idle computing resources and assist in optimizing data storage paths. When the CPU usage or memory usage is high, it means that the storage node is processing a large number of read and write requests, computing resources are insufficient, and the load is too high, affecting data read and write performance. It is necessary to migrate high-frequency access data to low-load storage nodes to alleviate computing resource pressure.
[0024] The network resource status is used to evaluate the data transmission performance between the storage node and the edge gateway, which helps to determine network congestion or access bottlenecks and decide whether data needs to be migrated to a storage node with higher bandwidth or lower latency. When the network bandwidth decreases or the network latency increases, it means that the data transmission speed and data access response speed of the storage node decrease, affecting the read and write performance. High-frequency access data needs to be migrated to storage nodes with lower network latency and higher network bandwidth to improve access efficiency. Low-frequency access data needs to be migrated to storage nodes with lower bandwidth to reduce the occupancy of core storage resources.
[0025] The storage node is a device or server responsible for data storage and management, with specific independent storage calculations (such as disk space), computing resources (including CPU and memory) and network interfaces; the edge gateway is a computing and communication device located between the storage node and the cloud platform, responsible for local processing, analysis and management of data, and plays a role in data filtering, processing and decision-making; the cloud platform is a remote data storage and processing system based on cloud computing architecture, with powerful computing, storage and management capabilities, responsible for global data management, cross-regional data migration, remote monitoring and resource scheduling; storage nodes, edge gateways and cloud platforms together constitute the edge computing architecture, and storage nodes and edge gateways usually use TCP / IP protocol or efficient data transmission protocol (such as gRPC, HTTP / HTTPS, MQTT, etc.) for data transmission; among them, the cloud platform is usually connected to multiple edge gateways, and one edge gateway is usually connected to multiple storage nodes to form a star or mesh topology. Under this edge computing architecture, data is dynamically migrated between storage nodes, which is completed by the edge gateway and the cloud platform in collaboration to ensure dynamic optimization of data storage paths.
[0026] S2: The edge gateway performs multi-dimensional performance analysis on the node feature data and uses anomaly detection algorithms to identify nodes to be optimized in the storage nodes.
[0027] The method for identifying a node to be optimized among storage nodes includes: Divide the used capacity in each set of node feature data by the corresponding total storage capacity to obtain the capacity utilization rate corresponding to each set of node feature data; replace the storage capacity status of each set of node feature data with the corresponding capacity utilization rate, and mark the replaced node feature data as replacement feature data; mark each set of replacement feature data obtained in real time as real-time feature data, and obtain historical feature data. The historical feature data is the replacement feature data of normal nodes obtained at historical moments. Normal nodes are storage nodes that are not marked as nodes to be optimized. The historical feature data is obtained through the built-in database of the edge gateway; divide each data in each set of real-time feature data into The data of the same type as the historical feature data are taken as a group of data sets, and the data sets correspond to the data in the real-time feature data one by one; the abnormal coefficient and coefficient threshold corresponding to each group of data sets are calculated in turn, and the abnormal coefficient corresponding to each group of data sets is compared with the corresponding coefficient threshold; the data sets with abnormal coefficients greater than or equal to the coefficient threshold are marked as abnormal sets, and the data sets with abnormal coefficients less than the coefficient threshold are not marked; the data in the real-time feature data corresponding to each abnormal set are marked as abnormal data, the real-time feature data corresponding to the abnormal data are marked as abnormal feature data, and the storage nodes corresponding to the abnormal feature data are taken as nodes to be optimized.
[0028] Methods for calculating the anomaly coefficient corresponding to a data set include: Calculate the Euclidean distance between each two data in the data set in turn and mark them as data distance; the calculation method of Euclidean distance is an existing technology, and the specific process is not repeated here; sort the data distance corresponding to each data from small to large, and generate a distance sorting table corresponding to each data; sort the data at the front in each distance sorting table The data distance of each is taken as the adjacent distance of the corresponding data. is an integer greater than 1, The specific value of is preset by those skilled in the art according to the actual situation; the adjacent distances of each data are added in sequence and then divided by , obtain the average distance corresponding to each data; mark the data corresponding to the real-time feature data in the data set as real-time data, and use the average distance of the real-time data as the abnormal coefficient of the data set.
[0029] Methods for calculating coefficient thresholds corresponding to a data set include Count the number of data in the data set and mark it as the number of data; add the average distances corresponding to the data set in sequence, and then divide it by the number of data to obtain the overall mean corresponding to the data set; subtract the overall mean from the average distance of each data, and then square it to obtain the square of the difference of each data; add the square of the difference of each data in sequence and divide it by the number of data, and then take the square root to obtain the overall dispersion; multiply the overall dispersion by , plus the overall mean, as the coefficient threshold of the data set, .
[0030] S3: The node to be optimized collects data feature information and sends it to the edge gateway.
[0031] Data characteristic information includes write frequency, read frequency, data size, and access time; The write frequency is the number of times data is written per unit time. The higher the write frequency, the more frequently the data is updated, and it is suitable for storage in storage nodes with high write performance. The read frequency is the number of times data is read per unit time. The higher the read frequency, the more frequently the data is accessed, and it is suitable for storage in low-latency, high-bandwidth storage nodes. The data size is the space occupied by the data in the storage node. The larger the data size, the more storage resources the data occupies, and it is suitable for storage in storage nodes with low capacity utilization. The access time is the last time the data was accessed, which is used to reflect the popularity of the data. The closer the access time is to the current time, the more frequently the data has been accessed recently, and it is suitable for storage in high-performance storage nodes (i.e., fast I / O read and write speeds, high CPU utilization, high network bandwidth, etc.).
[0032] S4: The edge gateway performs hierarchical analysis on the data feature information and combines it with hot data identification technology to accurately locate the data to be optimized in the node to be optimized.
[0033] The method for locating the data to be optimized in the node to be optimized includes: Get the current time, which is the time value of the current moment, obtained through the built-in clock of the edge gateway; mark all the data in the node to be optimized as data to be analyzed, subtract the current time from the access time of each data to be analyzed, and take the absolute value to obtain the time difference of each data to be analyzed; preset a weight set, which includes the weight coefficients corresponding to the write frequency, read frequency and time difference, and the weight set is pre-set by technical personnel in this field according to the actual situation; perform standardization processing (such as minimum-maximum standardization, Z-Score standardization, etc.) on the write frequency, read frequency and time difference of each data to be analyzed, and obtain standardized data , the standardized data includes the standard write frequency, the standard read frequency and the standard time difference; the standard write frequency, the standard read frequency and the inverse of the standard time difference of each data to be analyzed are multiplied by the corresponding weight coefficient respectively, and added in sequence to obtain the heat value of each data to be analyzed; it should be understood that the higher the standardized write frequency and read frequency, the higher the frequency at which the corresponding data is written and read equivalent to other data, the smaller the standardized time difference, the larger the inverse of the time difference, indicating that the corresponding data has been accessed recently; therefore, the larger the three factors related to the heat value, the higher the heat of the data, that is, the larger the heat value, and vice versa; Each data to be analyzed is hierarchically classified to obtain the hierarchical type of each data to be analyzed; according to the hierarchical type of each data to be analyzed, the type label of each data to be analyzed is obtained, the type label is a digital label corresponding to the hierarchical type, and the hierarchical labels corresponding to different hierarchical types are all different; the type label of each data to be analyzed and the node feature data of the corresponding node to be optimized are taken as a group of analysis data, and the analysis data corresponds to the data to be analyzed one by one; each group of analysis data is input into the trained migration analysis model respectively, and the corresponding judgment label is predicted, the judgment label is the digital label corresponding to the migration judgment result, and the migration judgment result includes migration and retention, and migration and retention correspond to different judgment labels; according to the judgment label, the migration judgment result corresponding to each group of analysis data is obtained, if the migration judgment result is migration, then the data to be analyzed corresponding to the corresponding analysis data is marked as data to be optimized, if the migration judgment result is retention, then the data to be analyzed corresponding to the corresponding analysis data is not marked.
[0034] The methods for obtaining the hierarchical type of each data to be analyzed include: The write frequency and the read frequency of each data to be analyzed are added to obtain the access frequency of each data to be analyzed; a classification standard set is preset, and the classification standard set includes three classification sets, which correspond to the access frequency, the heat value and the data size respectively; wherein each classification set includes a first threshold and a second threshold, and the first threshold is greater than the second threshold; based on the preset classification rules, the access frequency, the heat value and the data size of each data to be analyzed are respectively compared with the corresponding classification sets to obtain the frequency type, the heat type and the size type corresponding to each data to be analyzed; the frequency type, the heat type and the size type of each data to be analyzed are combined to obtain the hierarchical type of each data to be analyzed; the classification standard set is preset by a person skilled in the art according to actual conditions; The classification rules are: if the access frequency is greater than or equal to the first threshold, the corresponding frequency type is high frequency; if the access frequency is less than the first threshold and greater than the second threshold, the corresponding frequency type is medium frequency; if the access frequency is less than or equal to the second threshold, the corresponding frequency type is low frequency; if the heat value is greater than or equal to the first threshold, the corresponding heat type is high heat; if the heat value is less than the first threshold and greater than the second threshold, the corresponding heat type is medium heat; if the heat value is less than or equal to the second threshold, the corresponding heat type is low heat; if the data size is greater than or equal to the first threshold, the corresponding size type is big data; if the data size is less than the first threshold and greater than the second threshold, the corresponding size type is medium data; if the data size is less than or equal to the second threshold, the corresponding size type is small data; exemplarily, the data to be analyzed corresponds to high frequency, high heat and medium data, and the corresponding hierarchical type is high frequency-high heat-medium data.
[0035] The training process of the migration analysis model includes: Pre-collection Different sets of analytical data Each group of analysis data is assigned a corresponding judgment label. is an integer greater than 1, converting the analysis data and the corresponding judgment labels into a corresponding set of feature vectors; the judgment labels corresponding to the analysis data are collected by technicians in the process of historically locating the data to be optimized. Different groups of analysis data are analyzed, and each group of analysis data is analyzed in turn according to the actual situation to determine whether the data to be analyzed in the analysis data needs to be migrated from the corresponding node to be optimized. Set corresponding judgment labels for different analysis data in turn; Each set of feature vectors is used as the input of the migration analysis model. The migration analysis model takes a set of prediction judgment labels corresponding to each set of analysis data as output, and takes the actual judgment labels corresponding to each set of analysis data as the prediction target. The actual judgment labels are the pre-set judgment labels corresponding to the analysis data. The training goal is to minimize the sum of the prediction errors of all analysis data. The calculation formula of the prediction error is: ,in is the prediction error, is the group number of the eigenvector corresponding to the analyzed data, For the The prediction judgment label corresponding to the group analysis data, For the The actual judgment label corresponding to the group analysis data; the migration analysis model is trained until the sum of the prediction errors reaches convergence and the training is stopped.
[0036] S5: The edge gateway analyzes the node feature data, generates a data migration strategy, and determines whether the migration node in the data migration strategy is an adapter node. If the migration node is an adapter node, the storage path of the data to be optimized is dynamically optimized according to the data migration strategy.
[0037] Methods for generating data migration strategies include: Build Group migration node set, each group migration node set includes normal nodes, and the same normal nodes are allowed to exist in each set of migration nodes, that is, the same normal node is allowed to appear multiple times in the same set of migration nodes; is an integer greater than 1, is the number of data to be optimized; set an increasing numeric label for each set of migration nodes and mark it as a migration label. The range of the migration label is ; Randomly generated within the range of migration labels Candidate solutions, There is a one-to-one correspondence between candidate solutions and transition labels. ; Define the iterative process: for each candidate solution, two intervals are randomly generated The random coefficients between the candidate solutions are compared with the preset judgment coefficients to determine the update behavior of each candidate solution. The range of the judgment coefficients is ; Update each candidate solution according to the update behavior and obtain The best solution among the candidate solutions; The iterative process is executed in a loop. When the number of executions of the iterative process is greater than or equal to a preset iterative threshold, the iterative process is stopped and the best solution is obtained. The set of migration nodes corresponding to the best solution is marked as the best set, and the best set is used as the data migration strategy. The iterative threshold is preset by technicians in this field according to actual conditions.
[0038] Build Methods for group migration of node sets include: Randomly select from all normal nodes Normal nodes, according to Normal nodes build a set of migration nodes, Group migration node set, each group migration node set There is a one-to-one correspondence between normal nodes and data to be optimized; wherein, a normal node is allowed to be selected into a set of migration nodes for multiple times.
[0039] Methods for determining update behavior include: Mark one of the random coefficients as the first coefficient, and mark the other random coefficient as the second coefficient; compare the first coefficient and the second coefficient with the judgment coefficient respectively; if the first coefficient is less than the judgment coefficient, the update behavior is determined to be an unknown update, that is, the candidate solution is updated in the direction of the value that is not a candidate solution within the migration label range; if the first coefficient is greater than or equal to the judgment coefficient, and the second coefficient is less than the judgment coefficient, then the update behavior is determined to be a current update, that is, the candidate solution is updated in the direction of the value that is currently being used as a candidate solution within the migration label range; if the first coefficient and the second coefficient are both greater than or equal to the judgment coefficient, then the update behavior is determined to be a known update, that is, the candidate solution is updated in the direction of the value that is used as a candidate solution within the migration label range during the historical iteration process.
[0040] Methods for updating candidate solutions based on update behaviors include: When the update behavior is unknown update, from the interval A random unknown coefficient is generated in Subtract one, multiply by the unknown coefficient, and add one to obtain the unknown update value; update the candidate solution according to the unknown update value; When the update behavior is current update, from the interval A current coefficient is randomly generated, and the current coefficient is multiplied by the random solution to obtain a random update value, and the random solution is a random candidate solution; the current coefficient is subtracted by one, and then multiplied by the candidate solution to obtain a candidate update value; the random update value is added to the candidate update value to obtain the current update value; the candidate solution is updated according to the current update value; When the update behavior is a known update, from the interval A known coefficient is randomly generated, and the known coefficient is subtracted from one and multiplied by the candidate solution to obtain a candidate update value; the known coefficient is multiplied by the previous solution to obtain a previous update value, and the previous solution is the candidate solution in the previous iteration; the known coefficient is multiplied by the random solution to obtain a followed update value; the candidate update value, the previous update value and the followed update value are added in sequence to obtain a known update value; the candidate solution is updated according to the known update value.
[0041] Get Methods for finding the best solution among candidate solutions include: The migration node sets corresponding to the migration labels corresponding to each candidate solution are marked as the current set, and the normal nodes in each current set are obtained and marked as the current nodes; the access frequency, heat value and data size of each data to be optimized are taken as a set of parameters to be optimized, and the parameters to be optimized correspond to the data to be optimized one by one; the node feature data of the current node corresponding to each candidate solution, the node feature data of all nodes to be optimized and the parameters to be optimized of all data to be optimized are taken as evaluation data, and the evaluation data correspond to the candidate solution one by one; each set of evaluation data is input into the trained benefit evaluation model respectively to evaluate the corresponding migration benefit ratio, which is the ratio between the benefit brought by data migration and the migration cost; the benefits brought by migration include, for example, improved access performance, optimized storage node load balancing, and improved hit rate of data with high heat value; the migration cost includes, for example, data migration time and bandwidth consumed by data migration; the training process of the benefit evaluation model is consistent with that of the migration analysis model, and both are deep neural network models; the migration benefit ratios corresponding to each candidate solution are compared respectively, and the candidate solution with the largest migration benefit ratio is marked as the best solution.
[0042] The method for determining whether a migration node in a data migration strategy is an adaptation node includes: Obtain the storage nodes in the data migration strategy and mark them as migration nodes; obtain the data to be optimized corresponding to each migration node, and use the node feature data of each migration node and the parameters to be optimized corresponding to the data to be optimized as a set of prediction data, and the prediction data corresponds to the migration node one-to-one; input each set of prediction data into the trained feature prediction model to predict the corresponding migration feature data; the migration feature data is the node feature data corresponding to the migration node after the data to be optimized is migrated to the corresponding migration node; the training process of the feature prediction model is consistent with the training process of the migration analysis model, and both are deep neural network models; use the type label corresponding to each data to be optimized and the migration feature data of the corresponding migration node as a set of judgment data, and the judgment data corresponds to the data to be optimized one-to-one; input each set of judgment data into the trained migration analysis model to predict the corresponding judgment label and mark it as the prediction label; obtain the migration judgment result corresponding to each migration node according to the prediction label, if the migration judgment result is all retained, the corresponding migration node is marked as an adaptation node, if there is migration in the migration judgment result, the corresponding migration node is not marked.
[0043] S6: If the migration node is not an adapter node, the cloud platform obtains the global feature data and sends it to the edge gateway. The edge gateway conducts a comprehensive analysis of the global feature data, optimizes the data migration strategy, and dynamically optimizes the storage path of the data to be optimized according to the optimized data migration strategy.
[0044] The global feature data includes node feature data of all storage nodes; based on the global feature data, the migration node set is regenerated, and the data migration strategy is regenerated based on the regenerated migration node set; based on the regenerated data migration strategy, the data migration strategy generated in S5 is optimized.
[0045] This embodiment analyzes the node feature data of the storage nodes in real time and uses the anomaly detection algorithm to intelligently identify the storage nodes that need to be optimized; and analyzes the data feature information of the data in the nodes to be optimized to accurately locate the data to be optimized that should be migrated; at the same time, based on the evaluation of migration benefits, it independently generates data migration strategies, adaptively adjusts the migration of data between storage nodes, and realizes dynamic optimization of data storage paths; in addition, when the edge gateway cannot meet the optimization requirements, it can further optimize the data migration strategy with the help of the global perspective of the cloud, and use the synergistic advantages of the edge and the cloud to improve data storage efficiency and access performance; it gives full play to the advantages of edge computing and realizes intelligent and adaptive optimization of data storage paths, which helps to reduce the cloud load and improve the real-time and reliability of data access. At the same time, it optimizes the utilization rate of storage resources and solves the problem that traditional static configuration cannot adapt to dynamic access characteristics, thereby significantly improving the data management capabilities and overall performance in the edge computing environment.
[0046] Example 2 The present application also provides an electronic device. The electronic device may include one or more processors and one or more memories. The memory stores a computer-readable code, and when the computer-readable code is executed by one or more processors, the method for dynamically optimizing a data storage path based on edge computing as described above may be executed.
[0047] The method or system according to the implementation mode of the present application can also be implemented with the aid of the architecture of the electronic device shown in the present application. The electronic device may include a bus, one or more CPUs, ROM, RAM, a communication port connected to a network, input / output, a hard disk, etc. A storage device in an electronic device, such as a ROM or a hard disk, can store the method for dynamic optimization of data storage paths based on edge computing provided in the present application. Furthermore, the electronic device may also include a user interface. Of course, the architecture shown in the present application is only exemplary. When implementing different devices, one or more components in the electronic device shown in the present application may be omitted according to actual needs.
[0048] Example 3 As shown, one embodiment of the present application discloses a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by the processor, the method for dynamically optimizing the data storage path based on edge computing according to the embodiment of the present application described with reference to the above figures can be executed. The storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory (cache). Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0049] In addition, according to the implementation of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the present application provides a non-transitory machine-readable storage medium, which stores machine-readable instructions, and the machine-readable instructions can be executed by a processor to execute instructions corresponding to the method steps provided in the present application, such as: a data storage path dynamic optimization method based on edge computing. When the computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed.
[0050] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
[0051] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A data storage path dynamic optimization method based on edge computing, characterized in that: Applied in edge gateways, including: S1: Receive node feature data sent by the storage node; S2: Perform multi-dimensional performance analysis on node feature data and use anomaly detection algorithms to identify nodes to be optimized in storage nodes; S3: receiving data feature information sent by the node to be optimized; S4: Perform hierarchical analysis on data feature information and combine hotspot data identification technology to accurately locate the data to be optimized in the node to be optimized; S5: Analyze the node feature data, generate a data migration strategy, and determine whether the migration node in the data migration strategy is an adaptation node. If the migration node is an adaptation node, dynamically optimize the storage path of the data to be optimized according to the data migration strategy; S6: If the migration node is not an adapter node, the global feature data sent by the cloud platform is received, the global feature data is comprehensively analyzed, the data migration strategy is optimized, and the storage path of the data to be optimized is dynamically optimized according to the optimized data migration strategy.
2. The data storage path dynamic optimization method based on edge computing according to claim 1 is characterized in that: The node characteristic data includes storage capacity status, read / write performance status, load status and network resource status; wherein the storage capacity status includes total storage capacity and used capacity, the read / write performance status includes I / O read / write speed and IOPS, the load status includes CPU usage and memory usage, and the network resource status includes network bandwidth and network latency; The method for identifying a node to be optimized among storage nodes includes: Divide the used capacity in each set of node feature data by the corresponding total storage capacity to obtain the capacity utilization rate corresponding to each set of node feature data; replace the storage capacity state of each set of node feature data with the corresponding capacity utilization rate, and mark the replaced node feature data as replacement feature data; mark each set of replacement feature data obtained in real time as real-time feature data, and obtain historical feature data, which is the replacement feature data of normal nodes obtained at historical moments, and normal nodes are storage nodes that are not marked as nodes to be optimized; each data in each set of real-time feature data is respectively taken as a set of data sets with the same type of data in the historical feature data, and the data sets correspond to the data in the real-time feature data one by one; calculate the abnormal coefficient and coefficient threshold corresponding to each set of data sets in turn, and compare the abnormal coefficient corresponding to each set of data sets with the corresponding coefficient threshold; mark the data sets with abnormal coefficients greater than or equal to the coefficient threshold as abnormal sets, and do not mark the data sets with abnormal coefficients less than the coefficient threshold; mark the data in the real-time feature data corresponding to each set of abnormal sets as abnormal data, mark the real-time feature data corresponding to the abnormal data as abnormal feature data, and take the storage nodes corresponding to the abnormal feature data as nodes to be optimized.
3. The data storage path dynamic optimization method based on edge computing according to claim 2 is characterized in that: Methods for calculating the anomaly coefficient corresponding to a data set include: Calculate the Euclidean distance between each two data in the data set in turn and mark them as data distances; sort the data distances corresponding to each data from small to large to generate a distance sorting table corresponding to each data; sort the data at the front in each distance sorting table The data distance of each is taken as the adjacent distance of the corresponding data. is an integer greater than 1; add the adjacent distances of each data in turn and then divide by , obtain the average distance corresponding to each data; mark the data corresponding to the real-time feature data in the data set as real-time data, and use the average distance of the real-time data as the abnormal coefficient of the data set; Methods for calculating coefficient thresholds corresponding to a data set include Count the number of data in the data set and mark it as the number of data; add the average distances corresponding to the data set in sequence, and then divide it by the number of data to obtain the overall mean corresponding to the data set; subtract the overall mean from the average distance of each data, and then square it to obtain the square of the difference of each data; add the square of the difference of each data in sequence and divide it by the number of data, and then take the square root to obtain the overall dispersion; multiply the overall dispersion by , plus the overall mean, as the coefficient threshold of the data set, .
4. The data storage path dynamic optimization method based on edge computing according to claim 3 is characterized in that: The data characteristic information includes write frequency, read frequency, data size and access time, and the access time is the time when the data was last accessed; The method for locating the data to be optimized in the node to be optimized includes: Obtain the current time, mark all data in the node to be optimized as data to be analyzed, subtract the current time from the access time of each data to be analyzed, and take the absolute value to obtain the time difference of each data to be analyzed; preset a weight set, the weight set includes the write frequency, the read frequency and the weight coefficient corresponding to the time difference; standardize the write frequency, the read frequency and the time difference of each data to be analyzed, respectively, to obtain standardized data, the standardized data includes the standard write frequency, the standard read frequency and the standard time difference; multiply the reciprocal of the standard write frequency, the standard read frequency and the standard time difference of each data to be analyzed by the corresponding weight coefficient, and add them in sequence to obtain the heat value of each data to be analyzed; Each data to be analyzed is hierarchically classified to obtain the hierarchical type of each data to be analyzed; according to the hierarchical type of each data to be analyzed, the type label of each data to be analyzed is obtained, the type label is a digital label corresponding to the hierarchical type, and the hierarchical labels corresponding to different hierarchical types are all different; the type label of each data to be analyzed and the node feature data of the corresponding node to be optimized are taken as a group of analysis data, and the analysis data corresponds to the data to be analyzed one by one; each group of analysis data is input into the trained migration analysis model respectively, and the corresponding judgment label is predicted, the judgment label is the digital label corresponding to the migration judgment result, and the migration judgment result includes migration and retention, and migration and retention correspond to different judgment labels; according to the judgment label, the migration judgment result corresponding to each group of analysis data is obtained, if the migration judgment result is migration, then the data to be analyzed corresponding to the corresponding analysis data is marked as data to be optimized, if the migration judgment result is retention, then the data to be analyzed corresponding to the corresponding analysis data is not marked.
5. The data storage path dynamic optimization method based on edge computing according to claim 4 is characterized in that: The method for obtaining the hierarchical type of each data to be analyzed includes: The write frequency and the read frequency of each data to be analyzed are added to obtain the access frequency of each data to be analyzed; a classification standard set is preset, and the classification standard set includes three classification sets, which correspond to the access frequency, heat value and data size respectively; wherein each classification set includes a first threshold and a second threshold, and the first threshold is greater than the second threshold; based on the preset classification rules, the access frequency, heat value and data size of each data to be analyzed are compared with the corresponding classification sets respectively to obtain the frequency type, heat type and size type corresponding to each data to be analyzed; the frequency type, heat type and size type of each data to be analyzed are combined to obtain the hierarchical type of each data to be analyzed.
6. The method for dynamic optimization of data storage paths based on edge computing according to claim 5 is characterized in that: The method for generating a data migration strategy includes: Build Group migration node set, each group migration node set includes Normal nodes; among them, is an integer greater than 1, is the number of data to be optimized; set an increasing numeric label for each set of migration nodes and mark it as a migration label. The range of the migration label is ; Randomly generated within the range of migration labels Candidate solutions, There is a one-to-one correspondence between candidate solutions and transition labels. ; Define the iterative process: for each candidate solution, two intervals are randomly generated The random coefficients between the candidate solutions are compared with the preset judgment coefficients to determine the update behavior of each candidate solution. The range of the judgment coefficients is ; Update each candidate solution according to the update behavior and obtain The best solution among the candidate solutions; The iterative process is executed cyclically. When the number of executions of the iterative process is greater than or equal to a preset iteration threshold, the iterative process is stopped and the best solution is obtained. The migration node set corresponding to the best solution is marked as the best set, and the best set is used as the data migration strategy.
7. The method for dynamic optimization of data storage paths based on edge computing according to claim 6 is characterized in that: Build Methods for group migration of node sets include: Randomly select from all normal nodes Normal nodes, according to Normal nodes build a set of migration nodes, Group migration node set, each group migration node set There is a one-to-one correspondence between normal nodes and the data to be optimized; Methods for determining update behavior include: Mark one of the random coefficients as the first coefficient and the other random coefficient as the second coefficient; compare the first coefficient and the second coefficient with the judgment coefficient respectively; if the first coefficient is less than the judgment coefficient, determine that the update behavior is an unknown update; if the first coefficient is greater than or equal to the judgment coefficient, and the second coefficient is less than the judgment coefficient, determine that the update behavior is a current update; if both the first coefficient and the second coefficient are greater than or equal to the judgment coefficient, determine that the update behavior is a known update; Methods for updating candidate solutions based on update behaviors include: When the update behavior is unknown update, from the interval A random unknown coefficient is generated in Subtract one, multiply by the unknown coefficient, and add one to obtain the unknown update value; update the candidate solution according to the unknown update value; When the update behavior is current update, from the interval A current coefficient is randomly generated, and the current coefficient is multiplied by the random solution to obtain a random update value, and the random solution is a random candidate solution; the current coefficient is subtracted by one, and then multiplied by the candidate solution to obtain a candidate update value; the random update value is added to the candidate update value to obtain the current update value; the candidate solution is updated according to the current update value; When the update behavior is a known update, from the interval A known coefficient is randomly generated, and the known coefficient is subtracted from one and multiplied by the candidate solution to obtain a candidate update value; the known coefficient is multiplied by the previous solution to obtain a previous update value, and the previous solution is the candidate solution in the previous iteration; the known coefficient is multiplied by the random solution to obtain a followed update value; the candidate update value, the previous update value and the followed update value are added in sequence to obtain a known update value; the candidate solution is updated according to the known update value.
8. The data storage path dynamic optimization method based on edge computing according to claim 7 is characterized in that: Get Methods for finding the best solution among candidate solutions include: The migration node sets corresponding to the migration labels of each candidate solution are marked as the current set, and the normal nodes in each current set are obtained and marked as the current nodes; the access frequency, heat value and data size of each data to be optimized are taken as a set of parameters to be optimized, and the parameters to be optimized correspond to the data to be optimized one by one; the node feature data of the current node corresponding to each candidate solution, the node feature data of all nodes to be optimized and the parameters to be optimized of all data to be optimized are taken as evaluation data, and the evaluation data correspond to the candidate solutions one by one; each set of evaluation data is input into the trained benefit evaluation model respectively to evaluate the corresponding migration benefit ratio; the training process of the benefit evaluation model is consistent with the training process of the migration analysis model, and both are deep neural network models; the migration benefit ratios corresponding to each candidate solution are compared respectively, and the candidate solution with the largest migration benefit ratio is marked as the best solution.
9. The data storage path dynamic optimization method based on edge computing according to claim 8 is characterized in that: The method for determining whether a migration node in a data migration strategy is an adaptation node includes: Obtain the storage nodes in the data migration strategy and mark them as migration nodes; obtain the data to be optimized corresponding to each migration node, and use the node feature data of each migration node and the parameters to be optimized corresponding to the data to be optimized as a set of prediction data, and the prediction data corresponds to the migration node one-to-one; input each set of prediction data into the trained feature prediction model to predict the corresponding migration feature data; the migration feature data is the node feature data corresponding to the migration node after the data to be optimized is migrated to the corresponding migration node; the training process of the feature prediction model is consistent with the training process of the migration analysis model, and both are deep neural network models; use the type label corresponding to each data to be optimized and the migration feature data of the corresponding migration node as a set of judgment data, and the judgment data corresponds to the data to be optimized one-to-one; input each set of judgment data into the trained migration analysis model to predict the corresponding judgment label and mark it as the prediction label; obtain the migration judgment result corresponding to each migration node according to the prediction label, if the migration judgment result is all retained, the corresponding migration node is marked as an adaptation node, if there is migration in the migration judgment result, the corresponding migration node is not marked.
10. The data storage path dynamic optimization method based on edge computing according to claim 9 is characterized in that: The global characteristic data includes node characteristic data of all storage nodes; Regenerate a migration node set according to the global feature data, and regenerate a data migration strategy according to the regenerated migration node set; According to the regenerated data migration strategy, the data migration strategy generated in S5 is optimized.
Citation Information
Patent Citations
Big data-based marginalization calculation method
CN118349359A
Cited By
Data processing method of storage system and electronic equipment
CN120872262A
Distributed data storage dynamic optimization method based on edge computing
CN121193759A
Distributed data storage dynamic optimization method based on edge computing
CN121193759B