A Cloud Big Data Storage Management Method

By building a dynamic node network and real-time monitoring of terminal node access habits, the problems of waste of resources and inefficiency in cloud big data storage management are solved, and efficient and reliable data storage and backup are achieved.

CN119781690BActive Publication Date: 2025-07-11HEBEI XIONGAN TINGYUN INFORMATION TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411960366.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-07-11
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The existing cloud-based big data storage management methods cannot adapt to the dynamic changes in terminal node access habits, resulting in waste of storage resources and inefficient efficiency, and the frequent occurrence of failed storage and backups of non-online state nodes.

Method used

By collecting the multi-access habit feature vectors of terminal nodes, using the K-means clustering algorithm to group and build the first and second node networks, the target data and its traceability points are monitored in real time, the association processing instructions are generated, the data storage and backup strategies are dynamically adjusted, and the data block allocation is assigned to consider the active period and stable index of the terminal node.

Benefits of technology

It improves the utilization rate and reliability of storage resources, reduces the failed storage and backup of non-online state nodes, optimizes the time point of storage and backup, and ensures data security and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119781690B_ABST
    Figure CN119781690B_ABST
Patent Text Reader

Abstract

The present application discloses a cloud big data storage management method, and the method includes: classifying, by using a preset classification model, all terminal node's multi - access habit feature vectors to obtain several terminal node groups; constructing a first node network and a second node network based on all terminal node groups, where the first node network is used for distributed storage of data, and the second node network is used for distributed backup of data; real - time monitoring and obtaining target data and its traceability point, and generating an associated processing instruction for the target data according to a preset differential traceability processing decision, where the associated processing instruction includes storage, backup, and forwarding; periodically executing the foregoing steps at a preset time interval. Thereby, based on a dynamic node network, dynamic storage and backup of data are realized, adapting to changes in terminal node access habits, reducing failed storage of non - online state nodes, and improving the reliability of storage resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data management, and particularly to a method for storing and managing big data in the cloud. Background Art

[0002] With the rapid development of information network technology, the volume of data stored on the network is becoming increasingly large. Users can upload and download data by themselves to save, share, and obtain content, which leads to an increasing demand for data storage in the cloud. When a user downloads data and forwards it to other users, a copy of the data will be further generated, resulting in a waste of cloud storage resources. Moreover, the backup requirement for data in the cloud further squeezes the remaining cloud storage space, resulting in a low effective utilization rate of cloud storage resources.

[0003] The Chinese invention patent with the application number 202311237807.2 discloses a cloud big data storage management system. Based on the list of all terminal node devices connected to the cloud, terminal nodes are screened based on the historical connection period to establish a cloud storage network; forward and store data, and based on the data source differentiation storage strategy: when the data source is the user side, the data is segmented to generate multiple groups of backup data segments, and appropriate terminal nodes are selected in the cloud storage network for distributed backup storage to generate a data storage index table; when the data source is the cloud, forward the corresponding index link.

[0004] However, since the states of several terminal nodes connected to the cloud are dynamically changing, simply constructing a unified cloud storage network based on the historical connection period between the terminal nodes and the cloud cannot adapt to the changes in the cloud access habits of different user sides. In addition, due to the non-fixed time points for users to upload data, the time points for data storage and backup are unpredictable. Using the pre-unified cloud storage network for distributed storage management leads to limitations in the efficiency of cloud big data storage management, easily causing failed storage and backup of non-online state nodes, thus resulting in congestion of cloud storage resources. Summary of the Invention

[0005] The present application provides a method for storing and managing big data in the cloud, which realizes dynamic storage and backup of data based on a dynamic node network, adapts to changes in the access habits of terminal nodes, reduces failed storage of non-online state nodes, and improves the reliability of storage resources.

[0006] The present application provides a method for storing and managing big data in the cloud, including:

[0007] S101. Collect the communication records of all terminal nodes connected to the cloud at all time windows within a preset historical period. The communication records include communication time period, communication duration, communication data volume, and their stability indices, and generate a multi - access habit feature vector for each terminal node.

[0008] S102. Based on the multi - access habit feature vectors of all terminal nodes, use a preset classification model to classify them, obtain several terminal node groups, and label each terminal node group. The labeling content is set as the central feature vector.

[0009] S103. Based on all terminal node groups, construct a first node network and a second node network. The first node network is used for distributed storage of data, and the second node network is used for distributed backup of data.

[0010] S104. Real - time monitor and obtain target data and its traceability points, and generate an associated processing instruction for the target data according to a preset differential traceability processing decision. The associated processing instruction includes storage, backup, and forwarding.

[0011] S105. Periodically execute steps S101 to S104 at a preset time interval.

[0012] Preferably, the generation of the multi - access habit feature vector for each terminal node includes:

[0013] A1. Obtain the communication time periods in all communication records collected by this terminal node, draw a communication time distribution graph, and obtain the communication times with a density greater than the density threshold to form active time periods.

[0014] A2. Obtain the communication success rate of each communication data volume as the stability index of data communication.

[0015] A3. Obtain all communication durations, communication data volumes, and stability indices within the active time periods in the communication records, and respectively take the averages to obtain the average communication duration, average communication data volume, and average stability index.

[0016] A4. Based on the active time period, average communication duration, average communication data volume, and average stability index of each terminal node, construct a multi - access habit feature vector.

[0017] Preferably, the preset classification model is set as the K - means clustering algorithm, and S102 specifically includes:

[0018] Input all multi - access habit features into the classification model to generate k clustering results. Each cluster corresponds to a terminal node group, and use the cluster center of each cluster as the label of the corresponding terminal node group.

[0019] Preferably, the method for constructing the first node network includes: retrieving the central feature vectors of each terminal node group, and selecting the terminal node group whose active period and the current time node meet the adaptation condition as the first target object; wherein, the adaptation condition is set as: the current time node is within the active period; constructing the first node network based on all the terminal nodes in the first target object and the cloud;

[0020] The method for constructing the second node network includes:

[0021] retrieving the central feature vectors of each terminal node group, selecting the terminal node group whose stability index is greater than the stability threshold as the second target object, and determining the combination with the maximum active period coverage rate in the second target object; constructing the second node network based on the terminal nodes in all the terminal node groups in the combination with the maximum active period coverage rate and the cloud.

[0022] Preferably, according to the pre-set differential traceability processing decision, an association processing instruction for the target data is generated, specifically including:

[0023] S201, if the traceability point is a terminal node, use the first node network and the second node network to store and back up the target data respectively, generate index information of the target data and update it to the pre-set cloud index library; the index information includes storage index information and backup index information;

[0024] S202, if the traceability point is the cloud, retrieve the index information of the target data in the cloud index library and forward it to the target terminal node.

[0025] Preferably, the storage of the target data using the first node network includes:

[0026] B1. Based on the multi-access habit feature vectors of all the terminal nodes in the first node network, screen out the terminal nodes whose active period and the current time node meet the adaptation condition as distributed storage nodes;

[0027] B2. Based on the first allocation algorithm, cut the target data into several data blocks and allocate them to the corresponding distributed storage nodes for storage;

[0028] Among them, the first allocation algorithm includes:

[0029] S1. Calculate the size of all the data blocks cut from the target data block according to the following formula:

[0030]

[0031] is the size of the data block allocated to each distributed storage node, D is the size of the target data, is the stability index of the i-th distributed storage node, and N is the total number of distributed storage nodes;

[0032] S2. Based on the cutting size of the target data, cut it into N data blocks in a preset order, and store each data block in the corresponding distributed storage node according to its size;

[0033] B3. Label each data block with a unique identifier and record the corresponding distributed storage node to generate an index link. Determine the storage index information of the target data as the unique identifiers of all data blocks of the target data and the corresponding index links respectively. The unique identifier is used to combine all data blocks to obtain the original complete target data.

[0034] Preferably, the backup of the target data using the second node network includes:

[0035] C1. Based on the number of groups of terminal nodes in the combination of the maximum active period coverage rate, determine the number of backup sources of the target data. Each backup source is used to perform a backup of the target data;

[0036] C2. Based on each backup source, regard all the terminal nodes therein as distributed backup nodes;

[0037] C3. Based on the second allocation algorithm, cut the target data into several data blocks and allocate them to the corresponding distributed backup nodes for backup;

[0038] Among them, the second allocation algorithm includes:

[0039] S3. Calculate the size of all data blocks cut from the target data block in each backup source according to the following formula:

[0040]

[0041] is the size of the data block allocated to each distributed backup node in this backup source, D is the size of the target data, is the stability index of the j-th distributed backup node in this backup source, and M is the total number of distributed backup nodes in this backup source;

[0042] S4. For each backup source, based on the cutting size corresponding to the target data, divide it into M data blocks in a preset order, and store each data block in the corresponding distributed backup node according to its size;

[0043] C4. Generate all backup sources of the target data, and mark all data blocks in each backup source with a unique identifier and record the corresponding distributed backup nodes to generate index links. The backup source of the target data, the unique identifiers of all data blocks in the backup source, and the index links corresponding to the data blocks are determined as the backup index information of the target data. The unique identifier is used to combine all data blocks to obtain the original complete target data.

[0044] Preferably, in S2, the method of obtaining the preset order includes:

[0045] Perform content analysis on the target data and divide it into multiple sub-blocks, so that each sub-block contains a relatively independent data portion;

[0046] Calculate the important contribution factor of each sub-block to the target data: perform weighted summation of the content proportion of the sub-block in the target data, the business value, and the similarity with the historical uploaded data of its traceability point to obtain the important contribution factor of the sub-block;

[0047] Obtain the important contribution factor of each sub-block in turn according to the order of the sub-blocks in the target data to form an important contribution factor sequence;

[0048] Based on the sizes of all elements in the important contribution factor sequence, the comprehensive important trend is obtained, and the direction of the comprehensive important trend is determined as a preset order;

[0049] Among them, the methods for determining the comprehensive important trends include:

[0050] The average value of all important contribution factors is obtained as the important threshold; the position of the important contribution factors greater than the important threshold in the important contribution factor sequence is determined. If the position is at the front, the descending order is determined as the comprehensive important trend; if the position is at the back, the ascending order is determined as the comprehensive important trend.

[0051] Preferably, the S2 further includes:

[0052] S301, determining all cutting points for cutting the target data in a preset order as an initial cutting axis, on which all cutting points are sequentially arranged, and every two adjacent cutting points form a cutting interval segment, and each cutting interval segment corresponds to an initial data block;

[0053] S302, recording the sub-blocks falling into each cutting interval and their important contribution factors, and calculating the comprehensive important contribution degree of each cutting interval;

[0054] S303. Generate a mapping feature vector for the cutting interval segment based on the comprehensive important contribution degree of each cutting interval segment, the proportion of the initial data block size, and the stability index of the corresponding distributed storage node allocated, and input it into the pre-trained mapping relationship anomaly recognition model to output the recognition result of the cutting interval segment. The recognition result includes normal and abnormal and their abnormal types. The abnormal type is set to the over-standard of the comprehensive important contribution degree and its over-standard ratio, or the shortage of the comprehensive important contribution degree and its shortage ratio;

[0055] S304. Obtain the abnormal cutting interval segment based on the recognition result, and use the pre-set local adjustment strategy to update the size of the initial data block of the abnormal cutting interval segment to the target size to obtain the target data block;

[0056] S305. Update the cutting interval segment on the initial cutting point axis of the target data to obtain the target cutting point axis, and cut the target data according to the target cutting point axis to obtain N data blocks.

[0057] Preferably, the acquisition method of the pre-trained mapping relationship anomaly recognition model includes:

[0058] E1. Collect the mapping feature vectors of the cutting interval segments corresponding to a large number of data blocks allocated to the terminal nodes in history. Based on each historical mapping feature vector, perform label annotation on it, and the annotation content is set to normal, abnormal and their abnormal types;

[0059] E2. Use the annotated mapping feature vectors as the training set, and use the training set to train the pre-selected neural network structure, optimize the model parameters, and generate the final mapping relationship anomaly recognition model;

[0060] The pre-set local adjustment strategy includes:

[0061] D1. Obtain the abnormal type of the abnormal cutting interval segment. If the abnormal type is the over-standard of the comprehensive important contribution degree, use the first local adjustment algorithm to obtain the target size; if the abnormal type is the shortage of the comprehensive important contribution degree, use the second local adjustment algorithm to obtain the target size;

[0062] D2. Dynamically move the reference point of the cutting interval segment to adjust the size of the initial data block corresponding to the cutting interval segment to the target size to obtain the target data block.

[0063] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0064] By collecting and analyzing the multi - access habit feature vectors of terminal nodes, the adjustment node network is periodically constructed to analyze the access habits of terminal nodes; based on the active period and stability index of terminal nodes, the first node network for data storage and the second node network for data backup are respectively constructed, which not only stores data efficiently but also ensures the reliability of data backup, improving the dynamic adaptability of storage and backup; the target data and its traceability points are monitored in real - time, and the associated processing instructions are generated according to the differential traceability processing decision, realizing the dynamic storage and backup of data, adapting to the changes of terminal node access habits in real - time, optimizing the storage and backup time points, improving the utilization rate of storage resources and the security of data, reducing the failed storage and backup of non - online state nodes, and increasing the utilization rate of storage resources;

[0065] Traditional data distributed cutting storage adopts a uniform cutting method without considering the reliability index of terminal nodes in the corresponding node network. The first allocation algorithm allocates the data block size according to the stability index of terminal nodes during the active period, and the second allocation algorithm allocates the data block size in the backup source according to the stability index of distributed backup nodes during the corresponding active period. Considering the stability and active period of terminal nodes in the node network, it improves the rationality of data storage and backup. Terminal nodes with a high stability index are allocated larger data blocks, improving the efficiency and reliability of data storage and backup, and avoiding the resource waste and imbalance problems that may be caused by uniform allocation;

[0066] Through the node network constructed periodically for multi - location backup and data storage, it effectively solves the problems of insufficient dynamic adaptability and storage resource congestion existing in traditional cloud storage data;

[0067] By calculating the important contribution factor of sub - blocks and determining the cutting order according to its comprehensive important trend, the cutting of data blocks is made more reasonable; to the greatest extent, it avoids sub - blocks with a larger important contribution factor being allocated to terminal nodes with poor stability, thus reducing storage risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 It is a schematic flow chart of the cloud big data storage management method according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0069] To facilitate the understanding of the present invention, the present application will be described more comprehensively with reference to the relevant drawings; the preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein; on the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.

[0070] It should be noted that the terms "vertical", "horizontal", "upper", "lower", "left", "right" and similar expressions used in this article are for illustrative purposes only and do not represent the only implementation.

[0071] Unless otherwise defined, all technical and scientific terms used in this article have the same meaning as those commonly understood by those skilled in the technical field to which this invention belongs; the terms used in the description of this invention in this article are only for the purpose of describing specific implementations and are not intended to limit this invention; the term "and / or" used in this article includes any and all combinations of one or more of the related listed items.

[0072] Embodiment 1: Figure 1 It is a schematic flowchart of the cloud big data storage management method of the embodiment of the present invention.

[0073] As Figure 1 shown, a cloud big data storage management method includes the following steps:

[0074] S101, collect the communication records of all terminal nodes connected to the cloud in all time windows within a preset historical period. The communication records include, but are not limited to, communication time periods (connection time periods), communication durations, communication data volumes, and their stability indices, and generate a multi - access habit feature vector for each terminal node.

[0075] Among them, the preset historical period is set to the past week, and the time window is set to one day, which can be adjusted according to actual situations.

[0076] Specifically, generating a multi - access habit feature vector for each terminal node includes:

[0077] A1. Obtain the communication time periods in all the communication records collected by this terminal node, draw a communication time distribution diagram, obtain the communication times with a communication frequency greater than the density threshold, and form an active period;

[0078] Among them, step A1 specifically includes:

[0079] Obtain the communication time period: collect all the communication records of the terminal node within the preset historical period, extract the start and end times of each communication, and form a communication time series;

[0080] Draw the communication time distribution diagram: plot the communication time series as a time distribution diagram, with the horizontal axis being time (such as hours) and the vertical axis being the communication frequency;

[0081] Determine the density threshold: according to the communication time distribution diagram, set a density threshold for identifying time periods with frequent communication activities;

[0082] Identify the active period: Traverse the communication time distribution map, find the communication times with a communication frequency greater than the density threshold, and form the active period of this terminal node with these communication times.

[0083] Illustrative example: In the communication time distribution map of a certain terminal node in the past week, set the density threshold to 1.5 times the average communication frequency. It is found that the communication frequency of the communication time directly from 9:00 to 11:00 every morning is greater than the density threshold. Then the active period of this terminal node is from 9:00 to 11:00 in the morning.

[0084] A2. Obtain the communication success rate of the amount of communication data for each communication, that is, the integrity of the communication data during transmission (the size of the communication data received at the receiving end / the size of the communication data sent at the sending end, the matching degree between the communication data content sent and the communication data content received, and the two are weighted and summed) to represent the stability index of data communication of this terminal node;

[0085] A3. Obtain all the communication durations, communication data amounts, and stability indices within the active period in the communication records, and take the average respectively to obtain the average communication duration, average communication data amount, and average stability index;

[0086] A4. Based on the active period, average communication duration, average communication data amount, and average stability index of each terminal node, construct a multi - access habit feature vector.

[0087] S102. Based on the multi - access habit feature vectors of all terminal nodes, use a preset classification model to classify them, obtain several terminal node groups, and label each terminal node group. The labeling content is set as the central feature vector.

[0088] Among them, the preset classification model is set as the K - means clustering algorithm. Step S102 specifically includes:

[0089] Input all the multi - access habit features into the classification model to generate k clustering results. Each clustering corresponds to a terminal node group, and use the clustering center (central feature vector) of each clustering as the label of the corresponding terminal node group.

[0090] For example, in step 1, the model randomly selects k initial clustering centers, which are points in the multi - access habit feature space and represent the initial clustering hypotheses; in step 2, the multi - access habit feature vector of a terminal node is [x1, x2, ..., x4], calculate its distance from each clustering center, and then add it to the clustering with the minimum distance; in step 3, after all terminal nodes are assigned to a certain clustering, recalculate the center of each clustering (i.e., the average of the feature vectors of all terminal nodes in the clustering), for example, for clustering C, its new center feature vector is the mean vector of the feature vectors of all terminal nodes in C; in step 4, repeat steps 2 and 3 until the clustering centers no longer change significantly (i.e., reach the convergence condition), or reach the preset number of iterations. The convergence condition can be that the change in the clustering centers is less than a certain threshold, or the clustering result no longer changes; in step 5, finally, each clustering corresponds to a terminal node group, and use the center (center feature vector) of each clustering as the label of the corresponding terminal node group. Among them, x1 is the active period, x2 is the average communication duration, x3 is the average communication data volume, and x4 is the average stability index.

[0091] It should be noted that the specific classification process of the K - means clustering algorithm can refer to the relevant prior art content, and the present invention will not elaborate on this.

[0092] S103, based on all terminal node groups, construct a first node network and a second node network. The first node network is used for distributed storage of data, and the second node network is used for distributed backup of data.

[0093] Specifically, the construction method of the first node network includes:

[0094] Retrieve the center feature vectors of each terminal node group, and select the terminal node groups (at least one) whose active period and the current time node meet the adaptation condition as the first target objects; where the adaptation condition is set as: the current time node is within the active period; based on all terminal nodes in the first target objects and the cloud, construct the first node network.

[0095] Thus, the first node network ensures efficient storage of data during the active period.

[0096] Specifically, the construction method of the second node network includes:

[0097] Retrieve the central feature vectors of each group of terminal nodes, select the groups of terminal nodes whose stability index is greater than the stability threshold as the second target objects, and determine the maximum active period coverage combination (the combination of multiple active periods with the largest time proportion in a time window) among the second target objects; based on the terminal nodes within all the groups of terminal nodes in the maximum active period coverage combination and the cloud, construct a second node network.

[0098] Thus, by selecting stable nodes, the second node network ensures the reliability of data backup.

[0099] S104, Monitor and obtain the target data and its traceability points in real time, and generate an associated processing instruction for the target data according to the pre-set differential traceability processing decision. The associated processing instruction includes storage, backup, and forwarding.

[0100] In some embodiments, generating an associated processing instruction for the target data according to the pre-set differential traceability processing decision specifically includes:

[0101] S201, If the traceability point is a terminal node, indicating that the target data is the stored data uploaded by the terminal node, then use the first node network and the second node network to store and backup the target data respectively, generate index information for the target data and update it to the pre-set cloud index library. Among them, the index information includes storage index information and backup index information.

[0102] Specifically, using the first node network to store the target data includes:

[0103] B1. Based on the multi-access habit feature vectors of all terminal nodes in the first node network, screen out the terminal nodes whose active period and the current time node meet the adaptation conditions as distributed storage nodes;

[0104] B2. Based on the first allocation algorithm, cut the target data into several data blocks and allocate them to the corresponding distributed storage nodes for storage;

[0105] Among them, the first allocation algorithm includes:

[0106] S1. Calculate the size of all data blocks obtained by cutting the target data block according to the following formula:

[0107]

[0108] is the size of the data block allocated to each distributed storage node, D is the size of the target data, is the stability index of the i-th distributed storage node, and N is the total number of distributed storage nodes.

[0109] S2. Based on the cutting size of the target data (the size of the data block allocated to each distributed storage node), cut it into N data blocks in sequence according to the preset order, and store them in the corresponding distributed storage nodes based on the size of each data block.

[0110] B3. Label each data block with a unique identifier and record the corresponding distributed storage node to generate an index link. Determine the storage index information of all data blocks of the target data as the unique identifier and the corresponding index link of each data block. The unique identifier is used to combine all data blocks to obtain the original complete target data.

[0111] Specifically, using the second node network to back up the target data includes:

[0112] C1. Based on the number of terminal node groups in the combination of the maximum active period coverage rate, determine the number of backup sources of the target data. Each terminal node group (i.e., backup source) is used to perform a backup of the target data;

[0113] C2. Based on each backup source, use all the terminal nodes therein as distributed backup nodes;

[0114] C3. Based on the second allocation algorithm, cut the target data into several data blocks and allocate them to the corresponding distributed backup nodes for backup;

[0115] Among them, the second allocation algorithm includes:

[0116] S3. Calculate the size of all data blocks cut within each backup source for the target data block according to the following formula:

[0117]

[0118] is the size of the data block allocated to each distributed backup node in this backup source, D is the size of the target data, is the stability index of the jth distributed backup node in this backup source, and M is the total number of distributed backup nodes in this backup source.

[0119] S4. For each backup source, based on the cutting size corresponding to the target data (the size of the data block allocated to each distributed backup node in this backup source), divide it into M data blocks in sequence according to the preset order, and store them in the corresponding distributed backup nodes based on the size of each data block.

[0120] Among them, the preset order of division is the same as the principle in step S2, and the present invention will not elaborate on this.

[0121] C4. Generate all backup sources of the target data, label each data block in each backup source with a unique identifier, record the corresponding distributed backup nodes to generate index links, and determine the backup source of the target data, the unique identifiers of all data blocks in the backup source, and the index links corresponding to the data blocks respectively as the backup index information of the target data. The unique identifier is used to combine all data blocks to obtain the original complete target data.

[0122] S202. If the traceability point is the cloud, indicating that the target data is the data that the cloud needs to forward to the target terminal node, then retrieve the index information of the target data in the cloud index library and forward it to the target terminal node.

[0123] Specifically, when the traceability point is the cloud, retrieve the storage index information corresponding to the target data in the cloud index library, including the unique identifiers of all its data blocks and the corresponding index links respectively, and forward it to the target terminal node.

[0124] S105. Periodically execute steps S101 to S104 at a preset time interval to adapt to the changes in the access habits of different terminal nodes and avoid the limitations of a fixed unified node network in a dynamic network.

[0125] Among them, the preset time interval is set to 2 hours (online duration, active period), and it can also be adjusted according to the access situation of the terminal node.

[0126] The technical solutions in the embodiments of the present application described above have at least the following technical effects or advantages:

[0127] By collecting and analyzing the multivariate access habit feature vectors of the terminal nodes, periodically constructing and adjusting the node network, and analyzing the access habits of the terminal nodes; based on the active period and stability index of the terminal nodes, respectively constructing the first node network and the second node network for data storage and backup, while efficiently storing data, ensuring the reliability of data backup, improving the dynamic adaptability of storage and backup; real-time monitoring the target data and its traceability point, generating associated processing instructions according to the differential traceability processing decision, realizing the dynamic storage and backup of data, adapting to the changes in the access habits of the terminal nodes in real time, optimizing the storage and backup time points, and improving the utilization rate of storage resources and the security of data, reducing the failed storage and backup of non-online state nodes, and improving the utilization rate of storage resources;

[0128] Traditional data distributed cutting storage adopts a uniform cutting method, without considering the reliability index of the terminal node in the corresponding node network. The first allocation algorithm allocates the data block size according to the stability index of the terminal node in the active period. The second allocation algorithm allocates the data block size in the backup source according to the stability index of the distributed backup node in the corresponding active period. It considers the stability and active period of the terminal nodes in the node network, improves the rationality of data storage and backup, and allocates larger data blocks to terminal nodes with high stability index, improves the efficiency and reliability of data storage and backup, and avoids the waste of resources and imbalance problems that may be caused by uniform allocation.

[0129] Multi-location backup and data storage are performed through a periodically constructed node network, which effectively solves the problems of insufficient dynamic adaptability and storage resource congestion in traditional cloud storage data.

[0130] Embodiment 2: In the traditional scheme, the cutting order of data blocks is fixed or determined based on simple rules. Setting it based only on the order of a single data block size has great limitations. This may be due to ignoring the order of the importance of the target data content, resulting in sub-blocks with larger important contribution factors being allocated to terminal nodes with poor stability, increasing storage burden and hidden dangers.

[0131] Therefore, the embodiments of the present application are optimized to a certain extent based on the above embodiments.

[0132] In some embodiments, in step S2, the method of obtaining the preset order includes:

[0133] Perform content analysis on the target data and divide it into multiple sub-blocks, so that each sub-block contains a relatively independent data portion;

[0134] Calculate the important contribution factor of each sub-block to the target data: weighted sum the content proportion of the sub-block in the target data, the business value (which can be set based on expert experience), and the similarity of the historical uploaded data of the corresponding traceability point to obtain the important contribution factor of the sub-block; the weight factor is set according to the actual situation, mainly based on which parameter indicator the important contribution factor depends more on, the corresponding weight factor is larger;

[0135] Obtain the important contribution factor of each sub-block in turn according to the order of the sub-blocks in the target data (from the beginning to the end) to form an important contribution factor sequence;

[0136] Based on the size of all elements in the important contribution factor sequence, the comprehensive important trend is obtained, and the direction of the comprehensive important trend is determined as a preset order. The determination method of the comprehensive important trend includes:

[0137] The average value of all important contribution factors is obtained as the important threshold; the position of the important contribution factors greater than the important threshold in the important contribution factor sequence is determined. If the position is at the front, the descending order is determined as the comprehensive important trend; if the position is at the back, the ascending order is determined as the comprehensive important trend; the judgment standard for being at the front or back is set as: if it is before the middle position of the important contribution factor sequence, it is at the front, otherwise, it is at the back.

[0138] Therefore, the direction of dividing the data block size is the same as the comprehensive importance trend of the important contribution factor, which can avoid the situation where the data block is small but its important contribution factor is large and is allocated to a terminal node with poor stability, thereby increasing the storage burden and hidden dangers of the node. Because the important contribution factor is large and faces more severe and important content protection mechanisms, the node needs higher stability to bear greater storage pressure.

[0139] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages:

[0140] By calculating the important contribution factors of sub-blocks and determining the cutting order according to their comprehensive important trends, the cutting of data blocks is made more reasonable; sub-blocks with larger important contribution factors are avoided to the greatest extent possible from being allocated to terminal nodes with poor stability, thereby reducing storage risks;

[0141] By considering the content proportion, business value and similarity of historical uploaded data of the sub-blocks, the importance of the sub-blocks can be evaluated more accurately, and a sequence of important contribution factors of the target data can be generated according to the importance of the sub-blocks. The cutting order of the target data can be determined based on the comprehensive important trend of the sequence of important contribution factors, so as to reasonably allocate the storage data blocks and improve the storage efficiency and reliability. Based on the setting of the preset order, for sub-blocks with larger important contribution factors, it is ensured that they are allocated to terminal nodes with higher stability, which enhances the data protection mechanism, because data with higher importance requires higher stability to bear greater storage pressure.

[0142] Embodiment 3: In Embodiment 2, the preset order of the cutting size is limited, but it is only roughly determined based on the comprehensive and important trend of the important contribution factors. There must be some mismatches in the data block size, comprehensive importance, and stability index of the corresponding distributed storage nodes in some cutting intervals. This mismatch may cause data blocks with larger important contribution factors to be allocated to nodes with insufficient stability, increasing storage risks and node pressure.

[0143] Therefore, the embodiments of the present application are optimized to a certain extent based on the above embodiments.

[0144] In some embodiments, step S2 further includes:

[0145] S301. Determine all the cutting points for cutting the target data in the preset order as the initial cutting axes. All the cutting points are sequentially arranged on the initial cutting point axes. Each two adjacent cutting points form a cutting interval segment, and each cutting interval segment corresponds to an initial data block.

[0146] S302. Record the sub-blocks falling into each cutting interval segment and their important contribution factors, and calculate the comprehensive important contribution degree of each cutting interval segment (taking the average value of the important contribution factors of all the sub-blocks falling into this cutting interval segment).

[0147] S303. Generate the mapping feature vectors of the cutting interval segments based on the comprehensive important contribution degree of each cutting interval segment, the proportion of the initial data block size (the ratio of the initial data block size to the target data size), and the stability index of the corresponding allocated distributed storage node, and input them into the pre-trained mapping relationship anomaly recognition model to output the recognition results of the cutting interval segments. The recognition results include normal and abnormal and their abnormal types. The abnormal types are set as the over-standard of the comprehensive important contribution degree and its over-standard ratio, and the shortage of the comprehensive important contribution degree and its shortage ratio.

[0148] In some embodiments, the acquisition method of the pre-trained mapping relationship anomaly recognition model includes:

[0149] E1. Collect the mapping feature vectors of the cutting interval segments corresponding to a large number of data blocks allocated to the terminal nodes in history. Based on each historical mapping feature vector, perform label annotation on it, and the annotation content is set as normal, abnormal and their abnormal types.

[0150] Specifically, the annotation method is specifically as follows:

[0151] Obtain the change range of the stability index of the terminal node within the monitoring period (which can be set to 2 hours and adjusted according to the actual situation) after the data block corresponding to the historical mapping feature vector is allocated to the corresponding terminal node for storage. (The calculation method of the stability index can refer to step A2, which will not be elaborated in this invention. The change range of the stability index is obtained based on the stability index of each time node within the monitoring period, and the time node can be set to 5 minutes);

[0152] If the change range of the stability index is greater than the preset amplitude threshold, it is marked as abnormal, and its comprehensive important contribution degree and the stability index of the corresponding distributed storage node allocated are obtained, and the abnormal type is obtained based on the preset standard rules; wherein, the preset standard rules are set as follows: based on the stability index level interval, the corresponding important contribution degree interval of the acceptable data block is matched, the stability index level interval and its important contribution degree interval to which the stability index of the terminal node belongs during the monitoring period are obtained. If the comprehensive important contribution degree of the historical mapping feature vector is greater than the upper limit value of the important contribution degree interval, the abnormal type is determined to be the over-standard of the comprehensive important contribution degree and the over-standard ratio (the difference between the comprehensive important contribution degree and the upper limit value of the important contribution degree interval / the upper limit value). If the comprehensive important contribution degree of the historical mapping feature vector is less than the lower limit value of the important contribution degree interval, the abnormal type is determined to be the shortage of the comprehensive important contribution degree and the shortage ratio (the difference between the comprehensive important contribution degree and the lower limit value of the important contribution degree interval / the lower limit value). Otherwise, the label is changed to normal;

[0153] Otherwise, it is marked as normal.

[0154] It should be noted that the stability index level interval is preset according to experts or historical experience. Three level intervals of stability index are divided, corresponding to low stability, medium stability, and high stability respectively, and an important contribution degree interval of an acceptable data block is matched for each stability index level interval.

[0155] E2. Use the labeled mapping feature vector as the training set, and use the training set to train the pre-selected neural network structure, optimize the model parameters, and generate the final mapping relationship anomaly recognition model.

[0156] S304. Obtain the abnormal cutting interval segment based on the recognition result, and use the preset local adjustment strategy to update the size of the initial data block of the abnormal cutting interval segment to the target size to obtain the target data block.

[0157] Among them, the preset local adjustment strategy includes:

[0158] D1. Obtain the abnormal type of the abnormal cutting interval segment. If the abnormal type is the over-standard of the comprehensive important contribution degree, use the first local adjustment algorithm to obtain the target size. If the abnormal type is the shortage of the comprehensive important contribution degree, use the second local adjustment algorithm to obtain the target size.

[0159] Among them, if the comprehensive important contribution degree exceeds the standard, it indicates that when the data block size adapts to the stability index of the corresponding distributed storage node allocated, the stability index of the corresponding distributed storage node is too low relative to this comprehensive important contribution degree, and it is necessary to reduce the data block size to relieve the storage pressure and potential risks of the node; if the comprehensive important contribution degree is short, it indicates that when the data block size adapts to the stability index of the corresponding distributed storage node allocated, the stability index of the corresponding distributed storage node is too high relative to this comprehensive important contribution degree, and it is necessary to increase the data block size to make full use of the storage capacity of the node.

[0160] Among them, the first local adjustment algorithm is: multiplying the size of the initial data block by the excess ratio to obtain the redundant size, and subtracting the redundant size from the size of the initial data block to obtain the target size;

[0161] The second local adjustment algorithm is: multiplying the size of the initial data block by the shortage ratio to obtain the supplementary size, and adding the supplementary size to the size of the initial data block to obtain the target size.

[0162] D2. Dynamically move the reference point of the cutting interval segment (when the abnormal cutting interval segment is at the head end and the middle, it is set as the end cutting point of the cutting interval segment; when the abnormal cutting interval segment is at the tail end, it is set as the start cutting point of the cutting interval segment), so that the size of the initial data block corresponding to this cutting interval segment is adjusted to the target size to obtain the target data block.

[0163] S305. Update the cutting interval segments on the initial cutting point axis of the target data to obtain the target cutting point axis, and cut the target data according to the target cutting point axis to obtain N data blocks.

[0164] It should be noted that the updated abnormal cutting interval segment can be input into the mapping relationship anomaly recognition model to finally check whether the adjustment is successful, or the cutting interval segments in the target cutting point axis can be input into the mapping relationship anomaly recognition model for re-anomaly detection. Steps S304 and S305 are repeatedly executed until there is no abnormal cutting interval segment or the normal rate of the cutting interval segments in the target cutting point axis reaches a preset probability value (which can be set according to the actual situation and experience and is used to limit the overall segmentation reliability), and then the loop stops.

[0165] The technical solutions in the above embodiments of the present application at least have the following technical effects or advantages:

[0166] Through the mapping relationship anomaly recognition model, the abnormal results of the cutting interval segment can be accurately recognized, providing an accurate basis for subsequent adjustment; according to the type of anomaly, the local adjustment strategy is used to adjust the size of the abnormal cutting interval segment, making the matching between the data block size and the degree of important contribution and the stability index of the corresponding distributed storage node more balanced; by dynamically adjusting the position of the data block cutting point through anomaly detection, it is avoided that when the data block is small but its important contribution factor is large, it is allocated to a terminal node with poor stability, thus improving the storage efficiency and the stability of the node;

[0167] By introducing the mapping relationship anomaly recognition model and the local adjustment strategy, the refined adjustment of data block cutting is realized, which not only improves the rationality of data segmentation, but also enhances the stability and efficiency of data storage; solves the problem of mismatch between the data block size, the degree of important contribution of the content and the stability index of the distributed storage node; by dynamically adjusting the data block size, the adaptability of the data block to the node stability index is ensured, thus optimizing the allocation of storage resources; improves the reliability and performance of the data storage system; by accurately identifying and processing the abnormal cutting interval segment, the storage hidden danger and node pressure are reduced, and the reliability and performance of the entire storage system are improved.

[0168] Embodiment 4: In Embodiment 1, the periodic update of the node network has a preset fixed time interval, which may not accurately reflect the actual changes in the access habits of terminal nodes. The fixed time interval may cause the system to fail to respond in time to the changes in the access habits of terminal nodes, thus affecting the efficiency and reliability of storage and backup.

[0169] Therefore, the embodiment of the present application is optimized on the basis of the above embodiments.

[0170] In some embodiments, in step S105, the method for setting the time interval is specifically as follows:

[0171] Obtain the active period of the central feature vectors of all first target objects in the current first node network, select the duration of the minimum active period as the time interval, and periodically execute S101 to S104.

[0172] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages:

[0173] By selecting the minimum active period duration of the central feature vectors of all first target objects in the first node network as the time interval, it is possible to more dynamically adapt to changes in the access habits of terminal nodes, ensuring that the node network can be updated more frequently and reasonably to reflect the latest changes in the access habits of terminal nodes, thereby improving the dynamic adaptability of storage and backup, and being able to adapt to changes in the network environment and the behavior patterns of terminal nodes more quickly; it is possible to more accurately determine when to perform data storage and backup, thereby optimizing the time points for storage and backup, which helps to improve the utilization rate of storage resources and reduce storage and backup failures caused by terminal nodes being offline.

[0174] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A cloud big data storage management method, characterized in that, Including: S101, collect the communication records of all terminal nodes connected to the cloud in all time windows within a preset historical period. The communication records include communication time period, communication duration, communication data volume, and their stability indices, and generate a multi - access habit feature vector for each terminal node, including active time period, average communication duration, average communication data volume, and average stability index; S102, based on the multi - access habit feature vectors of all terminal nodes, use a preset classification model to classify them, obtain several terminal node groups, and label each terminal node group, with the labeling content set as the central feature vector; S103, based on all terminal node groups, retrieve the central feature vectors of each terminal node group, and select the terminal node groups whose active time periods match the current time node as the first target objects. The matching condition is: the current time node is within the active time period. Based on all terminal nodes in the first target objects and the cloud, construct the first node network; Select the terminal node groups whose stability indices are greater than the stability threshold as the second target objects, determine the maximum active - time - period coverage rate combination in the second target objects, and based on the terminal nodes in all terminal node groups within the maximum active - time - period coverage rate combination and the cloud, construct the second node network; The first node network is used for distributed storage of data, and the second node network is used for distributed backup of data; S104, monitor and obtain target data and its traceability points in real - time, and generate an associated processing instruction for the target data according to a preset differential traceability processing decision. The associated processing instruction includes storage, backup, and forwarding; Use the first node network to store the target data: based on the multi - access habit feature vectors of all terminal nodes in the first node network, screen out the terminal nodes whose active time periods match the current time node as distributed storage nodes; Based on the first allocation algorithm, cut the target data into several data blocks and allocate them to the corresponding distributed storage nodes for storage; label each data block with a unique identifier and record its corresponding distributed storage node to generate an index link. Determine the storage index information of all data blocks of the target data, including the unique identifiers of all data blocks and their corresponding index links. The unique identifier is used to combine all data blocks to obtain the original complete target data; Use the second node network to back up the target data: based on the number of terminal node groups in the maximum active - time - period coverage rate combination, determine the number of backup sources for the target data, and each backup source is used to perform a backup of the target data once; Based on each backup source, use all the terminal nodes therein as distributed backup nodes; Based on the second allocation algorithm, cut the target data into several data blocks and allocate them to the corresponding distributed backup nodes for backup; Generate all backup sources of the target data, label each data block in each backup source with a unique identifier, record the corresponding distributed backup nodes to generate an index link, and determine the backup source of the target data, the unique identifiers of all data blocks in the backup source, and the index links corresponding to the data blocks respectively as the backup index information of the target data. The unique identifier is used to combine all data blocks to obtain the original complete target data; The first allocation algorithm includes: S1. Calculate the sizes of all data blocks obtained by cutting the target data block according to the following formula: , is the size of the data block allocated to each distributed storage node, D is the size of the target data, is the stability index of the i-th distributed storage node, and N is the total number of distributed storage nodes; S2. Based on the cutting sizes of the target data, cut it sequentially in a preset order to obtain N data blocks, and allocate and store them in the corresponding distributed storage nodes according to the size of each data block; The second distribution algorithm includes: S3. Calculate the sizes of all data blocks obtained by cutting the target data block in each backup source according to the following formula: , is the size of the data block allocated to each distributed backup node in this backup source, D is the size of the target data, is the stability index of the j-th distributed backup node in this backup source, and M is the total number of distributed backup nodes in this backup source; S4. For each backup source, based on the cutting size corresponding to the target data, sequentially perform segmentation in a preset order to obtain M data blocks, and store them in the distributed backup nodes corresponding to the size based on the size of each data block; S105. Periodically execute steps S101 to S104 at a preset time interval.

2. The cloud big data storage management method according to claim 1, wherein The generation of the multi - access habit feature vector of each terminal node includes: A1. Obtain the communication time periods in all communication records collected by this terminal node, draw a communication time distribution diagram, obtain the communication times with a density greater than the density threshold, and form the active time periods; A2. Obtain the communication success rate of each communication data volume as the stability index of data communication; A3. Obtain all communication durations, communication data volumes, and stability indexes during the active time periods in the communication records, and respectively take the average to obtain the average communication duration, average communication data volume, and average stability index; A4. Based on the active time periods, average communication duration, average communication data volume, and average stability index of each terminal node, construct a multi - access habit feature vector.

3. The cloud big data storage management method according to claim 2, characterized in that, The preset classification model is set as the K - means clustering algorithm. The specific steps of S102 include: Input all multi - access habit features into the classification model to generate k clustering results. Each clustering corresponds to a terminal node group, and use the clustering center of each clustering as the label of the corresponding terminal node group.

4. The cloud big data storage management method according to claim 2, wherein Generate an association processing instruction for the target data according to a preset differential traceability processing decision, specifically including: S201. If the traceability point is a terminal node, use the first node network and the second node network to store and back up the target data respectively, generate index information of the target data and update it to a preset cloud index library; the index information includes storage index information and backup index information; S202. If the traceability point is the cloud, retrieve the index information of the target data in the cloud index library and forward it to the target terminal node.

5. The cloud big data storage management method according to claim 1, characterized in that In S2, the preset order acquisition method includes: Conduct content analysis on the target data and divide it into multiple sub - blocks so that each sub - block contains a relatively independent data part; Calculate the important contribution factor of each sub - block to the target data: perform a weighted sum of the content proportion of the sub - block in the target data, business value, and similarity with the historical upload data of its traceability point to obtain the important contribution factor of this sub - block; Obtain the important contribution factors of each sub - block in sequence according to the order of the sub - blocks in the target data to form an important contribution factor sequence; Based on the magnitudes of all elements in the important contribution factor sequence, obtain its comprehensive important trend, and determine the direction of the comprehensive important trend as the preset order; Among them, the determination method of the comprehensive important trend includes: The average value of all important contribution factors is obtained as the important threshold; the position of the important contribution factors greater than the important threshold in the important contribution factor sequence is determined. If the position is at the front, the descending order is determined as the comprehensive important trend; if the position is at the back, the ascending order is determined as the comprehensive important trend.

6. The cloud big data storage management method according to claim 5, wherein, The S2 further includes: S301, determining all cutting points for cutting the target data in a preset order as an initial cutting axis, on which all cutting points are sequentially arranged, and every two adjacent cutting points form a cutting interval segment, and each cutting interval segment corresponds to an initial data block; S302, recording the sub-blocks falling into each cutting interval and their important contribution factors, and calculating the comprehensive important contribution degree of each cutting interval; S303, based on the comprehensive important contribution degree of each cutting interval segment, the initial data block size ratio, and the stability index of the corresponding allocated distribution storage node, a mapping feature vector of the cutting interval segment is generated, and the feature vector is input into a pre-trained mapping relationship abnormality recognition model, and the recognition result of the cutting interval segment is output, the recognition result includes normal and abnormal and their abnormality types, and the abnormality type is set to the comprehensive important contribution degree exceeding the standard and the exceeding proportion, or the comprehensive important contribution degree shortage and the shortage proportion; S304, obtaining an abnormal cutting interval segment based on the recognition result, and using a preset local adjustment strategy to update the size of the initial data block of the abnormal cutting interval segment to a target size, thereby obtaining a target data block; S305, updating the cutting interval segment on the initial cutting point axis of the target data to obtain the target cutting point axis, and cutting the target data according to the target cutting point axis to obtain N data blocks.

7. The cloud big data storage management method according to claim 6, characterized in that The method for obtaining the pre-trained mapping relationship anomaly recognition model includes: E1. Collect the mapping feature vectors of the cut interval segments corresponding to the large number of data blocks assigned to the terminal node in history, and label each historical mapping feature vector based on it, with the label content set as normal, abnormal and its abnormal type; E2. Use the annotated mapping feature vector as a training set, use the training set to train the pre-selected neural network structure, optimize the model parameters, and generate the final mapping relationship anomaly recognition model; The preset local adjustment strategy includes: D1. Obtain the abnormal type of the abnormal cutting interval. If the abnormal type is that the comprehensive important contribution degree exceeds the standard, the first local adjustment algorithm is used to obtain the target size; if the abnormal type is that the comprehensive important contribution degree is insufficient, the second local adjustment algorithm is used to obtain the target size; D2. Dynamically move the reference point of the cutting interval segment so that the size of the initial data block corresponding to the cutting interval segment is adjusted to the target size to obtain the target data block.

Citation Information

Patent Citations

  • Cloud big data storage management method and system

    CN116974827A

  • Computer system data storage and reading method

    CN118377431A

  • Method for realizing network optimization and related device

    WO2020125716A1