Distributed storage system based on multi-copy management
By evaluating the data to be stored in a distributed storage system, dynamically determining the number of copies and storing them, the problem of low storage resource utilization efficiency in existing technologies is solved, and the rational utilization of storage resources and improvement of data reliability are achieved.
Patent Information
- Application Number
- CN202510670745.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-19
AI Technical Summary
The existing technology adopts a storage strategy of presetting the same number of copies in a distributed storage system, which leads to low efficiency in storage resource utilization and occupies a large amount of storage space.
The evaluation module evaluates the data to be stored, dynamically determines the number of copies based on the importance and access frequency of the data, generates multiple target copy data, and stores them in different storage nodes.
It achieves the rational use of storage resources, avoids resource waste, and improves data reliability and the overall performance of the storage system.
Smart Images

Figure CN120669905A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data storage, and in particular to a distributed storage system based on multi-copy management. Background Art
[0002] Distributed storage systems generally adopt a multi-copy distributed storage strategy to ensure data reliability through multi-copy redundant storage. For example, 3-copy storage can be used. After determining the node using a hash algorithm, one copy of the data is stored on the node (or machine), and the other two copies are stored on other nodes. When a node fails, the other two copies are still guaranteed to be accessible, and the failed copy can be repaired under appropriate conditions. In order to improve the performance of each node in the distributed storage system in providing business services to the outside world, each node can be sharded. Each data shard has a master copy that receives and responds to data requests and a slave copy that synchronizes data operations with the master copy. The master copy and its corresponding one or more slave copies are located on different nodes.
[0003] In the prior art, when implementing a multi-copy distributed storage strategy for data to be stored, a preset equal number of copies is often used for distributed storage without analyzing the data from a data perspective, resulting in the data to be stored occupying a large amount of storage space during distributed storage, leading to low storage efficiency.
[0004] Therefore, there is an urgent need for a distributed storage system based on multi-copy management to solve the above problems. Summary of the Invention
[0005] The present invention aims to at least partially address one of the technical problems in the above-mentioned technologies. To this end, the present invention proposes a distributed storage system based on multi-copy management. This system uses an evaluation module to evaluate the data to be stored and determines the corresponding number of copies based on the evaluation results. For highly important and frequently accessed data, more copies are allocated, while for less important data, the number of copies is reduced. This system achieves the rational utilization of storage resources, avoids resource waste, and improves the overall performance of the storage system while ensuring data reliability.
[0006] To achieve the above objectives, an embodiment of the present invention proposes a distributed storage system based on multi-copy management, including:
[0007] A first acquisition module, used to acquire data to be stored;
[0008] An evaluation module is used to evaluate the data to be stored and determine the data evaluation result;
[0009] A determination module is used to determine the number of copies of the data to be stored based on the data evaluation result, and obtain a number of target copy data;
[0010] A second acquisition module is used to obtain the storage node corresponding to the data to be stored;
[0011] The storage module is used to store the data to be stored and a plurality of target copy data in the storage node.
[0012] Preferably, it further includes: a data cleaning module, which is used to clean the data to be stored before the evaluation module performs data evaluation on the data to be stored.
[0013] Preferably, the data cleaning module includes:
[0014] The acquisition submodule is used to obtain the attribute information of the data to be stored;
[0015] a classification submodule, configured to perform cluster analysis on the attribute information of the data to be stored, and divide the data to be stored into a plurality of first sub-data to be stored based on the clustering result;
[0016] The first detection submodule is configured to:
[0017] Randomly select a first sub-data to be stored as target data;
[0018] Take any attribute value in the target data as the target attribute value;
[0019] Calculate the abnormality degree value corresponding to the target attribute value based on a first preset algorithm;
[0020] The abnormality level value is compared with the preset abnormality level threshold value, and the abnormality level value is determined to be
[0021] When the degree value is greater than or equal to the preset abnormality degree threshold, the target attribute value is used as the initial abnormal attribute value;
[0022] Traverse all attribute values in the target data to obtain several initial abnormal attribute values;
[0023] The second detection submodule is used to:
[0024] Take any initial abnormal attribute value;
[0025] Determine the target area with the initial abnormal attribute value as the center and the preset step size as the radius;
[0026] Calculate the correlation coefficients between the initial abnormal attribute values and other attribute values in the target area respectively to obtain several attribute correlation coefficients;
[0027] Sum up the correlation coefficients of several attributes to obtain the abnormal evaluation value;
[0028] Comparing the abnormality evaluation value with a preset abnormality evaluation threshold, and when it is determined that the abnormality evaluation value is greater than or equal to the preset abnormality evaluation threshold, using the initial abnormal attribute value as a target abnormal attribute value;
[0029] Traverse all initial abnormal attribute values to obtain several target abnormal attribute values;
[0030] Data cleaning submodule, used for:
[0031] Obtain data cleaning rules corresponding to target data;
[0032] Performing data cleaning on several target abnormal attribute values based on the data cleaning rules;
[0033] All the first sub-data to be stored are traversed to obtain the data to be stored after data cleaning.
[0034] Preferably, the first preset algorithm includes:
[0035]
[0036] Among them, T za represents the abnormality degree value corresponding to the a-th target attribute value in the z-th first sub-data to be stored; represents the mean of all attribute values in the zth first sub-data to be stored; m represents the total number of attribute values in the zth first sub-data to be stored; K zb represents the bth target attribute value in the zth first sub-data to be stored; norm() represents the linear normalization function; || represents the absolute value operation.
[0037] Preferably, the evaluation module includes:
[0038] A classification submodule, configured to input the data to be stored into a pre-trained data classification model for classification, and obtain a plurality of second sub-data to be stored;
[0039] Evaluation submodule, used to:
[0040] Evaluate each second sub-data to be stored based on a second preset algorithm to obtain an evaluation value corresponding to each second sub-data to be stored;
[0041] A set of evaluation values corresponding to all the second sub-data to be stored is used as a data evaluation result.
[0042] Preferably, the second preset algorithm includes:
[0043]
[0044] Among them, Q r represents the evaluation value corresponding to the rth second sub-data to be stored; ρr represents the range of the second r-th sub-data to be stored; σ r represents the variance of the rth second sub-data to be stored; θ r represents the standard deviation of the rth second sub-data to be stored; β r represents the data mean of the rth second sub-data to be stored; g r represents the number of read and write times of the data in the preset database that has the highest similarity to the rth second sub-data to be stored within the preset period; G represents the total number of read and write times of all data in the preset database within the preset period; d r represents the data size of the rth second sub-data to be stored; and D represents the total data size of the data to be stored.
[0045] Preferably, the determination module includes:
[0046] A comparison submodule, configured to compare an evaluation value corresponding to a second sub-data to be stored in any data evaluation result with a preset evaluation threshold;
[0047] Identify submodules for:
[0048] If the evaluation value corresponding to the second sub-data to be stored is greater than or equal to the preset evaluation threshold, determining that the number of copies corresponding to the second sub-data to be stored is the first number;
[0049] If the evaluation value corresponding to the second sub-data to be stored is less than the preset evaluation threshold, determining that the number of copies corresponding to the second sub-data to be stored is a second number; and the first number is greater than the second number;
[0050] Traverse all second sub-data to be stored, and obtain the number of copies corresponding to each second sub-data to be stored;
[0051] Generate submodules for:
[0052] generating a plurality of replica data corresponding to each second sub-data to be stored based on each second sub-data to be stored and the number of replicas corresponding to each second sub-data to be stored;
[0053] Traversing all the second sub-data to be stored, and obtaining a plurality of copy data corresponding to all the second sub-data to be stored;
[0054] The plurality of copy data corresponding to all the second sub-data to be stored are used as the plurality of target copy data corresponding to the data to be stored.
[0055] Preferably, the storage module includes:
[0056] Computing submodule, used for:
[0057] Select any storage node as the target storage node;
[0058] Calculate the eigenvalue of the data in the target storage node to obtain the target eigenvalue;
[0059] Query submodule, used to:
[0060] Based on the target characteristic value, the target characteristic value-replica combination storage template table is searched to determine the replica combination storage template corresponding to the target storage node;
[0061] Traverse all storage nodes to obtain several replica combination storage templates;
[0062] Storage submodule, used for:
[0063] The data to be stored and several target replica data are stored in the storage node based on several replica combination storage templates.
[0064] Preferably, the calculation submodule is used to select any storage node as a target storage node; calculate the characteristic value of the data in the target storage node, and obtain the target characteristic value, including:
[0065]
[0066] Among them, E i represents the target characteristic value of the i-th target storage node; B ij represents the data characteristics of the jth stored data in the i-th target storage node; M ij represents the amount of data stored in the jth target storage node; t ij S represents the access time of the jth storage data in the i-th target storage node; ij Indicates the length of the jth stored data in the i-th target storage node; C ij Indicates the data type of the jth stored data in the i-th target storage node; represents the rate of change of the data type C in the i-th target storage node with the access time t within a preset period; n represents the total amount of data stored in the i-th target storage node; Indicates the operation of the Laplace operator on the data type; P ij V represents the type error coefficient of the jth stored data in the i-th target storage node; ij Represents the semantic parameter of the i-th storage information in the i-th target storage node; || represents the absolute value operation; ln represents the natural logarithm.
[0067] Preferably, the determination module also includes: a synchronization sub-module, which is used to obtain the version numbers of each sub-data to be stored and the several copy data corresponding to each sub-data to be stored in real time, and when it is determined that the version number difference is not 0, each sub-data to be stored and the several copy data corresponding to each sub-data to be stored are synchronized based on the version information corresponding to the version number difference.
[0068] The present invention provides a distributed storage system based on multi-copy management, which evaluates the data to be stored through an evaluation module, determines the corresponding number of copies based on the evaluation results, and generates multiple target copy data. When a storage node fails, the data is damaged or lost, other copies can ensure the availability of the data, greatly reduce the risk of data loss, and ensure the reliability of the data; the data evaluation link can determine the number of copies based on the importance of the data, access frequency and other characteristics. For data with high importance and frequent access, more copies are allocated, and vice versa, the number of copies is reduced, thereby achieving the rational use of storage resources, avoiding resource waste, and improving the overall performance of the storage system while ensuring data reliability; multiple copies are distributed on different target storage nodes. When a user requests data, the system can select a storage node that is close to the user and has a low load to provide data based on the load balancing strategy, effectively shortening the data response time, improving data access efficiency, and improving user experience.
[0069] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.
[0070] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0072] Figure 1 is a block diagram of a distributed storage system based on multi-copy management according to an embodiment of the present invention;
[0073] Figure 2 is a block diagram of a data cleaning module according to one embodiment of the present invention;
[0074] Figure 3 is a block diagram of an evaluation module according to one embodiment of the present invention. DETAILED DESCRIPTION
[0075] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0076] Example 1
[0077] like Figure 1 As shown, a distributed storage system based on multi-copy management includes:
[0078] A first acquisition module, used to acquire data to be stored;
[0079] An evaluation module is used to evaluate the data to be stored and determine the data evaluation result;
[0080] A determination module is used to determine the number of copies of the data to be stored based on the data evaluation result, and obtain a number of target copy data;
[0081] A second acquisition module is used to obtain the storage node corresponding to the data to be stored;
[0082] The storage module is used to store the data to be stored and a plurality of target copy data in the storage node.
[0083] The working principle of the above technical solution is: the data to be stored is acquired through the first acquisition module; the evaluation module analyzes and evaluates the data to be stored, and the evaluation dimensions may include data importance, access frequency, update frequency, etc., so as to obtain the data evaluation results and provide a basis for subsequent decision-making; the determination module determines the number of copies of the data to be stored based on the data evaluation results, and obtains several target copy data to ensure that there are more copies of data with high importance or frequent access, so as to improve data availability and reading performance; the second acquisition module finds the storage nodes corresponding to the data to be stored, and these nodes constitute the various storage locations of distributed storage; the storage module stores the original data to be stored and the generated several target copy data in the selected storage nodes to complete the data storage process in the distributed system.
[0084] The beneficial effects of the above technical solution are: by determining and storing multiple copies, even if some storage nodes fail, data can still be obtained from other copies, reducing the risk of data loss; the number of copies is determined based on the data evaluation results, and multiple copies are set for important or frequently accessed data to improve reading efficiency, while avoiding excessive storage of ordinary data and rationally utilizing storage resources; different data can have different copy strategies based on evaluation, adapt to diversified data storage needs, and better cope with complex business scenarios; multiple copies are distributed on different storage nodes, which can realize parallel reading of data, reduce data access delay, and improve overall system performance.
[0085] Example 2
[0086] It also includes: a data cleaning module, which is used to clean the data to be stored before the evaluation module performs data evaluation on the data to be stored.
[0087] Example 3
[0088] like Figure 2 As shown, the data cleaning module includes:
[0089] The acquisition submodule is used to obtain the attribute information of the data to be stored;
[0090] a classification submodule, configured to perform cluster analysis on the attribute information of the data to be stored, and divide the data to be stored into a plurality of first sub-data to be stored based on the clustering result;
[0091] The first detection submodule is configured to:
[0092] Randomly select a first sub-data to be stored as target data;
[0093] Take any attribute value in the target data as the target attribute value;
[0094] Calculate the abnormality degree value corresponding to the target attribute value based on a first preset algorithm;
[0095] The abnormality level value is compared with the preset abnormality level threshold value, and the abnormality level value is determined to be
[0096] When the degree value is greater than or equal to the preset abnormality degree threshold, the target attribute value is used as the initial abnormal attribute value;
[0097] Traverse all attribute values in the target data to obtain several initial abnormal attribute values;
[0098] The second detection submodule is used to:
[0099] Take any initial abnormal attribute value;
[0100] Determine the target area with the initial abnormal attribute value as the center and the preset step size as the radius;
[0101] Calculate the correlation coefficients between the initial abnormal attribute values and other attribute values in the target area respectively to obtain several attribute correlation coefficients;
[0102] Sum up the correlation coefficients of several attributes to obtain the abnormal evaluation value;
[0103] Comparing the abnormality evaluation value with a preset abnormality evaluation threshold, and when it is determined that the abnormality evaluation value is greater than or equal to the preset abnormality evaluation threshold, using the initial abnormal attribute value as a target abnormal attribute value;
[0104] Traverse all initial abnormal attribute values to obtain several target abnormal attribute values;
[0105] Data cleaning submodule, used for:
[0106] Obtain data cleaning rules corresponding to target data;
[0107] Performing data cleaning on several target abnormal attribute values based on the data cleaning rules;
[0108] All the first sub-data to be stored are traversed to obtain the data to be stored after data cleaning.
[0109] In this embodiment, the attribute information of the data to be stored includes: data type (text, image, video, etc.), data format (.txt, .jpg, .mp4, etc.), and data size (the number of bytes of storage space occupied).
[0110] In this embodiment, clustering methods include but are not limited to K-Means algorithm, DBSCAN algorithm, etc.
[0111] In this embodiment, the correlation coefficient includes a Pearson correlation coefficient.
[0112] In this embodiment, the data cleaning rules include: checking required fields, handling missing values, data format checking, data value range checking, logical consistency checking, record uniqueness checking, etc.
[0113] The working principle of the above technical solution is: the acquisition submodule is responsible for obtaining the attribute information of the data to be stored. This attribute information is the basis for subsequent processing and covers various characteristics of the data, such as data type, value range, etc.; the classification submodule performs cluster analysis on the acquired attribute information of the data to be stored, and divides the data to be stored into several first sub-data to be stored based on the similarity of the data attributes. The purpose of this is to group the data so that more targeted processing can be performed based on the data characteristics of different groups; the first stage of abnormal attribute value detection: the first detection submodule first selects any one from all the first sub-data to be stored as the target data, and then selects any attribute value from the target data as the target attribute value; the abnormality degree value corresponding to the target attribute value is calculated based on the first preset algorithm. This algorithm will evaluate the abnormality degree of the attribute value based on factors such as the statistical characteristics of the data; the calculated abnormality degree value is compared with the preset abnormality degree threshold. If the abnormality degree value is greater than or equal to the preset abnormality degree threshold, the target attribute value is used as the initial abnormal attribute value; repeat the above steps, traverse all attribute values in the target data, and thus obtain several initial abnormal attribute values; the second stage of abnormal attribute value determination: the second detection submodule selects any one from several initial abnormal attribute values obtained from the first detection submodule; with the initial abnormal attribute value as the center and the preset step size as the radius, a target area is determined. The area contains other attributes within a certain range from the initial abnormal attribute value. value; calculate the correlation coefficient between the initial abnormal attribute value and other attribute values in the target area respectively, and obtain several attribute correlation coefficients, which reflect the degree of association between the attribute values; sum up these attribute correlation coefficients to obtain an abnormal evaluation value, which comprehensively reflects the association between the initial abnormal attribute value and the surrounding attribute values; compare the abnormal evaluation value with the preset abnormal evaluation threshold. If the abnormal evaluation value is greater than or equal to the preset abnormal evaluation threshold, the initial abnormal attribute value is used as the target abnormal attribute value; traverse all the initial abnormal attribute values to finally obtain several target abnormal attribute values; the data cleaning submodule first obtains the data cleaning rules corresponding to the target data, which include requirements in terms of completeness, accuracy, uniqueness, etc.; based on the obtained data cleaning rules, several target abnormal attribute values are cleaned, such as filling missing values, correcting erroneous values, etc.; repeat the above operations on the target data, traverse all the first sub-data to be stored, and finally obtain the data to be stored after data cleaning, completing the entire data cleaning process.
[0114] The beneficial effects of the above technical solution are: the first detection submodule calculates the abnormality degree value based on the preset algorithm, which can effectively identify attribute values that may be abnormal and use them as initial abnormal attribute values; the second detection submodule further analyzes the correlation between the attribute values to determine the target abnormal attribute value. This dual detection mechanism greatly improves the accuracy of identifying abnormal data, thereby laying the foundation for subsequent accurate data cleaning and improving the accuracy of the data; when determining the target abnormal attribute value, the second detection submodule considers the correlation between the attribute values, determines the target area with the initial abnormal attribute value as the center, and calculates the correlation coefficient. This method not only focuses on the abnormality degree of a single attribute value, but also considers its correlation with the surrounding attribute values, so that the final determined target abnormal attribute value is more in line with the internal logic of the data and enhances the rationality of the data; the classification submodule performs cluster analysis on the attribute information of the data to be stored, and divides the data into several first sub-data to be stored, so that subsequent data detection and cleaning can be carried out according to the data characteristics of different clusters, thereby improving the pertinence and efficiency of data cleaning. Data from different clusters may have different distribution characteristics and abnormal patterns. Processing them separately can better meet the cleaning needs of each type of data. The acquisition submodule obtains the attribute information of the data to be stored, providing a comprehensive data foundation for subsequent analysis. Because attribute information covers multiple characteristics of the data, the data cleaning module can adapt to different types of data. Whether it is numerical, character, or other types of data, it can be effectively cleaned by analyzing the attribute information, which has strong versatility. The data cleaning submodule performs cleaning operations according to the data cleaning rules corresponding to the target data, ensuring the standardization and consistency of the data cleaning process. These rules can cover multiple requirements such as completeness, accuracy, and uniqueness, ensuring that the cleaned data meets specific quality standards, which helps to improve the reliability of the data during subsequent storage and use. By traversing all the first sub-data to be stored and the attribute values therein, the module can comprehensively detect and process anomalies in the data, not missing any data points that may have problems, thereby effectively improving the overall quality of the data and reducing the impact of abnormal data on subsequent data analysis and application.
[0115] Example 4
[0116] The first preset algorithm includes:
[0117]
[0118] Among them, T za represents the abnormality degree value corresponding to the a-th target attribute value in the z-th first sub-data to be stored; represents the mean of all attribute values in the zth first sub-data to be stored; m represents the total number of attribute values in the zth first sub-data to be stored; K zbrepresents the bth target attribute value in the zth first sub-data to be stored; norm() represents the linear normalization function; || represents the absolute value operation.
[0119] The beneficial effect of the above technical solution is that it can accurately quantify the degree of deviation of each target attribute value relative to other attribute values, thereby obtaining a reasonable degree of abnormality value. The greater the deviation of the attribute value from the mean, the higher the degree of abnormality, which helps to quickly identify abnormal attribute values that may exist in the data. The calculation process uses the mean of all attribute values in the sub-data and the sum of the remaining attribute values, and comprehensively considers the overall distribution of attribute values in the sub-data. This means that rather than judging whether an attribute value is abnormal in isolation, it is evaluated based on the characteristics of the entire sub-data, making the judgment of the degree of abnormality more scientific and reasonable, and better reflecting the inherent laws of the data.
[0120] Example 5
[0121] like Figure 3 As shown, the evaluation module includes:
[0122] A classification submodule, configured to input the data to be stored into a pre-trained data classification model for classification, and obtain a plurality of second sub-data to be stored;
[0123] Evaluation submodule, used to:
[0124] Evaluate each second sub-data to be stored based on a second preset algorithm to obtain an evaluation value corresponding to each second sub-data to be stored;
[0125] A set of evaluation values corresponding to all the second sub-data to be stored is used as a data evaluation result.
[0126] In this embodiment, the data classification model training method includes:
[0127] Obtain data classification training dataset;
[0128] Inputting the data classification training data set into a neural network model for training to obtain an initial data classification model;
[0129] Get the data classification test dataset;
[0130] The initial data classification model is tested based on the data classification test data set. When the test results are qualified, a trained data classification model is obtained.
[0131] The working principle of the above technical solution is as follows: the classification submodule provides the data to be stored as input to a pre-trained data classification model. This data classification model is trained on a large amount of data, learning the data's characteristics and patterns and enabling effective classification of the input data. The data to be stored may contain information of various types and formats, such as text and numbers. This data is processed by the model according to its internal algorithm and the knowledge learned from training. The data classification model analyzes the input data to be stored based on its own algorithm and trained parameters. It extracts features from the data and then divides the data into different categories based on these features, thereby generating a number of second sub-data to be stored. For example, if classifying customer information data, the model may divide the customer information data into different groups based on characteristics such as age and consumption habits. Each group corresponds to a second sub-data to be stored. After the model's classification processing, the classification submodule outputs a number of second sub-data to be stored. These sub-data are divided based on data similarity or other classification criteria and serve as the basis for subsequent evaluation. The evaluation submodule evaluates each second sub-data to be stored obtained by the classification submodule using a second preset algorithm. The second preset algorithm is a predefined calculation rule. It will analyze and calculate the data from different angles according to the specific characteristics and requirements of the second sub-data to be stored, such as considering factors such as the integrity, accuracy, and importance of the data, and finally obtain the evaluation value corresponding to each second sub-data to be stored. This evaluation value can reflect the quality or characteristics of the sub-data in certain aspects; the evaluation submodule collects the evaluation values corresponding to all the second sub-data to be stored and forms a set. This set contains the evaluation results of all sub-data and is output as a data evaluation result. In this way, the evaluation results of each sub-data are integrated so that the data evaluation results can comprehensively reflect the overall situation of the data to be stored after classification, and provide a valuable reference for other subsequent operations based on the data evaluation results (such as determining the number of copies, etc.).
[0132] The beneficial effects of the above technical solution are as follows: the classification submodule uses a pre-trained data classification model to classify the data to be stored and divides it into several second sub-data to be stored. In this way, subsequent processing can be carried out in a targeted manner according to the characteristics of different sub-data. For example, a stricter storage strategy can be adopted for sub-data with high importance, and a more efficient storage method can be adopted for sub-data with smaller data volume, thereby improving the pertinence and efficiency of data processing; the evaluation submodule evaluates each second sub-data to be stored based on the second preset algorithm, and can perform a detailed analysis of the data from multiple dimensions to obtain the evaluation value corresponding to each sub-data. This refined evaluation method can more accurately reflect the quality, importance and other characteristics of the data, avoid the limitations of using a unified evaluation standard for all data, and provide a more reliable basis for data management and decision-making; the evaluation values corresponding to all second sub-data to be stored are combined into a set as the data evaluation result, so that the evaluation result can comprehensively cover all parts of the data to be stored. Decision makers can understand the evaluation status of different sub-data through this set, thereby having a clearer understanding of the overall data, which helps to formulate more reasonable storage, management and use strategies.
[0133] Example 6
[0134] The second preset algorithm includes:
[0135]
[0136] Among them, Q r represents the evaluation value corresponding to the rth second sub-data to be stored; ρ r represents the range of the second r-th sub-data to be stored; σ r represents the variance of the rth second sub-data to be stored; θ r represents the standard deviation of the rth second sub-data to be stored; β r represents the data mean of the rth second sub-data to be stored; g r represents the number of read and write times of the data in the preset database that has the highest similarity to the rth second sub-data to be stored within the preset period; G represents the total number of read and write times of all data in the preset database within the preset period; d r represents the data size of the rth second sub-data to be stored; and D represents the total data size of the data to be stored.
[0137] The beneficial effect of the above technical solution is that it comprehensively considers data characteristics from multiple dimensions to calculate the evaluation value. This includes the statistical characteristics of the data (range, variance, standard deviation, mean), which reflect the distribution and degree of dispersion of the data; it also considers characteristics related to the frequency of use of the data (the number of reads and writes of similar data in the preset database, the total number of reads and writes of all data in the preset database), and data volume characteristics (the amount of sub-data, the total amount of data to be stored). By integrating this multi-dimensional information, the value and importance of each second sub-data to be stored can be more comprehensively and accurately evaluated.
[0138] Example 7
[0139] Identify modules, including:
[0140] A comparison submodule, configured to compare an evaluation value corresponding to a second sub-data to be stored in any data evaluation result with a preset evaluation threshold;
[0141] Identify submodules for:
[0142] If the evaluation value corresponding to the second sub-data to be stored is greater than or equal to the preset evaluation threshold, determining that the number of copies corresponding to the second sub-data to be stored is the first number;
[0143] If the evaluation value corresponding to the second sub-data to be stored is less than the preset evaluation threshold, determining that the number of copies corresponding to the second sub-data to be stored is a second number; and the first number is greater than the second number;
[0144] Traverse all second sub-data to be stored, and obtain the number of copies corresponding to each second sub-data to be stored;
[0145] Generate submodules for:
[0146] generating a plurality of replica data corresponding to each second sub-data to be stored based on each second sub-data to be stored and the number of replicas corresponding to each second sub-data to be stored;
[0147] Traversing all the second sub-data to be stored, and obtaining a plurality of copy data corresponding to all the second sub-data to be stored;
[0148] The plurality of copy data corresponding to all the second sub-data to be stored are used as the plurality of target copy data corresponding to the data to be stored.
[0149] The working principle of the above technical solution is: the comparison submodule arbitrarily selects an evaluation value corresponding to the second sub-data to be stored from the data evaluation result (i.e., a set consisting of evaluation values corresponding to all second sub-data to be stored), and then compares the evaluation value with a preset evaluation threshold value. This step is the basis for the subsequent determination of the number of copies. By comparing the evaluation value and the threshold value, it provides a basis for judging the importance of the sub-data; the stator module makes a judgment based on the comparison result of the comparison submodule. If the evaluation value corresponding to a second sub-data to be stored is greater than or equal to the preset evaluation threshold value, it means that the sub-data has a higher value or importance in the evaluation system, and therefore the number of copies corresponding to the second sub-data to be stored is determined to be the first number; if the evaluation value corresponding to a second sub-data to be stored is less than the preset evaluation threshold value, it means that the sub-data is relatively unimportant. At this time, the number of copies corresponding to the second sub-data to be stored is determined to be the second number, and it is known that the first number is greater than the second number. In this way, different numbers of copies are allocated to different second sub-data to be stored according to the evaluation of the sub-data; the determination sub-module repeats the above-mentioned judgment and determination of the number of copies, and traverses all the second sub-data to be stored in the data evaluation result, thereby obtaining the number of copies corresponding to each second sub-data to be stored; the generation sub-module generates a corresponding number of copy data for each second sub-data to be stored based on each second sub-data to be stored and the number of copies corresponding to each second sub-data to be stored obtained by the determination sub-module. For example, if the number of copies corresponding to a second sub-data to be stored is n, the generation sub-module will generate n copies of the sub-data; the generation sub-module traverses all the second sub-data to be stored, repeats the above-mentioned operation of generating copy data, and obtains a number of copy data corresponding to all the second sub-data to be stored; finally, the generation sub-module integrates the number of copy data corresponding to all the second sub-data to be stored as a number of target copy data corresponding to the data to be stored, and these target copy data will be used for subsequent data storage and other operations.
[0150] The beneficial effect of the above technical solution is: according to the comparison result of the evaluation value of the second sub-data to be stored and the preset evaluation threshold, different numbers of copies are determined for different sub-data. For sub-data with high evaluation values and high importance, more copies (first number) are allocated, and for sub-data with low evaluation values and relatively unimportant, fewer copies (second number) are allocated. In this way, the same number of copies can be avoided from being allocated to all data, thereby making more rational use of storage resources, while ensuring the reliability of important data and reducing storage costs; for important data with evaluation values greater than or equal to the preset evaluation threshold, a larger number of copies (first number) is determined. When some storage nodes fail or data is lost, since there are more copies, the data can be restored with a greater probability, ensuring the availability and integrity of the data, and improving the reliability of important data; by evaluating each second sub-data to be stored and determining the corresponding number of copies, data management is made more refined. Administrators can better plan data storage, backup and recovery strategies based on the number of copies set, and improve the efficiency and pertinence of data management.
[0151] Example 8
[0152] Storage module, including:
[0153] Computing submodule, used for:
[0154] Select any storage node as the target storage node;
[0155] Calculate the eigenvalue of the data in the target storage node to obtain the target eigenvalue;
[0156] Query submodule, used to:
[0157] Based on the target characteristic value, the target characteristic value-replica combination storage template table is searched to determine the replica combination storage template corresponding to the target storage node;
[0158] Traverse all storage nodes to obtain several replica combination storage templates;
[0159] Storage submodule, used for:
[0160] The data to be stored and several target replica data are stored in the storage node based on several replica combination storage templates.
[0161] In this embodiment, the target feature value-replica combination storage template table is a data table of the mapping relationship between the preset target feature values and the replica combination storage templates, and the replica combination storage refers to storage schemes of different replica combinations.
[0162] The beneficial effects of the above technical solution are: by calculating the characteristic values of the data in each storage node and querying the corresponding replica combination storage template based on the characteristic values, it is possible to determine the appropriate storage method for different storage nodes. In this way, the storage location of the data to be stored and the replica data can be reasonably arranged according to the characteristics of the storage node and the existing data situation, the overall storage layout can be optimized, and the space utilization rate of the storage system can be improved; the storage submodule stores the data in the storage node based on the determined replica combination storage template, making the data storage process more standardized and orderly. It avoids the confusion and inefficiency caused by blind storage, reduces the time overhead in the data storage process, improves the efficiency of data storage, and can complete data storage operations more quickly; different storage nodes may have different performance, capacity and other characteristics. The replica combination storage template table matches the corresponding storage template for the storage node according to the data characteristic value of the storage node, and can perform data storage according to the specific situation of each storage node. For example, for storage nodes with higher performance, more important data or frequently accessed data copies can be allocated, thereby better leveraging the advantages of the storage node and improving data access performance.
[0163] Example 9
[0164] The calculation submodule is used to select any storage node as a target storage node; calculate the characteristic value of the data in the target storage node, and obtain the target characteristic value by:
[0165]
[0166] Among them, E i represents the target characteristic value of the i-th target storage node; B ij represents the data characteristics of the jth stored data in the i-th target storage node; M ij represents the amount of data stored in the jth target storage node; t ij S represents the access time of the jth storage data in the i-th target storage node; ij Indicates the length of the jth stored data in the i-th target storage node; C ij Indicates the data type of the jth stored data in the i-th target storage node; represents the rate of change of the data type C in the i-th target storage node with the access time t within a preset period; n represents the total amount of data stored in the i-th target storage node; Indicates the operation of the Laplace operator on the data type; P ij V represents the type error coefficient of the jth stored data in the i-th target storage node; ij Represents the semantic parameter of the i-th storage information in the i-th target storage node; || represents the absolute value operation; ln represents the natural logarithm.
[0167] The beneficial effects of the above technical solution are: it can comprehensively and meticulously reflect the comprehensive characteristics of the data in the target storage node, providing a rich and accurate basis for the formulation of subsequent storage strategies; it can quantify the correlation between different data characteristics. The interaction and influence between different characteristics are reflected through mathematical operations, so that the calculation of characteristic values is not just a simple listing of characteristics, but takes into account the inherent connections between them, more accurately reflecting the actual situation of the data in the target storage node; data that is accessed recently or frequently will be reflected accordingly in the characteristic value calculation, which helps the storage system to reasonably allocate storage resources according to the access situation of the data. Frequently accessed data can be placed in a storage location that is easier to read or more storage resources can be allocated, thereby improving data access efficiency.
[0168] Example 10
[0169] The determination module also includes: a synchronization sub-module, which is used to obtain the version numbers of each sub-data to be stored and the several copy data corresponding to each sub-data to be stored in real time. When it is determined that the version number difference is not 0, each sub-data to be stored and the several copy data corresponding to each sub-data to be stored are synchronized based on the version information corresponding to the version number difference.
[0170] The beneficial effects of the above technical solution are: obtaining the version number of each sub-data to be stored and its copy data in real time, and performing synchronization operations when the difference in the version numbers is not 0. This can ensure that in a distributed storage environment, all copies of the same sub-data to be stored are consistent in version. When a copy of data is updated, other copies of data can be updated to the same version in a timely manner through the comparison and synchronization mechanism of the version number, avoiding data reading errors or data conflicts caused by inconsistent data versions, thereby ensuring data consistency and accuracy; during the storage and use of data, due to various reasons (such as network delays, storage node failures, etc.), some copy data may not be updated to the latest version in a timely manner. By monitoring the version number, the synchronization submodule can detect and solve these problems in a timely manner to ensure that all copy data can reflect the latest information. This increases the redundancy and fault tolerance of the data. Even if problems occur in some copy data, other copy data can still provide accurate information, thereby improving the reliability and availability of the data.
[0171] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A distributed storage system based on multi-copy management, characterized in that: include: A first acquisition module, used to acquire data to be stored; An evaluation module is used to evaluate the data to be stored and determine the data evaluation result; A determination module is used to determine the number of copies of the data to be stored based on the data evaluation result, and obtain a number of target copy data; A second acquisition module is used to obtain the storage node corresponding to the data to be stored; The storage module is used to store the data to be stored and a plurality of target copy data in the storage node.
2. The distributed storage system based on multi-copy management according to claim 1, characterized in that: Also includes: The data cleaning module is used to clean the data to be stored before the evaluation module performs data evaluation on the data to be stored.
3. The distributed storage system based on multi-copy management according to claim 2, characterized in that: Data cleaning module, including: The acquisition submodule is used to obtain the attribute information of the data to be stored; a classification submodule, configured to perform cluster analysis on the attribute information of the data to be stored, and divide the data to be stored into a plurality of first sub-data to be stored based on the clustering result; The first detection submodule is configured to: Randomly select a first sub-data to be stored as target data; Take any attribute value in the target data as the target attribute value; Calculate the abnormality degree value corresponding to the target attribute value based on a first preset algorithm; Comparing the abnormality degree value with a preset abnormality degree threshold, and when it is determined that the abnormality degree value is greater than or equal to the preset abnormality degree threshold, using the target attribute value as the initial abnormal attribute value; Traverse all attribute values in the target data to obtain several initial abnormal attribute values; The second detection submodule is used to: Take any initial abnormal attribute value; Determine the target area with the initial abnormal attribute value as the center and the preset step size as the radius; Calculate the correlation coefficients between the initial abnormal attribute values and other attribute values in the target area respectively to obtain several attribute correlation coefficients; Sum up the correlation coefficients of several attributes to obtain the abnormal evaluation value; Comparing the abnormality evaluation value with a preset abnormality evaluation threshold, and when it is determined that the abnormality evaluation value is greater than or equal to the preset abnormality evaluation threshold, using the initial abnormal attribute value as a target abnormal attribute value; Traverse all initial abnormal attribute values to obtain several target abnormal attribute values; The data cleaning submodule is used to: Obtain data cleaning rules corresponding to target data; Performing data cleaning on several target abnormal attribute values based on the data cleaning rules; All the first sub-data to be stored are traversed to obtain the data to be stored after data cleaning.
4. The distributed storage system based on multi-copy management according to claim 3, characterized in that: The first preset algorithm includes: Among them, T za represents the abnormality degree value corresponding to the a-th target attribute value in the z-th first sub-data to be stored; represents the mean of all attribute values in the zth first sub-data to be stored; m represents the total number of attribute values in the zth first sub-data to be stored; K zb represents the bth target attribute value in the zth first sub-data to be stored; norm() represents the linear normalization function; || represents the absolute value operation.
5. The distributed storage system based on multi-copy management according to claim 1, characterized in that: Assessment modules, including: A classification submodule, configured to input the data to be stored into a pre-trained data classification model for classification, and obtain a plurality of second sub-data to be stored; Evaluation submodule, used to: Evaluate each second sub-data to be stored based on a second preset algorithm to obtain an evaluation value corresponding to each second sub-data to be stored; A set of evaluation values corresponding to all the second sub-data to be stored is used as a data evaluation result.
6. The distributed storage system based on multi-copy management according to claim 5, characterized in that: The second preset algorithm includes: Among them, Q r represents the evaluation value corresponding to the rth second sub-data to be stored; ρ r represents the range of the second r-th sub-data to be stored; σ r represents the variance of the rth second sub-data to be stored; θ r represents the standard deviation of the rth second sub-data to be stored; β r represents the data mean of the rth second sub-data to be stored; g r represents the number of read and write times of the data in the preset database that has the highest similarity to the rth second sub-data to be stored within the preset period; G represents the total number of read and write times of all data in the preset database within the preset period; d r represents the data size of the rth second sub-data to be stored; and D represents the total data size of the data to be stored.
7. The distributed storage system based on multi-copy management according to claim 5, characterized in that: Identify modules, including: A comparison submodule, configured to compare an evaluation value corresponding to a second sub-data to be stored in any data evaluation result with a preset evaluation threshold; Identify submodules for: If the evaluation value corresponding to the second sub-data to be stored is greater than or equal to the preset evaluation threshold, determining that the number of copies corresponding to the second sub-data to be stored is the first number; If the evaluation value corresponding to the second sub-data to be stored is less than the preset evaluation threshold, determining that the number of copies corresponding to the second sub-data to be stored is a second number; and the first number is greater than the second number; Traverse all second sub-data to be stored, and obtain the number of copies corresponding to each second sub-data to be stored; Generate submodules for: generating a plurality of replica data corresponding to each second sub-data to be stored based on each second sub-data to be stored and the number of replicas corresponding to each second sub-data to be stored; Traversing all the second sub-data to be stored, and obtaining a plurality of copy data corresponding to all the second sub-data to be stored; The plurality of copy data corresponding to all the second sub-data to be stored are used as the plurality of target copy data corresponding to the data to be stored.
8. The distributed storage system based on multi-copy management according to claim 7, characterized in that: Storage module, including: Computing submodule, used for: Select any storage node as the target storage node; Calculate the characteristic value of the data in the target storage node to obtain the target characteristic value; Query submodule, used to: Based on the target characteristic value, query the target characteristic value-replica combination storage template table to determine the replica combination storage template corresponding to the target storage node; Traverse all storage nodes to obtain several replica combination storage templates; Storage submodule, used for: The data to be stored and several target replica data are stored in the storage node based on several replica combination storage templates.
9. The distributed storage system based on multi-copy management according to claim 8, characterized in that: The calculation submodule is used to select any storage node as the target storage node; The method of calculating the characteristic value of the data in the target storage node to obtain the target characteristic value includes: Among them, E i represents the target characteristic value of the i-th target storage node; B ij represents the data characteristics of the jth stored data in the i-th target storage node; M ij represents the amount of data stored in the jth target storage node; t ij S represents the access time of the jth storage data in the i-th target storage node; ij Indicates the length of the jth stored data in the i-th target storage node; C ij Indicates the data type of the jth stored data in the i-th target storage node; represents the rate of change of the data type C in the i-th target storage node with the access time t within a preset period; n represents the total amount of data stored in the i-th target storage node; Indicates the operation of the Laplace operator on the data type; P ij V represents the type error coefficient of the jth stored data in the i-th target storage node; ij Represents the semantic parameter of the i-th storage information in the i-th target storage node; || represents the absolute value operation; ln represents the natural logarithm.
10. The distributed storage system based on multi-copy management according to claim 5, characterized in that: The determination module also includes: a synchronization sub-module, which is used to obtain the version numbers of each sub-data to be stored and the several copy data corresponding to each sub-data to be stored in real time. When it is determined that the version number difference is not 0, each sub-data to be stored and the several copy data corresponding to each sub-data to be stored are synchronized based on the version information corresponding to the version number difference.