Data storage method and device and related equipment
By identifying error patterns in data items and formulating a block-sharing strategy, data items with the same error pattern are grouped into the same data block, and error correction coding suitable for their error characteristics is applied. This solves the problem of low data storage reliability and improves data recovery efficiency and storage reliability.
Patent Information
- Application Number
- CN202511449037.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-01-16
AI Technical Summary
In existing technologies, data items are evenly distributed across different data blocks, resulting in low storage reliability. This fails to fully utilize the error correlation between data items, leading to low data recovery efficiency.
By identifying error patterns in data items, a block-segmentation strategy is formulated based on error correlation. Data items with the same error pattern are grouped into the same data block, and error correction codes suitable for their error characteristics are applied.
It improves data recovery efficiency, enhances data storage reliability and redundancy, reduces data loss and recovery time, and lowers operating costs.
Smart Images

Figure CN121349366A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a data storage method, apparatus and related equipment. Background Technology
[0002] In the field of data processing technology, data storage primarily relies on traditional data block partitioning. Specifically, data items are evenly distributed across different data blocks to ensure convenient, distributed storage and management. This allocation method aims to reduce the risk of data loss and improve overall data accessibility through distributed storage.
[0003] However, this data storage method, indiscriminately and uniformly distributing data items into different data blocks, leads to relatively low data storage reliability. Summary of the Invention
[0004] This invention provides a data storage method, apparatus, and related equipment to solve the technical problem in related technologies where data storage methods that uniformly distribute data items to different data blocks without differentiation lead to low data storage reliability.
[0005] In a first aspect, embodiments of the present invention provide a data storage method, the method comprising:
[0006] Obtain the dataset to be stored, which includes multiple data items;
[0007] Identify the error pattern for each data item, the error pattern being used to indicate the error correlation between the data item and other data items in the dataset over time;
[0008] Based on the error patterns, a partitioning strategy for the dataset is determined to obtain at least one data block; wherein data items with the same error patterns are partitioned into the same data block.
[0009] The dataset is stored based on at least one data block.
[0010] In a second aspect, embodiments of the present invention provide a data storage device, the device comprising:
[0011] The acquisition module is used to acquire the dataset to be stored, which includes multiple data items;
[0012] An identification module is used to identify the error pattern of each data item, wherein the error pattern is used to indicate the error correlation between the data item and other data items in the dataset in the time dimension;
[0013] a block module configured to determine a block strategy of the data set based on the error pattern, and obtain at least one data block; wherein data items of the same error pattern are divided into the same data block;
[0014] a storage module configured to store the data set based on the at least one data block.
[0015] In a third aspect, an embodiment of the present application provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, and when the computer program is executed by the processor, the steps of the data storage method described above are implemented.
[0016] In a fourth aspect, an embodiment of the present application provides a readable storage medium, and the readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data storage method described above are implemented.
[0017] In a fifth aspect, an embodiment of the present application provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the data storage method described above are implemented.
[0018] In the embodiment of the present application, by obtaining a data set to be stored, the data set includes a plurality of data items; identifying an error pattern of each data item, the error pattern is used to indicate the error correlation of the data item and other data items in the data set in the time dimension; determining a block strategy of the data set based on the error pattern, and obtaining at least one data block; wherein data items of the same error pattern are divided into the same data block; and storing the data set based on the at least one data block. In this way, the error pattern of the data items in the time window can be identified, and the data items of the same error pattern can be grouped into the same data block. Thus, in data recovery, the error correlation between the data items of the same error pattern can be used to improve the data recovery efficiency, thereby improving the storage reliability of the data. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the description of the embodiments of the present application will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0020] Figure 1 is a flowchart of the data storage method provided by the embodiment of the present application;
[0021] Figure 2 is a specific flowchart of the data storage method provided by an example of the present application.
[0022] Figure 3 This is a schematic diagram of the data storage device provided in an embodiment of the present invention;
[0023] Figure 4 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] Existing technologies in data storage and processing primarily rely on traditional data block partitioning and error-correcting coding methods. Specifically, data items are evenly distributed across different data blocks to ensure convenient distributed storage and management. This allocation method aims to reduce the risk of data loss and improve overall data accessibility through distributed storage.
[0026] However, because data items are uniformly distributed across different data blocks without differentiation, and errors may occur during data block storage, some data may be unrecoverable, resulting in relatively low data storage reliability.
[0027] To further enhance data reliability, the same error-correcting coding is applied to each data block. Error-correcting coding is a technique that detects and corrects errors that may occur during data transmission or storage by adding redundant information. In existing technologies, this coding method is widely used for data blocks to ensure that even if some data is damaged, the original data can be recovered through error-correcting coding.
[0028] However, while this method can improve data reliability and recovery efficiency to some extent, it does not consider the error correlation between data items. In practical applications, data items often have a certain correlation, and some data items may be more prone to simultaneous errors. For example, in a storage system, adjacent data items may be more susceptible to the same type of failure due to their physical proximity. Existing technologies fail to fully utilize this correlation between data items to improve the efficiency of error correction coding. Therefore, in related technologies, data items are uniformly distributed across different data blocks, and this data storage method leads to relatively low data recovery efficiency when data errors occur, resulting in relatively low data storage reliability.
[0029] Therefore, although the prior art has achieved certain success in data storage and processing, there is still room for improvement. Future technological development can consider more intelligently dividing data blocks and applying more refined error correction coding methods to fully utilize the correlation between data items. This will help further improve the reliability, recovery efficiency and overall storage performance of data. At the same time, with the continuous development of big data and artificial intelligence technology, more advanced data storage and processing technologies are expected to emerge to meet the growing demand for data storage and processing.
[0030] Through the analysis of the prior art, the prior art has the following defects:
[0031] The correlation between data items when they are erroneous is ignored. Since data items are evenly distributed to different data blocks, data items with similar error patterns may be scattered to different data blocks, resulting in that the recovery efficiency cannot be improved by fully utilizing the similarities during data recovery. In addition, the traditional error correction coding method may not be the optimal choice because it is not optimized for the error characteristics of data items.
[0032] The present application aims to solve the problem of ignoring the error correlation of data items in the prior art. By introducing error correlation analysis of data items and formulating a block strategy based on the results of error correlation analysis, the error patterns of data items within a time window can be identified, and data items with similar error patterns can be grouped into the same data block. In this way, during data recovery, the recovery efficiency can be improved by utilizing data items with similar error patterns.
[0033] The following first introduces the data storage method provided in the embodiments of the present application.
[0034] It should be noted that the embodiments of the present application can relate to the field of data processing technology, and can be applied to the operation and maintenance field of information technology (IT) systems. In some embodiments, the embodiments of the present application can be applied to credit systems, transaction systems, business systems and financial systems, etc., and can store data items such as transaction records, user behavior data, credit / amount change records in these systems. Among them, the data storage in the embodiments of the present application mainly relates to distributed storage, that is, the data items in the data set are stored in a distributed manner.
[0035] Figure 1 is a flowchart of the data storage method provided by the embodiments of the present application, as shown in Figure 1 The method comprises:
[0036] Step 101, obtaining a data set to be stored, the data set comprising a plurality of data items;
[0037] Step 102, identifying an error pattern of each of the data items, the error pattern being used to indicate an error correlation of the data item with other data items in the data set in a time dimension;
[0038] Step 103, determining a chunking strategy of the data set based on the error pattern, obtaining at least one data block; wherein the data items of the same error pattern are divided into the same data block;
[0039] Step 104, storing the data set based on the at least one data block.
[0040] In step 101, the data set can be composed of multiple data items, and the data types of the multiple data items in the data set can be one or multiple. In some embodiments, the data types of the data items in the data set are multiple, for example, taking the points system as an example, the data set can include transaction records, user behavior data, points change records, and other data items of multiple data types, and in some scenarios, the data items can be chunked according to the data types, that is, combining the data types and the error correlation of the data items in the time dimension, the data items of the same data type are chunked according to the error correlation. In some scenarios, all data items can be chunked according to the error correlation without distinguishing the data types.
[0041] In some embodiments, the data set to be processed can be collected, for example, taking the points system as an example, transaction records, user behavior data, points change records, and other data items in the points system can be collected to form a data set D. The data set D is composed of multiple data items, each data item has a unique identifier, which can be a transaction identity (Identity Document, ID), a user ID, etc. The data set D can be preprocessed, including cleaning, deduplication, formatting, etc., to ensure the quality of the data. The preprocessed data set is denoted as D', which can include transaction records, user behavior data, points change records, etc. in the points system. The preprocessed data set can be referred to as the data set to be stored.
[0042] In some embodiments, the data set to be processed can be collected from various sources, and these data can come from different business systems, log files, external interfaces, etc. For example, taking the points system as an example, transaction records, user behavior data, points change records, and other data items in the points system can be collected from various sources to form a data set D. The data set D is composed of multiple data items, each data item has a unique identifier, which can be a transaction ID, a user ID, etc., used to track and locate data in subsequent processing.
[0043] After collecting the data set D, a series of preprocessing operations need to be performed to ensure the quality and consistency of the data. The preprocessing operations mainly include the following steps:
[0044] Data cleaning: Check each data item in the data set D, remove invalid transactions, error records, abnormal score changes, etc., to ensure data accuracy and integrity.
[0045] Data deduplication: There may be duplicate data items in the data set D, such as duplicate score issuance records. These duplicate items may be caused by data entry errors, system problems, etc. By comparing the identifiers and other key fields of the data items, duplicate data items can be removed to ensure the uniqueness of the data.
[0046] Data formatting: Since the data in the data set D may come from different sources, the format of the data may not be consistent. The data needs to be formatted to a unified format to facilitate subsequent processing and analysis. For example, the date field of the data item can be unified to the format of "YYYY-MM-DD", and the score value field can be unified to the format of integer, etc.
[0047] After the above preprocessing operations are completed, the preprocessed data set is obtained, denoted as D'. The score system related data items in D' have been cleaned, deduplicated, and formatted, and the quality has been guaranteed, and subsequent data storage operations can be performed.
[0048] In step 102, the error mode can refer to the error correlation of the data item with other data items in the data set in the time dimension, wherein the error correlation can refer to the correlation between different data items, which can cause these data items to be more likely to fail simultaneously. For example, in a storage system, adjacent data items can be more susceptible to the same type of failure due to their close physical location, and can fail simultaneously. Data items of the same error mode can have similar error characteristics and error correlation in the time dimension, while data items of different error modes can not have error correlation in the time dimension.
[0049] In some embodiments, the error mode of the data item can be identified according to the physical location of the data item, such as clustering data items of adjacent physical locations into the same error mode, which have error correlation in the time dimension.
[0050] In some embodiments, since data items of the same error mode have error correlation in the time dimension, such as adjacent data items that can be more susceptible to the same type of failure due to their close physical location, and their errors are related, and such as different data items that can be more likely to fail simultaneously due to certain correlation between them, accordingly, data items of the same error mode have error similarity in the time dimension, therefore, the error correlation of different data items can be reflected by statistics of the error conditions of the data items, to identify the error mode of the data items.
[0051] That is, the error patterns of the data items can be identified by the error conditions of the data items, wherein the error conditions can include at least one of the error times and the error probabilities of the data items, and can represent the error characteristics of the data items, for example, the more the error times of a data item, the greater the error probability, and accordingly, the error characteristics of the data item indicate that it is more prone to errors. Wherein the data items in the same error pattern have similar error conditions, such as similar error times and / or similar error probabilities.
[0052] In some embodiments, the data items with different error conditions can be clustered by setting the clustering range of the error conditions, so as to cluster the data items with similar error conditions together, and these data items have the same error pattern, that is, the data items in the same error pattern can have error correlation in the time dimension.
[0053] In some embodiments, the clustering analysis method in machine learning, such as the K-means algorithm, can be used to effectively identify the error patterns of the data items within the preset time window. This method can cluster the data items according to their error conditions, aiming to divide the data items into clusters with similar error conditions, each cluster representing a specific error pattern, and the data items in each cluster can have error correlation in the time dimension.
[0054] In some embodiments, the step 102 specifically includes:
[0055] determining the error conditions of each of the data items within a preset time window, wherein the error conditions include at least one of the error times and the error probabilities of the data items within the preset time window;
[0056] clustering the plurality of data items based on the error conditions to obtain K error patterns of the plurality of data items; wherein K is a positive integer, and different data items in the same error pattern have similar error conditions.
[0057] In order to analyze the error correlation of the data items in the time dimension, a time window T, i.e., a preset time window, can be defined, which represents a fixed time period. Then, the data set D' can be traversed to record the error conditions of each data item within the time window T. For each data item i, the error times within the time window T can be recorded as .
[0058] Specifically, in order to analyze the error correlation of the data items in the time dimension in depth, a key concept, i.e., the time window T, can be defined. The time window T represents a fixed and continuous time period, and its length can be set according to the actual data characteristics and analysis requirements, such as 1 hour, 1 day or 1 week, etc.
[0059] The data set D' can be traversed to record the error details of each data item within the time window T, which can include the number of errors of the data item within the time window T. The implementation steps are as follows:
[0060] Initialize the error record, and create an error counter for each data item i in the data set D' , and set its initial value to 0. This counter will be used to record the number of errors of data item i within the time window T.
[0061] The data set D' can be traversed in time order for each data item in the data set D'. For each data item, it needs to be checked whether it is within the time window T and whether it is error.
[0062] Determine whether the data item is error, which can be determined by setting certain rules according to the actual characteristics of the data and the analysis requirements. For example, according to the business logic of the integral system, a certain rule is set to determine whether the record is error. For example, the integral change of the integral transaction record, the integral change of the integral system has continuity, and the integral change needs to be consistent with the user's integral balance. If the integral change of a certain integral transaction record is inconsistent with the user's integral balance, it can be determined that the record is error. For example, user ID needs to be encoded according to certain rules, if user ID is not consistent with the preset rule of encoding identification, it can be determined that the data item is error.
[0063] Update the error record, if the data item i is error within the time window T, then the value of the corresponding error counter is added by 1. In this way, after traversing the entire data set D', the number of errors of each data item within the time window T can be obtained.
[0064] Output results, the number of errors of each data item as the result output, so as to carry out subsequent analysis and processing.
[0065] Through the above steps, the error situation of each data item within the time window T can be effectively recorded, which provides strong support for subsequent data analysis and processing. And by counting the number of errors of the data item, the error mode of the data item can be identified, which makes the identification of error mode more simple and accurate.
[0066] In some embodiments, the error situation can include error probability, and the determination of the error situation of each data item within the preset time window includes:
[0067] Count the number of errors of each data item within the preset time window;
[0068] Based on the error number, the error probability of each data item in a preset time window is calculated.
[0069] Based on the error number, the error probability of each data item in a preset time window is calculated.
[0070] In some embodiments, the error probability of each data item in a time window T can be calculated, and the error probability of a data item i in a time window T is The error probability can be calculated by the following formula:
[0071]
[0072] Wherein, the denominator represents the total error number of all data items in the time window T. D' is a data set, and j is a data item, Indicates that the data item j belongs to the data set D'.
[0073] In this embodiment, the error mode of the data item is identified by counting the error probability of the data item, so that the identification of the error mode is more simple and accurate.
[0074] In some embodiments, different data items of error cases can be clustered by setting the clustering range of error cases, so as to cluster the data items of similar error cases together, such as clustering the data items with error number of 5-7 times together.
[0075] In some embodiments, in order to achieve the goal of identifying the error mode of the data item in the time window T, a clustering analysis method in machine learning, such as K-means clustering algorithm, can be used to effectively identify the error mode of the data item in the time window T, wherein the K-means clustering algorithm is a widely used clustering method, which is suitable for large data sets and relatively easy to implement. This method clusters according to the error probability and / or error number of the data item in the time window T, aiming to divide the data items into clusters with similar error cases, and each cluster represents a specific error mode. The K-means algorithm is selected because it can handle large-scale data sets and divide the data into a predetermined number of clusters through an iterative optimization process, and the center point of each cluster represents the average error case of the data items in the cluster, thereby facilitating the understanding and analysis of various error modes.
[0076] The specific process is as follows:
[0077] Prepare data, and the data set for clustering is the data set D'. Taking the point system as an example, the data set D' can include transaction records, user behavior data, point change records and other data items in the point system, and the error cases have been counted according to the time window T. Each data item i can have an error number and an error probability.
[0078] Selecting the number of clusters, determining the number of clusters K to be formed, can be determined by prior knowledge, data exploration or a cluster validity index (such as the silhouette coefficient).
[0079] Initializing the cluster centers, K data items can be randomly selected as initial cluster centers. These centers will represent the "prototype" of each cluster.
[0080] Assigning data items to clusters, for each data item i in the data set D', such as a single integral dispensing record data item, the distance of the error case of the data item to each cluster center is calculated, in some embodiments, the Euclidean distance can be used to calculate the distance of the error case of the data item to each cluster center. Data item i is assigned to the cluster represented by the nearest cluster center.
[0081] Updating the cluster centers, for each cluster, the average error case of all its member data items is calculated, such as the average error probability, the average number of errors, or the weighted average of error probability and error number. This average error case is taken as the new cluster center.
[0082] Repeating the iteration until the cluster centers no longer change significantly, or a predetermined number of iterations is reached.
[0083] Outputting the clustering results, the data items in the data set D' will be assigned to K cluster clusters, each cluster cluster represents an error mode.
[0084] By performing these refinement steps, the K-means clustering algorithm can be used to identify the error mode of the data items within the time window T, thereby providing a basis for subsequent block strategy formulation.
[0085] In this embodiment, the error cases of the data items are counted, and the clustering of multiple data items is performed based on the error cases of the data items, thereby identifying the error mode of each data item, which can make the identification of the error mode of the data items more simple and accurate.
[0086] In step 103, in some embodiments, a block strategy of the data set is determined based on the error mode, obtaining at least one data block.
[0087] In some embodiments, the data items in each error mode can be taken as a block to determine the block strategy of the data set, that is, the block strategy of the data set is formulated according to the error mode, all data items under an error mode can be taken as a block, which can assign the data items with error correlation in the time dimension to a data block, making the block more dense.
[0088] In some embodiments, the step 103 specifically includes:
[0089] calculating similarity between different error patterns of the plurality of data items to obtain a similarity matrix;
[0090] determining a chunking strategy of the data set based on the similarity matrix to obtain data chunks.
[0091] That is, the embodiment is further to identify similar error patterns and to assign all data items of similar error patterns to one data chunk, so that all data items of multiple error patterns with close internal relations, i.e., similar error patterns, can be assigned to one data chunk, thereby making the chunking more accurate.
[0092] In some embodiments, to group data items into the same data chunk, similarity between error patterns can be defined. Distance between different error pattern clusters can be calculated using distance metrics (e.g., Euclidean distance) according to error conditions of data items.
[0093] In some embodiments, distance between two different error patterns can be calculated by calculating distance between clustering centers of two error pattern clusters.
[0094] In some embodiments, for two error pattern clusters C k and C l , similarity S kl thereof can be calculated by the following formula:
[0095]
[0096] where P i,T and P j,T are error probabilities of data items i and j within a time window T, respectively, denotes Euclidean distance.
[0097] denotes maximum distance between error probabilities of all data items. After calculating similarity between all error pattern clusters, a similarity matrix S (S is composed of similarity S kl ) can be obtained.
[0098] In some embodiments, based on the similarity matrix, all data items in different error patterns with similarity greater than a preset threshold can be assigned to one data chunk, thereby formulating a chunking strategy of the data set to obtain multiple data chunks.
[0099] In some embodiments, the determining the chunking strategy of the data set based on the similarity matrix to obtain data chunks comprises:
[0100] constructing an undirected weighted graph based on the similarity matrix; wherein in the undirected weighted graph, each node represents a data item, and the edge weight between nodes represents the similarity of error patterns between data items;
[0101] identifying community structures in the undirected weighted graph to obtain at least one community, each of the communities including at least one node, and each of the communities corresponding to a data block.
[0102] In this embodiment, a specific blocking strategy can be formulated based on the similarity matrix S of error patterns. A community detection algorithm in graph theory (such as the Louvain algorithm) can be used to identify the community structure in the similarity matrix, and each community corresponds to a data block. Data items are assigned to the corresponding data block according to the error pattern cluster to which they belong, ensuring that similar error patterns are organized together.
[0103] The specific implementation process is as follows:
[0104] Constructing a similarity graph, a similarity matrix S can be used to construct an undirected weighted graph G, wherein each node represents a data item, and the edge weight between nodes represents the similarity of error patterns between data items. For example: assuming that the similarity matrix between integral transaction records has been calculated, through this matrix, an undirected weighted graph G can be constructed, in which: each node represents an integral transaction record, and the edge weight represents the similarity between two integral transaction records.
[0105] In order to identify the community structure in the undirected weighted graph G, a community detection algorithm in graph theory can be used, such as the Louvain algorithm. The Louvain algorithm is a community detection algorithm based on modularity optimization, which can effectively identify the community structure in the undirected weighted graph G, i.e. a set of nodes with tight internal connections and sparse external connections. In the integral system, the integral transaction records in the undirected weighted graph G can be divided into different communities by the Louvain algorithm, and each community represents a group, and the error patterns of the internal integral transaction records are more similar.
[0106] Through the Louvain algorithm, multiple communities in the undirected weighted graph G can be obtained, and each community corresponds to a data block. In this way, the data items are divided into different data blocks according to the similarity of their error patterns. Assuming that in the undirected weighted graph G of the integral system, the Louvain algorithm identifies two communities: community 1 contains integral transaction records A and B, and community 2 contains integral transaction records C and D. Then: community 1 becomes data block 1, which saves integral transaction records A and B. Community 2 becomes data block 2, which saves integral transaction records C and D.
[0107] According to the error pattern cluster to which each data item belongs, it can be allocated to the corresponding data block. Specifically, for each data item, find the error pattern cluster to which it belongs, and add the data item to the corresponding data block. In the credit system, after analyzing each credit transaction record, according to the data blocks determined in the previous paragraph: credit transaction records A and B are allocated to data block 1, and credit transaction records C and D are allocated to data block 2.
[0108] This embodiment is the core link of the block strategy, and the formulation of the block strategy is based on the similarity matrix of the error patterns of the data items such as the credit transaction records, and the community detection algorithm (such as Louvain algorithm) is used to deeply analyze the internal relationship between the data items such as the credit transaction records. Specifically, a non-directed weighted graph reflecting the similarity of the error patterns of the data items such as the credit transaction records can be constructed, and multiple communities with close internal relationships can be accurately identified and divided by the community detection algorithm, and each community corresponds to a data block. This block strategy based on error pattern similarity ensures that the data items such as the credit transaction records are more reasonably and efficiently laid out at the physical or logical storage level, providing strong support for subsequent data processing, system recovery and error correction coding, and further improving the data storage performance and overall reliability.
[0109] In step 104, error correction coding can be applied to each data block to increase the redundancy of the data, and then the encoded data blocks are distributedly stored.
[0110] In some embodiments, for each data block, the data block can be error correction coded in the same error correction coding manner.
[0111] In some embodiments, the step 104 specifically includes:
[0112] Based on the error characteristics of the data items in each of the data blocks, determine the error correction coding manner of each of the data blocks;
[0113] Based on the error correction coding manner of each of the data blocks, error correction code each of the data blocks;
[0114] Distribute the error correction coded data blocks.
[0115] Wherein, the error characteristics can refer to whether the data items are prone to errors, or the range of their error probability. The error correction coding manner can include error correction coding algorithm and error correction coding parameter.
[0116] The error correction coding algorithm can be selected based on the error characteristics of the data items in each data block. For example, Reed-Solomon code, Bose-Chaudhuri-Hocquenghem code, or Low Density Parity Check (LDPC) code can be selected as the error correction coding algorithm.
[0117] The detailed implementation process is as follows:
[0118] Selecting an error correction coding algorithm:
[0119] Based on the error characteristics of the data items in each data block and the storage requirements, an appropriate error correction coding algorithm can be selected. In some embodiments, Reed-Solomon code can be used as the error correction coding algorithm. Reed-Solomon code is a widely used error correction coding algorithm with strong error correction capability, suitable for handling burst errors and random errors.
[0120] After selecting the error correction coding algorithm, the error correction coding parameters of the data block, such as code length, information bit length, and check bit length, can be determined based on the error characteristics of the data items in the data block. These error correction coding parameters will determine the degree of redundancy and error correction capability of the encoding.
[0121] Based on the error correction coding algorithm and the error correction coding parameters, each data block is subjected to error correction coding. Specifically, the original data in the data block is taken as the information bit, and the corresponding check bit is generated according to the selected error correction coding algorithm and error correction coding parameters, and the information bit and check bit are combined into the error correction coded data block.
[0122] The error correction coded data block is stored in an appropriate storage medium. When storing, the integrity and consistency of the data block need to be ensured for subsequent data processing and recovery.
[0123] In this embodiment, data blocks that are prone to errors at the same time can be organized together to facilitate more efficient processing in subsequent storage and recovery processes. At the same time, by applying error correction coding and fragmentation technology, the redundancy, availability, and fault tolerance of the data can be further improved. This method has wide application prospects in distributed storage environments and can effectively solve the data integrity problem and reduce the recovery requirements of the data.
[0124] The following specific example elaborates the data storage method provided by the embodiment of the present application.
[0125] Figure 2is a specific flowchart of the data storage method provided by an example of the present application, as shown in Figure 2 The data storage method of the integral system includes the following steps:
[0126] Step 201: data collection and preprocessing, obtaining the data set of the integral system, which includes multiple data items of the integral system;
[0127] Step 202: defining a time window and counting the error times of the data items within the time window;
[0128] Step 203: based on the error times of the data items, counting the error probability of the data items;
[0129] Step 204: based on the error times and error probability, identifying the error mode of the data items;
[0130] Step 205: defining the similarity of the error mode and calculating the similarity matrix;
[0131] Step 206: based on the similarity matrix, formulating the block strategy and blocking the data items;
[0132] Step 207: applying error correction coding to the data blocks.
[0133] The core of this embodiment is the in-depth characteristic analysis of the data items. This is not just a simple data statistics, but a detailed identification and classification of the error mode of the data items within the time window. By analyzing the error times and error probability of the data items, and based on the error times and error probability, the error correlation of the data items in the time dimension is analyzed, which can more accurately understand the error behavior of the data items, and provide a strong basis for the subsequent block strategy formulation. This characteristic analysis method can better grasp the essential characteristics of the data, and provide a new perspective for data storage and recovery.
[0134] Another focus of this embodiment is the formulation of the block strategy. Unlike the traditional method of uniformly distributing data items to data blocks, the block strategy can be formulated according to the error mode similarity of the data items. This means that data items with similar error modes will be grouped into the same data block, so that the similarities can be fully utilized to improve the recovery efficiency during data recovery. This formulation of the block strategy not only improves the redundancy of the data, but also optimizes the storage structure of the data.
[0135] The embodiment also focuses on the optimization of error correction coding. Traditional error correction coding methods may not be the best choice because they are not optimized for the error characteristics of data items. The embodiment applies error correction coding that is more suitable for the error characteristics of each data block based on block division. This optimization makes error correction coding more adaptive to the actual error situation of data, thereby improving the reliability and recovery efficiency of data and bringing new possibilities for data storage and recovery.
[0136] In the embodiment, through data item characteristic analysis and block division strategy, data items with similar error patterns are effectively grouped into the same data block. When data errors occur, more targeted recovery operations can be performed using these similarities, thereby significantly improving the recovery efficiency of data. Compared with traditional methods, errors can be located and repaired faster, reducing the time and resources required for data recovery. Moreover, the embodiment applies error correction coding that is more suitable for the error characteristics of each data block based on block division, which further enhances the redundancy and reliability of data. Even if part of the data is erroneous or missing, it can be effectively recovered and reconstructed through error correction coding, ensuring the integrity and accuracy of the data. This enhanced redundancy makes it more robust in the face of various data failures.
[0137] The embodiment has broad commercial prospects.
[0138] In the era of big data, the demand for data storage and processing has surged. It is estimated that large enterprises can process up to 10 TB of data per day, with an error rate of 0.1%. Traditional recovery methods require several hours to several days. The embodiment of the present application can reduce the recovery time to one-tenth of the original by analyzing the error patterns of data items, significantly reducing business interruption time and improving operational efficiency.
[0139] For example, in a large data center, the annual data recovery and maintenance costs can reach millions of dollars. The embodiment of the present application can reduce costs by 30% to 50% through precise analysis and block division strategies, thereby saving costs. At the same time, it can improve data redundancy and reliability, reduce the risk of data loss and damage, and additional investment, with significant economic benefits.
[0140] In addition, the embodiment of the present application has wide application value and potential. It is not only suitable for industries with high reliability requirements such as telecommunications, finance, and healthcare, but also can be extended to emerging fields such as cloud computing and the Internet of Things, providing efficient and reliable data storage and processing services.
[0141] The data storage device provided by the embodiment of the present application will be described below.
[0142] Referring to Figure 3 , the structure of the data storage device provided by the embodiment of the present application is shown in the figure, as Figure 3 shown, the data storage device 300 comprises:
[0143] The acquisition module 301 is configured to acquire a data set to be stored, the data set comprising a plurality of data items;
[0144] The identification module 302 is configured to identify an error pattern of each data item, the error pattern being used to indicate error correlation of the data item with other data items in the data set in a time dimension;
[0145] The block module 303 is configured to determine a block strategy of the data set based on the error pattern, to obtain at least one data block; wherein data items with the same error pattern are divided into the same data block;
[0146] The storage module 304 is configured to store the data set based on the at least one data block.
[0147] In some embodiments, the identification module 302 is specifically configured to:
[0148] determine an error condition of each data item in a preset time window, the error condition comprising at least one of an error number of the data item in the preset time window and an error probability;
[0149] cluster the plurality of data items based on the error condition, to obtain K error patterns of the plurality of data items; wherein K is a positive integer, and different data items in the same error pattern have similar error conditions.
[0150] In some embodiments, the identification module 302 is further configured to:
[0151] count the error number of each data item in the preset time window;
[0152] calculate the error probability of each data item in the preset time window based on the error number.
[0153] In some embodiments, the block module 303 is specifically configured to:
[0154] calculate a similarity between different error patterns of the plurality of data items, to obtain a similarity matrix;
[0155] determine the block strategy of the data set based on the similarity matrix, to obtain a data block.
[0156] In some embodiments, the block module 303 is further configured to:
[0157] construct an undirected weighted graph based on the similarity matrix; wherein each node in the undirected weighted graph represents a data item, and an edge weight between nodes represents a similarity of error patterns between data items;
[0158] identifying a community structure in the undirected weighted graph, to obtain at least one community, each of the communities comprising at least one node, each of the communities corresponding to one data block.
[0159] In some embodiments, the storage module 304 is specifically configured to:
[0160] determining an error correction coding mode of each of the data blocks based on error characteristics of data items in each of the data blocks;
[0161] performing error correction coding on each of the data blocks based on the error correction coding mode of each of the data blocks;
[0162] performing error correction coding on each of the data blocks based on the error correction coding mode of each of the data blocks;
[0163] The data storage apparatus 400 can implement each process implemented in the above-mentioned data storage method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.
[0164] Referring to Figure 4 , a structural schematic diagram of an electronic device provided by an embodiment of the present application is shown. As Figure 4 shown, the electronic device 400 comprises a processor 401, a memory 402, a user interface 403 and a bus interface 404.
[0165] The processor 401 is configured to read a program in the memory 402 and perform the following processes:
[0166] obtaining a data set to be stored, the data set comprising a plurality of data items;
[0167] identifying an error pattern of each of the data items, the error pattern being used to indicate error correlation of the data item with other data items in the data set in a time dimension;
[0168] determining a block strategy of the data set based on the error pattern, to obtain at least one data block; wherein data items with the same error pattern are divided into the same data block;
[0169] performing storage of the data set based on the at least one data block.
[0170] In Figure 4In particular embodiments, the bus architecture can include any number of interconnecting buses and bridges, and the various circuitry representative of the processor(s) 401 and the memory 402 that can be linked through a variety of means, as is well known in the art. The bus architecture can also include various other circuitry that can be designed to coordinate the operation of the various components described above, such as one or more power management and voltage regulation circuits, which are well known and therefore will not be described in further detail. The bus interface 404 provides an interface to the bus architecture. The user interface 403 can also be an interface to other means that can be external or internal to the device, including but not limited to a keypad, a display, a speaker, a microphone, a joystick, etc.
[0171] The processor 401 is responsible for managing the bus architecture and general processing, including the execution of instructions. The memory 402 can be used for the storage of data and as a scratchpad for the processor 401 when executing operations.
[0172] In some embodiments, the processor 401 is further configured to:
[0173] determine an error condition of each of the data items within a preset time window, the error condition comprising at least one of a number of errors and a probability of errors of the data item within the preset time window;
[0174] cluster the plurality of data items based on the error condition to obtain error patterns of K categories of the plurality of data items; wherein K is a positive integer, and different data items in an error pattern of a same category have similar error conditions.
[0175] In some embodiments, the processor 401 is further configured to:
[0176] count a number of errors of each of the data items within a preset time window;
[0177] calculate a probability of errors of each of the data items within a preset time window based on the number of errors.
[0178] In some embodiments, the processor 401 is further configured to:
[0179] calculate a similarity between different error patterns of the plurality of data items to obtain a similarity matrix;
[0180] determine a partitioning strategy of the data set based on the similarity matrix to obtain data blocks.
[0181] In some embodiments, the processor 401 is further configured to:
[0182] construct an undirected weighted graph based on the similarity matrix; wherein each node in the undirected weighted graph represents a data item, and an edge weight between nodes represents a similarity of error patterns between data items;
[0183] Identifying a community structure in the undirected weighted graph, obtaining at least one community, each of the communities comprising at least one node, and each of the communities corresponding to one data block.
[0184] In some embodiments, the processor 401 is further configured to:
[0185] Determining an error correction coding mode of each of the data blocks based on error characteristics of data items in each of the data blocks;
[0186] Performing error correction coding on each of the data blocks based on the error correction coding mode of each of the data blocks;
[0187] Distributively storing the error correction coded data blocks.
[0188] Preferably, the embodiments of the present application further provide an electronic device 400, which comprises a processor 401, a memory 402, and a computer program stored in the memory 402 and capable of running on the processor 401. When the processor 401 executes the computer program, each process of the above-mentioned data storage method embodiments is implemented, and the same technical effects are achieved. To avoid repetition, details are not described herein.
[0189] The embodiments of the present application further provide a readable storage medium, which stores a computer program. When the processor executes the computer program, each process of the above-mentioned data storage method embodiments is implemented, and the same technical effects are achieved. To avoid repetition, details are not described herein. The readable storage medium is, for example, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk.
[0190] The embodiments of the present application further provide a computer program product, which comprises computer instructions. When the processor executes the computer instructions, each process of the above-mentioned data storage method embodiments is implemented, and the same technical effects are achieved. To avoid repetition, details are not described herein.
[0191] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solutions. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0192] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be described here.
[0193] In the embodiments provided by the present application, it should be understood that the disclosed system and method can be implemented in other manners. For example, the described system embodiments are merely schematic. The division of the units is merely a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.
[0194] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.
[0195] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit.
[0196] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, and various program codes that can be stored in the medium.
[0197] The above describes only the specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A data storage method, characterized by, The method comprises: acquiring a data set to be stored, the data set comprising a plurality of data items; identifying an error pattern of each of the data items, the error pattern being used to indicate an error correlation of the data item with other data items in the data set in a time dimension; determining a chunking strategy of the data set based on the error pattern, to obtain at least one data chunk; wherein data items of the same error pattern are divided into the same data chunk; performing storage of the data set based on the at least one data chunk.
2. The method of claim 1, wherein, The identification of the error pattern of each of the data items comprises: determining an error condition of each of the data items within a preset time window, the error condition comprising at least one of an error number of the data item within the preset time window and an error probability; performing clustering of the plurality of data items based on the error condition, to obtain K categories of error patterns of the plurality of data items; wherein K is a positive integer, and different data items in the same category of error pattern have similar error conditions.
3. The method of claim 2, wherein, The determination of the error condition of each of the data items within the preset time window comprises: counting the error number of each of the data items within the preset time window; calculating the error probability of each of the data items within the preset time window based on the error number.
4. The method of claim 1, wherein, The determination of the chunking strategy of the data set based on the error pattern, to obtain a data chunk, comprises: calculating a similarity between different error patterns of the plurality of data items, to obtain a similarity matrix; determining the chunking strategy of the data set based on the similarity matrix, to obtain a data chunk.
5. The method of claim 4, wherein, The determination of the chunking strategy of the data set based on the similarity matrix, to obtain a data chunk, comprises: constructing an undirected weighted graph based on the similarity matrix; wherein each node in the undirected weighted graph represents a data item, and an edge weight between nodes represents a similarity of error patterns between data items; identifying a community structure in the undirected weighted graph, to obtain at least one community, each of the communities comprising at least one node, and each community corresponding to a data chunk.
6. The method of claim 1, wherein, The storage of the data set based on the at least one data chunk comprises: determining an error correction coding mode of each of the data chunks based on error characteristics of data items in each of the data chunks; performing error correction coding on each of the data chunks based on the error correction coding mode of each of the data chunks; performing distributed storage of the data chunks after error correction coding.
7. A data storage device, characterized by The apparatus comprises: an acquisition module, configured to acquire a data set to be stored, the data set comprising a plurality of data items; an identification module, configured to identify an error pattern of each of the data items, the error pattern being used to indicate an error correlation of the data item with other data items in the data set in a time dimension; a chunking module, configured to determine a chunking strategy of the data set based on the error pattern, to obtain at least one data chunk; wherein data items of the same error pattern are divided into the same data chunk; a storage module, configured to perform storage of the data set based on the at least one data chunk.
8. An electronic device, comprising: The electronic device comprises a processor, a memory, a computer program stored on the memory and executable on the processor, the computer program implementing the steps of the data storage method according to any one of claims 1 to 6 when executed by the processor.
9. A readable storage medium, characterized by, The readable storage medium stores a computer program, the computer program implementing the steps of the data storage method according to any one of claims 1 to 6 when executed by a processor.
10. A computer program product, characterised in that, The computer program comprises computer instructions, the computer instructions implementing the steps of the data storage method according to any one of claims 1 to 6 when executed by a processor.