Erasure Code Compatible Read-Write Method and System Based on Bidirectional Data Access Agent
By adopting the erasure code-compatible read and write method based on the bidirectional data access agent in a distributed storage environment, the problems of heterogeneous storage resource perception and data feature understanding are solved, and efficient data writing and load balancing of the storage system are achieved.
Patent Information
- Application Number
- CN202510148387.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The existing technology has multiple challenges in distributed storage environments such as format compatibility, performance optimization and fault-tolerant recovery. It lacks intelligent perception and dynamic scheduling capabilities for heterogeneous storage resources, lacks deep understanding and adaptive learning capabilities for data characteristics, and lacks dynamic perception capabilities for data characteristics and storage environments.
The erasure code-compatible reading and writing method based on the two-way data access agent is adopted, heterogeneous resource portrait is constructed through multi-layer perception institutions, the optimal task allocation strategy is calculated using the multi-agent reinforcement learning model, the read proxy module and the write proxy module are deployed, the data format conversion model is constructed using the space-time graph convolution network, and the model parameters are distributed based on the federated learning framework, and a multi-level index structure is established through a hierarchical bundle filter, and the query performance is optimized in combination with the table hopping mechanism.
It realizes intelligent perception and dynamic scheduling of heterogeneous storage resources, improves deep understanding of data characteristics and adaptive learning ability, enhances dynamic perception of data characteristics and storage environment, improves the reliability and efficiency of the data writing process, and reduces the risk of load imbalance in the storage system.
Smart Images

Figure CN119620957B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed storage, and particularly to an erasure code compatible reading and writing method and system based on a bidirectional data access proxy. Background Art
[0002] With the advent of the big data era, data access and management in heterogeneous storage systems have become increasingly complex. Modern storage systems need to simultaneously process multiple data formats and storage architectures and ensure data reliability and consistency.
[0003] In a distributed storage environment, read and write operations of data face multiple challenges such as format compatibility, performance optimization, and fault tolerance and recovery. Traditional data access methods usually adopt a single proxy mode or a fixed coding scheme, and there are problems such as a lack of intelligent perception and dynamic scheduling capabilities for heterogeneous storage resources, a lack of in-depth understanding and adaptive learning capabilities for data characteristics, and a lack of dynamic perception capabilities for data characteristics and storage environments.
[0004] Therefore, there is an urgent need for a solution to solve the problems existing in the prior art. Summary of the Invention
[0005] An embodiment of the present invention provides an erasure code compatible reading and writing method and system based on a bidirectional data access proxy, which can at least solve some problems existing in the prior art.
[0006] In a first aspect of an embodiment of the present invention, an erasure code compatible reading and writing method based on a bidirectional data access proxy is provided, including:
[0007] Collect the computing capabilities and storage states of each storage node in the system, construct a heterogeneous resource portrait by encoding in combination with a multi-layer perceptron. Based on the heterogeneous resource portrait, input the node state sequence into a multi-agent reinforcement learning model, and obtain an optimal task allocation strategy through interactive iteration calculation of a policy network and a value network. According to the optimal task allocation strategy, deploy a read proxy module and a write proxy module in the system, construct a data format conversion model by using a spatio-temporal graph convolutional network, distribute model parameters to the storage nodes based on a federated learning framework, iteratively optimize the data format conversion model through local training and global aggregation to generate a data feature mapping rule library, establish a multi-level index structure for the data feature mapping rule library and erasure code configuration parameters through a hierarchical Bloom filter, and optimize the query performance by combining a skip list mechanism to obtain an initial system configuration table.
[0008] After receiving a data write request, perform multi-head attention analysis on the data to be written through an encoder based on the attention mechanism, extract the temporal and spatial features of the data to be written to generate a high-dimensional data feature vector, add the high-dimensional data feature vector to a capsule neural network with dynamic routing, calculate the optimal sharding strategy and adaptive coding parameters through iterative negotiation between capsules, perform hierarchical segmentation on the data to be written based on the optimal sharding strategy, generate multi-level data shards with redundancy markers and hierarchical verification information in combination with the adaptive coding parameters, calculate the data distribution scheme through the consistent hashing algorithm and the node selection strategy with load awareness, and send the multi-level data shards and the hierarchical verification information to the storage nodes based on the data distribution scheme and update the system initial configuration table;
[0009] After receiving a data read request, retrieve the target data storage location and coding parameters in the updated system initial configuration table, initiate multi-way parallel reading based on the storage location, obtain the multi-level data shards and hierarchical verification information in combination with the pipeline mechanism, add the read data to the data format conversion model, extract data features in combination with the graph convolutional layer and calculate the optimal conversion path according to the data feature mapping rule library, generate a conversion execution plan and perform format conversion at the same time to obtain an initial conversion result, verify the integrity of the initial conversion result through a double-layer spatial attention network to generate an integrity score. If the integrity score is lower than a pre-set integrity threshold, construct a data dependency graph through a graph attention network, locate the defect location in combination with the hierarchical verification information and perform repair, and repeat the reconstruction until the preset maximum number of reconstructions is reached to obtain the data in the target format.
[0010] In an alternative implementation,
[0011] Collect the computing power and storage status of each storage node in the system, construct a heterogeneous resource portrait through encoding in combination with a multi-layer perceptron. Based on the heterogeneous resource portrait, input the node state sequence into a multi-agent reinforcement learning model, and calculate the optimal task allocation strategy through the interaction and iteration of the policy network and the value network. According to the optimal task allocation strategy, deploy a read proxy module and a write proxy module in the system, construct a data format conversion model using a spatio-temporal graph convolutional network, distribute model parameters to the storage nodes based on the federated learning framework, and iteratively optimize the data format conversion model through local training and global aggregation to generate a data feature mapping rule library. Establish a multi-level index structure for the data feature mapping rule library and erasure code configuration parameters through a hierarchical Bloom filter, and optimize the query performance in combination with the skip list mechanism to obtain the system initial configuration table including:
[0012] Deploy distributed monitoring probes on each storage node, collect the computing power and storage status of the storage node through the distributed monitoring probes to obtain raw monitoring data, preprocess the raw monitoring data, eliminate the short-term fluctuations of the processor and memory usage data through the moving window averaging method, convert the storage capacity data into the ratio of used space to total space, and normalize the input / output throughput data based on the historical peak to obtain preprocessed data;
[0013] Use a multi-layer perceptron with a residual connection structure to extract features from the preprocessed data. The first hidden layer of the multi-layer perceptron extracts independent features of each index, the second hidden layer learns the combined features between the indexes through non-linear transformation, and the output layer generates a feature vector to construct a heterogeneous resource portrait;
[0014] Organize the feature vector into a node state sequence in a sliding window manner, and input the node state sequence into a multi-agent reinforcement learning model. Each agent maintains the corresponding node feature sequence, task queue status, and resource reservation situation, and processes the node state sequence through a hierarchical state encoder. The hierarchical state encoder includes a bottom encoder for processing continuous features, a middle encoder for extracting temporal patterns, and a top encoder for integrating information, and outputs a state encoding;
[0015] The policy network of the multi-agent reinforcement learning model includes an action network and an evaluation network. The action network outputs an action probability distribution based on the state encoding, the evaluation network evaluates the state value, constructs a dynamic communication topology between agents based on the task dependency relationship and resource sharing degree between nodes, processes the neighbor information sequence through a gated recurrent unit, and uses hierarchical reinforcement learning for decision decomposition to obtain an optimal task allocation strategy;
[0016] Deploy a read proxy module and a write proxy module according to the optimal task allocation strategy, and use a spatio-temporal graph convolutional network with a two-stream architecture to construct a data format conversion model. The spatial graph convolutional layer of the spatio-temporal graph convolutional network aggregates neighbor features through a message passing mechanism to update the node representation, the temporal convolutional layer uses dilated convolution to expand the receptive field to capture data dependency relationships, and a hierarchical attention mechanism introducing structural attention and temporal attention is used to process features to obtain a data format conversion model;
[0017] Distribute the model parameters of the data format conversion model based on the federated learning framework, adopt a hierarchical training strategy to locally train the basic feature extractor, globally cooperate to optimize the high-level features, and use progressive knowledge distillation to guide the training process with a pre-trained model to generate a data feature mapping rule library;
[0018] Construct a sharded hierarchical Bloom filter, divide the bitmap of the hierarchical Bloom filter into multiple independently maintained shards, establish a multi-level index structure for the data feature mapping rule library and the erasure code configuration parameters through the sharded hierarchical Bloom filter, determine the index level through a random layer generation algorithm according to the dynamic balance skip list mechanism and the multi-level index structure, maintain a data version chain in the index node to record the version evolution relationship, and adopt a hierarchical filtering strategy for query optimization to obtain the initial system configuration table.
[0019] In an alternative embodiment,
[0020] The policy network of the multi-agent reinforcement learning model includes an action network and an evaluation network. The action network outputs an action probability distribution based on the state encoding. The evaluation network evaluates the state value. A dynamic communication topology among agents is constructed based on the task dependency relationship and resource sharing degree among nodes. The neighbor information sequence is processed through a gated recurrent unit, and hierarchical reinforcement learning is used for decision decomposition. The optimal task allocation strategy obtained includes:
[0021] Construct the policy network of the multi-agent reinforcement learning model. The policy network includes an action network and an evaluation network. Receive state encoding data, which includes a computing resource dimension, a storage resource dimension, and a network resource dimension. The computing resource dimension records processor usage information, memory occupancy information, and process quantity information. The storage resource dimension records disk read / write information, storage space usage information, and storage queue length information. The network resource dimension records bandwidth usage information, network transmission delay information, and packet loss information;
[0022] Add the state encoding data to the action network, perform data normalization transformation and dimension reduction through the first layer to obtain standardized features, extract the correlation patterns in the standardized features and construct combined features through the second layer, and generate an action probability distribution according to the combined features;
[0023] Based on the evaluation network, process the state encoding data in parallel, analyze the mutual influence relationship between different dimension indicators through a feature interaction processing unit to obtain interaction features, and map the interaction features through a neural network to obtain the state value;
[0024] Statistically record the task transfer information between nodes during the system operation process, calculate the task dependency relationship between nodes according to the task transfer frequency, obtain the resource sharing degree according to the storage resource sharing state between nodes, and construct a dynamic communication topology among agents according to the task dependency relationship and the resource sharing degree;
[0025] Determine the set of neighbor nodes for each agent based on the dynamic communication topology, process the sequence of neighbor information in the set of neighbor nodes through a gated recurrent unit, calculate the information update weight using the information update control unit based on the change amplitude between the current state and the historical state, calculate the historical information retention weight using the historical information control unit based on the temporal correlation of the state information, and update the node state information according to the information update weight and the historical information retention weight;
[0026] Use hierarchical reinforcement learning for decision decomposition. Extract task feature information in the first-level decision, and obtain the task computing load feature and the task data access feature. The task computing load feature includes processor time-consuming information, memory usage information, and storage access information. The task data access feature includes data storage location information, data transmission scale information, and data access pattern information;
[0027] In the second-level decision, use the task feature information, the node state information, the action probability distribution, and the state value as inputs to determine the target execution node;
[0028] Based on the dynamic communication topology and the target execution node, train the policy network in a distributed manner. Each target execution node independently trains the network parameters based on the local task execution records, exchanges the network parameters among the target execution nodes with communication connections, and performs parameter integration to generate an optimal task allocation policy, where the optimal task allocation policy includes the correspondence between the task type and the target node.
[0029] In an alternative embodiment,
[0030] After receiving a data write request, perform multi-head attention analysis on the data to be written through an encoder based on the attention mechanism, extract the temporal feature and the spatial feature of the data to be written to generate a high-dimensional data feature vector, add the high-dimensional data feature vector to the capsule neural network with dynamic routing, calculate the optimal sharding strategy and the adaptive coding parameters through iterative negotiation between capsules, perform hierarchical segmentation on the data to be written based on the optimal sharding strategy, generate multi-level data shards with redundancy markers and hierarchical check information in combination with the adaptive coding parameters, calculate the data distribution scheme through the consistent hashing algorithm and the node selection strategy with load awareness, and send the multi-level data shards and the hierarchical check information to the storage nodes based on the data distribution scheme and update the system initial configuration table, including:
[0031] Receive a data write request, and perform multi-head attention analysis on the data to be written through an encoder based on the attention mechanism. Among them, the encoder includes multiple parallel attention processing heads. The first group of attention processing heads is responsible for analyzing temporal features, and the second group of attention processing heads is responsible for analyzing spatial features. Each data block generates a query vector, a key vector, and a value vector through a feature extraction layer, and combines them to obtain a data block sequence;
[0032] Adopt a sliding window mechanism to process the data block sequence, extract the temporal features of the data to be written, calculate the similarity between the query vector of the current data block and the key vectors of other data blocks within the sliding window and obtain the temporal attention weights through normalization processing. Adopt a grid partitioning method to process the data block sequence, extract the spatial features of the data to be written, organize adjacent data blocks into a two-dimensional grid structure and calculate the spatial correlation degree between the target data block and the surrounding data blocks to obtain the spatial attention weights;
[0033] Perform recursive feature fusion on the temporal features, and perform weighted combination on the features of adjacent time windows to obtain a temporal feature vector. Perform hierarchical feature fusion on the spatial features, and perform feature aggregation within the local grid and then layer-by-layer fusion to obtain a spatial feature vector. Concatenate the temporal feature vector and the spatial feature vector and map them through a fully connected layer to generate the high-dimensional data feature vector;
[0034] Add the high-dimensional data feature vector to a capsule neural network with dynamic routing. The capsule neural network includes a primary capsule layer, an intermediate capsule layer, and a high-level capsule layer. Detect local feature patterns through the primary capsule layer, reorganize and abstract the features of the primary capsule layer through the intermediate capsule layer, and represent different granularity data sharding strategies through the high-level capsule layer;
[0035] Connect the primary capsule layer, the intermediate capsule layer, and the high-level capsule layer through iterative negotiation between capsules. Calculate the feature similarity between the low-level capsules and the high-level capsules and update the connection weights in each round of the iterative negotiation. Extract the optimal sharding strategy and adaptive coding parameters from the high-level capsule layer;
[0036] Perform hierarchical segmentation on the data to be written based on the optimal sharding strategy, construct a multi-level data sharding including basic blocks, temporal shards, spatial shards, and aggregation shards, calculate the redundancy mark of the multi-level data sharding according to the adaptive coding parameters, generate the multi-level data sharding with the redundancy mark, and generate the hierarchical check information by calculating the parity check code of the basic block, the erasure code corresponding to different shards, and the checksum of the aggregation shard;
[0037] Construct a consistent hashing ring of virtual nodes through the consistent hashing algorithm and evenly distribute multiple virtual nodes for each storage node on the consistent hashing ring. Adopt a node selection strategy with load awareness to collect the load status of the storage nodes and calculate the load scores of the storage nodes. Calculate a data distribution scheme based on the load scores, and send the multi-level data shards and the hierarchical verification information to the storage nodes based on the data distribution scheme and update the system initial configuration table.
[0038] In an alternative embodiment,
[0039] Construct a consistent hashing ring of virtual nodes through the consistent hashing algorithm and evenly distribute multiple virtual nodes for each storage node on the consistent hashing ring. Adopt a node selection strategy with load awareness to collect the load status of the storage nodes and calculate the load scores of the storage nodes. Calculating a data distribution scheme based on the load scores includes:
[0040] Construct a consistent hashing ring through the consistent hashing algorithm. Perform the first hashing operation on the storage node identifier to obtain a reference hash value and determine the reference position of the virtual node. Perform the second hashing operation on the reference hash value to obtain a corrected hash value and determine the final position of the virtual node on the consistent hashing ring. Generate multiple virtual nodes for each storage node that include node attribute information, load monitoring parameters, and routing tags.
[0041] Calculate the node distribution dispersion of a local area on the consistent hashing ring in a sliding window manner, mark the area where the node distribution dispersion exceeds a preset threshold as an area to be optimized, adjust the distribution position of the virtual nodes within the area to be optimized, and optimize the distribution of the virtual nodes with a decreasing adjustment step until the virtual nodes are evenly distributed on the consistent hashing ring.
[0042] Adopt a node selection strategy with load awareness. Obtain the system resource usage status and business load data of the storage nodes through a load collection agent, store the system resource usage status and the business load data hierarchically according to a time span, perform data smoothing and anomaly filtering processing on the stored load data, and analyze the processed load data to obtain the load status of the storage nodes.
[0043] Calculate the system resource load score and the business load score of the storage node based on the load status of the storage node, determine the weight coefficients of the system resource load score and the business load score in combination with the system operation status, and combine the weighted system resource load score and the business load score to obtain the load score of the storage node.
[0044] Calculate the hash value of the data shard and determine the target position on the consistent hash ring. Select multiple virtual nodes within a fixed distance of the target position as candidate nodes. Sort the storage nodes corresponding to the candidate nodes based on the load score and select the storage node with the lowest score as the target storage node. Check the load status of the target storage node. If the load status exceeds the pre-set load limit, select the storage node with the second lowest score as the target storage node;
[0045] Calculate the difference in load scores between the storage nodes. Screen the storage node pairs whose load score differences exceed the preset threshold. Mark the storage node with the higher load score in the screened storage node pairs as the source node, and the storage node with the lower load score as the target node. Calculate the number of data shards that need to be migrated in the source node, and generate a data distribution plan including the source node, the target node, and the number of data shards.
[0046] In an alternative implementation,
[0047] After receiving a data reading request, retrieve the target data storage location and encoding parameters in the updated system initial configuration table. Based on the storage location, initiate multi-way parallel reading, and combine the pipeline mechanism to obtain the multi-level data shards and hierarchical verification information. Add the read data to the data format conversion model, combine the graph convolutional layer to extract data features and calculate the optimal conversion path according to the data feature mapping rule library. At the same time, generate a conversion execution plan and perform format conversion to obtain an initial conversion result. Verify the integrity of the initial conversion result through a double-layer spatial attention network to generate an integrity score. If the integrity score is lower than the pre-set integrity threshold, construct a data dependency graph through a graph attention network, combine the hierarchical verification information to locate the defect position and repair it, and repeat the reconstruction until the preset maximum number of reconstructions is reached to obtain the target format data, including:
[0048] Receive a data reading request, retrieve the storage location and encoding parameters of the target data in the updated system initial configuration table. Based on the storage location, initiate multi-way parallel reading, and combine the pipeline mechanism to obtain multi-level data shards and hierarchical verification information. The pipeline mechanism divides data processing into four processing stages: reading, decoding, verification, and assembly. Each processing stage maintains an independent task queue, and data is transferred between the processing stages through the producer-consumer mode;
[0049] Add the read data to the data format conversion model, construct a data structure diagram, set data fields as graph nodes, set the association relationships between fields as graph edges, assign weight values to the graph edges according to the business relevance between fields, extract data features by combining graph convolutional layers, extract local structural features in the first - layer graph convolutional processing, aggregate the information of directly connected nodes for each graph node, fuse the information of adjacent nodes in the second - layer graph convolutional processing, combine the features of the first - layer nodes and the features of second - order adjacent nodes, obtain global structural features in the third - layer graph convolutional processing, fuse the information of all nodes in the data structure diagram, and generate a data feature vector;
[0050] Retrieve conversion rules in the data feature mapping rule library according to the data feature vector, calculate the feature similarity between the feature vector and the preset conversion template, and select the conversion template with the highest similarity to calculate the optimal conversion path;
[0051] Generate a conversion execution plan based on the optimal conversion path and perform format conversion. In the structure reorganization stage, adjust the hierarchical structure of the source data format to the hierarchical structure of the target format. In the field mapping stage, establish the correspondence between source - format fields and target - format fields. In the data conversion stage, convert the field values to the data types required by the target format to obtain an initial conversion result;
[0052] Verify the integrity of the initial conversion result through a double - layer spatial attention network. In the local attention layer, check the existence, value range, and type matching of data fields, and calculate the field - level integrity score. In the global attention layer, verify the primary - foreign key relationship, uniqueness constraint, and business rules between data records, and calculate the record - level integrity score. Combine the field - level integrity score and the record - level integrity score to generate an integrity rating;
[0053] If the integrity rating is lower than the pre - set integrity threshold, construct a data dependency graph through a graph attention network, set data fields as graph nodes, set the dependency relationships between fields as graph edges, assign weight values to the graph edges based on the business impact degree, locate the defect positions by combining the hierarchical verification information, determine the repair priority based on the weights of the graph edges, sort the abnormal fields in descending order according to the influence range, repair the abnormal fields according to the sorting result, and perform data reconstruction in an iterative manner;
[0054] Recalculate the integrity rating after each round of reconstruction, repeat the reconstruction until the integrity rating reaches the integrity threshold or reaches the preset maximum number of reconstructions to obtain the data in the target format.
[0055] In an alternative embodiment,
[0056] Construct a data dependency graph through a graph attention network, set data fields as graph nodes, set the dependency relationships between fields as graph edges, assign weight values to the graph edges based on the business impact degree, locate defect positions in combination with the hierarchical verification information, determine the repair priority based on the weights of the graph edges, sort the abnormal fields in descending order according to the impact scope and repair the abnormal fields according to the sorting result. The data reconstruction in combination with the iterative method includes:
[0057] Construct a data dependency graph through a graph attention network, set data fields as graph nodes, parse the database constraint information to extract the dependency relationships between fields and construct reference dependency edges, parse triggers and stored procedures to construct calculation dependency edges, parse field update rules to construct temporal dependency edges, analyze data content to calculate the mutual information value of the field value distribution and construct association dependency edges, analyze the synchronous change pattern of field values to construct co-dependency edges, detect the hierarchical inclusion relationship of field values to construct nested dependency edges, and set the reference dependency edges, the calculation dependency edges, the temporal dependency edges, the association dependency edges, the co-dependency edges, and the nested dependency edges as the graph edges of the data dependency graph;
[0058] Assign weight values to the graph edges of the data dependency graph based on the business impact degree. Set the reference weight of the reference dependency edge as the first weight value in the dependency type layer, set the reference weight of the calculation dependency edge as the second weight value, set the reference weight of the temporal dependency edge as the third weight value, adjust the weight values of the graph edges based on the reference data volume ratio, calculation complexity, and time correlation strength in the data feature layer, and increase the weight values of the graph edges of the dependency relationships involved in the core business processes in the business importance layer;
[0059] Construct a three-level verification rule tree based on the hierarchical verification information. The first-level rule checks the non-emptiness, type matching, and length constraints of fields, marks the fields that violate the basic constraints as first-level abnormal nodes, the second-level rule checks the value range, enumeration value set, and field combination constraints of fields, marks the fields that violate the business rules as second-level abnormal nodes, and the third-level rule checks the consistency, timeliness, and accuracy standards of data, marks the fields that violate the quality requirements as third-level abnormal nodes;
[0060] In the data dependency graph, starting from each abnormal node, traverse all reachable nodes downstream along the graph edges, calculate the influence attenuation coefficient between nodes based on the cumulative graph edge weights of the dependency paths, sum the influence attenuation coefficients of the abnormal nodes on all reachable nodes to obtain the influence scope score, and sort the abnormal fields in descending order according to the influence scope score to obtain the repair priority sequence;
[0061] Repair the abnormal fields in sequence according to the above repair priority sequence. For numerical abnormal fields, generate values randomly based on the historical data of the time series. For categorical abnormal fields, infer the category based on the association rules between fields. For text abnormal fields, supplement the complete information based on string similarity matching, and update the repaired field values to the original data;
[0062] Combine with the iterative method for data reconstruction. Recalculate the node features and influence range scores on the data dependency graph, update the abnormal node set and calculate the data quality score. Compare the quality scores of two adjacent rounds. When the score improvement amplitude is greater than the pre-set score improvement threshold or reaches the preset maximum iteration number, output the repair result.
[0063] In the second aspect of the embodiments of the present invention, a Reed-Solomon code compatible read-write system based on a bidirectional data access proxy is provided, including:
[0064] The first unit is used to collect the computing capabilities and storage statuses of each storage node in the system, construct a heterogeneous resource portrait through encoding in combination with a multi-layer perceptron. Based on the heterogeneous resource portrait, input the node state sequence into a multi-agent reinforcement learning model, and obtain the optimal task allocation strategy through the interactive iteration of the policy network and the value network. According to the optimal task allocation strategy, deploy a read proxy module and a write proxy module in the system, construct a data format conversion model using a spatio-temporal graph convolutional network, distribute model parameters to the storage nodes based on the federated learning framework, and iteratively optimize the data format conversion model through local training and global aggregation to generate a data feature mapping rule library. Establish a multi-level index structure for the data feature mapping rule library and the Reed-Solomon code configuration parameters through a hierarchical Bloom filter, and optimize the query performance in combination with the skip list mechanism to obtain the initial configuration table of the system;
[0065] The second unit is used to, after receiving a data write request, perform multi-head attention analysis on the data to be written through an encoder based on the attention mechanism, extract the temporal features and spatial features of the data to be written to generate a high-dimensional data feature vector, add the high-dimensional data feature vector to a capsule neural network with dynamic routing, calculate the optimal sharding strategy and adaptive coding parameters through iterative negotiation between capsules, hierarchically split the data to be written based on the optimal sharding strategy, generate multi-level data shards with redundancy markers and hierarchical check information in combination with the adaptive coding parameters, calculate the data distribution scheme through the consistent hashing algorithm and the node selection strategy with load awareness, and send the multi-level data shards and the hierarchical check information to the storage nodes based on the data distribution scheme and update the initial configuration table of the system;
[0066] The third unit is used to retrieve the target data storage location and encoding parameters in the updated system initial configuration table after receiving a data reading request, initiate multi-channel parallel reading based on the storage location, obtain the multi-level data shards and hierarchical verification information in combination with the pipeline mechanism, add the read data to the data format conversion model, extract data features in combination with the graph convolutional layer and calculate the optimal conversion path according to the data feature mapping rule library, generate a conversion execution plan and perform format conversion at the same time to obtain an initial conversion result, verify the integrity of the initial conversion result through a double-layer spatial attention network, generate an integrity score, if the integrity score is lower than a preset integrity threshold, construct a data dependency graph through a graph attention network, locate the defect location in combination with the hierarchical verification information and perform repair, repeat the reconstruction until the preset maximum number of reconstructions is reached, and obtain the data in the target format.
[0067] In the third aspect of the embodiments of the present invention,
[0068] A kind of electronic device is provided, including:
[0069] A processor;
[0070] A memory for storing instructions executable by the processor;
[0071] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0072] In the fourth aspect of the embodiments of the present invention,
[0073] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0074] In the present invention, heterogeneous resource portraits and data format conversion models are constructed through multi-agent reinforcement learning and a federated learning framework, realizing the dynamic perception of the computing power and storage status of storage nodes, improving the system resource utilization efficiency and data processing adaptability, making the system have stronger scalability and compatibility, using an encoder based on the attention mechanism and a capsule neural network for data analysis and sharding strategy optimization, combining consistent hashing and load awareness to achieve balanced data distribution, significantly improving the reliability and efficiency of the data writing process, while reducing the risk of load imbalance in the storage system, performing data feature extraction and integrity verification based on a graph convolutional network and a double-layer spatial attention network, and constructing a data dependency relationship through a graph attention network for defect repair, effectively ensuring the accuracy of data reading and format conversion, and improving the fault tolerance and data availability of the system. Description of the Drawings
[0075] Figure 1Schematic flowchart of the erasure code compatible read and write method based on bidirectional data access proxy according to an embodiment of the present invention;
[0076] Figure 2 Schematic structural diagram of the erasure code compatible read and write system based on bidirectional data access proxy according to an embodiment of the present invention. Detailed implementation manners
[0077] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only some of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0078] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0079] Figure 1 Schematic flowchart of the erasure code compatible read and write method based on bidirectional data access proxy according to an embodiment of the present invention, as Figure 1 shown, the method includes:
[0080] S1. Collect the computing capabilities and storage states of each storage node in the system, construct a heterogeneous resource profile by combining a multi-layer perceptron, based on the heterogeneous resource profile, input the node state sequence into a multi-agent reinforcement learning model, obtain an optimal task allocation strategy through the interactive iteration calculation of a policy network and a value network, according to the optimal task allocation strategy, deploy a read proxy module and a write proxy module in the system, construct a data format conversion model by using a spatio-temporal graph convolutional network, distribute model parameters to the storage nodes based on a federated learning framework, iteratively optimize the data format conversion model through local training and global aggregation to generate a data feature mapping rule library, establish a multi-level index structure for the data feature mapping rule library and erasure code configuration parameters through a hierarchical Bloom filter, and optimize the query performance by combining a skip list mechanism to obtain an initial configuration table of the system;
[0081] The heterogeneous resource portrait is a technology for modeling and representing the characteristics and relationships of different types of resources (such as computing, storage, and network) for optimizing resource allocation and scheduling. The spatio-temporal graph convolutional network is a neural network model that combines spatial and temporal information for processing graph-structured data with spatio-temporal dependencies. The data feature mapping rule library is a knowledge base containing the mapping rules between data features and target models for guiding feature selection and feature engineering. The hierarchical Bloom filter is a hierarchical and efficient data storage and query method for quickly determining whether an element exists in a set. The skip list mechanism is a data structure that improves data retrieval speed through a multi-level linked list for implementing fast search and sorting operations.
[0082] In an alternative embodiment,
[0083] Collect the computing capabilities and storage statuses of each storage node in the acquisition system, combine with a multi-layer perceptron for encoding to construct a heterogeneous resource portrait. Based on the heterogeneous resource portrait, input the node state sequence into a multi-agent reinforcement learning model, and obtain the optimal task allocation policy through the interactive iteration of the policy network and the value network. According to the optimal task allocation policy, deploy a read proxy module and a write proxy module in the system, use a spatio-temporal graph convolutional network to construct a data format conversion model, distribute the model parameters to the storage nodes based on the federated learning framework, and iteratively optimize the data format conversion model through local training and global aggregation to generate a data feature mapping rule library. Establish a multi-level index structure for the data feature mapping rule library and erasure code configuration parameters through a hierarchical Bloom filter, and optimize the query performance by combining with the skip list mechanism to obtain the initial system configuration table including:
[0084] Deploy distributed monitoring probes on each storage node, collect the computing capabilities and storage statuses of the storage nodes through the distributed monitoring probes to obtain raw monitoring data, preprocess the raw monitoring data, eliminate the short-term fluctuations of the processor and memory usage data through the sliding window averaging method, convert the storage capacity data into the ratio of used space to total space, and normalize the input / output throughput data based on the historical peak to obtain preprocessed data;
[0085] Use a multi-layer perceptron with a residual connection structure to extract features from the preprocessed data. The first hidden layer of the multi-layer perceptron extracts independent features of each index, the second hidden layer learns the combined features between the indexes through non-linear transformation, and the output layer generates a feature vector to construct a heterogeneous resource portrait;
[0086] Organize the feature vectors into a node state sequence in a sliding window manner, and input the node state sequence into a multi-agent reinforcement learning model. Each agent maintains a corresponding node feature sequence, task queue status, and resource reservation situation, and processes the node state sequence through a hierarchical state encoder. The hierarchical state encoder includes a bottom encoder for processing continuous features, a middle encoder for extracting temporal patterns, and a top encoder for integrating information, and outputs a state encoding.
[0087] The policy network of the multi-agent reinforcement learning model includes an action network and an evaluation network. The action network outputs an action probability distribution based on the state encoding, and the evaluation network evaluates the state value. A dynamic communication topology between agents is constructed based on the task dependency relationship and resource sharing degree between nodes, the neighbor information sequence is processed through a gated recurrent unit, and hierarchical reinforcement learning is used for decision decomposition to obtain an optimal task allocation strategy.
[0088] Deploy a read proxy module and a write proxy module according to the optimal task allocation strategy, and construct a data format conversion model using a spatial-temporal graph convolutional network with a two-stream architecture. The spatial graph convolutional layer of the spatial-temporal graph convolutional network aggregates neighbor features through a message passing mechanism to update the node representation, the temporal convolutional layer uses dilated convolution to expand the receptive field to capture data dependency relationships, and a hierarchical attention mechanism introducing structural attention and temporal attention is used to process features to obtain a data format conversion model.
[0089] Distribute the model parameters of the data format conversion model based on the federated learning framework, adopt a hierarchical training strategy to locally train the basic feature extractor and globally co-optimize the high-level features, and use progressive knowledge distillation to guide the training process with a pre-trained model to generate a data feature mapping rule library.
[0090] Construct a sharded hierarchical Bloom filter, divide the bitmaps of the hierarchical Bloom filter into multiple independently maintained shards, establish a multi-level index structure for the data feature mapping rule library and erasure code configuration parameters through the sharded hierarchical Bloom filter, determine the index level according to the dynamic balance skip list mechanism and the multi-level index structure, maintain a data version chain in the index node to record the version evolution relationship, and adopt a hierarchical filtering strategy for query optimization to obtain the initial system configuration table.
[0091] The sliding window averaging method is a method that calculates the average value using a time window of a fixed size, and is used to smooth time series data or reduce noise. The hierarchical training strategy is a strategy that divides model training into multiple levels for gradual optimization, and is used to improve training efficiency and model performance. The progressive knowledge distillation is a technique that gradually transfers the knowledge of a complex model to a simple model, and is used to accelerate model deployment and inference. The index level refers to the distribution and organization method of different levels of indexes in a multi-level index structure, and is used to improve data retrieval efficiency. The hierarchical filtering strategy is a method that filters data in stages according to priorities and conditions, and is used to reduce the processing overhead of invalid data.
[0092] In a heterogeneous storage system, distributed monitoring probes are deployed to collect the status information of each node. The monitoring probes collect data every 10 seconds, including indicators such as CPU usage rate, memory occupancy rate, storage space usage, and network throughput. For CPU and memory usage data, a sliding window of 60 seconds is used for averaging to eliminate instantaneous fluctuations. The storage space data is converted into a usage rate. For example, if the total capacity of a certain node is 1TB and 800GB has been used, the usage rate is 0.8. The network throughput is normalized based on the historical peak of 10Gbps.
[0093] A four-layer perceptron network is used to extract features. The first hidden layer contains 256 neurons, which respectively process each original indicator; the second hidden layer has 128 neurons to learn the associations between indicators; the third hidden layer has 64 neurons for feature dimensionality reduction; the output layer represents the node resource profile with a 32-dimensional feature vector. Residual connections are added between each layer to alleviate the problem of gradient disappearance.
[0094] The sequence of feature vectors in the last 30 minutes is sampled by sliding with a 5-minute step and input into the multi-agent reinforcement learning model. Each agent corresponds to a storage node and maintains the status information such as the feature sequence of the last 6 time slices, the current task queue length, and the resource reservation situation of the node. The hierarchical state encoder first processes the continuous features with a one-dimensional convolutional network, then extracts the temporal patterns with a long short-term memory network, and finally integrates various types of information with a fully connected layer.
[0095] The policy network consists of two parts: an action generation network and a state evaluation network. The action network outputs the task allocation probability, and the evaluation network predicts the state value. An agent communication topology is constructed based on the data transmission volume and resource competition degree between nodes, and the information sequence from neighbor nodes is processed through a gated recurrent unit. A two-layer decision structure is adopted: the upper layer determines the task type allocation, and the lower layer refines the specific execution plan.
[0096] Deploy the read-write proxy module to execute the task allocation strategy. The data format conversion model adopts a two-stream graph convolutional network architecture. The spatial branch uses three layers of graph convolution for feature aggregation, and the temporal branch uses dilated convolution to expand the receptive field to 128 time steps. Combine the structural attention mechanism to focus on important node connections, and the temporal attention mechanism to capture long-term dependencies.
[0097] Use the federated learning framework to train the conversion model. Each node retains the original data and only exchanges model parameters. The underlying feature extractor is trained locally, and the high-level feature representation is optimized through model aggregation. Use the pre-trained model to guide the training process and improve the model performance through three-stage progressive knowledge distillation. Finally, generate a feature mapping rule library.
[0098] Construct a four-layer Bloom filter index structure, and each layer is divided into 8 independently maintained bitmap shards. Use a dynamically balanced skip list mechanism to optimize queries. The number of node layers is determined by a random algorithm and is maintained at 4 layers on average. The index node records the data version chain and supports querying historical versions by timestamp. When querying, adopt a top-down hierarchical filtering strategy to gradually narrow the search scope and obtain the system initial configuration table.
[0099] In this embodiment, through distributed monitoring and multi-layer perceptron feature extraction, an accurate portrait of heterogeneous resources is realized, which can effectively improve the resource utilization efficiency and system throughput. Based on the task scheduling strategy of multi-agent reinforcement learning, the load is adaptively balanced and the data locality is optimized, which helps to reduce the task response time and system energy consumption. Use federated learning and multi-level indexing to optimize data access, protect data privacy while improving query performance, and help improve the storage space utilization rate while reducing the query latency.
[0100] In an alternative embodiment,
[0101] The policy network of the multi-agent reinforcement learning model includes an action network and an evaluation network. The action network outputs an action probability distribution based on the state encoding, and the evaluation network evaluates the state value. Based on the task dependencies and resource sharing degrees between nodes, a dynamic communication topology between agents is constructed. The neighbor information sequence is processed through a gated recurrent unit, and hierarchical reinforcement learning is used for decision decomposition. The optimal task allocation strategy obtained includes:
[0102] Construct the policy network of the multi-agent reinforcement learning model. The policy network includes an action network and an evaluation network, and receives state encoding data. The state encoding data includes a computing resource dimension, a storage resource dimension, and a network resource dimension. The computing resource dimension records the processor usage information, memory occupancy information, and process number information. The storage resource dimension records the disk read-write information, storage space usage information, and storage queue length information. The network resource dimension records the bandwidth usage information, network transmission delay information, and packet loss information;
[0103] Add the state encoding data to the action network, perform data normalization transformation through the first layer and reduce the dimension to obtain standardized features, extract the correlation patterns in the standardized features through the second layer and construct combined features, and generate an action probability distribution according to the combined features;
[0104] Based on the evaluation network, process the state encoding data in parallel, analyze the mutual influence relationship between different dimensional indicators through a feature interaction processing unit to obtain interaction features, and map the interaction features through a neural network to obtain a state value;
[0105] Statistically record the task transfer information between nodes during the operation of the system, calculate the task dependence relationship between nodes according to the task transfer frequency, obtain the resource sharing degree according to the storage resource sharing state between nodes, and construct a dynamic communication topology between agents according to the task dependence relationship and the resource sharing degree;
[0106] Based on the dynamic communication topology, determine the neighbor node set of each agent, process the neighbor information sequence in the neighbor node set through a gated recurrent unit, calculate the information update weight by using an information update control unit based on the change amplitude between the current state and the historical state, calculate the historical information retention weight by using a historical information control unit based on the time correlation of the state information, and update the node state information according to the information update weight and the historical information retention weight;
[0107] Adopt hierarchical reinforcement learning for decision decomposition. Extract task feature information in the first-layer decision, and obtain task computing load features and task data access features. The task computing load features include processor time-consuming information, memory usage information, and storage access information. The task data access features include data storage location information, data transmission scale information, and data access pattern information;
[0108] In the second-layer decision, use the task feature information, the node state information, the action probability distribution, and the state value as inputs to determine the target execution node;
[0109] Based on the dynamic communication topology and the target execution node, train the policy network in a distributed manner. Each target execution node independently trains to obtain network parameters based on local task execution records, exchange the network parameters between the target execution nodes with communication connections, and perform parameter integration to generate an optimal task allocation strategy, where the optimal task allocation strategy includes the correspondence between task types and target nodes.
[0110] The task transfer frequency refers to the frequency at which tasks are scheduled and executed in the system, and is used to measure the efficiency of task flow. The resource sharing degree is an indicator that describes the sharing and utilization level of the same resource by multiple tasks or users, and is used to evaluate the resource allocation effect. The dynamic communication topology refers to the communication network structure that changes dynamically over time and is used to adapt to the real-time requirements in a distributed system.
[0111] Obtain status encoding data, which includes three main aspects: the computing resource dimension, the storage resource dimension, and the network resource dimension. The computing resource dimension records information such as the processor usage rate, the memory occupancy percentage, and the number of currently running processes. For example, the processor usage rate is 75%, the memory occupancy rate is 60%, and the number of running processes is 12. The storage resource dimension includes the disk read / write speed per second, the storage space usage, and the number of tasks waiting in the storage queue. For example, the disk read / write speed is 120MB / s, the storage space utilization rate is 45%, and the storage queue length is 8. The network resource dimension records indicators such as the bandwidth utilization rate, the network transmission delay time, and the packet loss rate. For example, the bandwidth utilization rate is 65%, the transmission delay is 25ms, and the packet loss rate is 0.1%.
[0112] After receiving the status encoding data, the action network performs data normalization processing through the first layer. For indicators with different dimensions, the maximum-minimum normalization method is used to transform them into a unified interval, and the original features are reduced in dimension through the principal component analysis method to obtain standardized features. The second layer uses a deep neural network to extract the correlation patterns in the standardized features, constructs combined features through a multi-layer perceptron, and the last layer uses the Softmax function to map the combined features into an action probability distribution.
[0113] The evaluation network processes the status encoding data using a parallel architecture. The feature interaction processing unit analyzes the mutual influence relationship between different dimension indicators through the attention mechanism, such as the correlation degree between the computing resource utilization rate and the storage access delay. The interaction features are mapped through a multi-layer neural network to obtain the state value evaluation result.
[0114] During the operation of the system, the task transfer information between nodes is statistically recorded, including the task submission time, the execution node, the completion time, etc. The task dependence relationship strength between nodes is calculated according to the task transfer frequency within a certain time window. At the same time, the storage resource sharing status between nodes is monitored, including information such as the shared storage capacity and the data transmission frequency, and the resource sharing degree is obtained accordingly. A dynamic communication topology between agents is constructed based on the task dependence relationship and the resource sharing degree.
[0115] For each agent, its set of neighbor nodes is determined according to the communication topology. A gated recurrent unit is used to process the sequence of state information from the neighbor nodes. The information update control unit calculates the change amplitude between the current state and the historical state to obtain the information update weight. The historical information control unit calculates the historical information retention weight based on the temporal correlation analysis of the state information. The node state information is updated according to the weights.
[0116] In the first-layer decision-making of hierarchical reinforcement learning, task feature information is extracted. The task computational load features include information such as the estimated processor time consumption, the expected memory occupancy size, and the storage access frequency. The task data access features include information such as the data file storage location, the data transfer size, and the data access time interval.
[0117] The second-layer decision-making takes the task feature information, the node state information, the action probability distribution, and the state value as inputs, and determines the target node most suitable for executing the task through a policy network. Based on the determined target execution node, the policy network is trained in a distributed manner. Each target execution node independently trains the network parameters based on the local task execution records, and the nodes with communication connections exchange and integrate the parameters regularly, and finally generate an optimal task allocation strategy.
[0118] In this embodiment, by constructing a policy network including an action network and an evaluation network, accurate perception of the system state and effective evaluation of action decisions are realized, the accuracy and efficiency of task allocation are improved, a dynamic communication topology is constructed based on the task dependence relationship and resource sharing degree among nodes, and a gated recurrent unit is used to process the neighbor information sequence, enhancing the cooperation ability among nodes and improving the overall resource utilization rate of the system. Hierarchical reinforcement learning is used for decision decomposition, decomposing the complex task allocation problem into two sub-problems of feature extraction and target node selection, reducing the complexity of the decision space, and improving the training efficiency and convergence speed of the model.
[0119] S2. After receiving a data write request, perform multi-head attention analysis on the data to be written through an encoder based on the attention mechanism, extract the temporal and spatial features of the data to be written to generate a high-dimensional data feature vector, add the high-dimensional data feature vector to the capsule neural network with dynamic routing, calculate the optimal sharding strategy and adaptive coding parameters through iterative negotiation between capsules, hierarchically split the data to be written based on the optimal sharding strategy, generate multi-level data shards with redundancy markers and hierarchical check information in combination with the adaptive coding parameters, calculate the data distribution scheme through the consistent hashing algorithm and the node selection strategy with load awareness, and send the multi-level data shards and the hierarchical check information to the storage nodes based on the data distribution scheme and update the system initial configuration table;
[0120] The capsule neural network with dynamic routing is a neural network model that combines a capsule network and a dynamic routing mechanism. It can achieve more accurate feature transmission and spatial relationship modeling by dynamically adjusting connection weights, and is widely used to process input data with complex structures. The iterative negotiation is a process of enabling multiple entities to reach an agreement or optimize a solution by repeatedly exchanging information and adjusting strategies, and is commonly used for decision-making and collaboration in multi-agent systems. The optimal sharding strategy refers to how to efficiently split data into multiple segments in data distributed storage or computing, aiming to optimize resource utilization and reduce access latency. The adaptive coding parameter refers to automatically adjusting the parameter settings in the coding process according to the input data or system state to improve compression efficiency or data transmission quality. The redundancy marking is a technique for marking or managing data redundancy, aiming to improve the fault tolerance and recovery ability of data, and is usually used in storage and transmission systems. The consistent hashing algorithm is an algorithm that ensures the balance of data distribution on nodes by hashing mapping data and nodes, and minimizes the reallocation of data when nodes are added or removed. The load-aware node selection strategy is a strategy for dynamically selecting task processing nodes according to the node load situation, and is used to optimize system performance and improve task execution efficiency.
[0121] In an alternative embodiment,
[0122] After receiving a data write request, perform multi-head attention analysis on the data to be written through an encoder based on the attention mechanism, extract the temporal features and spatial features of the data to be written to generate a high-dimensional data feature vector, add the high-dimensional data feature vector to the capsule neural network with dynamic routing, calculate the optimal sharding strategy and adaptive coding parameters through iterative negotiation between capsules, hierarchically split the data to be written based on the optimal sharding strategy, generate multi-level data shards with redundancy marking and hierarchical check information in combination with the adaptive coding parameters, calculate the data distribution scheme through the consistent hashing algorithm and the load-aware node selection strategy, and send the multi-level data shards and the hierarchical check information to the storage nodes based on the data distribution scheme and update the system initial configuration table, including:
[0123] Receive a data write request, and perform multi-head attention analysis on the data to be written through an encoder based on the attention mechanism. Among them, the encoder includes multiple parallel attention processing heads. The first group of attention processing heads is responsible for analyzing temporal features, and the second group of attention processing heads is responsible for analyzing spatial features. Each data block generates a query vector, a key vector, and a value vector through a feature extraction layer, and combines them to obtain a data block sequence;
[0124] The sliding window mechanism is adopted to process the data block sequence, extract the temporal features of the data to be written, calculate the similarity between the query vector of the current data block and the key vectors of other data blocks within the sliding window and obtain the temporal attention weights through normalization processing, adopt the grid partitioning method to process the data block sequence, extract the spatial features of the data to be written, organize adjacent data blocks into a two-dimensional grid structure and calculate the spatial correlation degree between the target data block and its surrounding data blocks to obtain the spatial attention weights;
[0125] Perform recursive feature fusion on the temporal features, perform weighted combination on the features of adjacent time windows to obtain the temporal feature vector, perform hierarchical feature fusion on the spatial features, perform feature aggregation within the local grid and then fuse layer by layer to obtain the spatial feature vector, splice the temporal feature vector and the spatial feature vector and generate the high-dimensional data feature vector through mapping by a fully connected layer;
[0126] Add the high-dimensional data feature vector to a capsule neural network with dynamic routing, the capsule neural network includes a primary capsule layer, an intermediate capsule layer and a high-level capsule layer, detect local feature patterns through the primary capsule layer, reorganize and abstract the features of the primary capsule layer through the intermediate capsule layer, and represent different granularity data sharding strategies through the high-level capsule layer;
[0127] Connect the primary capsule layer, the intermediate capsule layer and the high-level capsule layer through iterative negotiation between capsules, calculate the feature similarity between low-level capsules and high-level capsules and update the connection weights in each round of the iterative negotiation, and extract the optimal sharding strategy and adaptive coding parameters from the high-level capsule layer;
[0128] Perform hierarchical segmentation on the data to be written based on the optimal sharding strategy, construct a multi-level data sharding including base blocks, temporal shards, spatial shards and aggregated shards, calculate the redundancy marks of the multi-level data sharding according to the adaptive coding parameters, generate the multi-level data sharding with the redundancy marks, and generate the hierarchical check information by calculating the parity check code of the base block, the erasure codes corresponding to different shards and the checksum of the aggregated shards;
[0129] Construct a consistent hashing ring of virtual nodes through the consistent hashing algorithm and evenly distribute multiple virtual nodes on the consistent hashing ring for each storage node, adopt a node selection strategy with load awareness to collect the load status of the storage node and calculate the load score of the storage node, calculate the data distribution scheme according to the load score, and send the multi-level data sharding and the hierarchical check information to the storage node based on the data distribution scheme and update the system initial configuration table.
[0130] The grid partitioning is a technique that divides the data space or computational area into multiple regular grids, which is usually used to balance the load and improve the data access efficiency. The local feature pattern refers to the representative patterns or features extracted from data or signals within a local range, which is used for efficient data analysis and processing. The parity check code is a coding method used to detect data transmission errors. By adding parity check bits, it can determine whether errors occur in the data. The erasure code is a coding method that can tolerate a certain number of data losses or errors. It recovers the lost information through redundant data and is widely used in storage and network systems. The consistent hashing ring is a consistent hashing structure that manages data and nodes by mapping them to a circular space, which can effectively handle the dynamic addition and deletion of nodes.
[0131] When a data write request is received, it is necessary to preprocess and perform feature analysis on the data to be written. The original data is divided into basic data blocks according to a preset size (such as 4KB), and feature extraction is performed on each data block to obtain its basic feature representation. During the feature extraction process, a feature extraction network generates three vectors for each data block: a query vector (used to calculate the correlation with other data blocks), a key vector (used to be queried by other data blocks), and a value vector (carrying the actual feature information of the data block). These vectors together form the feature representation sequence of the data block.
[0132] When performing multi-head attention analysis, two parallel processing branches of temporal feature analysis and spatial feature analysis are started simultaneously. For temporal feature analysis, a sliding window mechanism is adopted, and the window size can be adjusted according to the data characteristics (such as 10 data blocks). Within each sliding window, the similarity between the query vector of the current data block and the key vectors of other data blocks within the window is calculated. The original similarity score is obtained through dot product operation, and then the Softmax function is used for normalization to obtain the temporal attention weights. This process continues as the window slides until all data blocks are processed.
[0133] At the same time, in the spatial feature analysis branch, the data block sequence is reorganized into a two-dimensional grid structure. The grid size can be dynamically adjusted according to the data scale (such as 4×4 or 8×8). In the grid structure, the set of neighboring data blocks for each data block is determined. By calculating the correlation between the feature vectors of the target data block and its surrounding data blocks, the original score reflecting the spatial correlation degree is obtained, and the spatial attention weights are obtained through normalization.
[0134] The feature fusion stage is divided into two steps: temporal feature fusion and spatial feature fusion. In temporal feature fusion, a recursive feature fusion method is adopted to weight and combine the features of adjacent time windows. Specifically, for each pair of adjacent time windows, the weighted coefficient is calculated according to the temporal attention weight, and the features of the two windows are weighted and summed to obtain the fused feature. This process is carried out recursively until the final temporal feature vector is obtained.
[0135] In spatial feature fusion, a hierarchical feature fusion strategy is adopted. Feature aggregation is carried out within the local grid range (such as a 2×2 sub-grid), and the weighted average of the data block features within the sub-grid is calculated. The local aggregated features are used as the input for the next-level fusion, and feature fusion is carried out layer by layer until a vector representing the spatial features of the entire dataset is obtained. The temporal feature vector and the spatial feature vector are concatenated together and processed through a fully connected layer for dimensionality reduction to generate the final high-dimensional data feature vector.
[0136] The high-dimensional data feature vector is input into a capsule neural network with dynamic routing for processing. This network consists of three layers of capsule structures: the primary capsule layer is responsible for detecting the local feature patterns of the data; the middle capsule layer reorganizes and abstracts these local features to form higher-level feature representations; the high-level capsule layer generates data sharding strategies with different granularities based on these feature representations.
[0137] In the dynamic routing process of the capsule network, connections between capsules at each layer are established through an iterative negotiation method. Specifically, in each iteration, the similarity between the output of the lower-layer capsules and the predictions of the higher-layer capsules is calculated, and the connection weights between the capsules are updated according to the similarity. After multiple rounds of iteration (usually 3 - 5 rounds), the optimal data sharding strategy and the corresponding adaptive coding parameters are extracted from the high-level capsule layer.
[0138] Based on the obtained sharding strategy, the data to be written is processed with multi-level sharding. First, the data is divided into basic blocks, then temporal shards are constructed according to temporal correlation, spatial shards are constructed according to spatial correlation, and finally the relevant shards are combined to form aggregated shards. For each level of sharding, the required redundancy is calculated according to the adaptive coding parameters, and the corresponding redundancy markers are added to the shards. At the same time, multi-level verification information is generated: parity check codes are calculated for the basic blocks, erasure codes are generated for shards at different levels, and checksums are calculated for the aggregated shards.
[0139] A consistent hashing ring is constructed, and multiple virtual nodes (usually 100 - 200 times the number of actual nodes) are created for each actual storage node on the hashing ring. These virtual nodes are evenly distributed on the hashing ring through a hash function, and real-time status information of each storage node, including metrics such as CPU usage, memory occupancy, and disk I / O, is collected through a load sensing mechanism, and the load score of each node is comprehensively calculated.
[0140] According to the node load scores and the distribution of the consistent hashing ring, the optimal data distribution scheme is calculated. According to this scheme, multi-level data shards and corresponding verification information are distributed to the selected storage nodes. After the data distribution is completed, the configuration table is updated to record the location information, redundancy information, and verification information of the data shards for subsequent data access.
[0141] In this embodiment, through the combination of the multi-head attention mechanism and the capsule network, the in-depth mining of the data temporal features and spatial features is realized, the adaptability and accuracy of the data sharding strategy are improved, the data storage is made more efficient and reliable, the multi-level data sharding structure and the adaptive redundancy coding mechanism are adopted, the utilization rate of the storage space is optimized while ensuring the data reliability, the fault tolerance and recovery efficiency of the system are improved, the dynamic load balancing of the storage nodes is realized based on the load-aware consistent hashing algorithm, the scalability and stability of the system are improved, and at the same time, the impact of node failures on the overall performance of the system is reduced.
[0142] In an alternative embodiment,
[0143] Constructing a consistent hashing ring of virtual nodes through the consistent hashing algorithm and evenly distributing multiple virtual nodes for each storage node on the consistent hashing ring, collecting the load status of the storage nodes and calculating the load scores of the storage nodes by using a load-aware node selection strategy, and calculating the data distribution scheme according to the load scores includes:
[0144] Constructing a consistent hashing ring through the consistent hashing algorithm, performing the first hashing operation on the storage node identifier to obtain a reference hash value and determining the reference position of the virtual node, performing the second hashing operation on the reference hash value to obtain a corrected hash value and determining the final position of the virtual node on the consistent hashing ring, and generating multiple virtual nodes for each storage node including node attribute information, load monitoring parameters, and routing tags;
[0145] Calculating the node distribution dispersion of the local area on the consistent hashing ring in a sliding window manner, marking the area where the node distribution dispersion exceeds the preset threshold as the area to be optimized, adjusting the distribution positions of the virtual nodes in the area to be optimized, and optimizing the distribution of the virtual nodes with a decreasing adjustment step until the virtual nodes are evenly distributed on the consistent hashing ring;
[0146] Adopt a load-aware node selection strategy. Obtain the system resource usage status and business load data of the storage node through a load collection agent, hierarchically store the system resource usage status and the business load data according to a time span, perform data smoothing and anomaly filtering on the stored load data, and analyze the processed load data to obtain the load status of the storage node;
[0147] Calculate the system resource load score and business load score of the storage node based on the load status of the storage node, determine the weight coefficients of the system resource load score and the business load score in combination with the system operation status, and combine the weighted system resource load score and the business load score to obtain the load score of the storage node;
[0148] Calculate the hash value of the data shard and determine the target position on the consistent hash ring. Select multiple virtual nodes within a fixed distance of the target position as candidate nodes, sort the storage nodes corresponding to the candidate nodes based on the load score, and select the storage node with the lowest score as the target storage node. Check the load status of the target storage node. If the load status exceeds the pre-set load limit, select the storage node with the second lowest score as the target storage node;
[0149] Calculate the difference in load scores between the storage nodes, screen the storage node pairs with the difference in load scores exceeding the preset threshold, mark the storage node with the higher load score in the selected storage node pairs as the source node, mark the storage node with the lower load score as the target node, calculate the number of data shards that need to be migrated in the source node, and generate a data distribution plan including the source node, the target node, and the number of data shards.
[0150] The reference hash value refers to the hash value used as a standard or reference in hash calculation for data comparison, retrieval, or verification. The node distribution dispersion refers to the distribution degree of nodes in space or topology in the network, which is used to measure the distance difference between nodes and affects network performance and load balancing.
[0151] Process the identification information of the storage node. The identification information of the storage node includes information such as the IP address and port number of the node. After combining this information into a string, perform the first hash operation using the SHA-256 hash algorithm to obtain the reference hash value. For example, if the IP address of a storage node is 192.168.1.100 and the port number is 8080, the combined identification string is "192.168.1.100:8080", and the reference hash value is obtained after SHA-256 hash operation.
[0152] After determining the reference position of the virtual node based on the reference hash value, a second hash operation is required to obtain the corrected hash value. The second hash operation uses a different hash seed and generates the hash values of multiple virtual nodes by appending an increasing serial number after the reference hash value. For example, if 100 virtual nodes are generated for each storage node, the serial numbers from 0 to 99 are appended after the reference hash value and then the second hash operation is performed. This can make the distribution of virtual nodes on the hash ring more uniform.
[0153] To achieve the uniform distribution of virtual nodes, it is necessary to evaluate the distribution of nodes on the hash ring. The sliding window method is adopted, and a window with a fixed size slides on the hash ring to count the number of virtual nodes within the window. The window size can be set to one-tenth of the hash ring space, and the sliding step is one-fourth of the window size. For each window position, calculate the deviation between the number of nodes within the window and the theoretical average value, and mark the area where the deviation exceeds 20% as the area to be optimized.
[0154] When adjusting the distribution position of virtual nodes within the area to be optimized, a decreasing adjustment step size is adopted. The initial adjustment step size can be set to half of the size of the area to be optimized, and the step size is reduced to half of the original after each adjustment. During each adjustment, move the virtual nodes within the area to be optimized to the left and right blank areas until the dispersion degree of the nodes within the area is lower than the preset threshold or the maximum number of adjustments is reached.
[0155] In the node selection strategy with load awareness, the load collection agent collects the system resource usage of storage nodes every 10 seconds, including indicators such as CPU usage rate, memory usage rate, and disk I / O. For business load data, such as the number of requests and response time, it is collected once a minute. The collected data is stored at three levels according to the time span: real-time data, hourly data, and daily data.
[0156] When processing the collected load data, the moving average method is used to smooth the data, and the window size is 5 sampling points. Outliers are identified through the box plot method, and the data exceeding 1.5 times the upper and lower quartile intervals is marked as abnormal and filtered out. The processed load data can more accurately reflect the actual load status of the storage node.
[0157] When calculating the load score of the storage node, the system resource load score is composed of three weighted indicators: CPU usage rate, memory usage rate, and disk I / O. The business load score is composed of the number of requests and the average response time weighted. When the system load is light, the weight of the business load score is larger; when the system load is heavy, the weight of the system resource load score increases accordingly.
[0158] When performing data sharding routing, assume that the key value of a certain data shard is "user_profile_1001", and its position on the hash ring is obtained through hash operation. Select the virtual nodes within a fixed distance in the clockwise direction from this position (such as 5% of the hash ring space) as candidate nodes. If 5 candidate nodes are selected and they respectively correspond to 3 actual storage nodes, then sort them according to the load scores of these 3 storage nodes.
[0159] During the generation process of the data migration plan, when it is found that the difference in load scores between two storage nodes exceeds a preset threshold (such as 30%), mark the node with a higher load as the source node and the node with a lower load as the target node. When calculating the number of data shards that need to be migrated from the source node, with the goal of balancing the loads of the two nodes, half of the difference can be used as the migration ratio.
[0160] In this embodiment, by constructing the virtual node distribution through the dual hash operation and the optimized method of decreasing step size, the uniformity of node distribution on the consistent hash ring is significantly improved, the uneven distribution of data among storage nodes is reduced, and the stability and reliability of the storage system are enhanced. By adopting a multi-level load data acquisition and processing mechanism, combined with the system resource usage status and business load data, an accurate evaluation of the load status of storage nodes is achieved, providing a reliable basis for load balancing decisions, effectively improving the system resource utilization efficiency. Based on the load-aware node selection strategy and the adaptive data migration plan, the dynamic balance of the storage system load is achieved, avoiding the situation of overloading a single node and improving the overall performance and service quality of the system.
[0161] S3. After receiving a data reading request, retrieve the target data storage location and encoding parameters in the updated system initial configuration table, start multi-way parallel reading based on the storage location, obtain the multi-level data shards and hierarchical verification information in combination with the pipeline mechanism, add the read data to the data format conversion model, extract data features in combination with the graph convolutional layer and calculate the optimal conversion path according to the data feature mapping rule library. At the same time, generate a conversion execution plan and execute the format conversion to obtain an initial conversion result. Verify the integrity of the initial conversion result through a double-layer spatial attention network to generate an integrity score. If the integrity score is lower than a pre-set integrity threshold, construct a data dependency graph through a graph attention network, locate the defect position in combination with the hierarchical verification information and repair it, and repeat the reconstruction until the preset maximum number of reconstructions is reached to obtain the data in the target format.
[0162] The multi-channel parallel reading is a technique for simultaneously reading data from multiple data sources or channels in parallel, aiming to improve data processing efficiency and system throughput. The pipeline mechanism is a mechanism for accelerating the calculation process and improving system performance by decomposing tasks into multiple consecutive stages and executing each stage in parallel. The optimal conversion path refers to selecting a path that can maximize efficiency or minimize cost among multiple possible paths, which is usually used in data conversion, network routing, or optimization problems. The data dependency graph is a graph structure that describes the dependency relationships between data elements, used to indicate which data must be processed first and which can be processed in parallel to optimize the calculation process. The hierarchical verification information refers to checking data through a multi-level verification mechanism to ensure the integrity and correctness of data at different levels, which is often used in fault tolerance and data verification.
[0163] In an alternative embodiment,
[0164] After receiving a data reading request, retrieve the target data storage location and encoding parameters in the updated system initial configuration table, start multi-channel parallel reading based on the storage location, obtain the multi-level data shards and hierarchical verification information in combination with the pipeline mechanism, add the read data to the data format conversion model, extract data features in combination with the graph convolutional layer and calculate the optimal conversion path according to the data feature mapping rule library, generate a conversion execution plan and execute format conversion at the same time to obtain an initial conversion result, verify the integrity of the initial conversion result through a double-layer spatial attention network to generate an integrity score. If the integrity score is lower than a pre-set integrity threshold, construct a data dependency graph through a graph attention network, locate the defect location in combination with the hierarchical verification information and repair it, repeat the reconstruction until the preset maximum number of reconstructions is reached, and the target format data obtained includes:
[0165] Receive a data reading request, retrieve the storage location and encoding parameters of the target data in the updated system initial configuration table, start multi-channel parallel reading based on the storage location, obtain multi-level data shards and hierarchical verification information in combination with the pipeline mechanism. The pipeline mechanism divides data processing into four processing stages: reading, decoding, verification, and assembly. Each processing stage maintains an independent task queue, and data is transferred between the processing stages through the producer-consumer mode;
[0166] Add the read data to the data format conversion model, construct a data structure diagram, set the data fields as graph nodes, set the association relationships between fields as graph edges, assign weight values to the graph edges according to the business relevance between fields, extract data features by combining graph convolutional layers, extract local structure features in the first-layer graph convolutional processing, aggregate the information of directly connected nodes for each graph node, fuse the information of adjacent nodes in the second-layer graph convolutional processing, combine the features of the first-layer nodes with the features of second-order adjacent nodes, obtain global structure features in the third-layer graph convolutional processing, fuse the information of all nodes in the data structure diagram, and generate a data feature vector;
[0167] Retrieve conversion rules in the data feature mapping rule library according to the data feature vector, calculate the feature similarity between the feature vector and a preset conversion template, and select the conversion template with the highest similarity to calculate the optimal conversion path;
[0168] Generate a conversion execution plan based on the optimal conversion path and perform format conversion. In the structure reorganization stage, adjust the hierarchical structure of the source data format to the hierarchical structure of the target format. In the field mapping stage, establish the correspondence between source format fields and target format fields. In the data conversion stage, convert the field values to the data types required by the target format to obtain an initial conversion result;
[0169] Verify the integrity of the initial conversion result through a double-layer spatial attention network. In the local attention layer, check the existence, value range, and type matching of data fields, and calculate the field-level integrity score. In the global attention layer, verify the primary-foreign key relationship, uniqueness constraint, and business rules between data records, and calculate the record-level integrity score. Combine the field-level integrity score and the record-level integrity score to generate an integrity rating;
[0170] If the integrity rating is lower than a pre-set integrity threshold, construct a data dependency graph through a graph attention network, set the data fields as graph nodes, set the dependency relationships between fields as graph edges, assign weight values to the graph edges based on the business impact degree, locate the defect positions by combining the hierarchical verification information, determine the repair priorities based on the weights of the graph edges, sort the abnormal fields in descending order according to the influence range, repair the abnormal fields according to the sorting result, and perform data reconstruction in an iterative manner;
[0171] Recalculate the integrity rating after each round of reconstruction, repeat the reconstruction until the integrity rating reaches the integrity threshold or reaches a preset maximum number of reconstructions to obtain the data in the target format.
[0172] The producer - consumer pattern is a design pattern commonly used in multi - thread or multi - process systems. Producers generate data, consumers process data, and they collaborate through a buffer to improve system efficiency. The conversion template is a template used to define data conversion rules or steps, which helps to standardize and automate the conversion process between different data formats. The field mapping phase is a step in the data conversion process. By corresponding source data fields with target data fields, it ensures the accuracy of data during the conversion process. The business impact degree refers to the magnitude of the impact of a certain business operation or decision on the system, users, or the entire business process, and is used to evaluate and optimize business strategies.
[0173] In the data reading phase, the system first receives a data reading request initiated by the user. The request contains the identification information of the target data. The system retrieves the storage location and encoding parameters corresponding to this identification information in the updated initial configuration table. The storage locations may be distributed on different storage nodes, such as different data nodes in a distributed file system. Taking medical data as an example, the basic patient information may be stored on node A, the test data on node B, and the imaging data on node C.
[0174] Based on the retrieved storage location information, the system starts a multi - path parallel reading mechanism. Each storage location is assigned an independent reading thread, and these threads execute the data reading tasks in parallel. The pipeline mechanism is used to divide the data processing into multiple stages: reading raw data, decoding data, data verification, and data assembly. An independent task queue is set for each processing stage, and data is passed between stages through the producer - consumer pattern.
[0175] In the data format conversion phase, the read data is input into the conversion model. First, a data structure graph is constructed, with each data field as a graph node and the association relationship between fields as graph edges. For example, in medical data, there is an association relationship between the patient ID and other information, and the system assigns weight values to the graph edges according to the business relevance of these fields.
[0176] Then, data features are extracted through a multi - layer graph convolutional network. The first - layer graph convolution processing mainly focuses on local structural features, aggregating the information of directly connected nodes for each node. The second - layer graph convolution fuses the information of adjacent nodes, combining the first - order node features with the second - order adjacent node features. The third - layer graph convolution obtains global structural features, fusing the information of all nodes in the entire data structure graph, and finally generating a feature vector.
[0177] The system retrieves conversion rules in the rule library based on the feature vector. By calculating the similarity between the feature vector and the preset conversion template, it selects the most similar conversion template to generate the optimal conversion path. For example, when converting XML format to JSON format, the system will select the most suitable conversion template.
[0178] The conversion execution stage includes three steps: restructuring, field mapping, and data conversion. Restructuring adjusts the hierarchical structure of the source data, field mapping establishes the field correspondence from the source format to the target format, and data conversion completes the specific data type conversion.
[0179] The integrity verification stage uses a two-layer spatial attention network for verification. The local attention layer checks the integrity of data fields, including whether the fields exist, whether the values are within the valid range, and whether the data types match. The global attention layer verifies the relational integrity between data records, such as the primary-foreign key relationship and uniqueness constraints. The verification results of the two layers are combined to generate the final integrity score.
[0180] When the integrity score is lower than the threshold, the system constructs a data dependency graph through a graph attention network to locate the defect positions. Weights are assigned to the dependencies according to the business impact degree to determine the repair priorities. The system will preferentially repair the abnormal fields with a larger impact range and perform data reconstruction in an iterative manner until the integrity score reaches the threshold or the maximum number of reconstructions is reached.
[0181] In this embodiment, through multi-way parallel reading and pipeline mechanism, the data processing efficiency is significantly improved, the time overhead of data reading and conversion is reduced, and the overall performance of the system is improved. The graph convolutional network and attention mechanism are used to extract data features, realizing more accurate data format conversion, reducing information loss during the conversion process, ensuring data quality. The two-layer spatial attention network is introduced for integrity verification, and the graph attention network is combined for defect repair, improving the reliability and accuracy of the data conversion result, and ensuring the integrity and consistency of the converted data.
[0182] In an alternative embodiment,
[0183] Construct a data dependency graph through a graph attention network, set data fields as graph nodes, set the dependencies between fields as graph edges, assign weight values to the graph edges based on the business impact degree, locate the defect positions in combination with the hierarchical verification information, determine the repair priorities based on the weights of the graph edges, sort the abnormal fields in descending order according to the impact range and repair the abnormal fields according to the sorting result. The data reconstruction in combination with the iterative method includes:
[0184] Construct a data dependency graph through a graph attention network, set data fields as graph nodes, parse database constraint information to extract the dependency relationship between fields to construct reference dependency edges, parse triggers and stored procedures to construct computational dependency edges, parse field update rules to construct temporal dependency edges, analyze data content to calculate the mutual information value of the field value distribution to construct association dependency edges, analyze the synchronous change pattern of field values to construct co-dependency edges, detect the hierarchical inclusion relationship of field values to construct nested dependency edges, and set the reference dependency edges, the computational dependency edges, the temporal dependency edges, the association dependency edges, the co-dependency edges, and the nested dependency edges as the graph edges of the data dependency graph;
[0185] Assign weight values to the graph edges of the data dependency graph based on the degree of business impact. Set the reference weight of the reference dependency edge as the first weight value at the dependency type layer, set the reference weight of the computational dependency edge as the second weight value, set the reference weight of the temporal dependency edge as the third weight value, adjust the weight values of the graph edges based on the reference data volume ratio, computational complexity, and time correlation strength at the data feature layer, and increase the weight values of the graph edges of the dependency relationships involved in the core business process at the business importance layer;
[0186] Construct a three-level verification rule tree based on hierarchical verification information. The first-level rule checks the non-emptiness, type matching, and length constraints of fields, marks the fields that violate the basic constraints as first-level abnormal nodes, the second-level rule checks the value range, enumeration value set, and field combination constraints of fields, marks the fields that violate the business rules as second-level abnormal nodes, and the third-level rule checks the consistency, timeliness, and accuracy standards of data, marks the fields that violate the quality requirements as third-level abnormal nodes;
[0187] In the data dependency graph, starting from each abnormal node, traverse all reachable nodes downstream along the graph edges, calculate the influence attenuation coefficient between nodes based on the cumulative graph edge weight of the dependency path, sum the influence attenuation coefficients of the abnormal node on all reachable nodes to obtain the influence range score, and sort the abnormal fields in descending order according to the influence range score to obtain the repair priority sequence;
[0188] Repair the abnormal fields in sequence according to the repair priority sequence. Randomly generate values for numerical abnormal fields based on historical data in the time series, infer categories for categorical abnormal fields based on the association rules between fields, supplement complete information for text abnormal fields based on string similarity matching, and update the repaired field values to the original data;
[0189] Perform data reconstruction in combination with an iterative method, recalculate the node features and influence range scores on the data dependency graph, update the abnormal node set and calculate the data quality score, compare the quality scores of adjacent two rounds, and output the repair result when the score improvement amplitude is greater than the pre-set score improvement threshold or reaches the preset maximum number of iterations.
[0190] The time-series dependency edge refers to an edge in time-series data that represents the dependency relationship between different time points and is used to model and analyze the causal relationship in time-series data. The influence attenuation coefficient is a coefficient used to describe the gradual weakening of influence over time or space and is often used in the model to control the scope and intensity of influence. The influence range score is a quantitative description of the influence range of a certain factor or event and is used to analyze and measure the influence or diffusion degree of different factors on the system.
[0191] Construct a data dependency graph through a graph attention network, use the fields in the data table as the nodes of the graph, and construct different types of dependency edges through multi-dimensional analysis. Specifically, when implementing, first parse the foreign key constraint information of the database to identify the reference relationship between fields. For example, the user ID field in the order table references the primary key ID of the user table, and construct this type of relationship as a reference dependency edge. Then analyze the triggers and stored procedures in the database, extract the field calculation logic therein, such as the total order price is calculated based on the order details, and construct this type of relationship as a calculation dependency edge. Then parse the data update rules to identify the sequential update order of field values, such as the status transition of the order from "pending payment" to "paid" to "shipped", and construct this type of relationship as a time-series dependency edge.
[0192] For the relevance analysis of data content, calculate the mutual information value of the value distributions between fields. Taking the user age and consumption amount as an example, by analyzing the value distributions of these two fields, the association pattern between age groups and consumption levels can be found, and construct an association dependency edge between the field pairs with significant associations. At the same time, observe the synchronous change law of field values. For example, the inventory quantity and sales quantity of goods often show a complementary change trend, and construct a collaborative dependency edge between such field pairs with synchronous changes. In addition, it is also necessary to detect the hierarchical inclusion relationship of field values, such as the geographical hierarchical relationship of provinces, cities, and districts and counties, and construct this type of inclusion relationship as a nested dependency edge.
[0193] After the graph construction is completed, it is necessary to assign reasonable weight values to the graph edges. At the dependency type level, the baseline weight of the reference dependency edge is set to 0.8, the calculation dependency edge is set to 0.6, and the timing dependency edge is set to 0.4. Weight adjustment is performed at the data feature level. For example, for the reference dependency edge, if the proportion of the referenced data volume is relatively large, the weight is appropriately increased; for the calculation dependency edge, the weight is adjusted according to the complexity of the calculation logic; for the timing dependency edge, the weight is adjusted according to the tightness of the time association. At the business importance level, for the dependency relationships involved in core business processes such as order processing and payment settlement, their weight values are increased by 20%.
[0194] Construct a three-level verification rule tree for data quality inspection. The first-level rules mainly check basic constraints, such as checking whether the username is empty and whether the mobile phone number is 11 digits. The second-level rules check business constraints, such as checking whether the age is within a reasonable range and whether the order status belongs to a predefined set of statuses. The third-level rules check high-level quality requirements, such as checking whether the order information is consistent with the payment information and whether the commodity price is updated in a timely manner.
[0195] After locating the abnormal nodes based on the inspection results, perform an impact scope analysis in the dependency graph. Starting from each abnormal node, find all reachable nodes through the graph traversal algorithm. During the traversal process, calculate the impact attenuation based on the edge weights on the path. For example, if an abnormal node is connected to a downstream node through an edge with a weight of 0.8, the impact on the downstream node is the original impact value multiplied by 0.8. Accumulate the impact values of the abnormal nodes on all reachable nodes to obtain the impact scope score of the abnormal node.
[0196] Sort the abnormal fields by priority and repair them according to the impact scope score. For numerical anomalies, such as abnormal commodity prices, reasonable values can be randomly generated based on the historical price fluctuation pattern of the commodity. For categorical anomalies, such as abnormal commodity classifications, the most likely classification can be inferred based on the values of related fields such as commodity names and descriptions through association rules. For text anomalies, such as incomplete commodity description information, complete the information by text similarity matching from the descriptions of similar commodities.
[0197] Adopt an iterative approach to continuously optimize the repair effect. Recalculate the feature values and impact scope of the nodes after each round of iteration, update the set of abnormal nodes, and calculate the data quality score. When the improvement rate of the quality score between two adjacent rounds is greater than 5% or reaches 10 iterations, terminate the iterative process and output the final repair result.
[0198] In this embodiment, by constructing a data dependency graph containing various dependency relationships, the association relationships between data fields are comprehensively characterized, improving the accuracy and integrity of data quality analysis. Based on the graph attention network and the multi-level weight allocation mechanism, the influence range of abnormal data is accurately quantified, realizing the precise positioning and priority sorting of anomaly repair, improving the efficiency of data repair. The iterative optimization strategy is adopted to dynamically adjust the repair plan to ensure continuous improvement of data quality. At the same time, through the constraints of multi-level verification rules, the accuracy and reliability of the repair results are guaranteed.
[0199] Figure 2 FIG. is a schematic structural diagram of an erasure code compatible read-write system based on a bidirectional data access proxy according to an embodiment of the present invention, as Figure 2 shown, the system includes:
[0200] The first unit is used to collect the computing capabilities and storage states of each storage node in the system, construct a heterogeneous resource portrait by encoding in combination with a multi-layer perceptron. Based on the heterogeneous resource portrait, the node state sequence is input into a multi-agent reinforcement learning model, and the optimal task allocation strategy is obtained through the interactive iteration of the policy network and the value network. According to the optimal task allocation strategy, a read proxy module and a write proxy module are deployed in the system. A data format conversion model is constructed by using a spatio-temporal graph convolutional network, and the model parameters are distributed to the storage nodes based on the federated learning framework. The data format conversion model is iteratively optimized through local training and global aggregation to generate a data feature mapping rule library. A multi-level index structure is established for the data feature mapping rule library and the erasure code configuration parameters through a hierarchical Bloom filter, and the query performance is optimized by combining the skip list mechanism to obtain an initial system configuration table;
[0201] The second unit is used to, after receiving a data write request, perform multi-head attention analysis on the data to be written through an encoder based on the attention mechanism, extract the temporal features and spatial features of the data to be written to generate a high-dimensional data feature vector. The high-dimensional data feature vector is added to a capsule neural network with dynamic routing, and the optimal sharding strategy and adaptive coding parameters are calculated through iterative negotiation between capsules. Based on the optimal sharding strategy, the data to be written is hierarchically segmented, and multi-level data shards with redundancy markers and hierarchical check information are generated in combination with the adaptive coding parameters. A data distribution scheme is calculated through a consistent hashing algorithm and a node selection strategy with load awareness. Based on the data distribution scheme, the multi-level data shards and the hierarchical check information are sent to the storage nodes and the initial system configuration table is updated;
[0202] The third unit is configured to retrieve the target data storage location and encoding parameters in the updated system initial configuration table after receiving a data reading request, initiate multi-channel parallel reading based on the storage location, obtain the multi-level data shards and hierarchical verification information in combination with the pipeline mechanism, add the read data to the data format conversion model, extract data features in combination with the graph convolutional layer and calculate the optimal conversion path according to the data feature mapping rule library, generate a conversion execution plan and perform format conversion at the same time to obtain an initial conversion result, verify the integrity of the initial conversion result through a double-layer spatial attention network to generate an integrity score. If the integrity score is lower than a pre-set integrity threshold, construct a data dependency graph through a graph attention network, locate the defect location in combination with the hierarchical verification information and repair it, repeat the reconstruction until the preset maximum number of reconstructions is reached, and obtain the data in the target format.
[0203] In the third aspect of the embodiments of the present invention,
[0204] a kind of electronic device is provided, including:
[0205] a processor;
[0206] a memory for storing instructions executable by the processor;
[0207] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0208] In the fourth aspect of the embodiments of the present invention,
[0209] a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0210] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.
[0211] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An erasure code compatible reading and writing method based on a bidirectional data access agent, characterized in that: include: The computing power and storage status of each storage node in the system are collected, and the heterogeneous resource portrait is constructed by encoding with a multi-layer perceptron. Based on the heterogeneous resource portrait, the node state sequence is input into the multi-agent reinforcement learning model, and the optimal task allocation strategy is obtained through interactive iterative calculation of the policy network and the value network. According to the optimal task allocation strategy, the read agent module and the write agent module are deployed in the system, and the data format conversion model is constructed by using a space-time graph convolutional network. The model parameters are distributed to the storage nodes based on the federated learning framework, and the data format conversion model is optimized through local training and global aggregation iterations to generate a data feature mapping rule base. The data feature mapping rule base and the erasure code configuration parameters are used to establish a multi-level index structure through a hierarchical Bloom filter, and the query performance is optimized in combination with the skip table mechanism to obtain the system initial configuration table; After receiving a data write request, a multi-head attention analysis is performed on the data to be written through an encoder based on an attention mechanism, and the temporal features and spatial features of the data to be written are extracted to generate a high-dimensional data feature vector, and the high-dimensional data feature vector is added to the capsule neural network of dynamic routing, and the optimal sharding strategy and adaptive coding parameters are calculated through iterative negotiation between capsules. The data to be written is layered and segmented based on the optimal sharding strategy, and multi-level data shards and layered verification information with redundancy marks are generated in combination with the adaptive coding parameters. The data distribution plan is calculated through a consistent hashing algorithm and a load-aware node selection strategy, and the multi-level data shards and the layered verification information are sent to the storage node based on the data distribution plan and the system initial configuration table is updated; After receiving a data read request, the target data storage location and encoding parameters are retrieved in the updated system initial configuration table, multi-channel parallel reading is started based on the storage location, the multi-level data segmentation and layered verification information are obtained in combination with the pipeline mechanism, the read data is added to the data format conversion model, the data features are extracted in combination with the graph convolution layer, and the optimal conversion path is calculated according to the data feature mapping rule library, and a conversion execution plan is generated and the format conversion is performed to obtain an initial conversion result, and the integrity of the initial conversion result is verified through a two-layer spatial attention network to generate an integrity score. If the integrity score is lower than a preset integrity threshold, a data dependency graph is constructed through a graph attention network, and the defect location is located and repaired in combination with the layered verification information. The reconstruction is repeated until the preset maximum number of reconstructions is reached to obtain the target format data.
2. The method according to claim 1, characterized in that The computing power and storage status of each storage node in the collection system are combined with a multi-layer perceptron for encoding to construct a heterogeneous resource portrait. Based on the heterogeneous resource portrait, the node state sequence is input into the multi-agent reinforcement learning model, and the optimal task allocation strategy is obtained through interactive iterative calculation of the policy network and the value network. According to the optimal task allocation strategy, the read agent module and the write agent module are deployed in the system, and the space-time graph convolution network is used to build a data format conversion model. The model parameters are distributed to the storage nodes based on the federated learning framework. The data format conversion model is optimized through local training and global aggregation iterations to generate a data feature mapping rule base. The data feature mapping rule base and the erasure code configuration parameters are used to establish a multi-level index structure through a hierarchical Bloom filter. The query performance is optimized in combination with the skip table mechanism. The system initial configuration table includes: Deploy distributed monitoring probes on each storage node, collect computing power and storage status of the storage node through the distributed monitoring probes to obtain raw monitoring data, pre-process the raw monitoring data, eliminate short-term fluctuations in processor and memory usage data through a sliding window averaging method, convert storage capacity data into a ratio of used space to total space, and normalize input and output throughput data based on historical peak values to obtain pre-processed data; A multilayer perceptron with a residual connection structure is used to extract features from the preprocessed data. The first hidden layer of the multilayer perceptron extracts independent features of various indicators. The second hidden layer learns the combined features between indicators through nonlinear transformation. The output layer generates feature vectors to construct heterogeneous resource portraits. Organizing the feature vector into a node state sequence by means of a sliding window, inputting the node state sequence into a multi-agent reinforcement learning model, each agent maintaining a corresponding node feature sequence, task queue status and resource reservation status, processing the node state sequence by means of a hierarchical state encoder, wherein the hierarchical state encoder includes a bottom encoder for processing continuous features, a middle encoder for extracting a temporal pattern and a top encoder for integrating information, and outputting a state code; The strategy network of the multi-agent reinforcement learning model includes an action network and an evaluation network. The action network outputs an action probability distribution based on the state encoding, and the evaluation network evaluates the state value. A dynamic communication topology between agents is constructed based on the task dependency relationship and resource sharing degree between nodes. Neighbor information sequences are processed through gated recurrent units, and hierarchical reinforcement learning is used for decision decomposition to obtain an optimal task allocation strategy. Deploy the reading proxy module and the writing proxy module according to the optimal task allocation strategy, and use the space-time graph convolution network with a dual-stream architecture to build a data format conversion model, wherein the space-time graph convolution layer of the space-time graph convolution network aggregates neighbor features through a message passing mechanism to update the node representation, and the temporal convolution layer uses a hole convolution to expand the receptive field to capture data dependencies, and introduces a hierarchical attention mechanism of structural attention and temporal attention to process features, thereby obtaining a data format conversion model; Distribute the model parameters of the data format conversion model based on the federated learning framework, adopt a hierarchical training strategy to locally train the basic feature extractor, globally collaboratively optimize high-level features, guide the training process with the pre-trained model through progressive knowledge distillation, and generate a data feature mapping rule base; Construct a sharded hierarchical Bloom filter, divide the bitmap of the hierarchical Bloom filter into multiple independently maintained shards, establish a multi-level index structure for the data feature mapping rule base and erasure code configuration parameters through the sharded hierarchical Bloom filter, determine the index level through a random layer number generation algorithm based on a dynamically balanced skip table mechanism and the multi-level index structure, maintain a data version chain in the index node to record the version evolution relationship, adopt a hierarchical filtering strategy to perform query optimization, and obtain the system initial configuration table.
3. The method according to claim 2, characterized in that The strategy network of the multi-agent reinforcement learning model includes an action network and an evaluation network. The action network outputs the action probability distribution based on the state encoding, and the evaluation network evaluates the state value. The dynamic communication topology between agents is constructed based on the task dependency relationship and resource sharing degree between nodes. The neighbor information sequence is processed by the gated recurrent unit, and the decision decomposition is performed by hierarchical reinforcement learning. The optimal task allocation strategy includes: Constructing a policy network of a multi-agent reinforcement learning model, the policy network includes an action network and an evaluation network, receiving state encoding data, the state encoding data includes a computing resource dimension, a storage resource dimension, and a network resource dimension, wherein the computing resource dimension records processor usage information, memory occupancy information, and process quantity information, the storage resource dimension records disk read and write information, storage space usage information, and storage queue length information, and the network resource dimension records bandwidth usage information, network transmission delay information, and data packet loss information; Adding the state encoding data to the action network, performing data normalization transformation and reducing the dimension through the first layer to obtain standardized features, extracting the correlation patterns in the standardized features and constructing combined features through the second layer, and generating action probability distribution according to the combined features; Based on the evaluation network, the state encoding data is processed in parallel, the mutual influence relationship between the indicators of different dimensions is analyzed by the feature interaction processing unit to obtain the interaction feature, and the interaction feature is mapped by the neural network to obtain the state value; Statistically record the task flow information between nodes during the operation of the system, calculate the task dependency between nodes according to the task flow frequency, obtain the resource sharing degree according to the storage resource sharing state between nodes, and construct the dynamic communication topology between intelligent agents according to the task dependency and the resource sharing degree; Determine a neighbor node set of each intelligent agent based on the dynamic communication topology, process a neighbor information sequence in the neighbor node set through a gated recurrent unit, calculate an information update weight based on the change amplitude between the current state and the historical state using an information update control unit, calculate a historical information retention weight based on the time correlation of the state information using a historical information control unit, and update the node state information according to the information update weight and the historical information retention weight; Hierarchical reinforcement learning is used to perform decision decomposition. Task feature information is extracted in the first-level decision to obtain task computing load features and task data access features. The task computing load features include processor time consumption information, memory usage information, and storage access information. The task data access features include data storage location information, data transmission scale information, and data access regularity information. In the second-level decision, the task feature information, the node state information, the action probability distribution and the state value are used as input to determine the target execution node; Based on the dynamic communication topology and the target execution nodes, the policy network is trained in a distributed manner. Each target execution node is independently trained to obtain network parameters based on local task execution records. The network parameters are exchanged and integrated between target execution nodes with communication connections to generate an optimal task allocation strategy, wherein the optimal task allocation strategy includes a correspondence between task types and target nodes.
4. The method according to claim 1, characterized in that: After receiving a data write request, a multi-head attention analysis is performed on the data to be written through an encoder based on an attention mechanism, and the temporal features and spatial features of the data to be written are extracted to generate a high-dimensional data feature vector, and the high-dimensional data feature vector is added to the capsule neural network of dynamic routing, and the optimal sharding strategy and adaptive coding parameters are calculated through iterative negotiation between capsules, and the data to be written is hierarchically segmented based on the optimal sharding strategy, and multi-level data shards and hierarchical verification information with redundancy marks are generated in combination with the adaptive coding parameters, and a data distribution scheme is calculated through a consistent hashing algorithm and a load-aware node selection strategy, and the multi-level data shards and the hierarchical verification information are sent to the storage node based on the data distribution scheme and the system initial configuration table is updated, including: Receive a data write request, and perform multi-head attention analysis on the data to be written through an encoder based on the attention mechanism, wherein the encoder includes multiple parallel attention processing heads, wherein the first group of attention processing heads is responsible for analyzing temporal features, and the second group of attention processing heads is responsible for analyzing spatial features. Each data block generates a query vector, a key vector, and a value vector through a feature extraction layer, and the combination obtains a data block sequence; The data block sequence is processed by a sliding window mechanism, the temporal features of the data to be written are extracted, the similarity between the query vector of the current data block and the key vectors of other data blocks in the window is calculated within the sliding window, and the temporal attention weight is obtained through normalization processing, the data block sequence is processed by a grid partitioning method, the spatial features of the data to be written are extracted, the adjacent data blocks are organized into a two-dimensional grid structure, and the spatial correlation between the target data block and the surrounding data blocks is calculated to obtain the spatial attention weight; Recursively perform feature fusion on the time series features, perform weighted combination of features of adjacent time windows to obtain a time series feature vector, perform hierarchical feature fusion on the spatial features, perform feature aggregation in a local grid and then fuse layer by layer to obtain a spatial feature vector, concatenate the time series feature vector and the spatial feature vector and generate the high-dimensional data feature vector through full connection layer mapping; Adding the high-dimensional data feature vector to a capsule neural network of dynamic routing, the capsule neural network comprises a primary capsule layer, an intermediate capsule layer and an advanced capsule layer, detecting local feature patterns through the primary capsule layer, reorganizing and abstracting the features of the primary capsule layer through the intermediate capsule layer, and representing data sharding strategies of different granularities through the advanced capsule layer; The primary capsule layer, the intermediate capsule layer and the high-level capsule layer are connected through iterative negotiation between capsules, the feature similarity between the low-level capsules and the high-level capsules is calculated and the connection weight is updated in each round of the iterative negotiation, and the optimal fragmentation strategy and adaptive coding parameters are extracted from the high-level capsule layer; Based on the optimal sharding strategy, the data to be written is divided into layers, a multi-level data shard including a basic block, a time-series shard, a space shard and an aggregate shard is constructed, the redundancy mark of the multi-level data shard is calculated according to the adaptive coding parameter, the multi-level data shard with the redundancy mark is generated, and the hierarchical check information is generated by calculating the parity check code of the basic block, the erasure code corresponding to different shards and the checksum of the aggregate shard; A consistent hash ring of virtual nodes is constructed through a consistent hash algorithm and multiple virtual nodes are evenly distributed on the consistent hash ring for each storage node. A load-aware node selection strategy is adopted to collect the load status of the storage node and calculate the load score of the storage node. A data distribution plan is calculated according to the load score. Based on the data distribution plan, the multi-level data shards and the hierarchical verification information are sent to the storage node and the system initial configuration table is updated.
5. The method according to claim 4, characterized in that A consistent hash ring of virtual nodes is constructed by a consistent hash algorithm, and a plurality of virtual nodes are evenly distributed on the consistent hash ring for each storage node. A load-aware node selection strategy is adopted to collect the load status of the storage node and calculate the load score of the storage node. The data distribution scheme is calculated according to the load score, including: A consistent hash ring is constructed by a consistent hash algorithm, a first hash operation is performed on the storage node identifier to obtain a reference hash value and determine the reference position of the virtual node, a second hash operation is performed on the reference hash value to obtain a modified hash value and determine the final position of the virtual node on the consistent hash ring, and a plurality of virtual nodes including node attribute information, load monitoring parameters and routing tags are generated for each storage node; The node distribution discreteness of the local area on the consistent hash ring is calculated by a sliding window method, and the area where the node distribution discreteness exceeds a preset threshold is marked as an area to be optimized, and the distribution position of the virtual node is adjusted in the area to be optimized, and the distribution of the virtual node is optimized by a decreasing adjustment step until the virtual nodes are evenly distributed on the consistent hash ring; Adopting a load-aware node selection strategy, obtaining the system resource usage status and business load data of the storage node through a load acquisition agent, hierarchically storing the system resource usage status and the business load data according to the time span, performing data smoothing and abnormal filtering on the stored load data, and analyzing the processed load data to obtain the load status of the storage node; Calculating a system resource load score and a business load score of the storage node based on the load status of the storage node, determining a weight coefficient of the system resource load score and the business load score in combination with the system operation status, and combining the weighted system resource load score and the business load score to obtain a load score of the storage node; Calculate the hash value of the data shard and determine the target position on the consistent hash ring, select multiple virtual nodes within a fixed distance from the target position as candidate nodes, sort the storage nodes corresponding to the candidate nodes based on the load scores and select the storage node with the lowest score as the target storage node, check the load status of the target storage node, and if the load status exceeds a preset load limit, select the storage node with the second lowest score as the target storage node; Calculate the difference in load scores between the storage nodes, screen the storage node pairs whose load score difference exceeds a preset threshold, mark the storage nodes with high load scores in the screened storage node pairs as source nodes, and mark the storage nodes with low load scores as target nodes, calculate the number of data shards that need to be migrated in the source nodes, and generate a data distribution plan that includes the source nodes, the target nodes, and the number of data shards.
6. The method according to claim 1, characterized in that After receiving the data read request, the target data storage location and encoding parameters are retrieved in the updated system initial configuration table, multi-channel parallel reading is started based on the storage location, the multi-level data sharding and layered verification information are obtained in combination with the pipeline mechanism, the read data is added to the data format conversion model, the data features are extracted in combination with the graph convolution layer, and the optimal conversion path is calculated according to the data feature mapping rule library, and a conversion execution plan is generated and the format conversion is performed to obtain the initial conversion result, and the integrity of the initial conversion result is verified by a two-layer spatial attention network to generate an integrity score. If the integrity score is lower than a preset integrity threshold, a data dependency graph is constructed through a graph attention network, and the defect position is located and repaired in combination with the layered verification information, and reconstruction is repeated until the preset maximum number of reconstructions is reached, and the target format data is obtained, including: Receive a data read request, retrieve the storage location and encoding parameters of the target data in the updated system initial configuration table, start multi-channel parallel reading based on the storage location, and obtain multi-level data sharding and layered verification information in combination with a pipeline mechanism. The pipeline mechanism divides data processing into four processing stages: reading, decoding, verification, and assembly. Each processing stage maintains an independent task queue, and data is transferred between the processing stages through a producer-consumer model; Add the read data to the data format conversion model, build a data structure graph, set the data fields as graph nodes, set the associations between the fields as graph edges, assign weight values to the graph edges according to the business relevance between the fields, extract data features in combination with the graph convolution layer, extract local structural features in the first layer of graph convolution processing, aggregate information of directly connected nodes for each graph node, fuse adjacent node information in the second layer of graph convolution processing, combine the first layer node features with the features of second-order adjacent nodes, obtain global structural features in the third layer of graph convolution processing, fuse the information of all nodes in the data structure graph, and generate a data feature vector; Retrieving conversion rules in a data feature mapping rule library according to the data feature vector, calculating feature similarity between the feature vector and a preset conversion template, and selecting the conversion template with the highest similarity to calculate an optimal conversion path; Generate a conversion execution plan based on the optimal conversion path and perform format conversion, adjust the hierarchical structure of the source data format to the hierarchical structure of the target format in the structural reorganization phase, establish a corresponding relationship between the source format field and the target format field in the field mapping phase, and convert the field value into the data type required by the target format in the data conversion phase to obtain an initial conversion result; Verify the integrity of the initial conversion result through a two-layer spatial attention network, check the existence, value range and type matching of the data field at the local attention layer, calculate the field-level integrity score, verify the primary and foreign key relationships, uniqueness constraints and business rules between data records at the global attention layer, calculate the record-level integrity score, and combine the field-level integrity score and the record-level integrity score to generate an integrity score; If the integrity score is lower than a preset integrity threshold, a data dependency graph is constructed through a graph attention network, data fields are set as graph nodes, dependencies between fields are set as graph edges, weight values are assigned to the graph edges based on the degree of business impact, the defect location is located in combination with the hierarchical verification information, the repair priority is determined based on the weight of the graph edge, the abnormal fields are sorted in descending order according to the impact range, and the abnormal fields are repaired according to the sorting results, and data reconstruction is performed in an iterative manner; The integrity score is recalculated after each round of reconstruction, and the reconstruction is repeated until the integrity score reaches the integrity threshold or reaches a preset maximum number of reconstruction times, thereby obtaining target format data.
7. The method according to claim 6, characterized in that A data dependency graph is constructed through a graph attention network, data fields are set as graph nodes, dependencies between fields are set as graph edges, weight values are assigned to the graph edges based on the degree of business impact, the defect location is located in combination with the hierarchical verification information, the repair priority is determined based on the weight of the graph edge, the abnormal fields are sorted in descending order according to the impact range, and the abnormal fields are repaired according to the sorting results. Data reconstruction in combination with an iterative method includes: Construct a data dependency graph through a graph attention network, set data fields as graph nodes, parse database constraint information to extract dependencies between fields to construct reference dependency edges, parse triggers and stored procedures to construct calculation dependency edges, parse field update rules to construct timing dependency edges, analyze data content to calculate the mutual information value of field value distribution to construct association dependency edges, analyze the synchronous change pattern of field values to construct collaborative dependency edges, detect the hierarchical inclusion relationship of field values to construct nested dependency edges, and set the reference dependency edges, the calculation dependency edges, the timing dependency edges, the association dependency edges, the collaborative dependency edges, and the nested dependency edges as graph edges of the data dependency graph; Based on the business impact, weight values are assigned to the graph edges of the data dependency graph. At the dependency type layer, the base weight of the reference dependency edge is set to a first weight value, the base weight of the calculation dependency edge is set to a second weight value, and the base weight of the timing dependency edge is set to a third weight value. At the data feature layer, the weight values of the graph edges are adjusted based on the reference data volume ratio, calculation complexity, and time correlation strength. At the business importance layer, the weight values of the graph edges of the dependency relationships involved in the core business process are increased. A three-level verification rule tree is constructed based on hierarchical verification information. The first-level rules check the non-emptiness, type matching, and length constraints of the field, and mark the fields that violate the basic constraints as the first-level abnormal nodes. The second-level rules check the value range, enumeration value set, and field combination constraints of the field, and mark the fields that violate the business rules as the second-level abnormal nodes. The third-level rules check the consistency, timeliness, and accuracy standards of the data, and mark the fields that violate the quality requirements as the third-level abnormal nodes. In the data dependency graph, starting from each abnormal node, traverse all reachable nodes downstream along the graph edge, calculate the impact attenuation coefficient between nodes based on the cumulative graph edge weight of the dependency path, sum the impact attenuation coefficients of the abnormal node on all reachable nodes to obtain the impact range score, and sort the abnormal fields in descending order according to the impact range score to obtain a repair priority sequence; The abnormal fields are repaired in sequence according to the repair priority sequence, the values of the numerical abnormal fields are randomly generated based on the historical data of the time series, the categories of the categorical abnormal fields are inferred based on the association rules between the fields, and the complete information of the text abnormal fields is supplemented based on the string similarity matching, and the repaired field values are updated to the original data; The data is reconstructed in an iterative manner, node features and impact range scores are recalculated on the data dependency graph, the abnormal node set is updated and the data quality score is calculated, and the quality scores of two adjacent rounds are compared. When the score improvement is greater than a preset score improvement threshold or reaches a preset maximum number of iterations, the repair result is output.
8. An erasure code compatible read / write system for a bidirectional data access agent, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to collect the computing power and storage status of each storage node in the system, and encode and construct a heterogeneous resource portrait in combination with a multi-layer perceptron. Based on the heterogeneous resource portrait, the node state sequence is input into the multi-agent reinforcement learning model, and the optimal task allocation strategy is obtained through interactive iterative calculation of the policy network and the value network. According to the optimal task allocation strategy, a read agent module and a write agent module are deployed in the system, and a data format conversion model is constructed using a space-time graph convolutional network. The model parameters are distributed to the storage nodes based on a federated learning framework, and the data format conversion model is optimized through local training and global aggregation iterations to generate a data feature mapping rule base. The data feature mapping rule base and the erasure code configuration parameters are used to establish a multi-level index structure through a hierarchical Bloom filter, and the query performance is optimized in combination with a skip table mechanism to obtain a system initial configuration table; The second unit is used for, after receiving a data write request, performing multi-head attention analysis on the data to be written through an encoder based on the attention mechanism, extracting the temporal features and spatial features of the data to be written to generate a high-dimensional data feature vector, adding the high-dimensional data feature vector to the capsule neural network of dynamic routing, calculating the optimal sharding strategy and adaptive coding parameters through iterative negotiation between capsules, performing hierarchical segmentation on the data to be written based on the optimal sharding strategy, generating multi-level data sharding and hierarchical verification information with redundancy marks in combination with the adaptive coding parameters, calculating the data distribution plan through a consistent hashing algorithm and a load-aware node selection strategy, sending the multi-level data sharding and the hierarchical verification information to the storage node based on the data distribution plan, and updating the system initial configuration table; The third unit is used to retrieve the target data storage location and encoding parameters in the updated system initial configuration table after receiving a data read request, start multi-channel parallel reading based on the storage location, obtain the multi-level data segmentation and layered verification information in combination with the pipeline mechanism, add the read data to the data format conversion model, extract data features in combination with the graph convolution layer and calculate the optimal conversion path according to the data feature mapping rule library, generate a conversion execution plan and perform format conversion to obtain an initial conversion result, verify the integrity of the initial conversion result through a two-layer spatial attention network, generate an integrity score, and if the integrity score is lower than a preset integrity threshold, construct a data dependency graph through the graph attention network, locate the defect location in combination with the layered verification information and repair it, repeat the reconstruction until the preset maximum number of reconstructions is reached, and obtain the target format data.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Federal learning method and system based on deep reinforcement learning
CN116486192A
Distributed real-time database intelligent management method and system based on big data
CN119356882A