Unstructured file parallel synchronization method and system
Through intelligent blocking and blockchain technology, the problems of inefficiency and difficulty in ensuring consistency in unstructured file transmission are solved, efficient parallel transmission and file reorganization are achieved, and the stability and reliability of the system are improved.
Patent Information
- Application Number
- CN202510091351.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-21
AI Technical Summary
When processing unstructured files, the prior art faces problems such as high bandwidth usage, complex data integrity verification, and disordered transmission sequence, resulting in low transmission efficiency and reduced system stability.
An unstructured file parallel synchronization method is adopted to intelligently process files through the pre-trained file chunking model, generate multiple data sub-blocks, and generate a double-layer cascade index number for each sub-block. These sub-blocks and indexes are distributedly stored and transmitted through the blockchain network, and have hash calculation and verification using data entropy characteristics, dynamically allocate transmission priorities and resource quotas to achieve efficient parallel transmission and file reorganization.
It improves the transmission efficiency and synchronization accuracy of unstructured files in distributed scenarios, ensures data integrity and consistency, and reduces the failure rate and operational costs of the system.
Smart Images

Figure CN119513057B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to file synchronization technology, and in particular to a method and system for parallel synchronization of unstructured files. Background Art
[0002] As an important form of data, unstructured files are widely present in various information systems, including text, images, audio and video. However, the data size of unstructured files is huge, the content is complex, and the structure is diverse. Traditional transmission and storage methods have many shortcomings in data processing efficiency and consistency assurance. In multi-node distributed scenarios, file transmission often faces problems such as high bandwidth usage, complex data integrity verification, and disordered transmission order. These problems not only affect transmission efficiency, but may also make file reorganization difficult, thereby reducing the stability and reliability of the system.
[0003] In the prior art, the block transmission methods for unstructured files are mostly based on fixed rules or simple features, which are difficult to dynamically adapt to the diversity and complexity of file content. At the same time, the index information after data segmentation often adopts a single-layer design, which cannot take into account the collaborative management of block level and file level, resulting in low efficiency in the transmission and verification process. In addition, in the process of distributed storage and reorganization of files, there is a lack of effective priority dynamic allocation and verification mechanism, making it difficult to fully guarantee the integrity and consistency of files.
[0004] In response to the above problems, there is an urgent need for an intelligent and dynamic unstructured file parallel synchronization method that can combine file content characteristics and network transmission status to achieve efficient block segmentation, priority allocation, integrity verification and file reorganization, ensuring efficient transmission and accurate synchronization of unstructured files in distributed scenarios. Summary of the invention
[0005] The embodiments of the present invention provide a method and system for parallel synchronization of unstructured files, which can solve the problems in the prior art.
[0006] According to a first aspect of the embodiments of the present invention,
[0007] A method for parallel synchronization of unstructured files is provided, comprising:
[0008] Obtain the target unstructured file in the source database, and perform intelligent block processing through a pre-trained file block model, wherein the file block model adaptively determines the optimal block size and number of blocks according to the content characteristics, data distribution characteristics, and file structure characteristics of the target unstructured file, generates multiple data sub-blocks, generates a double-layer cascade index number containing block-level information and file-level information for each data sub-block, calculates a data hash value based on the entropy value characteristics of the data sub-block and the double-layer cascade index number, combines the data sub-block, the double-layer cascade index number, and the data hash value to generate a data storage package, which is distributed and stored in multiple storage nodes in the blockchain network;
[0009] Reading data storage packages from each storage node in the blockchain network, building a data transmission priority model based on the double-layer cascade index number, the data transmission priority model comprehensively evaluates the data importance of the data sub-block, the network transmission status and the load of the target database, dynamically allocates transmission priority and network resource quota, and transmits the data sub-block to the target database through an adaptive parallel transmission channel according to the transmission priority. For the data sub-block that has completed the transmission, an asynchronous verification mechanism is used to calculate a verification hash value based on the entropy value characteristics of the data sub-block, and the verification hash value is compared with the data hash value in the data storage package. When the comparison results are consistent, a data verification identifier is generated for the corresponding data sub-block;
[0010] All data sub-blocks with data verification identifiers and their corresponding double-layer cascade index numbers are obtained, and the data sub-blocks are input into a pre-trained file reorganization model. The file reorganization model constructs a data block association map based on the file level information of the double-layer cascade index number, calculates the combination weights between the data sub-blocks according to the data block association map, hierarchically sorts the data sub-blocks based on the combination weights, and uses the content features of the data sub-blocks to perform boundary feature matching and sequence optimization to generate a reorganization sequence, splices and reorganizes the data sub-blocks according to the reorganization sequence to obtain a synchronization target file, generates blockchain confirmation information including reorganization sequence information, file integrity verification information and synchronization status information, stores the synchronization target file in a target-end database, and submits the blockchain confirmation information to the blockchain network. The blockchain confirmation information is synchronously updated in all storage nodes through the consensus mechanism of the blockchain network to complete the parallel synchronization of the target unstructured file.
[0011] In an optional embodiment,
[0012] Obtain the target unstructured file in the source database, and perform intelligent block processing through a pre-trained file block model. The file block model adaptively determines the optimal block size and number of blocks according to the content characteristics, data distribution characteristics and file structure characteristics of the target unstructured file, and generates multiple data sub-blocks including:
[0013] Obtain a target unstructured file from a source database, perform a bidirectional semantic scan on the target unstructured file, extract text content features, the text content features include semantic features and contextual association features, calculate the semantic association of adjacent text contents based on the semantic features and contextual association features, and take a position where the semantic association is lower than a preset association threshold as a first candidate block point;
[0014] Extract data distribution features of the target unstructured file, perform sliding window scanning according to a preset window size, calculate the local sensitive hash value of the data in each window, identify the data repetition pattern and density distribution features based on the local sensitive hash value, and take the position where the data repetition pattern mutates and the density distribution features change as the second candidate block point;
[0015] Analyze the structural features of the target unstructured file, extract the hierarchical features, format features and metadata features of the file, identify the structural separation position of the file based on the hierarchical features, format features and metadata features, and use the structural separation position as the third candidate block point;
[0016] The first candidate block point, the second candidate block point and the third candidate block point are combined to form a candidate block point set, and an evaluation feature vector is calculated for each candidate block point in the candidate block point set, wherein the evaluation feature vector includes a semantic integrity feature based on semantic association, a data balance feature based on data repetition pattern and density distribution feature, and a structural coherence feature based on structural feature;
[0017] The candidate block point set and its corresponding evaluation feature vector are input into a pre-trained file block model, the file block model calculates the optimization weight of each candidate block point according to the evaluation feature vector, screens and combines the candidate block point set based on the optimization weight, adaptively determines the optimal block size and number of blocks, and generates an initial block scheme;
[0018] The initial block partitioning scheme is input into a dynamic programming model, a cost function with semantic integrity features, data balance features and structural coherence features as optimization objectives is constructed, the cost function is input into the dynamic programming model, the optimal state transfer sequence is iteratively calculated through the dynamic programming model to obtain a final block partitioning scheme, the final block boundary position is determined according to the final block partitioning scheme, the target unstructured file is segmented according to the final block boundary position, and a plurality of data sub-blocks are generated.
[0019] In an optional embodiment,
[0020] Generate a double-layer cascade index number containing block-level information and file-level information for each data sub-block, calculate a data hash value based on the entropy value characteristics of the data sub-block and the double-layer cascade index number, combine the data sub-block, the double-layer cascade index number and the data hash value to generate a data storage package, and distribute and store it to multiple storage nodes in the blockchain network, including:
[0021] Extract the file name, creation time and file size of the file to which each data sub-block belongs to form a metadata information set, serialize the metadata information set to obtain a metadata byte stream, perform hash calculation on the metadata byte stream to obtain a file identifier, and simultaneously obtain the position number of the data sub-block in the original file, convert the position number into a binary code to obtain a block number code, fill the block number code with leading zeros to obtain a standardized block number, and connect the file identifier and the standardized block number through a separator to generate a double-layer cascade index number;
[0022] Divide the data sub-block into multiple data segments according to a preset size, calculate the number of occurrences of each byte value in each data segment to obtain a byte distribution feature, calculate the entropy value of each data segment based on the byte distribution feature, and form an entropy value feature vector with the entropy values of all data segments;
[0023] Convert the double-layer cascade index number into an index byte sequence, perform a weighted sum operation on the entropy value feature vector to obtain a feature fusion value, concatenate the index byte sequence and the feature fusion value to generate a mixed feature sequence, and perform a hash calculation on the mixed feature sequence to obtain a data fingerprint hash value;
[0024] The data sub-blocks, double-layer cascade index numbers, and data fingerprint hash values are encapsulated into a data storage package, a storage location is determined on the hash ring of the blockchain network based on the data fingerprint hash value, a preset number of storage nodes closest to the storage location in a clockwise direction are selected as a target storage node set, and the data storage package is sent to each storage node in the target storage node set for storage.
[0025] In an optional embodiment,
[0026] Reading data storage packages from each storage node in the blockchain network, building a data transmission priority model based on the double-layer cascade index number, the data transmission priority model comprehensively evaluates the data importance of the data sub-block, the network transmission status and the load of the target database, dynamically allocating the transmission priority and the network resource quota, and transmitting the data sub-block to the target database through the adaptive parallel transmission channel according to the transmission priority includes:
[0027] Read a data storage package from a storage node in the blockchain network, build an index directory tree based on the double-layer cascade index number in the data storage package, and extract data sub-blocks according to the index directory tree;
[0028] Perform feature matching on the identifiers in the double-layer cascade index number to obtain the inter-block correlation coefficient, calculate the data age decay value based on the timestamp, obtain the access intensity per unit time according to the access count, construct a data importance evaluation matrix with the inter-block correlation coefficient, age decay value, and access intensity, and obtain the data importance score through matrix operation;
[0029] Record the throughput change sequence, link load status sequence, and delay fluctuation sequence of the network transmission path, predict the future network status change trend through a sliding time window, and calculate the dynamic network score based on the network status change trend;
[0030] Collect resource usage indicators of the target database in real time, build a three-layer fuzzy rule set based on processor usage, memory usage, and input and output queue length, and calculate the comprehensive load score of the target end through rule reasoning;
[0031] The data importance score, dynamic network score, and target end comprehensive load score are normalized to construct a three-dimensional state vector, an action space matrix is generated according to historical transmission records, a transmission value function of each data sub-block is calculated based on the three-dimensional state vector and the action space matrix, an action sequence corresponding to the maximum transmission value is selected using a greedy strategy, the value function is iteratively updated until convergence, and the optimal action sequence after convergence is mapped to a transmission priority;
[0032] The data sub-blocks are prioritized according to the transmission priority, the data sub-blocks of adjacent priorities are combined and allocated to the same transmission channel to form a transmission channel group, and the throughput mean and variance of each transmission channel are calculated in real time. When the ratio of the variance to the mean exceeds a preset transmission threshold, the transmission efficiency score of each channel is calculated based on the information entropy principle, the channels with transmission efficiency scores lower than the efficiency threshold are dynamically reorganized, the reorganized transmission channels are reallocated to the transmission channel group, and the data sub-blocks are transmitted to the target database through the transmission channel group.
[0033] In an optional embodiment,
[0034] Obtain all data sub-blocks with data verification identifiers and their corresponding double-layer cascade index numbers, input the data sub-blocks into a pre-trained file reorganization model, the file reorganization model builds a data block association map based on the file level information of the double-layer cascade index numbers, calculates the combination weights between the data sub-blocks according to the data block association map, hierarchically sorts the data sub-blocks based on the combination weights, and uses the content features of the data sub-blocks to perform boundary feature matching and sequence optimization, and generates a reorganization sequence including:
[0035] Obtain a data sub-block with a data verification identifier and its corresponding double-layer cascade index number, extract file-level information and block-level position information in the double-layer cascade index number by segmented decoding, perform multi-scale sliding window scanning on the data sub-block to obtain a feature sequence, associate and map the feature sequence with the file-level information to construct a feature combination;
[0036] Inputting the feature combination into a pre-trained file reorganization model, using the file reorganization model to extract semantic features of data sub-blocks, calculating the association matrix between data sub-blocks based on a multi-head attention mechanism, and constructing a data block association map in combination with the block-level position information;
[0037] In the data block association graph, spatial attention and channel attention are integrated to extract local structural features of nodes, degree distribution entropy and clustering coefficient of nodes are calculated to obtain structural importance, and the semantic features are adaptively weighted integrated with the structural importance to generate a combined weight;
[0038] Using a gated recurrent neural network to perform time series analysis on the combination weights to generate a time series feature sequence, performing correlation clustering on the data sub-blocks according to the continuity of the time series feature sequence to obtain a hierarchical division result, calculating a local sorting order based on the semantic features and position information of the data sub-blocks within each hierarchy, and constructing an initial reorganization scheme including the hierarchical division result and the local sorting order;
[0039] Extracting boundary region features of adjacent data sub-blocks according to the initial reorganization scheme, calculating semantic similarity of boundary region features using a pre-built deep feature extraction network to obtain boundary matching, and when the boundary matching is lower than an adaptive matching threshold, adjusting the local sorting order based on the semantic relevance of the boundary region features to generate an optimized reorganization scheme;
[0040] The characteristic weights between the data sub-blocks are calculated according to the optimized reorganization scheme, the data sub-blocks are sequentially reorganized based on the characteristic weights to obtain a reorganized sequence, and after verifying the integrity of the reorganized sequence, the verification result is updated to the double-layer cascade index number.
[0041] In an optional embodiment,
[0042] According to the reorganization sequence, the data sub-blocks are spliced and reorganized to obtain a synchronization target file, and blockchain confirmation information including reorganization sequence information, file integrity verification information and synchronization status information is generated. The synchronization target file is stored in the target end database, and the blockchain confirmation information is submitted to the blockchain network. The blockchain confirmation information is synchronously updated in all storage nodes through the consensus mechanism of the blockchain network, and the parallel synchronization of the target unstructured file is completed, including:
[0043] A double-buffer streaming mechanism is used to splice and reorganize data sub-blocks, where the first buffer is used for reading data sub-blocks and verifying boundary feature matching, and the second buffer is used for sequential writing of verified data sub-blocks. When the amount of data in the second buffer reaches a preset threshold, the data sub-blocks are written in batches to the synchronization target file and the write position, data length and data block identifier are recorded to generate a file mapping table.
[0044] A multi-level hash verification tree is constructed based on the file mapping table, wherein the hash value and boundary feature information of the data block are stored in the leaf node, the subtree root hash value and the data block association degree are stored in the intermediate node, the root node stores the overall hash value and file metadata, and the file integrity verification information is generated through the verification relationship between the nodes;
[0045] Convert the multi-level hash verification tree into a compressed coding sequence, and package the compressed coding sequence, the recombined sequence information and the synchronization status information in a segmented coding manner to generate blockchain confirmation information with a hierarchical structure, each layer containing an independent digital signature and verification field;
[0046] Based on the consistent hashing algorithm, the blockchain confirmation information is mapped to the confirmation node ring of the blockchain network, multiple nodes in the clockwise direction of the mapping position are selected as the confirmation node group, and the two-phase commit protocol is used within the confirmation node group to complete the pre-confirmation;
[0047] The pre-confirmed blockchain confirmation information is sharded, and the confirmation information is divided into multiple verification shards according to the hierarchical structure of the multi-level hash verification tree, and the consensus algorithm based on practical Byzantine fault tolerance is executed in parallel in different shards, and the blockchain confirmation information is synchronously updated in all storage nodes through cross-shard verification;
[0048] After receiving the blockchain confirmation information, the storage node calculates the verification value layer by layer according to the multi-level hash verification tree, compares it with the verification field in the blockchain confirmation information, and generates a synchronization request including the inconsistent data block location and verification path;
[0049] A data synchronization channel is established based on the verification path of the synchronization request, and a selective retransmission mechanism is used to transmit inconsistent data blocks. The receiving node verifies the correctness of the data block through the multi-level hash verification tree, and the verified synchronization target file is updated to the target end database to complete the parallel synchronization of the target unstructured file.
[0050] In an optional embodiment,
[0051] Based on the consistent hashing algorithm, the blockchain confirmation information is mapped to the confirmation node ring of the blockchain network, multiple nodes in the clockwise direction of the mapping position are selected as the confirmation node group, and the two-phase commit protocol is used within the confirmation node group to complete the pre-confirmation, including:
[0052] Obtain the computing power value and storage capacity value of each confirmation node in the blockchain network, determine the number of virtual nodes to be allocated based on the computing power value and storage capacity value, use a hash algorithm to calculate the node identification hash value of the virtual node and map it to the hash ring, and build a confirmation node ring;
[0053] Calculate the confirmation information hash value of the blockchain confirmation information based on the consistent hashing algorithm and map it to the confirmation node ring, obtain the network delay value of each confirmation node in the confirmation node ring and the historical confirmation performance value in the node scoring table, and select multiple confirmation nodes as the confirmation node group in a clockwise direction from the mapping position of the confirmation information hash value based on the network delay value and the historical confirmation performance value;
[0054] The confirmation information hash value, timestamp and digital signature of the coordinating node are combined to generate a pre-confirmation request, and the node corresponding to the mapping position in the confirmation node group acts as a coordinating node to send the pre-confirmation request to the remaining confirmation nodes in the confirmation node group, and the confirmation nodes in the confirmation node group verify the legitimacy of the pre-confirmation request and return a preparation response;
[0055] The coordination node counts the number of ready responses of the confirmation node group, generates a submission request when the number of ready responses exceeds a preset ratio, sends the submission request to the confirmation node group, and the confirmation nodes in the confirmation node group perform local confirmation and generate a confirmation log;
[0056] A timeout detection mechanism is used to monitor the online status of the coordination node, and when a coordination node failure is detected, a new coordination node is selected from the confirmation node group, and confirmation authority is allocated to the new coordination node;
[0057] The confirmation operation information in the confirmation log is constructed into a verification tree, the root node value of the verification tree is combined with the confirmation signature of the confirmation node group to generate a pre-confirmation result, and the pre-confirmation result is sent to the remaining confirmation nodes in the blockchain network. After receiving the pre-confirmation result, each remaining confirmation node verifies the number of confirmation signatures and the proof path of the verification tree. After the verification is passed, the pre-confirmation result is recorded as a checkpoint to complete the pre-confirmation of the blockchain confirmation information.
[0058] According to a second aspect of the embodiments of the present invention,
[0059] Provided is a non-structured file parallel synchronization system, comprising:
[0060] The first unit is used to obtain the target unstructured file in the source database, and perform intelligent block processing through a pre-trained file block model, wherein the file block model adaptively determines the optimal block size and the number of blocks according to the content characteristics, data distribution characteristics and file structure characteristics of the target unstructured file, generates multiple data sub-blocks, generates a double-layer cascade index number containing block-level information and file-level information for each data sub-block, calculates a data hash value based on the entropy value characteristics of the data sub-block and the double-layer cascade index number, combines the data sub-block, the double-layer cascade index number and the data hash value to generate a data storage package, and distributes and stores it to multiple storage nodes in the blockchain network;
[0061] The second unit is used to read the data storage package from each storage node in the blockchain network, build a data transmission priority model based on the double-layer cascade index number, and the data transmission priority model comprehensively evaluates the data importance of the data sub-block, the network transmission status and the load of the target database, dynamically allocates the transmission priority and the network resource quota, and transmits the data sub-block to the target database through the adaptive parallel transmission channel according to the transmission priority. For the data sub-block that has completed the transmission, an asynchronous verification mechanism is used to calculate the verification hash value based on the entropy value characteristics of the data sub-block, and the verification hash value is compared with the data hash value in the data storage package. When the comparison results are consistent, a data verification identifier is generated for the corresponding data sub-block;
[0062] The third unit is used to obtain all data sub-blocks with data verification identifiers and their corresponding double-layer cascade index numbers, input the data sub-blocks into a pre-trained file reorganization model, the file reorganization model constructs a data block association map based on the file level information of the double-layer cascade index number, calculates the combination weights between the data sub-blocks according to the data block association map, hierarchically sorts the data sub-blocks based on the combination weights, and uses the content features of the data sub-blocks to perform boundary feature matching and sequence optimization to generate a reorganization sequence, splices and reorganizes the data sub-blocks according to the reorganization sequence to obtain a synchronization target file, generates blockchain confirmation information including reorganization sequence information, file integrity verification information and synchronization status information, stores the synchronization target file in a target end database, and submits the blockchain confirmation information to the blockchain network, synchronously updates the blockchain confirmation information in all storage nodes through the consensus mechanism of the blockchain network, and completes the parallel synchronization of the target unstructured file.
[0063] According to a third aspect of the embodiments of the present invention,
[0064] An electronic device is provided, comprising:
[0065] processor;
[0066] a memory for storing processor-executable instructions;
[0067] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0068] According to a fourth aspect of the embodiments of the present invention,
[0069] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0070] In this embodiment, the intelligent block model is used to dynamically determine the block size and number according to the content characteristics, data distribution characteristics and file structure characteristics of the unstructured file, and optimize the efficiency and accuracy of data block. The double-layer cascade index number contains both block-level and file-level information, supports more efficient transmission and accurate file reorganization, and provides strong support for subsequent processing. The data hash value is calculated in combination with the data entropy value characteristics, and is stored and managed through the blockchain network to ensure data integrity and immutability. Adaptive parallel transmission channels and dynamic priority models are used to adjust the transmission resource allocation in real time according to the network status and target load, significantly reduce transmission delay, and improve efficiency and reliability. In the asynchronous verification process, the hash value is calculated and verified by the entropy value characteristics of the data sub-block and the data hash value in the storage package is compared to achieve fast and efficient data integrity verification. The file reorganization link builds a data block association map based on the double-layer cascade index, optimizes the sorting and boundary matching in combination with the content characteristics, and ensures the integrity and consistency of the synchronized file. Finally, the confirmation information is updated synchronously through the blockchain network to further improve the security and credibility of data transmission and reorganization. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 A schematic diagram of a flow chart of a method for parallel synchronization of unstructured files according to an embodiment of the present invention;
[0072] Figure 2 The figure is a schematic diagram of the structure of the unstructured file parallel synchronization system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0073] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0074] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0075] Figure 1 FIG. 1 is a flow chart of a method for parallel synchronization of unstructured files according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0076] S101. Obtain the target unstructured file in the source database, and perform intelligent block processing through a pre-trained file block model, wherein the file block model adaptively determines the optimal block size and number of blocks according to the content characteristics, data distribution characteristics and file structure characteristics of the target unstructured file, generates multiple data sub-blocks, generates a double-layer cascade index number containing block-level information and file-level information for each data sub-block, calculates a data hash value based on the entropy value characteristics of the data sub-block and the double-layer cascade index number, combines the data sub-block, the double-layer cascade index number and the data hash value to generate a data storage package, and distributes and stores it to multiple storage nodes in the blockchain network;
[0077] S102. Read the data storage package from each storage node in the blockchain network, and build a data transmission priority model based on the double-layer cascade index number. The data transmission priority model comprehensively evaluates the data importance of the data sub-block, the network transmission status, and the load of the target database, and dynamically allocates the transmission priority and network resource quota. According to the transmission priority, the data sub-block is transmitted to the target database through an adaptive parallel transmission channel. For the data sub-block that has completed the transmission, an asynchronous verification mechanism is used to calculate the verification hash value based on the entropy value characteristics of the data sub-block, and the verification hash value is compared with the data hash value in the data storage package. When the comparison results are consistent, a data verification identifier is generated for the corresponding data sub-block;
[0078] S103. Obtain all data sub-blocks with data verification identifiers and their corresponding double-layer cascade index numbers, input the data sub-blocks into a pre-trained file reorganization model, the file reorganization model constructs a data block association map based on the file level information of the double-layer cascade index numbers, calculates the combination weights between the data sub-blocks according to the data block association map, hierarchically sorts the data sub-blocks based on the combination weights, and uses the content features of the data sub-blocks to perform boundary feature matching and sequence optimization to generate a reorganization sequence, splices and reorganizes the data sub-blocks according to the reorganization sequence to obtain a synchronization target file, generates blockchain confirmation information including reorganization sequence information, file integrity verification information and synchronization status information, stores the synchronization target file in a target-end database, and submits the blockchain confirmation information to the blockchain network, synchronously updates the blockchain confirmation information in all storage nodes through the consensus mechanism of the blockchain network, and completes the parallel synchronization of the target unstructured file.
[0079] Among them, the target unstructured files refer to those files in the source database that do not have a fixed data model or structure. The file segmentation model is a pre-trained machine learning model that is used to segment unstructured files into multiple smaller data units (data sub-blocks). The purpose of segmentation is to improve the efficiency of file transmission and storage while ensuring that each sub-block can be processed and verified independently. The two-level cascade index number is a numbering system used to identify data sub-blocks. It consists of two levels of indexes, representing file-level information and block-level information respectively. File-level information contains global information about the file, such as the start and end positions of the file, or the overall structural characteristics of the file. Block-level information contains information about specific data sub-blocks, such as the relative position of each sub-block in the file.
[0080] The data hash value is a unique value of a fixed length obtained by calculating the data through a hash algorithm, which is used to represent the "fingerprint" or unique identifier of the data. The entropy value feature is a measure of the uncertainty or amount of information in the data. In the process of file segmentation or data transmission, the data with a higher entropy value usually means that it contains more information and is more important.
[0081] The data transmission priority model is an intelligent model for evaluating and adjusting the order and priority of data transmission. The model takes into account multiple factors, including the importance of data sub-blocks, the current network transmission status (such as bandwidth, latency, etc.), and the load of the target database. The asynchronous verification mechanism is a verification method that does not require the verification results to be obtained immediately when each step is completed, but allows the verification process to be carried out in the background and compared and confirmed at a later time.
[0082] The file reorganization model is an intelligent model for recombining data split into multiple sub-blocks into a complete file. The model builds a data block association map based on the double-layer cascade index number of the data sub-blocks, and optimizes the order and boundary characteristics of the sub-blocks by combining weights.
[0083] In an optional implementation, a target unstructured file in a source database is obtained, and intelligent block processing is performed using a pre-trained file block model, wherein the file block model adaptively determines an optimal block size and a number of blocks according to content features, data distribution features, and file structure features of the target unstructured file, and generates multiple data sub-blocks, including:
[0084] Obtain a target unstructured file from a source database, perform a bidirectional semantic scan on the target unstructured file, extract text content features, the text content features include semantic features and contextual association features, calculate the semantic association of adjacent text contents based on the semantic features and contextual association features, and take a position where the semantic association is lower than a preset association threshold as a first candidate block point;
[0085] Extract data distribution features of the target unstructured file, perform sliding window scanning according to a preset window size, calculate the local sensitive hash value of the data in each window, identify the data repetition pattern and density distribution features based on the local sensitive hash value, and take the position where the data repetition pattern mutates and the density distribution features change as the second candidate block point;
[0086] Analyze the structural features of the target unstructured file, extract the hierarchical features, format features and metadata features of the file, identify the structural separation position of the file based on the hierarchical features, format features and metadata features, and use the structural separation position as the third candidate block point;
[0087] The first candidate block point, the second candidate block point and the third candidate block point are combined to form a candidate block point set, and an evaluation feature vector is calculated for each candidate block point in the candidate block point set, wherein the evaluation feature vector includes a semantic integrity feature based on semantic association, a data balance feature based on data repetition pattern and density distribution feature, and a structural coherence feature based on structural feature;
[0088] The candidate block point set and its corresponding evaluation feature vector are input into a pre-trained file block model, the file block model calculates the optimization weight of each candidate block point according to the evaluation feature vector, screens and combines the candidate block point set based on the optimization weight, adaptively determines the optimal block size and number of blocks, and generates an initial block scheme;
[0089] The initial block partitioning scheme is input into a dynamic programming model, a cost function with semantic integrity features, data balance features and structural coherence features as optimization objectives is constructed, the cost function is input into the dynamic programming model, the optimal state transfer sequence is iteratively calculated through the dynamic programming model to obtain a final block partitioning scheme, the final block boundary position is determined according to the final block partitioning scheme, the target unstructured file is segmented according to the final block boundary position, and a plurality of data sub-blocks are generated.
[0090] Exemplarily, first, a target unstructured file is obtained from a source database, such as a relational database MySQL or a non-relational database MongoDB. The target file can be a document of various types, such as a text file, a PDF file, an HTML file, etc. For example, a PDF file named "document.pdf" is read from a table in a MySQL database.
[0091] Next, a bidirectional semantic scan is performed on the acquired target unstructured file to extract text content features. The BERT model is used here to encode the text and extract semantic features and contextual features. For example, for the sentence "The catsat on the mat.", the BERT model generates a vector representation for each word, which captures the semantic and contextual information of the word. Then, the semantic relevance of adjacent text content is calculated based on these vectors, for example, using cosine similarity. If the semantic relevance of two sentences is lower than a preset relevance threshold, such as 0.5, the position between them is regarded as the first candidate block point.
[0092] At the same time, data distribution features are extracted for the target unstructured file. A sliding window mechanism is used, for example, the window size is set to 1024 bytes, to scan the file. In each window, the local sensitive hash value of the data is calculated, for example, using the SimHash algorithm. By comparing the local sensitive hash values of adjacent windows, data repetition patterns and density distribution features can be identified. If the data repetition pattern mutates, such as from high repetition to low repetition, and the density distribution features change, such as from high density to low density, then the position is considered as the second candidate block point. Assuming that at a certain location, the SimHash value has changed significantly and the data density has also changed, then this location can be used as the second candidate block point.
[0093] In addition, the structural features of the target unstructured file are analyzed. The hierarchical features of the file are extracted, such as chapters, paragraphs, titles, etc.; format features, such as fonts, font sizes, tables, etc.; and metadata features, such as author, creation time, file size, etc. For example, for a PDF file, its directory structure, title information, table location, etc. can be extracted. Based on these structural features, the structural separation positions of the file are identified, such as the end of a chapter, the boundary of a table, etc. These structural separation positions will be regarded as the third candidate block points.
[0094] The first candidate block point, the second candidate block point and the third candidate block point are combined to form a set of candidate block points. For each candidate block point in the set, its evaluation feature vector is calculated. The evaluation feature vector includes a semantic integrity feature based on semantic association, a data balance feature based on data repetition pattern and density distribution features, and a structural coherence feature based on structural features. For example, the evaluation feature vector of a candidate block point can be expressed as [0.8, 0.7, 0.9], representing its semantic integrity, data balance and structural coherence, respectively.
[0095] The set of candidate block points and their corresponding evaluation feature vectors are input into a pre-trained file block model. The model can be a deep learning-based model, such as a Transformer model, and is pre-trained using a large amount of text data. The model calculates the optimization weight of each candidate block point based on the evaluation feature vector. For example, the higher the weight of the candidate block point, the better its block effect. Based on these optimization weights, the set of candidate block points is screened and combined, the optimal block size and number of blocks are adaptively determined, and the initial block scheme is generated.
[0096] Input the initial block plan into the dynamic programming model. Construct a cost function with semantic integrity, data balance, and structural coherence as optimization targets. For example, the cost function can be defined as the weighted sum of the three features, and the weights are adjusted according to the specific application scenario. Input the cost function into the dynamic programming model, and obtain the final block plan by iteratively calculating the optimal state transition sequence. Determine the final block boundary position according to the final block boundary position, and divide the target unstructured file according to the final block boundary position to generate multiple data sub-blocks.
[0097] In this embodiment, semantics, data distribution and structural features are comprehensively considered, and the logical boundaries of files can be more accurately identified, thereby improving the accuracy of block segmentation. The block size and number of blocks can be adaptively adjusted according to different types of files and different application scenarios, with strong flexibility. Through intelligent block segmentation, large unstructured files can be decomposed into multiple smaller data sub-blocks, thereby improving the efficiency of subsequent processing, such as information retrieval, text analysis, etc.
[0098] In an optional embodiment, a double-layer cascade index number including block-level information and file-level information is generated for each data sub-block, a data hash value is calculated based on the entropy value characteristics of the data sub-block and the double-layer cascade index number, and the data sub-block, the double-layer cascade index number and the data hash value are combined to generate a data storage package, which is distributed and stored in multiple storage nodes in the blockchain network, including:
[0099] Extract the file name, creation time and file size of the file to which each data sub-block belongs to form a metadata information set, serialize the metadata information set to obtain a metadata byte stream, perform hash calculation on the metadata byte stream to obtain a file identifier, and simultaneously obtain the position number of the data sub-block in the original file, convert the position number into a binary code to obtain a block number code, fill the block number code with leading zeros to obtain a standardized block number, and connect the file identifier and the standardized block number through a separator to generate a double-layer cascade index number;
[0100] Divide the data sub-block into multiple data segments according to a preset size, calculate the number of occurrences of each byte value in each data segment to obtain a byte distribution feature, calculate the entropy value of each data segment based on the byte distribution feature, and form an entropy value feature vector with the entropy values of all data segments;
[0101] Convert the double-layer cascade index number into an index byte sequence, perform a weighted sum operation on the entropy value feature vector to obtain a feature fusion value, concatenate the index byte sequence and the feature fusion value to generate a mixed feature sequence, and perform a hash calculation on the mixed feature sequence to obtain a data fingerprint hash value;
[0102] The data sub-blocks, double-layer cascade index numbers, and data fingerprint hash values are encapsulated into a data storage package, a storage location is determined on the hash ring of the blockchain network based on the data fingerprint hash value, a preset number of storage nodes closest to the storage location in a clockwise direction are selected as a target storage node set, and the data storage package is sent to each storage node in the target storage node set for storage.
[0103] Exemplarily, first, extract the file name, creation time, and file size of the file to which the data sub-block belongs to form a metadata information set. For example, a file named "document.txt" has a creation time of "2024-07-2710:00:00" and a file size of 1024KB, then this information constitutes the metadata information set of the file. Serialize the metadata information set into a byte stream, for example, using UTF-8 encoding to obtain a byte sequence. Perform a hash calculation on the byte sequence, for example, using the SHA-256 algorithm, to obtain a hash value as the identifier of the file.
[0104] At the same time, obtain the position number of the data sub-block in the original file. For example, the position number of the first data sub-block is 1, the position number of the second data sub-block is 2, and so on. Convert the position number to binary code, for example, convert the position number 1 to binary code "0001", and convert the position number 2 to binary code "0010". Fill the binary code with leading zeros to make its length reach a preset value, for example, fill it to 8 bits, and obtain a standardized block number, such as "00000001" and "00000010". Connect the file identifier and the standardized block number through a separator, for example, use "-" as a separator to generate a double-layer cascade index number, such as "f78a58e5...-00000001" and "f78a58e5...-00000010".
[0105] Next, the data sub-block is divided into multiple data segments according to the preset size. For example, a 1KB data sub-block can be divided into 4 data segments according to the 256-byte size. The number of times each byte value appears in each data segment is calculated to obtain the byte distribution characteristics. For example, in a data segment, the byte value "0x00" appears 10 times, "0x01" appears 5 times, and so on. The entropy value of each data segment is calculated based on the byte distribution characteristics. The entropy values of all data segments constitute an entropy value feature vector.
[0106] Then, the double-layer cascade index number is converted into an index byte sequence, for example, using UTF-8 encoding. A weighted sum operation is performed on the entropy value feature vector to obtain a feature fusion value. The index byte sequence and the feature fusion value are concatenated to generate a mixed feature sequence. A hash calculation is performed on the mixed feature sequence, for example, using the SHA-256 algorithm, to obtain a data fingerprint hash value.
[0107] Finally, the data sub-blocks, double-layer cascade index numbers, and data fingerprint hash values are encapsulated into a data storage package. The storage location is determined on the hash ring of the blockchain network based on the data fingerprint hash value. The preset number of storage nodes closest to the storage location in the clockwise direction is selected as the target storage node set. The data storage package is sent to each storage node in the target storage node set for storage.
[0108] In this embodiment, by combining the double-layer cascade index number and the entropy value characteristics of the data sub-block to calculate the data hash value, data tampering can be effectively prevented and the integrity and security of the data can be guaranteed. The data storage package is distributedly stored on the hash ring of the blockchain network based on the data fingerprint hash value, which can balance the storage load and improve storage efficiency. The design of the double-layer cascade index number can quickly locate the storage node where the data sub-block is located, thereby speeding up data retrieval.
[0109] In an optional implementation, data storage packages are read from each storage node in the blockchain network, a data transmission priority model is constructed based on the double-layer cascade index number, the data transmission priority model comprehensively evaluates the data importance of the data sub-block, the network transmission status, and the load of the target database, and dynamically allocates the transmission priority and the network resource quota. According to the transmission priority, the data sub-block is transmitted to the target database through the adaptive parallel transmission channel, including:
[0110] Read a data storage package from a storage node in the blockchain network, build an index directory tree based on the double-layer cascade index number in the data storage package, and extract data sub-blocks according to the index directory tree;
[0111] Perform feature matching on the identifiers in the double-layer cascade index number to obtain the inter-block correlation coefficient, calculate the data age decay value based on the timestamp, obtain the access intensity per unit time according to the access count, construct a data importance evaluation matrix with the inter-block correlation coefficient, age decay value, and access intensity, and obtain the data importance score through matrix operation;
[0112] Record the throughput change sequence, link load status sequence, and delay fluctuation sequence of the network transmission path, predict the future network status change trend through a sliding time window, and calculate the dynamic network score based on the network status change trend;
[0113] Collect resource usage indicators of the target database in real time, build a three-layer fuzzy rule set based on processor usage, memory usage, and input and output queue length, and calculate the comprehensive load score of the target end through rule reasoning;
[0114] The data importance score, dynamic network score, and target end comprehensive load score are normalized to construct a three-dimensional state vector, an action space matrix is generated according to historical transmission records, a transmission value function of each data sub-block is calculated based on the three-dimensional state vector and the action space matrix, an action sequence corresponding to the maximum transmission value is selected using a greedy strategy, the value function is iteratively updated until convergence, and the optimal action sequence after convergence is mapped to a transmission priority;
[0115] The data sub-blocks are prioritized according to the transmission priority, the data sub-blocks of adjacent priorities are combined and allocated to the same transmission channel to form a transmission channel group, and the throughput mean and variance of each transmission channel are calculated in real time. When the ratio of the variance to the mean exceeds a preset transmission threshold, the transmission efficiency score of each channel is calculated based on the information entropy principle, the channels with transmission efficiency scores lower than the efficiency threshold are dynamically reorganized, the reorganized transmission channels are reallocated to the transmission channel group, and the data sub-blocks are transmitted to the target database through the transmission channel group.
[0116] Exemplarily, first, data storage packets are read from each storage node of the blockchain network. These data packets contain double-layer cascade index numbers, which are used to build an index directory similar to a tree structure. Through this index directory, the required data sub-blocks can be quickly located and extracted. For example, if the double-layer cascade index number of a data packet is "A-1-2", it means that the data sub-block is located at the second sub-index position under the first main index of area A.
[0117] Next, the importance of the data sub-blocks is evaluated. The identifiers in the double-layer cascade index numbers are analyzed to find the correlation between the data blocks, and this relationship is quantified to obtain the inter-block correlation coefficient. For example, if the index numbers of two data blocks are very similar, they may have a high correlation. At the same time, the data time decay value is calculated according to the timestamp of the data packet. The older the data, the lower its timeliness. In addition, the number of accesses to each data block per unit time is counted to obtain the access intensity. The three indicators of inter-block correlation coefficient, time decay value and access intensity are combined to construct a data importance evaluation matrix. By analyzing this matrix, the data importance score of each data sub-block can be obtained. For example, if a data block has a strong correlation with many other data blocks, has a high access frequency, and has a newer timestamp, its importance score will be higher.
[0118] At the same time, the status of the network transmission path is evaluated. The change sequence of indicators such as throughput, link load and latency of the network transmission path is continuously recorded. A time window, such as data from the past 10 minutes, is used to predict the change trend of the future network status. Based on the predicted trend of network status changes, a dynamic network score is calculated. For example, if the network throughput is predicted to decrease in the future, the dynamic network score will decrease.
[0119] In addition, the load of the target database is collected in real time. The processor usage, memory occupancy, and input / output queue length of the target database are continuously monitored. Based on these indicators, a fuzzy rule set is constructed, for example, "If the processor usage is high and the memory usage is high, the database load is high." Reasoning is performed through these rules to calculate the comprehensive load score of the target database. For example, if the processor usage of the database is 90%, the memory usage is 85%, and the input / output queue length is 1000, then according to the preset fuzzy rules, it can be inferred that the database load score is high.
[0120] Then, the data importance score, dynamic network score, and target end comprehensive load score are normalized so that their values range from 0 to 1. The three normalized scores are combined into a three-dimensional state vector to describe the current data transmission status. Based on the historical transmission records, an action space matrix is generated, which contains all possible transmission actions, such as increasing the transmission priority of a certain data sub-block. Based on the current three-dimensional state vector and action space matrix, the transmission value function of each data sub-block is calculated. Using a greedy strategy, select the action sequence that can maximize the transmission value function. Continuously iterate and update the value function until it converges. Map the converged optimal action sequence to the transmission priority of the data sub-block.
[0121] Finally, the data sub-blocks are sorted according to their transmission priorities, and data sub-blocks with similar priorities are assigned to the same transmission channel to form a transmission channel group. The throughput mean and variance of each transmission channel are monitored in real time. If the ratio of the throughput variance to the mean of a channel exceeds the preset threshold, it means that the transmission efficiency of the channel is unstable. At this time, the transmission efficiency score of each channel is calculated based on the information entropy principle, and the channels with scores lower than the efficiency threshold are dynamically reorganized, and the reorganized channels are reallocated to the transmission channel group. Through these transmission channel groups, the data sub-blocks are transmitted to the target database. For example, data sub-blocks with priorities 1 to 10 are assigned to channel 1, and data sub-blocks with priorities 11 to 20 are assigned to channel 2. If the throughput of channel 1 fluctuates greatly, it is split into two channels and the data sub-blocks are reallocated.
[0122] In this embodiment, by giving priority to the transmission of important data and dynamically adjusting the transmission strategy according to the network and database load, network resources can be used to the maximum extent and data transmission efficiency can be improved. By real-time monitoring and dynamic adjustment of the transmission channel, network congestion and data loss can be avoided, and the stability and reliability of data transmission can be guaranteed. By optimizing resource allocation and improving transmission efficiency, the consumption of network bandwidth and storage resources can be reduced, thereby reducing system operating costs.
[0123] In an optional implementation, all data sub-blocks with data verification identifiers and their corresponding double-layer cascade index numbers are obtained, and the data sub-blocks are input into a pre-trained file reorganization model, the file reorganization model builds a data block association map based on the file level information of the double-layer cascade index number, calculates the combination weights between the data sub-blocks according to the data block association map, hierarchically sorts the data sub-blocks based on the combination weights, and uses the content features of the data sub-blocks to perform boundary feature matching and sequence optimization, and generates a reorganization sequence, including:
[0124] Obtain a data sub-block with a data verification identifier and its corresponding double-layer cascade index number, extract file-level information and block-level position information in the double-layer cascade index number by segmented decoding, perform multi-scale sliding window scanning on the data sub-block to obtain a feature sequence, associate and map the feature sequence with the file-level information to construct a feature combination;
[0125] Inputting the feature combination into a pre-trained file reorganization model, using the file reorganization model to extract semantic features of data sub-blocks, calculating the association matrix between data sub-blocks based on a multi-head attention mechanism, and constructing a data block association map in combination with the block-level position information;
[0126] In the data block association graph, spatial attention and channel attention are integrated to extract local structural features of nodes, degree distribution entropy and clustering coefficient of nodes are calculated to obtain structural importance, and the semantic features are adaptively weighted integrated with the structural importance to generate a combined weight;
[0127] Using a gated recurrent neural network to perform time series analysis on the combination weights to generate a time series feature sequence, performing correlation clustering on the data sub-blocks according to the continuity of the time series feature sequence to obtain a hierarchical division result, calculating a local sorting order based on the semantic features and position information of the data sub-blocks within each hierarchy, and constructing an initial reorganization scheme including the hierarchical division result and the local sorting order;
[0128] Extracting boundary region features of adjacent data sub-blocks according to the initial reorganization scheme, calculating semantic similarity of boundary region features using a pre-built deep feature extraction network to obtain boundary matching, and when the boundary matching is lower than an adaptive matching threshold, adjusting the local sorting order based on the semantic relevance of the boundary region features to generate an optimized reorganization scheme;
[0129] The characteristic weights between the data sub-blocks are calculated according to the optimized reorganization scheme, the data sub-blocks are sequentially reorganized based on the characteristic weights to obtain a reorganized sequence, and after verifying the integrity of the reorganized sequence, the verification result is updated to the double-layer cascade index number.
[0130] Exemplarily, first, it is necessary to obtain a data sub-block with a data verification mark from the storage. Each data sub-block corresponds to a double-layer cascade index number, which consists of two parts: file-level information and block-level location information. The file-level information provides a description of the overall structure of the file, while the block-level location information indicates the specific location of the data sub-block in the file. In order to extract this information, a segmented decoding method is used to extract the file-level information and block-level location information from the double-layer cascade index number respectively.
[0131] Then, for each data sub-block, a multi-scale sliding window scanning technique is used to process it. This scanning method can extract the features of the data sub-block from different scales and dimensions, ensuring that the diverse patterns and information in the file are captured. The role of the multi-scale sliding window is to continuously extract the features of the data fragments by moving the window, thereby more comprehensively reflecting the structure and content of the file. Through mapping with file-level information, these feature sequences are combined into a comprehensive feature set, which includes not only the local information of each data sub-block, but also combines the global structure of the file.
[0132] These feature combinations are then input into a pre-trained file reorganization model. The core task of this model is to extract the semantic features of the data sub-blocks. Semantic features refer to the deep information carried by the data sub-blocks, such as the grammatical structure in the text, object recognition in the image, etc. The file reorganization model uses a multi-head attention mechanism to calculate the association matrix between the data sub-blocks. The association matrix reflects the mutual relationship and similarity between different data sub-blocks, indicating their possible combination methods in the file reorganization process. At this point, the relationship between the data sub-blocks is converted into a graph, where the nodes in the graph represent the data sub-blocks and the edges represent the strength of the association between them.
[0133] Based on the generated correlation matrix, combined with the location information of each data sub-block, the model constructs a data block correlation map. The data block correlation map shows the possible relationships between all data sub-blocks during the reorganization process. It groups sub-blocks with strong correlations together through the connection relationship between nodes and edges to form a reasonable reorganization structure. The generation of the map helps to determine which data sub-blocks should be prioritized during the reorganization process.
[0134] Next, the file reorganization model combines spatial attention and channel attention in the data block association graph to further extract the local structural features of the nodes. Spatial attention focuses on capturing the local features of the data block in space, while channel attention is weighted from multiple feature channels to help the model better understand the multi-dimensional properties of the data sub-blocks. Through these attention mechanisms, the model can calculate the structural importance of each data sub-block, which reflects the relative importance of the data sub-block in the file. The degree distribution entropy value and clustering coefficient are used to quantify the connection strength and organization between nodes (data sub-blocks), thereby determining their role in the entire file structure.
[0135] After weighted fusion of semantic features and structural importance, a combination weight is generated. The combination weight is used as the basis for sorting during the file reorganization process. A higher weight means that the data sub-block should be given priority in the reorganization. After obtaining the combination weight, the file reorganization model uses a gated recurrent neural network (GRU) to perform temporal analysis on the combination weight in order to capture the temporal characteristics of the data sub-block during the reorganization process. The temporal analysis helps the model identify those data sub-blocks with time dependencies and give priority to these sub-blocks with higher temporal continuity during the reorganization process.
[0136] Based on the results of time series analysis, the model performs correlation clustering on the data sub-blocks. The goal of correlation clustering is to assign those sub-blocks that are strongly correlated to the same level. The hierarchical division helps to clarify which data sub-blocks should be processed first and which should be processed later. Within each level, based on the semantic features and location information of the data sub-blocks, the model calculates the local sorting order to ensure that the data sub-blocks can be reorganized in the correct order at the same level.
[0137] After generating the initial reorganization plan, the boundary area features of adjacent data sub-blocks are extracted next. The boundary area refers to the connecting part between adjacent data sub-blocks, which is usually crucial for the correctness of file reorganization. Through the pre-built deep feature extraction network, the model calculates the semantic similarity of the boundary area features. The semantic similarity measures the degree of similarity between the boundary areas of two data sub-blocks, similar to the semantic similarity in text. If the calculated boundary matching degree is lower than the preset adaptive matching threshold, the model adjusts the local sorting order to ensure that the boundary docking between data sub-blocks is more natural and smooth.
[0138] Finally, the model calculates the feature weights between data sub-blocks based on the optimized reorganization scheme. These feature weights assign a priority to each data sub-block, and a higher feature weight indicates that the sub-block should be processed first during reorganization. The file reorganization model will sequentially reorganize the data sub-blocks based on these feature weights to ensure that the order of the files is correct and consistent with the structure and content of the original files. After the reorganization is completed, the model verifies the integrity of the reorganized sequence to ensure that the files are not lost or damaged. The verification results will be updated to the double-layer cascade index number to ensure that the file reorganization process can be traced and verified.
[0139] In this embodiment, through multi-scale feature extraction, semantic analysis and time series modeling and other technologies, the structural features and semantic information in the file can be intelligently captured to achieve accurate sorting and optimized reorganization of data sub-blocks. At the same time, combined with the attention mechanism and recurrent neural network, the relevance of the data blocks is deeply analyzed, thereby improving the accuracy and efficiency of the reorganization process while ensuring the integrity of the file content. This solution can significantly improve the parallelism and reliability of data transmission when processing large-scale unstructured files, and reduce the error rate and data loss risk during the transmission process.
[0140] In an optional implementation, the data sub-blocks are spliced and reassembled according to the reorganization sequence to obtain a synchronization target file, blockchain confirmation information including reorganization sequence information, file integrity verification information and synchronization status information is generated, the synchronization target file is stored in the target end database, and the blockchain confirmation information is submitted to the blockchain network. The blockchain confirmation information is synchronously updated in all storage nodes through the consensus mechanism of the blockchain network, and the parallel synchronization of the target unstructured file is completed, including:
[0141] A double-buffer streaming mechanism is used to splice and reorganize data sub-blocks, where the first buffer is used for reading data sub-blocks and verifying boundary feature matching, and the second buffer is used for sequential writing of verified data sub-blocks. When the amount of data in the second buffer reaches a preset threshold, the data sub-blocks are written in batches to the synchronization target file and the write position, data length and data block identifier are recorded to generate a file mapping table.
[0142] A multi-level hash verification tree is constructed based on the file mapping table, wherein the hash value and boundary feature information of the data block are stored in the leaf node, the subtree root hash value and the data block association degree are stored in the intermediate node, the root node stores the overall hash value and file metadata, and the file integrity verification information is generated through the verification relationship between the nodes;
[0143] Convert the multi-level hash verification tree into a compressed coding sequence, and package the compressed coding sequence, the recombined sequence information and the synchronization status information in a segmented coding manner to generate blockchain confirmation information with a hierarchical structure, each layer containing an independent digital signature and verification field;
[0144] Based on the consistent hashing algorithm, the blockchain confirmation information is mapped to the confirmation node ring of the blockchain network, multiple nodes in the clockwise direction of the mapping position are selected as the confirmation node group, and the two-phase commit protocol is used within the confirmation node group to complete the pre-confirmation;
[0145] The pre-confirmed blockchain confirmation information is sharded, and the confirmation information is divided into multiple verification shards according to the hierarchical structure of the multi-level hash verification tree, and the consensus algorithm based on practical Byzantine fault tolerance is executed in parallel in different shards, and the blockchain confirmation information is synchronously updated in all storage nodes through cross-shard verification;
[0146] After receiving the blockchain confirmation information, the storage node calculates the verification value layer by layer according to the multi-level hash verification tree, compares it with the verification field in the blockchain confirmation information, and generates a synchronization request including the inconsistent data block location and verification path;
[0147] A data synchronization channel is established based on the verification path of the synchronization request, and a selective retransmission mechanism is used to transmit inconsistent data blocks. The receiving node verifies the correctness of the data block through the multi-level hash verification tree, and the verified synchronization target file is updated to the target end database to complete the parallel synchronization of the target unstructured file.
[0148] Exemplarily, first, the target unstructured file is divided into blocks to generate multiple data sub-blocks, and a unique identifier and boundary feature information are generated for each data sub-block. For example, a 1GB video file is divided into 1000 data sub-blocks of 1MB in size, and each data sub-block has a unique identifier (e.g., a sub-block number) and boundary feature information (e.g., a number of bytes at the start position of the sub-block). At the same time, a recombined sequence containing the sequence information of the data sub-blocks is generated.
[0149] Then, a double-buffer streaming mechanism is used to splice and reassemble the data sub-blocks. The first buffer is used to receive and verify the data sub-blocks, and the verification content includes the data sub-block identification and boundary feature matching. The second buffer is used to cache the verified data sub-blocks. Assume that the first buffer receives data sub-blocks numbered 1, 2, and 3, and verifies their identification and boundary feature information. After the verification is passed, these data sub-blocks are placed in the second buffer. When the amount of data in the second buffer reaches a preset threshold, such as 100MB, the data sub-blocks in the buffer are written to the synchronization target file in the order specified by the reorganization sequence. And record the write position, data length and data block identification, and generate a file mapping table. For example, record that the write position of sub-block 1 is 0 and the length is 1MB; the write position of sub-block 2 is 1MB and the length is 1MB; the write position of sub-block 3 is 2MB and the length is 1MB, and so on.
[0150] Next, a multi-level hash verification tree is constructed based on the file mapping table. The hash value (e.g., SHA256 hash value) and boundary feature information of the data block are stored in the leaf node. The intermediate node stores the root hash value of the subtree and the data block association degree (e.g., the number of data blocks contained in the subtree). The root node stores the overall hash value and file metadata (e.g., file name, file size). File integrity verification information is generated through the verification relationship between nodes. For example, assuming that an intermediate node contains sub-block 1 and sub-block 2, the value stored in the node is the hash value of the hash value of sub-block 1 and sub-block 2, and the data block association degree 2.
[0151] Afterwards, the multi-level hash verification tree is converted into a compressed coding sequence, for example, using Huffman coding or other compression algorithms. The compressed coding sequence, reorganized sequence information, and synchronization status information (for example, synchronization progress) are packaged in a segmented coding manner to generate blockchain confirmation information with a hierarchical structure. Each layer contains independent digital signatures and verification fields (for example, timestamps).
[0152] Then, based on the consistent hashing algorithm, the blockchain confirmation information is mapped to the confirmation node ring of the blockchain network. Assuming there are 10 confirmation nodes, the blockchain confirmation information is mapped to node 3. Then nodes 3, 4, and 5 are selected as the confirmation node group. The two-phase commit protocol is used to complete the pre-confirmation within the confirmation node group. In the first stage, the blockchain confirmation information is sent to all confirmation nodes to request confirmation. In the second stage, the confirmation results of the confirmation nodes are collected. If all nodes confirm, the pre-confirmation passes.
[0153] The pre-confirmed blockchain confirmation information is sharded. The confirmation information is divided into multiple verification shards according to the hierarchical structure of the multi-level hash verification tree. For example, the root node information is used as one shard, and the next layer of node information is used as another shard. The consensus algorithm based on practical Byzantine fault tolerance is executed in parallel in different shards, and the blockchain confirmation information is synchronously updated in all storage nodes through cross-shard verification.
[0154] After receiving the blockchain confirmation information, the storage node calculates the verification value layer by layer according to the multi-level hash verification tree and compares it with the verification field in the blockchain confirmation information. For example, the hash value of the hash value of sub-block 1 and sub-block 2 is calculated and compared with the hash value stored in the corresponding intermediate node. A synchronization request containing the inconsistent data block location and verification path is generated.
[0155] Finally, a data synchronization channel is established based on the verification path of the synchronization request, and a selective retransmission mechanism is used to transmit inconsistent data blocks. The receiving node verifies the correctness of the data block through a multi-level hash verification tree, and updates the verified synchronization target file to the target database to complete the parallel synchronization of the target unstructured file.
[0156] In this embodiment, the file synchronization speed is significantly improved through data segmentation, parallel processing and selective retransmission mechanism, which is particularly suitable for the synchronization of large-scale unstructured data. Multi-level hash verification tree and blockchain technology ensure the integrity and consistency of data during transmission and storage, effectively preventing data damage or tampering. Distributed storage and consensus mechanism improve the fault tolerance and availability of the system, and can ensure the smooth completion of synchronization tasks even if some nodes fail.
[0157] In an optional implementation, the blockchain confirmation information is mapped to the confirmation node ring of the blockchain network based on a consistent hashing algorithm, a plurality of nodes in the clockwise direction of the mapping position are selected as a confirmation node group, and a two-phase commit protocol is used within the confirmation node group to complete the pre-confirmation, including:
[0158] Obtain the computing power value and storage capacity value of each confirmation node in the blockchain network, determine the number of virtual nodes to be allocated based on the computing power value and storage capacity value, use a hash algorithm to calculate the node identification hash value of the virtual node and map it to the hash ring, and build a confirmation node ring;
[0159] Calculate the confirmation information hash value of the blockchain confirmation information based on the consistent hashing algorithm and map it to the confirmation node ring, obtain the network delay value of each confirmation node in the confirmation node ring and the historical confirmation performance value in the node scoring table, and select multiple confirmation nodes as the confirmation node group in a clockwise direction from the mapping position of the confirmation information hash value based on the network delay value and the historical confirmation performance value;
[0160] The confirmation information hash value, timestamp and digital signature of the coordinating node are combined to generate a pre-confirmation request, and the node corresponding to the mapping position in the confirmation node group acts as a coordinating node to send the pre-confirmation request to the remaining confirmation nodes in the confirmation node group, and the confirmation nodes in the confirmation node group verify the legitimacy of the pre-confirmation request and return a preparation response;
[0161] The coordination node counts the number of ready responses of the confirmation node group, generates a submission request when the number of ready responses exceeds a preset ratio, sends the submission request to the confirmation node group, and the confirmation nodes in the confirmation node group perform local confirmation and generate a confirmation log;
[0162] A timeout detection mechanism is used to monitor the online status of the coordination node, and when a coordination node failure is detected, a new coordination node is selected from the confirmation node group, and confirmation authority is allocated to the new coordination node;
[0163] The confirmation operation information in the confirmation log is constructed into a verification tree, the root node value of the verification tree is combined with the confirmation signature of the confirmation node group to generate a pre-confirmation result, and the pre-confirmation result is sent to the remaining confirmation nodes in the blockchain network. After receiving the pre-confirmation result, each remaining confirmation node verifies the number of confirmation signatures and the proof path of the verification tree. After the verification is passed, the pre-confirmation result is recorded as a checkpoint to complete the pre-confirmation of the blockchain confirmation information.
[0164] Exemplarily, first, a confirmation node ring is constructed. The computing power value and storage capacity value of each confirmation node in the blockchain network are obtained. For example, the computing power value of node A is 100, and the storage capacity value is 500GB; the computing power value of node B is 150, and the storage capacity value is 1TB; the computing power value of node C is 80, and the storage capacity value is 256GB. Based on these values, the number of virtual nodes to be allocated is determined. The higher the computing power value and storage capacity value of the node, the more virtual nodes are allocated. For example, node A is allocated 2 virtual nodes, node B is allocated 3 virtual nodes, and node C is allocated 1 virtual node. Then, a hash algorithm (such as SHA-256) is used to calculate the node identification hash value of each virtual node, and these hash values are mapped to a ring-shaped hash space to form a confirmation node ring.
[0165] Next, map the blockchain confirmation information to the confirmation node ring. Use the consistent hashing algorithm to calculate the confirmation information hash value of the blockchain confirmation information. For example, the content of a confirmation information is "transaction A transfers 10 coins to transaction B", and its confirmation information hash value is calculated to be "hash_value_1". Map the hash value to the confirmation node ring.
[0166] Then, select a confirmation node group. Obtain the network delay value of each confirmation node in the confirmation node ring and the historical confirmation performance value in the node scoring table. For example, the network delay value of node A is 10ms, and the historical confirmation performance value is 95; the network delay value of node B is 5ms, and the historical confirmation performance value is 90; the network delay value of node C is 15ms, and the historical confirmation performance value is 85. Based on these values, select multiple confirmation nodes as the confirmation node group in a clockwise direction from the mapping position of the confirmation information hash value. For example, select nodes A and B with low network delay values and high historical confirmation performance values as the confirmation node group.
[0167] After that, two-phase commit pre-confirmation is performed. The confirmation information hash value "hash_value_1", the current timestamp (e.g., January 1, 2024 12:00:00) and the digital signature of the coordinating node (e.g., node A) are combined to generate a pre-confirmation request. Coordinating node A sends a pre-confirmation request to the remaining confirming nodes (node B) in the confirming node group. Node B verifies the legitimacy of the pre-confirmation request, such as verifying the validity of the digital signature, and returns a prepared response.
[0168] Coordinating node A counts the number of prepared responses of the confirmation node group. If the number of prepared responses exceeds a preset ratio (e.g., 50%), a commit request is generated and sent to the confirmation node group (node A and node B). The confirmation nodes in the confirmation node group perform local confirmation operations, such as writing confirmation information to a local database, and generate a confirmation log.
[0169] At the same time, a timeout detection mechanism is used to monitor the online status of the coordination node. If a failure of coordination node A is detected, a new coordination node (such as node B) is selected from the confirmation node group and confirmation authority is assigned to the new coordination node.
[0170] Finally, generate the pre-confirmation result. Construct the confirmation operation information in the confirmation log into a verification tree (e.g., a Merkle tree). Combine the root node value of the verification tree with the confirmation signature of the confirmation node group (node A and node B) to generate the pre-confirmation result. Send the pre-confirmation result to the remaining confirmation nodes (e.g., node C) in the blockchain network. After receiving the pre-confirmation result, the remaining confirmation nodes verify the number of confirmation signatures and the proof path of the verification tree. After the verification is passed, the pre-confirmation result is recorded as a checkpoint to complete the pre-confirmation of the blockchain confirmation information.
[0171] In this embodiment, the consistent hash algorithm is used to quickly locate the confirmation node group, and the two-phase commit protocol is used for pre-confirmation, which shortens the confirmation time and improves the confirmation efficiency. The timeout detection mechanism and the coordination node switching mechanism are used to effectively deal with coordination node failures and ensure the reliability of the pre-confirmation process. The integrity and non-tamperability of the confirmation information are ensured through the verification tree and confirmation signature mechanism, which improves the security of the confirmation.
[0172] Figure 2 FIG. 1 is a schematic diagram of a structure of a parallel synchronization system for unstructured files according to an embodiment of the present invention. Figure 2 As shown, the system comprises:
[0173] The first unit is used to obtain the target unstructured file in the source database, and perform intelligent block processing through a pre-trained file block model, wherein the file block model adaptively determines the optimal block size and the number of blocks according to the content characteristics, data distribution characteristics and file structure characteristics of the target unstructured file, generates multiple data sub-blocks, generates a double-layer cascade index number containing block-level information and file-level information for each data sub-block, calculates a data hash value based on the entropy value characteristics of the data sub-block and the double-layer cascade index number, combines the data sub-block, the double-layer cascade index number and the data hash value to generate a data storage package, and distributes and stores it to multiple storage nodes in the blockchain network;
[0174] The second unit is used to read the data storage package from each storage node in the blockchain network, build a data transmission priority model based on the double-layer cascade index number, and the data transmission priority model comprehensively evaluates the data importance of the data sub-block, the network transmission status and the load of the target database, dynamically allocates the transmission priority and the network resource quota, and transmits the data sub-block to the target database through the adaptive parallel transmission channel according to the transmission priority. For the data sub-block that has completed the transmission, an asynchronous verification mechanism is used to calculate the verification hash value based on the entropy value characteristics of the data sub-block, and the verification hash value is compared with the data hash value in the data storage package. When the comparison results are consistent, a data verification identifier is generated for the corresponding data sub-block;
[0175] The third unit is used to obtain all data sub-blocks with data verification identifiers and their corresponding double-layer cascade index numbers, input the data sub-blocks into a pre-trained file reorganization model, the file reorganization model constructs a data block association map based on the file level information of the double-layer cascade index number, calculates the combination weights between the data sub-blocks according to the data block association map, hierarchically sorts the data sub-blocks based on the combination weights, and uses the content features of the data sub-blocks to perform boundary feature matching and sequence optimization to generate a reorganization sequence, splices and reorganizes the data sub-blocks according to the reorganization sequence to obtain a synchronization target file, generates blockchain confirmation information including reorganization sequence information, file integrity verification information and synchronization status information, stores the synchronization target file in a target end database, and submits the blockchain confirmation information to the blockchain network, synchronously updates the blockchain confirmation information in all storage nodes through the consensus mechanism of the blockchain network, and completes the parallel synchronization of the target unstructured file.
[0176] According to a third aspect of the embodiments of the present invention,
[0177] An electronic device is provided, comprising:
[0178] processor;
[0179] a memory for storing processor-executable instructions;
[0180] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0181] According to a fourth aspect of the embodiments of the present invention,
[0182] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0183] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for parallel synchronization of unstructured files, characterized in that: include: Obtain the target unstructured file in the source database, and perform intelligent block processing through a pre-trained file block model, wherein the file block model adaptively determines the optimal block size and number of blocks according to the content characteristics, data distribution characteristics, and file structure characteristics of the target unstructured file, generates multiple data sub-blocks, generates a double-layer cascade index number containing block-level information and file-level information for each data sub-block, calculates a data hash value based on the entropy value characteristics of the data sub-block and the double-layer cascade index number, combines the data sub-block, the double-layer cascade index number, and the data hash value to generate a data storage package, which is distributed and stored in multiple storage nodes in the blockchain network; Reading data storage packages from each storage node in the blockchain network, building a data transmission priority model based on the double-layer cascade index number, the data transmission priority model comprehensively evaluates the data importance of the data sub-block, the network transmission status and the load of the target database, dynamically allocates transmission priority and network resource quota, and transmits the data sub-block to the target database through an adaptive parallel transmission channel according to the transmission priority. For the data sub-block that has completed the transmission, an asynchronous verification mechanism is used to calculate a verification hash value based on the entropy value characteristics of the data sub-block, and the verification hash value is compared with the data hash value in the data storage package. When the comparison results are consistent, a data verification identifier is generated for the corresponding data sub-block; All data sub-blocks with data verification identifiers and their corresponding double-layer cascade index numbers are obtained, and the data sub-blocks are input into a pre-trained file reorganization model. The file reorganization model constructs a data block association map based on the file level information of the double-layer cascade index number, calculates the combination weights between the data sub-blocks according to the data block association map, hierarchically sorts the data sub-blocks based on the combination weights, and uses the content features of the data sub-blocks to perform boundary feature matching and sequence optimization to generate a reorganization sequence, splices and reorganizes the data sub-blocks according to the reorganization sequence to obtain a synchronization target file, generates blockchain confirmation information including reorganization sequence information, file integrity verification information and synchronization status information, stores the synchronization target file in a target-end database, and submits the blockchain confirmation information to the blockchain network. The blockchain confirmation information is synchronously updated in all storage nodes through the consensus mechanism of the blockchain network to complete the parallel synchronization of the target unstructured file.
2. The method according to claim 1, characterized in that Obtain the target unstructured file in the source database, and perform intelligent block processing through a pre-trained file block model. The file block model adaptively determines the optimal block size and number of blocks according to the content characteristics, data distribution characteristics and file structure characteristics of the target unstructured file, and generates multiple data sub-blocks including: Obtain a target unstructured file from a source database, perform a bidirectional semantic scan on the target unstructured file, extract text content features, the text content features include semantic features and contextual association features, calculate the semantic association of adjacent text contents based on the semantic features and contextual association features, and take a position where the semantic association is lower than a preset association threshold as a first candidate block point; Extract data distribution features of the target unstructured file, perform sliding window scanning according to a preset window size, calculate the local sensitive hash value of the data in each window, identify the data repetition pattern and density distribution features based on the local sensitive hash value, and take the position where the data repetition pattern mutates and the density distribution features change as the second candidate block point; Analyze the structural features of the target unstructured file, extract the hierarchical features, format features and metadata features of the file, identify the structural separation position of the file based on the hierarchical features, format features and metadata features, and use the structural separation position as the third candidate block point; The first candidate block point, the second candidate block point and the third candidate block point are combined to form a candidate block point set, and an evaluation feature vector is calculated for each candidate block point in the candidate block point set, wherein the evaluation feature vector includes a semantic integrity feature based on semantic association, a data balance feature based on data repetition pattern and density distribution feature, and a structural coherence feature based on structural feature; The candidate block point set and its corresponding evaluation feature vector are input into a pre-trained file block model, the file block model calculates the optimization weight of each candidate block point according to the evaluation feature vector, screens and combines the candidate block point set based on the optimization weight, adaptively determines the optimal block size and number of blocks, and generates an initial block scheme; The initial block partitioning scheme is input into a dynamic programming model, a cost function with semantic integrity features, data balance features and structural coherence features as optimization objectives is constructed, the cost function is input into the dynamic programming model, the optimal state transfer sequence is iteratively calculated through the dynamic programming model to obtain a final block partitioning scheme, the final block boundary position is determined according to the final block partitioning scheme, the target unstructured file is segmented according to the final block boundary position, and a plurality of data sub-blocks are generated.
3. The method according to claim 1, characterized in that Generate a double-layer cascade index number containing block-level information and file-level information for each data sub-block, calculate a data hash value based on the entropy value characteristics of the data sub-block and the double-layer cascade index number, combine the data sub-block, the double-layer cascade index number and the data hash value to generate a data storage package, and distribute and store it to multiple storage nodes in the blockchain network, including: Extract the file name, creation time and file size of the file to which each data sub-block belongs to form a metadata information set, serialize the metadata information set to obtain a metadata byte stream, perform hash calculation on the metadata byte stream to obtain a file identifier, and simultaneously obtain the position number of the data sub-block in the original file, convert the position number into a binary code to obtain a block number code, fill the block number code with leading zeros to obtain a standardized block number, and connect the file identifier and the standardized block number through a separator to generate a double-layer cascade index number; Divide the data sub-block into multiple data segments according to a preset size, calculate the number of occurrences of each byte value in each data segment to obtain a byte distribution feature, calculate the entropy value of each data segment based on the byte distribution feature, and form an entropy value feature vector with the entropy values of all data segments; Convert the double-layer cascade index number into an index byte sequence, perform a weighted sum operation on the entropy value feature vector to obtain a feature fusion value, concatenate the index byte sequence and the feature fusion value to generate a mixed feature sequence, and perform a hash calculation on the mixed feature sequence to obtain a data fingerprint hash value; The data sub-blocks, double-layer cascade index numbers, and data fingerprint hash values are encapsulated into a data storage package, a storage location is determined on the hash ring of the blockchain network based on the data fingerprint hash value, a preset number of storage nodes closest to the storage location in a clockwise direction are selected as a target storage node set, and the data storage package is sent to each storage node in the target storage node set for storage.
4. The method according to claim 1, characterized in that: Reading data storage packages from each storage node in the blockchain network, building a data transmission priority model based on the double-layer cascade index number, the data transmission priority model comprehensively evaluates the data importance of the data sub-block, the network transmission status and the load of the target database, dynamically allocating the transmission priority and the network resource quota, and transmitting the data sub-block to the target database through the adaptive parallel transmission channel according to the transmission priority includes: Read a data storage package from a storage node in the blockchain network, build an index directory tree based on the double-layer cascade index number in the data storage package, and extract data sub-blocks according to the index directory tree; Perform feature matching on the identifiers in the double-layer cascade index number to obtain the inter-block correlation coefficient, calculate the data age decay value based on the timestamp, obtain the access intensity per unit time according to the access count, construct a data importance evaluation matrix with the inter-block correlation coefficient, age decay value, and access intensity, and obtain the data importance score through matrix operation; Record the throughput change sequence, link load status sequence, and delay fluctuation sequence of the network transmission path, predict the future network status change trend through a sliding time window, and calculate the dynamic network score based on the network status change trend; Collect resource usage indicators of the target database in real time, build a three-layer fuzzy rule set based on processor usage, memory usage, and input and output queue length, and calculate the comprehensive load score of the target end through rule reasoning; The data importance score, dynamic network score, and target end comprehensive load score are normalized to construct a three-dimensional state vector, an action space matrix is generated according to historical transmission records, a transmission value function of each data sub-block is calculated based on the three-dimensional state vector and the action space matrix, an action sequence corresponding to the maximum transmission value is selected using a greedy strategy, the value function is iteratively updated until convergence, and the optimal action sequence after convergence is mapped to a transmission priority; The data sub-blocks are prioritized according to the transmission priority, the data sub-blocks of adjacent priorities are combined and allocated to the same transmission channel to form a transmission channel group, and the throughput mean and variance of each transmission channel are calculated in real time. When the ratio of the variance to the mean exceeds a preset transmission threshold, the transmission efficiency score of each channel is calculated based on the information entropy principle, the channels with transmission efficiency scores lower than the efficiency threshold are dynamically reorganized, the reorganized transmission channels are reallocated to the transmission channel group, and the data sub-blocks are transmitted to the target database through the transmission channel group.
5. The method according to claim 1, characterized in that Obtain all data sub-blocks with data verification identifiers and their corresponding double-layer cascade index numbers, input the data sub-blocks into a pre-trained file reorganization model, the file reorganization model builds a data block association map based on the file level information of the double-layer cascade index numbers, calculates the combination weights between the data sub-blocks according to the data block association map, hierarchically sorts the data sub-blocks based on the combination weights, and uses the content features of the data sub-blocks to perform boundary feature matching and sequence optimization, and generates a reorganization sequence including: Obtain a data sub-block with a data verification identifier and its corresponding double-layer cascade index number, extract file-level information and block-level position information in the double-layer cascade index number by segmented decoding, perform multi-scale sliding window scanning on the data sub-block to obtain a feature sequence, associate and map the feature sequence with the file-level information to construct a feature combination; Inputting the feature combination into a pre-trained file reorganization model, using the file reorganization model to extract semantic features of data sub-blocks, calculating the association matrix between data sub-blocks based on a multi-head attention mechanism, and constructing a data block association map in combination with the block-level position information; In the data block association graph, spatial attention and channel attention are integrated to extract local structural features of nodes, degree distribution entropy and clustering coefficient of nodes are calculated to obtain structural importance, and the semantic features are adaptively weighted integrated with the structural importance to generate a combined weight; Using a gated recurrent neural network to perform time series analysis on the combination weights to generate a time series feature sequence, performing correlation clustering on the data sub-blocks according to the continuity of the time series feature sequence to obtain a hierarchical division result, calculating a local sorting order based on the semantic features and position information of the data sub-blocks within each hierarchy, and constructing an initial reorganization scheme including the hierarchical division result and the local sorting order; Extracting boundary region features of adjacent data sub-blocks according to the initial reorganization scheme, calculating semantic similarity of boundary region features using a pre-built deep feature extraction network to obtain boundary matching, and when the boundary matching is lower than an adaptive matching threshold, adjusting the local sorting order based on the semantic relevance of the boundary region features to generate an optimized reorganization scheme; The characteristic weights between the data sub-blocks are calculated according to the optimized reorganization scheme, the data sub-blocks are sequentially reorganized based on the characteristic weights to obtain a reorganized sequence, and after verifying the integrity of the reorganized sequence, the verification result is updated to the double-layer cascade index number.
6. The method according to claim 1, characterized in that According to the reorganization sequence, the data sub-blocks are spliced and reorganized to obtain a synchronization target file, and blockchain confirmation information including reorganization sequence information, file integrity verification information and synchronization status information is generated. The synchronization target file is stored in the target end database, and the blockchain confirmation information is submitted to the blockchain network. The blockchain confirmation information is synchronously updated in all storage nodes through the consensus mechanism of the blockchain network, and the parallel synchronization of the target unstructured file is completed, including: A double-buffer streaming mechanism is used to splice and reorganize data sub-blocks, where the first buffer is used for reading data sub-blocks and verifying boundary feature matching, and the second buffer is used for sequential writing of verified data sub-blocks. When the amount of data in the second buffer reaches a preset threshold, the data sub-blocks are written in batches to the synchronization target file and the write position, data length and data block identifier are recorded to generate a file mapping table. A multi-level hash verification tree is constructed based on the file mapping table, wherein the hash value and boundary feature information of the data block are stored in the leaf node, the subtree root hash value and the data block association degree are stored in the intermediate node, the root node stores the overall hash value and file metadata, and the file integrity verification information is generated through the verification relationship between the nodes; Convert the multi-level hash verification tree into a compressed coding sequence, and package the compressed coding sequence, the recombined sequence information and the synchronization status information in a segmented coding manner to generate blockchain confirmation information with a hierarchical structure, each layer containing an independent digital signature and verification field; Based on the consistent hashing algorithm, the blockchain confirmation information is mapped to the confirmation node ring of the blockchain network, multiple nodes in the clockwise direction of the mapping position are selected as the confirmation node group, and the two-phase commit protocol is used within the confirmation node group to complete the pre-confirmation; The pre-confirmed blockchain confirmation information is sharded, and the confirmation information is divided into multiple verification shards according to the hierarchical structure of the multi-level hash verification tree, and the consensus algorithm based on practical Byzantine fault tolerance is executed in parallel in different shards, and the blockchain confirmation information is synchronously updated in all storage nodes through cross-shard verification; After receiving the blockchain confirmation information, the storage node calculates the verification value layer by layer according to the multi-level hash verification tree, compares it with the verification field in the blockchain confirmation information, and generates a synchronization request including the inconsistent data block location and verification path; A data synchronization channel is established based on the verification path of the synchronization request, and a selective retransmission mechanism is used to transmit inconsistent data blocks. The receiving node verifies the correctness of the data block through the multi-level hash verification tree, and the verified synchronization target file is updated to the target end database to complete the parallel synchronization of the target unstructured file.
7. The method according to claim 6, characterized in that Based on the consistent hashing algorithm, the blockchain confirmation information is mapped to the confirmation node ring of the blockchain network, multiple nodes in the clockwise direction of the mapping position are selected as the confirmation node group, and the two-phase commit protocol is used within the confirmation node group to complete the pre-confirmation, including: Obtain the computing power value and storage capacity value of each confirmation node in the blockchain network, determine the number of virtual nodes to be allocated based on the computing power value and storage capacity value, use a hash algorithm to calculate the node identification hash value of the virtual node and map it to the hash ring, and build a confirmation node ring; Calculate the confirmation information hash value of the blockchain confirmation information based on the consistent hashing algorithm and map it to the confirmation node ring, obtain the network delay value of each confirmation node in the confirmation node ring and the historical confirmation performance value in the node scoring table, and select multiple confirmation nodes as the confirmation node group in a clockwise direction from the mapping position of the confirmation information hash value based on the network delay value and the historical confirmation performance value; The confirmation information hash value, timestamp and digital signature of the coordinating node are combined to generate a pre-confirmation request, and the node corresponding to the mapping position in the confirmation node group acts as a coordinating node to send the pre-confirmation request to the remaining confirmation nodes in the confirmation node group, and the confirmation nodes in the confirmation node group verify the legitimacy of the pre-confirmation request and return a preparation response; The coordination node counts the number of ready responses of the confirmation node group, generates a submission request when the number of ready responses exceeds a preset ratio, sends the submission request to the confirmation node group, and the confirmation nodes in the confirmation node group perform local confirmation and generate a confirmation log; A timeout detection mechanism is used to monitor the online status of the coordination node, and when a coordination node failure is detected, a new coordination node is selected from the confirmation node group, and confirmation authority is allocated to the new coordination node; The confirmation operation information in the confirmation log is constructed into a verification tree, the root node value of the verification tree is combined with the confirmation signature of the confirmation node group to generate a pre-confirmation result, and the pre-confirmation result is sent to the remaining confirmation nodes in the blockchain network. After receiving the pre-confirmation result, each remaining confirmation node verifies the number of confirmation signatures and the proof path of the verification tree. After the verification is passed, the pre-confirmation result is recorded as a checkpoint to complete the pre-confirmation of the blockchain confirmation information.
8. A parallel synchronization system for unstructured files, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to obtain the target unstructured file in the source database, and perform intelligent block processing through a pre-trained file block model, wherein the file block model adaptively determines the optimal block size and the number of blocks according to the content characteristics, data distribution characteristics and file structure characteristics of the target unstructured file, generates multiple data sub-blocks, generates a double-layer cascade index number containing block-level information and file-level information for each data sub-block, calculates a data hash value based on the entropy value characteristics of the data sub-block and the double-layer cascade index number, combines the data sub-block, the double-layer cascade index number and the data hash value to generate a data storage package, and distributes and stores it to multiple storage nodes in the blockchain network; The second unit is used to read the data storage package from each storage node in the blockchain network, build a data transmission priority model based on the double-layer cascade index number, and the data transmission priority model comprehensively evaluates the data importance of the data sub-block, the network transmission status and the load of the target database, dynamically allocates the transmission priority and the network resource quota, and transmits the data sub-block to the target database through the adaptive parallel transmission channel according to the transmission priority. For the data sub-block that has completed the transmission, an asynchronous verification mechanism is used to calculate the verification hash value based on the entropy value characteristics of the data sub-block, and the verification hash value is compared with the data hash value in the data storage package. When the comparison results are consistent, a data verification identifier is generated for the corresponding data sub-block; The third unit is used to obtain all data sub-blocks with data verification identifiers and their corresponding double-layer cascade index numbers, input the data sub-blocks into a pre-trained file reorganization model, the file reorganization model constructs a data block association map based on the file level information of the double-layer cascade index number, calculates the combination weights between the data sub-blocks according to the data block association map, hierarchically sorts the data sub-blocks based on the combination weights, and uses the content features of the data sub-blocks to perform boundary feature matching and sequence optimization to generate a reorganization sequence, splices and reorganizes the data sub-blocks according to the reorganization sequence to obtain a synchronization target file, generates blockchain confirmation information including reorganization sequence information, file integrity verification information and synchronization status information, stores the synchronization target file in a target end database, and submits the blockchain confirmation information to the blockchain network, synchronously updates the blockchain confirmation information in all storage nodes through the consensus mechanism of the blockchain network, and completes the parallel synchronization of the target unstructured file.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Big data disperse uplink method and system
CN112767110A
Data asset automatic generation and management system based on data operation technology
CN118941395A