An Information Storage Optimization Method Based on Blockchain Technology

By introducing off-chain storage optimization and consensus efficiency improvement methods into blockchain technology, the problems of low node storage capacity and low consensus efficiency are solved, and the system operation efficiency and transaction throughput are improved.

CN114691032BActive Publication Date: 2025-06-17NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210132340.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-14
Publication Date
2025-06-17
Estimated Expiration
2042-02-14

AI Technical Summary

Technical Problem

The existing blockchain technology has problems such as low node storage capacity and low consensus efficiency, resulting in insufficient system operation efficiency and transaction throughput.

Method used

Through off-chain storage optimization and reducing the frequency of view replacement, an information storage optimization method based on blockchain technology is proposed, including user registration and login, data index storage on the blockchain platform, off-chain cloud storage large amounts of data, setting data nodes in the blockchain network to improve consensus efficiency, adding contribution factors to the block header of blockchain nodes, etc.

Benefits of technology

The system's operating efficiency and transaction throughput are improved, and through off-chain storage optimization and consensus efficiency improvement, the problems of low node storage capacity and low consensus efficiency are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114691032B_ABST
    Figure CN114691032B_ABST
Patent Text Reader

Abstract

The present invention provides an information storage optimization method based on blockchain technology. In this method, a new data node is added to the blockchain network to store the information of the blockchain network and the status information of the nodes in the network. The structure of the block header of the blockchain node is changed, and fields such as contribution factor, malicious times, and response time are added, and corresponding weights are assigned to each field. High-quality consensus nodes are obtained through a reward and punishment function and weight calculation to reduce the frequency of view change. Then, due to the limited storage capacity of the block node, an off-chain storage scheme is proposed. The stored data is preprocessed, and then the text data and non-text data are stored separately to improve the storage efficiency of the system. Finally, the concept of dynamic lightning network is proposed, and resources are dynamically allocated according to the historical transaction situations between nodes. The present invention can improve the transaction throughput and operation efficiency of the system and reduce the latency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information storage optimization method based on blockchain technology and belongs to the technical field of data storage optimization. Background Art

[0002] With the development of information technology and network technology, online business handling has become increasingly common. However, traditional information management systems rely heavily on servers and are characterized by instability and lack of security. Therefore, new technical solutions are needed to address the above problems and drawbacks. Blockchain technology is essentially a distributed shared database for storage. Different from traditional relational databases, it has advantages such as decentralization and tamper-proofing. Applying blockchain technology to the system can enhance the security, stability, and immutability of data. However, existing blockchain technologies have the disadvantages of low node storage capacity and low consensus efficiency. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to overcome the problems of reaching the upper limit of blockchain node storage capacity and low consensus efficiency in the prior art, and provide an information storage optimization method based on blockchain technology, which improves the operation efficiency and transaction throughput of the system by means of off-chain storage optimization and reducing the frequency of view changes.

[0004] The present invention provides an information storage optimization method based on blockchain technology, including the following steps:

[0005] Step 1: The user registers and logs in to the system;

[0006] Step 2: After the user logs in to the system, for the data submitted to the system, the blockchain platform only stores the data index, and a large amount of data is stored in off-chain cloud storage;

[0007] Step 3: When the data submitted to the system is packaged and uploaded to the blockchain platform, consensus needs to be reached. The consensus nodes can be grouped by setting data nodes in the blockchain network to improve the consensus efficiency;

[0008] Step 4: By adding contribution factors, malicious nodes, and response time to the block header of the blockchain node and attaching corresponding weights, high-quality nodes can be distinguished to improve the consensus efficiency;

[0009] Step 5: For the file information data involved in the system, these data are first preprocessed; for the transaction information data involved in the system, it is transferred to Step 8 for processing;

[0010] Step 6: The preprocessed data is stored and processed. At the same time, when querying data, the high-frequency index module is accessed first. If the query fails, the storage module is accessed;

[0011] Step 7: When performing storage processing, divide the file data into text data and non-text data, and store the text data and non-text data separately. When storing non-text data, preprocess the non-text data small files.

[0012] Step 8: Propose the concept of a dynamic Lightning Network. The resource allocation of RSMC and HTLC is static. For the transaction information data involved in the system, use a configuration file to record the historical transaction records between two nodes to dynamically maintain RSMC and HTLC.

[0013] In the present invention, the user first performs system registration and login operations. After logging in to the system, upload data. After the nodes in the blockchain network reach a consensus, the data can be packaged to the blockchain platform Fisco. Since the storage capacity of the blockchain storage container is limited, off-chain cloud storage is performed for massive data. When querying the system data, first search in the high-frequency index module to improve the retrieval efficiency.

[0014] First of all, in the present invention, in order to improve the operation efficiency and throughput of the system, optimize the storage model of the information management system based on blockchain technology, and propose to add a data node in the blockchain network to store the information of the blockchain network and the status information of the nodes in the network. Secondly, in order to achieve privacy protection for the information and operation records in the system and prevent malicious nodes from obtaining and tampering with data. Then, due to the limited storage capacity of the block nodes, a off-chain storage scheme is proposed to preprocess the stored data, and then store the text data and non-text data separately to improve the storage efficiency of the system. Finally, propose the concept of a dynamic Lightning Network and dynamically allocate resources according to the historical transaction situation between nodes.

[0015] In addition, in order to reduce the frequency of view replacement and improve the operation efficiency of the system. Change the structure of the block header of the blockchain node, add fields such as contribution factor, malicious times, and corresponding time, and assign corresponding weights to each field. Calculate high-quality consensus nodes through the reward and punishment function and weights to reduce the frequency of view replacement. At the same time, classify the consensus nodes into consensus nodes, standby nodes, ordinary nodes, and verification nodes, and each type of node is responsible for different functions to improve the operation efficiency of the system.

[0016] The further optimized technical solution of the present invention is as follows:

[0017] In the said Step 1, the specific operation of user registration is as follows:

[0018] New users register through the user registration module and submit the registration information to the background administrator account for review. After passing the review, they can perform login operations to prevent malicious nodes from joining the blockchain network.

[0019] The user registration module of the present invention is mainly responsible for the registration of new users. When the user selects the type of registered account, the system needs to submit the registration information to the background administrator account for review. The review mechanism can prevent malicious nodes from joining the blockchain network.

[0020] The specific operations of step 4 are as follows:

[0021] Step 4.1, Contribution factor reward rule: After a consensus node of the blockchain generates a block, the value of the contribution factor of this node will increase. After a standby node completes a vote election, the activity value of this node will increase. Let n be the total number of view switches currently, An be the contribution factor value before the node obtains the reward, and A n+1 be the contribution factor value after the node obtains the reward, Consensusn be the total number of consensus times of the consensus node, MaxValue be the maximum value of the node contribution factor value, then

[0022]

[0023] Step 4.2, Contribution factor penalty rule: Whenever a consensus node of the blockchain times out and fails to respond, a penalty of reducing the contribution factor value will be executed. A′ n+1 is the contribution factor value after the node is penalized, A′n is the contribution factor value before obtaining the reward, Count is the number of times the node fails to respond within the specified time, In(t) is the penalty factor, and MinValue is the minimum value of the node contribution factor value. Then

[0024]

[0025] Step 4.3, The reward and penalty values of the malicious times and response time fields are between MinValue and MaxValue. The more malicious times, the closer it is to MinValue, and the shorter the response time, the closer the value is to MaxValue. The weight values of the malicious times and response time fields of the k-th node in the n + 1-th round are W n+1 ,k, and the corresponding current reward and penalty values are A n+1 ,k. j is the number of consensus nodes in the blockchain network. Finally, the final result Result of each round is calculated according to the reward and penalty values of each field and the corresponding weights. n+1 When the malicious times of the node exceed N times, the node will be changed to an ordinary node and will no longer participate in the consensus process. Finally, the final result of each round is calculated according to the weights of each field to reduce the view switching probability.

[0026]

[0027] The present invention sets up data nodes in a blockchain network and optimizes the block headers of blockchain nodes, adding fields such as contribution factors, malicious times, and response times, and assigning corresponding weights to each field. Through a reward and punishment mechanism and weight calculation, high-quality nodes are distinguished to reduce the probability of view replacement, avoid malicious nodes during the consensus process, and improve the operating efficiency and transaction throughput of the system.

[0028] In summary, the present invention uses a reward and punishment mechanism to avoid the appearance of malicious nodes, reduce the frequency of view switching, and improve the operating efficiency of the system. The improved reward and punishment function linearly decreases the reward value as the number of rewards increases, and can be applied to the optimization scheme.

[0029] The specific operation of step 5 is as follows:

[0030] The data is cleared using regular expressions and the built-in string library of Python, and at the same time, the data is tokenized.

[0031] The text data of the present invention is input by the user himself, so invalid information needs to be processed, including removing useless symbols and tokenizing Chinese. The rich text data type initially submitted by the user contains various web page tags and redundant punctuation marks, which will cause a burden on storage and trouble in text processing. Regular expressions and the built-in string library of Python can be used for clearing. Tokenizing the text data can better extract and utilize text features, and the jieba library of Python can be used for tokenizing text data.

[0032] In step 6, the high-frequency index module is accessed to query the data on the off-chain storage module. If the query is successful, the most recent query time and the total number of queries are updated. If the query fails, it is searched and updated in the distributed storage module, and the data with the fewest query times is preferentially replaced in the high-frequency index module; considering the different I / O input and output service rates between different storage devices in the distributed storage, different sizes of caches need to be dynamically allocated to different performance storage devices, so that the load size of the storage device matches its reading and writing performance, ultimately improving the operation and retrieval efficiency of the notarization system. The distributed storage module uses a random cache allocation load algorithm to allocate caches, and the steps of this algorithm are as follows:

[0033] T = P * T1+(1 - P) * T2 (4)

[0034] Wherein, T is the average service delay time of the storage device, P1, P2, P3 are the hit rates of the caches of device 1, device 2, and device 3 respectively, T 1,1 T 1,2 T 1,3 are the access cache delay times of device 1, device 2, and device 3 respectively, T2,1 T 2,2 T 2,3 are the access times of Device 1, Device 2, and Device 3 to the storage device respectively; Q represents the total area of random access, M represents the current storage size, and M k is the storage size of the kth device. Ignoring the cache latency time, substituting the formula P = M / Q into the formula gives a cache allocation scheme that makes the access latency times of the storage equal. Then

[0035]

[0036] The high-frequency index module of the present invention adopts a multi-level structure, which can improve the data search efficiency. The data on the off-chain storage module will add the fields of the most recent query time and the total number of query times in the high-frequency index module. If the query is successful, the most recent query time and the total number of query words are updated. If the query fails, it is searched in the distributed storage module, and the high-frequency index module is updated. In the high-frequency index module, the data with the fewest query times is preferentially replaced. If the number of query times is the same, the data with the oldest most recent query time is replaced. In view of the different I / O rates between different devices in the distributed storage, different sizes of caches need to be allocated to storage devices with different performances, so that the load of the storage device matches its performance and the operating efficiency of the system is improved. The above requirements can be achieved by adopting a random cache allocation load algorithm.

[0037] In summary, the present invention optimizes the off-chain storage model, introduces data preprocessing and high-frequency index mechanisms to accelerate the information retrieval efficiency of the system. By screening and classifying the information for storage and separately processing small files of unstructured data, reserving and merging the small files and then storing them in HDFS, the read and write efficiency of the system can be improved and the operating pressure on the memory can be reduced.

[0038] In step 7, the file data in the blockchain node is divided into text data and non-text data. Among them, the text data is structured data and the non-text data is unstructured data. The text data is stored in Hbase, and the non-text data is stored in HDFS; in the optimization scheme, when storing the non-text data, the non-text data file will be preprocessed first; because the sizes of files such as videos or recordings are different and most of them are within a few megabytes. In the Hadoop platform distributed system HDFS, processing a large number of small files will cause a sharp drop in performance, and the operating pressure on the DataNodes will become very large. The specific operations are as follows:

[0039] Step 7.1. Therefore, an improved model is proposed to address the existing problems, and the storage of small files is optimized. In the initial HDFS system architecture, an intermediate layer is added, which is specifically responsible for the processing of small files. The intermediate layer includes a small file reservation module and a small file merging module;

[0040] Step 7.2: When the system needs to process and store non-text data, first judge the size of the file. Determine whether it is a small file. If it is determined that the file size is greater than 1 MB, it is directly stored in the HDFS distributed system. If it is determined that the file size is less than 1 MB, it enters the intermediate layer module between the system and HDFS. In the intermediate layer module, first temporarily store the file obtained from the system in the small file reservation module. When the storage capacity in the small file reservation module reaches the upper limit, merge the data in the small file reservation module and store the merged file in the HDFS distributed storage.

[0041] Step 7.3: When a user needs to query non-text data files such as videos or recordings from the system, first query the size of the file to be searched in the configuration file. Determine whether the file to be searched is a small file. If the file to be searched is a large file, it is directly stored in HDFS. If the file to be searched is a small file smaller than 1 MB, first search in the small file reservation module in the intermediate layer. If the search fails in the small file reservation module, it means that the file has been merged and stored in HDFS, and the relevant file needs to be searched in the HDFS distributed storage system.

[0042] By optimizing the storage model solution, the present invention adds an intermediate layer between the system client and HDFS, which can effectively alleviate the problem of overloaded operation and excessive memory of the NameNode node and DataNodes nodes in the distributed storage file system caused by frequent reading and writing of a large number of small files in HDFS, thereby improving the retrieval and operation efficiency of the system.

[0043] In the said Step 8, the main idea of Segregated Witness is to extract a series of useless information such as witness data from the transaction information, which can reduce the storage burden of the block. The Segregated Witness technology promotes the development of the Lightning Network. The main idea of the Lightning Network is to conduct a large number of transactions outside the blockchain. The core concepts are RSMC (Revocable Sequential Maturity Contract) and HTLC (Hashed Timelocked Contract). The configuration file contains information about Node 1, Node 2, the recent transaction time, the recent transaction frequency, the recent transaction data volume, etc.

[0044] The present invention proposes an off-chain storage model for improving the information management system based on blockchain technology. Through the above steps, the innovative information storage optimization method based on blockchain technology in this patent can be basically realized. The main core content is to optimize the blockchain layer and the off-chain storage structure, improving the operation efficiency and throughput of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1This is the overall model framework diagram of the present invention.

[0046] Figure 2 The flowchart of the off-chain storage model in the present invention.

[0047] Figure 3 This is the distribution diagram of blockchain network nodes in the present invention.

[0048] Figure 4 This is the schematic diagram of the improved block header storage solution in the present invention.

[0049] Figure 5 This is the system distributed storage structure diagram in the present invention.

[0050] Figure 6 This is the file optimized storage diagram in the present invention. Detailed implementation manners

[0051] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings: This embodiment is implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0052] Embodiment 1

[0053] This embodiment provides an optimized method for storing notarization information based on blockchain technology. As Figure 1 and Figure 2 shown, it includes the following steps:

[0054] Step 1: User registration.

[0055] The user registration module is mainly responsible for the registration of new users. The user identities are divided into administrator accounts, ordinary users, notary institutions, and public security, procuratorial, and judicial institutions. Ordinary users need to register their accounts in a real-name system. Notary institutions need to provide business licenses when registering their accounts. Public security, procuratorial, and judicial institutions need to provide corresponding certificates when registering their accounts. The system needs to submit the registration information to the background administrator account for review. The notarization system with a review mechanism can prevent malicious nodes from joining the blockchain network.

[0056] Step 2: Upload the notarization information data to the blockchain platform.

[0057] Step 3: First, set data nodes in the blockchain network. The data nodes store the basic information of the blockchain network and the node status information in the blockchain network. The node status information includes information such as node type and node creation time. As Figure 3 shown.

[0058] Step 4: Add contribution factors, malicious node, response time and other fields to the block header of the blockchain node and attach corresponding weights. The reward and punishment mechanism is used to avoid the emergence of malicious nodes, reduce the frequency of view switching, and improve the operating efficiency of the system. The improved reward and punishment function decreases linearly with the increase of the number of rewards and can be applied to the optimization scheme. As Figure 4 shown.

[0059] Step 4.1: Contribution factor reward rule: After a consensus node completes the generation of a block, the contribution factor value of the node will increase. After a standby node completes a voting election, the activity value of the node will increase. An+1 is the contribution factor value of the node after obtaining the reward. Consensusn is the total number of consensus times of the consensus node, and MaxValue is the maximum value of the node contribution factor value.

[0060]

[0061] Step 4.2: Contribution factor punishment rule: Whenever a consensus node times out and fails to respond, a punishment for reducing the contribution factor value will be executed. An+1 is the contribution factor value of the node after punishment, An is the contribution factor value before obtaining the reward, Count is the number of times the node fails to respond within the specified time, and MinValue is the minimum value of the node contribution factor value.

[0062]

[0063] Step 4.3: The reward and punishment values of the malicious times and response time fields are between MinValue and MaxValue. The more malicious times, the closer it is to MinValue, and the shorter the response time, the closer the value is to MaxValue. When the malicious times of a node exceed N times, the node will be changed to an ordinary node and will no longer participate in the consensus process. Finally, the final result of each round is calculated according to the weights of each field, reducing the probability of view switching.

[0064]

[0065] Step 5: Preprocess the data uploaded to the storage module. The text data is input by the user himself, so invalid information needs to be processed, including removing useless symbols and Chinese word segmentation, etc. The rich text data type initially submitted by the user contains various web page tags and redundant punctuation marks, which will cause a burden on storage and trouble in text processing. Regular expressions and the built-in string library of Python can be used for clearing. Performing word segmentation on the text data can better extract and utilize text features, and Python and the jieba library can be used for word segmentation of text data.

[0066] Step 6: Then access the high-frequency index module. The high-frequency index module adopts a multi-level structure, which can improve the data search efficiency. The data on the off-chain storage module will add the fields of the most recent query time and the total number of query times in the high-frequency index module. If the query is successful, the most recent query time and the total number of query words are updated. If the query fails, it is searched in the distributed storage module and the high-frequency index module is updated. In the high-frequency index module, the data with the fewest query times is preferentially replaced. If the number of query times is the same, the data with the oldest most recent query time is replaced. Considering the different I / O rates between different devices in the distributed storage, different sizes of caches need to be allocated to storage devices with different performances, so that the load of the storage device matches its performance and the operating efficiency of the system is improved. The above requirements can be achieved by adopting the random cache allocation load algorithm. The algorithm steps are as follows:

[0067] T = P * T1+(1 - P) * T2 (4)

[0068] Where T is the average latency time for the storage device to provide services, P is the cache hit rate, T1 is the latency time for accessing the cache, and T2 is the time for accessing the storage device. Q represents the total area of random access, and M represents the current storage size. Ignoring the cache latency time, substituting the formula P = M / Q into the formula gives the cache allocation scheme, making the access latency times of the storage equal.

[0069]

[0070] Step 7: Divide the file data in the blockchain node into text data and non-text data. The text data is structured data. The non-text data is unstructured data. The text data is stored in Hbase, and the non-text data is stored in HDFS. As Figure 5 shown.

[0071] In the optimization plan, when storing non-text data, the non-text data files are first preprocessed. Since the sizes of notarized video or audio files vary and most are within a few megabytes. In the Hadoop platform distributed system HDFS, processing a large number of small files will cause a sharp decline in performance, and the operating pressure on the DataNodes nodes will become very large.

[0072] Step 8.1: Therefore, an improved model is proposed for the existing problems to optimize the storage of small files. In the initial HDFS system architecture, an intermediate layer is added, which is specifically responsible for the processing of small files. The storage optimization scheme model is as Figure 6 shown. The intermediate layer includes a small file reservation module and a small file merging module in total.

[0073] Step 8.2: When the system needs to process and store non-text data, first, it determines the size of the file to check if it is a small file. If it is determined that the file size is greater than 1 MB, it is directly stored in the HDFS distributed system. If it is determined that the file size is less than 1 MB, it enters the intermediate layer module between the system and HDFS. In the intermediate layer module, the file obtained from the notarization system is temporarily stored in the small file reservation module. When the storage capacity in the small file reservation module reaches the upper limit, the data in the small file reservation module is merged, and the merged file is stored in the HDFS distributed storage.

[0074] Step 8.3: When a user wants to query non-text data files such as videos or recordings from the notarization system, first, it queries the size of the file to be searched in the configuration file to determine if the file to be searched is a small file. If the file to be searched is a large file, it is directly stored in HDFS. If the file to be searched is a small file smaller than 1 MB, it first searches in the small file reservation module in the intermediate layer. If the search fails in the small file reservation module, it means the file has been merged and stored in HDFS, and the relevant file needs to be searched in the HDFS distributed storage system.

[0075] By optimizing the storage model solution and adding an intermediate layer between the system client and HDFS, it can effectively alleviate the problem of overloading the memory of the NameNode and DataNodes in the distributed storage file system due to frequent reading and writing of a large number of small files in HDFS, thereby improving the retrieval and operation efficiency of the system. As Figure 6 shown.

[0076] Step 8: The main idea of Segregated Witness is to extract a series of useless information such as witness data from transaction information, which can reduce the storage burden of the block. The Segregated Witness technology promotes the development of the Lightning Network. The main idea of the Lightning Network is to conduct a large number of transactions outside the blockchain. The core concepts are RSMC (Revocable Sequential Maturity Contract) and HTLC (Hashed Time-Locked Contract).

[0077] The concept of a dynamic Lightning Network is proposed. The resource allocation of RSMC and HTLC is static. The historical transaction records between two nodes can be recorded in a configuration file to dynamically maintain RSMC and HTLC. The configuration file contains information about Node 1, Node 2, the most recent transaction time, the most recent transaction frequency, the most recent transaction data volume, etc.

[0078] In summary, an off-chain storage model for improving the notarization information management system based on blockchain technology is proposed. Through the above steps, the innovative method for optimizing the storage of notarization information based on blockchain technology in this patent can be basically realized. The main core content is to layer the blockchain and optimize the off-chain storage structure to improve the operation efficiency and throughput of the system.

[0079] As described above, it is only the specific implementation manner in the present invention, but the protection scope of the present invention is not limited thereto. Any transformation or replacement that can be understood and conceived by those familiar with the technology within the technical scope disclosed by the present invention should be covered within the scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An information storage optimization method based on blockchain technology, characterized in that, Including the following steps: Step 1: The user registers and logs in to the system; Step 2: After the user logs in to the system, for the data submitted to the system, the blockchain platform only stores the data index, and a large amount of data is stored off-chain in the cloud; Step 3: When packing and uploading the data submitted to the system to the blockchain platform, consensus needs to be reached. The consensus nodes are grouped by setting data nodes in the blockchain network to improve the consensus efficiency; Step 4: Add contribution factors, malicious nodes and response time to the block header of the blockchain node, and distinguish high-quality nodes by attaching corresponding weights to improve the consensus efficiency; The specific operations are as follows: Step 4.1, Contribution Factor Reward Rule: After a consensus node of the blockchain completes the generation of a block, the value of the contribution factor of this node will increase. After a standby node completes a vote election, the activity value of this node will increase. Let n be the total number of view switches currently, A n be the contribution factor value before the node obtains the reward. Let A n+1 be the contribution factor value after the node obtains the reward. Consensus n be the total number of consensus of the consensus nodes, MaxValue be the maximum value of the node contribution factor value, then Step 4.2, Contribution Factor Penalty Rule: Whenever a consensus node of the blockchain times out and fails to respond, a penalty of reducing the contribution factor value will be executed, A′ n+1 is the contribution factor value after node penalty, A′n is the contribution factor value before obtaining the reward, Count is the number of times the node fails to respond within the specified time, In(t) is the penalty factor, MinValue is the minimum value of the node contribution factor value, then Step 4.3: The reward and punishment values of the malicious count and response time fields are between MinValue and MaxValue. The more malicious counts, the closer to MinValue, and the shorter the response time, the closer to MaxValue. The weight values of the malicious count and response time fields of the k-th node in the (n + 1)-th round are W n+1,k , and the corresponding current reward and punishment values are A n+1,k , where j is the number of consensus nodes in the blockchain network; finally, the final result Result of each round is calculated according to the reward and punishment values and corresponding weights of each field n+1 ; when the malicious count of a node exceeds N times, the node is changed to an ordinary node and no longer participates in the consensus process. Finally, the final result of each round is calculated according to the weights of each field to reduce the probability of view switching Step 5: For the file information data involved in the system, first preprocess this data; for the transaction information data involved in the system, go to Step 8 for processing; Step 6: Store and process the preprocessed data. When querying data, first access the high-frequency index module. If the query fails, access the storage module; Step 7: When performing storage processing, divide the file data into text data and non-text data, store the text data and non-text data separately. When storing non-text data, preprocess the non-text data small files; Step 8: For the transaction information data involved in the system, use a configuration file to record the historical transaction records between two nodes to dynamically maintain RSMC and HTLC.

2. The information storage optimization method based on blockchain technology according to claim 1, characterized in that, In the above Step 1, the specific operations for user registration are as follows: New users register through the user registration module and submit the registration information to the background administrator account for review to prevent malicious nodes from joining the blockchain network.

3. The information storage optimization method based on blockchain technology according to claim 1, characterized in that, The specific operations of the above Step 5 are as follows: Use regular expressions and the built-in string library of Python to clean the data, and at the same time perform word segmentation on the data.

4. The information storage optimization method based on blockchain technology according to claim 1, characterized in that, In the above Step 6, access the high-frequency index module to query the data on the off-chain cloud storage in the high-frequency index module. If the query is successful, update the recent query time and the total query times. If the query fails, search and update the high-frequency index module in the distributed storage module. The data with the least query times is preferentially replaced in the high-frequency index module; Considering the different I / O input and output service rates between different storage devices in the distributed storage, it is necessary to dynamically allocate different sizes of caches to different performance storage devices so that the load size of the storage device matches its read and write performance, and finally improve the operation and retrieval efficiency of the notarization system; The distributed storage module uses a random cache allocation load algorithm to allocate caches. The steps of this algorithm are as follows: T = P * T1 + (1-P)* T2 (4) Among them, T is the average latency of the storage device to provide services, and P1, P2, and P3 are the hit rates of the caches of device 1, device 2, and device 3 respectively, T 1,1 T 1,2 T 1,3 are the latency times for device 1, device 2, and device 3 to access the cache respectively, T 2,1 T 2,2 T 2,3 are the times for device 1, device 2, and device 3 to access the storage device respectively; M k is the storage size of the k-th device; Q represents the total area of random access, M represents the current storage size. Ignoring the cache latency time, substituting the formula P = M / Q into the formula gives the cache allocation scheme, making the access latency times of the storage equal.

5. The information storage optimization method based on blockchain technology according to claim 1, characterized in that, In the above Step 7, divide the file data in the blockchain node into text data and non-text data. Among them, the text data is structured data, and the non-text data is unstructured data. The text data is stored in Hbase, and the non-text data is stored in HDFS; In the optimization plan, when storing non-text data, first preprocess the non-text data files; The specific operations are as follows: Step 7.1: An improved model is proposed for the existing problems, and the storage of small files is optimized. In the initial HDFS system architecture, an intermediate layer is added, which is specifically responsible for the processing of small files. The intermediate layer includes a small file reservation module and a small file merging module in total; Step 7.2: When the system needs to process and store non-text data, first, the size of the file is judged to determine whether it is a small file. If it is determined that the size of the file is greater than 1MB, it is directly stored in the HDFS distributed system. If it is determined that the size of the file is less than 1MB, it enters the intermediate layer module between the system and HDFS; In the intermediate layer module, the file obtained from the system is temporarily stored in the small file reservation module first; When the storage capacity in the small file reservation module reaches the upper limit, the data in the small file reservation module is merged, and the merged file is stored in the HDFS distributed storage; Step 7.3: When the user needs to query non-text data files such as videos or recordings from the system, first, the size of the file to be searched is queried in the configuration file to determine whether the file to be searched is a small file. If the file to be searched is a large file, it is directly stored in HDFS. If the file to be searched is a small file smaller than 1MB, it is first searched in the small file reservation module of the intermediate layer. If the search fails in the small file reservation module, it means that the file has been merged and stored in HDFS, and the relevant file needs to be searched in the HDFS distributed storage system.

6. The information storage optimization method based on blockchain technology according to claim 1, characterized in that, In step 8, the configuration file contains information about Node 1, information about Node 2, the most recent transaction time, the most recent transaction frequency, and the most recent transaction data volume.

Citation Information

Patent Citations

  • Digital evidence storage platform and evidence storage method based on block chain

    CN110912937A

  • Blockchain consensus optimization method based on ring signature and aggregation signature

    CN112003820A