A blockchain-based method and apparatus for storing experimental data in civil engineering.

By adjusting the Huffman coding table and optimizing the coding length according to the importance and access frequency of civil engineering test data, the problem that traditional Huffman coding cannot simultaneously meet the requirements of data compression and access speed is solved, thus achieving efficient data storage and fast access.

CN120216726BActive Publication Date: 2026-05-26河南省栾卢高速公路建设有限公司 +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
河南省栾卢高速公路建设有限公司
Filing Date
2025-03-10
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Traditional Huffman coding cannot simultaneously meet the requirements of data compression and access speed in civil engineering data storage, and is not suitable for the data storage compression needs of the current scenario.

Method used

By determining the importance, access frequency, and repetition rate of character data in civil engineering test data, the Huffman coding table is adjusted to prioritize important and frequently accessed data, assigning them short codes, while less important data is assigned long codes, thereby achieving efficient data compression and access.

Benefits of technology

It achieves improved access speed while compressing data, ensuring fast reading of high-volume data and high compression of data with high repetition rate, thereby enhancing data security and transparency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216726B_ABST
    Figure CN120216726B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data processing technology, specifically to a blockchain-based method and apparatus for storing civil engineering test data. The method includes: acquiring civil engineering test data stored on the blockchain and real-time collected civil engineering test data, and preprocessing it to obtain analysis data; determining the importance, access frequency, and final repetition rate of character data based on the analysis data, and obtaining the encoding priority of the character data, adjusting the Huffman coding table; compressing all civil engineering test data according to the adjusted Huffman coding table, and re-storing it in the blockchain; and comprehensively adjusting the Huffman coding table based on the obtained encoding priority of each character data, compressing the data, and re-storing it in the blockchain. This method reduces the amount of data while maintaining access speed, thereby improving the reading speed for high-access data during subsequent data access.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to a blockchain-based method and apparatus for storing experimental data in civil engineering. Background Technology

[0002] As the requirements for data accuracy and security in the field of civil engineering increase, traditional experimental data storage and management methods are facing many challenges, including data tampering, leakage, and lack of traceability. To solve these problems, blockchain technology is introduced. As a decentralized and tamper-proof distributed ledger technology, blockchain technology provides a way to enhance data security and transparency, and has significant advantages for the storage and management of civil engineering experimental data.

[0003] Before storing civil engineering data on the blockchain, data compression is performed to reduce storage requirements and transaction costs. Generally, Huffman coding is used for lossless compression, which reduces the amount of data while ensuring data integrity. However, traditional Huffman coding defines the code length based solely on data frequency. In civil engineering data, data collected at different times has different importance, and there are different access frequencies when accessing this data later. Furthermore, the code length also affects the access speed. It cannot simultaneously satisfy both data compression and access speed, making traditional Huffman coding not entirely suitable for the data storage and compression needs of the current scenario. Summary of the Invention

[0004] To address the technical problem that existing Huffman coding methods cannot simultaneously satisfy data compression and access speed, the present invention aims to provide a blockchain-based method for storing experimental data in civil engineering. The specific technical solution adopted is as follows:

[0005] The system acquires civil engineering test data stored on the blockchain and real-time collected civil engineering test data, and preprocesses it to obtain analysis data, which includes character data.

[0006] Based on the analysis data, the importance, access frequency and final repetition rate of character data are determined in sequence, and the encoding priority of character data is obtained, and the encoding table of Huffman coding is adjusted.

[0007] Based on the adjusted Huffman coding table, all civil engineering test data are compressed and re-stored in the blockchain.

[0008] Furthermore, civil engineering test data stored on the blockchain and real-time collected civil engineering test data are acquired and preprocessed to obtain analytical data, including:

[0009] The first dataset is formed by collecting civil engineering test data in real time through testing equipment and storing the first dataset on the blockchain. The second dataset is formed by extracting the civil engineering test data stored on the blockchain and the first dataset.

[0010] The second dataset was filtered to obtain the analysis data.

[0011] Furthermore, the importance of character data is determined based on the analyzed data, including:

[0012] Based on the analysis data, the collection information of civil engineering test data in the same area at different time periods is obtained. The collection information includes the number of civil engineering test data and the acquisition time of civil engineering test data.

[0013] The number of civil engineering test data for each region is obtained by numbering the civil engineering test data, and the time interval is obtained by subtracting the acquisition time of adjacent civil engineering test data.

[0014] The formula for determining the importance of a region is as follows:

[0015]

[0016] Among them, Q i The value of region i is represented by Δt; N represents the number of civil engineering test data for region i; Δt n This represents the time interval between the nth civil engineering test data collection and the subsequent collection; and the importance of the current region is the importance of the character data in the current region.

[0017] Furthermore, the access frequency of character data is determined based on the analyzed data, including:

[0018] Based on the analysis data, the number of accesses for each analysis data session and the total number of accesses for all analysis data sessions are calculated.

[0019] The formula for determining the access frequency of the analyzed data is as follows:

[0020]

[0021] Among them, W j Indicates the access frequency of the j-th data analysis; m j This represents the number of times the data was accessed in the j-th analysis; M represents the total number of accesses to all analyzed data; and the access frequency of the current analyzed data is the access frequency of its corresponding character data.

[0022] Furthermore, the final repetition rate of the character data is determined based on the analyzed data, including:

[0023] Based on the analysis data, determine the repetition rate of character data in each analysis data, and determine the first repetition rate of character data in the same area;

[0024] The overall repetition rate is obtained from all character data, and the final repetition rate is obtained by correcting the first repetition rate with the overall repetition rate.

[0025] Furthermore, based on the analyzed data, the repetition rate of character data in each analysis is determined, and the first repetition rate of character data within the current same region is determined, including:

[0026] The formula for calculating the repetition rate of character data in each analysis is as follows:

[0027]

[0028] Among them, C j,k x represents the repetition rate of the k-th character in the j-th analysis data; k X represents the number of occurrences of the k-th character in the current analysis data; j This represents the total number of character data points in the j-th analysis.

[0029] The formula for determining the first repetition rate of character data within the same current region is:

[0030]

[0031] Among them, C1 k This represents the first repetition rate of the k-th character data within the current region; σ represents the mean repetition rate of the k-th character data within the current region; σ represents the standard deviation; F represents the number of times the k-th character data contains the same character data within its region; C j,k This represents the repetition rate of the k-th character in the j-th analysis data within the same region; norm() represents the normalization function.

[0032] Furthermore, the overall repetition rate is obtained based on all character data. The final repetition rate is obtained by correcting the first repetition rate using the overall repetition rate, including:

[0033] The formula for calculating the overall repetition rate based on all character data is as follows:

[0034]

[0035] Where, p k This represents the overall repetition rate of the k-th character data; l k This represents the number of times the k-th character appears in all regions; L represents the total number of characters.

[0036] The formula for calculating the final repetition rate by correcting the first repetition rate based on the overall repetition rate is as follows:

[0037]

[0038] Among them, C2 k This represents the final repetition rate of the k-th character data; This represents the average of the first repetition rates of the k-th character data across different regions; p k This represents the overall repetition rate of the k-th character data.

[0039] Furthermore, based on the determined importance, access frequency, and final repetition rate of the character data, the encoding priority of the character data is obtained, and the corresponding calculation formula is as follows:

[0040] Y k =Q1′ k *W1 k +(1-Q1′ k )*C2 k

[0041] Q1′ k =norm(Q1) k )

[0042] Among them, Y k Indicates the encoding priority of the k-th character data; Q1 k W1 indicates the importance of the k-th character data; k Indicates the access frequency of the k-th character data; C2 k Q1′ represents the final repetition rate of the k-th character data. k Indicates Q1 k The normalized value; norm() represents the normalization function.

[0043] Furthermore, the Huffman coding table is adjusted, including: determining the coding priority of each character data in the analysis data according to the coding priority of the character data, sorting them from high to low, assigning low coding lengths to character data with high coding priority, until all character data is encoded to obtain the Huffman coding table.

[0044] To solve the above-mentioned technical problems, the present invention provides another technical solution as follows: a blockchain civil engineering test data storage device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of a blockchain civil engineering test data storage method as described in any of the preceding claims.

[0045] The present invention has the following beneficial effects:

[0046] 1. By determining the importance, access frequency, and repetition rate of character data in civil engineering test data, and comprehensively determining the encoding priority of each character data, the Huffman coding table is adjusted, the data is compressed and stored in the blockchain; this reduces the amount of data while ensuring its access speed, thereby improving the reading speed of high-access data in subsequent data reading and access processes, while also ensuring a high degree of compression for data with high repetition rate.

[0047] 2. The blockchain civil engineering test data storage device provided by this invention has the same beneficial effects as the blockchain civil engineering test data storage method provided by this invention, and will not be described in detail here. Attached Figure Description

[0048] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is a flowchart illustrating the steps of a blockchain-based civil engineering experimental data storage method according to an embodiment of the present invention. Detailed Implementation

[0050] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a blockchain-based civil engineering experimental data storage method and apparatus proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0052] The specific scheme of the blockchain-based civil engineering experimental data storage method provided by the present invention will be described in detail below with reference to the accompanying drawings.

[0053] The lossless compression of civil engineering data using Huffman coding, which defines the coding length solely based on data frequency, affects access speed and cannot simultaneously satisfy both data compression and access speed. One embodiment of this invention provides a blockchain-based civil engineering test data storage method. This method determines the importance, access frequency, and repetition rate of character data in the civil engineering test data, comprehensively determines the coding priority of each character data, adjusts the Huffman coding table, compresses the data, and stores it in the blockchain. To implement this blockchain-based civil engineering test data storage method, a blockchain-based civil engineering test data storage device is provided. This device is essentially a software system composed of various devices that implement corresponding functions. The specific steps of this method are described in detail below.

[0054] Please see Figure 1 The diagram illustrates a flowchart of a blockchain-based civil engineering experimental data storage method according to an embodiment of the present invention, the method comprising:

[0055] Step S1: Obtain civil engineering test data stored on the blockchain and real-time collected civil engineering test data, and preprocess them to obtain analysis data, which includes character data;

[0056] Step S2: Based on the analysis data, determine the importance, access frequency and final repetition rate of the character data in sequence, obtain the encoding priority of the character data, and adjust the encoding table of Huffman coding;

[0057] Step S3: Based on the adjusted Huffman coding table, compress all civil engineering test data and re-store them in the blockchain.

[0058] The explanation is that blockchain, with its decentralized and tamper-proof characteristics, constructs a distributed ledger system. This system not only enhances data security and transparency but also provides significant advantages for the storage and management of civil engineering test data. Through blockchain technology, civil engineering test data can be recorded and verified in a decentralized network, ensuring the authenticity and integrity of the data. This allows stakeholders to trust and rely on this data more, thereby improving the efficiency and reliability of the entire engineering project.

[0059] Civil engineering test data refers to various types of data collected and recorded in experiments and tests conducted in the field of civil engineering. These data typically include detailed information on multiple aspects such as material performance testing, structural bearing capacity testing, foundation stability analysis, concrete strength testing, and soil mechanical property determination. By analyzing and processing this test data, engineers can evaluate the performance of materials and structures, ensure that the quality of the project meets the standard requirements, and thus provide a scientific basis for the design, construction, and maintenance of civil engineering projects.

[0060] Huffman coding is a widely used encoding method in data compression. Its core idea is to assign different length codes to each character based on its frequency of occurrence in the text to be encoded. High-frequency characters use shorter codes, and low-frequency characters use longer codes. This variable-length encoding effectively reduces the overall code length, achieving data compression. Specifically, Huffman coding first counts the frequency of each character in the text to be encoded. Then, it constructs a Huffman tree based on these frequencies. In this tree, each leaf node represents a character, and each non-leaf node represents a merged frequency value. The tree construction process starts with the two nodes with the lowest frequencies, merging them into a new node whose frequency is the sum of the frequencies of the two nodes. This process continues until only one node remains. Finally, the code for each character is the path from the root node to the leaf node of that character, with the left branch representing 0 and the right branch representing 1. Huffman coding can significantly reduce the storage space required for data, especially when processing large amounts of repetitive data. It also features lossless compression, meaning that the original data can be completely recovered during decompression without any information loss.

[0061] Preferably, step S1 includes:

[0062] The first dataset is formed by collecting civil engineering test data in real time through testing equipment and storing the first dataset on the blockchain. The second dataset is formed by extracting the civil engineering test data stored on the blockchain and the first dataset.

[0063] The second dataset was filtered to obtain the analysis data.

[0064] The testing equipment includes strain gauges, displacement gauges, etc., which can effectively acquire the new civil engineering test data needed at present to form the first dataset. Among them, a strain gauge is an instrument used to measure the minute deformation of a material or structure when it is subjected to force; a displacement gauge is used to measure the positional change of an object in space. Through these high-precision testing devices, engineers and researchers can acquire data in real time to ensure the safe operation of equipment and structures.

[0065] It can be explained that the first dataset is stored on the blockchain, and the civil engineering test data stored on the blockchain and the first dataset are extracted to form the second dataset. That is, the second dataset is formed by extracting data from the first dataset and the civil engineering test data originally stored on the blockchain.

[0066] In this embodiment, the second dataset is filtered to obtain analysis data, in order to select character data that is beneficial for Huffman coding and is related to civil engineering.

[0067] Understandably, civil engineering test data is typically collected from different regions, and multiple collections with different focuses are conducted within the same region. For key regions, data collection with different focuses is carried out gradually. By analyzing the frequency of user access to this data, we can understand which data is frequently used during the later stages of project implementation. Access behavior usually includes discussion, research, and analysis of the data to reflect its actual application value. Through access frequency, we can infer the importance of this data in actual engineering applications. For data with high access frequency, its reading rate needs to be increased accordingly. Then, based on the number of times the data appears repeatedly, a trade-off analysis can determine the encoding priority of character data to adjust the Huffman coding accordingly.

[0068] Preferably, step S2, which determines the importance of character data based on the analyzed data, includes:

[0069] Based on the analysis data, the collection information of civil engineering test data in the same area at different time periods is obtained. The collection information includes the number of civil engineering test data and the acquisition time of civil engineering test data.

[0070] The number of civil engineering test data for each region is obtained by numbering the civil engineering test data, and the time interval is obtained by subtracting the acquisition time of adjacent civil engineering test data.

[0071] The formula for determining the importance of a region is as follows:

[0072]

[0073] Among them, Q i The value of region i is represented by Δt; N represents the number of civil engineering test data for region i; Δt n This represents the time interval between the nth civil engineering test data collection and the subsequent collection; and the importance of the current region is the importance of the character data in the current region.

[0074] The definition of importance refers to the value and influence of character data in civil engineering test data, reflecting which character data is important to the results and which is relatively minor, thus facilitating the evaluation and optimization of character data.

[0075] Understandably, when conducting civil engineering test data for a specific area, the geological conditions of that area are often highly complex, potentially involving critical infrastructure or important building projects. To ensure the accuracy and reliability of the data, frequent testing and data collection are usually required. This involves collecting data with different focuses and conducting long-term monitoring of the current area. Based on these circumstances, the civil engineering test data for the current area is analyzed to determine the importance of the corresponding data, identifying which data are critical and which are secondary. Critical data refers to key factors affecting the safety and stability of the entire project, such as soil bearing capacity and groundwater level changes, while secondary data may include auxiliary information such as ambient temperature and humidity.

[0076] It should be noted that the number of times civil engineering test data is collected refers to each actual collection, regardless of time and region. For example, if a test was conducted in the same area this morning and afternoon, that would be two separate sets of civil engineering test data. In other words, each set of civil engineering test data is collected from the same area.

[0077] Specifically, based on the analysis data, the collection information of civil engineering test data at different time periods within the same region is obtained. The collection information includes the number of the civil engineering test data and the acquisition time of the civil engineering test data. The number of the civil engineering test data refers to the numbering of each collection of civil engineering test data in the analysis data as 1, 2, 3, etc. Through these codes, the number of civil engineering test data in the same region can be obtained, and similarly, the number of civil engineering test data in all regions can be obtained. The acquisition time of the civil engineering test data refers to the acquisition time of each collection of civil engineering test data in the analysis data. The time interval is obtained by subtracting the acquisition time of adjacent collections of civil engineering test data. In this embodiment, a whole day is used as the unit of time. Based on the above corresponding data, the importance of the current region is determined, and the importance of each region can be obtained similarly.

[0078] Based on the above, the importance of the current region is taken as the importance of the total character data in the current region. It can be noted that the importance of character data may vary between different analysis data due to different regions. In this case, the average value of the character data collected in multiple collection processes is taken as the importance of the character data to avoid data errors.

[0079] Understandably, analyzing data to determine access frequency indicates that the relevant data is involved in key engineering decisions, quality control, or safety assessments in actual operation, suggesting that the data has a high level of attention; while determining the repetition rate is to determine the repeatability of civil engineering test data.

[0080] Preferably, step S2, which determines the access frequency of character data based on the analyzed data, includes:

[0081] Based on the analysis data, the number of accesses for each analysis data session and the total number of accesses for all analysis data sessions are calculated.

[0082] The formula for determining the access frequency of the analyzed data is as follows:

[0083]

[0084] Among them, W j Indicates the access frequency of the j-th data analysis; m j This represents the number of times the data was accessed in the j-th analysis; M represents the total number of accesses to all analyzed data; and the access frequency of the current analyzed data is the access frequency of its corresponding character data.

[0085] Understandably, while the access frequency of each analysis data on the blockchain can be determined, the character data within it is difficult to define. Therefore, the access frequency of the accessed analysis data is obtained and used as the access frequency of the character data. It can be noted that there is a large amount of character data in the analysis data, and the access frequency of a single character data may differ in different analysis data. In this case, the average access frequency of the character data is calculated across multiple analysis data and used as the access frequency of the character data.

[0086] Preferably, in step S2, determining the final repetition rate of character data based on the analyzed data includes:

[0087] Step S21: Determine the repetition rate of character data in each analysis based on the analysis data, and determine the first repetition rate of character data in the same current region;

[0088] Step S22: Obtain the overall repetition rate based on all character data, and then correct the first repetition rate to obtain the final repetition rate.

[0089] Understandably, the calculation of the repetition rate of character data may be subject to chance, so a comprehensive analysis is required.

[0090] Preferably, step S21 includes:

[0091] The formula for calculating the repetition rate of character data in each analysis is as follows:

[0092]

[0093] Among them, C j,k x represents the repetition rate of the k-th character in the j-th analysis data; k X represents the number of occurrences of the k-th character in the current analysis data;j This represents the total number of character data points in the j-th analysis.

[0094] The formula for determining the first repetition rate of character data within the same current region is:

[0095]

[0096] Among them, C1 k This represents the first repetition rate of the k-th character data within the current region; σ represents the mean repetition rate of the k-th character data within the current region; σ represents the standard deviation; F represents the number of times the k-th character data contains the same character data within its region; C j,k This represents the repetition rate of the k-th character in the j-th analysis data within the same region; norm() represents the normalization function.

[0097] Understandably, for all the analyzed data, the repetition rate determined by the above steps is determined by combining the data from multiple analyses within the same region. Due to the different amounts of data collected each time, the repetition rates of similar characters in the above results may appear close, but if the two are directly compared with all the data, they may show different character repetition rates. For example, for character data 'a' and character data 'b', the first repetition rate of the character data determined based on their respective regions may both be close to 0.2, which seems very close. However, this closeness is based on the analysis within their respective regions, not on the analysis of all civil engineering test data. In this case, it only plays a predictive role.

[0098] Therefore, the first repetition rate determined above is used as the feature weight. The overall repetition rate is determined by comparing each character data with the overall analysis data, and the first repetition rate is corrected by the overall repetition rate to obtain the final repetition rate of the character data.

[0099] Preferably, step S21 includes:

[0100] The formula for calculating the overall repetition rate based on all character data is as follows:

[0101]

[0102] Where, p k This represents the overall repetition rate of the k-th character data; l k This represents the number of times the k-th character appears in all regions; L represents the total number of characters.

[0103] The formula for calculating the final repetition rate by correcting the first repetition rate based on the overall repetition rate is as follows:

[0104]

[0105] Among them, C2 k This represents the final repetition rate of the k-th character data; This represents the average of the first repetition rates of the k-th character data across different regions; p k This represents the overall repetition rate of the k-th character data.

[0106] It can be explained that, The average value represents the first repetition rate of the k-th character data across different regions. By calculating the average value, the complex and potentially scattered data points of the character data across different regions can be condensed into a single value, making it easier to understand and analyze the overall characteristics of the data. Specifically, the average value can eliminate the influence of extreme values ​​or random fluctuations of individual data points, simplify the data, and thus grasp the overall trend of the data.

[0107] Understandably, based on the above steps, the importance, access frequency, and repetition rate of character data are determined one by one. The importance of character data is obtained through the importance of character data. When the importance of character data is high, it is necessary to reduce access latency. At this time, the access frequency is given more importance, and the encoding should be short to facilitate the retrieval and decoding speed of civil engineering test data. When the importance of character data is low, it is compressed and stored based on the final repetition rate of character data, thereby saving storage space.

[0108] Preferably, in step S2, the encoding priority of the character data is obtained based on the determined importance, access frequency, and final repetition rate of the character data, and the corresponding calculation formula is as follows:

[0109] Y k =Q1′ k *W1 k +(1-Q1′ k )*C2 k

[0110] Q1′ k =norm(Q1) k )

[0111] Among them, Y k Indicates the encoding priority of the k-th character data; Q1 k W1 indicates the importance of the k-th character data; k Indicates the access frequency of the k-th character data; C2 k Q1′ represents the final repetition rate of the k-th character data. k Indicates Q1 k The normalized value; norm() represents the normalization function.

[0112] To clarify, encoding priority refers to the ordering of character data during encoding, which helps to allocate development resources reasonably and enable the most critical functions to be implemented first; that is, in the encoding process, some character data has a higher encoding priority and needs to be encoded first.

[0113] Preferably, in step S2, adjusting the Huffman coding table includes:

[0114] The encoding priority of each character in the analysis data is determined based on the encoding priority of the character data, and then sorted from high to low. Character data with high encoding priority is assigned a low encoding length until all character data is encoded to obtain the Huffman encoding table.

[0115] It can be explained that the encoding priority of each character in all analyzed data is determined according to the calculation formula; then all character data are sorted from high to low encoding priority to ensure that characters with higher encoding priority are processed first during the encoding process, thereby improving encoding efficiency and compression ratio; after sorting, characters with high encoding priority are assigned shorter encoding lengths, while characters with low encoding priority are assigned longer encoding lengths; in this way, it is ensured that the encoding length of each character is inversely proportional to the calculated value of the encoding priority, that is, high-frequency characters use shorter codes, and low-frequency characters use longer codes, so as to achieve optimal encoding efficiency overall, reduce the total length after encoding, and achieve effective data compression; after all character data is encoded, a complete Huffman coding table is obtained, which records in detail the code corresponding to each character and its encoding length.

[0116] As explained, in step S3, all civil engineering test data are compressed and re-stored into the blockchain according to the adjusted Huffman coding table.

[0117] By using an adjusted Huffman coding table, all civil engineering test data are compressed efficiently and without loss, saving significant resources during data transmission and storage. The data is then re-stored in the blockchain to ensure the security and transparency of the civil engineering test data. This not only improves the access speed to high-frequency civil engineering test data but also reduces the overall data volume.

[0118] Understandably, by determining the importance, access frequency, and repetition rate of character data in civil engineering test data, and comprehensively determining the encoding priority of each character data, the Huffman coding table is adjusted, the data is compressed and stored in the blockchain; this reduces the amount of data while ensuring its access speed, thereby improving the reading speed of high-access data in subsequent data reading and access processes, while also ensuring a high degree of compression for data with high repetition rates.

[0119] One embodiment of the present invention also proposes a blockchain civil engineering test data storage device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of a blockchain civil engineering test data storage method provided in one embodiment of the present invention. This device has the same beneficial effects as the aforementioned blockchain civil engineering test data storage method, and will not be described in detail here.

[0120] Understandably, when the various devices of a blockchain civil engineering test data storage device are in operation, they need to utilize the blockchain civil engineering test data storage method provided in the aforementioned embodiments. Therefore, whether the memory, processor, and computer program are integrated or different hardware is configured to produce functions similar to those achieved by the present invention, they all fall within the protection scope of the present invention.

[0121] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0122] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A blockchain civil engineering test data storage method, characterized by, The method includes: The system acquires civil engineering test data stored on the blockchain and real-time collected civil engineering test data, and preprocesses it to obtain analysis data, which includes character data. Based on the analyzed data, the importance, access frequency, and final repetition rate of character data are determined sequentially, and the encoding priority of character data is obtained, including: Based on the analysis data, the collection information of civil engineering test data in the same area at different time periods is obtained. The collection information includes the number of civil engineering test data and the acquisition time of civil engineering test data. The number of civil engineering test data for each region is obtained by numbering the civil engineering test data, and the time interval is obtained by subtracting the acquisition time of adjacent civil engineering test data. The formula for determining the importance of a region is as follows: wherein, represents the importance degree of the th region; represents the number of civil engineering test data of the th region; represents the time interval between the civil engineering test data obtained the th time and the next time; and the importance degree of the current region is the importance degree of the character data in the current region; Based on the analysis data, the number of accesses for each analysis data session and the total number of accesses for all analysis data sessions are calculated. The formula for determining the access frequency of the analyzed data is as follows: in, Indicates the first The frequency of access to the analyzed data; Indicates the first The number of times the data was accessed during the analysis; This represents the total number of accesses to all analyzed data; and the access frequency of the current analyzed data is the access frequency of its corresponding character data. Based on the determined importance, access frequency, and final repetition rate of the character data, the encoding priority of the character data is obtained, and the corresponding calculation formula is as follows: in, Indicates the first Encoding priority of individual character data; Indicates the first The importance of each character's data; Indicates the first The access frequency of each character data; Indicates the first The final repetition rate of the character data; Indicates to Normalized values; Represents the normalization function; Adjust the Huffman coding table; Based on the adjusted Huffman coding table, all civil engineering test data are compressed and re-stored in the blockchain.

2. The blockchain-based civil engineering experimental data storage method as described in claim 1, characterized in that, Acquire civil engineering test data stored on the blockchain and real-time collected civil engineering test data, and preprocess them to obtain analytical data, including: The first dataset is formed by collecting civil engineering test data in real time through testing equipment and storing the first dataset on the blockchain. The second dataset is formed by extracting the civil engineering test data stored on the blockchain and the first dataset. The second dataset was filtered to obtain the analysis data.

3. The blockchain-based civil engineering experimental data storage method as described in claim 1, characterized in that, The final repetition rate of character data is determined based on the analyzed data, including: Based on the analysis data, determine the repetition rate of character data in each analysis data, and determine the first repetition rate of character data in the same area; The overall repetition rate is obtained from all character data, and the final repetition rate is obtained by correcting the first repetition rate with the overall repetition rate.

4. The blockchain-based civil engineering experimental data storage method as described in claim 3, characterized in that, Based on the analyzed data, the repetition rate of character data in each analysis is determined, and the first repetition rate of character data within the current same region is determined, including: The formula for calculating the repetition rate of character data in each analysis is as follows: in, Indicates the first The first analysis of data The repetition rate of individual character data; Indicates the first [number]th [item] in the current analysis data. The number of times each character data appears; Indicates the first The total number of character data in each analysis; The formula for determining the first repetition rate of character data within the same current region is: in, Indicates the current number within the same region. The first repetition rate of the character data; Indicates the first The average repetition rate of each character within the same current region; Indicates standard deviation; Indicates the first The number of times the same character data exists in the same area; Indicates the current number within the same region The first analysis of data The repetition rate of individual character data; This represents the normalization function.

5. The blockchain-based civil engineering experimental data storage method as described in claim 4, characterized in that, The overall repetition rate is obtained based on all character data. The final repetition rate is obtained by adjusting the first repetition rate based on the overall repetition rate, including: The formula for calculating the overall repetition rate based on all character data is as follows: in, Indicates the first The overall repetition rate of individual character data; Indicates the first The number of times each character appears in all regions; Indicates the total number of characters in the data; The formula for calculating the final repetition rate by correcting the first repetition rate based on the overall repetition rate is as follows: in, Indicates the first The final repetition rate of the character data; Indicates the first The data for each character is based on the average of its first repetition rate in different regions; Indicates the first The overall repetition rate of individual character data.

6. The blockchain-based civil engineering experimental data storage method as described in claim 1, characterized in that, Adjusting the Huffman coding table includes: determining the coding priority of each character data in the analysis data according to the coding priority of the character data, sorting them from high to low, assigning low coding lengths to character data with high coding priority, until all character data is encoded to obtain the Huffman coding table.

7. A blockchain-based civil engineering test data storage device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of a blockchain-based civil engineering experimental data storage method as described in any one of claims 1 to 6.