Block chain-based geological data secure storage method and system, and storage medium

Through blockchain-based technology, the shortcomings of geological data storage and sharing in the existing technology are solved, efficient storage, flexible access and trusted circulation of geological data are achieved, and the security and reliability of data are ensured.

CN119989404AActive Publication Date: 2025-05-13MINERAL RESOURCES EXPLORATION CENT OF HENAN PROVINCIAL GEOLOGICAL BUREAU

Patent Information

Application Number
CN202510052068.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

The existing technology has shortcomings in the distributed storage and secure sharing of massive heterogeneous geological data, including the lack of unified and standardized data organization, insufficient scalability, lack of flexible access control mechanisms, lack of trusted verification mechanisms for the entire life cycle of data, and trusted data sharing mechanisms for multiple parties.

Method used

The blockchain-based geological data security storage method is adopted to achieve efficient storage, flexible access and trusted circulation of geological data through standardized geological data organization, layered network topology construction, distributed data block storage, fine-grained dynamic access control, multi-copy fault-tolerant backup of data and full-process behavior security auditing.

Benefits of technology

It effectively solves the problem of standardized management of massive heterogeneous geological data, ensures efficient storage and dynamic expansion of geological data, realizes flexible and diverse hierarchical authorized access, ensures traceability and tamper-proof of the entire life cycle of data, reduces the security risks of data leakage and illegal utilization, and supports trusted cross-domain geological data sharing applications from multiple parties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989404A_ABST
    Figure CN119989404A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a block chain-based geological data secure storage method and system, and a storage medium. The method comprises the steps of collecting geological environment data, establishing a standardized data set, storing distributed data blocks through a block chain network, constructing a hierarchical access authority structure, and realizing fine-grained access control. Polling voting, data verification and the like are utilized to guarantee storage consistency and reliability, and abnormal operation is audited through technologies such as time sequence analysis, behavior feature extraction and the like. According to the method, efficient storage, flexible access and credible circulation of the geological data are realized through key technologies such as a standardized geological data organization form, hierarchical network topology construction, distributed data block storage, fine-grained dynamic access control, data multi-copy fault-tolerant backup and whole-process behavior safety auditing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular to a blockchain-based geological data secure storage method, system and storage medium. Background Art

[0002] With the continuous deepening of geological exploration, the scale and types of geological data are becoming increasingly large. The traditional centralized geological data storage method mainly stores the data in the server or storage device of the data center, and manages the data access rights through usernames, passwords, access control lists, etc. This method is simple to operate, but it has many shortcomings in the distributed storage and secure sharing of massive heterogeneous data. In addition, more and more geophysical exploration, remote sensing and other technologies are being used in geological surveys, generating a large amount of unstructured data such as images and videos, which further increases the heterogeneity and complexity of geological data. How to effectively organize, manage and securely store these geological big data, and support flexible access and trusted sharing for different business needs, has become an urgent problem to be solved.

[0003] The existing centralized geological data storage methods have the following major deficiencies in dealing with massive heterogeneous data: lack of a unified and standardized data organization form, geological data from different sources and types vary greatly in format, coordinate system, resolution, etc., making it difficult to effectively integrate; the centralized storage architecture lacks scalability and cannot effectively cope with the exponential growth of geological data scale, and there are performance bottlenecks and single point failure risks; lack of flexible fine-grained data access control mechanism, mainly relying on static access control strategies, unable to meet differentiated user permission requirements; lack of a trusted verification mechanism for the entire life cycle of data, making it difficult to trace the source and evolution process of data, and there are security risks of data tampering; lack of a multi-party trusted data sharing mechanism, data exchange and value circulation between different stakeholders are difficult to reflect, affecting the data scale effect. Summary of the invention

[0004] The present application provides a blockchain-based geological data security storage method, system and storage medium, which is used to achieve efficient storage, flexible access and trusted circulation of geological data through key technologies such as standardized geological data organization, hierarchical network topology construction, distributed data block storage, fine-grained dynamic access control, multi-copy fault-tolerant backup of data, and full-process behavioral security audit.

[0005] In the first aspect, the present application provides a method for secure storage of geological data based on blockchain, which includes: collecting data on the topography, mining environment, hydrological environment and soil environment of the study area, establishing spatiotemporal correlation mapping of the collected data through zoning geocoding, and standardizing the mapping data through spatial reference conversion to obtain a standardized geological environment data set; performing hierarchical credit evaluation and dynamic load distribution on network nodes based on the standardized geological environment data set, constructing a network topology of the nodes through a minimum spanning tree algorithm to obtain a blockchain basic network; and performing geospatial mapping on the standardized geological environment data set in the blockchain basic network according to a geographic grid. The rows are partitioned into quadtrees, and the partitioned data is stored in adjacent blocks through data association evaluation to obtain distributed data blocks; a three-layer access permission structure is constructed based on the distributed data blocks, the permission inheritance relationship is calculated through data sensitivity quantification, and the user permissions are verified through environmental parameter constraints to obtain an access control chain; the nodes are grouped and numbered according to the access control chain, the data consistency is verified through polling voting, and the backup data is backed up through data check codes to obtain verification records; the verification records are analyzed in time series, the operation modes are classified through behavioral feature extraction, and abnormal behaviors are judged through access frequency statistics and spatiotemporal feature matching to obtain an audit report.

[0006] In a second aspect, the present application provides a blockchain-based geological data security storage system, the blockchain-based geological data security storage system comprising:

[0007] The acquisition module is used to collect data on the topography, mining environment, hydrological environment and soil environment of the study area, establish spatiotemporal correlation mapping of the collected data through zoning geocoding, and standardize the mapped data through spatial reference conversion to obtain a standardized geological environment data set;

[0008] A construction module is used to perform hierarchical credit assessment and dynamic load distribution on network nodes according to the standardized geological environment data set, and to construct a network topology of nodes through a minimum spanning tree algorithm to obtain a blockchain basic network;

[0009] A partitioning module, used to perform quadtree partitioning on the standardized geological environment data set in the blockchain basic network according to the geographic grid, and perform proximity group storage on the partitioned data through data association evaluation to obtain distributed data blocks;

[0010] A calculation module is used to construct a three-layer access permission structure based on the distributed data block, calculate the permission inheritance relationship through data sensitivity quantification, verify the user permission through environmental parameter constraints, and obtain an access control chain;

[0011] A backup module, used to group and number nodes according to the access control chain, verify data consistency through polling and voting, and back up the data to be backed up through a data check code to obtain a verification record;

[0012] The classification module is used to perform time series analysis on the verification records, classify the operation modes by extracting behavioral features, determine abnormal behaviors by access frequency statistics and spatiotemporal feature matching, and obtain an audit report.

[0013] The third aspect of the present application provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computer, it enables the computer to execute the above-mentioned blockchain-based geological data security storage method.

[0014] In the technical solution provided by this application, data is collected on the topography, mining environment, hydrological environment and soil environment of the study area to establish a standardized geological environment data set, and the distributed data blocks are stored in the blockchain-based network generated by the hierarchical network topology. A three-layer access permission structure is constructed to achieve fine-grained access control, and the consistency and reliability of data storage and backup are guaranteed by means of polling voting, data verification code, etc., and abnormal operation behaviors are audited in combination with time series analysis, behavioral feature extraction and other technologies. This method can effectively solve the problem of standardized management of massive heterogeneous geological data, ensure the efficient storage and dynamic expansion of geological data, realize flexible and diverse hierarchical authorized access to geological data, ensure the traceability and tamper-proof of geological data throughout its life cycle, reduce the security risks of data leakage and illegal use, support multi-party trusted cross-domain geological data sharing applications, and lay a solid foundation for the in-depth analysis and mining of geological big data. In addition, the full-process abnormal behavior audit and warning can timely discover and curb the data abuse of internal personnel and improve the security protection capabilities of the entire system. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0016] Figure 1 This is a schematic diagram of an embodiment of a method for securely storing geological data based on blockchain in an embodiment of the present application;

[0017] Figure 2 This is a schematic diagram of an embodiment of a geological data security storage system based on blockchain in an embodiment of the present application. DETAILED DESCRIPTION

[0018] Embodiments of the present application provide a method, system and storage medium for secure storage of geological data based on blockchain. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0019] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiments of the present application, an embodiment of the method for secure storage of geological data based on blockchain includes:

[0020] Step S101, collect data on the topography, mining environment, hydrological environment and soil environment of the study area, establish spatiotemporal correlation mapping of the collected data through zoning geocoding, and standardize the mapped data through spatial reference conversion to obtain a standardized geological environment data set;

[0021] Step S102: Perform hierarchical credit assessment and dynamic load distribution on network nodes based on the standardized geological environment data set, and construct network topology of nodes using the minimum spanning tree algorithm to obtain a blockchain basic network;

[0022] Step S103, the standardized geological environment data set in the blockchain basic network is partitioned into quadtrees according to the geographic grid, and the partitioned data is stored in proximity blocks through data association evaluation to obtain distributed data blocks;

[0023] Step S104: construct a three-layer access permission structure based on distributed data blocks, calculate the permission inheritance relationship through data sensitivity quantification, verify the user permission through environmental parameter constraints, and obtain an access control chain;

[0024] Step S105: group and number the nodes according to the access control chain, verify the data consistency through polling and voting, and back up the data to be backed up through the data verification code to obtain a verification record;

[0025] Step S106: Perform time series analysis on the verification records, classify the operation modes by extracting behavioral features, determine abnormal behaviors through access frequency statistics and spatiotemporal feature matching, and obtain an audit report.

[0026] It is understandable that the execution subject of this application can be a blockchain-based geological data security storage system, or a terminal or a server, which is not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.

[0027] Specifically, the geological environment data of the study area is collected using the sky-ground integration technology. The topographic image data is obtained through remote sensing equipment. The remote sensing data contains information such as surface morphology, slope, and slope direction. The mining environment parameters are obtained through field sampling, including data such as mine type, geological structure, and hydrogeological conditions. Long-term monitoring is carried out through hydrological monitoring stations to collect water environment parameters such as water quality, water level, water temperature, and flow rate. Soil sampling is used to analyze soil composition, obtain soil acidity, heavy metal content, and other soil environmental indicators, and form original environmental data. In the data collection stage, the topographic data mainly include elevation values, slope values, surface cover types, etc., the mining environment parameters include lithology, structure, weathering degree, etc., the water environment monitoring covers indicators such as pH value, dissolved oxygen, and conductivity, and the soil sampling analysis includes parameters such as organic matter content and heavy metal concentration.

[0028] The collected original environmental data are divided according to the administrative area boundaries, and the area is evenly divided using geographic grids. Each grid unit is assigned a unique geographic code to facilitate data location and retrieval. The geographic code adopts a hierarchical coding method, including the area code, grid number and data type identifier. A timestamp is added to each data record to mark the specific time node of data collection and establish the time-space correspondence of the data. The timestamp record is accurate to seconds to ensure the time sequence and traceability of the data. When collecting data from a certain monitoring point, the record contains information such as the monitoring point number, sampling time, longitude, latitude, and environmental parameter values. All data items have clear units and value ranges. In the process of spatial reference conversion, the seven-parameter conversion method is used to unify the data of different coordinate systems into the CGCS2000 coordinate system. The seven parameters include easting translation parameters, northing translation parameters, elevation translation parameters, horizontal axis rotation parameters, vertical axis rotation parameters, vertical axis rotation parameters, and scale ratio parameters. After the spatial reference conversion, the data format is standardized, the data structure and attribute fields are unified, the missing data are supplemented, and the outliers are eliminated. During the data normalization process, the units and decimal places of numerical data are unified, and the encoding and length of character data are unified to form a standardized geological environment data set.

[0029] The blockchain network nodes are constructed based on the standardized geological environment data set. The capacity of the network nodes is evaluated, and the capacity of the nodes is graded by calculating the data flow per unit time to evaluate the processing capacity of the nodes. Each node is divided into three capacity levels: high, medium, and low according to its processing capacity. The high-capacity node is responsible for processing large amounts of data and complex computing tasks, the medium-capacity node processes general data and daily computing tasks, and the low-capacity node is mainly used for data backup and simple computing tasks. The node trust evaluation is based on the historical data transmission success rate and the node online time, and a trust score is assigned to each node. Task scheduling is performed according to the node response time and load balancing factor to ensure network load balancing. The minimum spanning tree algorithm is used to construct the network topology. By calculating the connection weights between nodes, the optimal connection scheme is selected, and finally the communication rules between nodes are formulated to complete the construction of the blockchain basic network. The standardized geological environment data set is stored in quadtree partitions. First, the spatial scope is defined, and the boundaries of the study area are demarcated according to geographic coordinates. The geographic grid is used for uniform division, and then the quadtree partitioning method is used to hierarchically subdivide the geographic grid. The spatial relationship between partitions is calculated, adjacent partitions are identified, and the degree of association between partitions is quantified. The association of partitioned data is calculated through data content analysis, and similarity is calculated based on the temporal and spatial attributes and content characteristics of the data. Data blocks with high similarity are stored adjacently. Storage nodes are selected through data block size control and load balancing to form distributed data blocks. During data partition storage, data in adjacent areas is preferentially stored in the same node or adjacent nodes to reduce the overhead of cross-node access.

[0030] Build an access rights system based on distributed data blocks. Divide management levels, including administrator permission level, business user permission level, and general user permission level. Perform sensitivity analysis on data content, evaluate data value by keyword frequency, and divide sensitivity levels according to business types. Different permission levels correspond to different data access scopes and operation permissions. The administrator permission level has the highest operation permission, the business user permission level has read and write permissions for specific business data, and the general user permission level only has read permissions for basic data. Set permission inheritance paths and transfer rules to ensure the rationality and security of permission allocation. Perform environmental parameter detection, limit operation time windows, restrict login location ranges, and prevent illegal access and data leakage. Group nodes and verify data consistency. Each node is grouped according to its service capabilities and network location, and nodes in the group maintain data consistency through a polling voting mechanism. The voting process uses a weighted voting method, and the voting weight of high-trust nodes is larger. Nodes regularly verify and back up stored data, generate verification codes, and store multiple copies. The backup strategy considers the importance and access frequency of data, and backs up important data and hot data more frequently. Detailed operation logs are recorded during the backup process, including backup time, data version, storage location and other information.

[0031] Finally, data access behavior analysis is performed. Each data access operation of the user is recorded, including access time, access content, operation type and other information. The user's operation mode is identified through time series analysis, and the access frequency and access rules are counted. Abnormal access behaviors are marked and analyzed, such as frequent access to sensitive data, access at unconventional times, and remote login. Combined with the access location and time characteristics, a user behavior profile is established for subsequent access control and security audits. When generating an audit report, abnormal behaviors are classified according to risk levels, and detailed event descriptions and processing suggestions are provided. For example, taking hydrological monitoring data as an example, the water level, water quality and other parameters collected by the monitoring station are partitioned and stored according to geographical location after spatial benchmark conversion and data normalization. Permission verification is required when accessing data, and access logs are recorded. By analyzing the access records, it can be found whether there is a risk of data leakage, such as a user downloading a large amount of hydrological data in a short period of time, or accessing monitoring data in sensitive areas during non-working hours. At the same time, the data verification and backup mechanism ensures the security and reliability of the data. Even if some nodes fail, they can be restored through the backup data of other nodes.

[0032] In the embodiment of the present application, data is collected on the topography, mining environment, hydrological environment and soil environment of the study area, a standardized geological environment data set is established, and the distributed data blocks are stored in the blockchain-based network generated by the hierarchical network topology. A three-layer access permission structure is constructed to achieve fine-grained access control, and the consistency and reliability of data storage and backup are guaranteed by means of polling voting, data verification code, etc., and abnormal operation behaviors are audited in combination with time series analysis, behavioral feature extraction and other technologies. This method can effectively solve the standardized management problem of massive heterogeneous geological data, ensure the efficient storage and dynamic expansion of geological data, realize flexible and diverse hierarchical authorized access to geological data, ensure the traceability and tamper-proof of the entire life cycle of geological data, reduce the security risks of data leakage and illegal use, support multi-party trusted cross-domain geological data sharing applications, and lay a solid foundation for the in-depth analysis and mining of geological big data. In addition, the full-process abnormal behavior audit and warning can timely discover and curb the data abuse of internal personnel and improve the security protection capabilities of the entire system.

[0033] In a specific embodiment, the process of executing step S101 may include the following steps:

[0034] (1) Obtain image data of topography through remote sensing equipment, collect parameters of mining environment through field sampling, conduct long-term monitoring of hydrological environment through hydrological monitoring stations, and analyze the composition of soil environment through soil sampling to obtain original environmental data;

[0035] (2) Based on the original environmental data, the administrative area boundaries are divided, the area is divided according to the geographic grid, and each grid cell is assigned a unique geographic code through coordinate mapping to obtain the zoning code data;

[0036] (3) Mark the collection time based on the partition coding data, record the data collection node through the timestamp, establish the time-space correspondence through data association analysis, and obtain time-space association data;

[0037] (4) The coordinate system of the spatiotemporal correlation data is converted, the spatial position is offset by the easting translation parameter, the northing translation parameter, and the elevation translation parameter, the direction is adjusted by the horizontal axis rotation parameter, the vertical axis rotation parameter, and the vertical axis rotation parameter, and the scale is unified by the scale ratio parameter to obtain the benchmark conversion data;

[0038] (5) Normalize the data format according to the benchmark conversion data, unify the data structure through attribute field standardization, supplement the missing values ​​through data integrity check, and obtain standardized data;

[0039] (6) Perform quality inspection on standardized data, eliminate duplicate data through data consistency analysis, mark abnormal data through outlier detection, and obtain a standardized geological environment data set.

[0040] Specifically, the integrated sky-ground acquisition technology combines data acquisition methods at three levels: satellite remote sensing, aerial remote sensing, and ground monitoring. Remote sensing equipment includes multispectral imagers and hyperspectral imagers to collect images of topography. Multispectral imagers collect surface reflection information through different bands, including visible light, near-infrared, and mid-infrared bands, to obtain surface morphological characteristics; hyperspectral imagers can obtain more detailed spectral information for identifying surface material composition. For the collection of mine environmental parameters, the layout of field sampling points follows the principle of representativeness. Sampling points are set up in mining areas of different geological units and different mining stages to collect rock samples, groundwater samples, etc., and record geological structural characteristics, rock integrity, water content and other parameters. The layout of hydrological monitoring stations takes into account the distribution characteristics of the water system. Fixed monitoring stations are set up at the confluence of major rivers and tributaries and key monitoring areas, and automatic monitoring equipment is used to continuously collect parameters such as water level, flow, and water quality. Soil sampling adopts the grid distribution method, setting sampling points at a certain interval in the study area, collecting surface and profile soil samples, and analyzing indicators such as heavy metal content, organic matter content, and pH value.

[0041] Based on the collected original environmental data, regional division and coding are carried out. First, the first-level division is carried out according to the administrative boundaries to ensure the unity of data management and administrative management. On the basis of administrative division, regular grids are used for secondary division. The grid size is determined according to the area of ​​the study area and the data density to ensure that each grid unit contains enough sampling points. The geocoding adopts a hierarchical structure, with the administrative division code as the prefix, the grid number as the intermediate code, and the data type identifier as the suffix to form a unique geocoding. The coordinates of each grid unit are determined by the longitude and latitude of its lower left corner point, which is convenient for spatial positioning and retrieval. The data with geocoding are time-stamped to generate a timestamp in a unified format. The timestamp contains year, month, day, hour, minute, and second information, and is recorded in the international standard time format. For continuous monitoring data such as data from hydrological monitoring stations, the start and end time of monitoring is recorded; for discrete sampling data such as soil samples, the sampling time is recorded. Data association analysis establishes the connection between data through the two dimensions of time and space. Data at the same time and different spaces reflect the spatial distribution characteristics, and data at the same space and different times reflect the laws of temporal evolution. For example, the temporal correspondence between water quality monitoring data and soil monitoring data within a certain grid unit can reveal the mutual influence between surface water and soil environment.

[0042] The three translation parameters of easting, northing and elevation are used to adjust the position of the coordinate origin, the three rotation parameters of horizontal axis, vertical axis and vertical axis are used to adjust the direction of the coordinate axis, and the scale parameter is used to unify the scale of different coordinate systems. Through the combined transformation of these seven parameters, the conversion of data from different sources to a unified coordinate system is realized. The conversion process first determines the common control points, which have known coordinates in both coordinate systems, then calculates the conversion parameters, and finally converts all data points uniformly. The focus of standardization is to unify the data structure and data format. Attribute field standardization includes the normalization of field names, the unification of field types, and the standardization of field value ranges. For example, for water quality monitoring data, pH values ​​are uniformly retained to one decimal place, conductivity is uniformly used in μS / cm, and dissolved oxygen is uniformly used in mg / L. Data integrity checks supplement missing values ​​through contextual information, such as interpolation using data from adjacent time points, or estimation using data from adjacent spatial points. Data duplication often occurs in the process of multi-source data integration, and duplicate records are identified by comparing timestamps and spatial locations. Outlier detection is based on statistical methods, calculating the mean and standard deviation of the data, marking the value that deviates from the mean by more than three times the standard deviation as a suspected anomaly, and then combining professional knowledge to determine whether it is a true outlier. The quality-inspected data forms a standardized geological environment data set, which is the basis for subsequent blockchain storage.

[0043] For example: the monitoring station continuously collects water quality data, and each monitoring point records parameters such as water temperature, pH value, and dissolved oxygen every ten minutes. These raw data are first encoded according to the administrative region and grid unit where the monitoring station is located. For example, monitoring point A is located in the 25th grid in a certain area, and is encoded as "regional code-025-W" (W represents hydrological data). Each record carries an accurate timestamp, such as "2024-01-0610:30:00". The spatial coordinates of the monitoring station are converted and unified to the CGCS2000 coordinate system. After the data format is standardized, the water temperature is uniformly retained to one decimal place, and the pH value is retained to two decimal places. Abnormal data where the pH value remains above 9.5 for twelve consecutive hours is marked, and the abnormal value caused by equipment failure is confirmed by checking the instrument calibration record.

[0044] In a specific embodiment, the process of executing step S102 may include the following steps:

[0045] (1) Statistics were collected on the data storage volume in the standardized geological environment dataset, network nodes were classified according to the data flow per unit time, and node processing capacity was evaluated according to resource occupancy rate to obtain node capacity data;

[0046] (2) Calculate the node trust score based on the node capacity data, evaluate the node stability based on the historical data transmission success rate, and quantify the node reliability based on the node online time to obtain hierarchical credit assessment data;

[0047] (3) Perform node task allocation and scheduling based on hierarchical credit assessment data, sort task priorities according to node response time, adjust resource allocation based on load balancing factors, and obtain dynamic load distribution data;

[0048] (4) Calculate the connection weights between nodes based on the dynamic load distribution data, quantify the communication overhead between nodes according to the distance matrix, estimate the connection cost according to the communication quality index, and obtain the node connection data;

[0049] (5) Calculate the minimum cost based on the node connection data, optimize the node connection order according to the minimum spanning tree algorithm, select the connection scheme according to the path cost, and obtain the network topology data;

[0050] (6) Establish inter-node communication rules based on network topology data, standardize the data transmission method according to the communication protocol, configure the block generation rules through the node consensus mechanism, and obtain the blockchain basic network.

[0051] Specifically, the capacity of the standardized geological environment data set is evaluated. The processing requirements of the network nodes are determined by statistical data storage, including indicators such as data file size, number of data records, and data update frequency. The data flow per unit time is calculated using the following formula:

[0052]

[0053] Among them, F represents the data flow per unit time, V i represents the size of the i-th data file, T i represents the data transmission time, n represents the total number of data files, α is the data redundancy coefficient, D r Indicates the data update rate. The node processing capacity is quantified by calculating the resource occupancy rate:

[0054]

[0055] Among them, C represents the node processing capacity index, M u and M t are used memory and total memory respectively, P u and P t are CPU usage and total processing capacity, S u and S t are the used storage space and total storage space respectively, and β1, β2, and β3 are weight coefficients.

[0056] The node trust is calculated based on the node capacity data, taking into account the historical data transmission success rate, node online time and node response time. The data transmission success rate reflects the stability of the node. The total number of data transmissions and the number of successful data transmissions for each node are recorded during the calculation process. The node online time statistics include two dimensions: continuous online time and cumulative online time. The node reliability index is obtained by weighted average. The node trust score adopts a cumulative evaluation method. The initial trust of the newly added node is low, and the trust gradually increases with the increase in the number of successful interactions. For the load distribution process, the hierarchical credit evaluation data is first used as the basis, and the tasks are prioritized according to the node response time. The determination of task priority needs to consider the importance of data, processing time limit and resource requirements, and high-priority tasks are assigned to nodes with high trust and good performance. The load balancing factor is used to adjust the uniformity of task distribution to avoid overloading some nodes while other nodes are idle. The load distribution process is dynamically adjusted, and the task allocation strategy is updated in time when the node performance or network status changes.

[0057] When determining the connection relationship between nodes, the connection weight between nodes needs to be calculated. The connection weight takes into account factors such as the physical distance between nodes, network bandwidth, and communication delay. The number of network hops between each node is recorded through the distance matrix. The fewer the hops, the lower the communication overhead. Communication quality indicators include network delay, packet loss rate, and bandwidth utilization, which together determine the connection cost. The calculation of the connection cost needs to balance multiple factors, ensuring both communication efficiency and the rational use of network resources. When constructing the network topology based on node connection data, the minimum spanning tree algorithm is used for optimization. The minimum spanning tree algorithm starts with the edge with the lowest cost and gradually builds a loop-free network structure connecting all nodes. During the execution of the algorithm, the edge with the lowest cost is first selected to join the spanning tree, and then the edge with the lowest cost that will not form a loop is selected from the remaining edges until all nodes are connected. The final network topology not only ensures the connectivity of the network, but also minimizes the overall communication overhead.

[0058] After the network topology is determined, it is necessary to establish communication rules between nodes. The communication protocol specifies the format, process and verification method of data transmission to ensure the reliability and security of data transmission. The node consensus mechanism is used to maintain data consistency. When a new data block needs to be added to the blockchain, the consensus mechanism is used to ensure that each node reaches an agreement on the validity of the data.

[0059] For example, in a geological environment monitoring network, there are four monitoring nodes responsible for data collection and processing in different areas. First, the data processing status of each node is counted. Node A mainly processes terrain data, with a large amount of data but a low update frequency; Node B is responsible for hydrological data, with a medium amount of data and needs to be updated in real time; Node C processes soil data, with a small amount of data but requires frequent sampling; Node D, as a backup node, needs to store important data of other nodes. Based on these characteristics, the nodes are capacity graded and trust evaluated. The connection method of the nodes adopts the minimum spanning tree structure. Considering the physical location and network status between the nodes, the ABDC backbone connection method is finally formed, in which Node B is directly connected to other nodes as the central node. In the process of data transmission, the real-time requirements of hydrological data are high, so the task of Node B is given a higher priority. When Node B detects abnormal water quality, it can quickly pass the alarm information to other nodes to achieve timely response.

[0060] The communication rules between nodes stipulate the generation and verification process of data blocks. When a node collects new monitoring data, it first packages the data into a standard format, adds a timestamp and digital signature, and then broadcasts it to other nodes for verification. After receiving the data, other nodes verify the integrity and validity of the data. Only when more than 2 / 3 of the nodes confirm that the data is valid will the new data block be added to the blockchain. The blockchain network established in this way not only ensures the safe storage of geological environment monitoring data, but also realizes the efficient sharing and processing of data. The network structure has good scalability. When new monitoring nodes need to be added, the network topology can be dynamically adjusted through the minimum spanning tree algorithm to enable new nodes to smoothly access the existing network. At the same time, the node-based trust evaluation and load balancing mechanism ensure the stability and efficiency of network operation.

[0061] In a specific embodiment, the process of executing step S103 may include the following steps:

[0062] (1) Defining the spatial scope of the standardized geological environment dataset in the blockchain basic network, defining the boundaries of the study area according to geographic coordinates, dividing the regional geomorphic units according to terrain features, and obtaining regional scope data;

[0063] (2) Gridding the study area based on the regional range data, evenly dividing the area through the geographic grid, and hierarchically subdividing the geographic grid according to the quadtree partitioning method to obtain quadtree partitioning data;

[0064] (3) The spatial relationship of partitions is calculated through quadtree partition data, adjacent partitions are marked according to spatial positions, and the degree of association between partitions is quantified according to geographic distance to obtain spatial relationship data;

[0065] (4) Analyze the partition data content based on the spatial relationship data, calculate the partition content similarity through data correlation evaluation, quantify the data correlation according to the attribute correlation coefficient, and obtain correlation evaluation data;

[0066] (5) The partitioned data is merged into blocks based on the association evaluation data, the block range is defined by the proximity threshold, and the size of the block is controlled according to the data block size to obtain block storage data;

[0067] (6) Allocate nodes for block storage data, select storage nodes through load balancing, determine backup nodes through data redundancy strategy, and obtain distributed data blocks.

[0068] Specifically, to define the boundaries of the standardized geological environment dataset in the blockchain infrastructure network, it is necessary to first determine the geographic coordinate range of the study area, including the coordinates of the four vertices of the minimum circumscribed rectangle. The division of terrain features considers the integrity of the geomorphic unit and divides the study area into different geomorphic units such as mountains, hills, and plains. Each geomorphic unit has similar terrain and environmental characteristics. The division of geomorphic units needs to consider factors such as elevation, slope, and surface cover to ensure that the divided units have internal consistency and external differences.

[0069] For the data in the designated area, the grid partitioning method is used for spatial subdivision. The geographic grid is a regular spatial division method that divides the study area evenly into grid cells of equal size. The quadtree partitioning method is an adaptive spatial division strategy that recursively subdivides the grid according to data density and spatial characteristics. Starting from the root node, when the amount of data or data heterogeneity in a grid cell exceeds the threshold, the grid is divided into four subgrids. This partitioning method can form a finer grid in data-dense areas and maintain a larger grid size in data-sparse areas.

[0070] Spatial relationship calculation involves the location relationship and degree of association between partitions. The spatial relationship of partitions is calculated by the following formula:

[0071]

[0072] Among them, R ij represents the spatial association between partition i and partition j, Q ij is the shared boundary length, L ij is the distance between the center points of the two partitions, H ij is the elevation difference, E max is the maximum elevation difference, G ij is the landform type similarity, B max is the maximum similarity, ω1, ω2, ω3 are weight coefficients.

[0073] Based on spatial relationship data, similarity analysis is performed on the partition data content. Content similarity considers the attribute characteristics of the data, such as geological parameters, environmental indicators, etc. The more similar the data content of adjacent partitions is, the greater the possibility of merging. For the data attributes of each partition, the attribute correlation coefficient is calculated to quantify the degree of correlation between the data. The correlation evaluation needs to consider the time series characteristics, spatial distribution characteristics and attribute value distribution characteristics of the data.

[0074] During the block merging process, the proximity threshold is set to control the merging range. When the spatial relationship and data similarity of adjacent partitions exceed the threshold, these partitions are merged into one block. Block size control ensures that the merged data block is neither too large to affect processing efficiency nor too small to cause storage fragmentation. For the node allocation of block storage data, the load balancing calculation formula is:

[0075]

[0076] Among them, N k is the load balancing index of node k, Y k is the current storage capacity, X total is the total storage capacity, W k is the processing power, Z max is the maximum processing capacity, U k is the current task number, J max is the maximum number of tasks, γ1, γ2, and γ3 are balance factors.

[0077] For example, in a mining area environmental monitoring project, the study area includes an open pit, a spoil dump, and a surrounding buffer zone. First, the boundaries of the study area are determined based on satellite images and terrain data, and the area is divided into three main geomorphic units: mining area, waste storage area, and ecological restoration area. A geographic grid with an initial grid size of 500 meters is used for uniform division, and then quadtree partitioning is performed based on data characteristics. In areas with intensive mining activities, the grid is subdivided to 125 meters; in areas with drastic environmental changes, it is subdivided to 62.5 meters; in buffer zones, a larger grid size is maintained. Spatial relationship calculations show that the partitions within the mining area have a high degree of spatial correlation, which is due to common terrain features and environmental impacts. Data content analysis found that adjacent monitoring points often have similar trends in pollutant concentration changes. Based on these characteristics, spatially adjacent and data-similar partitions are merged into data blocks. The final data blocks reflect the spatial distribution characteristics of environmental impacts, such as the spread of pollutants and the direction of groundwater flow. In the process of node allocation, considering the characteristics of frequent data updates and high processing requirements in the mining area, these data blocks are allocated to high-performance nodes. At the same time, to ensure data security, data copies are saved on geographically dispersed nodes. This distributed storage method not only ensures efficient access to data, but also provides reliable disaster recovery backup.

[0078] For each data block, the node allocation strategy is further refined. Load balancing not only considers the storage capacity and processing power of the node, but also the geographical distribution of the node. Data blocks in adjacent areas are preferentially stored on nodes with closer network distances to reduce network latency for data access. The data redundancy strategy adopts a multi-copy mechanism, and important environmental monitoring data are stored in at least three copies, distributed in different physical locations. The spatial distribution of geological environmental data shows obvious clustering characteristics. For example, in the groundwater monitoring data around the mining area, the water level and water quality data of adjacent monitoring wells often show strong correlation. This spatial correlation directly affects the degree of refinement of data partitioning. Areas with strong correlation require more detailed grid division to capture data change characteristics. At the same time, the temporal resolution of the data also affects the storage strategy. Different storage solutions are used for real-time monitoring data and periodic sampling data. The size control of data blocks requires balancing multiple factors. Too large data blocks will increase the storage pressure and processing burden of a single node, while too small data blocks will increase management overhead and network communication costs. Through practice, it is found that it is most appropriate to control the size of data blocks within the range that can support efficient query and analysis. The physical size and logical size of a data block need to be considered separately. The physical size affects storage efficiency, while the logical size affects query performance.

[0079] After the node allocation is completed, it is necessary to establish a mapping relationship between data blocks and nodes. This mapping relationship is not static, but is dynamically adjusted as the amount of data and the state of the node change. When the load of a node is too high, the system will trigger the migration of data blocks and transfer some data blocks to nodes with lower load. Similarly, when new nodes are added to the network, data blocks will be redistributed to optimize the storage balance of the entire network. The distributed data block network has good scalability and fault tolerance. Through reasonable data partitioning and node allocation, the efficiency of data access is guaranteed while the integrity and security of data are maintained. When a node fails, the replica data on other nodes can ensure the continuity of the service. At the same time, the system will automatically start the data recovery process and rebuild the lost data copy on the new node. This distributed storage architecture is particularly suitable for processing data with obvious spatial characteristics such as geological environment monitoring. It can effectively support spatial query and analysis, such as finding the trend of environmental parameter changes in a specific area, or analyzing the diffusion range of pollutants. At the same time, the spatial distribution characteristics of data also provide optimization space for data compression and indexing. Data blocks with high similarity can use more efficient compression algorithms, and data blocks with strong correlation can establish spatial indexes to accelerate queries.

[0080] In a specific embodiment, the process of executing step S104 may include the following steps:

[0081] (1) Based on the distributed data blocks, the management levels are divided, the access rights are graded according to the importance of the data, and the user groups are divided according to the permission levels to obtain a three-layer access rights structure; wherein the three-layer access rights structure includes an administrator permission layer, a business user permission layer, and a general user permission layer;

[0082] (2) Analyze the data content in the three-layer access permission structure, assess the data value based on the frequency of keyword occurrence, and classify the data sensitivity level by business type to obtain sensitivity quantitative data;

[0083] (3) Sort out the permission transfer relationship through sensitive quantitative data, set the permission inheritance path through the user group level, and limit the permission transfer rules according to the parent-child node relationship to obtain the permission inheritance data;

[0084] (4) Calculate the permission group to which the user belongs based on the permission inheritance data, determine the access level based on the user identity type, restrict the user's operation scope based on the permissions within the group, and obtain permission calculation data;

[0085] (5) Perform environmental detection on the permission calculation data, limit the operation window according to the access time, and constrain the login scope according to the access location to obtain environmental constraint data;

[0086] (6) The user's permission status is judged based on the environmental constraint data, the user's legitimacy is verified through identity authentication, and the access request is authorized through permission evaluation to obtain the access control chain.

[0087] Specifically, the management level is divided for distributed data blocks. According to the importance of data, data is divided into three categories: core data, business data and basic data. Core data includes key information such as geological disaster warning information and important mineral resource data; business data includes work data such as daily monitoring data and analysis reports; basic data includes general information such as historical records and reference materials. Based on this data classification, a three-layer access permission structure is divided accordingly: the administrator permission layer is responsible for the management and maintenance of the entire system and has the highest data access and operation permissions; the business user permission layer is mainly for professional and technical personnel, with access and processing permissions for specific business data; the general user permission layer is mainly for the query and use of basic data. After determining the three-layer access permission structure, it is necessary to conduct an in-depth analysis of the data content. The data value is determined by keyword frequency analysis. Keywords include professional terms such as "fault", "displacement" and "water content". The higher the frequency of these words, the higher the professionalism and importance of the data. At the same time, the data sensitivity is divided according to the business type. For example, data related to geological disaster warning has the highest sensitivity level, requiring real-time updates and strict confidentiality; while conventional geological parameter measurement data has a relatively low sensitivity level and allows a wider range of access and sharing.

[0088] The sorting out of permission transfer relationships is an important link based on sensitive quantitative data. The user group hierarchy is divided into system management group, business processing group and data query group according to responsibilities, and each group can be subdivided into different subgroups. Permission inheritance adopts a top-down transmission method. The permissions of the upper group include all permissions of the lower group, but the lower group cannot obtain the special permissions of the upper group. The permission transfer between parent and child nodes follows the principle of least privilege, ensuring that each user group can only obtain the minimum set of permissions required to complete its work.

[0089] The calculation of the permission group to which the user belongs is based on the permission inheritance data. First, the basic access level of the user is determined according to the user's identity type, such as system administrator, business specialist, ordinary user, etc. Then, the specific permission group to which the user belongs is determined according to the user's role and responsibilities in the organization. Each permission group has a clearly defined scope of operation, including the type of data that can be accessed, the type of operation allowed (such as reading, modifying, deleting, etc.), and the time limit for operation. Environmental detection is a dynamic permission control process. Environmental parameters are detected for each access request, including access time, login location, device characteristics, etc. The limitation of the operation time window ensures that sensitive operations can only be performed within the specified time period. For example, the modification of certain important data can only be performed during working hours. The constraint of the login location is implemented through IP address and physical location verification to prevent illegal access from other places.

[0090] The determination of permission status is the final security control link. The legitimacy of the user's identity is verified through multi-factor identity authentication, which includes password verification, digital certificates, biometrics, etc. Permission evaluation comprehensively considers the user's identity, environmental conditions, and operation requests to determine whether to grant access rights. The entire process forms a complete access control chain, and each access request must be verified by this control chain before it can be authorized.

[0091] For example, monitoring projects involve multiple aspects such as surface deformation monitoring, groundwater monitoring and soil environment monitoring. Surface deformation monitoring data include GPS measurement data, tilt measurement data, etc. These data are directly related to geological disaster warning and belong to core data; groundwater monitoring data include water level, water quality and other parameters, which belong to business data; soil environment monitoring data include conventional physical and chemical indicators, which belong to basic data.

[0092] Analyze the keyword frequency of each data item. Taking the surface deformation monitoring report as an example, the frequency of keywords such as "deformation", "displacement" and "settlement" is much higher than other terms, indicating that this part of the data is closely related to geological disaster warning. Through text analysis of a large number of historical monitoring reports, a keyword frequency statistical table is established as an important basis for data value assessment. At the same time, combined with the timeliness requirements and update frequency of the data, the sensitivity level is further refined. The sensitivity level of real-time monitoring data is higher than that of regular monitoring data, and the sensitivity level of abnormal data is higher than that of normal data. The setting of the permission inheritance path adopts a tree structure. At the administrator permission level, set up the system management group and the security audit group; at the business user permission level, set up the surface monitoring group, hydrological monitoring group and environmental monitoring group according to the monitoring type; at the ordinary user permission level, set up the data query group and the report statistics group. The permission inheritance relationship between the groups is clear and unambiguous. For example, the surface monitoring group inherits all the permissions of the data query group and has the permission to process surface monitoring data.

[0093] In the process of allocating user permissions, work needs and security requirements are considered. The person in charge of the monitoring project is assigned to the corresponding business group and has data processing and analysis permissions; the field technicians are assigned to specific monitoring groups and have data collection and entry permissions; the project managers are assigned to the management group and have global management and supervision permissions. The refined management of permissions ensures the controllability of data access. The setting of environmental constraints takes into account the actual work scenario. For monitoring data collected on-site, it is required to upload data within the specified monitoring point location range; for data analysis and processing, it is required to be carried out in a specified office network environment; for system maintenance operations, it is required to be performed on a specific management terminal. These environmental constraints are closely integrated with the workflow, which ensures data security without affecting normal work.

[0094] In specific operations, each access request must go through a complete control chain verification. For example, when a technician from the surface monitoring team needs to modify monitoring data, first verify his digital certificate to confirm his identity; then check whether he is at the specified working hours and working location; finally, determine whether he has the data modification permission based on the permission group he belongs to. Only when all verifications are passed can the data modification operation be performed.

[0095] In a specific embodiment, the process of executing step S105 may include the following steps:

[0096] (1) Evaluate the node service capability based on the access control chain, group the nodes by processing performance, and sort the grouped nodes by number according to the network distance to obtain node group data;

[0097] (2) Voting rules are set based on node grouping data, voting order is arranged based on node response time, node decision importance is differentiated based on voting weight, and round-robin voting data is obtained;

[0098] (3) Compare the storage content of each node based on the polling voting data, verify the data content through the data block feature value, mark the data differences according to the consistency judgment rules, and obtain consistency verification data;

[0099] (4) The data to be backed up is screened based on the consistency verification data, the backup priority is sorted based on the importance of the data, the backup content is confirmed based on the data integrity, and a list of data to be backed up is obtained;

[0100] (5) Redundantly encode the data in the backup list, divide the backup data into blocks through data sharding, calculate the data check code through check bit generation, and obtain data check data;

[0101] (6) The data backup operation is recorded through data verification data, the data version is identified by the backup time, and the data copy is located according to the backup location to obtain a verification record.

[0102] Specifically, the service capability of the node is comprehensively evaluated based on the access control chain. The node service capability evaluation includes three key indicators: processing performance, storage capacity, and network bandwidth. Processing performance is measured by CPU usage and memory occupancy, storage capacity includes available space and read and write speed, and network bandwidth considers uplink and downlink rates. Based on these indicators, the nodes are divided into high-performance group, medium-performance group, and basic performance group. Within each performance group, the nodes are numbered and sorted according to the network distance between nodes. The network distance is calculated by the communication delay and number of hops between nodes.

[0103] For the data after the nodes are grouped and sorted, specific voting rules are set. The voting process adopts a hierarchical voting mechanism, and high-performance group nodes have higher voting weights because these nodes have better data processing capabilities and stability. The voting order is arranged according to the node response time, and nodes with fast response time have priority in voting. The importance of node decision-making is related to its historical performance, including indicators such as the number of successfully processed transactions and the accuracy of data synchronization. This voting mechanism based on performance and reputation ensures the reliability of voting results.

[0104] During the consistency verification process, the data content stored in each node is compared. The calculation of the data block feature value uses a multiple hash algorithm to comprehensively calculate information such as data content, timestamp, and version number. The consistency judgment rule stipulates the allowable range of data differences, and marks the differences that exceed the allowable range. The marking information includes the difference location, difference type, and difference degree, which are used for subsequent data repair.

[0105] Based on the results of consistency verification, the data to be backed up is screened and prioritized. The evaluation of data importance takes into account the frequency of use, update time, and business relevance of the data. For example, the backup priority of real-time monitoring data is higher than that of historical archived data, and the backup priority of abnormal data is higher than that of normal data. Data integrity checks include structural integrity and content integrity to ensure that the data to be backed up is not damaged or missing.

[0106] Redundant coding is performed on the selected data to be backed up. The data sharding adopts a dynamic sharding strategy, and the number and size of shards are determined according to the size and importance of the data. Each data shard generates a corresponding check code, which is generated using erasure code technology, which can detect data errors and repair damaged data to a certain extent. The check code is stored together with the original data shard to form a complete backup unit.

[0107] Finally, the data backup operation is recorded in detail. Each backup operation records the backup time, data version number, and backup location information. The backup time is used to track the historical version of the data, the version number is used to distinguish the data status in different periods, and the backup location information is used to quickly locate and restore the data. These records constitute a complete verification chain, ensuring the traceability and recoverability of data backup.

[0108] In the specific implementation process, the backup verification of geological environment monitoring data in a mining area is used as an example to illustrate the entire process. The monitoring data includes three core data types: surface deformation monitoring, groundwater monitoring, and soil environment monitoring. The performance of the nodes responsible for storing these data is evaluated, and the performance indicators of the nodes are collected during the evaluation process: since surface deformation monitoring data requires real-time processing, it is assigned to high-performance group nodes for processing. The CPU usage of these nodes remains at a low level, the memory and storage space are sufficient, and the network bandwidth is sufficient to support real-time data transmission; groundwater monitoring data and soil environment monitoring data are assigned to medium-performance group nodes, which can meet the needs of regular data processing.

[0109] After determining the node grouping, the data consistency verification process begins. The verification starts with the high-performance group nodes because these nodes store the most critical real-time monitoring data. Each data block generates a feature value, which contains a comprehensive encoding of information such as data content, acquisition time, and monitoring point location. By comparing the feature values ​​of the same data block on different nodes, data inconsistencies are found. For example, if the data of a surface deformation monitoring point differs on different nodes, the correct data version needs to be determined through a voting mechanism.

[0110] During the voting process, the voting weight of high-performance group nodes is larger because the data of these nodes is updated more timely and the data reliability is higher. When data inconsistency is found, the node with the fastest response time will first initiate the vote, and other nodes will participate in the vote in turn according to their respective response times. For each controversial data block, multiple rounds of voting are used to determine the final correct version. The voting results are recorded in the block, forming an unalterable decision proof.

[0111] The priority of data backup fully considers the business value of the data. For example, the data backup of data points where abnormal deformation of the surface is detected will be given the highest priority because this data is directly related to geological disaster warning. Similarly, the backup priority of monitoring point data that detects abnormal water quality will also be increased. The data integrity check not only verifies the integrity of the data itself, but also verifies the integrity of related metadata, such as monitoring time, monitoring point location, instrument parameters and other information.

[0112] The data sharding strategy is flexibly adjusted according to the characteristics of the data. For monitoring data with strong time series, sharding is performed according to the time dimension to facilitate subsequent data recovery by time period; for data with strong spatial correlation, sharding is performed according to spatial location to facilitate the recovery of regional data. Each shard generates a corresponding checksum. The design of the checksum takes into account the characteristics of the data and adopts stronger error correction capabilities for important fields.

[0113] The backup records contain complete data traceability information. Each record clearly records the source, processing process and storage location of the data. For example, for surface deformation monitoring data, the records include timestamps such as the original acquisition time, data processing time, data verification time, backup time, and the storage location of the data on each node. These records form a complete data life cycle chain, and any data anomalies can be traced back to the source through this chain.

[0114] In a specific embodiment, the process of executing step S106 may include the following steps:

[0115] (1) Sort the verification records by timestamp, divide the verification records into intervals by time window, and segment the data sequence according to the continuity of the operation to obtain time series analysis data;

[0116] (2) Extract user operation behaviors based on time series analysis data, classify the behavior sequences according to the operation types, describe the operation process through the data access path, and obtain behavior feature data;

[0117] (3) Clustering user behavior patterns through behavior feature data, summarizing behavior rules according to operation sequence, grouping operation patterns through similarity calculation, and obtaining operation pattern data;

[0118] (4) Counting access behaviors based on operation mode data, counting the number of accesses per unit time, analyzing the operation frequency according to the access target, and obtaining access frequency data;

[0119] (5) The user operation locations are distributed and counted based on the access frequency data, the access locations are marked using geographic coordinates, and the spatiotemporal features are associated with the access periods to obtain spatiotemporal feature data;

[0120] (6) Perform anomaly comparison on spatiotemporal feature data, identify abnormal operations through normal behavior templates, classify abnormal behaviors according to risk levels, and obtain an audit report.

[0121] Specifically, the verification records are processed in time series. Each verification record contains a timestamp accurate to milliseconds. The division of time windows is based on the characteristics of business operations. For example, a day is divided into multiple operation periods: morning shift (8:00-16:00), evening shift (16:00-24:00) and night shift (0:00-8:00). In each time window, the data sequence is segmented according to the continuity of the operation. The criteria for continuity include the operation time interval, the relevance of the operation object and the relevance of the operation type. The in-depth processing of the time series analysis data focuses on the characteristics of the user's operation behavior. The operation types include basic operations such as data query, data modification, and data deletion, as well as complex operations such as data analysis and report generation. The data access path records the complete operation chain of the user from login to logout, including information such as the type of data accessed, the order of access, and the duration of stay. By analyzing this information, a characteristic description of the user's operation behavior is formed.

[0122] Behavioral clustering analysis uses a time-series-based clustering algorithm. First, the operation sequence is converted into a feature vector. The feature dimensions include operation type, operation frequency, operation duration, etc. The analysis of the operation sequence focuses on the order of operations. For example, data viewing operations usually precede data modification operations. The similarity calculation uses a dynamic time warping algorithm, which can handle operation sequences of different lengths and consider the elasticity of operation timing. When performing access behavior statistics based on operation mode data, focus on the operation density and operation distribution per unit time. The statistical window can be hours, shifts, or days, and different statistical granularities are set for different types of operations. The analysis of access targets includes data type distribution, access level distribution, and operation permission distribution. These statistical results reflect the user's job responsibilities and operating habits.

[0123] The geographic location distribution analysis of user operations is an important means of discovering abnormal access. Each access request carries geographic coordinate information, which is converted into actual location information through geocoding. The correlation analysis of spatiotemporal features combines access time and location to form a spatiotemporal trajectory of user activities. For example, access requests occurring in different locations within a short period of time often indicate potential security risks.

[0124] The identification of abnormal behavior is based on the construction of normal behavior templates. Normal behavior templates are extracted from historical access records, including typical operation modes, common access times and locations, reasonable operation frequencies and other features. For example, the normal operation mode of geological data collectors is to upload data at fixed monitoring points, the operation time is concentrated in working hours, and the operation content is mainly data upload and simple data verification. The normal operation mode of data analysts includes more data query and analysis operations, and the operation location is usually in a fixed office. The division of risk levels considers abnormal characteristics in multiple dimensions. The highest risk level corresponds to behaviors that clearly violate security policies, such as unauthorized data modification and batch download of sensitive data; the medium risk level corresponds to behaviors that deviate from the normal mode but have not yet posed a direct threat, such as data access at unconventional times, login requests at abnormal locations, etc.; the low risk level corresponds to some minor anomalies, such as small fluctuations in operation frequency and atypical operation sequences.

[0125] For example, a geological environment monitoring project involves data collection and analysis at multiple monitoring points, and users include on-site monitoring personnel, data analysts, and system managers. A time series analysis of the verification records within a week shows that on-site monitoring personnel regularly upload data between 9 a.m. and 3 p.m. every day, and the data upload interval for each monitoring point is about 2 hours, which constitutes the basic operation time series characteristics.

[0126] When extracting user operation behaviors, we found that the operation sequence of monitoring personnel is usually: log in to the system - select monitoring points - upload data - data verification - exit the system, and the whole process lasts between 5 and 10 minutes. The typical operation sequence of data analysts includes: log in to the system - data query - data analysis - report generation - data export, and the operation time span is longer, which may last for several hours. These operation sequences constitute the behavioral characteristics of different user roles.

[0127] Cluster analysis of these behavior characteristics has led to the formation of three main operation modes: data collection mode, data analysis mode, and system management mode. Each mode has its own specific operation sequence and time characteristics. For example, the operation of the data collection mode is usually linear, short-term, and repetitive; the operation of the data analysis mode is more complex, involving multiple iterations of query and analysis.

[0128] The statistics of access frequency show obvious time patterns. The number of visits on weekdays is significantly higher than that on weekends, and the access characteristics of the morning and evening shifts are also significantly different. Through the analysis of access targets, it is found that the data access tendencies of different user groups are: monitoring personnel mainly access raw data, analysts access more statistical data and historical data, and managers' access is distributed in various types of data.

[0129] The spatiotemporal feature analysis pays special attention to off-site access and cross-regional access. Under normal circumstances, the access location of the monitoring personnel should match the location of the monitoring point, and the access location of the data analyst should be within the office area. Any access that deviates from this pattern will be marked as a potential anomaly. For example, the data upload location of the monitoring personnel is too far away from the monitoring point, or the analyst accesses sensitive data from an abnormal location during non-working hours. The generated audit report records in detail the abnormal behaviors found and their risk assessment results. For high-risk anomalies, the report gives the specific time, location, operation content and violation type; for medium and low-risk anomalies, the report provides relevant background information and suggested improvement measures.

[0130] In a specific embodiment, the process of performing the step of clustering user behavior patterns through behavior feature data may include the following steps:

[0131] (1) Sequence segmentation is performed on the behavior feature data, the behavior segments are segmented according to the operation time interval, and the behavior sequence is marked according to the minimum operation unit to obtain the behavior sequence data;

[0132] (2) Calculate the operation correlation through the behavior sequence data, perform correlation analysis on the continuous behaviors through the operation combination, quantify the behavior relationship according to the operation co-occurrence frequency, and obtain the behavior correlation data;

[0133] (3) Construct a behavior distance matrix based on the behavior association data, evaluate the similarity of behavior sequences by calculating the edit distance, merge the behavior patterns according to the distance threshold, and obtain the behavior distance data;

[0134] (4) Perform hierarchical clustering on the behavior distance data, divide the behavior categories by the inter-cluster distance, and classify the user behavior patterns according to the cluster center to obtain the behavior cluster data;

[0135] (5) Extracting behavioral feature vectors based on behavioral clustering data, describing behavioral regularities through operation transition probabilities, characterizing behavioral patterns according to temporal relationships, and obtaining behavioral regularity data;

[0136] (6) The operation modes are grouped according to the behavior pattern data, similar behaviors are merged according to the similarity threshold, and the operation modes are identified according to the behavior characteristics to obtain the operation mode data.

[0137] Specifically, the segmentation of the behavior sequence is based on the operation time interval. When the time interval between adjacent operations exceeds the preset threshold, the interval between the two operations is regarded as the segmentation point of a behavior segment. The minimum operation unit is an indivisible basic operation, such as data query, data modification, data upload, etc. Each minimum operation unit contains basic attributes such as operation type, operation object, and operation time. The marking of the behavior sequence adopts string encoding. Each operation type corresponds to a unique character code, and a complete behavior sequence description is formed by string concatenation.

[0138] The calculation of operation correlation focuses on the relationship between consecutive operations. Operation combination analysis considers the order and time interval of operations. For example, "data query-data modification-data upload" constitutes a typical operation combination. Operation co-occurrence frequency indicates the probability of two operations co-occurring in the same behavior segment. The higher the frequency, the stronger the correlation between the two operations. By constructing an operation correlation matrix, the correlation strength between all operation pairs is recorded.

[0139] The behavior distance matrix is ​​constructed using the edit distance calculation method. The edit distance refers to the minimum number of operations required to transform one behavior sequence into another behavior sequence, including insertion, deletion, and substitution. For two behavior sequences, the similarity of the sequences can be quantified by calculating the edit distance between them. The distance threshold is set based on specific business needs. When the edit distance of two behavior sequences is less than the threshold, the two sequences are considered to belong to the same behavior pattern.

[0140] Hierarchical cluster analysis adopts a bottom-up aggregation strategy. First, each behavior sequence is regarded as an independent cluster, and then the closest clusters are gradually merged until the preset number of clusters is reached or the distance between clusters exceeds the threshold. The distance between clusters is calculated using the average linkage method, that is, the average distance between all sequence pairs in two clusters is calculated. The cluster center refers to the sequence with the smallest average distance to other sequences in the same cluster. The cluster center can represent the typical characteristics of this type of behavior pattern.

[0141] The extraction of behavioral feature vectors focuses on the statistical and structural characteristics of the sequence. The operation transition probability describes the conversion relationship between different operations and is represented by constructing a Markov transition matrix. For example, in geological data processing, the probability of data query operations being followed by data analysis operations is high, while data modification operations are usually followed by data verification operations. The characterization of temporal relationships not only considers the order of operations, but also the time interval distribution of operations.

[0142] The grouping of operation modes requires setting a reasonable similarity threshold. The calculation of similarity comprehensively considers multiple dimensions such as the edit distance of the behavior sequence, the probability of operation transfer, and time characteristics. In the process of merging similar behaviors, the most representative features are retained to form a standard mode for this type of behavior. The identification of behavior features includes information such as mode number, typical sequence, and main features.

[0143] Taking the geological environment monitoring of a mining area as an example, the entire analysis process is explained: the monitoring work includes multiple links such as on-site data collection, data processing, and data analysis. By analyzing the operation logs within a day, it is found that the typical operation sequence of data collectors includes: logging into the system (L)-selecting monitoring points (S)-data collection (C)-data upload (U)-data verification (V)-exiting the system (E), which can be expressed as "LSCUVE". When the time interval between two collection operations exceeds 30 minutes, it is divided into different behavior segments.

[0144] When calculating the operation correlation, it is found that the operation pairs such as "data collection-data upload" and "data upload-data verification" have a high co-occurrence frequency, while the operation pairs such as "data collection-data analysis" rarely appear at the same time. This correlation relationship reflects the rationality of the business process. Through the edit distance calculation, similar operation modes can be identified. For example, the collection operations of different monitoring points may only differ in the step of monitoring point selection.

[0145] As cluster analysis proceeds, the operation sequences gradually form different categories. Each category has its own specific characteristics. For example, the operation sequences of data collection are usually short and regular, while the operation sequences of data analysis are long and contain multiple iterations. The cluster center represents the most typical operation sequence of each category, which can be used as a reference standard for behavioral pattern recognition.

[0146] In the process of extracting behavioral rules, special attention is paid to the temporal characteristics of operations. For example, data collection operations are usually performed at fixed time points, while the time distribution of data analysis operations is relatively scattered. These temporal characteristics, together with the operation transition probability, constitute a complete description of the behavior pattern. Finally, by setting an appropriate similarity threshold, behavior sequences with similar characteristics are merged into the same operation mode to form a behavior classification system.

[0147] The above describes the geological data security storage method based on blockchain in the embodiment of the present application. The following describes the geological data security storage system based on blockchain in the embodiment of the present application. Figure 2 In the embodiments of the present application, an embodiment of a geological data security storage system based on blockchain includes:

[0148] The acquisition module is used to collect data on the topography, mining environment, hydrological environment and soil environment of the study area, establish spatiotemporal correlation mapping of the collected data through zoning geocoding, and standardize the mapped data through spatial reference conversion to obtain a standardized geological environment data set;

[0149] The construction module is used to perform hierarchical credit assessment and dynamic load distribution on network nodes based on the standardized geological environment data set, and to construct the network topology of nodes through the minimum spanning tree algorithm to obtain the blockchain basic network;

[0150] A partitioning module, used to perform quadtree partitioning on the standardized geological environment data set in the blockchain basic network according to the geographic grid, and perform proximity group storage on the partitioned data through data association evaluation to obtain distributed data blocks;

[0151] A calculation module is used to construct a three-layer access permission structure based on the distributed data block, calculate the permission inheritance relationship through data sensitivity quantification, verify the user permission through environmental parameter constraints, and obtain an access control chain;

[0152] A backup module, used to group and number nodes according to the access control chain, verify data consistency through polling and voting, and back up the data to be backed up through a data check code to obtain a verification record;

[0153] The classification module is used to perform time series analysis on the verification records, classify the operation modes by extracting behavioral features, determine abnormal behaviors by access frequency statistics and spatiotemporal feature matching, and obtain an audit report.

[0154] Through the collaboration of the above components, data collection on the topography, mining environment, hydrological environment and soil environment of the study area is carried out to establish a standardized geological environment data set. The distributed data blocks are stored in the blockchain-based network generated by the hierarchical network topology, and a three-layer access permission structure is constructed to achieve fine-grained access control. The consistency and reliability of data storage and backup are guaranteed by means of polling voting and data checksums, and abnormal operation behaviors are audited by combining time series analysis, behavioral feature extraction and other technologies. This method can effectively solve the problem of standardized management of massive heterogeneous geological data, ensure the efficient storage and dynamic expansion of geological data, realize flexible and diverse hierarchical authorized access to geological data, ensure the traceability and tamper-proof of geological data throughout its life cycle, reduce the security risks of data leakage and illegal use, support multi-party trusted cross-domain geological data sharing applications, and lay a solid foundation for the in-depth analysis and mining of geological big data. In addition, the full-process abnormal behavior audit and warning can timely discover and curb the data abuse of internal personnel and improve the security protection capabilities of the entire system.

[0155] The present application also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the blockchain-based geological data secure storage method.

[0156] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0157] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.

[0158] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for secure storage of geological data based on blockchain, characterized in that: The blockchain-based geological data security storage method includes: Data collection was conducted on the topography, mining environment, hydrological environment and soil environment of the study area. Spatiotemporal correlation mapping was established for the collected data through zoning geocoding, and the mapped data was standardized by spatial reference conversion to obtain a standardized geological environment data set. Based on the standardized geological environment data set, the network nodes are evaluated for hierarchical credit and dynamically allocated for load, and the network topology of the nodes is constructed through the minimum spanning tree algorithm to obtain the blockchain basic network; The standardized geological environment data set in the blockchain basic network is partitioned into quadtrees according to the geographic grid, and the partitioned data is stored in proximity blocks through data association evaluation to obtain distributed data blocks; Building a three-layer access permission structure based on the distributed data block, calculating the permission inheritance relationship through data sensitivity quantification, verifying the user permission through environmental parameter constraints, and obtaining an access control chain; The nodes are grouped and numbered according to the access control chain, data consistency is verified by polling and voting, and the data to be backed up is backed up via a data check code to obtain a verification record; The verification records are subjected to time series analysis, the operation modes are classified by extracting behavioral features, abnormal behaviors are determined by access frequency statistics and spatiotemporal feature matching, and an audit report is obtained.

2. The method for secure storage of geological data based on blockchain according to claim 1, characterized in that: The data of the topography, mining environment, hydrological environment and soil environment in the study area are collected, and the spatial and temporal correlation mapping of the collected data is established through zoning geocoding, and the mapping data is standardized by spatial reference conversion to obtain a standardized geological environment data set, including: The image data of the topography is acquired through remote sensing equipment, and the parameters of the mine environment are collected through field sampling. The hydrological environment is monitored for a long time through the hydrological monitoring station, and the soil environment is analyzed through soil sampling to obtain the original environmental data; Divide administrative area boundaries according to the original environmental data, divide the area according to geographic grids, assign a unique geographic code to each grid unit through coordinate mapping, and obtain partition code data; Marking the collection time based on the partition coding data, recording the data collection node through the timestamp, establishing the time-space correspondence through data association analysis, and obtaining time-space association data; The coordinate system of the spatiotemporal correlation data is converted, the spatial position is offset by an easting translation parameter, a northing translation parameter, and an elevation translation parameter, the direction is adjusted by a horizontal axis rotation parameter, a vertical axis rotation parameter, and a vertical axis rotation parameter, and the scale is unified by a scale ratio parameter to obtain benchmark conversion data; The data format is normalized according to the benchmark conversion data, the data structure is unified by attribute field standardization, and missing values ​​are supplemented by data integrity check to obtain standardized data; The standardized data is quality checked, duplicate data is eliminated through data consistency analysis, and abnormal data is marked through outlier detection to obtain the standardized geological environment data set.

3. The method for secure storage of geological data based on blockchain according to claim 1, characterized in that: The method of performing hierarchical credit assessment and dynamic load distribution on network nodes based on the standardized geological environment dataset, and constructing a network topology of the nodes via a minimum spanning tree algorithm to obtain a blockchain basic network includes: Statistics are collected on the data storage volume in the standardized geological environment data set, network nodes are classified according to the data flow per unit time, and node processing capabilities are evaluated according to resource occupancy rates to obtain node capacity data; The node trust is scored based on the node capacity data, the node stability is evaluated based on the historical data transmission success rate, and the node reliability is quantified based on the node online time to obtain hierarchical credit assessment data; Perform node task allocation and scheduling based on the hierarchical credit assessment data, sort the task priorities according to the node response time, adjust the resource allocation according to the load balancing factor, and obtain dynamic load distribution data; Calculate the connection weights between nodes based on the dynamic load distribution data, quantify the communication overhead between nodes according to the distance matrix, estimate the connection cost according to the communication quality index, and obtain the node connection data; The minimum cost is calculated based on the node connection data, the node connection sequence is optimized according to the minimum spanning tree algorithm, and the connection scheme is selected according to the path cost to obtain the network topology data; The node-to-node communication rules are established according to the network topology data, the data transmission method is standardized according to the communication protocol, and the block generation rules are configured through the node consensus mechanism to obtain the blockchain basic network.

4. The method for secure storage of geological data based on blockchain according to claim 1, characterized in that: The standardized geological environment data set in the blockchain basic network is partitioned into quadtrees according to geographic grids, and the partitioned data is stored in proximity blocks through data association evaluation to obtain distributed data blocks, including: Defining the spatial scope of the standardized geological environment dataset in the blockchain basic network, demarcating the boundaries of the study area according to geographic coordinates, dividing the regional geomorphic units according to terrain features, and obtaining regional scope data; Gridding the study area according to the regional range data, evenly dividing the area through the geographic grid, hierarchically subdividing the geographic grid according to the quadtree partitioning method, and obtaining quadtree partitioning data; The spatial relationship of partitions is calculated by using the quadtree partition data, adjacent partitions are marked according to spatial positions, and the degree of association between partitions is quantified according to geographic distance to obtain spatial relationship data; Analyze the partition data content based on the spatial relationship data, calculate the partition content similarity through data association evaluation, quantify the data association according to the attribute association coefficient, and obtain association evaluation data; The partition data is merged into blocks according to the association evaluation data, the block range is defined by a proximity threshold, and the size of the block is controlled according to the data block size to obtain block storage data; Node allocation is performed on the block storage data, storage nodes are selected through load balancing, and backup nodes are determined through a data redundancy strategy to obtain the distributed data block.

5. The method for secure storage of geological data based on blockchain according to claim 1, characterized in that: The three-layer access permission structure is constructed based on the distributed data block, the permission inheritance relationship is calculated by quantifying the data sensitivity, the user permission is verified through the environmental parameter constraint, and the access control chain is obtained, including: Based on the distributed data blocks, the management levels are divided, the access rights are graded according to the importance of the data, and the user groups are divided according to the permission levels to obtain a three-layer access permission structure; wherein the three-layer access permission structure includes an administrator permission layer, a business user permission layer, and a general user permission layer; Analyze the data content in the three-layer access permission structure, assess the data value according to the frequency of keyword occurrence, divide the data sensitivity level according to the business type, and obtain sensitivity quantitative data; The permission transfer relationship is sorted out through the sensitivity quantified data, the permission inheritance path is set through the user group level, and the permission transfer rules are limited according to the parent-child node relationship to obtain the permission inheritance data; The permission inheritance data is used to calculate the permission group to which the user belongs, the access level is determined according to the user identity type, the user operation scope is constrained according to the permissions within the group, and permission calculation data is obtained; Performing environmental detection on the authority calculation data, limiting the operation window according to the access time, and constraining the login scope according to the access location to obtain environmental constraint data; The user authority status is judged according to the environmental constraint data, the legitimacy of the user is verified through identity authentication, and the access request is authorized through authority evaluation to obtain the access control chain.

6. The method for secure storage of geological data based on blockchain according to claim 1, characterized in that: The nodes are grouped and numbered according to the access control chain, data consistency is verified by polling and voting, and the backup data is backed up by a data verification code to obtain a verification record, including: Evaluate the node service capability according to the access control chain, group and divide the nodes according to processing performance, and sort the grouped nodes by number according to network distance to obtain node grouping data; The voting rules are set through the node grouping data, the voting order is arranged through the node response time, the node decision importance is distinguished according to the voting weight, and the polling voting data is obtained; Compare the storage content of each node based on the polling voting data, verify the data content through the data block feature value, mark the data difference according to the consistency judgment rule, and obtain consistency verification data; The data to be backed up is screened by the consistency verification data, the backup priority is sorted by the importance of the data, the backup content is confirmed according to the data integrity, and a list to be backed up is obtained; Redundantly encode the data in the to-be-backed-up list, divide the backup data into blocks by data sharding, calculate the data check code by check bit generation, and obtain data check data; The data backup operation is recorded via the data verification data, the data version is identified by the backup time, and the data copy is located according to the backup location to obtain the verification record.

7. The method for secure storage of geological data based on blockchain according to claim 1, characterized in that: The verification records are analyzed in time series, the operation modes are classified by extracting behavioral features, abnormal behaviors are determined by access frequency statistics and spatiotemporal feature matching, and an audit report is obtained, including: The verification records are sorted according to timestamps, the verification records are divided into intervals by time windows, and the data sequence is segmented according to operation continuity to obtain time series analysis data; Extracting user operation behaviors based on the time series analysis data, classifying the behavior sequences according to the operation types, describing the operation processes through data access paths, and obtaining behavior feature data; Clustering user behavior patterns through the behavior feature data, summarizing behavior rules according to operation sequence, grouping operation patterns through similarity calculation, and obtaining operation pattern data; The access behavior is counted according to the operation mode data, the number of accesses per unit time is counted, and the operation frequency is analyzed according to the access target to obtain the access frequency data; The user operation locations are distributed and counted based on the access frequency data, the access locations are marked by geographic coordinates, and the spatiotemporal features are associated according to the access time periods to obtain spatiotemporal feature data; The spatiotemporal feature data is compared for anomalies, abnormal operations are identified through normal behavior templates, abnormal behaviors are graded according to risk levels, and the audit report is obtained.

8. The method for secure storage of geological data based on blockchain according to claim 7, characterized in that: The user behavior patterns are clustered via the behavior feature data, the behavior rules are summarized according to the operation sequence, and the operation patterns are grouped by similarity calculation to obtain the operation pattern data, including: Sequentially segmenting the behavior feature data, dividing the behavior segments according to the operation time interval, marking the behavior sequence according to the minimum operation unit, and obtaining behavior sequence data; Calculating the operation correlation degree through the behavior sequence data, performing correlation analysis on the continuous behaviors through the operation combination, quantifying the behavior relationship according to the operation co-occurrence frequency, and obtaining behavior correlation data; Based on the behavior association data, a behavior distance matrix is ​​constructed, the behavior sequence similarity is evaluated by editing distance calculation, and the behavior patterns are merged according to the distance threshold to obtain behavior distance data; Performing hierarchical clustering on the behavior distance data, dividing the behavior categories by inter-cluster distance, and classifying the user behavior patterns according to the cluster centers to obtain behavior clustering data; Extracting behavior feature vectors according to the behavior clustering data, describing the behavior law through operation transition probability, characterizing the behavior pattern according to the time sequence relationship, and obtaining behavior law data; The operation modes are grouped according to the behavior rule data, similar behaviors are merged according to a similarity threshold, and the operation modes are identified according to the behavior characteristics to obtain the operation mode data.

9. A blockchain-based geological data security storage system, used to implement the blockchain-based geological data security storage method as described in any one of claims 1 to 8, characterized in that: The blockchain-based geological data security storage system includes: The acquisition module is used to collect data on the topography, mining environment, hydrological environment and soil environment of the study area, establish spatiotemporal correlation mapping of the collected data through zoning geocoding, and standardize the mapped data through spatial reference conversion to obtain a standardized geological environment data set; A construction module is used to perform hierarchical credit assessment and dynamic load distribution on network nodes according to the standardized geological environment data set, and to construct a network topology of nodes through a minimum spanning tree algorithm to obtain a blockchain basic network; A partitioning module, used to perform quadtree partitioning on the standardized geological environment data set in the blockchain basic network according to the geographic grid, and perform proximity group storage on the partitioned data through data association evaluation to obtain distributed data blocks; A calculation module is used to construct a three-layer access permission structure based on the distributed data block, calculate the permission inheritance relationship through data sensitivity quantification, verify the user permission through environmental parameter constraints, and obtain an access control chain; A backup module, used to group and number nodes according to the access control chain, verify data consistency through polling and voting, and back up the data to be backed up through a data check code to obtain a verification record; The classification module is used to perform time series analysis on the verification records, classify the operation modes by extracting behavioral features, determine abnormal behaviors by access frequency statistics and spatiotemporal feature matching, and obtain an audit report.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the blockchain-based geological data secure storage method as described in any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Block chain-based geological data secure storage method and system

    CN118586043A

  • Enterprise sensitive data security access management method and system

    CN118656870A

  • Block chain-based geological data security sharing system and method

    CN118965413A

Cited By

  • Agricultural production process block chain evidence storage method and system

    CN121071940A

  • Trusted terminal network security protection method and system

    CN121262002A

  • Geographic information data analysis method and system based on block chain and federal learning, electronic equipment and storage medium

    CN121388039A

  • Mobile terminal off-line storage method and system based on mountainous area

    CN121418941A