A data quality evaluation method and system under a distributed storage architecture

By combining a distributed storage architecture and a three-dimensional sharding strategy with a federated learning differential privacy framework and blockchain technology, the problems of data silos and multimodal data fusion in waterway port facilities have been solved, enabling real-time quality assessment and trend prediction, and improving the data processing efficiency and assessment accuracy of port facilities.

CN122045184BActive Publication Date: 2026-07-24TIANJIN RES INST FOR WATER TRANSPORT ENG M O T
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIANJIN RES INST FOR WATER TRANSPORT ENG M O T
Filing Date
2026-04-16
Publication Date
2026-07-24

Smart Images

  • Figure CN122045184B_ABST
    Figure CN122045184B_ABST
Patent Text Reader

Abstract

The application provides a data quality evaluation method and system under a distributed storage architecture, the method comprising: collecting structured data sets and unstructured data fingerprints of water transportation port facilities; storing the structured data according to facility types, spatial grids and time windows; training GNN and a Transformer model based on the unstructured data fingerprints and the stored data, and generating an evaluation result verification proof through a blockchain and federated learning; constructing a real-time three-dimensional twin body fused with BIM and GIS, integrating data and performing dynamic quality evaluation, and generating facility quality scores, defect heat maps and trend prediction results; and automatically generating a repair plan based on the results. The application improves data processing efficiency through distributed storage and three-dimensional fragmentation, realizes data fusion through a federated learning differential privacy framework, enhances evaluation accuracy through spatio-temporal correlation analysis and a hybrid prediction model, and realizes verifiable evaluation and intelligent repair.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data quality assessment technology, and in particular to a data quality assessment method and system under a distributed storage architecture. Background Technology

[0002] Currently, the quality assessment of waterway port facilities data mainly adopts a traditional model combining centralized storage architecture and manual inspection. In practice, raw data (such as sensor monitoring values, video surveillance streams, equipment operation and maintenance records, etc.) is usually stored in a single database or file system, and data quality is judged through periodic sampling inspections or human experience. For example, terminal structural health monitoring relies on manual comparison of design parameters with sensor readings, channel dredging data is assessed for completeness through offline statistical mean and variance, and security monitoring videos require manual annotation of abnormal events. In terms of data fusion, simple splicing methods are often used to process structured and unstructured data, lacking spatiotemporal correlation analysis; the quality assessment dimensions are limited to single-point accuracy verification, without systematically integrating multi-dimensional indicators such as completeness and timeliness. Although some advanced scenarios have introduced basic machine learning models, they are mostly static threshold alarms or univariate predictions, which are difficult to adapt to the dynamic characteristics of port facilities.

[0003] Existing technologies have significant limitations: First, centralized storage is prone to creating data silos, resulting in low efficiency in cross-node collaboration and making it difficult to support the real-time data processing needs of large-scale distributed port facilities; Second, the ability to fuse multimodal data is weak, and the semantic parsing of unstructured data (such as monitoring videos and work order texts) relies on manual intervention, leading to low efficiency in extracting spatiotemporal correlation features; Third, the generalization ability of quality assessment models is insufficient, and traditional statistical methods or shallow machine learning models are unable to capture complex spatiotemporal correlation patterns such as wharf settlement and channel displacement, resulting in delayed or misjudged assessment results. Summary of the Invention

[0004] This invention aims to at least address the technical problems existing in the prior art, and in particular, innovatively proposes a data quality assessment method and system under a distributed storage architecture.

[0005] To achieve the above-mentioned objectives of this invention, this invention provides a data quality assessment method under a distributed storage architecture, characterized in that the method includes: S1. Collect multimodal raw data of waterway port facilities and perform real-time preprocessing to generate preprocessed structured datasets and unstructured data fingerprints; S2. Based on the preprocessed structured dataset, dynamically segment and store the data according to the facility type, spatial grid, and time window three-dimensional segmentation strategy to obtain segmented storage data. S3. Based on the unstructured data fingerprint and sharded storage data, the quality assessment rules are encoded into an executable contract on the blockchain, and the GNN spatiotemporal correlation analysis model and the Transformer semantic parsing model are trained across nodes through a federated learning differential privacy collaborative training framework. The evaluation results are verified by using zk-SNARKs zero-knowledge proof. S4. Based on the evaluation results, verify and prove that a real-time three-dimensional twin of the water transport facility integrating BIM+GIS is constructed, integrating the segmented storage data and real-time monitoring stream data, and performing multi-dimensional dynamic quality evaluation through a dynamic quality evaluation sandbox to generate facility quality scores, defect distribution heat maps and trend prediction results. S5. Based on the facility quality score, defect distribution heat map and trend prediction results, compare the BIM model with the real-time monitoring data to automatically generate a repair plan.

[0006] In another aspect, the present invention also provides a data quality assessment system under a distributed storage architecture, the system including a processor and a memory for storing processor-executable instructions; wherein the processor is configured to implement the data quality assessment method under the distributed storage architecture when executing the executable instructions.

[0007] The beneficial effects of this invention are as follows: This invention overcomes the limitations of data silos through a distributed storage architecture and a three-dimensional sharding strategy, enabling real-time collaborative processing across nodes. It dynamically shards and stores data according to facility type, spatial grid, and time window, supporting parallel processing of multi-dimensional data such as sliding windows for wharf structures and fixed windows for waterway dredging, significantly improving the efficiency of real-time data processing for large-scale port facilities. Addressing the weakness in multimodal data fusion, it employs a federated learning differential privacy framework to collaboratively train a GNN spatiotemporal correlation model and a Transformer semantic parsing model, automatically extracting spatial correlation features of wharf pile foundation settlement and parsing semantic entities from security monitoring text. By integrating unstructured data fingerprinting technology to achieve deep fusion of structured and unstructured data, manual intervention is reduced and the efficiency of spatiotemporal feature extraction is improved. Through GNN-TCN spatiotemporal correlation analysis and Prophet-ARIMA hybrid prediction model, the generalization ability of the quality assessment model is enhanced. It can capture complex spatiotemporal correlation patterns such as wharf structure settlement and waterway displacement, and can also achieve accurate trend prediction through seasonal identification and confidence interval calculation. This effectively solves the problem of misjudgment and lag in traditional statistical methods. Finally, dynamic assessment results including facility quality scores, defect heat maps and trend predictions are generated, realizing the verifiability of assessment results and intelligent generation of remediation plans.

[0008] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0009] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a flowchart of a data quality assessment method under a distributed storage architecture according to the present invention. Detailed Implementation

[0010] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0011] Example 1 like Figure 1 As shown, a data quality assessment method under a distributed storage architecture is characterized in that the method includes: S1. Collect multimodal raw data of waterway port facilities and perform real-time preprocessing to generate preprocessed structured datasets and unstructured data fingerprints; In step S1, it is necessary to explain in detail that the multimodal raw data covers monitoring data of various types of facilities such as wharves, waterways, port machinery, and security monitoring. Specifically, it includes: stress and strain sensor data of wharf structures (such as pile foundation strain values ​​and beam deflection data), displacement monitoring data (such as settlement and horizontal displacement values); waterway depth measurement data, mud surface elevation data, and water flow velocity and direction data; port machinery operating status data (such as crane lifting capacity, working radius, slewing angle, and operating speed), electrical parameter data (such as motor current, voltage, and power), and fault alarm data; security monitoring video stream data, image data (such as facial recognition images and abnormal behavior capture images), and access control record data; as well as unstructured text data such as facility design drawings, maintenance work orders, and inspection reports. The real-time preprocessing process first cleans the structured data, using the Laida criterion (3σ principle) to remove outliers from sensor data such as stress, strain, and displacement. Missing data is filled using time-series-based interpolation algorithms (such as linear interpolation and spline interpolation). At the same time, the data format and units are standardized, and the sampling frequency of different sensors is standardized to the minute level. For unstructured data, keyframes are extracted from video stream data (using inter-frame difference method combined with SIFT feature matching to determine keyframes), image data is normalized and grayscale is processed, and text data is segmented, stop words are removed, and TF-IDF feature vectorization is performed. This process generates a preprocessed structured dataset and an unstructured data fingerprint containing data type and feature summary information.

[0012] S2. Based on the preprocessed structured dataset, dynamically segment and store the data according to the facility type, spatial grid, and time window three-dimensional segmentation strategy to obtain segmented storage data. S3. Based on the unstructured data fingerprint and sharded storage data, the quality assessment rules are encoded into an executable contract on the blockchain, and the GNN spatiotemporal correlation analysis model and the Transformer semantic parsing model are trained across nodes through a federated learning differential privacy collaborative training framework. The evaluation results are verified by using zk-SNARKs zero-knowledge proof. S4. Based on the evaluation results, verify and prove that a real-time three-dimensional twin of the water transport facility integrating BIM+GIS is constructed, integrating the segmented storage data and real-time monitoring stream data, and performing multi-dimensional dynamic quality evaluation through a dynamic quality evaluation sandbox to generate facility quality scores, defect distribution heat maps and trend prediction results. S5. Based on the facility quality score, defect distribution heat map and trend prediction results, compare the BIM model with the real-time monitoring data to automatically generate a repair plan.

[0013] Optionally, obtaining fragmented storage data includes: S201. Based on the facility type dimension, the preprocessed structured dataset is divided into four major infrastructure data modules: wharf structure, waterway dredging, port machinery, and security monitoring. Each infrastructure data module is associated with the corresponding facility ID code and industry standard parameter threshold. In step S201, the classification of facility types is based on the core functional modules of waterway port facilities. The wharf structure data module covers monitoring data of the main structures such as wharf piles, beams, and panels, and is associated with the pile bearing capacity thresholds specified in the "Port Engineering Pile Foundation Code" (such as the standard value of the vertical ultimate bearing capacity of a single reinforced concrete cast-in-place pile) and the structural safety level parameters in the "Unified Standard for Reliability Design of Port Engineering Structures." The waterway dredging data module contains data related to navigation conditions, such as waterway depth and mud surface elevation, and is associated with the minimum water depth requirements for different grades of waterways in the "Waterway Engineering Design Code" (such as a minimum water depth of not less than 4.0m for Class I waterways navigable by 3000-ton vessels) and dredging data. The permissible deviation value for the flatness of dredging projects; the port machinery data module integrates the operating parameters of machinery such as cranes and conveying equipment, the working level classification standards in the "Crane Design Specification" (such as A1 to A8 levels) and the motor temperature rise limit in the electrical safety regulations (such as the temperature rise of F-class insulated motors not exceeding 155K); the security monitoring data module includes security-related data such as video streams and access control records, and is associated with the video image quality indicators in the "Technical Requirements for Information Transmission, Exchange and Control of Security Video Surveillance Network System" (such as resolution not less than 1920×1080 pixels, frame rate not less than 25fps) and the response time threshold of the access management system (such as identity verification response time not greater than 1 second). Each data module is associated with a physical entity through a facility ID code. The coding rule uses an 18-bit string, with the first 6 bits being the facility type code (e.g., "MT0001" for wharf structures and "HD0002" for dredging channels), the middle 8 bits being the spatial location code (generated based on Gauss-Kruger projection coordinate transformation), and the last 4 bits being the equipment serial number, ensuring the uniqueness of data traceability.

[0014] S202. Using spatial grid partitioning rules, spatial coordinate mapping is performed on each infrastructure data module to generate a spatial index table containing grid number, facility center point coordinates, and neighborhood correlation degree. In step S202, this embodiment adopts a regular grid division method based on the UTM projection coordinate system, spatially gridding the port area into 10m × 10m cells. The grid number adopts the string format of "X-axis coordinate - Y-axis coordinate", for example, "352000-4485000" represents the grid cell at X=352000m and Y=4485000m in the UTM coordinate system. For the facility entities in each infrastructure data module, the three-dimensional coordinates of their center points are extracted from their design drawings or real-time monitoring data, and mapped to the corresponding grid cells through a coordinate transformation algorithm to establish the membership relationship between the facility entities and the grid cells. At the same time, the Euclidean distance between the center point of each facility entity and the facility entities in the surrounding 8 adjacent grids is calculated, and a neighborhood correlation matrix is ​​generated by the inverse distance weighting method. The matrix element values ​​range from [0,1], and the larger the value, the stronger the spatial correlation. Finally, a spatial index table is formed that includes the grid number, the three-dimensional coordinates of the facility center point (accurate to the centimeter level), the neighborhood correlation matrix, and the corresponding facility ID list.

[0015] S203. Based on the spatial index table, a dynamic time window mechanism is set up. The data of wharf structure is set up using an a-minute sliding window, the data of waterway dredging is set up using a b-minute fixed window, the data of port machinery is set up using a c-minute real-time window, and the data of security monitoring is set up using a d-minute aggregated window. Each time window data is accompanied by a timestamp and a data collection frequency label. In step S203, the parameter settings of the dynamic time window mechanism need to be combined with the timeliness requirements and change frequency of different facility types of data. For wharf structure data, parameters such as settlement and stress change slowly but require long-term monitoring. A 60-minute sliding window (i.e., a=60) is used, with the window sliding forward every 15 minutes to ensure that it can capture slow changing trends while avoiding data redundancy. Waterway dredging data is related to the navigation cycle, so a 1440-minute (24-hour) fixed window (i.e., b=1440) is used. A slice of dredging data for the previous 24 hours is generated at 0:00 every day, which meets the daily cycle management needs of waterway maintenance. Port machinery data, such as crane operating parameters, needs to be monitored in real time. A 5-minute real-time window (i.e., c=5) is used, with the window continuously scrolling without overlap to ensure the timeliness of fault warnings. In security monitoring data, video stream keyframe extraction and access control records have periodic patterns. A 30-minute aggregation window (i.e., d=30) is used to sort the unstructured data fingerprints in the window by timestamp and generate feature summaries to reduce storage pressure. Each time window of data is accompanied by a UTC timestamp accurate to the millisecond level, and is labeled with the data collection frequency (e.g., "1 time / minute" for wharf structures, "10 times / second" for port machinery).

[0016] S204. By cross-linking the infrastructure data module, spatial index table, and time window data through a three-dimensional sharding algorithm, sharded data units containing facility type, spatial grid, and three-dimensional time window identifiers are generated, and sharded storage data is obtained. Each sharded data unit contains a data check code, storage node address, and data update version number.

[0017] In step S204, the 3D sharding algorithm first establishes a 3D Cartesian product index system of facility type, spatial grid, and time window, linking the four infrastructure data modules from step S201, the spatial index table from step S202, and the dynamic time window data from step S203 in multiple dimensions. Specifically, a three-level hash index table is constructed using the facility type code as the first-level index key, the spatial grid number as the second-level index key, and the time window start timestamp as the third-level index key. For each specific data record within an infrastructure data module, the corresponding spatial grid number is parsed from its associated facility ID code, and combined with the time window to which the data acquisition timestamp belongs, the unique position of the record in the 3D index system is determined. Subsequently, the data set under each 3D identifier is sharded and packaged to generate sharded data units containing data volumes, metadata, and verification information. The data verification code is generated by hashing the fragmented data body using the SHA-256 algorithm to ensure data integrity. The storage node address is calculated based on the three-dimensional identifier using a consistent hashing algorithm to achieve uniform data distribution among distributed nodes. The data update version number uses an incrementing sequence based on timestamps, automatically incrementing by 1 with each data update for conflict detection and version control. The final generated fragmented storage data is stored in the distributed database in key-value pairs. The key is a three-dimensional identifier string (formatted as "facility type code-spatial grid number-time window start timestamp"), and the value is a fragmented data unit containing the data verification code, storage node address, data update version number, and actual data record.

[0018] Optionally, the verification proof for generating the evaluation results includes: S301. Based on the unstructured data fingerprint, a three-segment data fingerprint containing data type identifier, hash value, and timestamp is generated using the SHA-3 algorithm, and the fingerprint hash value and access address are anchored on the blockchain to generate an on-chain data fingerprint anchoring result. In step S301, the generation process of unstructured data fingerprints must ensure uniqueness and immutability. First, core features are extracted for different types of unstructured data (video keyframes, images, and text): video keyframes are generated into 512-dimensional feature vectors using the ORB feature point extraction algorithm; image data is processed through the penultimate layer of the ResNet-50 network to output 2048-dimensional feature vectors; and text data is encoded into 768-dimensional feature vectors using the BERT model. Each feature vector is assigned a 1-byte type identifier according to its data type (video / image / text) (0x01 for video keyframes, 0x02 for images, and 0x03 for text). Then, a 64-byte SHA-3-512 hash value (calculated by hashing the original feature vectors) and an 8-byte UTC timestamp (accurate to the second) are concatenated to form a three-segment data fingerprint structure (total length 1 + 64 + 8 = 73 bytes). Subsequently, the data fingerprint is written to the blockchain (using the Hyperledger Fabric consortium blockchain) via a smart contract call interface. The blockchain stores the fingerprint hash value (the 73-byte fingerprint is hashed again using SHA-256) and the access address of the corresponding unstructured data in the distributed storage system (formatted as "node IP:port / data UUID"). This generates an on-chain data fingerprint anchoring result containing the block height, transaction ID, fingerprint hash value, and access address, thus achieving traceability and tamper-proofing of unstructured data.

[0019] S302. Based on the on-chain data fingerprint anchoring results, the quality assessment rules are encoded into an executable contract on the blockchain, and on-chain executable logic for accuracy verification, integrity verification, and timeliness monitoring is defined to generate a set of on-chain contracts for quality assessment rules. In step S302, the coding process of the on-chain contract set for the quality assessment rules needs to combine the characteristics of waterway facility data with industry standards, and the smart contract should be developed using the Solidity language. First, the accuracy verification logic defines comparison rules between sensor data and industry standard thresholds for structured data. For example, in wharf structure data, when the pile foundation stress value exceeds 80% of the ultimate bearing capacity standard value specified in the "Port Engineering Pile Foundation Code," an early warning flag is automatically triggered. If the waterway depth data is lower than the minimum water depth requirement of the "Waterway Engineering Design Code" for three consecutive time windows, the contract performs accuracy downgrading processing. For unstructured data, accuracy verification is achieved by comparing the cosine similarity between the on-chain anchored feature fingerprint and the real-time extracted fingerprint. When the ORB feature vector similarity of the video keyframe is lower than 0.9, it is determined to be data distortion. The integrity verification logic includes data coverage and correlation checks. Coverage is calculated by comparing the actual amount of data collected with the theoretically required amount of data collected in each facility type's data module (e.g., port machinery data needs to cover more than 95% of the operating parameter sensors). Correlation is checked using the neighborhood correlation matrix in the spatial index table to verify the synchronization of related facility data (e.g., wharf panel displacement data and adjacent pile foundation stress data need to be collected within the same time window). The timeliness monitoring logic sets timeout thresholds based on the time window parameters of different data modules. For example, if security monitoring data has not completed feature extraction and fingerprint anchoring within a 30-minute aggregation window, the contract automatically records the data delay timestamp and updates the quality score decay coefficient. The contract set also includes a rule dynamic update interface, allowing authorized nodes (e.g., port management departments) to modify the evaluation threshold through a multi-signature mechanism. Each update generates a new contract version and associates it with the old version's hash value, ensuring the traceability of rule iterations. The final generated on-chain quality assessment rule contract set includes a main contract (responsible for rule scheduling) and four sub-contracts (corresponding to the four major infrastructure data modules). Each contract pushes intermediate evaluation results to the off-chain system through an event mechanism.

[0020] S303. Construct a federated learning differential privacy collaborative training framework based on the on-chain contract set of the quality assessment rules; In step S303, the federated learning framework constructed in this embodiment adopts a two-tier architecture of "central server - edge nodes". The central server is deployed in the port data center and is responsible for the aggregation and distribution of global model parameters; the edge nodes are distributed in various infrastructure monitoring areas (such as dock operation areas, waterway monitoring stations, and machinery control centers), retaining local data and performing model training. The core components of the framework include a differential privacy module, a model training module, a parameter exchange module, and a security authentication module. The differential privacy module uses a Laplacian mechanism to add noise to the model gradients uploaded by edge nodes, with the noise intensity ε set to 0.1. The gradient norm is controlled within the range of [0,10] using a Clip operation to prevent gradient leakage of original data features. The model training module supports multi-task parallel training, allocating independent training processes for the GNN spatiotemporal correlation analysis model and the Transformer semantic parsing model, and employing the Adam optimizer (initial learning rate 0.001, weight decay coefficient 1e-5). The parameter exchange module implements encrypted parameter transmission based on the Secure Multi-Party Computation (SMPC) protocol, using the Elliptic Curve Cryptography (ECC) algorithm to encrypt gradient parameters. After receiving the encrypted gradients, the central server performs aggregation through homomorphic addition, and then encrypts and distributes the aggregated global parameters to each edge node. The security authentication module adopts a node authentication mechanism based on digital certificates. Each edge node must submit an identity certificate issued by the port CA to the central server. Only after verification through the certificate chain can it join the federated learning network, preventing malicious nodes from participating in training. The framework also features a dynamic node addition / exit mechanism. When a new monitoring node is added, the current global model parameters are automatically synchronized and the local training state is initialized. When a node is offline for more than two training cycles (each cycle is 24 hours), the central server temporarily removes it from the federated node list. Training is resumed through an incremental synchronization mechanism after reconnection.

[0021] S304. Based on the fragmented storage data, a GNN spatiotemporal correlation analysis model is trained using a federated learning differential privacy collaborative training framework, and spatial correlation features of settlement between wharf pile foundation and breakwater are extracted. The TCN network is combined to analyze the time series trend of channel displacement and generate spatiotemporal correlation analysis model output. In step S304, the input data for the GNN spatiotemporal correlation analysis model consists of structured monitoring data closely related to spatial location from the fragmented storage data generated in step S204. This primarily includes time-series data such as wharf pile settlement, stress values, inclination angles, breakwater settlement, horizontal displacement, and channel bottom elevation changes. The model architecture employs a spatiotemporal graph convolutional network (ST-GCN). In the spatial dimension, a graph structure based on the topological relationships of facility entities is constructed. The nodes of the graph are the monitoring points of each infrastructure (such as specific pile numbers and breakwater monitoring section numbers). The edge weights are determined by the neighborhood correlation matrix generated in step S202; that is, the larger the weighted average of the inverse Euclidean distance between two monitoring points in space, the higher the edge weight, indicating a stronger spatial correlation. In the temporal dimension, the monitoring data within each time window (such as a 60-minute sliding window for wharf structures) is taken as a time step and input into the temporal convolutional layer. The specific training process is as follows: First, the fragmented storage data is preprocessed, including outlier detection (using the 3σ rule) and imputation (using K-nearest neighbor-based interpolation) to ensure the integrity of the input data. Then, the preprocessed data is divided into training and validation sets in an 8:2 ratio. For spatial feature extraction, a two-layer graph convolutional network (GCN) is used, with each layer containing 64 hidden units and the activation function being ReLU. By aggregating the features of each node and its neighboring nodes, the spatial dependency between the wharf pile foundation and the breakwater is captured. For example, uneven settlement of the pile foundation in a certain area may lead to stress concentration in adjacent breakwaters. For temporal feature extraction, a three-layer temporal convolutional network (TCN) is used, with kernel sizes of 3, 5, and 7, and a stride of 1. Causal convolution ensures that the model only uses historical data for prediction, focusing on analyzing the changing trends of channel displacement at different time scales, such as short-term fluctuations caused by tides and long-term settlement trends caused by siltation. The loss function of the model is mean squared error (MSE), and the optimizer is the Adam optimizer defined in step S303. Within the federated learning framework, each edge node (such as the wharf area edge node and the breakwater monitoring edge node) trains its model using local sharded data. After every 5 epochs of local training, the model gradients with added differential privacy noise are uploaded to the central server. The central server aggregates the gradients of all edge nodes (using a weighted average, with weights equal to the proportion of data volume for each node), updates the global model parameters, and distributes the updated parameters back to each edge node to begin the next round of local training. This process iterates for 50 epochs until the model's MSE converges on the validation set (a change of less than 1e-4 for 5 consecutive epochs). After training, the GNN spatiotemporal correlation analysis model can output a spatial correlation feature vector (128 dimensions, including correlation strength, influence range, etc.) between the wharf pile foundation and breakwater settlement, as well as a time-series trend prediction curve for channel displacement over the next three time windows.

[0022] S305. Based on the unstructured data fingerprint, a Transformer semantic parsing model is trained using a federated learning differential privacy collaborative training framework. Semantic entity recognition and relation extraction are performed on the text descriptions of port security monitoring video frames and the unstructured text of equipment fault work orders to generate semantic parsing model output. In step S305, the input data for the Transformer semantic parsing model is the original unstructured text data associated with the unstructured data fingerprint generated in step S301. This mainly includes manually annotated text descriptions of port security monitoring video frames (e.g., "abnormal swing of the crane boom in Terminal A area"), descriptions of fault phenomena in equipment fault work orders (e.g., "abnormal noise from the gantry crane's traveling mechanism, startup delay"), and natural language text in equipment maintenance records. The model architecture adopts a bidirectional encoder structure based on a BERT pre-trained model, performing domain-adaptive fine-tuning on the BERT-base model (12-layer Transformer, 768-dimensional hidden states, 12 attention heads). The specific training process is as follows: First, the unstructured text data is preprocessed, including Chinese word segmentation (using the Jieba word segmentation tool), stop word removal (based on the port industry stop word list), and entity annotation (using the BIO annotation system, defining 12 types of entity labels such as equipment entities (e.g., "crane" and "gantry crane"), fault entities (e.g., "abnormal noise" and "swing"), and location entities (e.g., "terminal A area" and "traveling mechanism"). Subsequently, a training set containing 500,000 labeled samples was constructed (70% from historical work order data and 30% from manually labeled surveillance video description text), divided into training and validation sets in a 9:1 ratio. The semantic entity recognition (NER) task of the model uses a CRF layer as the output layer, maximizing the conditional probability of the label sequence to achieve accurate identification of entity boundaries and types. The relation extraction task adopts a sentence pair classification architecture, using the identified entity pairs as input and capturing semantic associations between entities through the Transformer's self-attention mechanism, outputting eight predefined relation categories such as "fault-equipment" and "fault-location". Under the federated learning framework, edge nodes responsible for processing security video text (such as the dock monitoring center node) and edge nodes processing work order text (such as the equipment management center node) are trained using local data. To meet the privacy protection requirements of unstructured text, the differential privacy module adds Gaussian noise (noise standard deviation σ=0.05) to the attention weight parameters of the Transformer model and prevents gradient information leakage through gradient pruning (pruning threshold set to 5.0). After each edge node completes 10 epochs of local training, it uploads the encrypted model parameters (encrypted using the ECC algorithm) to the central server. The central server aggregates the parameters using a federated averaging (FedAvg) algorithm, assigning higher aggregation weights to nodes with larger amounts of text data (such as work order processing nodes) (weight coefficient = node sample size / total sample size). After the global model parameters are updated, they are distributed to each edge node for further training. After 30 iterations, the entity recognition F1 score on the validation set (entity recognition F1 = 2) is used as the benchmark. Accuracy Training stops when the ratio of recall to (precision + recall) stabilizes above 0.92. After training, the Transformer semantic parsing model can output structured semantic triples (such as (crane, fault location, boom), (gantry crane, fault phenomenon, abnormal noise)) and sentiment scores of the text description (used to judge the urgency of the fault, ranging from 0 to 1, the closer to 1, the higher the urgency).

[0023] S306. Integrate the output of the spatiotemporal correlation analysis model and the output of the semantic parsing model, and perform multi-dimensional quality assessment through the on-chain contract set of quality assessment rules to generate preliminary assessment results including data accuracy score, integrity index, and timeliness deviation value. In step S306, the multi-dimensional quality assessment process employs a three-stage workflow of "feature fusion - rule matching - quantitative scoring". First, feature fusion is performed, mapping the 128-dimensional spatial correlation feature vector output by the GNN spatiotemporal correlation analysis model to the semantic triples output by the Transformer semantic parsing model. This constructs a fused feature matrix containing spatiotemporal features (such as pile settlement correlation strength and channel displacement trend slope) and semantic features (such as fault entity confidence and relation extraction accuracy). During the fusion process, an attention mechanism is used to assign dynamic weights to different features. For example, when the "emergency fault" sentiment score is ≥0.8 in the semantic parsing results, the timeliness weight coefficient of the corresponding equipment monitoring data is automatically increased (adjusted from the base value of 1.0 to 1.5).

[0024] Next, the on-chain contract set of quality assessment rules calls the rule scheduling interface of the main contract, triggering the execution of sub-contracts in the order of "accuracy → completeness → timeliness". In the accuracy assessment phase, the main contract first calls the structured data sub-contract to compare the pile foundation stress value predicted by the GNN model with the threshold of the "Port Engineering Pile Foundation Code". If the predicted value of a certain monitoring point exceeds the threshold by 80% for two consecutive time windows, the individual score is calculated according to the built-in scoring function of the contract (accuracy score = 100 - (percentage exceeding the threshold × 100)). At the same time, the unstructured data sub-contract is called to calculate the cosine similarity between the real-time extracted video frame ORB feature fingerprint and the on-chain anchor fingerprint. When the similarity is ≥0.9, the score is 90 points. For every 0.01 decrease in similarity, 5 points are deducted. When it is below 0.8, the minimum score of 40 points is triggered. In the integrity assessment phase, the main contract calls the coverage verification sub-contract to calculate the ratio of the actual data collection volume to the theoretically required data collection volume for each data module. For example, a 95% coverage rate for port machinery data earns 100 points, with 2 points deducted for every 1% decrease. Simultaneously, the correlation verification sub-contract is triggered to check synchronization through a neighborhood correlation matrix. If the data collection time difference for related facilities exceeds 5 minutes, points are deducted proportionally to the time difference within the time window (e.g., a 10-minute difference within a 30-minute window deducts 33 points). In the timeliness assessment phase, the main contract calls the corresponding sub-contract based on the data type. For security monitoring data, fingerprint anchoring is completed within a 30-minute aggregation window, earning 100 points, with 2 points deducted for every minute of delay. For equipment work order data, a 15-minute response threshold is set, and after a timeout, the deviation value is calculated using the exponential decay formula (timeliness deviation = delay minutes × 0.05).

[0025] Finally, the contract set performs a weighted fusion of the three dimensions of scoring, with the weights allocated as follows: accuracy 40%, completeness 35%, and timeliness 25%, generating preliminary evaluation results. The results include specific scores (out of 100, with 60 or above considered passing), detailed deductions for each dimension (e.g., "Accuracy deduction 15 points: pile foundation stress exceeds threshold by 15%) and anomaly markers (e.g., "Completeness warning: channel depth data coverage is only 88%)".

[0026] S307. Using the zk-SNARKs zero-knowledge proof algorithm, with the preliminary evaluation result as input, construct circuit constraints and generate evaluation result verification proof.

[0027] In step S307, the core of the zk-SNARKs zero-knowledge proof algorithm lies in transforming the verification logic of the preliminary evaluation result into polynomial circuit constraints and generating a publicly verifiable proof. The specific implementation process is as follows: First, define the circuit's public input and private input. The public input includes the total score, scores for each dimension (accuracy score, completeness index, timeliness deviation value), and anomaly marker encoding from the preliminary evaluation result; the private input consists of intermediate calculation parameters generated during the contract set execution process in step S306 (such as threshold comparison results, similarity calculation values, coverage ratios, etc.).

[0028] The construction of circuit constraints must cover the core logic of the three evaluation dimensions. Taking accuracy evaluation as an example, the circuit needs to verify whether the scoring function of the structured data sub-contract is executed correctly: Let the predicted stress value be P, the specification threshold be T, and the percentage exceeding the threshold be R=(PT) / T×100% (when P>0.8T), then the accuracy score Sacc=100-R×100. The circuit needs to construct constraints R=(PT) / T×100% and Sacc=100-R×100, and ensure that P and T are non-negative real numbers, and R is in the range of [0,20] (corresponding to a score of 80-100 points). For the ORB feature fingerprint similarity Ssim of unstructured data, the circuit constraints are as follows: if Ssim ≥ 0.9, then the score Sorb = 90; if 0.8 ≤ Ssim < 0.9, then Sorb = 90 - 5 × (0.9 - Ssim) / 0.01; if Ssim < 0.8, then Sorb = 40. At the same time, it is necessary to verify whether the cosine distance formula for similarity calculation is correct.

[0029] The integrity assessment circuit needs to constrain the coverage ratio C = actual data collected / theoretical data collected, with a score Scom = 100 - 2 × (1 - C) × 100 (Scom = 100 when C ≥ 0.95, Scom = 60 when C < 0.8), and verify that the numerator and denominator of C are non-negative integers. In the correlation verification, the time difference Δt must satisfy Δt ≤ 5 minutes to avoid deduction; otherwise, a deduction D = (Δt / time window) × 100 is applied. The circuit needs to constrain the numerical relationship between Δt and the time window, and the calculation logic of D.

[0030] The timeliness assessment circuit sets constraints for different data types: security monitoring data delay △t1, score Stim1=100-2×△t1 (△t1≤30 minutes); equipment work order data delay △t2, timeliness deviation value Btim2=△t2×0.05. The circuit needs to verify the timing accuracy of △t1 and △t2 and the calculation rules of score / deviation value.

[0031] The circuit also needs to include a total score fusion constraint: Total = 0.4 × Sacc + 0.35 × Scom + 0.25 × (100 - Btim² × 200) (the timeliness deviation value is converted into a 0-100 score and then used for weighting), and ensures that the sum of the weights of each dimension is 1. In addition, the triggering conditions for anomaly marking (such as integrity warning when C < 0.9) also need to be converted into circuit constraints and the condition judgment is implemented through logic gates.

[0032] After completing the circuit construction, the Groth16 proof system is used to generate zero-knowledge proofs. The specific steps are as follows: 1. During the trusted setup phase, a common reference string (CRS) is generated based on the circuit, including the prover key and the verifier key; 2. Prover uses private inputs and CRS to generate proof π; 3. Verifier can verify whether the preliminary assessment results meet the preset quality assessment rules by only publicly providing the input and proof π, without needing to know the specific intermediate calculation process. This ensures the reliability of the assessment results while protecting data privacy. The proof generation time is controlled within 100ms, and the verification time is no more than 20ms, meeting the performance requirements of real-time assessment of port data.

[0033] Optionally, the generated facility quality score, defect distribution heatmap, and trend prediction results include: S401. Based on the evaluation results, verify and prove that a real-time three-dimensional twin integrating BIM and GIS is constructed; After verifying the validity of the evaluation results in step S401, a real-time 3D twin is constructed using the port facility's BIM (Building Information Modeling) model as the geometric core and data carrier, and GIS (Geographic Information System) as the spatial positioning and environmental framework. First, preprocessing is performed to integrate BIM and GIS data. The BIM model must include detailed geometric information (such as the diameter, material, and spatial coordinates of wharf piles, the structural layers and concrete strength grade of breakwaters, and the 3D models and key component parameters of equipment such as gantry cranes and hoists), attribute information (such as construction time, design service life, and maintenance records), and the correlation with monitoring sensors (such as the precise coordinates of sensor installation locations in the BIM model, and the mapping between sensor numbers and equipment IDs). The GIS data provides topographic data of the port area (a digital elevation model (DEM) with an accuracy of 0.5 meters), hydrological data (such as historical tide curves and water flow velocity vector maps), and administrative divisions and surrounding environmental information (such as channel boundaries, anchorage locations, and the distribution of nearby buildings). By employing coordinate system one (using the WGS84 coordinate system and employing a seven-parameter coordinate transformation model to accurately register the local coordinate system of the BIM model with the geodetic coordinate system of the GIS, with the registration error controlled within ±0.1 meters) and data format conversion (converting the BIM's IFC format file to the GIS-supported CityGML or 3DTiles format, preserving the model's topological relationships and attribute data), accurate overlay and spatial association of the BIM model in the GIS environment are achieved. Next, a real-time data access channel is established, pushing fragmented historical monitoring data (such as pile foundation settlement data and equipment vibration data from the past three months) and real-time monitoring stream data verified through quality assessment (such as current crane operation status data updated every second and video surveillance ORB feature fingerprints) to the 3D twin platform in real time via a data interface (such as the WebSocket protocol). The platform has a built-in data caching and update mechanism. Static attribute data (such as facility materials) is synchronized periodically (once a day), while dynamic monitoring data (such as stress and displacement) is processed in real time (delay ≤ 500ms). The data update priority is dynamically adjusted based on the timeliness deviation value in the data quality assessment results. For example, data with a deviation value ≥ 0.5 (corresponding to a delay ≥ 10 minutes) is marked as "low priority" and updates are temporarily suspended when system resources are tight.Finally, a visualization rendering engine is built to support the display of multi-level of detail (LOD) models. When the user zooms in on the view, the model precision is automatically switched (a simplified model is displayed in the distance, and a high-precision model containing details such as bolts and welds is displayed in the foreground). The status of the facilities is displayed intuitively through color coding, dynamic annotation and other methods. For example, piles with a settlement correlation strength greater than 0.8 output by the GNN model are highlighted in red, and "emergency fault" equipment identified by the Transformer semantic parsing model is marked with a flashing icon and semantic triple information (such as "gantry crane - fault phenomenon - abnormal noise") is displayed floating next to the model.

[0034] S402. Based on the real-time 3D twin, integrate the fragmented storage data and real-time monitoring stream data to generate an integrated dataset containing structured data indexes, unstructured data fingerprints, and real-time data stream processing rules. In step S402, the construction of the integrated dataset needs to achieve unified management and efficient access to historical and real-time data. For historical data stored in shards, a structured data indexing system is first established. B+ and tree indexes are used to index key fields such as device ID, monitoring timestamp, and data type. For example, pile foundation stress data is indexed by "device ID + date" to support querying stress change curves for single or multiple days by device. For unstructured data (such as video recordings and maintenance work order images), the ORB feature fingerprint and text semantic triple generated in step S306 are used as index keys and stored in a distributed hash table (DHT). Similar video clips or related work orders can be quickly located by fingerprint comparison. Real-time monitoring streaming data is indexed in real time using a streaming processing engine (such as Flink). For sensor data updated at the second level (such as crane amplitude and channel depth), real-time statistical features (mean, variance, peak value) are generated according to a sliding time window (5-minute window) and an in-memory index is established to ensure that the query response time is ≤100ms.

[0035] During data integration, unified data processing rules must be established: structured data needs to undergo format standardization (e.g., unifying the stress unit of different sensor manufacturers to MPa) and outlier filtering (filtering data with scores <60 based on the accuracy score in S306); unstructured data needs to be preprocessed, with keyframes extracted from video data and fingerprint chains generated (one ORB fingerprint generated every 30 seconds), and text data having semantic triples extracted using the Transformer model and associated with the corresponding device BIM model ID; real-time data streams need to have a priority processing mechanism set up, dynamically adjusting the processing queue according to the sentiment score in S306, with emergency fault data with a score ≥0.8 entering the high-priority queue (processing delay ≤1 second), and ordinary data entering the regular queue (processing delay ≤5 seconds). The final integrated dataset contains a three-level data directory: the first-level directory is divided by facility type (dock, waterway, equipment), the second-level directory is classified by data dimension (structured monitoring data, unstructured multimedia data, semantic parsing results), and the third-level directory uses index keys (such as device ID, timestamp, fingerprint hash) to achieve precise data location.

[0036] S403. Based on the integrated dataset, construct a dynamic quality assessment sandbox environment and generate a sandbox execution configuration that includes malicious behavior detection rules, code execution trajectory capture logic, and zero-knowledge proof verification circuit. In step S403, the construction of the dynamic quality assessment sandbox environment aims to provide an isolated, controllable, and traceable execution space for the quality assessment of the integrated dataset, preventing abnormal behavior during the assessment process from affecting the main system and ensuring the accuracy and security of the assessment logic. First, the sandbox environment uses lightweight virtualization technology (such as Docker containers) to build independent running instances. Each instance is allocated independent computing resources (CPU, memory, storage) and network space, and interacts with external systems through pre-defined secure interfaces, prohibiting direct access to underlying hardware or unauthorized system resources. The sandbox is pre-installed with the runtime environment required for assessment, including the zk-SNARKs proof verification library, the GNN model inference engine, the Transformer semantic parsing tool, and the streaming data processing framework. Version control mechanisms ensure the consistency and compatibility of each component version.

[0037] The rules for detecting malicious behavior need to cover potential risks such as data injection, logical tampering, and resource abuse. Specific rules include: data input validation rules, which validate the format of the integrated dataset entering the sandbox (e.g., validation of structured data field types, lengths, and value ranges; validation of the validity of unstructured data fingerprint hash values). If format anomalies are detected (e.g., negative stress values ​​or inconsistent fingerprint hash lengths), an alarm is triggered and the batch of data is rejected; behavior pattern recognition rules, which establish a baseline of normal behavior by analyzing the system call sequences of the evaluation process (e.g., file read / write, network connections, and memory usage). When frequent abnormal file access (e.g., attempts to read sensitive files outside the sandbox), abnormal network connections (sending large amounts of data to unknown IPs), or memory overflow occurs, it is determined as malicious behavior and the evaluation process is terminated; and access control rules, which grant processes within the sandbox only the minimum permissions required to complete the evaluation task, such as permission to read the integrated dataset and permission to execute the evaluation algorithm, prohibiting high-risk operations such as modifying system configurations and creating new processes.

[0038] The code execution trajectory capture logic records key execution steps during the evaluation process, enabling full traceability of the evaluation behavior. Specifically, a dynamic instrumentation tool is deployed in a sandbox environment to instrument the key function entry and exit points of the evaluation algorithm code (such as the GNN model inference function, the Transformer semantic parsing module, and the zero-knowledge proof verification circuit), capturing function call parameters, return values, execution time, and intermediate variable states. For example, when the GNN model calculates the settlement correlation strength, the input node feature matrix, adjacency matrix, output correlation strength value, and activation function output of each layer are recorded; during the zero-knowledge proof verification process, the public input, proof π, verification result, and verification time are recorded. The captured execution trajectory data is stored in an immutable log file in timestamp order. The log file uses a linked storage structure, with each log block containing the hash value of the previous block, ensuring that the log data is not tampered with. Simultaneously, the execution trajectory can be retrieved by evaluation task ID, time range, function name, and other dimensions, facilitating problem localization and auditing.

[0039] The integration of the zero-knowledge proof verification circuit ensures the reliability and privacy protection of the evaluation results. The verification circuit and corresponding verifier key generated in step S307 are pre-loaded in the sandbox environment. When the integrated dataset enters the sandbox, the verification circuit is first invoked to verify the evaluation result verification proof π attached to the data. The verification process strictly follows the verification procedure of the Groth16 proof system: the public input (total score, scores for each dimension, and anomaly marker encoding) and proof π are input into the verification circuit. The circuit performs logical verification according to preset constraints (such as the total score weighting formula and scoring rules for each dimension). If the verification passes, the quality evaluation result of the dataset is confirmed as valid, allowing it to enter the subsequent dynamic evaluation process; if the verification fails, the dataset is marked as "verification failed," and the reason for the failure is recorded (such as invalid proof π, public input not conforming to circuit constraints), and it is rejected from participating in the dynamic evaluation.

[0040] Through the configuration of the aforementioned malicious behavior detection rules, code execution trajectory capture logic, and zero-knowledge proof verification circuit, the dynamic quality assessment sandbox environment can provide a secure, reliable, and traceable execution platform for multi-dimensional dynamic quality assessment of integrated datasets.

[0041] S404. Based on the sandbox execution configuration, a multi-dimensional weighted scoring model and dynamic adjustment mechanism are used to perform a multi-dimensional dynamic quality assessment, generating a facility quality score that includes the entity inspection score of structural components, the weight coefficient of decoration and renovation projects, and a real-time update strategy. The core of the multi-dimensional weighted scoring model in step S404 lies in constructing a quality assessment index system covering the entire life cycle of port and waterway facilities, and achieving accurate characterization of facility status through dynamic weight adjustment. The model's index system comprises two main categories: basic dimensions and dynamic dimensions. Basic dimensions cover structural safety (initial weight 0.4), functional integrity (initial weight 0.3), operational efficiency (initial weight 0.2), and environmental adaptability (initial weight 0.1). Structural safety is further subdivided into sub-indicators such as concrete strength (0.35%), steel structure corrosion (0.25%), foundation settlement (0.2%), and structural cracks (0.2%). Functional integrity includes equipment operating status (0.4%), system interoperability (0.3%), and emergency response capability (0.3%). Operational efficiency includes throughput (0.5%), equipment utilization rate (0.3%), and energy consumption indicators (0.2%). Environmental adaptability includes typhoon resistance level (0.4%), corrosion resistance performance (0.3%), and ecological impact index (0.3%). Each sub-indicator is quantified by integrating structured monitoring data (such as concrete rebound hammer readings and stress sensor data), semantic parsing results of unstructured data (such as the text "crack width 3mm" in maintenance work orders), and statistical features of real-time data streams (such as the crane's mean time between failures), and converted into a standardized score of 0-100.

[0042] The dynamic adjustment mechanism achieves adaptive weight updates by introducing real-time feedback factors and historical trend factors. The real-time feedback factor (range 0-0.5) is dynamically generated based on the quality assessment results of real-time monitoring data in the sandbox environment. For example, if the accuracy score of a facility's structural crack sub-index is below 60 for three consecutive sliding windows (15 minutes), or if the integrity score C < 0.9 triggers an anomaly flag, the system automatically increases the weight of the structural safety dimension by 0.1 (maximum not exceeding 0.6) while decreasing the weight of the environmental adaptability dimension by 0.05. The historical trend factor (range 0-0.3) is calculated based on the trend of assessment data over the past three months. If the functional integrity score shows a continuous downward trend (monthly average decrease > 5%), its weight is increased by 0.05. Weight adjustments are implemented through preset constraint functions to ensure that the sum of the weights of each dimension remains 1 after adjustment, and that the change in the weight of a single dimension does not exceed 0.1 at a time, avoiding drastic fluctuations in the assessment results.

[0043] In the process of generating facility quality scores, the scores of each sub-indicator are first weighted and summed to obtain the basic dimension scores (e.g., structural safety score = concrete strength score × 0.35 + steel structure corrosion score × 0.25 + ...). Then, the comprehensive score is calculated by combining the dynamically adjusted dimension weights (comprehensive score = structural safety score × dynamic weight + ...). Simultaneously, for the physical inspection scores of structural components, a stratified sampling strategy is adopted: 20% of the structural components (e.g., wharf pile foundations, beams, panels) are randomly selected from the integrated dataset, and their key parameters (e.g., pile foundation verticality deviation, beam deflection) are compared with the monitoring data to calculate the inspection score (inspection score = ∑(deviation rate between inspection data and monitoring data × component weight)), which is used as a correction item for the comprehensive score (comprehensive score = comprehensive score × 0.9 + inspection score × 0.1). The weighting coefficient for decoration and renovation projects is dynamically set based on the facility's service life: 0.05 for facilities less than 5 years old, 0.1 for facilities 5-10 years old, and 0.15 for facilities over 10 years old. This weighting coefficient is multiplied by the decoration and renovation project score (such as wall integrity rate and signage clarity) and then included in the overall score. Regarding the real-time update strategy, the system updates sub-indicator scores hourly based on the latest monitoring data stream, updates dynamic weights daily at 2 AM based on the day's accumulated data, and calibrates the overall score on the 1st of each month based on the results of physical inspections, ensuring that the scoring results reflect changes in the facility's quality status in real time.

[0044] S405. Based on the facility quality score, using the Grad-CAM visualization method and DBSCAN clustering analysis algorithm, generate a defect distribution heat map that includes the thermal value normalization formula, the density index calculation logic, and the spatial autocorrelation assessment results. In step S405, the Grad-CAM visualization method is used to convert the contribution of each sub-indicator in the facility quality assessment to the overall score into a visual heatmap value, intuitively locating key areas affecting facility quality. Specifically, the facility quality score generated in step S404 is first used as input. The overall score is then backpropagated to each basic dimension and sub-indicator, calculating the gradient contribution value of each sub-indicator to the overall score. For example, for the "foundation settlement" sub-indicator under the structural safety dimension, its gradient weight in the overall score is calculated using the Grad-CAM algorithm. This weight reflects the degree of influence of the sub-indicator on the final score. Subsequently, the gradient contribution value of the sub-indicator is associated with the corresponding spatial coordinates of the facility BIM model. A heatmap value normalization formula is used to map the gradient contribution value to a pixel value range of 0-255: Heatmap value = (gradient contribution value - minimum contribution value) / (maximum contribution value - minimum contribution value) × 255. Here, the minimum contribution value is the minimum of the gradient contribution values ​​of all sub-indicators, and the maximum contribution value is the maximum value, ensuring that the heatmap values ​​are comparable globally. The normalized thermal values ​​are superimposed on the corresponding structural component surface of the real-time 3D twin, forming a preliminary defect-affected thermal distribution. The red area (high thermal value) indicates that the sub-index in this area has a greater negative impact on the quality score, while the blue area (low thermal value) indicates that the impact is smaller.

[0045] The DBSCAN clustering analysis algorithm is used to identify densely populated defect areas from heatmaps and quantify the spatial distribution characteristics of defects. The algorithm first extracts pixels with thermal values ​​above a set threshold (e.g., 180, corresponding to the top 30% of gradient contributions) from the heatmap as a set of candidate defect points. Each candidate point contains three-dimensional spatial coordinates (x, y, z) and thermal value attributes. Then, clustering is performed based on the core parameters of the DBSCAN algorithm (ε-neighborhood radius and minimum number of contained points): the ε-neighborhood radius is dynamically set according to the facility type, 5 meters for dock areas, 10 meters for waterway areas, and 2 meters for equipment areas; the minimum number of contained points is uniformly set to 5. The algorithm identifies core points, boundary points, and noise points by calculating the number of contained points within the ε-neighborhood of each candidate point, and aggregates densely connected core points and boundary points into defect clusters. For each cluster, its density index is calculated using the formula: Density = (Total thermal value within the cluster) / (Cluster volume × Number of points within the cluster), where the cluster volume is calculated using the bounding box method (by taking the extreme values ​​of the x, y, and z coordinates of all points within the cluster and calculating the volume of a cuboid). A higher density index indicates a more concentrated distribution of defects in the region and a greater degree of impact.

[0046] Spatial autocorrelation assessment results were obtained through Moran's I index calculation to analyze the spatial clustering or discrete patterns of defect distribution. The facility's three-dimensional space was divided into 10m × 10m × 5m cubic grid cells, and the number of candidate defect points within each grid cell was counted to form a spatial weight matrix. After constructing the weight matrix using the Queen adjacency rule (grids sharing edges or vertices are adjacent cells), the global Moran's I index (ranging from -1 to 1) was calculated. If I > 0 and is significant, it indicates that the defects are spatially clustered; if I < 0 and is significant, it indicates a discrete distribution; and if I is close to 0, it indicates a random distribution. Simultaneously, local Moran's I index calculations identified "high-high clustering" (high number of defect points in both the cell itself and surrounding grids) and "low-low clustering" (low number of defect points in both the cell itself and surrounding grids), which were highlighted in red and blue respectively in the heatmap to further emphasize the spatial correlation characteristics of the defects. The final defect distribution heatmap integrates the contribution visualization of Grad-CAM, the dense region identification of DBSCAN, and the spatial autocorrelation analysis results.

[0047] S406. Based on the aforementioned defect distribution heatmap, using the Prophet model, ARIMA model, and grid search algorithm, generate trend prediction results that include seasonality identification logic, confidence interval calculation rules, and multi-scenario integrated strategies.

[0048] Optionally, the generated trend prediction results include: S4061. Standardize the defect distribution heatmap to generate a standardized spatiotemporal defect dataset; The standardization process of the defect distribution heatmap in step S4601 aims to transform spatialized defect information into a structured dataset suitable for time series prediction models. The specific steps are as follows: First, extract the core features of each defect cluster in the heatmap, including the three-dimensional coordinates of the cluster center, density index, maximum heat value within the cluster, and cluster volume. For discrete noise points that do not form clusters (identified by the DBSCAN algorithm), if their heat value is still higher than the global average, record their coordinates and heat value separately as isolated defect points. Second, spatially divide the facility BIM model according to functional areas (such as the wharf front area, yard area, approach bridge area, and equipment area), with each area serving as an independent prediction unit. For each prediction unit, count the number of defect clusters, the sum of the total density index, the average heat value, and the number of isolated defect points within a preset time window (such as the past month, on a daily basis) to form preliminary time series features. Next, Z-score standardization is performed on each feature, using the formula: Standardized value = (Original value - Feature mean) / Feature standard deviation. The feature mean and standard deviation are calculated based on historical defect data from the past year, ensuring comparability of features of different magnitudes. Finally, the standardized regional features are arranged in timestamp order to construct a six-dimensional standardized spatiotemporal defect dataset containing region ID, timestamp, number of standardized clusters, total standardized density, average standardized heatmap value, and number of standardized isolated defect points.

[0049] S4062. Based on the standardized spatiotemporal defect dataset, extract seasonal, trend and residual components, identify the main frequency cycle through Fourier transform, and generate a seasonal feature vector containing multi-scale seasonal identification logic of year, season and month. In step S4062, seasonal feature extraction employs time series decomposition, decomposing the regional feature sequences (such as the number of standardized clusters) in the standardized spatiotemporal defect dataset into a trend term (T), a seasonal term (S), and a residual term (R), i.e., sequence value = T + S + R. The trend term is processed using a moving average method (window size set to 30 days) to remove short-term fluctuations, and the residual term is the difference between the original sequence and the trend and seasonal terms. The core of the seasonality identification logic lies in capturing periodic fluctuations in the sequence through Fourier transform. Specifically, a Fast Fourier Transform (FFT) is performed on the decomposed seasonal term S to convert the time-domain signal into a frequency-domain spectrum, and the power spectral density (PSD) of each frequency component is calculated. The top three frequencies with the highest PSD values ​​are selected as the dominant frequencies, and multi-scale seasonal patterns are identified based on the corresponding period (period = sampling frequency / dominant frequency). For example, if the standardized cluster number sequence of a certain region shows dominant frequency periods of 365 days, 90 days, and 30 days in the frequency domain analysis, then it is determined that the region exhibits seasonality at three scales: annual, seasonal, and monthly. For each scale of seasonality, a feature vector is constructed: the annual scale features include the sine / cosine values ​​of the month (e.g., sin(2π×month / 12), cos(2π×month / 12)), the quarterly scale features include the sine / cosine values ​​of the quarter (sin(2π×quarter / 4), cos(2π×quarter / 4)), and the monthly scale features include the sine / cosine values ​​of the date (sin(2π×date / 30), cos(2π×date / 30)). Simultaneously, considering the industry characteristics of port and waterway facilities, special time-marked features are introduced, such as binary marker variables for typhoon season (July-September) and dry season (December-February of the following year) (1 indicates the period is in that period, 0 indicates no). This ultimately forms a seasonal feature vector with 12 dimensions (2 dimensions for the year + 2 dimensions for the quarter + 2 dimensions for the month + 6 marker variables for special periods).

[0050] S4063. Based on the seasonal feature vector, the initial range of the p, d, and q parameters of the ARIMA model is determined by ACF analysis. The parameter combination is cross-validated by a grid search algorithm and optimized by the AIC criterion to generate the optimal ARIMA parameter configuration set. In step S4063, ACF analysis (autocorrelation function analysis) is used to determine the initial parameter ranges for the autoregressive term (p) and moving average term (q) in the ARIMA model. Specifically, firstly, a stationarity test (ADF test) is performed on the target feature sequences (such as standardized total density) of each region in the standardized spatiotemporal defect dataset. If the sequence is non-stationary (p-value > 0.05), it is stationary through differencing (d=1 or d=2, usually not exceeding order 2), with the differencing order d fixed as the initial parameter. Subsequently, the autocorrelation function (ACF) and partial autocorrelation function (PACF) plots of the stationary sequences are plotted: in the ACF plot, the maximum lag value exceeding the confidence interval (usually at a 95% confidence level) is used as the upper limit of q; in the PACF plot, the maximum lag value exceeding the confidence interval is used as the upper limit of p. For example, if the ACF plot falls into the confidence interval for the first time at lag 3, and the PACF plot falls into the confidence interval for the first time at lag 2, then the initial range of p is set to 0-2, and the initial range of q is set to 0-3.

[0051] The grid search algorithm is used to find the optimal parameter combination within the initial parameter range. A parameter grid is formed by combining possible values ​​of p (autoregressive order), d (difference order), and q (moving average order), where p∈[0,pmax], d∈[0,dmax], and q∈[0,qmax] (pmax, dmax, and qmax are determined based on ACF / PACF analysis results, and typically each dimension does not exceed 5). For each parameter combination, the first 80% of the time series data is used as the training set, and the last 20% as the validation set. The root mean square error (RMSE) of the model on the validation set is calculated using rolling window cross-validation (the window size is 10% of the training set length). Simultaneously, the Akaike Information Criterion (AIC) is used to penalize model complexity; a smaller AIC value indicates a higher goodness of fit to the data and lower complexity. Finally, the parameter combination with the smallest cross-validation RMSE and the smallest AIC value is selected as the optimal ARIMA parameter configuration set. For example, for a standardized cluster number sequence in a certain region, if the grid search results show that when p=2, d=1, q=1, the RMSE is 5.2 and the AIC is 320.5, which are better than other combinations, then (2,1,1) is determined as the optimal ARIMA parameter for this sequence.

[0052] S4064. Based on the optimal ARIMA parameter configuration set, construct the Prophet-ARIMA hybrid prediction model, and use a weighted average fusion strategy to integrate the prediction results to generate trend prediction results that include seasonality identification logic, confidence interval calculation rules, and multi-scenario integration strategies.

[0053] The Prophet-ARIMA hybrid prediction model employs a "parallel modeling-result fusion" architecture, fully leveraging the Prophet model's ability to capture nonlinear trends and strong seasonality, and the ARIMA model's advantage in accurately predicting linear stationary sequences. First, for each regional feature sequence (e.g., standardized total density) in the standardized spatiotemporal defect dataset, the Prophet and ARIMA models are trained independently: the Prophet model uses the seasonal feature vector generated in step S4062 as input, and performs predictions through a built-in additive model (trend term + seasonal term + holiday effect), where the trend term uses either a Logistic growth model or a linear model (automatically selected based on the sequence's growth characteristics), and sets multi-scale seasonal parameters for years, seasons, and months; the ARIMA model is constructed based on the optimal parameter configuration set (p, d, q) determined in step S4063, and performs linear autoregression and moving average predictions on the stationary sequence.

[0054] In the prediction result fusion stage, a dynamic weighted average strategy is adopted, with the weight coefficients dynamically adjusted based on the prediction error of the model on the historical validation set. Regarding the confidence interval calculation rules, the Prophet model generates 80% and 95% confidence intervals for the predicted values ​​through Monte Carlo simulation, while the ARIMA model calculates confidence intervals based on the assumption of a normal distribution of the residuals. The confidence intervals of the hybrid model adopt an "interval merging" strategy: the minimum lower bound of the confidence intervals of the two models is taken as the lower bound of the hybrid interval, and the maximum upper bound is taken as the upper bound of the hybrid interval, ensuring coverage of potential extreme fluctuations. The multi-scenario integration strategy sets scenario parameters for different operating states of port facilities (such as normal operation, maintenance period, and peak period). By introducing scenario-specific holiday effects (such as maintenance period marker variables) into the Prophet model and adjusting the difference order d in the ARIMA model (d=2 when the maintenance period sequence fluctuates significantly), specific prediction results for the corresponding scenarios are generated. These results, together with the basic prediction results, form a trend prediction report, including predicted values, confidence intervals, and scenario comparison analyses for the number of defect clusters, density indicators, and the number of isolated defect points for the next 1, 3, and 6 months.

[0055] Optionally, the generated remediation plan includes: S501. Based on the facility quality score, defect distribution heat map and trend prediction results, use a three-dimensional geometric difference detection algorithm to compare the spatial coordinate deviations in the BIM model design parameters and real-time monitoring data, and generate a geometric difference analysis report containing component-level geometric deviation values, position offsets and spatial overlap. In step S501, firstly, the component geometric data (including design coordinates, dimensions, and topological relationships) of the BIM model and the three-dimensional coordinates (x, y, z) of defect points in the real-time monitoring data are processed to achieve coordinate system unification. A seven-parameter coordinate transformation method (including 3 translation parameters, 3 rotation parameters, and 1 scaling parameter) is used to eliminate systematic errors under different coordinate systems. Then, for key components of the facility (such as wharf panels, approach bridge main beams, equipment foundations, etc.), rapid coarse matching is performed using spatial bounding boxes: using the smallest bounding cube of the component in the BIM model as a reference, the point set of the real-time monitoring point cloud data within the corresponding cube range is calculated, and outliers that are more than a preset threshold (e.g., 0.5 meters) away from the component design boundary are excluded. Next, an improved Iterative Closest Point (ICP) algorithm is used for fine matching: the vertices of the triangular facets on the surface of the BIM model component are used as the target point set, and the candidate defect points in the real-time monitoring point cloud are used as the source point set. Geometric alignment is achieved by minimizing the root mean square error (RMSE) between point sets. During the iteration process, the weight factor is dynamically adjusted, and defect points with a density index higher than the threshold are given higher weights to ensure that difference detection focuses on areas with concentrated defects. After fine matching, the shortest distance (distance in the normal direction) from each monitoring point to the surface of the BIM model component is calculated as the geometric deviation value. A positive deviation value indicates that the monitoring point is outside the component (such as surface depressions caused by concrete spalling), while a negative value indicates internal intrusion (such as exposed rebar or foreign object attachment). The positional offset is obtained by comparing the design coordinates of the component's centroid with the fitted centroid coordinates of the monitoring point set, and the offset components in the three-dimensional directions (x, y, z axes) are calculated separately. Spatial overlap is measured by calculating the percentage of the intersection volume of the three-dimensional convex hull of the monitoring point set and the BIM model component to the total volume of the component. An overlap of less than 90% indicates that there may be missing components or severe deformation. Finally, the geometric difference analysis report is categorized by component type, listing the maximum / minimum geometric deviation value, three-dimensional position offset, spatial overlap, and corresponding defect cluster density for each component.

[0056] S502. Based on the geometric difference analysis report and trend prediction results, use the Markov chain model to predict the defect propagation path and impact range, and generate a defect propagation prediction map that includes the defect propagation probability, the level of the impact area and the time window. In step S502, the Markov chain model divides the facility into several state spaces, each representing a potential defective or defect-free state. The transition probability between states is determined based on historical defect propagation data and current defect characteristics. First, according to the defect type (e.g., cracks, corrosion, deformation) and severity (minor, moderate, severe), the facility surface is divided into three-dimensional grid cells with a side length of 0.5 meters, with each grid cell serving as a state node. State variables include whether the cell has a defect, the defect type, the current severity, and the density index. Second, a state transition matrix is ​​constructed: by analyzing defect propagation cases of similar facilities over the past three years, the transition probability of different types of defects between adjacent grid cells is statistically analyzed (e.g., the probability of a severe crack cell propagating to an adjacent cell, the probability of a moderately corroded cell escalating to severe corrosion, etc.). The calculation of the transition probability introduces a spatiotemporal decay factor; the closer to the current defective cell and the higher the defect density, the greater the transition probability. For example, for a severe crack cell, the probability of its adjacent cells developing cracks within one month is set to 0.35, while the probability decreases to 0.12 when the cell is two grid cells away. Simultaneously, based on the trend prediction results generated in step S4064, the transition probability is dynamically adjusted: if the trend prediction shows that the number of defects in a certain area is increasing over the next 3 months, the transition probability of all units in that area is multiplied by a growth factor of 1.2. The initial state vector of the Markov chain is determined based on the current defect distribution heatmap, i.e., the initial defect state of each grid unit. Through iterative calculation (the number of iterations corresponds to the prediction time window, such as 24 iterations per week for the next 6 months), the probability distribution of the defect state of each grid unit at each time step is obtained. Based on this, a defect diffusion probability map is generated (different colors indicate the probability of defects occurring in each unit during the prediction period), and the influence area is divided into levels according to the probability values ​​(high-risk area: probability > 0.7, medium-risk area: 0.3 ≤ probability ≤ 0.7, low-risk area: probability < 0.3), while marking the time window in which significant defect diffusion is expected to occur in each area (e.g., diffusion is expected to occur in the high-risk area within the next 4-6 weeks).

[0057] S503. Based on the defect propagation prediction map, construct a repair priority evaluation matrix, and combine the weight coefficients in the facility quality score with the density index in the defect distribution heat map to generate a three-level repair priority ranking table containing emergency repair items, routine repair items, and preventive maintenance items. In step S503, the repair priority assessment matrix is ​​constructed using "impact level - urgency level" as the core two-dimensional index. The impact level dimension integrates the weight coefficients of the facility quality score and the defect density index, while the urgency level dimension is based on the diffusion probability and time window in the defect diffusion prediction map. The specific construction steps are as follows: First, determine the impact level score (0-10 points): Multiply the weight coefficient of each component in the facility quality score (e.g., 0.3 for the wharf panel and 0.25 for the approach bridge main beam) by the density index of the defect cluster of that component (0-1 after standardization), and then multiply by 10 to obtain the base score. If the component is on the critical path (e.g., the wharf area where ships dock in the main channel), an additional 2 points are added. The higher the impact level score, the greater the impact of the defect on the overall function of the facility. Secondly, the urgency score (0-10 points) is determined: based on the defect propagation prediction map, the urgency score for high-risk areas (probability > 0.7) with a time window within 1 month is 9-10 points; for medium-risk areas (0.3 ≤ probability ≤ 0.7) with a time window within 1-3 months, the score is 5-8 points; and for low-risk areas (probability < 0.3) with a time window greater than 3 months, the score is 1-4 points. If the defect type is rapidly spreading (such as alkali-aggregate reaction cracks in concrete), the score is increased by 1-2 points within the corresponding range. Subsequently, the impact and urgency scores are used as the horizontal and vertical axes of a matrix to divide the area into 9 assessment quadrants: the (high impact - high urgency) quadrant corresponds to emergency repair items, which must be initiated within 72 hours; the (medium impact - high urgency) and (high impact - medium urgency) quadrants correspond to routine repair items, which must be developed and implemented within 1-2 weeks; and the (low impact - low urgency), (medium impact - low urgency), and (low impact - medium urgency) quadrants correspond to preventative maintenance items, which are included in the quarterly maintenance plan. Finally, the repair items within the same priority are further sorted: emergency repair items are sorted in descending order by the product of "impact score × urgency score", while routine repair items and preventive maintenance items are fine-tuned based on the operating frequency of the component where the defect is located (e.g., approach bridges that are passed daily are prioritized over equipment foundations that are inspected once a week). The final result is a three-level repair priority sorting table that includes the repair item name, the component where it is located, the defect type, the impact score, the urgency score, the priority level, and the recommended processing time.

[0058] S504. Based on the repair priority ranking table, using historical repair case data in the expert experience database, matching similar defect scenarios through case reasoning algorithm, extracting key measures, resource requirements and time costs in the repair plan, and generating a repair plan that includes a repair step sequence, resource allocation list and time nodes.

[0059] In step S504, it is necessary to explain in detail that the core of the case reasoning algorithm is to retrieve the most similar historical cases to the current defect scenario from the expert experience database and generate a new repair plan based on the repair solutions of similar cases. First, a case representation model is constructed, quantifying the features of the current defect scenario into a multi-dimensional vector containing 7 core dimensions. Historical cases in the expert experience database are also stored in this vector structure and contain corresponding repair solutions. Second, weighted Euclidean distance is used to calculate the case similarity. Weights are assigned to the feature dimensions, and the weighted distance between the current scenario vector and the case vectors in the database is calculated. A similarity threshold is set to filter out the top 3 most similar cases. If the matching degree of the highest similarity case exceeds 0.95, the repair solution is directly reused; if it is below 0.7, a manual intervention process is triggered. For similar cases that meet the conditions, a voting method is used to extract key repair measures, and the top 5 core measures with the highest frequency in the 3 cases are counted as the basic steps. Resource requirements are dynamically adjusted based on the list of similar case resources and the current defect scale. If the current defect density is 1.5 times that of the cases, the number of resources is increased by 30% proportionally. The time cost estimation is based on the logical relationship of the repair steps and historical case work time data. The critical path method is used to calculate the total project duration, with a 15% buffer time. The final repair plan includes the repair step sequence, resource allocation list, time schedule, and feasibility assessment.

[0060] Example 2 A data quality assessment system under a distributed storage architecture includes: a processor and a memory for storing processor-executable instructions; wherein the processor is configured to implement a data quality assessment method under a distributed storage architecture when executing the executable instructions. It should be noted that the computer device includes a processor and memory, and may also include one or more of a multimedia component, an input / output (I / O) interface, and a communication component. The processor controls the overall operation of the computer device, completing all or part of the steps of the data quality assessment method under the distributed storage architecture. The memory stores various types of data to support device operation and can be implemented by volatile or non-volatile storage devices or combinations thereof, such as SRAM, EEPROM, etc. The multimedia component includes a screen (such as a touch screen) and an audio component. The audio component has a microphone to receive external audio signals and includes at least one speaker to output audio signals. The I / O interface provides an interface for the processor and other interface modules (such as a keyboard, mouse, buttons, etc.). The communication component is used for wired or wireless communication between devices. Wireless communication includes Wi-Fi, Bluetooth, etc., and the communication component includes a Wi-Fi module, etc. As a preferred embodiment, the computer device can be implemented using electronic components such as ASICs and DSPs to execute the data quality assessment method under a distributed storage architecture.

[0061] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A data quality assessment method under a distributed storage architecture, characterized in that, The method includes: Multimodal raw data of waterway port facilities are collected and preprocessed in real time to generate preprocessed structured datasets and unstructured data fingerprints; the unstructured data fingerprints contain data type and feature summary information. Based on the preprocessed structured dataset, dynamic fragmentation and storage are performed according to the three-dimensional fragmentation strategy of facility type, spatial grid and time window to obtain fragmented storage data; Based on the unstructured data fingerprint and sharded storage data, the quality assessment rules are encoded into an executable contract on the blockchain. The GNN spatiotemporal correlation analysis model and the Transformer semantic parsing model are trained across nodes through a federated learning differential privacy collaborative training framework. The evaluation results are verified by using zk-SNARKs zero-knowledge proof. Based on the evaluation results, it is verified that a real-time three-dimensional twin of water transport facilities integrating BIM+GIS is constructed, integrating the segmented storage data and real-time monitoring stream data, and performing multi-dimensional dynamic quality assessment through a dynamic quality assessment sandbox to generate facility quality scores, defect distribution heat maps and trend prediction results. Based on the facility quality score, defect distribution heat map and trend prediction results, the BIM model is compared with real-time monitoring data to automatically generate a repair plan; Obtaining fragmented storage data includes: Based on the facility type dimension, the preprocessed structured dataset is divided into four major infrastructure data modules: wharf structure, waterway dredging, port machinery, and security monitoring. Each infrastructure data module is associated with a corresponding facility ID code and industry standard parameter threshold. Using spatial grid partitioning rules, spatial coordinate mapping is performed on each infrastructure data module to generate a spatial index table containing grid number, facility center point coordinates, and neighborhood correlation degree; Based on the spatial index table, a dynamic time window mechanism is set up. A sliding window of a minute is used for wharf structure data, a fixed window of b minutes is used for waterway dredging data, a real-time window of c minutes is used for port machinery data, and an aggregated window of d minutes is used for security monitoring data. Each time window data is accompanied by a timestamp and a data collection frequency label. The infrastructure data module, spatial index table, and time window data are cross-linked by a 3D sharding algorithm to generate sharded data units containing facility type, spatial grid, and 3D time window identifiers, and obtain sharded storage data. Each sharded data unit contains a data check code, storage node address, and data update version number. The verification proof for generating the evaluation results includes: Based on the unstructured data fingerprint, a three-segment data fingerprint containing data type identifier, hash value, and timestamp is generated using the SHA-3 algorithm, and the fingerprint hash value and access address are anchored on the blockchain to generate on-chain data fingerprint anchoring results. Based on the on-chain data fingerprint anchoring results, the quality assessment rules are encoded into an executable contract on the blockchain, defining on-chain executable logic for accuracy verification, integrity verification, and timeliness monitoring, and generating a set of on-chain contracts for quality assessment rules. A federated learning differential privacy collaborative training framework is constructed based on the on-chain contract set of the quality assessment rules. Based on the fragmented storage data, a GNN spatiotemporal correlation analysis model is trained using a federated learning differential privacy collaborative training framework, and spatial correlation features of wharf pile foundation and breakwater settlement are extracted. Combined with TCN network analysis of channel displacement time series trend, spatiotemporal correlation analysis model output is generated. Based on the unstructured data fingerprint, a Transformer semantic parsing model is trained using a federated learning differential privacy collaborative training framework. This model performs semantic entity recognition and relation extraction on the text descriptions of port security monitoring video frames and the unstructured text of equipment fault work orders, generating the semantic parsing model output. By integrating the output of the spatiotemporal correlation analysis model and the output of the semantic parsing model, a multi-dimensional quality assessment is performed through the on-chain contract set of quality assessment rules to generate preliminary assessment results including data accuracy score, completeness index, and timeliness deviation value. Using the zk-SNARKs zero-knowledge proof algorithm, the circuit constraints are constructed with the preliminary evaluation results as input, and the evaluation results verification proof is generated.

2. The data quality assessment method under a distributed storage architecture as described in claim 1, characterized in that, The generated facility quality score, defect distribution heatmap, and trend prediction results include: Based on the evaluation results, it is verified that a real-time 3D twin embody integrating BIM and GIS has been constructed. Based on the real-time 3D twin, the fragmented storage data and real-time monitoring stream data are integrated to generate an integrated dataset containing structured data indexes, unstructured data fingerprints, and real-time data stream processing rules. Based on the integrated dataset, a dynamic quality assessment sandbox environment is constructed, generating a sandbox execution configuration that includes malicious behavior detection rules, code execution trajectory capture logic, and zero-knowledge proof verification circuitry. Based on the sandbox execution configuration, a multi-dimensional weighted scoring model and dynamic adjustment mechanism are used to perform a multi-dimensional dynamic quality assessment, generating a facility quality score that includes the entity inspection score of structural components, the weight coefficient of decoration and renovation projects, and a real-time update strategy. Based on the facility quality score, the Grad-CAM visualization method and DBSCAN clustering analysis algorithm are used to generate a defect distribution heat map that includes the thermal value normalization formula, the density index calculation logic and the spatial autocorrelation assessment results. Based on the aforementioned defect distribution heatmap, the Prophet model, ARIMA model, and grid search algorithm are used to generate trend prediction results that include seasonality identification logic, confidence interval calculation rules, and multi-scenario integrated strategies.

3. The data quality assessment method under a distributed storage architecture as described in claim 2, characterized in that, The generated trend prediction results, which include seasonality identification logic, confidence interval calculation rules, and multi-scenario integration strategies, include: The defect distribution heatmap is standardized to generate a standardized spatiotemporal defect dataset; Based on the standardized spatiotemporal defect dataset, seasonal, trend and residual components are extracted, and the main frequency cycle is identified by Fourier transform to generate a seasonal feature vector containing multi-scale seasonal identification logic of year, season and month. Based on the seasonal feature vector, the initial range of the p, d, and q parameters of the ARIMA model is determined by ACF analysis. The parameter combination is cross-validated by a grid search algorithm and optimized by the AIC criterion to generate the optimal ARIMA parameter configuration set. Based on the optimal ARIMA parameter configuration set, a Prophet-ARIMA hybrid prediction model is constructed. A weighted average fusion strategy is used to integrate the prediction results, generating trend prediction results that include seasonality identification logic, confidence interval calculation rules, and multi-scenario integration strategies.

4. The data quality assessment method under a distributed storage architecture as described in claim 1, characterized in that, The generated remediation plan includes: Based on the facility quality score, defect distribution heat map and trend prediction results, a three-dimensional geometric difference detection algorithm is used to compare the spatial coordinate deviations in the BIM model design parameters and real-time monitoring data to generate a geometric difference analysis report containing component-level geometric deviation values, position offsets and spatial overlap. Based on the geometric difference analysis report and trend prediction results, the defect propagation path and impact range are predicted using the Markov chain model, and a defect propagation prediction map containing defect propagation probability, impact area level and time window is generated. Based on the defect propagation prediction map, a repair priority evaluation matrix is ​​constructed. Combining the weight coefficients in the facility quality score with the density index in the defect distribution heat map, a three-level repair priority ranking table containing emergency repair items, routine repair items, and preventive maintenance items is generated. Based on the repair priority ranking table, and using historical repair case data in the expert experience database, a case reasoning algorithm is used to match similar defect scenarios, extract key measures, resource requirements and time costs in the repair plan, and generate a repair plan that includes a repair step sequence, resource allocation list and time nodes.

5. A data quality assessment system based on a distributed storage architecture, characterized in that, The system includes: processor; Memory used to store processor-executable instructions; The processor is configured to implement the data quality assessment method under the distributed storage architecture of any one of claims 1 to 4 when executing the executable instructions.