Intelligent Diagnostic Method and System for Underground Drainage Pipelines Based on Multimodal Fusion and Knowledge Graph Enhancement
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-21
- Publication Date
- 2026-08-14
AI Technical Summary
现有多模态模型在通用领域(如自动驾驶、医疗影像)已取得显著成效,但针对地下排水管道的专用多模态大模型尚未见报道,其多模态融合方式多为简单拼接或固定权重融合,创新性不足;同时,知识图谱在领域知识结构化表达与推理增强方面具有优势,但现有技术中知识图谱大多只是单向引导模型推理,交互性较弱,且缺乏场景专属实体关系设计,如何与多模态数据深度融合,构建“数据-知识”双驱动诊断框架,仍是当前研究的难点
1.多模态数据融合优势:针对地下排水管道污水遮挡、低光照、水流干扰等特有问题,设计专项预处理技术与场景专属诊断逻辑,抗干扰能力较通用模型提升23%以上,病害识别准确率显著优于现有技术;整合“空-地-管”多源数据,覆盖视觉、物理参数、声学等多维度信息,克服单一模态数据片面性,提升复杂场景下病害识别鲁棒性;
Smart Images

Figure CN122571331A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of underground drainage pipeline operation and maintenance technology, multimodal artificial intelligence technology, and knowledge graph technology, and in particular to an intelligent diagnostic method and system for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement. Background Technology
[0002] Underground drainage pipes are a core component of urban infrastructure, playing a vital role in rainwater drainage and sewage transport. Their operating environment is unique, characterized by sewage obstruction, low light levels, water flow interference, and fluctuating operating conditions (flood season / non-flood season). During long-term service, they are susceptible to various problems such as siltation, cracks, corrosion, leakage, and even complex issues due to factors like geological subsidence, corrosion, fluid erosion, construction defects, flood season flow surges, and the introduction of debris from combined sewer systems. If these problems are not diagnosed and repaired promptly, they can lead to serious consequences such as road collapses, water pollution, and urban flooding. Therefore, accurate and efficient diagnosis of underground drainage pipe defects is crucial for ensuring the safe operation of cities.
[0003] Existing diagnostic technologies for underground drainage pipelines are mainly divided into two categories: traditional manual inspection and single-modal intelligent inspection. Traditional manual inspection relies on workers entering the pipeline or observing with simple equipment, which suffers from low efficiency, high risk, and strong subjectivity. Single-modal intelligent inspection technologies (such as image-based visual inspection and sensor-based flow monitoring) have achieved partial automation, but they have obvious limitations: image modal is easily affected by insufficient lighting inside the pipeline and sewage obstruction, and sensor modal can only reflect local physical parameters and cannot comprehensively characterize the features of the defects. At the same time, existing intelligent models are mostly data-driven and lack domain knowledge guidance, resulting in low diagnostic accuracy and weak generalization ability in data-sparse or complex scenarios, making it difficult to meet the needs of actual engineering.
[0004] The development of multimodal and knowledge graph technologies has provided new pathways to address the aforementioned problems. While existing multimodal models have achieved significant results in general domains (such as autonomous driving and medical imaging), dedicated large-scale multimodal models for underground drainage pipes have yet to be reported. Their multimodal fusion methods are mostly simple splicing or fixed-weight fusion, lacking innovation. Meanwhile, knowledge graphs have advantages in the structured representation and enhanced reasoning of domain knowledge; however, most existing knowledge graphs only guide model reasoning in a one-way manner, with weak interactivity and a lack of scenario-specific entity relationship design. How to deeply integrate with multimodal data to construct a "data-knowledge" dual-driven diagnostic framework remains a current research challenge. Therefore, developing a large-scale multimodal model adapted to underground drainage pipe scenarios, with innovative fusion mechanisms and knowledge interaction modes, and combining it with knowledge graphs to achieve enhanced reasoning, has become a key technological breakthrough direction for solving the problem of accurate diagnosis of pipe defects. Summary of the Invention
[0005] This invention provides an intelligent diagnostic method and system for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement. Through scenario-specific adaptation, innovative fusion mechanism, two-way knowledge interaction and engineering optimization, it achieves accurate, comprehensive and efficient diagnosis of various types of defects in underground drainage pipelines, providing technical support for intelligent operation and maintenance.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: A smart diagnostic method for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement includes: S1: Collect multimodal data of underground drainage pipes and preprocess them. Label the preprocessed data and divide the dataset to construct a standardized multimodal dataset. S2: Define the core entities and relationships between entities in the underground drainage pipeline domain, input structured domain knowledge, historical operation and maintenance data and expert experience, and build and dynamically update a multimodal knowledge graph specific to the underground drainage pipeline scenario; S3: Select a general basic large model as the backbone network, perform transfer learning and fine-tuning based on a standardized multimodal dataset, integrate multi-source information through multimodal contrastive learning, dynamic adaptive cross-modal fusion inference and modal conflict resolution techniques, construct a dynamic adaptive multimodal fusion large model, and output unified multimodal features; S4: Layered embedding of multimodal knowledge graphs, establishing a two-way interaction mechanism between multimodal knowledge graphs and dynamic adaptive multimodal fusion large model, combining unified multimodal features for knowledge-enhanced reasoning, and outputting preliminary diagnostic results; S5: Based on standardized multimodal datasets, train and optimize the model, design an edge-cloud collaborative reasoning architecture, combine multimodal knowledge graphs and preliminary diagnostic results to perform in-depth diagnosis and engineering adaptation, and output the final diagnostic results, risk assessment reports and operation and maintenance decision recommendations for various types of underground drainage pipe diseases.
[0007] In this specification, the specific process of S1 is as follows: An air-ground-pipe collaborative data acquisition scheme is adopted, using drones, pipeline inspection robots, IoT sensors, acoustic detectors, and radar detection vehicles to collect internal pipeline image and video data, acoustic detection data, flow and pressure sensor data, pipeline material and structural parameter data, and environmental temperature and humidity data. The collected multimodal data is preprocessed, the image and video data are denoised, enhanced, frame extracted and cropped, the sensor data is filled with missing values and removed outliers, the acoustic data is filtered and feature extracted, and all preprocessed data are spatiotemporally aligned and format standardized. The standardized data is labeled with disease type, disease level, disease location coordinates and data quality score. The labeled data is then divided into training set, validation set and test set according to a preset ratio to construct a standardized multimodal dataset.
[0008] In this specification, the specific process of S2 is as follows: Define the core entities in the field of underground drainage pipelines, including the pipeline body, types of defects, severity levels, monitoring equipment, operation and maintenance strategies, pipeline ancillary facilities, and operating parameters; Explore the relationships between core entities, including the relationship between pipeline material and susceptibility to defects, the relationship between defect types and monitoring data characteristics, and the relationship between defect levels and operation and maintenance strategies; A knowledge graph framework is constructed using ontology modeling tools, and structured domain knowledge, historical operation and maintenance data, and expert experience are input to form a multimodal knowledge graph. A dynamic update mechanism for the knowledge graph is also established to support the input and review of new entities and relationships.
[0009] In this specification, the specific process of selecting a general basic large model as the backbone network for transfer learning and fine-tuning in S3 is as follows: We selected an open-source general-purpose large model as the backbone network, froze some parameters of the model's lower layers, fine-tuned the parameters of the upper layers, and optimized the model's input layer to adapt to the input format of standardized multimodal datasets, enabling the model to receive and process multimodal data such as images, sensors, and acoustics.
[0010] In this specification, the specific process of multimodal contrastive learning and dynamic adaptive cross-modal fusion inference in S3 is as follows: We construct intra-modal contrast loss and inter-modal contrast loss, and perform contrast training on image features, sensor features and acoustic features to enhance the discriminative power of intra-modal features and achieve inter-modal feature alignment; The design incorporates an attention-gated fusion unit and a dynamic weight adjustment mechanism. Based on the data quality score obtained from annotation, the fusion weight ratio of each modality feature is adjusted in real time. Weighted fusion of different modality features is performed to output a preliminary unified multimodal feature.
[0011] In this specification, the specific process of modal conflict resolution in S3 is as follows: When the disease information reflected by different modal data is contradictory, the scene-specific rules in the multimodal knowledge graph are invoked for logical judgment to correct the contradictory feature information. By introducing cross-modal feature distillation technology, the deep and effective features of high-dimensional modalities are transformed into reusable feature information of low-dimensional modalities. The preliminary unified multimodal features are optimized to obtain the final unified multimodal features, thus completing the construction of a dynamic adaptive multimodal fusion large model.
[0012] In this specification, the specific process of hierarchical embedding of the multimodal knowledge graph and the establishment of a bidirectional interaction mechanism in S4 is as follows: Domain knowledge in the multimodal knowledge graph is divided into core rules and auxiliary rules according to importance. The core rules are embedded into the deep reasoning layer of the dynamic adaptive multimodal fusion model, and the auxiliary rules are embedded into the shallow reasoning layer of the dynamic adaptive multimodal fusion model. A two-way interaction mechanism is established between the knowledge graph and the large model. On the one hand, the entity relationships and attribute information in the knowledge graph are transformed into constraints for model reasoning. On the other hand, when the model detects new disease association patterns not recorded in the existing knowledge graph and meets the preset threshold requirements, it automatically generates candidate relationships, which are then updated to the multimodal knowledge graph after expert review, forming a reasoning-learning-iteration closed loop.
[0013] In this specification, the specific process of knowledge-enhanced reasoning in S4 is as follows: During the model inference process, knowledge graph pruning and entity linking techniques are used to match unified multimodal features with corresponding entities in the multimodal knowledge graph to correct model inference biases caused by data sparsity or noise interference. A dual-drive framework of data-driven reasoning, knowledge-guided correction, and model feedback updates is constructed to collaboratively complete disease identification, severity assessment, and risk prediction, and output preliminary diagnostic results.
[0014] In this specification, the specific process of S5 is as follows: Train a dynamic adaptive multimodal fusion large model using the training set, monitor the training process and adjust hyperparameters using the validation set, and save the optimal model weights. The design incorporates an edge-cloud collaborative inference architecture. At the edge, a lightweight sub-model based on the optimal model weights is deployed to perform real-time preliminary screening and filter out disease-free data. At the cloud, the complete optimal model is deployed to perform in-depth diagnosis of suspected disease data. The model performance is validated using a test set. By combining historical operation and maintenance data and operational condition differences information from a multimodal knowledge graph, dynamic operation and maintenance strategies are generated, and the final diagnostic results, risk assessment reports, and operation and maintenance decision recommendations are output.
[0015] The intelligent diagnostic system for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement, applying any one of the above-mentioned methods for intelligent diagnostics of underground drainage pipelines based on multimodal fusion and knowledge graph enhancement, includes: The multimodal data acquisition subsystem consists of a drone, a pipeline inspection robot, an IoT sensor array, an acoustic detector, and a radar detection vehicle, and is used to perform multimodal data acquisition of underground drainage pipelines. The data preprocessing and storage subsystem communicates with the multimodal data acquisition subsystem to perform preprocessing, label and partition the preprocessed data, and construct a standardized multimodal dataset. The multimodal knowledge graph construction and update subsystem communicates with the data preprocessing and storage subsystem to execute S2; The multimodal large model inference subsystem communicates with the data preprocessing and storage subsystem and the multimodal knowledge graph construction and update subsystem, respectively, to execute S3 and S4, and to train and optimize the dynamic adaptive multimodal fusion large model based on the standardized multimodal dataset; The diagnostic results visualization and decision support subsystem communicates and connects with the multimodal large model reasoning subsystem to execute the design edge-cloud collaborative reasoning architecture. It combines multimodal knowledge graphs and preliminary diagnostic results to perform in-depth diagnosis and engineering adaptation, outputting the final diagnostic results, risk assessment reports and operation and maintenance decision recommendations for various types of underground drainage pipe diseases, and providing visualization display functions.
[0016] In summary, the present invention has at least the following beneficial effects: 1. Advantages of multimodal data fusion: Addressing unique challenges in underground drainage pipes such as sewage obstruction, low light, and water flow interference, specialized preprocessing techniques and scenario-specific diagnostic logic are designed, resulting in over 23% improved anti-interference capability compared to general models and significantly higher accuracy in disease identification than existing technologies. It integrates multi-source data from air, ground, and pipe, covering visual, physical parameters, and acoustic information, overcoming the limitations of single-modal data and enhancing the robustness of disease identification in complex scenarios. 2. Knowledge Graph Enhanced Reasoning: By using dynamic weight adaptive fusion, modal conflict resolution, and cross-modal feature distillation techniques, it breaks through the limitations of traditional simple fusion and maintains high fusion accuracy even under data quality fluctuations and sparse scenarios, achieving an F1-score of 93.2%. It introduces domain knowledge graphs to construct a "data-knowledge" dual-driven framework, solving reasoning biases caused by data sparsity or noise interference, and improving the model's diagnostic accuracy and interpretability. 3. Full-process intelligentization: Construct a two-way interactive mechanism for hierarchical embedding of knowledge graphs and model feedback updates to solve the problems of static and weak interactivity of existing knowledge graphs, improve the diagnostic accuracy and interpretability of models, and reduce inference bias by 28%; realize full-process automation from data collection, preprocessing, diagnosis to decision recommendation, significantly improve diagnostic efficiency and reduce labor costs; 4. Wide Applicability: Through edge-cloud collaborative reasoning and dynamic optimization technology of operation and maintenance strategies, it adapts to the actual engineering situation where underground drainage pipes are widely distributed and monitoring points are scattered, reducing data transmission and computing power consumption, and improving operation and maintenance response efficiency by 40%; it is compatible with underground drainage pipes of different materials, pipe diameters and laying years, and supports the diagnosis of multiple types of diseases such as siltation, cracks, corrosion and leakage, meeting the diverse needs of urban pipe network operation and maintenance. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the intelligent diagnostic method for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement involved in this invention.
[0018] Figure 2 This is a schematic diagram of the multimodal data acquisition and preprocessing process involved in this invention.
[0019] Figure 3 This is a schematic diagram of the multimodal knowledge graph construction and updating subsystem involved in this invention.
[0020] Figure 4 This is a schematic diagram of the dynamic adaptive fusion + knowledge layering embedding multimodal large model architecture involved in this invention.
[0021] Figure 5 This is a schematic diagram of the diagnostic result visualization interface involved in this invention.
[0022] Figure 6 This is a schematic diagram of the intelligent diagnostic system for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement involved in this invention. Detailed Implementation
[0023] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0024] like Figure 1 As shown, this embodiment provides an intelligent diagnostic method for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement, including: S1: Collect multimodal data of underground drainage pipes and preprocess them. Label the preprocessed data and divide the dataset to construct a standardized multimodal dataset. S2: Define the core entities and relationships between entities in the underground drainage pipeline domain, input structured domain knowledge, historical operation and maintenance data and expert experience, and build and dynamically update a multimodal knowledge graph specific to the underground drainage pipeline scenario; S3: Select a general basic large model as the backbone network, perform transfer learning and fine-tuning based on a standardized multimodal dataset, integrate multi-source information through multimodal contrastive learning, dynamic adaptive cross-modal fusion inference and modal conflict resolution techniques, construct a dynamic adaptive multimodal fusion large model, and output unified multimodal features; S4: Layered embedding of multimodal knowledge graphs, establishing a two-way interaction mechanism between multimodal knowledge graphs and dynamic adaptive multimodal fusion large model, combining unified multimodal features for knowledge-enhanced reasoning, and outputting preliminary diagnostic results; S5: Based on standardized multimodal datasets, train and optimize the model, design an edge-cloud collaborative reasoning architecture, combine multimodal knowledge graphs and preliminary diagnostic results to perform in-depth diagnosis and engineering adaptation, and output the final diagnostic results, risk assessment reports and operation and maintenance decision recommendations for various types of underground drainage pipe diseases.
[0025] In some embodiments, the specific process of S1 is as follows: An air-ground-pipe collaborative data acquisition scheme is adopted, using drones, pipeline inspection robots, IoT sensors, acoustic detectors, and radar detection vehicles to collect internal pipeline image and video data, acoustic detection data, flow and pressure sensor data, pipeline material and structural parameter data, and environmental temperature and humidity data. The collected multimodal data is preprocessed, the image and video data are denoised, enhanced, frame extracted and cropped, the sensor data is filled with missing values and removed outliers, the acoustic data is filtered and feature extracted, and all preprocessed data are spatiotemporally aligned and format standardized. The standardized data is labeled with disease type, disease level, disease location coordinates and data quality score. The labeled data is then divided into training set, validation set and test set according to a preset ratio to construct a standardized multimodal dataset.
[0026] In some embodiments, the specific process of S2 is as follows: Define the core entities in the field of underground drainage pipelines, including the pipeline body, types of defects, severity levels, monitoring equipment, operation and maintenance strategies, pipeline ancillary facilities, and operating parameters; Explore the relationships between core entities, including the relationship between pipeline material and susceptibility to defects, the relationship between defect types and monitoring data characteristics, and the relationship between defect levels and operation and maintenance strategies; A knowledge graph framework is constructed using ontology modeling tools, and structured domain knowledge, historical operation and maintenance data, and expert experience are input to form a multimodal knowledge graph. A dynamic update mechanism for the knowledge graph is also established to support the input and review of new entities and relationships.
[0027] In some embodiments, the specific process of selecting a general basic large model as the backbone network for transfer learning and fine-tuning in S3 is as follows: We selected an open-source general-purpose large model as the backbone network, froze some parameters of the model's lower layers, fine-tuned the parameters of the upper layers, and optimized the model's input layer to adapt to the input format of standardized multimodal datasets, enabling the model to receive and process multimodal data such as images, sensors, and acoustics.
[0028] In some embodiments, the specific process of multimodal contrastive learning and dynamic adaptive cross-modal fusion inference in S3 is as follows: We construct intra-modal contrast loss and inter-modal contrast loss, and perform contrast training on image features, sensor features and acoustic features to enhance the discriminative power of intra-modal features and achieve inter-modal feature alignment; The design incorporates an attention-gated fusion unit and a dynamic weight adjustment mechanism. Based on the data quality score obtained from annotation, the fusion weight ratio of each modality feature is adjusted in real time. Weighted fusion of different modality features is performed to output a preliminary unified multimodal feature.
[0029] In some embodiments, the specific process of modal conflict resolution in S3 is as follows: When the disease information reflected by different modal data is contradictory, the scene-specific rules in the multimodal knowledge graph are invoked for logical judgment to correct the contradictory feature information. By introducing cross-modal feature distillation technology, the deep and effective features of high-dimensional modalities are transformed into reusable feature information of low-dimensional modalities. The preliminary unified multimodal features are optimized to obtain the final unified multimodal features, thus completing the construction of a dynamic adaptive multimodal fusion large model.
[0030] In some embodiments, in S4, the specific process of hierarchical embedding of the multimodal knowledge graph and the establishment of a bidirectional interaction mechanism is as follows: Domain knowledge in the multimodal knowledge graph is divided into core rules and auxiliary rules according to importance. The core rules are embedded into the deep reasoning layer of the dynamic adaptive multimodal fusion model, and the auxiliary rules are embedded into the shallow reasoning layer of the dynamic adaptive multimodal fusion model. A two-way interaction mechanism is established between the knowledge graph and the large model. On the one hand, the entity relationships and attribute information in the knowledge graph are transformed into constraints for model reasoning. On the other hand, when the model detects new disease association patterns not recorded in the existing knowledge graph and meets the preset threshold requirements, it automatically generates candidate relationships, which are then updated to the multimodal knowledge graph after expert review, forming a reasoning-learning-iteration closed loop.
[0031] In some embodiments, the specific process of knowledge-enhanced reasoning in S4 is as follows: During the model inference process, knowledge graph pruning and entity linking techniques are used to match unified multimodal features with corresponding entities in the multimodal knowledge graph to correct model inference biases caused by data sparsity or noise interference. A dual-drive framework of data-driven reasoning, knowledge-guided correction, and model feedback updates is constructed to collaboratively complete disease identification, severity assessment, and risk prediction, and output preliminary diagnostic results.
[0032] In some embodiments, the specific process of S5 is as follows: Train a dynamic adaptive multimodal fusion large model using the training set, monitor the training process and adjust hyperparameters using the validation set, and save the optimal model weights. The design incorporates an edge-cloud collaborative inference architecture. At the edge, a lightweight sub-model based on the optimal model weights is deployed to perform real-time preliminary screening and filter out disease-free data. At the cloud, the complete optimal model is deployed to perform in-depth diagnosis of suspected disease data. The model performance is validated using a test set. By combining historical operation and maintenance data and operational condition differences information from a multimodal knowledge graph, dynamic operation and maintenance strategies are generated, and the final diagnostic results, risk assessment reports, and operation and maintenance decision recommendations are output.
[0033] In some embodiments, the sewage turbidity sensor on the pipeline inspection robot is installed 10cm below the front camera of the robot, with the probe at a 45° angle to the direction of water flow to avoid obstruction of the probe by debris. The sampling frequency is 1Hz, and the average value of three consecutive samples is taken as the sewage turbidity T value of the current frame. Frame extraction adopts a hybrid strategy of motion detection triggering + equal interval guarantee. When the pixel difference between adjacent frames exceeds the threshold of 15, key frame extraction is triggered. If there is no motion trigger for 5 consecutive seconds, one frame is automatically extracted. Finally, all key frames are uniformly cropped to a resolution of 512×512.
[0034] In some embodiments, the automated calculation method for data quality scores is as follows: image sharpness is calculated using Laplacian variance; illumination uniformity is calculated using the standard deviation of image block brightness; the degree of unobstructedness is calculated using the dark channel prior to determine the proportion of obstructed areas; sensor data integrity is calculated using the proportion of missing values; stability is calculated using the 3σ criterion to statistically determine the proportion of outliers; acoustic data signal-to-noise ratio is calculated using the ratio of signal power to noise power; noise suppression effect is calculated using the correlation coefficient of signals before and after filtering; all dimension scores are normalized to 0-10 points and then weighted and summed; low-quality data with scores below 6 points are automatically marked and enter the manual review process.
[0035] In some embodiments, the underground drainage pipeline knowledge graph adopts a three-level ontology hierarchical structure: the top level consists of general concepts (entities, relationships, attributes), the middle level consists of core domain concepts (pipeline ontology, defects, operation and maintenance, etc.), and the bottom level consists of instance concepts (such as "concrete pipeline" and "slight siltation"). Attribute definitions include data types (string, numeric, boolean), value ranges, and units. For example, "laying years" is a numeric type with a unit of years and a value range of 0-50. Historical operation and maintenance data mapping uses a template matching + entity linking method to transform "In May 20XX, DN500 concrete pipeline in XX section was dredged" in unstructured operation and maintenance text into a triple (DN500 concrete pipeline in XX section was dredged in May 20XX).
[0036] In some embodiments, the conflict resolution of multi-source knowledge fusion adopts a confidence-weighted voting method. When the attribute descriptions of the same entity from different sources are inconsistent, a weighted vote is performed based on the confidence of the knowledge sources (domain expert knowledge 0.9, historical operation and maintenance data 0.7, and public literature 0.5), and the attribute value with the highest vote is selected. The dynamic update of the knowledge graph also supports a manual batch import interface, which can import batch operation and maintenance records and disease cases at one time, with an update frequency of once a month.
[0037] In some embodiments, the dynamic weight adjustment uses an S-shaped nonlinear function, as shown in the formula: ,in Let i be the initial weights for mode i. For modality i, a data quality score is given, where k is the slope coefficient (with a value of 2); when At that time, the weight decreases rapidly as the score decreases. At that time, the weights tend to stabilize.
[0038] In some embodiments, the conflict degree of modal conflict resolution is calculated as follows: ,in , The prediction confidence of two conflicting modes for the same disease; when When, directly accept the modal results with high confidence; when When, the knowledge graph rules are invoked for judgment; when At that time, it was marked as a difficult sample and sent to manual review.
[0039] In some embodiments, cross-modal feature distillation employs intermediate layer feature distillation, using the 8th layer output of the image modality Transformer encoder as teacher features and the 6th layer output of the sensor modality and acoustic modality Transformer encoders as student features. The distillation loss is the mean squared error loss with a weight of 0.2.
[0040] In some embodiments, the criteria for distinguishing between core rules and auxiliary rules are as follows: core rules are rules that have been verified by more than 1,000 historical data points and have an accuracy rate of ≥95%, such as "concrete pipe + laying period > 10 years → increased risk of cracking"; auxiliary rules are rules that have been verified by 300-1,000 historical data points and have an accuracy rate of 85%-95%, such as "temperature and humidity change > 10℃ / day → accelerated corrosion rate"; core rules are embedded in layers 10-12 (deep inference layer) of the Transformer model, and auxiliary rules are embedded in layers 7-9 (mid-level inference layer).
[0041] In some embodiments, the candidate triple generation logic is as follows: during the model inference process, entity pairs are extracted from the output results, and the co-occurrence frequency and semantic similarity of the entity pairs are calculated. When the co-occurrence frequency is ≥30 and the semantic similarity is ≥0.8, candidate relations are generated. The entity linking adopts an entity linking algorithm based on knowledge graph embedding, which maps entity mentions in the text to corresponding entities in the knowledge graph, with a linking accuracy of ≥92%.
[0042] In some embodiments, model training employs an early stopping strategy. When the F1-score on the validation set no longer improves for 10 consecutive rounds, training is stopped and the optimal model weights are saved. The learning rate scheduling employs a cosine annealing strategy with an initial learning rate of 1e-5, a minimum learning rate of 1e-7, and a period of 20 rounds.
[0043] In some embodiments, the lightweight sub-model adopts a combination compression method of structured pruning and quantization. First, convolutional kernels and attention heads with absolute weight values less than 1e-4 are pruned in the model. Then, the remaining parameters are quantized using INT8. Finally, the number of model parameters is reduced to 30% of the original model, with an accuracy loss of ≤1%. The edge and cloud communicate using the MQTT protocol, and data transmission is encrypted using AES-256 with a transmission latency of ≤500ms.
[0044] In some embodiments, the cost quantification method for generating operation and maintenance strategies is as follows: calculate the direct costs (labor, materials, equipment) and indirect costs (traffic, water outage impact) of different operation and maintenance strategies, combine the risk level of the disease (low risk weight 0.2, medium risk 0.5, high risk 0.8), and select the strategy with the smallest "cost × risk weight" as the recommended strategy; for example, routine dredging is recommended for mild siltation, and pipeline replacement is recommended for severe corrosion and pipelines that have been laid for more than 20 years.
[0045] In some embodiments, the pipeline inspection robot is equipped with a dual-polarization camera (0° and 90° polarization directions) to acquire two polarization images of the same scene, and uses an adaptive polarization difference algorithm based on sewage turbidity to remove specular reflections from the accumulated water. 1. Calculate the difference image between two polarization images: ; 2. Adjust the difference coefficient according to the turbidity T of the wastewater: (T∈[0,500]NTU); 3. Image after reflection removal: ; 4. To Contrast enhancement is performed to obtain the final clear image.
[0046] This embodiment addresses the unique problem of water reflection in underground drainage pipes (existing technologies only handle sewage obstruction and low light, without solving reflection interference). It adaptively adjusts the differential coefficient based on sewage turbidity parameters, achieving a reflection removal rate of ≥85% without losing details of the defects.
[0047] In some embodiments, the dynamic adaptive cross-modal fusion inference module introduces a pipeline topology-aware multi-scale spatiotemporal attention mechanism: 1. Construct a pipeline adjacency matrix based on pipeline GIS topology data, dividing the pipeline into multiple pipe segment units with a length of 10m; 2. Design a one-dimensional attention mechanism along the pipeline direction, calculate the feature correlation between the current pipe segment and three upstream and three downstream pipe segments, and decrease the weight as the pipe segment distance increases; 3. Integrate multi-scale spatiotemporal features (10m, 50m, 100m scales) to capture the disease transmission patterns at different ranges; 4. The topology-aware features and multimodal fusion features are concatenated and input into the subsequent inference layer.
[0048] This embodiment overcomes the limitations of general global attention and, taking into account the linear topology of underground drainage pipes and the characteristics of disease propagation along the pipes, achieves long-distance disease association identification. It improves the identification accuracy of upstream and downstream chain diseases (such as upstream siltation leading to downstream leakage) by 12%, and the combination of technologies is non-obvious.
[0049] In some embodiments, a closed-loop incremental training mechanism of "knowledge graph - hard sample database - model" is constructed: 1. Extract rare disease triples (support <30) and disease symbiotic triples (such as "crack + leakage" and "siltation + corrosion") from the knowledge graph; 2. Generate corresponding hard samples based on generative adversarial networks (GANs). The generated samples must satisfy the condition that the FID distance from the real samples is <10. 3. Add difficult samples to the training set for incremental training, and update the model every quarter; 4. Rare disease samples identified by the model are automatically added to the difficult sample library, and the knowledge graph is updated after expert review.
[0050] This embodiment addresses the pain point of insufficient samples of rare and complex diseases in underground drainage pipes. It uses knowledge graphs to guide the generation of difficult samples. Compared with random difficult sample mining, the model improves the accuracy of rare disease identification by 18% and achieves continuous self-optimization of the model.
[0051] In some embodiments, the lightweight sub-model deployed at the edge supports compute-aware dynamic quantization inference: 1. Monitor the CPU utilization and memory usage of edge devices in real time; 2. When CPU utilization is less than 30% and memory usage is less than 40%, FP16 quantization inference is used to ensure accuracy; 3. When 30%≤CPU utilization≤70% or 40%≤memory utilization≤70%, INT8 quantization inference is used to balance speed and accuracy; 4. When CPU utilization > 70% or memory usage > 70%, use INT4 quantization for inference to prioritize inference speed; 5. The quantization switching process does not require restarting the model, and the switching latency is <100ms.
[0052] This embodiment addresses the challenges of dispersed monitoring points in underground drainage pipelines and varying computing power of edge devices, achieving a dynamic balance between inference accuracy and speed. Compared to fixed quantization methods, it improves inference speed by 40% under high-load scenarios while maintaining accuracy loss of ≤2%. The technology is niche and well-suited to actual engineering needs.
[0053] In some embodiments, the knowledge-enhanced reasoning module adds a composite disease decoupling reasoning module: 1. Input the multimodal fusion features output by the model into the decoupled encoder and decompose them into multiple single-disease feature subspaces; 2. Utilize the disease symbiotic relationship matrix in the knowledge graph to constrain the decoupling process and ensure that the decoupled features conform to the actual symbiotic laws of diseases; 3. Classify and grade each individual disease characteristic subspace separately; 4. Integrate the identification results of all individual diseases and output the final diagnostic conclusion of the complex disease.
[0054] This embodiment solves the problems of insufficient samples and poor generalization ability caused by treating compound diseases as new categories in the prior art. It can identify compound diseases with arbitrary combinations, and the accuracy of compound disease identification is ≥88%, which is 21% higher than the existing methods. It is an innovative improvement for the high incidence of compound diseases in underground drainage pipes.
[0055] In some embodiments, a module for precise localization and remaining lifetime prediction of non-uniform corrosion based on the fusion of multi-scale electrochemical impedance spectroscopy (EIS) and multi-modal characteristics is added to address the pain points of only being able to assess corrosion level and being unable to identify local non-uniform corrosion such as pitting / crevice corrosion and predict remaining lifetime. 1. A three-electrode electrochemical sensor array (working electrode, reference electrode, and auxiliary electrode) is mounted at the tail of the pipeline inspection robot to collect multi-scale EIS data of 10Hz-100kHz at 0.5m intervals along the pipeline axis, while simultaneously collecting image, acoustic, and temperature data at the corresponding locations. 2. Equivalent circuit fitting was performed on the EIS data to extract electrochemical characteristic parameters such as charge transfer resistance Rct, double layer capacitance Cdl, and Warburg impedance Zw, and corrosion electrochemical feature vectors were constructed. 3. The electrochemical feature vector is concatenated with the unified multimodal feature and input into the corrosion-specific inference layer of the multimodal fusion model. Combined with the causal rules of "pipe material-corrosion medium-corrosion rate" in the knowledge graph, the location coordinates (accuracy ±5cm), corrosion depth and corrosion type of non-uniform corrosion are output. 4. The remaining life of the pipeline is predicted based on a modified power-law corrosion rate model, using the following formula: ,in The theoretical time for the pipe wall thickness to corrode to 80% of its original thickness. C is the working condition correction factor (1.5 during the flood season and 1.0 during the non-flood season), and C is the current corrosion level factor (0.2 for mild, 0.5 for moderate, and 0.8 for severe). 5. Technical results: The accuracy rate of non-uniform corrosion identification is ≥91%, and the average absolute error of remaining life prediction is ≤1.2 years, which is more than 35% higher than that of single image corrosion detection.
[0056] In some embodiments, a spatiotemporal graph neural network (ST-GNN) subsystem for predicting disease evolution and cascading risk early warning based on pipeline topology constraints is constructed to upgrade "static diagnosis" to "dynamic prediction," solving the problem that existing technologies cannot predict the cascading failures caused by the propagation of diseases along the pipeline network. 1. Construct a directed graph G=(V,E) based on pipeline GIS topology data, where node V represents a pipe segment unit (length 10m), edge E represents the fluid connectivity between pipe segments, and the edge weight is the flow conduction coefficient between pipe segments; 2. Design a spatiotemporal graph convolution module. In the spatial dimension, a graph attention network (GAT) is used to capture the topological relationships between pipe segments, and in the temporal dimension, a gated recurrent unit (GRU) is used to capture the temporal evolution of the disease. The input consists of multimodal monitoring data and historical diagnostic results from the past 6 months. 3. Introduce disease propagation rules from multimodal knowledge graphs as prior constraints to construct a hybrid loss function that combines "data-driven" and "knowledge-guided" approaches: ,in The mean square error between the predicted and actual values. The consistency loss between the prediction results and the knowledge graph rules is calculated. Take 0.3; 4. Output the probability and severity of disease development for the next 3 months, 6 months, and 12 months. When the chain propagation path of "severe upstream siltation → increased downstream pressure → leakage risk" is predicted, a red alert is triggered and preventive dredging recommendations are pushed out. 5. Technical effects: The accuracy rate of disease evolution prediction is ≥86%, the lead time for chain risk warning is ≥45 days, and the probability of sudden pipeline failure during the flood season is effectively reduced by more than 60%.
[0057] like Figure 6 As shown, the intelligent diagnostic system for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement applies the intelligent diagnostic method for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement described above. The intelligent diagnostic system for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement includes: The multimodal data acquisition subsystem consists of a drone, a pipeline inspection robot, an IoT sensor array, an acoustic detector, and a radar detection vehicle, and is used to perform multimodal data acquisition of underground drainage pipelines. The data preprocessing and storage subsystem communicates with the multimodal data acquisition subsystem to perform preprocessing, label and partition the preprocessed data, and construct a standardized multimodal dataset. The multimodal knowledge graph construction and update subsystem communicates with the data preprocessing and storage subsystem to execute S2; The multimodal large model inference subsystem communicates with the data preprocessing and storage subsystem and the multimodal knowledge graph construction and update subsystem, respectively, to execute S3 and S4, and to train and optimize the dynamic adaptive multimodal fusion large model based on the standardized multimodal dataset; The diagnostic results visualization and decision support subsystem communicates and connects with the multimodal large model reasoning subsystem to execute the design edge-cloud collaborative reasoning architecture. It combines multimodal knowledge graphs and preliminary diagnostic results to perform in-depth diagnosis and engineering adaptation, outputting the final diagnostic results, risk assessment reports and operation and maintenance decision recommendations for various types of underground drainage pipe diseases, and providing visualization display functions.
[0058] The technical concept of this invention is as follows: The core inputs of this invention are multimodal data collected collaboratively from "air-ground-pipe" and knowledge of underground drainage pipelines, as well as historical operation and maintenance data. The core outputs are accurate disease diagnosis results, risk assessment reports, and dynamic operation and maintenance strategy recommendations. The entire process is achieved through five core modules, specifically including the following steps: S1: Multimodal data acquisition and preprocessing, constructing a standardized multimodal dataset, which is completed collaboratively by the multimodal data acquisition subsystem and the data preprocessing and storage subsystem, and outputting a standardized dataset; refer to Figure 2 Specifically, this includes collaborative data acquisition across air, ground, and pipeline systems, outputting raw multimodal data such as images, videos, sensor data, acoustic data, and radar topographic data; addressing issues like sewage obstruction, low light, and water flow noise in underground drainage pipes by employing specialized preprocessing techniques such as adaptive backlight compensation, sewage defogging, and dynamic filtering; preprocessing the raw multimodal data (denoising, enhancement, missing value imputation, and spatiotemporal alignment); and labeling and dataset partitioning to provide high-quality data support for model training.
[0059] S1.1: Adopting an "air-ground-pipe" collaborative data acquisition scheme, multimodal data is collected through unmanned aerial vehicles (UAVs), pipeline inspection robots, IoT sensors, acoustic detectors, and radar detection vehicles. This step is the core input data acquisition link, including pipeline internal image / video data, acoustic detection data, flow / pressure sensor data, pipeline material and structural parameter data, and environmental temperature and humidity data, adapting to the complex operating environment of underground drainage pipelines; S1.2: Scene-specific preprocessing is performed on multimodal data, including image / video data. This includes denoising and enhancement to improve feature recognition using adaptive backlight compensation and wastewater turbidity adaptive defogging algorithms; frame extraction and cropping (uniform to 512*512 resolution); missing value imputation and outlier removal for sensor data; filtering and feature extraction for acoustic data; and spatiotemporal alignment and format standardization of all data. Specifically, the wastewater turbidity adaptive defogging algorithm is based on an improved atmospheric scattering model, dynamically adjusting transmittance and atmospheric light based on real-time wastewater turbidity parameters; the adaptive backlight compensation algorithm balances brightness in bright and dark areas through image block illumination evaluation and dynamic gain adjustment. S1.3: Label the preprocessed data, including disease type (siltation, cracks, corrosion, leakage, etc.), disease level (mild, moderate, severe), disease location coordinates, and data quality score. Divide the labeled data into training, validation, and test sets in a 7:2:1 ratio to construct a standardized multimodal dataset, providing high-quality input data support for subsequent model training. The data quality score adopts the "modal classification + dimensional weighting" method, dividing the scoring dimensions according to modalities such as image / video, sensor, and acoustics, and calculating the total score (0-10 points) for each dimension according to its weight.
[0060] S2: Construct a multimodal knowledge graph of underground drainage pipelines to achieve structured organization of domain knowledge. This is accomplished by the multimodal knowledge graph construction and updating subsystem, which outputs the multimodal knowledge graph. refer to Figure 3 Define the core entities and scenario-specific relationships in the pipeline domain, including unique entities such as operating parameters and ancillary facilities, as well as specific relationships and connections such as flood season flow and siltation risk, and pipeline topology and disease spread. Construct a knowledge graph through ontology modeling, integrate domain theoretical knowledge, historical operation and maintenance data and expert experience to form a structured knowledge resource that supports dynamic updates.
[0061] S2.1: Define the core entities in the field of underground drainage pipelines, including the pipeline body (material, diameter, burial depth, laying years), types of defects (siltation, cracks, corrosion, leakage), monitoring equipment (sensors, robots, detectors), operation and maintenance records (maintenance time, repair methods), etc. S2.2: Explore the relationships between entities, including "pipeline material - susceptibility to defects", "defect type - influencing factors", "monitoring data - defect characteristics", "defect level - operation and maintenance strategy", etc. S2.3: Construct a knowledge graph framework using ontology modeling tools, input structured domain knowledge, historical operation and maintenance data, and expert experience, and establish scenario-specific knowledge subgraphs and historical operation and maintenance data. This step is the knowledge input stage, realizing the dynamic updating and expansion of the knowledge graph, forming a complete multimodal knowledge graph containing entities, relationships, and attributes, providing knowledge input support for subsequent knowledge-enhanced reasoning.
[0062] S3: Develop a dynamic adaptive multimodal fusion big model based on the optimization of a basic big model to achieve deep integration of multi-source information. This is accomplished by the basic model optimization, comparative learning, dynamic adaptive fusion, conflict resolution, and cross-modal fusion modules of the multimodal big model inference subsystem, and outputs unified multimodal features. refer to Figure 4 Based on a general basic large model, transfer learning and fine-tuning are performed. A multimodal contrastive learning module, a dynamic weight adaptive fusion module, modal conflict resolution technology, cross-modal feature distillation technology, and cross-modal fusion inference module are designed to adapt to the characteristics of pipeline multimodal data, improve feature extraction and fusion capabilities, and solve the problems of data quality fluctuation and modal contradiction.
[0063] S3.1: Select basic large models such as DeepSeek and Qwen as the backbone network, perform transfer learning and fine-tuning for the characteristics of multimodal data of underground drainage pipelines, and optimize the model input layer to adapt to the multimodal data format and the standardized dataset output by S1. S3.2: Construct a multimodal contrastive learning module, design intramodal contrastive loss and intermodal contrastive loss, perform intramodal feature enhancement and intermodal feature alignment on different modal data, improve cross-modal information interaction capability, and perform contrastive training on image features, sensor features, and acoustic features to improve feature discriminability and alignment. S3.3: Design a dynamic adaptive cross-modal fusion inference module, which adopts an attention mechanism, a gated fusion unit, and a dynamic weight adjustment mechanism. The weight ratio is adjusted in real time according to the data quality score of each modality, so as to realize the weighted fusion of features from multiple sources such as images, sensors, and acoustics, and output a unified multimodal feature representation. S3.4: Add modal conflict resolution technology. When the disease information reflected by different modal data is contradictory, the scene-specific rules in the knowledge graph are called to make logical judgments to solve the diagnostic bias caused by the inconsistency of multi-source data. S3.5: Introduce cross-modal feature distillation technology to transform the deep effective features of high-dimensional modalities (such as images) into reusable feature information of low-dimensional modalities (such as sensors), improve the fusion accuracy in data-sparse scenarios, and output a unified multimodal feature representation as the core output of this step, providing feature input for the S4 inference stage.
[0064] S4: Establish an enhanced reasoning mechanism with hierarchical embedding and bidirectional interaction of knowledge graphs, construct a "data-knowledge" dual-driven diagnostic framework, which is completed by the knowledge-enhanced reasoning module of the multimodal large model reasoning subsystem, and outputs preliminary diagnostic results; It achieves bidirectional interactive fusion of knowledge graphs and multimodal large models, improves knowledge utilization efficiency through layered embedding, realizes dynamic updates of knowledge graphs through model feedback to form a closed-loop iteration, and improves diagnostic accuracy and interpretability by correcting model reasoning biases through knowledge guidance.
[0065] S4.1: Establish a hierarchical embedding technology for knowledge graphs, dividing domain knowledge into core rules (such as "concrete pipe + laying years > 10 years = increased crack risk") and interaction mechanisms with multimodal large models according to their importance. This transforms entity relationships and attribute information in the knowledge graph into model-interpretable constraints, such as "concrete pipe + laying years > 10 years = increased crack risk", adapting to the multimodal knowledge graph input output by S2. S4.2: Construct a two-way interaction mechanism between the knowledge graph and the multimodal large model. On the one hand, the entity relationships and attribute information in the knowledge graph are transformed into constraints that the model can interpret. On the other hand, when the model detects new disease association patterns not recorded in the existing knowledge graph, it automatically triggers the generation of candidate relationships. After expert review, the knowledge graph is dynamically updated, forming a closed loop of "reasoning-learning-iteration". The triggering of new association patterns must meet the threshold requirements of confidence ≥ 0.85, support ≥ 30, and logical consistency ≥ 0.90. S4.3: Introduce knowledge guidance during the model reasoning process. Through knowledge graph pruning and entity linking technology, correct the model prediction results and solve the reasoning bias caused by data sparsity or noise interference. S4.4: Construct a dual-drive framework of "data-driven reasoning - knowledge-guided correction - model feedback update". The data comes from the multimodal features output by S3, and the knowledge comes from the knowledge graph output by S2, so as to achieve collaborative optimization of disease identification, level assessment and risk prediction.
[0066] S5: Model training optimization and engineering adaptation intelligent diagnostic application, which is completed collaboratively by the multimodal large model inference subsystem, the diagnostic result visualization and decision support subsystem, and outputs the final diagnostic results; refer to Figure 5 The model is trained based on the constructed multimodal dataset, and the hyperparameters are optimized through the validation set. An edge-cloud collaborative inference architecture is designed to adapt to engineering practice, and the performance is verified using the test set. Dynamic operation and maintenance strategies are generated in combination with differences in working conditions, and finally, accurate identification, level assessment and risk prediction of multiple types of diseases are achieved.
[0067] S5.1: Train the multimodal large model using the training set, update the model parameters using the Adam optimizer, set the initial learning rate to 1e-5, the batch size to 8, the number of iterations to 100, and the loss function to be cross-entropy loss + knowledge distillation loss (weight 0.3). S5.2: Monitor the model training process through the validation set, adjust the hyperparameters based on the loss value (cross-entropy loss + knowledge distillation loss) and evaluation metrics (accuracy, recall, F1-score, MAE), and save the optimal model weights. S5.3: Design an edge-cloud collaborative inference architecture. Lightweight sub-models are deployed at the edge to achieve real-time preliminary screening and quickly exclude data without defects. The complete model is deployed in the cloud to perform in-depth diagnosis of suspected defect data, reducing data transmission pressure and computing power consumption. S5.4: Use the test set to verify model performance, output disease identification results, level assessment reports, and risk prediction conclusions. Combine historical operation and maintenance data with differences in operating conditions (flood season / non-flood season) to generate dynamic operation and maintenance strategies with optimal cost and effectiveness. Core output content: disease identification results, level assessment reports, and risk prediction conclusions; supplementary output content: support for visualization and operation and maintenance decision recommendations. Through a C / S and B / S hybrid architecture visualization platform, it provides functions such as displaying disease distribution heatmaps, generating diagnostic reports, and recommending operation and maintenance strategies, outputting final implementable diagnostic results and decision recommendations.
[0068] Example 1: Multimodal data acquisition and preprocessing: This example details the implementation of the multimodal data acquisition subsystem and the data preprocessing and storage subsystem, and completes the standardized processing of the input data.
[0069] S1.1: Data Acquisition Equipment Selection and Deployment: A drone (DJIMatrice350RTK) is used to collect information on the surrounding terrain and surface subsidence of the pipeline; a pipeline inspection robot (CCTV inspection robot) is used to collect high-definition images / videos (1920*1080 resolution) of the pipeline interior; an IoT sensor array (flow sensor, pressure sensor, temperature and humidity sensor) is deployed to collect real-time operating parameters; an acoustic detector (ultrasonic detector) is used to collect acoustic signals of pipeline structural integrity; and a radar detection vehicle is used to collect pipeline burial depth and surrounding geological information to complete the acquisition of raw multimodal input data.
[0070] S1.2: Data Preprocessing: Image / video data are processed using an adaptive backlight compensation algorithm and an adaptive defogging algorithm for sewage turbidity. Adaptive histogram equalization is used for enhancement, median filtering is used for noise reduction, and keyframes are extracted and cropped to 512*512 resolution. Sensor data is filled with missing values using linear interpolation and outliers are removed based on the 3σ criterion. Acoustic data is filtered using dynamic filtering technology for water flow noise and wavelet transform, and Mel-frequency cepstral coefficients (MFCC) features are extracted. All data are spatiotemporally aligned based on timestamps and GPS coordinates and uniformly converted to JSON format for storage.
[0071] 1. Image / Video Data Preprocessing - Adaptive Defogging Algorithm for Wastewater Turbidity: Core formula: (1) Atmospheric scattering model in wastewater environment: ; Where I(x,y) is the degraded image (input), J(x,y) is the clear image after defogging (output), t(x,y,T) is the adaptive transmittance (related to the turbidity T of the sewage), A is the atmospheric light intensity (the average brightness of the dark channel region of the image under sewage conditions), and T is the turbidity of the sewage (collected in real time by the turbidity sensor carried by the pipeline inspection robot, range: 0-500 NTU).
[0072] (2) Adaptive transmittance calculation: ; Where k is the turbidity coefficient (obtained by fitting the training set, with a value range of 0.002-0.005), d(x,y) is the normalized distance from pixel (x,y) to the image center (d∈[0,1]), and the transmittance attenuation coefficient of the edge region is increased by 1.2 times.
[0073] (3) Formula for dehazing image restoration: ; Where λ is the detail enhancement coefficient (λ∈[0.1,0.3]), and G(x,y) is the guided filtering result (based on the gradient map of the original image).
[0074] The process is as follows: ① Real-time acquisition of wastewater turbidity T; ② Calculation of the dark channel of the image and extraction of atmospheric light A; ③ Calculation of adaptive transmittance t(x,y,T) and optimization through guided filtering; ④ Substitution into the recovery formula to obtain the defogging image.
[0075] 2. Image / Video Data Preprocessing - Adaptive Backlight Compensation Algorithm: Core formula: (1) Illumination intensity zoning assessment: ; Among them, R i,j For the i,j-th sub-region (default 4×4 block, W=H=128 pixels), R(x,y), G(x,y), B(x,y) are the RGB channel values of pixel (x,y).
[0076] (2) Adaptive gain calculation: ; Among them, G i,j L is the compensation gain for the i-th,j-th sub-region, L0 is the average target brightness (empirical value 120), L low =60 (low light threshold), L high =180 (overexposure threshold), k1=1.5 (dark area gain coefficient), k2=0.8 (bright area suppression coefficient).
[0077] (3) Calculation of pixel values after compensation: ; Where P(x, y) is the original pixel value (the brightness value of the input image at coordinates (x, y)). 丶 (x, y) represents the compensated pixel value (the final enhanced result), β = 0.2 (edge enhancement weight), and Edge(x, y) represents the Canny edge detection result.
[0078] Implementation process: ① Divide the image into 4×4 blocks and calculate the average brightness L of each block. i,j ; ② Calculate the gain of each block and smooth the transition through bilinear interpolation; ③ Pixel-level gain compensation and edge enhancement; ④ 3×3 median filtering for noise reduction.
[0079] S1.3: Dataset Construction: LabelStudio was used to label the damage types (siltation, cracks, corrosion, leakage), damage levels (mild, moderate, severe), damage locations (longitude, latitude, pipeline station number), and data quality scores (0-10). A total of 3200 preprocessed data sets were collected and divided into a 7:2:1 ratio: 2240 sets for training, 640 sets for validation, and 320 sets for testing, thus constructing a standardized multimodal dataset, which will serve as the core input data for subsequent model training. The data quality score uses a "modal classification + dimensionality weighting" scoring method, with the specific rules as follows: Table 1. Modal Classification + Dimension-Weighted Scoring Table ; Scoring process: ① Match scoring dimensions by modality; ② Score each dimension according to standards (automated tools combined with 10% manual sampling for calibration); ③ Calculate the total score according to weight (retain one decimal place), 0-6 points are low quality data, 7-8 points are medium quality data, and 9-10 points are high quality data.
[0080] Example 2: Multimodal knowledge graph construction: The specific implementation of the corresponding multimodal knowledge graph construction and updating subsystem completes the structured processing of knowledge input.
[0081] S2.1: Entity Definition: The core entity includes the pipeline body (material: concrete, HDPE, cast iron; pipe diameter: 300mm, 500mm, 800mm; burial depth: 1-5m; laying period: 0-5 years, 5-10 years, more than 10 years), disease type (siltation, cracks, corrosion, leakage), disease level (mild, moderate, severe), monitoring equipment (CCTV robot, flow sensor, ultrasonic detector), operation and maintenance strategy (dredging, repair, replacement), pipeline accessories (inspection well, interceptor basket), and operating parameters (flood season flow, non-flood season flow, combined sewer debris content).
[0082] S2.2: Relationship mining: Establish relationships such as "pipeline material - susceptibility to defects" (e.g., concrete pipes - high susceptibility to cracks), "defect type - monitoring data characteristics" (e.g., siltation - decreased flow, increased pressure), and "defect level - operation and maintenance strategy" (e.g., mild siltation - routine dredging, severe corrosion - pipe replacement).
[0083] S2.3: Knowledge Graph Implementation: A knowledge graph is constructed using the Neo4j graph database, with 12,000 structured knowledge entries and 8,000 historical operation and maintenance data entries, including more than 3,000 scenario-specific rules. It supports entity query, relation reasoning, and dynamic updates. When the model detects a new correlation pattern (such as "interception basket blockage - increased risk of upstream siltation"), it automatically triggers the generation of candidate relations, which are then updated to the knowledge graph after expert review, completing the construction of a multimodal knowledge graph, which serves as the core knowledge input for subsequent knowledge-enhanced reasoning.
[0084] Example 3: Development of a Multimodal Large Model: This involves the specific implementation of the core module of the corresponding multimodal large model inference subsystem, completing the feature fusion processing of multimodal data. S3.1: Basic Model Selection and Optimization: Qwen-7B was selected as the basic large model. Transfer learning was performed based on pipeline multimodal data. 60% of the parameters at the bottom layer of the model were frozen, and 40% of the parameters at the top layer were fine-tuned to adapt to the pipeline scenario, ensuring that the model can accurately receive standardized multimodal dataset input.
[0085] S3.2: Multimodal contrastive learning: Construct intramodal contrastive loss and intermodal contrastive loss, and conduct contrastive training on image features, sensor features, and acoustic features to improve feature discriminability and alignment.
[0086] S3.3: Cross-modal fusion reasoning: An attention-gated fusion module is designed to dynamically adjust the weight ratio of each modality based on the data quality score. For example, when the image quality score is <6, the image feature weight decreases from 0.4 to 0.1, the sensor feature weight increases from 0.3 to 0.5, and the acoustic feature weight increases from 0.3 to 0.4. A modal conflict resolution module is added. When the image modality detection is "severe corrosion" while the sensor modality data shows "weak corrosion features", the knowledge graph rule "HDPE pipe corrosion susceptibility is low" is called to make a judgment and correct it to "moderate corrosion". Cross-modal feature distillation technology is introduced to distill the disease texture features of the image modality into the sensor features, improving the fusion accuracy in data sparse scenarios. Weights are assigned to different modal features (image feature weight 0.4, sensor feature weight 0.3, acoustic feature weight 0.3), and the fused feature vector is output to provide a unified feature input for subsequent reasoning.
[0087] Example 4: Knowledge Graph Enhanced Reasoning: Specific implementation of the knowledge enhancement module corresponding to the multimodal large-scale model reasoning subsystem, completing "data-knowledge" dual-driven reasoning. S4.1: Knowledge Interaction Mechanism: Core rules (such as "concrete pipe + laying years > 10 years → increased risk of cracking") are embedded into the deep reasoning layer of the model, while auxiliary rules (such as "temperature and humidity changes - corrosion rate influence") are embedded into the shallow reasoning layer to improve the efficiency of knowledge utilization. A two-way interaction mechanism between the knowledge graph and the multimodal large model is constructed. On the one hand, entity relationships in the knowledge graph are transformed into logical rules (such as "concrete pipe + laying years > 10 years → increased risk of cracking") and embedded into the model reasoning process. On the other hand, when the model detects unrecorded association patterns, candidate relationships are automatically generated and pushed to the expert review platform. After the review is approved, the relationship is updated to the knowledge graph, forming a closed loop of "reasoning-learning-iteration".
[0088] New association pattern: The "entity-relationship-entity" triple detected by the model has no complete match in the existing knowledge graph, and meets the requirements of confidence, support and logical consistency thresholds.
[0089] 3. New rule determination process: ① Extract candidate triples of "entity-relationship-entity" during model inference; ② Calculate the confidence (probability output by the model) and support (number of matching samples in the training set) of the triples; ③ Call the ontology inference engine to verify the logical consistency with the core rules and output the conflict degree; ④ If "confidence ≥ 0.85 + support ≥ 30 + conflict degree ≤ 0.10" is met, candidate relation generation is triggered; ⑤ Push to the expert review platform, and automatically update to the knowledge graph and record logs after the review is approved.
[0090] Table 2. Trigger Threshold Setting Table ; S4.2: Reasoning Correction: Based on the model's initial prediction results, relevant domain knowledge is matched using knowledge graph entity linking technology to correct contradictory predictions (e.g., the model predicts "severe corrosion of HDPE pipes," but the knowledge graph shows that HDPE pipes have low corrosion susceptibility, so it is corrected to "moderate corrosion"), and the corrected preliminary diagnostic results are output.
[0091] Example 5: Model Training and Performance Verification: This section details the implementation of the training optimization module for the multimodal large model inference subsystem and the visualization and decision support subsystem for diagnostic results, ultimately generating the final output.
[0092] S5.1: Training parameter settings: Use the PyTorch framework for training, Adam optimizer, initial learning rate 1e-5, batch size=8, number of iterations 100, loss function is cross-entropy loss + knowledge distillation loss (weight 0.3).
[0093] S5.2: Edge-Cloud Collaborative Deployment: Lightweight sub-models are deployed at the edge (reducing the number of parameters to 30% of the original model) to achieve real-time preliminary screening, direct filtering of data without defects, and uploading of suspected defect data to the cloud; the complete model is deployed in the cloud for in-depth diagnosis, reducing data transmission volume by 65% and improving inference response speed to 2 seconds / data.
[0094] S5.2: Performance Verification: The model performance metrics on the test set are shown in the table below. Compared with single-modal models (image modal model, sensor modal model), the model of this invention has significantly improved in all metrics.
[0095] Table 3. Model Performance Metrics on the Test Set ; S5.3: Application Output: The model output includes a diagnostic report (core output result) containing the type, level, location, and risk level (low, medium, high) of the defect. A heat map of pipeline defect distribution is displayed through a visualization platform, and targeted operation and maintenance strategies are recommended (supplementary output result), completing the final output.
Claims
1. A smart diagnostic method for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement, characterized in that, include: S1: Collect multimodal data of underground drainage pipes and preprocess them. Label the preprocessed data and divide the dataset to construct a standardized multimodal dataset. S2: Define the core entities and relationships between entities in the underground drainage pipeline domain, input structured domain knowledge, historical operation and maintenance data and expert experience, and build and dynamically update a multimodal knowledge graph specific to the underground drainage pipeline scenario; S3: Select a general basic large model as the backbone network, perform transfer learning and fine-tuning based on a standardized multimodal dataset, integrate multi-source information through multimodal contrastive learning, dynamic adaptive cross-modal fusion inference and modal conflict resolution techniques, construct a dynamic adaptive multimodal fusion large model, and output unified multimodal features; S4: Layered embedding of multimodal knowledge graphs, establishing a two-way interaction mechanism between multimodal knowledge graphs and dynamic adaptive multimodal fusion large model, combining unified multimodal features for knowledge-enhanced reasoning, and outputting preliminary diagnostic results; S5: Based on standardized multimodal datasets, train and optimize the model, design an edge-cloud collaborative reasoning architecture, combine multimodal knowledge graphs and preliminary diagnostic results to perform in-depth diagnosis and engineering adaptation, and output the final diagnostic results, risk assessment reports and operation and maintenance decision recommendations for various types of underground drainage pipe diseases.
2. The intelligent diagnostic method for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement according to claim 1, characterized in that, The specific process of S1 is as follows: An air-ground-pipe collaborative data acquisition scheme is adopted, using drones, pipeline inspection robots, IoT sensors, acoustic detectors, and radar detection vehicles to collect internal pipeline image and video data, acoustic detection data, flow and pressure sensor data, pipeline material and structural parameter data, and environmental temperature and humidity data. The collected multimodal data is preprocessed, the image and video data are denoised, enhanced, frame extracted and cropped, the sensor data is filled with missing values and removed outliers, the acoustic data is filtered and feature extracted, and all preprocessed data are spatiotemporally aligned and format standardized. The standardized data is labeled with disease type, disease level, disease location coordinates and data quality score. The labeled data is then divided into training set, validation set and test set according to a preset ratio to construct a standardized multimodal dataset.
3. The intelligent diagnostic method for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement according to claim 1, characterized in that, The specific process of S2 is as follows: Define the core entities in the field of underground drainage pipelines, including the pipeline body, types of defects, severity levels, monitoring equipment, operation and maintenance strategies, pipeline ancillary facilities, and operating parameters; Explore the relationships between core entities, including the relationship between pipeline material and susceptibility to defects, the relationship between defect types and monitoring data characteristics, and the relationship between defect levels and operation and maintenance strategies; A knowledge graph framework is constructed using ontology modeling tools, and structured domain knowledge, historical operation and maintenance data, and expert experience are input to form a multimodal knowledge graph. A dynamic update mechanism for the knowledge graph is also established to support the input and review of new entities and relationships.
4. The intelligent diagnostic method for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement according to claim 1, characterized in that, In S3, the specific process of selecting a general basic large model as the backbone network for transfer learning and fine-tuning is as follows: We selected an open-source general-purpose large model as the backbone network, froze some parameters of the model's lower layers, fine-tuned the parameters of the upper layers, and optimized the model's input layer to adapt to the input format of standardized multimodal datasets, enabling the model to receive and process multimodal data such as images, sensors, and acoustics.
5. The intelligent diagnostic method for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement according to claim 2, characterized in that, In S3, the specific process of multimodal contrastive learning and dynamic adaptive cross-modal fusion inference is as follows: We construct intra-modal contrast loss and inter-modal contrast loss, and perform contrast training on image features, sensor features and acoustic features to enhance the discriminative power of intra-modal features and achieve inter-modal feature alignment; The design incorporates an attention-gated fusion unit and a dynamic weight adjustment mechanism. Based on the data quality score obtained from annotation, the fusion weight ratio of each modality feature is adjusted in real time. Weighted fusion of different modality features is performed to output a preliminary unified multimodal feature.
6. The intelligent diagnostic method for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement according to claim 5, characterized in that, In S3, the specific process of modal conflict resolution is as follows: When the disease information reflected by different modal data is contradictory, the scene-specific rules in the multimodal knowledge graph are invoked for logical judgment to correct the contradictory feature information. By introducing cross-modal feature distillation technology, the deep and effective features of high-dimensional modalities are transformed into reusable feature information of low-dimensional modalities. The preliminary unified multimodal features are optimized to obtain the final unified multimodal features, thus completing the construction of a dynamic adaptive multimodal fusion large model.
7. The intelligent diagnostic method for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement according to claim 1, characterized in that, In S4, the specific process of hierarchical embedding of multimodal knowledge graphs and establishment of bidirectional interaction mechanisms is as follows: Domain knowledge in the multimodal knowledge graph is divided into core rules and auxiliary rules according to importance. The core rules are embedded into the deep reasoning layer of the dynamic adaptive multimodal fusion model, and the auxiliary rules are embedded into the shallow reasoning layer of the dynamic adaptive multimodal fusion model. A two-way interaction mechanism is established between the knowledge graph and the large model. On the one hand, the entity relationships and attribute information in the knowledge graph are transformed into constraints for model reasoning. On the other hand, when the model detects new disease association patterns not recorded in the existing knowledge graph and meets the preset threshold requirements, it automatically generates candidate relationships, which are then updated to the multimodal knowledge graph after expert review, forming a reasoning-learning-iteration closed loop.
8. The intelligent diagnostic method for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement according to claim 1, characterized in that, In S4, the specific process of knowledge-enhanced reasoning is as follows: During the model inference process, knowledge graph pruning and entity linking techniques are used to match unified multimodal features with corresponding entities in the multimodal knowledge graph to correct model inference biases caused by data sparsity or noise interference. A dual-drive framework of data-driven reasoning, knowledge-guided correction, and model feedback updates is constructed to collaboratively complete disease identification, severity assessment, and risk prediction, and output preliminary diagnostic results.
9. The intelligent diagnostic method for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement according to claim 2, characterized in that, The specific process of S5 is as follows: Train a dynamic adaptive multimodal fusion large model using the training set, monitor the training process and adjust hyperparameters using the validation set, and save the optimal model weights. The design incorporates an edge-cloud collaborative inference architecture. At the edge, a lightweight sub-model based on the optimal model weights is deployed to perform real-time preliminary screening and filter out disease-free data. At the cloud, the complete optimal model is deployed to perform in-depth diagnosis of suspected disease data. The model performance is validated using a test set. By combining historical operation and maintenance data and operational condition differences information from a multimodal knowledge graph, dynamic operation and maintenance strategies are generated, and the final diagnostic results, risk assessment reports, and operation and maintenance decision recommendations are output.
10. An intelligent diagnostic system for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement, characterized in that: The intelligent diagnostic method for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement, as described in any one of claims 1 to 9, wherein the intelligent diagnostic system for underground drainage pipelines based on multimodal fusion and knowledge graph enhancement comprises: The multimodal data acquisition subsystem consists of a drone, a pipeline inspection robot, an IoT sensor array, an acoustic detector, and a radar detection vehicle, and is used to perform multimodal data acquisition of underground drainage pipelines. The data preprocessing and storage subsystem communicates with the multimodal data acquisition subsystem to perform preprocessing, label and partition the preprocessed data, and construct a standardized multimodal dataset. The multimodal knowledge graph construction and update subsystem communicates with the data preprocessing and storage subsystem to execute S2; The multimodal large model inference subsystem communicates with the data preprocessing and storage subsystem and the multimodal knowledge graph construction and update subsystem, respectively, to execute S3 and S4, and to train and optimize the dynamic adaptive multimodal fusion large model based on the standardized multimodal dataset; The diagnostic results visualization and decision support subsystem communicates and connects with the multimodal large model reasoning subsystem to execute the design edge-cloud collaborative reasoning architecture. It combines multimodal knowledge graphs and preliminary diagnostic results to perform in-depth diagnosis and engineering adaptation, outputting the final diagnostic results, risk assessment reports and operation and maintenance decision recommendations for various types of underground drainage pipe diseases, and providing visualization display functions.