A global hot water system fault diagnosis and early warning method and system based on multi-modal perception fusion and knowledge graph reasoning
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-08-11
AI Technical Summary
[0009]本发明提供了一种基于多模态感知融合与知识图谱推理的全域热水系统故障诊断与预警方法及系统,以解决现有技术中存在的因感知维度单一、数据分析孤立,而导致的对热水系统早期隐蔽性故障预警能力不足、故障发生后根源定位困难、报警精准度低以及运维模式被动等技术问题
[0013]相较于现有技术,本发明通过引入振动、声学、红外热成像等多模态数据,能够从多个物理维度全面感知设备的微观运行状态,这些数据对于常规温压流等过程参数无法反映的设备早期性能衰退迹象,如轴承的初期磨损、换热器内部的轻微结垢、阀门的微小内漏等,具有极高的敏感度。这就使得本发明能够将故障的发现窗口期大幅度提前,实现了真正意义上的预测性维护,从根本上提升了早期预警的能力。
Smart Images

Figure CN121502602B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of intelligent operation and maintenance and Internet of Things (IoT) technology, and in particular to an intelligent monitoring and fault early warning technology applied to a whole-domain hot water supply network. More specifically, this invention discloses a method and system for fault diagnosis and early warning of a whole-domain hot water system based on multimodal perception fusion and knowledge graph reasoning. Background Technology
[0002] Urban centralized heating systems and regional hot water supply networks are crucial infrastructure in modern society. Their stable, efficient, and safe operation directly impacts people's livelihoods and energy conservation. Traditionally, the operation and maintenance of such systems have relied primarily on Supervisory Control and Alarm (SCADA) systems based on setpoints and fixed thresholds, along with periodic manual inspections. SCADA systems remotely monitor macroscopic operating parameters of the system by deploying sensors such as temperature, pressure, and flow at key nodes. When monitored values exceed preset safety ranges, the system triggers an alarm, notifying maintenance personnel for intervention. However, as hot water supply networks become increasingly large and complex, this traditional operation and maintenance model has revealed increasingly serious limitations.
[0003] In recent years, with the development of big data and machine learning technologies, several improvement schemes have emerged aimed at enhancing the intelligence of heating systems. For example, Chinese invention patent publication number CN118331184A discloses a smart heating system and control method. This scheme collects status information of the controlled equipment through a data acquisition unit and feeds it back to the field controlled equipment. Intelligent temperature regulating valves also transmit signal data to the field controlled equipment via a network. This data is then transmitted to a management platform through communication between the field controlled equipment and the management platform. The management platform uses intelligent modeling and intelligent optimization algorithms to integrate, allocate, and calculate the collected data. Based on the calculated results, it automatically generates operational suggestions and prompts. The field controlled equipment then controls the status of its own equipment, thereby achieving smart heating and realizing process management and operational management of the entire heating system, thus improving the management methods of the heating system.
[0004] However, a deeper analysis of existing technologies, including the aforementioned patents, reveals several deep-seated technical pain points that remain unresolved when addressing early, subtle, and gradually changing faults, as well as system-level interlocking faults. First, the perception dimension is limited, resulting in a severe deficiency in early warning capabilities for "soft faults." Existing technologies primarily rely on conventional process parameters such as temperature, pressure, flow rate, and energy consumption. These parameters are effective for responding to already occurring, severe "hard faults," such as pipe bursts or sudden pump shutdowns. However, they are extremely insensitive to the numerous, more prevalent, and equally damaging progressive "soft faults." For example, initial pitting or wear in circulating water pump bearings, slow scaling inside heat exchangers leading to decreased heat exchange efficiency, and minute, undetectable leaks in the pipeline network—these faults, in their nascent stages, cause very slight changes in macroscopic temperature and pressure parameters, often masked by normal operating fluctuations, seasonal variations, or the randomness of user loads, resulting in extremely low signal-to-noise ratios. When traditional monitoring systems finally trigger alarms due to parameters exceeding thresholds, equipment damage has often progressed to a relatively severe stage, resulting not only in high repair costs but also potentially leading to unplanned downtime and significant economic losses. Current technologies generally lack the ability to perceive and analyze multimodal and multidimensional data that can directly and sensitively characterize the mechanical health of equipment (e.g., vibration characteristics), fine-grained fluid dynamics (e.g., acoustic characteristics), and the microscopic distribution of thermal conditions (e.g., infrared thermal imaging).
[0005] Second, there is a lack of system-level fault correlation analysis and root cause tracing capabilities. Existing monitoring and diagnostic systems typically treat each sensor and each device as an independent monitoring unit for analysis. When an abnormal alarm occurs at a monitoring point in the system, the system itself cannot intelligently and automatically analyze the root cause behind the anomaly. For example, a monitoring point in the pipeline reports a sudden pressure drop, which could be caused by multiple reasons: the pressure sensor at that point may have malfunctioned or drifted; the output power of the upstream booster pump may have decreased or cavitation may have occurred; or a large downstream user may have suddenly opened a bypass valve. Under the current technological framework, the system cannot effectively distinguish between these. The data from various devices and subsystems form isolated "data silos," lacking systematic modeling of their physical connections (such as the upstream and downstream topology of pipelines), logical control relationships (such as the linkage strategy between pumps and valves), and functional impact relationships. Therefore, once an alarm occurs, experienced operation and maintenance engineers still need to intervene, relying on their personal knowledge and historical experience to conduct extensive manual troubleshooting and data comparison across multiple subsystems. The fault location process is time-consuming, labor-intensive, and inefficient. This problem is particularly severe in large-scale hot water supply systems with complex pipe networks and tens of thousands of measuring points.
[0006] Third, the alarm accuracy is low, with both false alarms and missed alarms existing. Due to the limited scope of perception and the isolation of analytical capabilities, the alarm mechanisms of existing systems are often too crude and rigid. Alarm methods based on fixed thresholds are easily affected by factors such as normal operating condition switching, seasonal changes between winter and summer, and sudden changes in user water usage patterns during holidays, resulting in a large number of invalid false alarms. These false alarms not only distract maintenance personnel but also lead to alarm fatigue in the long run, reducing their sensitivity and trust in alarm signals. Consequently, they may react slowly when real faults occur and ignore important alarm information. At the same time, there is a serious risk of missed alarms for the aforementioned "soft faults." This coexistence of false alarms and missed alarms results in a very low overall signal-to-noise ratio for existing early warning systems, significantly reducing their practical application value.
[0007] Fourth, the operation and maintenance model is passive, lacking proactive predictive maintenance capabilities. Generally speaking, whether based on simple threshold-based alarms or fault diagnosis based on historical data mining, existing technical solutions mostly remain at the level of "post-event response" or "critical alarm." They are essentially still passive operation and maintenance models, i.e., waiting for problems to occur or be about to occur before taking action. The system cannot effectively predict the future health status of equipment based on the evolution trends of multi-dimensional equipment status data, nor can it accurately estimate its remaining effective lifespan. This leads to a waste of maintenance resources and cannot fundamentally prevent catastrophic failures, making it difficult to effectively guarantee the overall operational reliability of the system.
[0008] Therefore, a novel technical solution is needed in this field to effectively solve the above-mentioned series of technical problems and realize a new intelligent operation and maintenance model for the whole-area hot water supply system, from point monitoring to network insight, from passive response to proactive prediction, and from superficial alarms to root cause location. Summary of the Invention
[0009] This invention provides a method and system for fault diagnosis and early warning of a global hot water system based on multimodal perception fusion and knowledge graph reasoning, in order to solve the technical problems in the prior art, such as insufficient early warning capability for early hidden faults in hot water systems, difficulty in locating the root cause after the fault occurs, low alarm accuracy, and passive operation and maintenance mode, which are caused by the single perception dimension and isolated data analysis.
[0010] A first aspect of the present invention provides a method for fault diagnosis and early warning of a hot water system based on multimodal perception fusion and knowledge graph reasoning, the method comprising the following steps. First, a multimodal data acquisition step is performed, acquiring multimodal operational data of the hot water system based on multiple sensors deployed at key nodes of the system. The multimodal operational data includes at least: vibration data characterizing the mechanical state of the equipment, acoustic data characterizing the fluid state, infrared thermal imaging data characterizing the thermal condition of the equipment, and process parameter data characterizing the system's operating state. Second, a state feature vector generation step is performed, preprocessing and feature engineering the multimodal operational data to extract multiple features characterizing the health state of key nodes, and fusing these features into a multimodal state feature vector. Third, a single-point anomaly detection step is performed, using a preset unsupervised anomaly detection model to analyze the multimodal state feature vector in real time, calculating an anomaly score. When the anomaly score exceeds a preset threshold, the key node where the anomaly has occurred is identified. Then, the root cause reasoning step is executed. Once the critical node where the anomaly occurred is identified, a multi-hop related entity retrieval is performed in the pre-constructed hot water system knowledge graph, starting from the critical node. For each retrieved related entity, a candidate fault hypothesis is generated, and multimodal operational data of the related entities is actively invoked for cross-validation. Based on the validation results, the corresponding confidence level is calculated for each candidate fault hypothesis. Finally, the early warning information generation step is executed. Based on the candidate fault hypothesis with the highest confidence level, the root cause and type of the fault are determined, and a structured early warning information document containing fault root cause location, diagnostic evidence chain, and potential impact domain analysis is generated.
[0011] In another aspect, the present invention provides a fault diagnosis and early warning system for a hot water system based on multimodal perception fusion and knowledge graph reasoning. The system includes: a multimodal perception module, whose function is to collect multimodal operating data of key nodes in the hot water system, wherein the multimodal operating data includes at least vibration data characterizing the mechanical state of the equipment, acoustic data characterizing the fluid state, infrared thermal imaging data characterizing the thermal condition of the equipment, and process parameter data characterizing the system operating state; a data processing and fusion module, whose function is to perform preprocessing and feature engineering, extract multiple features characterizing the health state of key nodes from the multimodal operating data, and fuse them into a multimodal state feature vector; and an anomaly detection module, configured to use a preset unsupervised anomaly detection model to detect multimodal anomalies. The system performs real-time analysis of state feature vectors to calculate anomaly scores and identifies key nodes where anomalies occur when the anomaly scores exceed a preset threshold. A fault reasoning engine module, connected to a pre-built knowledge graph of a hot water system, is configured to respond to the identification of key nodes by performing multi-hop related entity retrieval within the knowledge graph, starting from that node. It then generates candidate fault hypotheses for the retrieved related entities, actively calls multimodal operational data of these related entities for cross-validation, and finally calculates the confidence level for each candidate fault hypothesis. Finally, an application service module is configured to determine the root cause and type of the fault based on the candidate fault hypothesis with the highest confidence level, and generates structured early warning information including fault root cause location, diagnostic evidence chain, and impact domain analysis.
[0012] The technical problem this invention aims to solve is how to achieve precise perception of the fine-grained health status of key equipment in a complex and vast global hot water supply network, and on this basis, realize early and accurate warnings of various hidden faults, while also enabling rapid and accurate intelligent source tracing and location of fault roots across the entire network. To address this technical problem, this invention first establishes a deep perception system capable of finely characterizing the microscopic mechanical and fluid dynamic states of equipment by integrating multimodal data such as vibration, acoustics, and infrared thermal imaging, fundamentally improving the ability to detect the emergence of "soft faults." Building upon this, this invention does not stop at analyzing single-point data, but constructs a digital twin knowledge graph depicting the physical topology, logical control, and functional relationships of the entire hot water supply network. This integrates previously discrete equipment and data into a computable and reasoning-enabled intelligent network, completely breaking down "data silos." Most importantly, when the perception layer detects an anomaly, this invention does not issue an isolated threshold alarm, but automatically launches a multi-hop association reasoning engine based on a knowledge graph. This engine simulates the diagnostic thinking of domain experts, generating fault hypotheses, cross-validating multi-source evidence, and accurately tracing the root cause along the graph path, thereby improving the early warning and accuracy of fault diagnosis.
[0013] Compared to existing technologies, this invention, by introducing multimodal data such as vibration, acoustics, and infrared thermal imaging, can comprehensively perceive the microscopic operating status of equipment from multiple physical dimensions. This data is highly sensitive to early signs of performance degradation that conventional process parameters such as temperature, pressure, and flow cannot reflect, such as initial wear of bearings, minor scaling inside heat exchangers, and minute internal leaks in valves. This allows the invention to significantly advance the fault detection window, achieving true predictive maintenance and fundamentally improving early warning capabilities.
[0014] Secondly, this invention, by constructing and utilizing a knowledge graph of the hot water system, connects previously isolated devices and data points into an intelligent network with topological structure and logical relationships, completely breaking down the "data silos." When a fault occurs, the associative reasoning capability based on the knowledge graph enables the system to perform logically clear and evidence-rich chain analysis and tracing, just like an experienced domain expert. It can accurately distinguish whether the problem is caused by a problem with the sensor itself, a fault in the actuator component, or a conductive anomaly caused by other related devices, thus improving the accuracy and efficiency of fault root cause location.
[0015] Furthermore, the multi-source data cross-validation mechanism in this invention effectively reduces the false alarm rate of the system, thereby enhancing its reliability. This mechanism avoids false alarms caused by instantaneous fluctuations in data from a single sensor or switching between normal system operating conditions. The system only generates a high-confidence warning when features from multiple independent sensor data sources from different dimensions, along with the system topology, collectively point to a potential fault. This decision-making approach based on multi-evidence fusion improves the signal-to-noise ratio of alarms, allowing maintenance personnel to place greater trust in the system's warning results. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the overall architecture of a global hot water system intelligent fault diagnosis and early warning system according to an embodiment of the present invention.
[0018] Figure 2 This is a general flowchart of a method for intelligent fault diagnosis and early warning of a whole-area hot water system according to an embodiment of the present invention.
[0019] Figure 3This is a schematic diagram of the hardware structure of a multimodal sensing terminal for the sensing layer according to an embodiment of the present invention.
[0020] Figure 4 This is a structural example diagram of a specific hot water supply system knowledge graph according to an embodiment of the present invention.
[0021] Figure 5 This is a schematic diagram of the core fault tracing and reasoning process based on a knowledge graph, as described in an embodiment of the present invention.
[0022] Figure 6 This is a schematic diagram illustrating a specific application scenario of the present invention in monitoring circulating water pumps in a heat exchange station, according to one embodiment of the present invention. Figure 7 This is a schematic diagram comparing the performance of the method of the present invention with that of a conventional method, according to an embodiment of the present invention. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0024] Please refer to the following first. Figure 1 This illustrates the overall architecture of a preferred embodiment of the intelligent fault diagnosis and early warning system for a global hot water system according to the present invention. The system adopts a collaborative design concept of "end-edge-cloud," and from bottom to top comprises: a perception layer 100, a network transmission layer 200, a cloud platform layer 400, and an application service layer 500.
[0025] (1) Perception layer 100 The sensing layer 100 is the data foundation of the entire system, responsible for comprehensive, accurate, and multi-dimensional data collection on the status of various key nodes and equipment in the hot water supply system. Its core deployment is the multimodal sensing terminal 300 (its specific structure will be discussed later). Figure 3 (Details in the text) This terminal integrates a series of sensors, which may include: 1. Conventional process parameter sensors: Temperature sensors (such as PT100 platinum resistance temperature sensors): deployed at boiler outlet / return water inlets, heat exchanger outlets at each stage, main pipes, and user inlets, etc., to measure fluid temperature.
[0026] Pressure sensors (such as piezoresistive pressure sensors): deployed at locations such as water pump inlets / outlets, main pipelines, and before and after pressure reducing valves in high-rise buildings to monitor pipeline pressure.
[0027] Flow sensors (such as electromagnetic flow meters or ultrasonic flow meters): deployed on the main pipeline and key branches of the pipeline network to measure the flow rate of hot water in real time.
[0028] 2. Deployed multimodal state perception sensors: Vibration accelerometer: Preferably a triaxial IEPE (Integrated Electronics Piezo-Electric) piezoelectric accelerometer. It is tightly coupled and installed near the bearing housing of the motor or pump body in critical rotating equipment (such as primary / secondary circulation pumps, make-up water pumps). Its core function is to capture the microscopic mechanical vibration signals of the equipment in real time. By analyzing the spectrum of the vibration signal, mechanical "soft faults" such as bearing wear, rotor imbalance, equipment misalignment, and foundation loosening can be diagnosed at a very early stage. These faults are not initially reflected in conventional temperature and pressure parameters.
[0029] High-frequency acoustic sensors / hydrophones: preferably contact-type, wide-bandwidth (e.g., 1Hz-100kHz) acoustic emission sensors or hydrophones. They are installed on the outer wall of critical pipe sections. Their core function is to passively listen to the acoustic signals of the fluid inside the pipe. Normal fluid flow produces a smooth background noise, while minor leaks in the pipe, incomplete closure or opening of valves, and cavitation phenomena in water pumps all generate unique, high-frequency abnormal acoustic signals. By analyzing these acoustic signatures, precise location of leaks in the pipe network and diagnosis of abnormal fluid conditions can be achieved.
[0030] Infrared thermal imaging cameras are deployed in key indoor equipment areas such as boiler rooms and heat exchange stations. Their core function is to acquire the two-dimensional temperature field distribution of equipment and pipelines non-contactly. Through intelligent image analysis, they can effectively identify thermal anomalies such as motor overheating, poor contact of electrical cabinet contacts, and damage to pipeline insulation layers, achieving a "thermal health" status scan of specific areas.
[0031] Water quality sensors, such as total dissolved solids (TDS) sensors and conductivity sensors, are deployed at the water inlet or in the circulation pipeline. Their core function is to monitor the chemical properties of the water. Changes in water quality (such as increased hardness) are one of the root causes of problems such as heat exchanger scaling and pipe corrosion. Incorporating water quality data into the analysis can provide a fundamental explanation for abnormal data from other sensors.
[0032] (2) Network Transport Layer 200 This layer is responsible for the stable and efficient transmission of the massive amounts of heterogeneous data collected by the perception layer 100 to the cloud platform. Depending on the characteristics of different scenarios and data types, multiple communication technologies can be used in combination. For scenarios with centralized data and stable power supply, such as heat exchange stations, industrial Ethernet can be used for high-bandwidth, low-latency data transmission. For pipeline monitoring points distributed throughout the city where cabling is difficult, LPWAN (Low Power Wide Area Network) technologies, such as LoRa or NB-IoT, can be used to transmit low-frequency, small data packet signals such as temperature, pressure, and flow. Image / video stream data generated by the infrared thermal imaging camera 103 needs to be transmitted via 4G / 5G wireless networks or fiber optics.
[0033] (3) Cloud platform layer 400 The cloud platform layer 400 is responsible for data storage, processing, analysis, and decision-making. Deployed on public or private cloud servers, it consists of multiple closely cooperating functional modules: 401. Data Access and Preprocessing Module: Responsible for receiving data from the network transport layer 200 via protocols such as MQTT. It performs cleaning (removing outliers and filling missing values), time alignment (ensuring data from different sensors are synchronized in timestamps), and normalization on the raw data. More importantly, this module performs preliminary feature engineering, such as: performing Fast Fourier Transform (FFT) on the time-domain waveform signals acquired by vibration sensors to extract frequency domain features such as amplitude and energy at characteristic frequencies like rotational and harmonic frequencies; and performing wavelet transform on signals from acoustic sensors to extract transient impact features reflecting leakage or cavitation.
[0034] 402. Knowledge Graph Construction and Management Module: The core function of this module is to digitize and model the physical world's hot water supply network into a knowledge graph, such as... Figure 4 As shown.
[0035] Entities (Nodes / Entities): Define all critical objects in the system. For example, each water pump (Pump_ID:P001), each valve (Valve_ID:V089), each pipe section (Pipe_ID:Seg123), each heat exchange station (Station_ID:S05), and each user building is defined as an entity. Each entity has rich static and dynamic attributes, such as: equipment model, installation date, rated power, last maintenance time, real-time temperature, etc.
[0036] Edges / Relations: Define various relationships between entities. For example: Physical connection (connectsTo): such as Pipe_ID:Seg123 connectsTo Pump_ID:P001, and may have the attribute of direction: downstream.
[0037] Control relationships (controls): such as PLC_ID:C01 controls Valve_ID:V089.
[0038] BelongsTo: For example, Pump_ID:P001 belongsTo Station_ID:S05.
[0039] Network topology construction: The network topology has diverse data sources. It can automatically build the basic network topology by batch importing structured data such as equipment ledgers (BOM list), BIM (Building Information Modeling) data, and P&ID (Pipes and Instrumentation Diagram). Then, domain experts can supplement and verify the knowledge rules.
[0040] Preferably, the knowledge graph construction process may include the following automated and semi-automated steps: First, data extraction and entity recognition steps are performed. This involves developing adapter programs for different data sources (such as Excel files of equipment ledgers, IFC files of BIM models, and XML or DWG files of P&ID diagrams) to automatically extract equipment information, pipeline connection information, and control logic. Named Entity Recognition (NER) technology in Natural Language Processing (NLP) can be used to identify entities such as equipment names and fault phenomena from unstructured maintenance logs or equipment manuals. Second, relationship extraction and knowledge fusion steps are performed. Based on the extracted entities and topological information in the data sources, predefined mapping rules or association algorithms are used to automatically construct relationships such as 'physical connections' and 'belongings' between entities. For example, by parsing the connection information of the P&ID diagram, the connectsTo relationship between pipes and pumps / valves is automatically generated. For the same entity from different data sources (e.g., the same pump existing in both the BIM model and the equipment ledger), knowledge fusion is performed using entity alignment technology (such as methods based on attribute similarity calculation) to avoid information redundancy. Secondly, in the knowledge injection step of executing expert rules, the system provides a visual human-computer interaction interface. Domain experts can input their diagnostic knowledge and experience (e.g., 'If the vibration kurtosis index of a water pump is greater than 4 and its outlet pressure is stable, then it is highly likely to be a bearing failure') into the system on this interface, either in the form of 'IF-THEN' natural language or by dragging and dropping links. The system backend automatically converts these rules into reasoning paths in the knowledge graph or SWRL (Semantic Web Rule Language) rules, serving as the knowledge foundation for the fault reasoning engine.
[0041] 403. Multimodal Sensing Fusion Module: This module is responsible for processing and fusing heterogeneous data from different sensors to form a more comprehensive and robust device health assessment than any single data source. Its operation can employ a layered fusion strategy: Feature layer fusion: For the same monitored object (e.g., a water pump), the frequency domain feature vectors extracted from vibration sensors (such as the energy of each harmonic and kurtosis index), the temperature change rate obtained from temperature sensors, the outlet pressure fluctuation coefficient obtained from pressure sensors, and the sound pressure level (SPL) extracted from acoustic sensors are concatenated into a high-dimensional, unified feature vector. This vector can be regarded as a multimodal health snapshot of the water pump at the current moment.
[0042] Decision-level fusion: As an alternative or supplementary approach, a single model (such as an SVM classifier for vibration data or a threshold model for temperature data) can be used to derive a preliminary health assessment. Then, a weighted voting method or DS evidence theory can be employed to fuse these preliminary assessments, resulting in a final health evaluation conclusion. This approach is logically clear and easy to interpret. In this embodiment, feature-level fusion is preferred because it better captures the inherent correlations between different modal data.
[0043] 404. Fault Reasoning and Early Warning Engine: Responsible for the entire process from anomaly detection to fault root cause localization. When a device's multimodal health snapshot is identified as abnormal by an upstream machine learning model (such as an autoencoder or an isolation forest), this engine does not immediately issue an alarm. Instead, it initiates an automated reasoning process based on a knowledge graph (402). For the core workflow of this process, please refer to [link / reference needed]. Figure 5 A brief description is as follows: 1. Abnormal event instantiation: The detected abnormality (such as the vibration kurtosis index of Pump_ID:P001 > 4.5) is temporarily added to the knowledge graph as an "abnormal event" node.
[0044] 2. Initiate multi-hop path search: Starting from the "abnormal event" node, a graph algorithm (such as bidirectional breadth-first search) is used to perform multi-hop traversal on knowledge graph 402 to find first-order, second-order, or even higher-order related entities that have predefined relationships (such as connectsTo, poweredBy, controls) with Pump_ID:P001. For example, it will find the upstream power switch Switch_ID:S23, the downstream pipe Pipe_ID:Seg123, and the logically related frequency converter VFD_ID:V01.
[0045] 3. Hypothesis Generation and Evidence Gathering: For each related entity found, the engine generates a possible causal hypothesis for the failure. For example: Assumption 1 (self-fault): Pump_ID:P001 has a damaged bearing.
[0046] Assumption 2 (upstream transmission): The downstream Pipe_ID:Seg123 is blocked, causing the pump outlet pressure to rise and resulting in abnormal vibration.
[0047] Assumption 3 (Control system failure): The output frequency of the associated VFD_ID:V01 is unstable, causing the pump speed to fluctuate and resulting in abnormal vibration.
[0048] 4. Cross-validation and hypothesis pruning: The engine will immediately and proactively call upon the multimodal real-time data of these related entities for cross-validation. For example, to verify hypothesis 2, it will check whether the pressure sensor readings along Pipe_ID:Seg123 have indeed increased, and whether the acoustic sensors on that pipe segment have captured eddy noise caused by blockage. If both the pressure and acoustic data are normal, then the confidence level of hypothesis 2 will be greatly reduced (pruning).
[0049] 5. Output high-confidence conclusions: After a series of hypothesis generation and pruning, the engine will eventually output one or more fault diagnosis conclusions with the highest confidence, such as: "Diagnosis conclusion: Pump_ID:P001 is the root cause of the fault. Evidence: Its vibration kurtosis exceeds the standard, and its upstream power supply, downstream pipeline and control system data are all normal. Inference: It is highly likely (95%) that it is early wear of the bearing." (4) Application Service Layer 500 This layer serves as the human-computer interaction window, responsible for transforming the analysis results from the cloud platform layer 400 into valuable information and instructions for operations and maintenance personnel. Its functions may include: Visualized Monitoring Center: This center dynamically displays the operational status of the entire hot water supply network on a large screen or web interface, using a 2D or 3D GIS / BIM model. When an alert is triggered, it highlights the root cause device of the fault, the predicted fault propagation path, and the affected user range on the model.
[0050] Intelligent Alarms and Push Notifications: Detailed and clearly defined early warning messages are pushed to the mobile devices (mobile app, smartwatch) of maintenance personnel, rather than simple "XX exceeds limits" alarms. The information includes: faulty equipment location, possible causes of the fault, confidence level, relevant data evidence (such as abnormal vibration spectrum diagrams), and system-recommended handling plans.
[0051] Automatic work order generation and dispatch: For high-confidence alerts, the system can link with the operation and maintenance management system (such as EAM / CMMS) to automatically create maintenance work orders and intelligently dispatch the work orders to the most suitable engineers based on the fault type, location, and the skills and schedules of the operation and maintenance personnel.
[0052] Health Status Assessment and Life Prediction Reports: Regularly generate health status assessment reports for each critical piece of equipment and predict its remaining effective life (RUL) based on historical data trends using algorithms such as Long Short-Term Memory (LSTM) networks, providing data support for developing long-term maintenance and equipment replacement plans.
[0053] Example 1 This embodiment details the implementation process of a fault diagnosis and early warning method for hot water systems based on multimodal perception fusion and knowledge graph reasoning. Please refer to... Figure 2 The method is Figure 1 The functional implementation and dynamic execution process of the system architecture shown mainly includes the following steps.
[0054] Step S201 involves performing multimodal heterogeneous data acquisition and transmission. This step forms the data foundation for the entire method and corresponds to... Figure 1 The continuous operation of the perception layer 100 and the network transmission layer 200. In this embodiment, the system deploys such as key nodes selected in the hot water supply network, such as the main supply and return water pipes of the primary pipeline, the primary and secondary equipment groups in the heat exchange station, and key branches and end-user entrances of the secondary pipeline. Figure 3 The multimodal sensing terminal 300 is shown. These terminals collect data in parallel and continuously according to a preset sampling strategy. Specifically, the sampling parameters differ for different types of data. For process parameters with a slow rate of change, such as water temperature measured by a PT100 platinum resistance temperature sensor or water quality measured by a total dissolved solids (TDS) sensor, the sampling frequency can be set to a lower level, for example, collecting and reporting data once per minute, which can meet the monitoring requirements.
[0055] However, for dynamic parameters that characterize the microscopic health status of equipment, high-frequency sampling is required. Preferably, for triaxial IEPE piezoelectric accelerometers deployed near the bearing housing of a secondary circulation pump, such as the pump casing, the sampling frequency must be much higher than the pump's rotational frequency and its main fault characteristic frequencies. For example, for a pump with a rated speed of 2950 rpm (approximately 49.17 Hz), in order to accurately capture the high-frequency impact signals generated by early damage to components such as the bearing inner and outer rings and balls, the sampling frequency of its vibration signal should follow the Nyquist theorem with sufficient margin, preferably set to 10 kHz to 50 kHz, and in this embodiment, it is set to 40.96 kHz. Similarly, for contact-type, wideband (e.g., 1 Hz-100 kHz) acoustic emission sensors installed on the outer wall of critical pipe sections, in order to capture the high-frequency "hissing" sound generated by minor leaks inside the pipe or the cavitation bubble bursting sound generated by pump cavitation, the sampling frequency also needs to be set above 20 kHz. For infrared thermal imaging cameras deployed in heat exchange stations, their operating modes can be set according to actual needs. For example, in regular inspection mode, it can be set to take a panoramic thermal image of all key equipment (such as water pump motors, frequency converters, and power distribution cabinet contacts) within the monitored area every 30 minutes and upload it. In linkage trigger mode, when other sensors (such as vibration sensors) detect an anomaly, the system can automatically instruct the camera to point at the abnormal equipment and take continuous pictures at a higher frequency (such as once per minute) to capture the thermal characteristic changes during the fault development process. All raw data collected will be appended with accurate GPS / NTP synchronization timestamps, unique device IDs, and sensor ID information, and will be aggregated to the cloud platform layer 400 in real time or near real time through the corresponding network transmission layer 200 (e.g., low-frequency process parameter data via NB-IoT or LoRa, high-frequency vibration and acoustic data via industrial Ethernet, and image and video data via 5G or fiber optics) for subsequent processing.
[0056] Step S202 involves data preprocessing and feature engineering. This step is performed in the data access and preprocessing module 401 of the cloud platform. Its purpose is to transform the raw, multi-source, heterogeneous data streams into standardized, more information-dense structured feature data, providing high-quality input for subsequent intelligent analysis models. When the cloud platform receives data via message queue protocols such as MQTT, it first performs a series of data preprocessing operations. During the data cleaning phase, the system uses methods such as the 3-sigma principle or box plots to identify and remove obvious outliers (outliers) caused by instantaneous sensor interference or network transmission errors. For missing data points, methods such as linear interpolation, polynomial interpolation, or mean-based filling based on adjacent time series are used to complete the data, ensuring data integrity. During the time alignment phase, the system performs strict timestamp alignment of all sensor data streams using a unified time base. This step is crucial, ensuring that at any given analysis point in time, a complete data snapshot containing readings from all sensors at that moment is obtained, a prerequisite for multimodal fusion analysis.
[0057] The next crucial step is feature engineering, which is the core step of transforming raw data into meaningful information. Specifically, for time-domain waveform signals, such as vibration and acoustic data, the system calculates a series of time-domain statistical characteristics. These characteristics include, but are not limited to: mean, variance, standard deviation, peak value, RMS, kurtosis, margin, kurtosis, and skewness. Among these, the RMS reflects the overall energy of the signal, while higher-order statistics such as kurtosis and margin are particularly sensitive to the impact components in the signal and are effective indicators for diagnosing impact failures in components such as bearings and gears.
[0058] Furthermore, the system performs frequency domain analysis on the time-domain waveform signal. Preferably, this is done by performing a Fast Fourier Transform (FFT) to convert the signal from the time domain to the frequency domain. Then, key frequency domain features are extracted from the spectrum. For example, for the vibration signal of a water pump, the system focuses on its rotational frequency (1X), second harmonic (2X), blade pass frequency (BPF), and the energy or amplitude of its harmonic components. Rotor imbalance usually leads to a significant increase in the rotational frequency component, misalignment causes an increase in the second harmonic component, and blade damage is related to the blade pass frequency. For bearing failures, the system also automatically identifies theoretical fault characteristic frequencies such as the inner race fault frequency (BPFI), outer race fault frequency (BPFO), ball fault frequency (BSF), and cage fault frequency (FTF), and monitors the occurrence and amplitude changes of these frequencies and their sidebands. Alternatively, for non-stationary signals, wavelet transform (WT) or wavelet packet decomposition (WPT) can be used for time-frequency analysis to capture the local characteristics of the fault signal in time and frequency more precisely, such as transient impact sound waves caused by minute leaks.
[0059] For thermal images captured by infrared thermal imaging cameras, the system extracts features using image processing algorithms. First, image segmentation techniques (such as threshold-based or edge-detection methods) are used to automatically identify the outlines of key equipment in the image, such as motors, pumps, and power distribution cabinets. Then, for each equipment area, key thermal characteristics are calculated, including maximum temperature, minimum temperature, average temperature, temperature gradient, temperature variance, and the temperature difference from the ambient background temperature. These features can effectively identify thermal anomalies such as equipment overheating, insulation damage, and poor electrical connections.
[0060] After step S202, the raw data streams from different sensors, with varying sampling frequencies and data formats, are successfully converted into a structured, relatively uniform multimodal feature vector stream. Each vector represents a comprehensive health profile of a key node at a given moment.
[0061] Step S203 involves performing device-level single-point anomaly detection. This step, also executed at the cloud platform layer, aims to perform preliminary screening and assessment of the health status of each individual device, quickly identifying potential anomalies. Considering the scarcity of data samples with clear and diverse fault labels in the early stages of system operation, training a fault diagnosis model using traditional supervised learning methods (such as classifiers) is impractical. Therefore, in this embodiment, unsupervised or semi-supervised anomaly detection algorithms are preferably used. These algorithms only require training using data from a large amount of equipment under normal operating conditions.
[0062] A preferred model is a deep learning-based autoencoder neural network. This model consists of an encoder and a decoder. During training, the model is fed a large number of multimodal feature vectors under normal operating conditions. The encoder compresses these vectors into a low-dimensional latent representation, and the decoder attempts to reconstruct the original input vector from this latent representation. The optimization objective of the model is to minimize the reconstruction error between the input vector and the reconstructed vector. In this way, the autoencoder learns the inherent patterns and structures of "normal" data. During real-time monitoring, the latest multimodal feature vectors of the device are input into this pre-trained model. If the device is in a normal state, its feature vector patterns are similar to the training data, and the model can reconstruct it with a small reconstruction error. Conversely, if the device malfunctions, its feature vector patterns will deviate from the normal range, and the model will be unable to reconstruct it well, resulting in a reconstruction error significantly higher than normal. The magnitude of this reconstruction error can be directly used as an anomaly score to quantify the degree to which the current device state deviates from its normal pattern.
[0063] A preferred autoencoder model is the stacked autoencoder, whose network structure can be specifically designed as follows: Input layer: The number of neurons is equal to the dimension of the multimodal feature vector generated in step S202, for example, d_in = 128.
[0064] The encoder consists of multiple sequential, fully connected layers with decreasing numbers of neurons. For example, it can be designed as three layers: 128 -> 64 -> 32. Each layer uses a non-linear activation function, such as the rectified linear unit (ReLU), to learn complex non-linear patterns in the data.
[0065] Latent representation layer: The last layer of the encoder is a low-dimensional representation of the compressed data, for example, with a dimension of d_latent = 16.
[0066] Decoder: The structure is symmetrical to the encoder, consisting of fully connected layers with an increasing number of neurons, designed to reconstruct the original input from the latent representation. For example, it might be designed as a three-layer system: 16 -> 32 -> 64 -> 128.
[0067] Output layer: The number of neurons is the same as that of the input layer, and a linear activation function is usually used.
[0068] The training data requirements, parameter settings, and optimization methods for the model are as follows: Training data: It is necessary to collect multimodal data of the equipment under various normal operating conditions (covering different seasons and different loads) for a sufficiently long period of time, and process it into a large number of normal feature vector samples, such as at least 100,000 samples, to ensure that the model can fully learn normal patterns.
[0069] Loss function: The optimization goal of the model is to minimize the reconstruction error between the input and the output. Therefore, the loss function is usually the mean squared error (MSE).
[0070] Optimizer: The Adaptive Moment Estimation (Adam) optimizer is employed, which combines the advantages of AdaGrad and RMSProp optimization algorithms. It can adaptively adjust the learning rate of each parameter, and features fast convergence and stable performance. The initial learning rate can be set to 0.001.
[0071] Training process: Mini-batch gradient descent is used for training, with a batch size of, for example, 256. To prevent overfitting, dropout layers can be added between fully connected layers, and an early stopping strategy can be used, i.e., training is automatically stopped when the model's performance on the validation set no longer improves for several consecutive epochs.
[0072] During the real-time monitoring phase, the quantified reconstruction error can be obtained by calculating the MSE between the input vector X and the reconstructed vector X' output by the model, and this error can be used as the anomaly score. Through the specific model structure design, parameter settings, and training strategies described above, a high-performance, robust unsupervised anomaly detection model can be constructed, providing reliable input for subsequent accurate inference.
[0073] Another alternative model is the Isolation Forest algorithm. This algorithm is based on an intuitive principle: outliers, due to their "sparse and distinct" nature, are generally easier to isolate than normal points. The algorithm constructs multiple binary trees by randomly selecting features and split points. In these trees, outliers are typically located closer to the root node, meaning they require fewer splits to be assigned to a single leaf node. Therefore, the average length of an isolated path for a data point can be used as a measure of its anomaly score; the shorter the path, the greater the likelihood of it being an anomaly. The Isolation Forest algorithm is computationally efficient, insensitive to high-dimensional data, and well-suited for the real-time detection requirements in this scenario.
[0074] During real-time operation, the system deploys a corresponding anomaly detection model for each critical device and continuously calculates its anomaly score. Based on historical data statistical analysis or expert experience, the system sets a dynamic or fixed anomaly score threshold. When a device's anomaly score consistently and significantly exceeds this threshold—for example, if the anomaly score is above the 95th percentile of the threshold for five consecutive sampling periods—the system determines that the device has experienced a "device-level single point of failure" and pushes the anomaly event (including device ID, anomaly time, anomaly score, and key characteristics triggering the anomaly) to the next step for in-depth analysis.
[0075] Step S204 triggers and initiates knowledge graph-based relational reasoning. This is the most innovative core step that distinguishes the entire method from all existing technologies, corresponding to the work of the fault reasoning and early warning engine 404. Unlike traditional systems that immediately push simple alarms to users upon detecting an anomaly, the method of this invention introduces an intermediate layer of "intelligent diagnosis." Once a single point of failure is determined in step S203, the system does not immediately alert the user, but automatically initiates a deep, relational reasoning task based on the knowledge graph 402.
[0076] The query entry point for this task is the anomalous event identified in the previous step, namely the device entity that experienced the anomaly (e.g., entity Pump_ID:P001) and its related anomalous features (e.g., the attribute vibration kurtosis_Score is 0.92). The fault reasoning and early warning engine 404 first starts from the node Pump_ID:P001 in the knowledge graph and uses graph database query languages (such as Cypher for Neo4j) or graph traversal algorithms (such as Breadth-First Search (BFS) or Depth-First Search (DFS)) to perform multi-hop retrieval of associated entities along various predefined "relationship" edges in the knowledge graph. For example, it retrieves all entities with a physical connection (connectsTo) to Pump_ID:P001, such as the upstream pipe Pipe_ID:Seg122 and the downstream pipe Pipe_ID:Seg123; it retrieves entities with logical control relationships, such as the programmable logic controller PLC:C01 and the frequency converter VFD_ID:V01 that control it; and it retrieves entities with functional dependencies (belongsTo), such as the heat exchange station it resides in:S05. This retrieval can be extended to second-order or even higher-order neighbor entities, thereby constructing a local system association subgraph centered on anomalies.
[0077] Furthermore, in order to ensure computational efficiency and reasoning focus, the multi-hop associated entity retrieval process sets explicit boundary conditions and termination conditions.
[0078] A preferred boundary condition is to set a maximum number of hops. For example, the system can preset the maximum number of hops to 3. This means that the inference engine, starting from an anomalous node, will at most retrieve related entities within its third order of influence. This hop count setting is based on a trade-off of domain knowledge; typically, most direct and indirect causal chains of failures in a hot water system fall within this range. A larger hop count provides limited information gain, but the computational cost increases exponentially.
[0079] Another optional or supplementary boundary condition is semantic constraints based on entity type. The system can pre-define a list of "terminating entity types," and the extension of the search path will terminate when it encounters entities of these types. For example, it can be set to stop searching upstream when entities such as "main heating boiler," "city main water supply network interface," or "substation main transformer," which are located at the system boundary, are retrieved.
[0080] The termination condition, besides reaching the maximum number of hops, can also be dynamic. For example, an upper limit can be set on the number of associated entities. When the total number of entities in the constructed associated subgraph reaches a preset value (e.g., 50), the retrieval stops to prevent the retrieval scope from being excessively expanded due to a "super node" (a node connected to a large number of other devices). Furthermore, the retrieval path can also be terminated early when all newly retrieved entities no longer have available cross-validation sites with the multimodal sensors described in this invention. Through the combined application of these boundary and termination conditions, it is ensured that the reasoning process of the knowledge graph always proceeds within a controllable range highly relevant to the problem. After acquiring this correlation subgraph, the engine automatically generates a series of possible fault causal chain hypotheses that conform to engineering logic, based on the subgraph's structure and the expert rule base preset in the graph. These rules can be stored in the form of "IF-THEN," for example: Rule 1, IF (Pipe.downstream.Pressure INCREASES) AND (Pump.Vibration INCREASES) THEN (CAUSE MAY BE Pipe.downstream.Blockage), indicates that downstream pipe blockage may cause pump vibration and pressure to increase simultaneously; Rule 2, IF (Sensor.Value ABNORMAL) AND (ALL other modalities of the same equipment NORMAL) THEN (CAUSE MAY BE Sensor.Fault), indicates that if only one sensor reading is abnormal, while other physical dimension data of the same equipment are normal, it is likely that the sensor itself is faulty. By matching the features of the current abnormal event and the correlation subgraph with the preconditions in the rule base, the system can generate a series of candidate fault hypotheses to be verified.
[0081] Please see Figure 4 This example illustrates the structure of a specific knowledge graph for a hot water supply system. In this example, 'Pump P001', 'Valve V089', and 'Pipe Seg123' are physical device entities; 'Controller PLC01' is a logical entity; 'Heat Exchange Station S05' is a location entity; and 'Abnormal Event E01' is an event entity dynamically created during inference. The relationships between them clearly demonstrate the system's inherent connections: for example, the `connectsTo` relationship indicates that 'Valve V089' is upstream of 'Pipe Seg123', and 'Pump P001' is downstream; the `belongsTo` relationship specifies that these devices all belong to 'Heat Exchange Station S05'; and the `controls` relationship defines the logical control of 'Controller PLC01' over 'Pump P001'. When an abnormal event occurs, as shown in the figure, the `occurredOn` relationship associates the event with a specific device. The inference engine can then traverse along these relational paths. For example, starting from 'abnormal event E01', it can find 'pump P001' through occurredOn, and then find 'pipe Seg123' through connectsTo, thus constructing a complete analysis chain.
[0082] Step S205 involves performing multi-source data cross-validation and fault root cause localization. This step is the evidence collection and decision-making stage of the inference process, aiming to find the hypothesis that best matches the current multi-source data evidence among numerous candidate hypotheses. For each hypothesis generated in the previous step, the inference engine will immediately take action, proactively and in real-time retrieving the multimodal data features of all relevant entities (devices) in the hypothesis chain within a time window before and after the time of the anomaly occurrence from the time-series database.
[0083] Then, the engine begins comparing and scoring the evidence. For example, to verify the hypothesis that "the blockage in downstream pipe Seg123 caused pump P001 to overload," the engine checks for the existence of expected evidence. Expected evidence is that all pressure sensor readings on pipe Seg123 should show a synchronous and significant upward trend, and the acoustic sensor data for that pipe segment may show specific vortex noise characteristics indicating blockage. The engine compares the real-time retrieved data with this expected pattern. If the data shows that the pressure and acoustic data for pipe Seg123 are normal, it indicates that key evidence supporting the hypothesis is missing, and the confidence score for the hypothesis will be assigned a very low value, or it will be directly "pruned" and excluded. Conversely, if the data shows that other characteristics of Pump_ID:P001 itself (such as the motor temperature monitored by the infrared camera) also show a synchronous abnormal increase, and the status data of all external related entities (such as the power supply system, upstream and downstream pipelines, and control units) show normal, then the confidence score for the hypothesis of "pump malfunction" will be assigned a very high value.
[0084] This process is essentially an application of Bayesian reasoning or evidence theory, updating the level of confidence in each hypothesis by continuously introducing new, independent evidence. Ultimately, after a systematic, data-driven evaluation of all candidate hypotheses, the engine arrives at one or a few fault diagnosis conclusions with the highest confidence scores.
[0085] A preferred implementation is a scoring model based on evidence matching. In this model, the system pre-defines a set of strongly relevant positive evidences and a set of strongly relevant negative evidences for each candidate fault hypothesis. For example, for the hypothesis of 'downstream pipe blockage', the positive evidence might include: {Evidence 1: Significantly increased downstream pressure; Evidence 2: Significantly decreased downstream flow; Evidence 3: Eddy noise characteristics in the acoustic signal}. During cross-validation, the inference engine checks whether these pieces of evidence appear in the real-time data one by one and assigns a matching score based on their significance (e.g., 1 point for a perfect match, 0.5 points for a partial match, and 0 points for no match). Finally, the confidence score of the hypothesis can be calculated using a weighted formula, for example: CS = (∑(w_i * S_pi) - ∑(w_j * S_nj)) / ∑w_i Where w_i is the weight of the i-th expected piece of evidence, and S_pi is its matching score; w_j is the weight of the j-th expected piece of disproving evidence, and S_nj is its matching score. The weight w can be preset by domain experts based on experience; for example, the weight of the evidence of increased pressure can be set higher than that of acoustic noise.
[0086] Another alternative implementation is to employ a simplified Bayesian inference network. Each candidate fault hypothesis can be considered as a root node event (H) to be inferred, while the multimodal data features retrieved from related entities serve as observed evidence nodes (E1, E2, ... En). The system needs to pre-estimate the prior probability P(H) of each hypothesis, and the conditional probability P(Ei|H) of observing specific evidence given that a hypothesis is true, using expert knowledge or historical data statistics. When new evidence is observed in real time, the system can iteratively update the posterior probability of the hypothesis using Bayes' theorem and use it as a confidence score. P(H|E1, E2) = α * P(E2|H, E1) * P(E1|H) * P(H) Here, α is the normalization factor. In this way, with each new cross-validation piece of evidence introduced, the hypothesis's confidence level is dynamically and logically updated, ultimately resulting in a quantitative confidence assessment with clear probabilistic significance. Please refer to the following for details. Figure 5 This paper details the fault tracing and reasoning process based on a knowledge graph. The process begins with 'receiving a single point of failure event,' corresponding to the output of step S203. Subsequently, the process moves to 'locating the anomalous entity in the knowledge graph,' i.e., finding the object where the event occurred. The core reasoning process occurs in two steps: 'multi-hop retrieval of related entities to construct a related subgraph' and 'matching with the expert rule base to generate candidate fault hypotheses,' which is the core implementation of step S204. Next, 'actively retrieving multimodal data for cross-validation for each hypothesis' is the specific manifestation of step S205, verifying the truth or falsity of the hypothesis through data evidence. After 'calculating the confidence level of each hypothesis,' the system uses a decision node 'Is a unique high-confidence root cause determined?' to make a judgment. If 'yes,' it directly 'outputs the diagnostic conclusion'; if 'no,' it 'outputs multiple possible causes and their confidence ranking,' providing maintenance personnel with a more comprehensive decision-making reference. Finally, the process ends. This flowchart fully demonstrates the intelligent logic of how the method of this invention transforms anomalous information into accurate diagnostic conclusions.
[0087] Step S206: Generate a visual early warning report and operation and maintenance suggestions. After the inference engine has completed the evaluation of all hypotheses and taken the hypothesis (or set of hypotheses) with the highest confidence score as the final fault diagnosis conclusion, the system enters the information output and decision support stage.
[0088] The system automatically generates a structured, graphically-rich early warning report. Unlike traditional alarms that only indicate "XX equipment XX parameter exceeds limits," this report is detailed and logically clear. The report clearly points out: root cause location, identifying which device or component is the root cause of the problem; fault type inference, inferring the specific fault mode (such as "bearing inner ring wear," "motor winding insulation degradation," or "valve internal leakage") based on abnormal characteristic combinations (e.g., abnormal specific frequency components in the vibration spectrum or overall temperature increase); evidence chain presentation, listing all data evidence supporting the diagnosis, along with excluded hypotheses and their reasons, using charts (such as vibration spectrum comparison before and after the anomaly) combined with textual explanations, enabling maintenance personnel to fully understand and accept the diagnosis; and impact domain analysis, automatically analyzing the potential water supply areas and user ranges affected if the fault continues to develop by tracing downstream paths on a knowledge graph, providing decision support for emergency plan development and resource allocation.
[0089] Finally, this comprehensive early warning report will be accurately sent to the designated maintenance personnel responsible for the region or type of equipment via various methods, including mobile app push notifications, SMS, email, or internal enterprise instant messaging tools, through the application service layer 500. Furthermore, for high-confidence early warnings requiring immediate action, the system can seamlessly integrate with the enterprise's maintenance management system (such as EAM / CMMS) to automatically create and dispatch electronic repair work orders containing all the detailed diagnostic information mentioned above, thereby automating the entire process from intelligent fault discovery to closed-loop maintenance tasks.
[0090] Through the continuous and cyclical execution of the above steps S201 to S206, the method proposed in this invention can realize closed-loop intelligent management of the entire hot water system, from fine-grained multi-dimensional state perception to high-level systemic and correlated fault tracing, thereby elevating the operation and maintenance mode to a whole new level.
[0091] Example 2 This embodiment describes the specific structure of a hot water system fault diagnosis and early warning system based on multimodal perception fusion and knowledge graph reasoning. Please refer to... Figure 1This system is the physical or logical carrier of the method described in Embodiment 1. It adopts an overall "end-edge-cloud" collaborative architecture, comprising, from bottom to top, a multimodal perception module, a data processing and fusion module, an anomaly detection module, a fault inference engine module, and an application service module. These modules logically belong to the perception layer 100, the cloud platform layer 400, and the application service layer 500.
[0092] Multimodal sensing module, corresponding to Figure 1 The sensing layer 100 is the data source for the entire system. Its core function is to comprehensively, accurately, and multi-dimensionally collect data on the status of various key nodes and equipment in the hot water supply system. This module consists of a large number of distributed multimodal sensing terminals 300. Each sensing terminal is physically a hardware unit integrating multiple sensors. Please refer to [link to relevant documentation]. Figure 3 It may include a microcontroller unit (MCU), multiple sensor interfaces, a power management unit, and a wireless or wired communication module. In this embodiment, a typical multimodal sensing terminal 300 integrates different combinations of sensors depending on its deployment location and the object being monitored.
[0093] Specifically, for sensing terminals installed on critical rotating equipment such as primary or secondary circulation pumps and makeup water pumps, the core sensor is at least one triaxial accelerometer 101 coupled and installed near the bearing housing of the motor or pump body to collect vibration data. Simultaneously, to more comprehensively assess the equipment status, the terminal can also integrate a temperature sensor to monitor bearing temperature and a non-contact current sensor to monitor motor operating current. For sensing terminals deployed on critical pipe sections, the core is at least one contact-type broadband acoustic sensor 102 (or acoustic emission sensor) installed on the outer wall of the pipe section to collect fluid acoustic data and monitor phenomena such as leakage and cavitation. This terminal can also integrate pressure and temperature sensors. For sensing terminals deployed in concentrated equipment areas such as heat exchange stations or boiler rooms, it may be a node centered on an industrial computer (IPC), connected to all conventional process parameter sensors (temperature, pressure, flow rate, water meters, etc.) in the area, and controlling at least one infrared thermal imaging camera 103 to perform timed or triggered thermal condition scans of critical equipment.
[0094] The data processing and fusion module, logically located at the cloud platform layer 400, processes the heterogeneous data collected by the multimodal sensing module and fuses it into a feature vector that uniformly represents the health status of the equipment. In its implementation, this module can consist of a series of microservices or software components. First, a data access service (such as one based on Apache Kafka or MQTT Broker) is responsible for receiving and buffering data streams from the sensing layer. Subsequently, a data preprocessing service is responsible for performing data cleaning, alignment, and standardization. At its core is a feature engineering service containing an algorithm library with feature extraction algorithms for different data types. For example, it includes a signal processing component based on Fast Fourier Transform (FFT) for extracting frequency domain features from vibration data; a component based on wavelet transform for extracting transient impact features from acoustic data; and a component based on a computer vision library (such as OpenCV) for performing image segmentation and temperature field analysis from infrared thermal imaging images to extract thermal features. Finally, a fusion component combines the standardized features extracted from various data sources into a high-dimensional multimodal state feature vector according to preset rules (e.g., simply concatenating or weighted fusion), and publishes it to the subsequent processing flow.
[0095] More specifically, the feature layer fusion method, which integrates multiple features into a multimodal state feature vector, can be implemented by the following steps. First, the features extracted from different modal data are normalized to eliminate the influence of differences in units and numerical ranges between different features. Preferably, the Z-score normalization method can be used to convert each feature into a standard normal distribution with a mean of 0 and a standard deviation of 1. Second, feature selection and dimensionality reduction are performed. Considering that the fused vector may have excessively high dimensionality and contain redundant or irrelevant features, algorithms such as principal component analysis (PCA) or mutual information-based feature selection (MIFS) can be used to reduce the dimensionality or filter all normalized features, selecting the subset of features that contributes most to the health status assessment.
[0096] Finally, feature fusion is performed. A simple fusion method is to directly concatenate the selected feature vectors to form a higher-dimensional unified feature vector. A more preferred method is weighted fusion. In this method, features (or feature groups) of different modalities are assigned different weights w_k, which reflect the importance of different modal data in specific operating conditions or specific fault diagnosis tasks. The final multimodal feature vector V_fused can be expressed as: V_fused = [w_1*V_1, w_2*V_2, ..., w_k*V_k, ...] Where V_k is the normalized feature vector from the k-th modality. The weight w_k can be determined statically, i.e., pre-set by domain experts based on experience; or dynamically, for example, by training an attention mechanism network so that the model automatically learns and assigns weights most suitable for the current state based on real-time input data. For example, in a suspected leakage scenario, the model might automatically assign higher weights to acoustic features.
[0097] The anomaly detection module, logically located at the cloud platform layer 400, is configured to use a pre-defined unsupervised anomaly detection model to perform real-time analysis and scoring of feature vectors generated by the data processing and fusion module. In its implementation, this module can be an independent machine learning model service. This service maintains one or more pre-trained anomaly detection models for each type of critical equipment in the system (e.g., a specific model of water pump, a specific type of heat exchanger). These models, such as the previously mentioned autoencoders or isolated forests, are trained offline on a large amount of historical normal operating data of the equipment and stored in a model library in file format (e.g., HDF5 or PMML format). During real-time monitoring, this module subscribes to the feature vector stream, loads the corresponding model based on the device ID, performs inference calculations, and derives an anomaly score. When the anomaly score exceeds a dynamically adjusted or statically configured threshold, the module generates an anomaly event and encapsulates this event along with all relevant contextual information (such as device ID, timestamp, original feature vector, anomaly score, etc.) into a message, which is then sent to the fault inference engine module.
[0098] The fault reasoning engine module essentially involves deep integration with a pre-built knowledge graph of the hot water system. In its implementation, this module can consist of the following parts: a graph database, such as Neo4j or JanusGraph, for storing and managing the hot water system knowledge graph. This knowledge graph is a graph-structured data model containing entities, relationships, and attributes. Entities represent various physical objects (such as pumps, valves, and pipe sections) and logical objects (such as controllers and control algorithms) in the hot water system; relationships are represented by directed edges, such as physical connections (connectsTo) and contains (contains), logical controls (controls), and functional effects (affects). A rule engine, such as Drools, is used to store and execute diagnostic knowledge rules from domain experts. Finally, a reasoning coordination service subscribes to anomalous events published by the anomaly detection module. Once an event is received, it initiates the reasoning process: First, it sends a query to the graph database to retrieve the local subgraphs related to the anomalous entity; then, it submits the subgraphs and anomalous information to the rule engine to generate candidate failure hypotheses; next, it calls a time-series data query service based on the hypotheses to obtain the multimodal data required for cross-validation; finally, it executes a confidence evaluation algorithm to score each hypothesis and outputs the final diagnostic conclusion.
[0099] The application service module serves as the system's human-computer interaction interface and external service window. In its implementation, this module can be a web application server and a series of API interfaces. It provides maintenance personnel with a visual monitoring center. The front-end interface of this center can be based on GIS (Geographic Information System) and BIM (Building Information Modeling) technologies to intuitively display the topology of the hot water supply network and the real-time status of equipment in two-dimensional or three-dimensional form. Upon receiving diagnostic conclusions from the fault inference engine module, it highlights alarms on the interface and displays detailed warning reports through pop-ups and other means. Simultaneously, this module provides a push notification service, proactively delivering warning information to relevant personnel by integrating with SMS gateways, email servers, or mobile application push platforms. Furthermore, it provides standard RESTful API interfaces, allowing integration with existing enterprise maintenance management systems (CMMS), asset management systems (EAM), or work order systems to achieve automatic work order creation, dispatch, and status tracking, forming a closed-loop management system.
[0100] Example 3 This embodiment will use a specific application scenario—the fault diagnosis and early warning of early bearing wear in the secondary circulating water pump of a heat exchange station in a large residential community—to explain in detail the actual application process, parameter settings, and technical effects of the technical solution described in this invention.
[0101] The application scenario is described as follows: The monitored object is a secondary circulating water pump P001 in heat exchange station S05. Its equipment model is ABC-150, it uses rolling bearings, and its rated speed is 2950 rpm, corresponding to a rotational frequency f_r of approximately 49.17 Hz. This pump has been running continuously for over 4000 hours, and according to the equipment maintenance manual, it has entered a potentially high-risk period for bearing wear. The fault scenario is the appearance of tiny, invisible fatigue pitting on the raceway of the pump's drive-end bearing, which is the initial stage of bearing wear.
[0102] Hardware Selection and Deployment: In this scenario, the deployment scheme for the multimodal sensing terminal is as follows: A single-axis IEPE piezoelectric accelerometer is installed vertically and horizontally on the drive-end bearing housing of water pump P001 to collect vibration data. The sensor range is ±50g, and the frequency response range is 0.5 Hz to 15 kHz. A contact acoustic emission sensor with a frequency response range of 20 kHz to 100 kHz is installed on the outer wall of the secondary side main water supply pipeline at the pump outlet. Inside the heat exchange station, facing the equipment group including P001, an industrial-grade infrared thermal imaging camera with a resolution of 384x288 pixels is installed. Simultaneously, the system also connects to the existing PT100 temperature sensor and piezoresistive pressure sensor deployed at the pump inlet and outlet, as well as the electromagnetic flow meter on the pipeline.
[0103] Parameter settings and data acquisition: The sampling frequency of the vibration sensor is set to 40.96 kHz. The sampling frequency of the acoustic sensor is set to 250 kHz. The infrared camera is set to capture a thermal image every 15 minutes. Conventional temperature, pressure, and flow parameters are acquired once per minute.
[0104] In the early stages of a failure, such as the first few days after bearing pitting, it has almost no impact on the pump's macroscopic performance. Data collected by temperature and pressure sensors deployed at the pump's inlet and outlet, as well as data from flow sensors on the pipeline, all fall within normal seasonal fluctuations, and their intraday fluctuation curves show no visible anomalies compared to historical data. At this time, a traditional SCADA system that relies solely on these process parameters will not generate any alarms.
[0105] However, the technical solution of this invention continuously collects high-frequency vibration waveform data through a vibration sensor installed on the bearing housing. After performing FFT transformation on these data, the data processing and fusion module found that although the energy amplitude of the spectrum at the low-frequency turnaround (49.17 Hz) and second harmonic (98.34 Hz) is still within the normal range, previously unseen impact response and harmonic components with weak energy but clear signal-to-noise ratio have begun to appear on the background noise baseline in the mid-to-high frequency range (e.g., 3000 Hz to 8000 Hz). By comparing with the theoretical fault characteristic frequencies (BPFO, BPFI, etc.) of this type of bearing, the system found that these newly appearing frequency components highly match the fault characteristic frequencies of the bearing outer ring. At the same time, the system calculated that the time-domain kurtosis index of the vibration signal is 3.8, which slightly exceeds the reference value of 3 during normal stable operation. These subtle changes are captured and quantified by the data processing and fusion module and input as key features into the multimodal state feature vector.
[0106] The feature vector was fed into the pre-trained autoencoder anomaly detection model for pump P001. Because the pattern of this vector differed significantly from the "normal pattern" learned by the model from a large amount of health data (mainly contributed by the aforementioned kurtosis index and high-frequency spectral energy features), the model generated a reconstruction error two standard deviations higher than normal when reconstructing the vector. This error was converted into an anomaly score as high as 0.88, exceeding the preset alarm threshold of 0.8 for the first time. Based on this, the system determined that a single point of failure had occurred in device Pump_ID:P001.
[0107] At this point, the fault reasoning and early warning engine was triggered and began executing knowledge graph-based reasoning, as described above. After ruling out hypotheses such as "downstream pipe blockage" (because the pressure and flow data were normal) and "sensor malfunction" (because the infrared thermogram showed a slight temperature increase of 0.5°C near the bearing housing, synchronized with abnormal vibration, providing circumstantial evidence), the engine finally identified the root cause of the fault with a high confidence level of 98% as "mechanical failure of Pump_ID:P001 itself, with the fault mode highly suspected to be early bearing wear." The system then generated a detailed early warning report and pushed it to the maintenance engineer responsible for the heat exchange station.
[0108] Please see Figure 6 The figure illustrates the application scenario of this embodiment. Figure 6The diagram clearly depicts the circulating water pump P001 within heat exchange station S05 and the various sensors deployed around it, as well as the data flow from data transmission via network to the cloud platform, ultimately presenting early warning information on the monitoring terminals of maintenance personnel. In the left-hand block diagram of 'Heat Exchange Station S05 Site', the core monitoring object 'circulating water pump P001' is surrounded by a 'multimodal sensor group (101, 102, 103, etc.)' deployed around it. After these sensors collect data, the data is aggregated and preliminarily processed locally via an edge gateway. Subsequently, the data is uploaded to the cloud analysis platform 400 on the right-hand side via high-speed networks such as 5G / fiber optics. In the cloud, the fault reasoning and early warning engine 404 is the core processing unit, which performs fault diagnosis by querying / reasoning on the knowledge graph database 402. Finally, the accurate early warning information derived from the analysis is pushed to the monitoring terminals 500 of maintenance personnel.
[0109] Please see Figure 7 This figure illustrates a comparison of the fault early warning capabilities between the method of this invention and traditional monitoring methods based on process parameter thresholds. The diagram uses time as the horizontal axis and the quantitative indicators of the health status of key equipment as the vertical axis, showcasing the entire evolution process from fault initiation to final failure. In the figure, curve C1 represents the anomaly score of the multimodal state feature vector calculated by the method of this invention, and curve C2 represents a single process parameter (e.g., bearing temperature) monitored by the traditional method. Time T0 represents the moment when initial physical damage (e.g., bearing pitting) occurs inside the equipment. As shown, within a very short time after time T0, due to the high sensitivity of the high-frequency dynamic signals such as vibration and acoustics used in the method of this invention to weak impact-type fault characteristics, curve C1 responds rapidly and continues to rise, significantly exceeding the preset anomaly judgment threshold TH1 at time T1, thus triggering a precise early warning. In contrast, curve C2, representing traditional monitoring methods, shows almost no change for a considerable period from T0 to T2 (the fault latency period) due to the extremely slow response of macroscopic process parameters such as temperature to early damage. Only at time T2, when the damage has accumulated to a relatively severe level, does the temperature rise sharply due to frictional heat, exceeding the alarm threshold TH2. Time T3 indicates that the equipment has experienced functional failure. This figure clearly shows that the fault detection lead time provided by the method of this invention is ΔT = T2 - T1. In this embodiment, this lead time can reach several days to several weeks, providing a valuable window for implementing predictive maintenance.
[0110] In summary, this invention first establishes a deep sensing system capable of precisely depicting the microscopic mechanical and fluid dynamic states of equipment by integrating multimodal data such as vibration, acoustics, and infrared thermal imaging, fundamentally improving the ability to detect the nascent stages of "soft faults." Building upon this foundation, the invention goes beyond the analysis of single-point data; instead, it constructs a digital twin knowledge graph depicting the physical topology, logical control, and functional relationships of the entire hot water supply network. This integrates previously discrete equipment and data into a computationally and logically-driven intelligent network, completely breaking down "data silos." Most importantly, when the sensing layer detects an anomaly, the invention does not issue an isolated threshold alarm but automatically activates a multi-hop associative reasoning engine based on the knowledge graph. This engine simulates the diagnostic thinking of domain experts, generating fault hypotheses, cross-validating multi-source evidence, and accurately tracing the root cause along the graph path. Ultimately, this achieves a shift from the traditional passive, appearance-based alarm mode to a proactive, root-cause-based predictive maintenance paradigm, significantly improving the early warning and accuracy of fault diagnosis.
[0111] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer programs. When the computer program is loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer program can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program can be transferred from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).
[0112] Those skilled in the art will understand that the various numerical designations such as "first," "second," etc., used in this disclosure are merely for the convenience of description and are not intended to limit the scope of the embodiments of this disclosure, nor do they indicate the order of events.
[0113] At least one of the features described in this disclosure can also be described as one or more, and multiple features can be two, three, four or more, and this disclosure does not impose any limitations. In the embodiments of this disclosure, for a technical feature, the technical features in that technical feature are distinguished by "first", "second", "third", "A", "B", "C" and "D", etc., and there is no sequential order or size order among the technical features described by "first", "second", "third", "A", "B", "C" and "D".
[0114] The correspondences shown in the tables of this disclosure can be configured or predefined. The values of the information in each table are merely examples and can be configured to other values; this disclosure is not limiting. When configuring the correspondences between information and parameters, it is not necessarily required to configure all the correspondences shown in each table. For example, the correspondences shown in some rows of the tables in this disclosure may not be configured. Furthermore, appropriate modifications and adjustments can be made based on the above tables, such as splitting, merging, etc. The names of the parameters shown in the headers of the above tables can also use other names that the communication device can understand, and the values or representations of the parameters can also be other values or representations that the communication device can understand. In the implementation of the above tables, other data structures can also be used, such as arrays, queues, containers, stacks, linear lists, pointers, linked lists, trees, graphs, structures, classes, heaps, hash tables, or hash tables, etc.
[0115] The predefined terms in this disclosure can be understood as defined, pre-defined, stored, pre-stored, pre-negotiated, pre-configured, solidified, or pre-burned. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0116] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. The above descriptions are merely specific embodiments of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A hot water system fault diagnosis and early warning method based on multi-modal perception fusion and knowledge graph reasoning, characterized in that, Includes the following steps: S1. Based on multiple sensors deployed at key nodes of the hot water system, collect multimodal operating data of the hot water system. The multimodal operating data includes at least: vibration data characterizing the mechanical state of the equipment, acoustic data characterizing the fluid state, infrared thermal imaging data characterizing the thermal condition of the equipment, and process parameter data characterizing the operating state of the system. S2. The multimodal operating data is preprocessed and feature-engineered to extract multiple features that characterize the health status of key nodes, and fused into a multimodal state feature vector; wherein, the feature engineering includes: performing a fast Fourier transform on the vibration data to extract frequency domain features; performing a wavelet transform on the acoustic data to extract transient impact features; and performing image segmentation and temperature field analysis on the infrared thermal imaging data to extract thermal features; S3. Using a preset unsupervised anomaly detection model, the multimodal state feature vector is analyzed in real time to calculate anomaly scores. When the anomaly scores exceed a preset threshold, the key node where the anomaly occurs is determined. The preset unsupervised anomaly detection model is an autoencoder model, specifically a deep stacked autoencoder. S4. In response to identifying the key node where the anomaly occurred, a multi-hop related entity retrieval is performed in the pre-constructed hot water system knowledge graph, starting from the key node of the anomaly; a candidate fault hypothesis is generated for each retrieved related entity, and the multimodal operating data of the related entity is actively invoked for cross-validation; and the confidence level is calculated for each candidate fault hypothesis based on the validation results. The hot water system knowledge graph includes entities, relationships, and attributes; the entities represent physical and logical objects in the hot water system; the relationships represent the physical topological connections, logical control relationships, and functional dependencies between the entities; and the multi-hop associated entity retrieval involves traversing the graph along the relationships. The cross-validation includes: for the candidate fault hypothesis of downstream pipeline blockage, verifying whether the pressure data of the downstream pipeline has a synchronous upward trend and whether the acoustic data has eddy current noise characteristics; for the candidate fault hypothesis of sensor itself, verifying whether the modal data of other physical dimensions of the abnormal key node have synchronous changes consistent with the anomaly. S5. Based on the candidate fault hypothesis with the highest confidence level, determine the root cause and type of the fault, and generate structured early warning information that includes fault root cause location, diagnostic evidence chain and influence domain analysis.
2. The method of claim 1, wherein, In step S1, the vibration data is collected by an acceleration sensor deployed on the housing of a rotating or reciprocating device, which includes a water pump or valve; the acoustic data is collected by a contact acoustic sensor deployed on the outer wall of a key pipe section; and the infrared thermal imaging data is collected by an infrared camera deployed in the concentrated equipment area of a heat exchange station or boiler room.
3. The method of claim 1, wherein, In step S4, cross-validation is performed by comparing the real-time data of the retrieved associated entities with the expected data patterns of the candidate fault hypotheses, thereby eliminating hypotheses that do not match the actual data and determining the candidate fault hypotheses with the highest confidence.
4. A hot water system fault diagnosis and early warning system based on multi-modal perception fusion and knowledge graph reasoning, characterized in that, include: The multimodal sensing module is used to collect multimodal operating data of key nodes in the hot water system. The multimodal operating data includes at least: vibration data characterizing the mechanical state of the equipment, acoustic data characterizing the fluid state, infrared thermal imaging data characterizing the thermal condition of the equipment, and process parameter data characterizing the operating state of the system. The data processing and fusion module is used to preprocess and feature-engineer the multimodal operating data, extract multiple features that can characterize the health status of key nodes, and fuse them into a multimodal state feature vector. The data processing and fusion module is further configured to execute fast Fourier transform, wavelet transform, and image segmentation algorithms to extract frequency domain features, transient impact features, and thermal features from the vibration data, acoustic data, and infrared thermal imaging data, respectively. An anomaly detection module is configured to use a preset unsupervised anomaly detection model to perform real-time analysis on the multimodal state feature vector, calculate an anomaly score, and determine the key node where an anomaly occurs when the anomaly score exceeds a preset threshold; the preset unsupervised anomaly detection model is an autoencoder model, specifically a deep stacked autoencoder. A hot water system knowledge graph storage module is used to store a hot water system knowledge graph containing entities, relationships, and attributes; the entities represent physical and logical objects in the hot water system, including water pumps, valves, pipe sections, and controllers; the relationships represent the physical topological connections, logical control relationships, and functional dependencies between the entities; The fault reasoning engine module, connected to the hot water system knowledge graph storage module, is configured to, in response to identifying a key node where an anomaly has occurred, perform a multi-hop related entity retrieval in the hot water system knowledge graph, starting from the key node where the anomaly occurred and following the relationship; generate candidate fault hypotheses for each retrieved related entity, and actively call the multimodal operational data of the related entities for cross-validation; and calculate the confidence level for each candidate fault hypothesis based on the validation results. The application service module is configured to determine the root cause and type of the fault based on the candidate fault hypothesis with the highest confidence level, and generate structured early warning information that includes fault root cause location, diagnostic evidence chain and impact domain analysis.
5. The system of claim 4, wherein, The multimodal sensing module includes: at least one triaxial accelerometer coupled to the housing of a water pump or valve; at least one contact broadband acoustic sensor mounted on the outer wall of a key pipe section; and at least one infrared thermal imaging camera deployed in the area where equipment is concentrated.
6. The system of claim 4, wherein, The fault reasoning engine module is further configured to: perform cross-validation by comparing the real-time data of the retrieved related entities with the expected data patterns of the candidate fault hypotheses, thereby eliminating hypotheses that do not match the actual data and determining the candidate fault hypotheses with the highest confidence.
Citation Information
Patent Citations
Intelligent heat supply system and control method
CN118331184A
Intelligent equipment fault diagnosis and reasoning method and system based on unsupervised learning
CN119807959A