A wall-mounted stove fault diagnosis method and system based on multi-source data analysis
By constructing the perturbation characteristic sequence and electrical network topology of the wall-hung boiler, and combining it with environmental humidity data, accurate diagnosis of wall-hung boiler faults was achieved. This solved the problems of insufficient data processing capabilities and poor environmental adaptability in existing technologies, and improved the accuracy and efficiency of diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AIWO (SHENZHEN) INTELLIGENT ENVIRONMENTAL TECH CO LTD
- Filing Date
- 2025-12-09
- Publication Date
- 2026-07-28
AI Technical Summary
Existing fault diagnosis technologies for wall-hung boilers suffer from insufficient data processing capabilities, poor environmental adaptability, and strong reliance on subjective judgment. They are unable to capture early fault signals, are prone to misjudging environmentally induced faults as hardware defects, and have low diagnostic efficiency and poor consistency.
By acquiring data on gas flow fluctuations and water pump speed changes, a perturbation feature sequence is constructed. Combined with environmental humidity data, an electrical network topology is built, the propagation path of the perturbation features is analyzed, and a standardized diagnostic report is generated.
It enables accurate early warning of faults, accurately distinguishes faults caused by environmental factors, reduces reliance on human experience, and ensures the consistency and efficiency of diagnostic logic in different scenarios.
Smart Images

Figure CN121580136B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wall-hung boiler fault diagnosis technology, and in particular to a wall-hung boiler fault diagnosis method and system based on multi-source data analysis. Background Technology
[0002] Currently, wall-hung boilers are the core equipment for heating and hot water supply in modern homes, and their stable operation is crucial for ensuring quality of life. With the popularization of smart homes, the demand for fault diagnosis of wall-hung boilers is increasing, especially in complex usage scenarios. With the help of cloud-edge collaboration technology, edge devices can quickly collect and analyze operating data, while linking with the massive fault knowledge base and intelligent algorithms in the cloud, enabling the equipment to locate faults more quickly and accurately and provide solutions.
[0003] In one existing technology, the operating status of a wall-hung boiler is monitored by comparing preset fixed rules and parameter thresholds. The system collects sensor data in real time and compares the data with preset conditions in the rule base. If the conditions are met, the corresponding fault code is output. If the conditions are not met, the system determines that the operation is normal and does not generate any prompt information, and continues to execute the next round of data collection and comparison.
[0004] However, existing technologies only focus on instantaneous data values, ignoring subtle fluctuations in parameters such as gas flow rate and water pump speed over time (i.e., time-series perturbations), thus failing to capture early fault signals. Furthermore, their fixed rules make it difficult to adapt to complex scene changes (such as fluctuations in ambient humidity), easily misdiagnosing environment-induced faults as hardware defects. Moreover, the diagnostic process relies on human experience and lacks standardized procedures, resulting in low efficiency and poor consistency. In summary, existing technologies suffer from insufficient data processing capabilities, poor environmental adaptability, and strong reliance on subjective judgment. Summary of the Invention
[0005] This invention provides a method and system for diagnosing wall-hung boiler faults based on multi-source data analysis, in order to solve the problems of insufficient data processing capabilities, poor environmental adaptability and strong subjective dependence in existing technologies.
[0006] In a first aspect, to solve the above-mentioned technical problems, the present invention provides a method for diagnosing wall-hung boiler faults based on multi-source data analysis, comprising: The raw data, exhaust temperature curve, and ambient humidity data are acquired, and the gas flow fluctuation and water pump speed change data in the raw data are processed to obtain a perturbation feature sequence. Features are extracted from the perturbation feature sequence and combined with the environmental humidity data to construct a relationship graph, thereby obtaining the electrical network topology. The fluctuation value of the perturbation feature sequence is calculated. When the fluctuation value exceeds a preset fluctuation threshold, the connection weights and abnormal nodes are extracted from the electrical network topology to obtain a set of abnormal nodes. Traverse each node in the set of abnormal nodes to find the fault origin and obtain the initial fault origin. Based on the initial fault initiation point, the peak shift of the exhaust temperature curve is analyzed to determine the propagation influence range, and the nodes within the propagation influence range are extracted to obtain the root cause identification node; If the root cause identification node contains multiple nodes, then the optimized fault link is obtained by optimizing the path according to the electrical network topology. Based on the optimized fault link, a diagnostic report template is generated, and the final diagnostic report output is obtained.
[0007] Secondly, the present invention provides a wall-hung boiler fault diagnosis system based on multi-source data analysis, comprising: The data acquisition module is used to acquire raw data, exhaust temperature curve and ambient humidity data, and process the gas flow fluctuation and water pump speed change data in the raw data to obtain a perturbation feature sequence. The data construction module is used to extract features from the perturbation feature sequence and combine them with the environmental humidity data to construct a relationship graph to obtain the electrical network topology. The data calculation module is used to calculate the fluctuation value of the perturbation feature sequence. When the fluctuation value exceeds a preset fluctuation threshold, the connection weights and abnormal nodes are extracted from the electrical network topology to obtain a set of abnormal nodes. The data traversal module is used to traverse each node in the abnormal node set, find the fault starting point, and obtain the initial fault starting point. The data analysis module is used to analyze the peak shift of the exhaust temperature curve based on the initial fault initiation point, determine the propagation influence range, extract the nodes within the propagation influence range, and obtain the root cause identification node. The data optimization module is used to optimize the path according to the electrical network topology if the root cause identification node contains multiple nodes, thereby obtaining an optimized fault link. The data generation module is used to generate a diagnostic report template based on the optimized fault link and obtain the final diagnostic report output result.
[0008] Thirdly, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the wall-hung boiler fault diagnosis method based on multi-source data analysis as described above.
[0009] Fourthly, the present invention also provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the wall-hung boiler fault diagnosis method based on multi-source data analysis described above.
[0010] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention generates a perturbation feature sequence by processing the fluctuation of gas flow and the change of water pump speed. This can transform the dynamic correlation of multiple parameters into quantifiable time-series features, thereby capturing subtle system-level anomalies that cannot be identified by traditional threshold judgment, and finally achieving accurate early warning of early faults, thus solving the problem of missed detection of time-series perturbations.
[0011] (2) By combining environmental humidity data to construct the electrical network topology, this invention can dynamically reveal the correlation between components such as "humidity sensor → ignition electrode → burner", thereby quantifying environmental factors and incorporating them into the fault diagnosis logic, and ultimately achieving accurate differentiation of humidity-induced faults, avoiding misjudging environmental problems as hardware defects.
[0012] (3) By establishing a standardized process from data preprocessing to report generation, this invention can automatically process data using algorithms such as IQR and linear interpolation, reduce subjective biases caused by human intervention, thereby ensuring the consistency of diagnostic logic in different scenarios, and ultimately achieving efficient and structured diagnostic report output, significantly reducing reliance on human experience. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of the fault diagnosis method for wall-hung boilers based on multi-source data analysis provided in the first embodiment of the present invention; Figure 2 This is a schematic diagram of the wall-hung boiler fault diagnosis system based on multi-source data analysis provided in the second embodiment of the present invention. Detailed Implementation
[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] Reference Figure 1 The first embodiment of the present invention provides a fault diagnosis method for wall-hung boilers based on multi-source data analysis, including the following steps: S11, acquire raw data, exhaust temperature curve and ambient humidity data, and process the gas flow fluctuation and water pump speed change data in the raw data to obtain a perturbation feature sequence. S12, extract features from the perturbation feature sequence and construct a relationship graph by combining the environmental humidity data to obtain the electrical network topology; S13, calculate the fluctuation value of the perturbation feature sequence. When the fluctuation value exceeds a preset fluctuation threshold, extract the connection weights and abnormal nodes from the electrical network topology to obtain a set of abnormal nodes. S14, Traverse each node in the abnormal node set, find the fault starting point, and obtain the initial fault starting point; S15, Analyze the peak shift of the exhaust temperature curve based on the initial fault starting point, determine the propagation influence range, extract the nodes within the propagation influence range, and obtain the root cause identification node; S16, If the root cause identification node contains multiple nodes, then optimize the path according to the electrical network topology to obtain the optimized fault link; S17. Generate a diagnostic report template based on the optimized fault link to obtain the final diagnostic report output result.
[0016] It should be noted that the steps of this method are coordinated through algorithms (such as LSTM time series analysis and IQR data cleaning) to ultimately achieve the following: Figure 1 The standardized operation of the process shown ensures consistent diagnostic logic. Its core lies in mapping the multi-source time-series data of the wall-hung boiler operation (especially ambient humidity) into a dynamic electrical network topology for the first time. By analyzing the propagation path of the perturbation features in this topology, accurate early fault diagnosis and root cause differentiation can be achieved.
[0017] In step S11, raw data, exhaust temperature curve, and ambient humidity data are acquired, and the gas flow fluctuation and water pump speed change data in the raw data are processed to obtain a perturbation feature sequence, including: The raw data is acquired, and the gas flow fluctuation and water pump speed change data in the raw data are analyzed in time series to obtain the ignition signal timing. By comparing the deviation between the ignition signal timing and the pre-established standard timing template, the characteristic value of the deviation is determined, and a perturbation characteristic sequence is obtained.
[0018] It should be noted that the raw data is acquired through the distributed data acquisition module of the wall-hung boiler. This module includes a gas flow sensor and a water pump speed sensor, which collect data from the sensors at a frequency of 1 second per acquisition, storing the raw dataset in the format of "timestamp-gas flow value-water pump speed value". Temperature data is collected by an exhaust temperature sensor at the same frequency as the distributed data acquisition module, and these data are concatenated along the time dimension to form an exhaust temperature curve, synchronously linked to the timestamp of the raw data. Ambient humidity is acquired using an environmental humidity sensor, which collects data at the same frequency and stores it in the format of "timestamp-humidity value", perfectly aligned with the time dimension of the raw data and the exhaust temperature curve.
[0019] A Long Short-Term Memory (LSTM) network was used to perform time-series analysis on the raw data to uncover the time dependency between gas flow fluctuations, water pump speed changes, and ignition signals, generating ignition signal time series data. The construction and training process of the LSTM network was as follows: the input layer was fed with time-series data of gas flow and water pump speed (each input vector consisted of 10 consecutive data points, corresponding to 1 second of operation); the hidden layer consisted of 2 layers of LSTM units (32 neurons per layer), trained based on simulated data, laboratory controlled fault injection, or public datasets, with 1000 training iterations and a validation set accuracy of 95%; irrelevant instantaneous fluctuations were removed by a "forget gate," while the "input gate" retained key ignition-related features such as sudden drops in gas flow (e.g., <1.8 m³ / h) and sudden increases in water pump speed (e.g., >1300 rpm); the output layer output the "trigger timestamp" of the ignition signal (i.e., the actual trigger time of the ignition electrode, accurate to the millisecond level), thus obtaining the ignition signal time series data.
[0020] A standard timing template for ignition signals is pre-established. Based on historical data of the wall-hung boiler under normal operating conditions (humidity 40%-60%RH, gas pressure 1.5-2.0kPa), the standard trigger interval of the ignition signal is statistically analyzed to form the standard timing template. A fixed time window alignment method is used to align the ignition signal timing with the standard template in terms of time dimension (to resolve the length difference caused by fluctuations in operating conditions). The window duration is set according to the sensor acquisition frequency, and the time axes of both the actual timing and the standard template are divided into continuous windows. Then, within each window, the trigger time of the actual timing and the trigger time of the standard template are extracted, and the deviation value of each trigger time is calculated (the deviation value is equal to the actual trigger time minus the standard trigger time). If the deviation value is within the allowable range of the standard template (e.g., the deviation value is less than 3ms when the humidity is greater than 60%RH), it is judged as normal fluctuation and no feature is extracted. If the deviation value exceeds the allowable range (e.g., the deviation value is greater than 3ms when the humidity is greater than 60%RH, this threshold is based on statistical data from fault simulation experiments of 100 wall-hung boilers under different humidity environments), then the feature value of the deviation (including deviation amplitude, number of consecutive deviations, and deviation duration) is extracted. All deviation feature values exceeding the threshold are concatenated in chronological order to form a perturbation feature sequence.
[0021] In step S12, features are extracted from the perturbation feature sequence and combined with the environmental humidity data to construct a relationship graph, thereby obtaining the electrical network topology, including: The first feature dataset is obtained by extracting features associated with the exhaust temperature curve from the perturbation feature sequence; The second feature dataset is obtained by fusing the environmental humidity data with the first feature dataset; Based on the temperature and humidity data in the second feature dataset, a relationship graph is constructed using edges to obtain the component relationship graph; The component relationship diagram is optimized, and the coupling relationship between nodes is updated to obtain the electrical network topology.
[0022] It should be noted that, through time-series correlation analysis, the covariance between perturbation features and exhaust temperature curves is explored, and correlation features are extracted. The perturbation feature sequence and exhaust temperature curve are divided into 1-second windows, allowing both to be analyzed within the same time unit. The perturbation-temperature correlation coefficient of the "mean perturbation deviation amplitude" and "exhaust temperature fluctuation amplitude" within each window is calculated. The perturbation-temperature correlation coefficient is used to measure the degree of linear correlation between two variables (such as the mean perturbation deviation amplitude and the exhaust temperature fluctuation amplitude). Let the two variables be X (representing the mean perturbation deviation amplitude) and Y (representing the exhaust temperature fluctuation amplitude), each containing n samples (corresponding to n time windows). The mean of X is calculated. and the mean of Y Substitute into the correlation coefficient calculation formula: Extract the time difference between the time of the perturbation and the time of the peak temperature shift in the exhaust gas (e.g., the temperature peak drops 0.5 seconds after the perturbation occurs); count the number of consecutive perturbations and the duration of the exhaust gas temperature being below the threshold (e.g., 175℃) (e.g., 3 consecutive perturbations correspond to a temperature below the threshold for 2 seconds); sort the above-mentioned related features by time window in the format of "window ID-correlation coefficient-time difference-duration duration" to obtain the first feature dataset.
[0023] Introducing ambient humidity sensor data, the impact of humidity on the correlation between perturbation and temperature is quantified. The average humidity value within each window is calculated, and the regression coefficient between "window average humidity" and "perturbation-temperature correlation coefficient" is calculated. Let independent variable A (window average humidity (%RH)) and dependent variable B (perturbation-temperature correlation coefficient). The regression coefficient is solved using the least squares method: a Extract the "perturbation features when humidity is higher than 60%RH"; add humidity features to the first feature dataset, in the format of "window ID-correlation coefficient-time difference-duration-average humidity-regression coefficient", to obtain the second feature dataset.
[0024] Using the core components of the wall-hung boiler as nodes, and the temperature and humidity correlation features in the second feature dataset as edge weights, an initial relationship graph is constructed. The core components of the wall-hung boiler include four types of nodes: ignition electrode (perturbation source), burner (temperature generation), exhaust sensor (temperature detection), and humidity sensor (environmental factor). The edge weights between nodes are calculated: the edge weight between ignition electrode and burner is equal to the correlation coefficient between perturbation and temperature (reflecting the influence of ignition perturbation on combustion, e.g., 0.85); the edge weight between burner and exhaust sensor is equal to the exhaust temperature fluctuation amplitude (reflecting the influence of combustion state on temperature detection, e.g., ±3℃); the edge weight between humidity sensor and ignition electrode is equal to the regression coefficient of humidity-Pearson correlation coefficient (reflecting the influence of humidity on ignition perturbation, e.g., 0.1 / 5%RH); the edge weight between humidity sensor and burner is equal to the correlation coefficient between humidity and the duration of exhaust temperature below the threshold (the calculation process is the same as the calculation of the correlation coefficient between perturbation and temperature, calculating the mean of humidity and the duration of exhaust temperature below the threshold respectively, and inputting it into the correlation coefficient calculation formula); the nodes are connected by directed edges to obtain the component relationship graph.
[0025] By dynamically adjusting the node coupling relationship, the initial relationship graph is optimized, and weakly related edges with a weight lower than 0.3 are removed (such as the humidity sensor → exhaust sensor with a weight of 0.2, which has no actual electrical connection and is therefore deleted); the output is a topology graph containing nodes (ignition electrode, burner, exhaust sensor, humidity sensor, control module), edge weights, and connection directions.
[0026] In step S13, the fluctuation value of the perturbation feature sequence is calculated. When the fluctuation value exceeds a preset fluctuation threshold, connection weights and abnormal nodes are extracted from the electrical network topology to obtain a set of abnormal nodes, including: Get the error code log; Extract abnormal records related to the environmental humidity data from the fault code log to obtain the abnormal pattern; The fluctuation value of the perturbation feature sequence is calculated. When the fluctuation value exceeds a preset fluctuation threshold, the connection weights and abnormal nodes associated with the abnormal pattern are extracted from the electrical network topology to obtain an abnormal node set.
[0027] It should be noted that historical fault code logs exported from the wall-hung boiler's control system are stored in a structured format of "timestamp-fault code-fault description-environmental parameter snapshot". Through humidity threshold filtering and feature summarization, typical patterns of humidity-sensitive faults are extracted from the logs. Humidity thresholds are set (e.g., >60%RH is considered high humidity), based on fault simulation experimental data from 300 wall-hung boilers under different humidity environments. All fault records occurring in high humidity environments are filtered out, and the frequency of each fault code in high humidity environments is calculated. For humidity-sensitive faults, their associated "perturbation features" and "temperature features" are extracted. Anomaly patterns are generated in the format of "fault code-humidity condition-perturbation feature-temperature feature". By statistically analyzing the dispersion of perturbation features, their fluctuation intensity is quantified. Using the "deviation amplitude" in the perturbation feature sequence as the main indicator, the "sliding window standard deviation" method is employed: the window size is set to 5 consecutive perturbation data points, and the standard deviation of the deviation amplitude within each window is calculated (the larger the standard deviation, the more severe the fluctuation). The average of all window standard deviations is taken as the final fluctuation value. The preset fluctuation threshold needs to be set in conjunction with the perturbation characteristics during normal equipment operation. Based on historical normal data (humidity 40%-60%RH), the fluctuation threshold is set as three times the standard deviation of the statistical perturbation deviation amplitude, with a specific value of 0.15ms.
[0028] When the fluctuation value exceeds the preset fluctuation threshold, in the electrical network topology, the nodes associated with the abnormal mode are searched: ignition electrode (micro-disturbance source, corresponding to deviation amplitude characteristics); burner (corresponding to exhaust temperature characteristics); humidity sensor (environmental factor node); the connection weights between the above nodes are extracted (e.g., the weight of ignition electrode → burner is 0.85). If the weight is > 0.7 (strong correlation), it is marked as an abnormal associated node; all marked nodes are summarized to obtain the abnormal node set, in the format of "node name - abnormal feature - connection weight".
[0029] In step S14, each node in the abnormal node set is traversed to find the fault starting point and obtain the initial fault starting point, including: Traverse the connection relationships of each node in the abnormal node set to obtain an initial propagation path set; wherein, the initial propagation path set includes humidity-sensitive paths; The initial propagation path set is sorted according to the preset repair priority to obtain a priority sorted dataset; If the priority of the humidity-sensitive path in the priority ranking dataset is higher than the preset priority threshold, then the propagation path is extracted from the priority ranking dataset, and the propagation path is pattern matched to obtain a set of fault propagation patterns. The initial fault origin is obtained by locating the fault origin based on the set of fault propagation patterns.
[0030] It should be noted that, from the electrical network topology, the input / output connections of each node in the abnormal node set are extracted (e.g., the input connection of the abnormal node "ignition electrode" is "control module", and the output connection is "burner"). The path chain is traversed in the direction of "upstream node → current node → downstream node". For example: control module → ignition electrode → burner → exhaust sensor (containing 2 abnormal nodes: ignition electrode and burner); humidity sensor → ignition electrode → burner (containing 2 abnormal nodes: ignition electrode and burner). If the path contains a "humidity sensor" node or the connection weight is affected by humidity (e.g., the weight of humidity sensor → ignition electrode > 0.6), it is marked as a "humidity-sensitive path". The set format is "path ID - node sequence - whether it is humidity-sensitive", thus obtaining the initial propagation path set.
[0031] Based on the impact of a path on equipment safety and operational efficiency, repair priority rules are set: the more abnormal nodes a path contains, the higher its priority; the more times a path appears in the fault log in terms of historical fault frequency, the higher its priority. Paths are sorted from highest to lowest priority and stored in the format "Path ID-Priority-Node Sequence" to obtain a priority-ranked dataset. From the priority-ranked dataset, data on paths with scores ≥ threshold and that are humidity-sensitive are extracted. If the priority of a humidity-sensitive path is higher than the preset priority threshold, the selected paths are matched with the abnormal patterns in step S13, and the data is stored in the format "Pattern ID-Path-Matched Fault Code" to generate a fault propagation pattern set.
[0032] In the fault propagation path, the upstream abnormal node (i.e., the first node in the path) is identified. For example, in M1 mode, the path is "humidity sensor → ignition electrode → burner," and the upstream node is either "humidity sensor" or "ignition electrode." If the detection data of the "humidity sensor" is abnormal (e.g., the actual humidity is 60% but it displays 80%), it is determined to be the fault starting point. If the humidity sensor is normal, but the perturbation characteristic of the "ignition electrode" (deviation of 3.2ms) is the earliest abnormality in the path (0.5 seconds earlier than the burner temperature abnormality), then the "ignition electrode" is determined to be the fault starting point, and the initial fault starting point is obtained.
[0033] In step S15, based on the initial fault initiation point, the peak shift of the exhaust temperature curve is analyzed to determine the propagation influence range. Nodes within the propagation influence range are extracted to obtain root cause identification nodes, including: Get the node identifier; Calculate the fluctuation amplitude of the perturbation feature sequence. If the fluctuation amplitude is higher than the preset feature amplitude threshold, extract the temperature curve data from the exhaust temperature curve, perform peak analysis on the temperature curve data, and obtain the peak offset dataset. Determine the propagation influence range of the peak offset dataset and the perturbation feature sequence to obtain the propagation influence range dataset; A correlation analysis is performed between the dataset of the propagation impact range and the node identifier to obtain the root cause identifier node.
[0034] It should be noted that a unique identifier for all components is extracted from the wall-hung boiler components to establish a mapping relationship between nodes and physical components. The identifier is defined in the format of "component type + number", for example, ignition electrode → IE-01, burner → B-01, humidity sensor → HS-01, exhaust sensor → ES-01. The comprehensive fluctuation amplitude is calculated using the "number of consecutive deviations" and "duration of deviation" in the perturbation feature sequence as indicators: the fluctuation amplitude is equal to the number of consecutive deviations divided by the normal consecutive number threshold (set to 5 times, based on the statistical analysis of the maximum probability distribution of the continuous occurrence of perturbations during normal operation) multiplied by 0.6 plus (the duration of deviation divided by the normal duration threshold (set to 2 seconds, which can be set based on the experimental data analysis of the thermal inertia characteristics of the wall-hung boiler)) multiplied by 0.4. The feature amplitude threshold is preset (e.g., 1.0, based on the fluctuation amplitude statistics of 500 wall-hung boilers during normal operation, the upper limit of the 95% confidence interval is set as the feature amplitude threshold). If the fluctuation amplitude (1.08) > the threshold, the exhaust temperature curve analysis is triggered, and the temperature curve data within the corresponding time range (e.g., the exhaust temperature 5 seconds before and after the perturbation occurs) is extracted.
[0035] By comparing temperature peaks under normal and abnormal conditions, the degree of deviation is quantified. Based on the exhaust temperature curve of the wall-hung boiler during normal operation (humidity 40%-60%RH, no fault codes), the peak time (e.g., 2 seconds after ignition) and peak temperature (e.g., 185℃) are statistically analyzed as benchmarks. From the temperature curve data, the current peak time (e.g., 2.5 seconds after ignition) and peak temperature (e.g., 175℃) are identified, and the peak deviation is calculated. The peak deviation includes time deviation and temperature deviation. The time deviation is equal to the current peak time minus the benchmark peak time (positive deviation, indicating peak delay); the temperature deviation is equal to the current peak temperature minus the benchmark peak temperature (negative deviation, indicating peak reduction). The data is stored in the format of "time deviation - temperature deviation - corresponding perturbation time" to generate a peak deviation dataset.
[0036] The spatiotemporal relationship between the peak offset and the perturbation features is correlated to define the scope of fault propagation. The perturbation features corresponding to the peak offset time (16:00:02.000) are identified (e.g., continuous deviation occurs at 16:00:01.500), and the time difference (0.5 seconds) is determined, which is the time it takes for the fault to propagate from the perturbation node to the temperature detection node. In the electrical network topology, with the initial fault origin (e.g., IE-01) as the center, the propagation speed is multiplied by the time difference (preset propagation speed is 0.2 nodes / second; based on the experimental determination of the transmission delay of electrical signals in the wall-hung boiler control system, the average transmission delay is found to be 5 seconds / node, so the propagation speed is set to 1 node divided by 5 seconds equals 0.2 nodes / second) to define the affected nodes (e.g., B-01 and ES-01 can be affected within 0.5 seconds). The data is stored in the format of "affected node identifier - time offset - temperature offset correlation degree" (the higher the correlation degree, the greater the impact of the peak offset on the node), generating a propagation influence range dataset.
[0037] By analyzing the abnormal correlations of nodes within the affected area, the core root cause node is located. The node identifiers (e.g., B-01) in the affected area dataset are matched with the node identifier set to obtain their physical attributes (e.g., the normal temperature range of B-01). The physical attributes of a node refer to the inherent parameters, functional characteristics, and normal operating thresholds of each component of the wall-hung boiler (e.g., ignition electrode, burner, sensors, etc.). These attributes need to be obtained through a combination of "extraction from the equipment manual + real-time data calibration." The technical manual of the wall-hung boiler will clearly specify the core parameters of each component. For example: Ignition electrode (IE-01): Normal ignition voltage range (15- 20kV), trigger interval standard value (2ms±0.5ms), humidity tolerance upper limit (85%RH); due to equipment aging or environmental changes, some physical properties will drift over time and need to be corrected by real-time data collection: for example, the "normal exhaust temperature range" of the burner is 170-190℃ when the machine is new, and can be calibrated after 2 years of use based on historical data of fault-free operation (such as 165-185℃); the calibration logic is to take the 95% confidence interval of the normal operation data of the past 3 months as the new physical property, bind the extracted physical property with the node identifier, and store it as "node identifier-physical property key-value pair".
[0038] For each affected node, the anomaly contribution is calculated as the node's anomaly degree multiplied by the propagation path weight. The node with the highest anomaly contribution is taken as the root cause identification node. The node's anomaly degree is a quantitative indicator that measures the deviation of a single node from its physical properties, reflecting the severity of the node's failure. The anomaly degree of each node needs to be determined in conjunction with its core functional parameters. For example, for the ignition electrode (IE-01): the anomaly dimensions are "deviation amplitude" (the difference between the actual trigger interval and the standard value) and "number of consecutive deviations". For each anomaly dimension, the "standardized anomaly value" is calculated (converting the deviation between the actual value and the physical property threshold into a score of 0-10, with higher scores indicating more severe anomalies): the anomaly value is equal to (actual value minus physical property) divided by (physical property multiplied by 0.2) multiplied by 10. The average of the multiple anomaly dimensions of the node is taken to obtain the comprehensive anomaly degree.
[0039] In step S16, if the root cause identification node contains multiple nodes, then the optimized fault link is obtained according to the electrical network topology optimization path, including: Get data on the frequency of interactions between components; Feature extraction is performed on the interaction frequency data to obtain an interaction frequency feature sequence; Calculate the fluctuation amplitude of the interaction frequency feature sequence. If the fluctuation amplitude is higher than a preset frequency amplitude threshold, generate the interaction strength data between nodes to obtain the link weight dataset. Based on the environmental humidity data, the link weight dataset is matched with the propagation path of the electrical network topology to obtain the humidity-induced path dataset; If the root cause identification node contains multiple nodes, then the humidity-induced path dataset is correlated with the node identification to obtain the optimized fault link.
[0040] It should be noted that the real-time interaction frequency between nodes (components) in the electrical network is collected to quantify the component cooperation strength. Signal transmission between components (such as the trigger signal from the control module to the ignition electrode) and data exchange (such as the humidity sensor to the control module for humidity data) are all considered as one interaction. The interaction frequency of each node pair is counted in a 1-minute window. For example, humidity sensor (HS-01) → ignition electrode (IE-01): 120 interactions within 1 minute (humidity data is transmitted once every 0.5 seconds). The data is stored in the format of "node pair - time window - number of interactions" to generate interaction frequency data. A 5-minute sliding window (1-minute step) is used to calculate the interaction frequency statistics within each window. The average number of interactions within the time window is calculated (e.g., HS-01→IE-01 has an average of 115 interactions / minute within a 5-minute window). The frequency volatility is obtained by dividing the average number of interactions by the maximum number of interactions within the window (e.g., 120→110→115→125→105, volatility = (125-105) / 115≈17.4%). The Pearson correlation coefficient between the interaction frequency and the average humidity within the window is calculated (the calculation process is the same as the calculation process for the correlation coefficient between perturbation and temperature, calculating the mean of interaction frequency and humidity respectively, and inputting them into the correlation coefficient calculation formula). The features of each window are concatenated in chronological order, in the format of "node pair-window sequence-mean-volatility-humidity correlation", to obtain the interaction frequency feature sequence.
[0041] Using frequency volatility as the core indicator, the overall volatility of the characteristic sequence is calculated. The volatility is equal to the average volatility of each window (e.g., the volatility of the five windows from HS-01 to IE-01 are 17.4%, 18.2%, 16.8%, 19.5%, and 17.1%, respectively, with an average volatility of 17.8%). A preset frequency volatility threshold (e.g., 20%, based on the upper limit of the 95% confidence interval of historical fault-free data) can be used to collect historical interaction frequency data of the wall-hung boiler under fault-free and standard conditions, and calculate the "volatility distribution" of each component's interaction frequency (e.g., the interaction frequency volatility between the ignition electrode and the burner falls within 15% for 95% of cases). The upper limit of the 95% confidence interval is taken as the basic threshold (e.g., 15% → core component threshold), combined with the interaction frequency under different humidity environments. Inter-frequency characteristics were used to establish a "humidity-volatility correction coefficient": Experimental data showed that for every 10% increase in RH, the average volatility of the interaction frequency of humidity-sensitive components increased by 5% (e.g., volatility was 15% at 60% humidity and increased to 25% at 80% humidity). Therefore, the frequency amplitude threshold was equal to the base threshold plus (the proportion of humidity exceeding 60% divided by 10%) multiplied by 5%. If the volatility amplitude (17.8%) was less than the threshold, it indicated stable interaction, and the weight was directly calculated using "frequency mean multiplied by humidity correlation". If the volatility amplitude was greater than the threshold, a "volatility penalty coefficient" was added (e.g., for every 5% increase in volatility, the weight decreased by 10%). In this case, the link weight was calculated using the weight calculation formula as (frequency mean divided by the maximum frequency mean) multiplied by (humidity correlation plus 1) multiplied by the penalty coefficient (normalized to 0-1). The data was stored in the format of "node pair-link weight-humidity correlation" to obtain the link weight dataset.
[0042] Select node pairs with a humidity correlation greater than 0.5 from the link weight dataset. In the electrical network topology, starting from the root cause identifier node, construct paths based on node pairs with link weights less than 0.6 (strong correlation). For example, when the root cause identifier nodes are HS-01, IE-01, and B-01, the matching path is: HS-01→IE-01→B-01 (link weights are 0.82 and 0.75 respectively, both greater than 0.6). Calculate the humidity sensitivity of each node pair in the path. Humidity sensitivity is equal to the link weight multiplied by the humidity correlation. If the total path sensitivity (the sum of the sensitivities of each node pair) is greater than 2.0, it is marked as a humidity-induced path. Store the data in the format of "path ID-node sequence-total humidity sensitivity-link weight list" to obtain the humidity-induced path dataset.
[0043] When there are multiple root cause identification nodes, select the path that covers all root cause identification nodes from the humidity-induced path dataset (e.g., HS-01→IE-01→B-01 covers 3 nodes, which meets the condition); remove redundant nodes in the path that are not related to the root cause identification (e.g., the "control module" in the path HS-01→control module→IE-01→B-01 is not a root cause node, which is simplified to HS-01→IE-01→B-01); sort the paths in descending order of "total humidity sensitivity" and select the path with the highest score as the optimized fault link.
[0044] In step S17, a diagnostic report template is generated based on the optimized fault link to obtain the final diagnostic report output, including: Link data is extracted from the optimized faulty link, and the link data and the environmental humidity data are preprocessed to obtain a cleaning dataset; The link data and humidity data in the cleaning dataset are grouped to obtain a humidity-related path dataset; Calculate the path coverage rate of the humidity-related path dataset. If the path coverage rate is higher than a preset coverage threshold, analyze the accuracy of the link data and the node identifier to obtain an accuracy dataset. The accuracy dataset is then fused with a preset report template to obtain the final diagnostic report output.
[0045] It should be noted that the faulty link data and environmental humidity data were cleaned and standardized. Key information was extracted from the faulty links, including node sequences (e.g., HS-01→IE-01→B-01), link weights between nodes (0.82, 0.75), total humidity sensitivity (1.02), and corresponding timestamps. Environmental humidity data matching the link data timestamps (e.g., humidity values of 70%, 72%, and 75% from 9:00 to 9:05) were extracted. Extreme values in the link weights were removed using the IQR (interquartile range) method (e.g., if a node's weight suddenly drops to 0.1, it is determined to be a data transmission error and is removed). Sort the target data (such as link weights) in ascending order and calculate the first quartile (the median of the top 50% of the data, i.e., the 25th percentile) and the third quartile (the median of the bottom 50% of the data, i.e., the 75th percentile). The interquartile range (IQR) is equal to the third quartile minus the first quartile. The lower limit is equal to the first quartile minus 1.5 times the IQR, and the upper limit is equal to the third quartile plus 1.5 times the IQR. All data below the lower limit or above the upper limit are considered outliers. For missing values in the humidity data, linear interpolation is used to complete the missing data to obtain a clean dataset. Linear interpolation estimates missing values based on the linear relationship between two adjacent known data points. Let the previous known point of the missing value be (x0, y0) and the next known point be (x1, y1). The linear interpolation formula is y = (x0, y0) / y1. .
[0046] The link data is grouped according to humidity ranges, dividing the humidity data into three ranges: low humidity (≤40%RH), medium humidity (41%-60%RH), and high humidity (≥61%RH), to adapt to the operating characteristics of wall-hung boilers under different humidity conditions. The link data in the clean data set is assigned to the corresponding humidity range according to the timestamp (e.g., if the timestamp of a link data corresponds to 75% humidity, it is assigned to the high humidity group). The link data within the same range is merged according to the node sequence, and the average link weight and the average total humidity sensitivity are calculated. The data is stored in the format of "humidity range-node sequence-average link weight-average total humidity sensitivity" to obtain the humidity association path dataset.
[0047] Path coverage rate refers to the proportion of root cause identification nodes in optimized fault link coverage to the total number of root cause identification nodes within a certain humidity range. Path coverage rate is equal to (the number of root cause nodes included in the link divided by the total number of root cause nodes) multiplied by 100%. A preset coverage rate threshold is set (e.g., 80%, based on statistics from 200 fault simulation experiments, when the coverage rate is ≥80%, the fault location accuracy can be stably above 90%). If the path coverage rate (100%) of a certain humidity range is greater than the threshold, it indicates that the link in that range can effectively cover the root cause nodes, and it enters the accuracy analysis; otherwise, it is judged as "insufficient data" and is not included in the subsequent analysis. The number of root cause nodes confirmed by repair is extracted from the actual fault repair records, and the accuracy rate is calculated by dividing the number of accurately identified root cause nodes by the number of root cause nodes included in the link. For example, if the link contains 3 nodes, of which 2 are confirmed by repair, the accuracy rate = 2 / 3 ≈ 66.7%. The data is stored in the format of "humidity range - path coverage rate - location accuracy rate" to obtain the accuracy dataset.
[0048] The preset report template structure includes fixed modules: fault link description, humidity impact analysis, root cause node location results, accuracy assessment, and maintenance recommendations. The fault link description module is populated with the "node sequence" and "average link weight" from the humidity-related path dataset. The humidity impact analysis module is populated based on the "total humidity sensitivity mean" (e.g., "total humidity sensitivity 1.02, indicating that high humidity has a significant impact on this link"). The location accuracy module is populated with the "location accuracy" from the accuracy dataset, and the amount of verification data is noted (e.g., "location accuracy 66.7%, verified based on 15 maintenance records"). The maintenance recommendations module combines root cause nodes and humidity impact to generate targeted recommendations (e.g., "prioritize checking the humidity protection of IE-01, then calibrate the combustion parameters of B-01"). A combination of text and graphics is used, including a humidity-fault frequency trend chart (drawn based on the humidity-related path dataset), and the final output is a diagnostic report in PDF or text format.
[0049] In summary, this invention generates a perturbation feature sequence by processing gas flow fluctuations and water pump speed changes, transforming the dynamic correlation of multiple parameters into analyzable features, capturing system-level anomalies, and constructing an electrical network topology by combining environmental humidity data. This clarifies cross-component relationships such as "humidity sensor → ignition electrode → burner," revealing the indirect impact of environmental factors on multiple parameters and avoiding misdiagnosis of environmentally induced faults as hardware defects. Using "environmental humidity data" as the core clue, this invention dynamically adjusts the analysis logic to adapt diagnostic results to different environments, clarifies the quantitative relationship between humidity and faults, and provides "humidity protection" recommendations in the report, solving the problem of traditional solutions that "only repair hardware, not protect against the environment." A standardized process is established (from data preprocessing to report generation), and algorithms such as IQR and linear interpolation are used to eliminate human subjective bias, ensuring consistent diagnostic logic across different scenarios. A structured report containing "link weights, accuracy, and humidity impact" is generated, transforming implicit experience into explicit data and reducing reliance on human experience.
[0050] Reference Figure 2 The second embodiment of the present invention provides a wall-hung boiler fault diagnosis system based on multi-source data analysis, comprising: The data acquisition module is used to acquire raw data and ambient humidity data, and process the gas flow fluctuation and water pump speed change data in the raw data to obtain a perturbation feature sequence. The data construction module is used to extract features from the perturbation feature sequence and combine them with the environmental humidity data to construct a relationship graph to obtain the electrical network topology. The data calculation module is used to calculate the fluctuation value of the perturbation feature sequence. When the fluctuation value exceeds a preset fluctuation threshold, the connection weights and abnormal nodes are extracted from the electrical network topology to obtain a set of abnormal nodes. The data traversal module is used to traverse each node in the abnormal node set, find the fault starting point, and obtain the initial fault starting point. The data analysis module is used to analyze the peak shift of the exhaust temperature curve in the perturbation feature sequence based on the initial fault initiation point, determine the propagation influence range, extract the nodes within the propagation influence range, and obtain the root cause identification node. The data optimization module is used to optimize the path according to the electrical network topology if the root cause identification node contains multiple nodes, thereby obtaining an optimized fault link. The data generation module is used to generate a diagnostic report template based on the optimized fault link and obtain the final diagnostic report output result.
[0051] It should be noted that the wall-hung boiler fault diagnosis system based on multi-source data analysis provided in this embodiment of the invention is used to execute all the process steps of the wall-hung boiler fault diagnosis method based on multi-source data analysis in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.
[0052] This invention also provides an electronic device. The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a fault diagnosis program. When the processor executes the computer program, it implements the steps described in the various embodiments of the wall-hung boiler fault diagnosis method based on multi-source data analysis, for example... Figure 1 The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above-described device embodiments, such as the data generation module.
[0053] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.
[0054] The electronic device may be a desktop computer, laptop, handheld computer, or smart tablet, etc. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the above components are merely examples of electronic devices and do not constitute a limitation on the electronic device. It may include more or fewer components than described above, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.
[0055] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the electronic device, connecting various parts of the electronic device through various interfaces and lines.
[0056] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital card (SD card), flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0057] If the modules / units integrated in the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0058] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0059] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A fault diagnosis method for wall-hung boilers based on multi-source data analysis, characterized in that, include: The raw data, exhaust temperature curve, and ambient humidity data are acquired, and the gas flow fluctuation and water pump speed change data in the raw data are processed to obtain a perturbation feature sequence. Features are extracted from the perturbation feature sequence and combined with the environmental humidity data to construct a relationship graph, thereby obtaining the electrical network topology. The fluctuation value of the perturbation feature sequence is calculated. When the fluctuation value exceeds a preset fluctuation threshold, the connection weights and abnormal nodes are extracted from the electrical network topology to obtain a set of abnormal nodes. Traverse each node in the set of abnormal nodes to find the fault origin and obtain the initial fault origin. Based on the initial fault initiation point, the peak shift of the exhaust temperature curve is analyzed to determine the propagation influence range, and the nodes within the propagation influence range are extracted to obtain the root cause identification node; If the root cause identification node contains multiple nodes, then the optimized fault link is obtained by optimizing the path according to the electrical network topology. Based on the optimized fault link, a diagnostic report template is generated to obtain the final diagnostic report output result; The step of extracting features from the perturbation feature sequence and constructing a relationship graph based on the environmental humidity data to obtain the electrical network topology includes: extracting features associated with the exhaust temperature curve from the perturbation feature sequence to obtain a first feature dataset; fusing the environmental humidity data with the first feature dataset to obtain a second feature dataset; constructing a relationship graph using the temperature and humidity data in the second feature dataset as edges to obtain a component relationship graph; and optimizing the component relationship graph by updating the coupling relationships between nodes to obtain the electrical network topology. The step of analyzing the peak shift of the exhaust temperature curve based on the initial fault initiation point, determining the propagation influence range, extracting nodes within the propagation influence range, and obtaining root cause identification nodes includes: Obtain node identifiers; calculate the fluctuation amplitude of the perturbation feature sequence; if the fluctuation amplitude is higher than a preset feature amplitude threshold, extract temperature curve data from the exhaust temperature curve, perform peak analysis on the temperature curve data to obtain a peak offset dataset; determine the propagation influence range of the peak offset dataset and the perturbation feature sequence to obtain a propagation influence range dataset; perform correlation analysis between the propagation influence range dataset and the node identifiers to obtain root cause identifier nodes.
2. The wall-hung boiler fault diagnosis method based on multi-source data analysis according to claim 1, characterized in that, The process of acquiring raw data and processing the gas flow fluctuation and water pump speed change data in the raw data to obtain a perturbation feature sequence includes: The raw data is acquired, and the gas flow fluctuation and water pump speed change data in the raw data are analyzed in time series to obtain the ignition signal timing. By comparing the deviation between the ignition signal timing and the pre-established standard timing template, the characteristic value of the deviation is determined, and a perturbation characteristic sequence is obtained.
3. The wall-hung boiler fault diagnosis method based on multi-source data analysis according to claim 1, characterized in that, The calculation of the fluctuation value of the perturbation feature sequence, when the fluctuation value exceeds a preset fluctuation threshold, extracts connection weights and abnormal nodes from the electrical network topology to obtain an abnormal node set, including: Get the error code log; Extract abnormal records related to the environmental humidity data from the fault code log to obtain the abnormal pattern; The fluctuation value of the perturbation feature sequence is calculated. When the fluctuation value exceeds a preset fluctuation threshold, the connection weights and abnormal nodes associated with the abnormal pattern are extracted from the electrical network topology to obtain an abnormal node set.
4. The wall-hung boiler fault diagnosis method based on multi-source data analysis according to claim 1, characterized in that, The process of traversing each node in the set of abnormal nodes to find the fault starting point and obtain the initial fault starting point includes: Traverse the connection relationships of each node in the abnormal node set to obtain an initial propagation path set; wherein, the initial propagation path set includes humidity-sensitive paths; The initial propagation path set is sorted according to the preset repair priority to obtain a priority sorted dataset; If the priority of the humidity-sensitive path in the priority ranking dataset is higher than the preset priority threshold, then the propagation path is extracted from the priority ranking dataset, and the propagation path is pattern matched to obtain a set of fault propagation patterns. The initial fault origin is obtained by locating the fault origin based on the set of fault propagation patterns.
5. The wall-hung boiler fault diagnosis method based on multi-source data analysis according to claim 1, characterized in that, If the root cause identifier node contains multiple nodes, then the optimized fault link is obtained by optimizing the path according to the electrical network topology, including: Get data on the frequency of interactions between components; Feature extraction is performed on the interaction frequency data to obtain an interaction frequency feature sequence; Calculate the fluctuation amplitude of the interaction frequency feature sequence. If the fluctuation amplitude is higher than a preset frequency amplitude threshold, generate the interaction strength data between nodes to obtain the link weight dataset. Based on the environmental humidity data, the link weight dataset is matched with the propagation path of the electrical network topology to obtain the humidity-induced path dataset; If the root cause identification node contains multiple nodes, then the humidity-induced path dataset is correlated with the node identification to obtain the optimized fault link.
6. The wall-hung boiler fault diagnosis method based on multi-source data analysis according to claim 1, characterized in that, The step of generating a diagnostic report template based on the optimized fault link to obtain the final diagnostic report output includes: Link data is extracted from the optimized faulty link, and the link data and the environmental humidity data are preprocessed to obtain a cleaning dataset; The link data and humidity data in the cleaning dataset are grouped to obtain a humidity-related path dataset; Calculate the path coverage rate of the humidity-related path dataset. If the path coverage rate is higher than a preset coverage threshold, analyze the accuracy of the link data and the node identifier to obtain an accuracy dataset. The accuracy dataset is then fused with a preset report template to obtain the final diagnostic report output.
7. A fault diagnosis system for wall-hung boilers based on multi-source data analysis, characterized in that, For implementing the method as described in any one of claims 1-6, comprising: The data acquisition module is used to acquire raw data, exhaust temperature curve and ambient humidity data, and process the gas flow fluctuation and water pump speed change data in the raw data to obtain a perturbation feature sequence. The data construction module is used to extract features from the perturbation feature sequence and combine them with the environmental humidity data to construct a relationship graph to obtain the electrical network topology. The data calculation module is used to calculate the fluctuation value of the perturbation feature sequence. When the fluctuation value exceeds a preset fluctuation threshold, the connection weights and abnormal nodes are extracted from the electrical network topology to obtain a set of abnormal nodes. The data traversal module is used to traverse each node in the abnormal node set, find the fault starting point, and obtain the initial fault starting point. The data analysis module is used to analyze the peak shift of the exhaust temperature curve based on the initial fault initiation point, determine the propagation influence range, extract the nodes within the propagation influence range, and obtain the root cause identification node. The data optimization module is used to optimize the path according to the electrical network topology if the root cause identification node contains multiple nodes, thereby obtaining an optimized fault link. The data generation module is used to generate a diagnostic report template based on the optimized fault link and obtain the final diagnostic report output result.