Intelligent fault diagnosis and early warning method for electrical automatic composting process system
By combining sliding window filtering, hidden Markov models, and density peak clustering with XGBoost classification models, the problem of fault diagnosis in composting systems under dynamic environments is solved. This enables stable operation of the composting system and accurate fault identification, supports strong and weak early warning, and reduces the impact of noise data.
Patent Information
- Application Number
- CN202511018980.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies cannot effectively adapt to fluctuations in the composition of compost raw materials and interference from environmental temperature, leading to sensor data drift, affecting clustering results, failing to accurately identify composting system malfunctions, and causing interruptions in the fermentation process and environmental pollution.
High-frequency noise is removed by sliding window filtering, composting stages are divided by hidden Markov model, online K-means++ algorithm and density peak clustering technology are used, and real-time fault diagnosis is performed by XGBoost classification model to generate enhanced feature vectors and provide early warning.
It achieves stable operation of the composting system, reduces the sensitivity to noise data, accurately distinguishes between normal fluctuations and abnormal faults, supports strong and weak early warning, and reduces the risk of fermentation interruption and material spoilage.
Smart Images

Figure CN120850107A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent monitoring technology, specifically to a method for intelligent diagnosis and early warning of faults in an electrically automated composting process system. Background Technology
[0002] With the acceleration of urbanization, the output of organic waste such as agricultural waste (such as straw and livestock manure) and urban kitchen waste has surged. Traditional landfill or incineration methods are prone to causing secondary pollution (such as methane emissions and leachate pollution). Composting technology achieves resource utilization through microbial degradation and has become one of the core solutions in the environmental protection field. Once the composting system malfunctions, it may lead to problems such as interruption of fermentation, material decay, and excessive odor emissions, which not only affect the quality of compost but also pollute the environment. Therefore, it is necessary to monitor the operating status of the equipment, detect potential faults in a timely manner, and issue early warnings to ensure the stable operation of the system.
[0003] In existing technologies, traditional classification models are mostly trained offline, which cannot adapt to dynamic scenarios such as fluctuations in compost raw material composition and environmental temperature interference. Moreover, clustering methods are sensitive to noisy data, and compost sensors are easily affected by dust and humidity, leading to data drift and affecting the clustering effect. The overall model adaptability is poor. Therefore, how to process dynamic noise data through adaptive clustering analysis, identify dynamic changes in data distribution, and distinguish between normal fluctuations and abnormal faults to achieve fault diagnosis and early warning for electrically automated composting processes is the problem that this invention aims to solve. To this end, an intelligent fault diagnosis and early warning method for electrically automated composting process systems is proposed. Summary of the Invention
[0004] The purpose of this invention is to provide an intelligent fault diagnosis and early warning method for an electrically automated composting process system, so as to solve the problems mentioned in the background art.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0006] A method for intelligent fault diagnosis and early warning in an electrically automated composting process system includes the following steps:
[0007] S1. Collect sensor data such as temperature, humidity and oxygen content in the composting process, use sliding window filtering to remove high-frequency noise, and synchronize timestamps to align multi-dimensional signals;
[0008] S2. Based on the Hidden Markov Model (HMM), the composting stages are divided into heating / high temperature / maturation stages, and time-domain and frequency-domain features are extracted for different working conditions.
[0009] S3. The cluster centers are initialized using the online K-means++ algorithm. The intra-cluster distance is calculated in real time based on the new data stream, and the center position is iteratively optimized to adapt to the time-varying characteristics of composting parameters and realize dynamic data classification.
[0010] S4. Density peak clustering is introduced to calculate the local density and minimum distance of data points. Outliers are filtered out by local density and distance thresholds, and effective clusters are retained. This reduces the impact of noisy data on the clustering results and improves the clustering accuracy.
[0011] S5. The clustering results after density peak clustering are assigned to the corresponding cluster centers, cluster labels are generated and concatenated with the original features to construct an enhanced feature vector, which is used as the input feature of the classification model. The XGBoost classification model is then input, and the contribution of the feature to the fault is quantified using the SHAP value.
[0012] S6. Monitor the operation data of the composting system in real time, identify normal fluctuations and abnormal faults through the XGBoost classification model, issue early warning signals and provide feedback to the operators.
[0013] A further improvement to the technical solution of the present invention is that: S1 specifically includes:
[0014] Sensors, including temperature, humidity, oxygen concentration, and equipment current sensors, are arranged in layers inside the composting reactor of the composting system. The sampling frequency is set to 1Hz to ensure coverage of dynamic changes and to collect sensor data during the composting process. The sensor data is aggregated to the edge gateway via the Modbus RTU protocol (RS485 bus). At the same time, CRC check is performed on each data, and if the check fails, a retransmission mechanism is triggered (up to 3 retries). Data that times out is marked as invalid to ensure data integrity.
[0015] A sliding window filter is applied to the time series data of each sensor, with a window length of N=5 (corresponding to 5 seconds of data, covering short-term fluctuations in compost parameters), and boundary processing is performed. For the first N-1 points, linear recursion is used to fill the gaps, and then the signals before and after filtering are compared through spectrum analysis to remove high-frequency noise.
[0016] All sensors are connected to the edge gateway via PTP (Precision Time Protocol). The master clock source uses a GPS timing module to ensure an initial time deviation of <1μs. The edge gateway sends synchronization messages to the sensors every minute, corrects the transmission delay based on the IEEE1588 protocol, and constructs a unified time axis with a resolution of 100ms. Cubic spline interpolation is performed on the unsynchronized data points to finally generate an aligned data matrix.
[0017] A further improvement to the technical solution of the present invention lies in that: the sensors arranged in layers inside the composting reactor are specifically:
[0018] The temperature sensor is a PT100 resistance temperature detector (accuracy ±0.1℃), which is arranged in three layers inside the composting reactor at distances of 0.3m / 1.0m / 1.7m from the top to monitor the vertical temperature gradient.
[0019] The humidity sensor is a capacitive hygrometer (range 0-100%RH), arranged on the same layer as the temperature sensor with a spacing of ≥0.5m to avoid thermal interference;
[0020] The oxygen concentration sensor is an electrochemical sensor (resolution 0.1% Vol), located near the ventilation opening in the middle of the reactor, to reflect the aerobic fermentation efficiency;
[0021] The current sensor of the equipment is a Hall closed-loop sensor (range 0-50A), which is connected in series in the power supply circuit of the turner / ventilator to monitor the operating status of the equipment.
[0022] A further improvement to the technical solution of the present invention is that: S2 specifically includes:
[0023] Collect sensor data (at least 3 complete batches) for historical composting cycles. Based on the typical characteristics of the composting process, use a Hidden Markov Model (HMM) to divide the composting process into stages, namely the heating period, the high temperature period, and the maturation period. Mark the start and end times of the heating period (temperature first exceeds 35℃ to 50℃), the high temperature period (50-70℃ for ≥3 days), and the maturation period (temperature ≤40℃ and fluctuation <2℃ within 24 hours).
[0024] The real-time sensor data is segmented into time windows, and each segment is treated as an observation sequence. The data is then input into a trained Hidden Markov Model (HMM) and the optimal state path is calculated using the Viterbi algorithm. These paths are identified as State 1, State 2, and State 3, where State 1 represents the warming period, State 2 represents the high-temperature period, and State 3 represents the composting period. Moving average filtering is then applied to the decoding results to eliminate short-term fluctuations, and the label and confidence level of the current composting stage are output.
[0025] For different composting stages, time-domain and frequency-domain features are extracted separately. For time-domain features, they are divided into general features and stage-specific features. The general features are the mean, standard deviation, and peak factor. Among the stage-specific features, the temperature rise rate during the heating period reflects the microbial initiation speed, the oxygen concentration fluctuation range during the high-temperature period indicates ventilation efficiency, and the temperature drop slope during the maturation period assesses the maturation stability. For the extraction of frequency-domain features, the temperature / current signals are subjected to FFT transformation to extract the main frequency components and the concentrated spectral energy range. Then, the time-domain and frequency-domain features are spliced together, classified and stored according to stage labels, and a feature matrix is generated.
[0026] A further improvement to the technical solution of the present invention is that: S3 specifically includes:
[0027] Starting from the composting initiation stage, the first 200 sets of sensor data were collected, covering the initial fluctuation characteristics of the warming period. The data were preprocessed and normalized to the [0, 1] interval to eliminate dimensional differences. The first center point was randomly selected, and the Euclidean distance between the remaining data points and the first center point was calculated. The next center point was selected according to the distance-weighted probability. This process was repeated until three centers were selected, corresponding to the warming period, the high-temperature period, and the maturation period, respectively. If the initial center point distribution was uneven, resampling was performed until the minimum distance threshold was met, i.e., the center point spacing was <0.3 normalized units.
[0028] The real-time sensor data is segmented into time windows of 5 minutes in length. Each segment is treated as a new sample. The distance between the new sample and the current three center points is calculated, and the sample is assigned to the nearest cluster. The historical center points of the cluster are assigned a weighted time decay coefficient. The new center points are then updated. If the distance between the new sample and all center points is greater than the threshold of 0.5 normalized units, it is marked as an outlier and does not participate in the center update. The standard deviation of the distance within the cluster is calculated every hour. If the standard deviation of the distance within a cluster suddenly increases by 50%, the center point is reinitialized.
[0029] The samples are assigned stage labels based on their cluster affiliation. The heating stage is defined as a temperature <50℃ and a rate of increase >0.5℃ / h; the high-temperature stage is defined as a temperature of 50-70℃ and an oxygen fluctuation <5%Vol; and the maturation stage is defined as a temperature <40℃ and a 24-hour fluctuation <2℃. A correlation analysis of time-varying characteristics is performed, and the current stage label and intra-cluster distance are output in real time (reflecting the degree of data deviation from the center; a distance >0.3 triggers process adjustment suggestions). The classification results are stored in a time-series database, and a composting cycle stage change curve is generated for subsequent process review.
[0030] A further improvement to the technical solution of this invention lies in the following: the process of correlation analysis of the time-varying features is as follows:
[0031] During the warming period, monitor the rate of temperature rise at the center of the cluster. If the rate decreases by 20% for two consecutive hours, it indicates a decrease in microbial activity.
[0032] During periods of high temperature, monitor the central value of cluster oxygen concentration. If the central value is <15%Vol and continues for 6 hours, issue a warning of insufficient ventilation.
[0033] During the decomposition period, observe the slope of the temperature drop in the cluster. If the slope is less than 0.1℃ / day, the decomposition is considered complete.
[0034] A further improvement to the technical solution of the present invention is that: S4 specifically includes:
[0035] For each data point of the compost sensor, calculate its Euclidean distance to all other data points, count the number of neighbors whose distance is less than the cutoff distance, and use this as the local density. The cutoff distance is set to 10% of the average distance between data points. For each data point, select all data points with higher local density than it, calculate the minimum distance between the selected data point and the data point. If the data point is already the highest density point, the minimum distance is set to the maximum distance with the second highest density point to avoid misjudging isolated points. Then, draw a scatter plot of local density and minimum distance for all data points. Outliers are abnormal values with low local density and high minimum distance.
[0036] Calculate the mean and standard deviation of the local density of all data points, set a low density threshold, and remove data points smaller than the low density threshold. Simultaneously, calculate the mean of the minimum distance among all data points. A high distance threshold is set, and data points exceeding the high distance threshold are removed. Data points that meet both the low density threshold and the high distance threshold removal criteria are marked as outliers and removed from the dataset.
[0037] In the local density and minimum distance scatter plot, select high-density and high-distance data points as cluster centers, and assign the remaining points to the nearest high-density centers. Calculate the distance between the cluster boundary point and the adjacent cluster boundary point. If it is less than the cutoff distance, merge the clusters (to avoid over-segmentation due to noise).
[0038] The composting stage labels are assigned based on the range of cluster center parameters. For newly collected 5-minute window data, the local density and minimum distance are calculated. If the current cluster characteristics are met, a label is assigned; otherwise, it is marked as a transitional state. If three consecutive window data are removed as outliers, a sensor fault alarm is triggered.
[0039] A further improvement to the technical solution of this invention lies in the following: the specific process of cluster center selection and cluster merging is as follows:
[0040] In the local density and minimum distance scatter plot, data points with high density and high distance are manually selected as cluster centers. For data points that are not selected as cluster centers, their density ratio with all cluster centers is calculated, and they are assigned to the cluster corresponding to the center with the largest density ratio. That is, they are given priority to be assigned to the high-density neighborhood. If the density ratio of a data point with all centers is lower than the preset threshold, it is marked as a point to be determined and is not assigned to any cluster for the time being.
[0041] For each cluster, data points with local density lower than the cluster mean but belonging to the cluster are selected as boundary points. If the proportion of boundary points in a cluster exceeds 30%, the local density threshold is lowered to avoid over-segmentation. At the same time, the minimum distance between the boundary points of adjacent clusters is calculated. If the distance is less than the cutoff distance, the two clusters are merged, prioritizing the merging of clusters with small local density differences to avoid mistakenly merging low-density noise clusters into the main cluster. Then, the local density and minimum distance of the undetermined point are recalculated. If the ratio of its density to the center density of a merged cluster is >0.9, it is assigned to that cluster; otherwise, it is marked as an outlier.
[0042] For each cluster, the silhouette coefficient of all points is calculated to evaluate the intra-cluster compactness and inter-cluster separation. The mean of the silhouette coefficients of all points is taken. If the score is <0.5, it indicates that the clustering effect is not good. Then, the cutoff distance is optimized or the local density threshold is adjusted. Specifically, for optimizing the cutoff distance, if the silhouette coefficient is low, the cutoff distance is increased by 5% step size to expand the neighbor range and improve the robustness of density calculation. For adjusting the density threshold, if the silhouette coefficient of a certain cluster is significantly lower than that of other clusters, the local density threshold of that cluster is reduced to absorb more boundary points. When the silhouette coefficient is ≥0.5 and the adjustment range is <0.05 for two consecutive times, the iteration is stopped and the final cluster partitioning result is output.
[0043] A further improvement to the technical solution of the present invention is that S5 specifically includes:
[0044] For each data point after density peak clustering screening, calculate its Euclidean distance to all cluster centers and assign it to the cluster corresponding to the nearest center. If the distance difference between a data point and multiple cluster centers is less than 20% of the cutoff distance, it is marked as a transition point and assigned to a cluster with higher local density to avoid ambiguity in stage division. Then, the cluster label is converted into a numerical feature (heating period = 1, high temperature period = 2, decay period = 3), which is concatenated with the original sensor features to form an enhanced feature vector.
[0045] The enhanced feature vectors constructed based on historical sensor data are divided into a training set (70%) and a test set (30%). The training set is input into the XGBoost classification model, and the objective function is set to multi-class classification. The classification objectives include normal operation, sensor drift, and dust interference. The tree depth, learning rate, and subsample ratio are optimized by grid search. The trained XGBoost classification model is evaluated on the test set, and the accuracy and F1 score are calculated. If the F1 score is lower than 0.8, it indicates that the model performance is not up to standard, and it is necessary to return to adjust the clustering or re-optimize the model parameters.
[0046] For the trained XGBoost classification model, calculate the SHAP value of each feature (reflecting the average influence of the feature on the model output), and draw a global importance ranking map. If the SHAP value of the cluster label ranks in the top 3, it indicates that the division of composting stages is crucial for fault diagnosis. If the temperature standard deviation contributes significantly, then abnormal temperature fluctuations need to be given special attention. For misclassified samples, calculate the SHAP value of each feature, locate the key features that caused the error, and adjust the local density threshold of the cluster center based on the local analysis results. Then, generate a visualization report to show the top 3 features of the fault class.
[0047] A further improvement to the technical solution of the present invention is that: S6 specifically includes:
[0048] The system collects sensor data in real time during the operation of the composting system, obtains key parameters of the composting system, timestamps and synchronizes data from multiple sensors, and then analyzes and obtains enhanced feature vectors as model input.
[0049] Load the pre-trained XGBoost classification model, predict each enhanced feature vector based on real-time sensor data, and output the returned category label and corresponding probability. If an anomaly warning is required, further analyze whether it is a strong or weak warning. If the XGBoost classification model predicts the same fault class 3 times in a row and the probability is >0.8, a strong warning is issued, an alarm is triggered immediately and pushed to the operation terminal. If the probability fluctuates between 0.6 and 0.8, a weak warning is issued, an anomaly log is recorded and marked as pending confirmation for operators to review.
[0050] For the issued early warning signals, the abnormal location is highlighted on the monitoring screen in the composting control room, accompanied by an audible and visual alarm. At the same time, the warning details are pushed to the maintenance personnel via WeChat / SMS. The maintenance personnel can upload on-site photos or videos through the mobile APP, and the system automatically links them to the corresponding warning record. After the maintenance is completed, the operator marks the warning as handled and fills in the handling result, forming a closed-loop management.
[0051] Due to the adoption of the above technical solution, the technical progress achieved by this invention compared to the prior art is as follows:
[0052] 1. This invention provides an intelligent fault diagnosis and early warning method for an electrically automated composting process system. By collecting multi-dimensional sensor data such as temperature, humidity, and oxygen content in the composting process in real time, and using the online K-means++ algorithm and density peak clustering technology, it dynamically adapts to the time-varying characteristics of composting parameters, effectively filters outliers caused by sensor drift or dust interference, significantly reduces the model's sensitivity to noise data, ensures that the composting system can still operate stably under complex working conditions, and reduces the risk of fermentation interruption or material spoilage caused by abnormal parameters.
[0053] 2. This invention provides an intelligent fault diagnosis and early warning method for an electrically automated composting process system. Based on the combination of composting stage division using a hidden Markov model and an XGBoost classification model, it accurately distinguishes between normal fluctuations and abnormal faults. By quantifying the feature contribution through SHAP values, it can locate key fault causes and supports a graded mechanism for strong and weak early warnings, enabling operators to prioritize the handling of high-risk faults. Attached Figure Description
[0054] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0055] Figure 1 This is a schematic diagram of the workflow of the present invention;
[0056] Figure 2 Schematic diagram of the method of the present invention. Detailed Implementation
[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0058] Example 1, as Figure 1 , Figure 2 As shown, this invention provides an intelligent fault diagnosis and early warning method for an electrically automated composting process system, comprising the following steps:
[0059] S1. Collect sensor data such as temperature, humidity, and oxygen content during the composting process. Use a sliding window filter to remove high-frequency noise and synchronize timestamps to align multi-dimensional signals. Arrange sensors in layers inside the composting reactor of the composting system, including temperature sensors, humidity sensors, oxygen concentration sensors, and equipment current sensors. Set the sampling frequency to 1Hz to ensure coverage of dynamic changes, collect sensor data during the composting process, and transmit the data via Modbus. The RTU protocol (RS485 bus) aggregates sensor data to the edge gateway. A circular buffer of 1024 records (each record contains a timestamp, 4 sensor values, and a CRC field) is created in the edge gateway. Head pointers / tail pointers manage read and write operations, temporarily storing unprocessed raw data packets. Simultaneously, a CRC check is performed on each data record; if the check fails, a retransmission mechanism is triggered (maximum 3 retries). Data that times out is marked as invalid, ensuring data integrity. A sliding window filter is applied to the time-series data of each sensor, with a window length of N=5 (corresponding to 5 seconds of data, covering short-term fluctuations in compost parameters), and boundary processing is performed. Linear recursion is used to fill the first N-1 points. Then, spectral analysis is used to compare the signals before and after filtering, removing high-frequency noise (>0.2Hz) with an amplitude reduction of ≥15dB, while retaining low-frequency effective components (<0.05Hz). All sensors are connected to the edge gateway via PTP (Precision Time Protocol). The master clock source uses a GPS timing module to ensure an initial time deviation of <1μs. The edge gateway sends synchronization messages to the sensors every minute, based on IEEE... The 1588 protocol corrects transmission delays and constructs a unified time axis with a resolution of 100ms. Cubic spline interpolation is performed on unsynchronized data points to finally generate an aligned data matrix.
[0060] The sensors arranged in layers inside the composting reactor are as follows:
[0061] The temperature sensor is a PT100 resistance temperature detector (RTD) with an accuracy of ±0.1℃, arranged in three layers inside the compost reactor at distances of 0.3m, 1.0m, and 1.7m from the top, respectively, to monitor the vertical temperature gradient. The humidity sensor is a capacitive hygrometer (range 0-100%RH), arranged in the same layer as the temperature sensor with a spacing of ≥0.5m to avoid thermal interference. The oxygen concentration sensor is an electrochemical sensor (resolution 0.1%Vol), located near the ventilation opening in the middle of the reactor, to reflect the aerobic fermentation efficiency. The equipment current sensor is a Hall effect closed-loop sensor (range 0-50A), connected in series in the power supply circuit of the turner / ventilator to monitor the equipment operating status.
[0062] The formula for calculating sliding window filtering is as follows:
[0063] ;
[0064] In the formula, The output variable is the output value after moving average filtering, representing the value at time point [time]. At that time, the smoothed result of the sensor data reflects both the current moment and the past. The average value of the data at each time point is used to remove high-frequency noise. The input variable is the raw time-series data collected by the sensor, representing the time point. At that time, the measurement value of a certain sensor, , where is the window length, representing the size of the sliding window, i.e., the number of consecutive data points involved in the averaging calculation. The larger the value, the stronger the smoothing effect, but it may lose rapidly changing features. The smaller the value, the more dynamic details are preserved, but the noise reduction capability is weakened. This is a time index, representing the current point in time. At that time, the output of the moving average filter The calculation steps are as follows: Select a window, starting from the current time... Backtracking At each time point, the data sequence is obtained. Calculate the sum of all data within the window. Divide the sum by the window size. , obtain smoothed value ;
[0065] S2. Based on a Hidden Markov Model (HMM), the composting process is divided into three stages: heating, high temperature, and maturation. Time-domain and frequency-domain features are extracted for different operating conditions. Sensor data from historical composting cycles (at least three complete batches) are collected. According to the typical characteristics of the composting process, the HMM is used to divide the composting process into three stages: heating, high temperature, and maturation. The start and end timestamps of the heating stage (temperature first exceeding 35℃ to 50℃), the high temperature stage (50-70℃ for ≥3 days), and the maturation stage (temperature ≤40℃ with fluctuations <2℃ within 24 hours) are marked. Observational variables are extracted for each stage: temperature, oxygen concentration, and turning mechanism. The flow is defined with three hidden states corresponding to three composting stages. The initial state probability distribution is set to uniform distribution. A Gaussian mixture model (GMM) is used to fit the observation probability matrix. Each hidden state corresponds to two Gaussian components covering normal fluctuations and abnormal disturbances. The Baum-Welch algorithm is then used to iteratively optimize the transition probability matrix and the observation probability matrix, with a maximum of 100 iterations and a convergence threshold of 1e-4. The phase division accuracy is tested on the validation set using Viterbi decoding, with a target of ≥90%. If the target is not met, the number of Gaussian components is adjusted or training data is increased. The real-time sensor data is segmented according to time windows, and each segment is used as an observation sequence. Input the trained Hidden Markov Model and use the Viterbi algorithm to decode and calculate the optimal state path. These are states 1, 2, and 3, with state 1 representing the warming period. When triggered, state 2 is the high-temperature period, which must be met for 3 consecutive hours. State 3 is the ripening stage. The process continues for 24 hours. Moving average filtering is then applied to the decoding results to eliminate short-term jitter, outputting the current composting stage label L∈{1,2,3} (1=heating period, 2=high temperature period, 3=maturation period) and confidence level. For different composting stages, time-domain and frequency-domain features are extracted. Time-domain features are divided into general features and stage-specific features. General features include mean, standard deviation, and peak factor. Stage-specific features include the temperature rise rate during the heating period (reflecting microbial initiation speed), the oxygen concentration fluctuation range during the high temperature period (indicating ventilation efficiency), and the temperature drop slope during the maturation period (assessing maturation stability). For frequency-domain feature extraction, FFT transformation is performed on the temperature / current signals to extract the dominant frequency component and the concentrated spectral energy range. For example, a turning machine fault warning occurs if the dominant frequency of the current signal shifts to 50Hz (power frequency interference) during the high temperature period, and ventilation blockage detection occurs if the oxygen signal spectral energy is concentrated in <0.1Hz during the maturation period, indicating a slow response of the ventilation system. The time-domain and frequency-domain features are then concatenated, categorized and stored according to stage labels, generating a feature matrix.
[0066] S3. Initialize cluster centers using the online K-means++ algorithm. Calculate intra-cluster distances in real-time based on new data streams, iteratively optimize center positions, adapt to the time-varying characteristics of composting parameters, and achieve dynamic data classification. Starting from the composting initiation stage, collect the first 200 sets of sensor data, covering the initial fluctuation characteristics of the warming period. Preprocess the data, normalize it to the [0, 1] interval, eliminate dimensional differences, randomly select the first center point, calculate the Euclidean distance between the remaining data points and the first center point, and select the next center point according to distance-weighted probability. Repeat this process until three centers are selected, corresponding to the warming period, high-temperature period, and maturation period, respectively. If the initial center point distribution is uneven, resample until the minimum distance threshold is met, i.e., the center point spacing < 0.3 normalized units. Divide the real-time collected sensor data into segments with a time window of 5 minutes in length, and treat each segment as a new sample. Calculate the distance between the new sample and the current 3... The distance to each centroid is used to assign it to the nearest cluster, and a weighted time decay coefficient is assigned to the historical centroids of the clusters. The new centroids are then updated. If the distance between a new sample and all centroids is greater than the threshold of 0.5 normalized units, it is marked as an outlier and does not participate in the centroid update. The standard deviation of the distance within a cluster is calculated every hour. If the standard deviation of the distance within a cluster suddenly increases by 50%, centroid reinitialization is triggered. Stage labels are assigned according to the cluster to which the sample belongs. The heating period is defined as temperature <50℃ and rising rate >0.5℃ / h, the high temperature period is defined as temperature 50-70℃ and oxygen fluctuation <5%Vol, and the maturation period is defined as temperature <40℃ and 24-hour fluctuation <2℃. Correlation analysis of time-varying features is performed, and the current stage label and the distance within the cluster are output in real time (reflecting the degree of data deviation from the center; when >0.3, process adjustment suggestions are triggered). The classification results are stored in a time-series database, and a composting cycle stage change curve is generated for subsequent process review.
[0067] The process of correlation analysis of time-varying features is as follows:
[0068] During the warming period, monitor the rate of temperature rise at the center of the cluster. If the rate decreases by 20% for 2 consecutive hours, it indicates reduced microbial activity. During the high-temperature period, track the center value of oxygen concentration in the cluster. If the center value is <15%Vol and lasts for 6 hours, it is an early warning of insufficient ventilation. During the decomposition period, observe the slope of temperature drop in the cluster. If the slope is <0.1℃ / day, decomposition is considered complete.
[0069] The formula for updating the new center point is as follows:
[0070] ;
[0071] In the formula, For the first Clusters in time The updated center point, i.e., the coordinates of the cluster center, at that time. The time decay coefficient controls the historical center point. Contribution percentage to the current update The larger the value, the slower the center point updates (it relies more on historical data). The smaller the value, the more sensitive the center point is to updates (better adapted to new data). For the first Clusters in time The historical center point of time, that is, the cluster center of the previous moment. For the first Clusters in time The number of samples at any given time, i.e., the number of data points currently assigned to this cluster, is used to calculate the mean of the samples within the cluster, avoiding center point shifts due to different cluster sizes. For the first All samples within a cluster Summing the coordinates (adding each dimension). For the first The sample mean of each cluster, i.e., the centroid update formula of the traditional K-means, is used to calculate the geometric center of all samples within the current cluster. This represents the contribution of the current cluster sample mean to the center point update.
[0072] S4. Density peak clustering is introduced to calculate the local density and minimum distance of data points. Outliers caused by sensor drift or dust interference are filtered out by local density and distance thresholds, retaining effective clusters, reducing the impact of noise data on clustering results, and improving clustering accuracy.
[0073] S5. The clustering results after density peak clustering are assigned to the corresponding cluster centers, cluster labels are generated and concatenated with the original features to construct an enhanced feature vector, which is used as the input feature of the classification model. The XGBoost classification model is then input, and the contribution of the feature to the fault is quantified using the SHAP value.
[0074] S6. Monitor the operation data of the composting system in real time, identify normal fluctuations and abnormal faults through the XGBoost classification model, issue early warning signals and provide feedback to the operators.
[0075] Example 2, as Figure 1 , Figure 2 As shown, based on Embodiment 1, the present invention provides a technical solution: preferably, S4 specifically includes:
[0076] For each data point of the compost sensor, calculate its Euclidean distance to all other data points. Count the number of neighbors whose distance is less than the cutoff distance, and use this as the local density. The cutoff distance is set to 10% of the average distance between data points. For each data point, select all data points with higher local densities. Calculate the minimum distance between this data point and the selected data points. If the data point is already the highest density point, the minimum distance is set to the maximum distance to the second highest density point to avoid misjudging isolated points. Then, draw a scatter plot of the local density and minimum distance of all data points. Outliers are characterized by low local density and high minimum distance. Calculate the mean of the local density of all data points. with standard deviation And set a low density threshold. Data points smaller than the low density threshold are removed. Simultaneously, the mean of the minimum distances among all data points is calculated. Set a high distance threshold Data points exceeding the high distance threshold are removed. Data points that simultaneously meet both the low density threshold and the high distance threshold removal criteria are marked as outliers and removed from the dataset. High density is selected from the local density and minimum distance scatter plot. And high distance ( The data points are used as cluster centers, and the remaining points are assigned to the nearest high-density centers. The distance between the cluster boundary point and the adjacent cluster boundary point is calculated. If it is less than the cutoff distance, the clusters are merged (to avoid over-segmentation due to noise). The composting stage labels are assigned according to the cluster center parameter range. For the newly collected 5-minute window data, its local density and minimum distance are calculated. If it meets the current cluster characteristics, a label is assigned. Otherwise, it is marked as a transitional state. If three consecutive window data are removed as outliers, a sensor fault alarm is triggered.
[0077] Furthermore, the specific process for cluster center selection and cluster merging is as follows:
[0078] In the local density and minimum distance scatter plot, high-density data points with high distances are manually selected as cluster centers. For data points not selected as cluster centers, their density ratios with all cluster centers are calculated, and they are assigned to the cluster corresponding to the center with the highest density ratio, i.e., preferentially assigned to a high-density neighborhood. If the density ratio of a data point with all centers is lower than a preset threshold, it is marked as a pending point and not assigned to any cluster. For each cluster, data points with local densities lower than the cluster mean but belonging to that cluster are selected as boundary points. If the proportion of boundary points in a cluster exceeds 30%, the local density threshold is lowered to avoid over-segmentation. At the same time, the minimum distance between adjacent cluster boundary points is calculated. If the distance is less than the cutoff distance, the two clusters are merged, prioritizing the merging of clusters with small local density differences to avoid mistakenly merging low-density, noisy clusters into the main cluster. Recalculate the local density and minimum distance of the undetermined point. If the ratio of its density to the center density of a merged cluster is >0.9, it is assigned to that cluster; otherwise, it is marked as an outlier. For each cluster, calculate the silhouette coefficient of all points to evaluate the intra-cluster compactness and inter-cluster separation. Take the mean of the silhouette coefficients of all points. If the score is <0.5, it indicates poor clustering performance. Then, optimize the cutoff distance or adjust the local density threshold. For optimizing the cutoff distance, if the silhouette coefficient is low, increase the cutoff distance by 5% step size to expand the neighbor range and improve the robustness of density calculation. For adjusting the density threshold, if the silhouette coefficient of a cluster is significantly lower than that of other clusters, decrease the local density threshold of that cluster to absorb more boundary points. When the silhouette coefficient is ≥0.5 and the adjustment range is <0.05 for two consecutive times, stop the iteration and output the final cluster partitioning result.
[0079] S5 specifically includes:
[0080] For each data point after density peak clustering screening, its Euclidean distance to all cluster centers is calculated, and it is assigned to the cluster corresponding to the nearest center. If the distance difference between a data point and multiple cluster centers is less than 20% of the cutoff distance, it is marked as a transition point and assigned to a cluster with higher local density to avoid ambiguity in stage division. Then, the cluster labels are converted into numerical features (warming period = 1, high temperature period = 2, decay period = 3), which are concatenated with the original sensor features to form an enhanced feature vector. The enhanced feature vector constructed based on historical sensor data is divided into a training set (70%) and a test set (30%). The training set is input into the XGBoost classification model, and the objective function is set to multi-class classification. The classification objectives include normal operation, sensor drift, and dust interference. The tree depth (3-8), learning rate (0.01-0.2), and subsample ratio (0.6-1.0) are optimized through grid search. During the optimization process, the model complexity and generalization ability are balanced. To avoid overfitting or underfitting, ensure the model performs well on the training set and has high generalization ability on the test set. Evaluate the trained XGBoost classification model on the test set, calculate accuracy and F1 score. If the F1 score is below 0.8, the model performance is substandard and needs to be adjusted by re-optimizing the clustering or re-optimizing the model parameters. For the trained XGBoost classification model, calculate the SHAP value of each feature (reflecting the average influence of the feature on the model output) and draw a global importance ranking chart. If the SHAP value of the cluster label ranks in the top 3, it indicates that the composting stage division is crucial for fault diagnosis. If the temperature standard deviation contributes significantly, then abnormal temperature fluctuations need to be focused on. For misclassified samples, calculate the SHAP value of each feature, locate the key features that caused the error, and adjust the local density threshold of the cluster center based on the local analysis results. Then, generate a visualization report showing the top 3 features of the fault class.
[0081] S6 specifically includes:
[0082] The system collects sensor data in real time during the operation of the composting system to obtain key parameters. It synchronizes data from multiple sensors using timestamps to ensure parameter matching at the same moment. If a sensor's data is missing for three consecutive minutes, it is filled with the average of the previous three minutes to avoid abnormal interruptions. If the missing data exceeds 10 minutes, a sensor health check is triggered, and enhanced feature vectors are obtained. These enhanced feature vectors are then used as input to a pre-trained XGBoost classification model. The model predicts each enhanced feature vector based on real-time sensor data, outputting the class label and corresponding probability. If an anomaly warning is needed, further analysis is performed to determine whether a strong or weak warning is required. If the XGBoost classification model predicts the same fault class three times consecutively with a probability > 0.8, a strong warning is issued. The system immediately triggers an alarm and pushes it to the operation terminal. If the probability fluctuates between 0.6 and 0.8, a weak warning is issued, the anomaly log is recorded and marked as pending confirmation for operator review. In addition, the classification threshold is dynamically adjusted based on the historical false alarm rate. If the false alarm rate of a certain category is >10%, its probability threshold is reduced at intervals of 0.05 until the false alarm rate is <5%. For the issued warning signal, the abnormal location is highlighted on the monitoring screen in the composting control room, accompanied by an audible and visual alarm. At the same time, the warning details are pushed to the maintenance personnel via WeChat / SMS. The maintenance personnel upload on-site photos or videos through the mobile APP, and the system automatically links them to the corresponding warning record. After the maintenance is completed, the operator marks the warning as handled and fills in the handling result, forming a closed-loop management.
[0083] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for intelligent fault diagnosis and early warning in an electrically automated composting process system, characterized in that, Includes the following steps: S1. Collect sensor data in the composting process, use sliding window filtering to remove high-frequency noise, and synchronize timestamps to align multi-dimensional signals; S2. Based on the Hidden Markov Model, the composting stages are divided into heating / high temperature / maturation stages, and time-domain and frequency-domain features are extracted for different working conditions. S3. The cluster centers are initialized using the online K-means++ algorithm. The intra-cluster distance is calculated in real time based on the new data stream, and the center position is iteratively optimized to adapt to the time-varying characteristics of composting parameters. S4. Density peak clustering is introduced to calculate the local density and minimum distance of data points. Outliers are removed by threshold filtering, and effective clusters are retained. S5. Assign the clustering results after density peak clustering to the corresponding cluster centers, construct enhanced feature vectors, input them into the XGBoost classification model, and use SHAP values to quantify the contribution of features to the fault. S6. Monitor the operation data of the composting system in real time, identify normal fluctuations and abnormal faults through the XGBoost classification model, issue early warning signals and provide feedback to the operators.
2. The intelligent fault diagnosis and early warning method for an electrically automated composting process system according to claim 1, characterized in that: S1 specifically includes: Sensors, including temperature sensors, humidity sensors, oxygen concentration sensors, and equipment current sensors, are arranged in layers inside the composting reactor of the composting system to collect sensor data in the composting process. The sensor data is aggregated to the edge gateway via the Modbus RTU protocol. At the same time, CRC check is performed on each data. If the check fails, a retransmission mechanism is triggered. If the timeout occurs, the data is marked as invalid. A sliding window filter is applied to the time series data of each sensor, with a window length of N=5, and boundary processing is performed. For the first N-1 points, linear recursion filling is used, and then the signals before and after filtering are compared through spectrum analysis to remove high-frequency noise. All sensors are connected to the edge gateway via PTP. The master clock source uses a GPS timing module. The edge gateway sends synchronization messages to the sensors every minute, corrects transmission delays based on the IEEE 1588 protocol, constructs a unified time axis, performs cubic spline interpolation on unsynchronized data points, and finally generates an aligned data matrix.
3. The intelligent fault diagnosis and early warning method for an electrically automated composting process system according to claim 2, characterized in that: The sensors arranged in layers inside the composting reactor are specifically as follows: The temperature sensor is a PT100 resistance temperature detector (RTD), which is arranged in three layers inside the composting reactor at distances of 0.3m, 1.0m, and 1.7m from the top to monitor the vertical temperature gradient. The humidity sensor is a capacitive hygrometer, arranged on the same layer as the temperature sensor, with a spacing of ≥0.
5. The oxygen concentration sensor is an electrochemical sensor and is located near the ventilation opening in the middle of the reactor. The current sensor of the equipment is a Hall closed-loop sensor, connected in series in the power supply circuit of the turner / ventilator to monitor the operating status of the equipment.
4. The intelligent fault diagnosis and early warning method for an electrically automated composting process system according to claim 1, characterized in that: S2 specifically includes: Collect sensor data from historical composting cycles. Based on the typical characteristics of the composting process, use a hidden Markov model to divide the composting process into stages, namely the heating period, the high temperature period, and the maturation period, and mark the start and end timestamps of the heating period, the high temperature period, and the maturation period. The real-time sensor data is segmented into time windows, and each segment is used as an observation sequence. The data is then input into a trained Hidden Markov Model and the optimal state path is calculated using the Viterbi algorithm. These paths are state 1, state 2, and state 3, where state 1 is the warming period, state 2 is the high-temperature period, and state 3 is the composting period. The decoding results are then filtered using a moving average to output the label and confidence level of the current composting stage. For different composting stages, time-domain and frequency-domain features are extracted separately. For time-domain features, they are divided into general features and stage-specific features. The general features are the mean, standard deviation, and peak factor. For stage-specific features, the temperature rise rate is used during the heating period, the oxygen concentration fluctuation range is used during the high-temperature period, and the temperature drop slope is used during the maturation period. For frequency-domain feature extraction, the temperature / current signal is subjected to FFT transformation to extract the main frequency component and the concentrated range of spectral energy. Then, the time-domain and frequency-domain features are spliced together, classified and stored according to stage labels, and a feature matrix is generated.
5. The intelligent fault diagnosis and early warning method for an electrically automated composting process system according to claim 4, characterized in that: S3 specifically includes: Starting from the composting initiation stage, the first 200 sets of sensor data were collected, covering the initial fluctuation characteristics of the warming period. The data were preprocessed and normalized to the [0, 1] interval. The first center point was randomly selected, and the Euclidean distance between the remaining data points and the first center point was calculated. The next center point was selected according to the distance-weighted probability. This process was repeated until three centers were selected, corresponding to the warming period, the high-temperature period, and the maturation period, respectively. If the initial center point distribution was uneven, resampling was performed until the minimum distance threshold was met. The real-time sensor data is segmented into time windows of 5 minutes in length. Each segment is treated as a new sample. The distance between the new sample and the current three center points is calculated, and the sample is assigned to the nearest cluster. The historical center points of the cluster are assigned a weighted time decay coefficient. The new center points are then updated. If the distance between the new sample and all center points is greater than the threshold of 0.5 normalized units, it is marked as an outlier and does not participate in the center update. The standard deviation of the distance within the cluster is calculated every hour. If the standard deviation of the distance within a cluster suddenly increases by 50%, the center point is reinitialized. The samples are assigned stage labels based on their cluster affiliation. The warming stage is defined as a temperature <50℃ and a rate of increase >0.5℃ / h, the high-temperature stage as a temperature of 50-70℃ and an oxygen fluctuation <5%Vol, and the maturation stage as a temperature <40℃ and a 24-hour fluctuation <2℃. The system performs correlation analysis of time-varying characteristics, outputs the current stage label and intra-cluster distance in real time, stores the classification results in a time-series database, and generates a composting cycle stage change curve.
6. The intelligent fault diagnosis and early warning method for an electrically automated composting process system according to claim 5, characterized in that: The process of correlation analysis of the time-varying features is as follows: During the warming period, monitor the rate of temperature rise at the center of the cluster. If the rate decreases by 20% for two consecutive hours, it indicates a decrease in microbial activity. During periods of high temperature, monitor the central value of cluster oxygen concentration. If the central value is <15%Vol and continues for 6 hours, issue a warning of insufficient ventilation. During the decomposition period, observe the slope of the temperature drop in the cluster. If the slope is less than 0.1℃ / day, the decomposition is considered complete.
7. The intelligent fault diagnosis and early warning method for an electrically automated composting process system according to claim 6, characterized in that: S4 specifically includes: For each data point of the compost sensor, calculate its Euclidean distance to all other data points, count the number of neighbors whose distance is less than the cutoff distance, and use this as the local density. The cutoff distance is set to 10% of the average distance between data points. For each data point, select all data points with higher local density than it, calculate the minimum distance between the selected data point and the data point. If the data point is already the highest density point, the minimum distance is set to the maximum distance with the second highest density point. Then, draw a scatter plot of the local density and minimum distance of all data points. Calculate the mean and standard deviation of the local density of all data points, and set a low density threshold to remove data points smaller than the low density threshold. At the same time, calculate the mean of the minimum distance of all data points, set a high distance threshold to remove data points larger than the high distance threshold. Data points that meet both the low density threshold and the high distance threshold removal criteria are marked as outliers and removed from the dataset. In the local density and minimum distance scatter plot, select high-density and high-distance data points as cluster centers, assign the remaining points to the nearest high-density centers, calculate the distance between cluster boundary points and adjacent cluster boundary points, and merge clusters if the distance is less than the cutoff distance; The composting stage labels are assigned based on the range of cluster center parameters. For newly collected 5-minute window data, the local density and minimum distance are calculated. If the current cluster characteristics are met, a label is assigned; otherwise, it is marked as a transitional state. If three consecutive window data are removed as outliers, a sensor fault alarm is triggered.
8. The intelligent fault diagnosis and early warning method for an electrically automated composting process system according to claim 7, characterized in that: The specific process for selecting cluster centers and merging clusters is as follows: In the local density and minimum distance scatter plot, data points with high density and high distance are manually selected as cluster centers. For data points that are not selected as cluster centers, their density ratio with all cluster centers is calculated, and they are assigned to the cluster corresponding to the center with the largest density ratio. That is, they are given priority to be assigned to the high-density neighborhood. If the density ratio of a data point with all centers is lower than the preset threshold, it is marked as a point to be determined and is not assigned to any cluster for the time being. For each cluster, data points whose local density is lower than the cluster mean but belong to the cluster are selected as boundary points. If the proportion of boundary points in a cluster exceeds 30%, the local density threshold is lowered. At the same time, the minimum distance between the boundary points of adjacent clusters is calculated. If the distance is less than the cutoff distance, the two clusters are merged, prioritizing the merging of clusters with small local density differences. Then, the local density and minimum distance of the undetermined point are recalculated. If the ratio of its density to the center density of a merged cluster is >0.9, it is assigned to that cluster; otherwise, it is marked as an outlier. For each cluster, calculate the silhouette coefficient of all points to evaluate the intra-cluster compactness and inter-cluster separation. Take the mean of the silhouette coefficients of all points. If the score is <0.5, it indicates that the clustering effect is not good. Then optimize the cutoff distance or adjust the local density threshold. Specifically, for optimizing the cutoff distance, if the silhouette coefficient is low, increase the cutoff distance by a step size of 5%. For adjusting the density threshold, if the silhouette coefficient of a certain cluster is significantly lower than that of other clusters, decrease the local density threshold of that cluster. When the silhouette coefficient is ≥0.5 and the adjustment range is <0.05 for two consecutive times, stop the iteration and output the final cluster partitioning result.
9. The intelligent fault diagnosis and early warning method for an electrically automated composting process system according to claim 8, characterized in that: S5 specifically includes: For each data point after density peak clustering screening, calculate its Euclidean distance to all cluster centers and assign it to the cluster corresponding to the nearest center. If the distance difference between a data point and multiple cluster centers is less than 20% of the cutoff distance, it is marked as a transition point and assigned to a cluster with higher local density. Then, the cluster label is converted into a numerical feature and concatenated with the original sensor features to form an enhanced feature vector. The enhanced feature vectors constructed based on historical sensor data are divided into training and test sets. The training set is input into the XGBoost classification model, and the objective function is set to multi-class classification. The classification objectives include normal operation, sensor drift, and dust interference. The tree depth, learning rate, and subsample ratio are optimized through grid search. The trained XGBoost classification model is evaluated on the test set, and the accuracy and F1 score are calculated. If the F1 score is lower than 0.8, it indicates that the model performance is not up to standard, and it is necessary to go back to adjust the clustering or re-optimize the model parameters. For the trained XGBoost classification model, calculate the SHAP value of each feature and draw a global importance ranking map. If the SHAP value of the cluster label ranks in the top 3, it indicates that the division of composting stages is crucial for fault diagnosis. If the temperature standard deviation contributes significantly, then abnormal temperature fluctuations need to be given special attention. For misclassified samples, calculate the SHAP value of each feature, locate the key features that cause the error, and adjust the local density threshold of the cluster center based on the local analysis results. Then, generate a visualization report to show the top 3 features of the fault class.
10. The intelligent fault diagnosis and early warning method for an electrically automated composting process system according to claim 9, characterized in that: S6 specifically includes: The system collects sensor data in real time during the operation of the composting system, obtains key parameters of the composting system, timestamps and synchronizes data from multiple sensors, and then analyzes and obtains enhanced feature vectors as model input. Load the pre-trained XGBoost classification model, predict each enhanced feature vector based on real-time sensor data, and output the returned category label and corresponding probability. If an anomaly warning is required, further analyze whether it is a strong or weak warning. If the XGBoost classification model predicts the same fault class 3 times in a row and the probability is >0.8, a strong warning is issued, an alarm is triggered immediately and pushed to the operation terminal. If the probability fluctuates between 0.6 and 0.8, a weak warning is issued, an anomaly log is recorded and marked as pending confirmation for operators to review. For the issued early warning signals, the abnormal location is highlighted on the monitoring screen in the composting control room, accompanied by an audible and visual alarm. At the same time, the warning details are pushed to the maintenance personnel via WeChat / SMS. The maintenance personnel can upload on-site photos or videos through the mobile APP, and the system automatically links them to the corresponding warning record. After the maintenance is completed, the operator marks the warning as handled and fills in the handling result, forming a closed-loop management.
Citation Information
Cited By
Cherry planting growth environment data management platform
CN121544417A
Measuring and testing method for intelligent low-voltage distribution box line
CN121741387A
Photovoltaic fault diagnosis method and system based on multi-mode adaptive weighting
CN122173886A