Vehicle-mounted mobile pump station abnormal data detection method and system
By clustering and quantifying the anomaly probability of real-time vibration and pressure fluctuation data of vehicle-mounted mobile pump stations, and dynamically adjusting the model, the problems of low anomaly detection accuracy and poor fault early warning reliability of vehicle-mounted mobile pump stations in complex environments are solved, and accurate anomaly identification and rapid fault early warning are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG SANYUAN IND CONTROL AUTOMATION CO LTD
- Filing Date
- 2026-04-14
- Publication Date
- 2026-07-07
Smart Images

Figure CN122346701A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial data processing technology, and in particular to a method and system for detecting abnormal data from a vehicle-mounted mobile pumping station. Background Technology
[0002] Currently, vehicle-mounted mobile pumping stations are important equipment in emergency rescue and industrial production, and their operational stability is directly related to the success or failure of critical tasks.
[0003] In one existing technology, vibration parameters and pressure fluctuation data are mainly collected by sensors. A fixed-mode signal processing scheme is used for filtering and threshold judgment. When the monitored value exceeds the preset threshold, a fault alarm is triggered. However, the operating environment of vehicle-mounted mobile pump stations is complex and variable. The fluctuations in the equipment's own state and external environmental interference are frequently intertwined. The data collected by the sensors has significant time-varying characteristics and multi-source heterogeneity. Because the existing technology relies on fixed-mode processing, it lacks the ability to dynamically model and adaptively analyze industrial data, making it difficult to accurately distinguish between internal equipment anomalies and external disturbances. This leads to frequent misjudgments and missed judgments, and the source of anomalies is unclear.
[0004] Therefore, existing technologies suffer from low accuracy in anomaly detection and poor reliability in fault warning in complex environments due to insufficient industrial data processing capabilities. Summary of the Invention
[0005] This invention provides a method and system for detecting abnormal data in vehicle-mounted mobile pumping stations, in order to solve the problems in the prior art where the detection accuracy of abnormalities in vehicle-mounted mobile pumping stations in complex environments is low and the reliability of fault early warning is poor due to insufficient industrial data processing capabilities.
[0006] Firstly, in order to solve the above-mentioned technical problems, the present invention provides a method for detecting abnormal data in a vehicle-mounted mobile pumping station, comprising: The real-time vibration sequence and pressure fluctuation sequence of the vehicle-mounted mobile pump station are acquired, and the real-time vibration sequence and pressure fluctuation sequence are preprocessed to obtain a preliminary feature vector. Clustering is performed based on the preliminary feature vectors to generate grouped data clusters; Anomaly probability scores are obtained by performing probability mapping on the grouped data clusters using a pre-trained classifier. If the anomaly probability score exceeds a preset score threshold, the contribution of each feature in the corresponding classification result is extracted, and the features are sorted in descending order of contribution to obtain a high-significance feature sequence. The high-significance feature sequence is then matched with a preset mapping dictionary to obtain a preliminary anomaly label. Calculate the confidence level of the preliminary anomaly label. If the confidence level exceeds a preset confidence threshold, refine the preliminary anomaly label into anomaly cause labels according to a preset level, and use them as refined anomaly source labels. Calculate the feature distribution density of the refined anomaly source tag, obtain the real-time data stream and compare it with the feature distribution density to obtain the feedback deviation value. If the feedback deviation value exceeds the preset deviation threshold, adjust the grouping parameters in the clustering process and re-cluster the initial feature vector to obtain optimized grouped data clusters. The optimized grouped data clusters are used to score the new real-time data stream for anomalies and obtain the attribution probability. If the attribution probability exceeds a preset warning threshold, a fault warning signal is generated.
[0007] Secondly, the present invention provides an abnormal data detection system for vehicle-mounted mobile pumping stations, comprising: The feature extraction module is used to acquire the real-time vibration sequence and pressure fluctuation sequence of the vehicle-mounted mobile pump station, and to preprocess the real-time vibration sequence and pressure fluctuation sequence to obtain a preliminary feature vector. The clustering and grouping module is used to perform clustering processing based on the preliminary feature vectors to generate grouped data clusters; An anomaly detection module is used to perform probability mapping on the grouped data clusters using a pre-trained classifier to obtain anomaly probability scores; The source initial judgment module is used to extract the contribution of each feature in the corresponding classification result if the anomaly probability score exceeds a preset score threshold, sort them in descending order of contribution to obtain a high significance feature sequence, and match the high significance feature sequence with a preset mapping dictionary to obtain a preliminary anomaly label. The source refinement module is used to calculate the confidence level of the preliminary anomaly label. If the confidence level exceeds a preset confidence threshold, the preliminary anomaly label is refined into anomaly cause labels according to a preset level, which are used as refined anomaly source labels. The model optimization module is used to calculate the feature distribution density of the refined anomaly source label, obtain the real-time data stream and compare it with the feature distribution density to obtain the feedback deviation value. If the feedback deviation value exceeds the preset deviation threshold, the grouping parameters in the clustering process are adjusted and the initial feature vector is re-clustered to obtain the optimized grouped data cluster. The early warning triggering module is used to score the new real-time data stream for anomalies using the optimized grouped data clusters to obtain the attribution probability. If the attribution probability exceeds the preset early warning threshold, a fault early warning signal is generated.
[0008] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention uses a clustering algorithm to group and aggregate the initial feature vectors, calculates the convergence value within the cluster to distinguish between stable operating condition clusters and transient disturbance clusters, and establishes a state recognition library that can accurately identify the state of the equipment itself and the interference of the external environment. This avoids the limitation of traditional fixed threshold methods that are difficult to adapt to complex operating conditions and solves the problem of misjudgment and omission caused by the inability to distinguish between internal equipment abnormalities and external disturbances in the prior art.
[0009] (2) This invention uses support vector machine to quantify the anomaly probability of grouped data clusters, and maps the geometric distance to an intuitive anomaly probability score, so as to realize the accurate quantitative assessment of the degree of anomaly; then, combined with the random forest algorithm to analyze the feature contribution, a highly significant feature sequence is generated and matched with the preset mapping dictionary to quickly locate the source of anomaly, which solves the problem of unclear anomaly source location and difficulty in implementing targeted maintenance in the prior art.
[0010] (3) This invention cross-validates and assesses the confidence of preliminary labels by integrating historical environmental factor data, and uses the fusion matrix to quantify the correlation between environmental changes and the occurrence of anomalies, effectively eliminating misjudgments caused by environmental interference; then, combined with the classification mapping table, the high-confidence labels are refined hierarchically in terms of physical system, component location, and inducing mechanism, accurately identifying the specific causes of anomalies, thus solving the problem that the sources of anomalies cannot be refined and classified in the prior art.
[0011] (4) This invention dynamically calculates the feedback deviation between the feature distribution density and the real-time data stream through a real-time feedback mechanism, adaptively adjusts the distance metric function and grouping parameters of the grouping model, so that the detection model can track the dynamic changes in the operating status of the equipment; then, based on the optimization model, it performs anomaly quantification analysis on the real-time data stream and triggers blocking logic, thereby realizing full-link closed-loop control from anomaly detection to real-time early warning, which solves the problems of poor model adaptability and insufficient reliability of fault early warning in the prior art. Attached Figure Description
[0012] Figure 1 This is a schematic flowchart of an abnormal data detection method for a vehicle-mounted mobile pumping station provided in the first embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of an abnormal data detection system for a vehicle-mounted mobile pumping station provided in the second embodiment of the present invention. Detailed Implementation
[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] Reference Figure 1 The first embodiment of the present invention provides a method for detecting abnormal data of a vehicle-mounted mobile pumping station, comprising the following steps: S11, acquire the real-time vibration sequence and pressure fluctuation sequence of the vehicle-mounted mobile pump station, preprocess the real-time vibration sequence and pressure fluctuation sequence to obtain a preliminary feature vector; S12, perform clustering processing based on the preliminary feature vectors to generate grouped data clusters; S13, perform probability mapping on the grouped data clusters using a pre-trained classifier to obtain anomaly probability scores; S14, if the anomaly probability score exceeds a preset score threshold, the contribution of each feature in the corresponding classification result is extracted, and the highly significant feature sequence is obtained by arranging them in descending order of contribution. The highly significant feature sequence is then matched with a preset mapping dictionary to obtain a preliminary anomaly label. S15, calculate the confidence level of the preliminary anomaly label. If the confidence level exceeds a preset confidence threshold, refine the preliminary anomaly label into anomaly cause labels according to a preset level, and use them as refined anomaly source labels. S16, calculate the feature distribution density of the refined anomaly source tag, obtain the real-time data stream and compare it with the feature distribution density to obtain the feedback deviation value. If the feedback deviation value exceeds the preset deviation threshold, adjust the grouping parameters in the clustering process and re-cluster the preliminary feature vector to obtain the optimized grouped data cluster. S17, use the optimized grouped data cluster to perform anomaly scoring on the new real-time data stream to obtain the attribution probability. If the attribution probability exceeds the preset warning threshold, generate a fault warning signal.
[0015] In step S11, the real-time vibration sequence and pressure fluctuation sequence of the vehicle-mounted mobile pump station are acquired. The real-time vibration sequence and pressure fluctuation sequence are preprocessed to obtain a preliminary feature vector, including: Real-time vibration and pressure fluctuation sequences of the vehicle-mounted mobile pump station are acquired through a pre-deployed sensor network. The real-time vibration sequence and the pressure fluctuation sequence are filtered to obtain a clean vibration sequence and a clean pressure sequence. The pure vibration sequence and the pure pressure sequence are truncated by time windows to obtain a piecewise matrix; If the local variance of the segmented matrix is greater than a preset variance threshold, then a fitting process is performed on the segmented matrix to determine the time series feature set; The time-series feature set is subjected to dimensionality reduction processing to obtain preliminary feature vectors.
[0016] It should be noted that the sensor network is deployed at key monitoring nodes of the vehicle-mounted mobile pumping station, including vibration sensors and pressure sensors, which are used to collect mechanical vibration data on the equipment surface and pressure change data of the fluid inside the pipeline, respectively. Due to electromagnetic interference from motors and ambient noise in industrial environments, the raw signals collected often contain a large amount of clutter. Bandpass filtering technology can filter out high-frequency electromagnetic noise and low-frequency environmental interference, thereby obtaining pure vibration and pressure sequences, effectively improving the accuracy of subsequent feature extraction and preventing noise from masking the true characteristics of equipment failure.
[0017] To capture transient changes at different operational stages of the pumping station, time windows need to be applied to the pure vibration and pure pressure sequences. The time window length is determined based on statistical analysis of the duration of typical transient events at the pumping station. Statistical analysis of the duration of typical transient events such as cavitation, water hammer, and valve switching in historical fault data reveals that these events typically last between 1 and 3 seconds, with the characteristic changes occurring mainly concentrated within 2 seconds. Therefore, a time window length of 2 seconds is chosen to fully encompass the entire process of a single transient event. The window step size is determined by considering both computational efficiency and temporal resolution. An excessively large step size leads to inaccurate event boundary capture, while an excessively small step size increases the computational burden. Optimization analysis shows that a step size of half the window length, i.e., 1 second, ensures sufficient overlap between adjacent windows to capture event continuity while keeping the computational load within a reasonable range. Therefore, the continuous one-dimensional time series is divided into multiple segmented matrices containing both vibration and pressure data. This matrix processing method transforms isolated single-point data into structured data with spatiotemporal correlation.
[0018] It is worth noting that not all data from all time periods are valuable for analysis. When the pump station is operating smoothly, data fluctuations are minimal, and feature extraction is not very meaningful. Therefore, the local variance of the segmented matrix is calculated to measure the degree of data fluctuation within that time period. The preset variance threshold is set based on statistical analysis of a large amount of historical normal operation data. Specifically, vibration and pressure data of the pump station under rated operating conditions and different load conditions are collected during smooth operation. The local variance of each segmented matrix is calculated to obtain the local variance distribution under normal operating conditions. Statistical analysis shows that the 95th percentile of the local variance value under normal operating conditions is approximately 0.5, meaning that 95% of the normal operating data has a local variance of no more than 0.5. Therefore, the preset variance threshold is set to 0.5.
[0019] When the local variance of a certain segment of the matrix is greater than 0.5, it indicates that the data fluctuations within that time period exceed the statistical range of normal operating conditions, possibly due to events such as cavitation, water hammer, or valve switching that cause drastic fluctuations. In this case, polynomial fitting is performed on the segmented matrix to extract dynamic features that reflect transient waveforms, such as the rate of change of peaks and troughs, forming part of the time-series feature set. When the local variance of a certain segment of the matrix is less than or equal to 0.5, it indicates that the pump station is operating smoothly within that time period, and the data fluctuations are within the normal range. In this case, complex fitting is not required; instead, the statistical features of the segmented matrix are directly extracted, including the mean, variance, maximum, and minimum values of vibration amplitude, as well as the mean and variance of pressure data. These statistical features are also included in the time-series feature set. In other words, regardless of whether the local variance is greater than the threshold, each segmented matrix will generate a set of feature vectors, ensuring the completeness of the input for subsequent clustering and classification steps. This ensures both high-precision feature extraction in abnormal fluctuation ranges and data integrity in stable operating ranges.
[0020] The time-series feature sets obtained through fitting are often high-dimensional and contain redundant information. Principal component analysis (PCA) can be used for dimensionality reduction, projecting the original high-dimensional features onto orthogonal directions with the largest variance, thus retaining the principal component components that contribute most to data variability. Specifically, the time-series feature set is first centered, and the covariance matrix is calculated. Then, eigenvalue decomposition is performed on the covariance matrix, selecting the top k eigenvectors with a cumulative contribution rate of over 95% to form the projection matrix. Finally, the original high-dimensional features are multiplied by the projection matrix to obtain a low-dimensional preliminary feature vector. For example, compressing the original 50-dimensional time-series feature set to a 10-dimensional preliminary feature vector not only retains the core information reflecting the pump station's operating status but also significantly improves the inference speed and generalization ability of subsequent state recognition models.
[0021] For example, in a flood control and disaster relief operation, a sensor network collected vibration data from the surface of a pumping station and pressure data from inside the pipelines in real time. The raw signal contained 50Hz high-frequency noise generated by electromagnetic interference from the motor and low-frequency interference from environmental vibrations. After filtering out these, a clean signal was obtained. Continuous signals were truncated in a 2-second window with a 1-second step size to obtain multiple segmented matrices. Statistical analysis determined a preset variance threshold of 0.5. The local variance of a certain segmented matrix was calculated to be 0.8, which is greater than 0.5, indicating that a cavitation event occurred at the pumping station during that period. Polynomial fitting was performed on the matrix to extract time-series features such as the peak change rate of the vibration waveform (0.35) and the trough change rate of the pressure fluctuation (0.28), forming a 50-dimensional time-series feature set. Principal component analysis was used for dimensionality reduction. The covariance matrix was calculated and the eigenvalues were solved. The cumulative contribution rate of the first 10 principal components reached 96%. Therefore, the first 10 principal components were used to form a projection matrix, which mapped the 50-dimensional features into a 10-dimensional preliminary feature vector. This vector retained the core fault feature information and provided an efficient input for subsequent state recognition.
[0022] In step S12, clustering is performed based on the preliminary feature vectors to generate grouped data clusters, including: The initial feature vectors are mapped to a high-dimensional feature space, and the local density value of each initial feature vector is calculated in the high-dimensional feature space. The adjacent preliminary feature vectors are aggregated based on the local density values to generate the original data clusters; Calculate the centroid of each of the original data clusters, and calculate the Euclidean distance from each of the preliminary feature vectors within the cluster to the corresponding centroid, to obtain a set of Euclidean distances. Determine the intra-cluster convergence value based on the set of Euclidean distances. If the convergence value within the cluster is less than a preset convergence threshold, the original data cluster is marked as a stable operating condition cluster; otherwise, the original data cluster is marked as a transient disturbance cluster. The labeled stable condition cluster and transient disturbance cluster are output as grouped data clusters.
[0023] It should be noted that the DBSCAN density clustering algorithm is used for cluster analysis. This algorithm sets a cutoff distance parameter and counts the number of neighboring nodes within that cutoff distance for each feature vector, using this as the local density distribution value for that vector. The cutoff distance is set based on statistical analysis of the distance between sample points in the feature space, typically taking the 2% to 5% quantile of the Euclidean distance between sample points to ensure that dense regions can be identified while effectively filtering out noisy points. Using the vector with the highest local density as the core, nearby vectors are aggregated to generate multiple original data clusters. This density-based clustering method can automatically identify clusters of arbitrary shapes and has good robustness to noisy data.
[0024] For each generated original data cluster, the average value of all preliminary feature vectors within the cluster is calculated across all dimensions. This average value is defined as the cluster centroid. Then, the Euclidean distance from each vector point within the cluster to the centroid is calculated, forming a set of Euclidean distances. By calculating the average of all distances in this set, the cluster convergence value is obtained. If the average Euclidean distance of a cluster is small, it indicates that the data points within the cluster are closely distributed around the centroid, indicating high data consistency; if the average is large, it indicates that the data points are widely distributed and exhibit high volatility.
[0025] It is worth noting that the data distribution pattern of the pump station varies significantly under different operating conditions. The preset convergence threshold is based on statistical analysis of historical normal operation data. Specifically, historical operating data of the pump station under various stable operating conditions are collected, and the intra-cluster convergence values of the corresponding clusters for each stable operating condition are calculated to obtain the convergence distribution under stable operating conditions. Statistical analysis shows that the 95th percentile of the convergence values under stable operating conditions is approximately 0.3, meaning that 95% of the stable operating condition data has an intra-cluster convergence of no more than 0.3. Therefore, the preset convergence threshold is set to 0.3. When the calculated intra-cluster convergence value is less than 0.3, it indicates that the fluctuation pattern of the data segment corresponding to that cluster is highly consistent, and it is marked as a stable operating condition cluster, which usually corresponds to the stable operation phase of the pump station at rated speed. Conversely, if the intra-cluster convergence value is greater than 0.3, it indicates that the data point distribution is divergent, and it is marked as a transient disturbance cluster, which often corresponds to the stage where air is mixed into the pump station's suction end or water hammer occurs in the pipeline. The labeled stable condition cluster and transient disturbance cluster are output as grouped data clusters.
[0026] For example, suppose we obtain a preliminary 10-dimensional feature vector with 1000 sample points, and map it to a 10-dimensional high-dimensional feature space. After statistical analysis of the pairwise Euclidean distances between all sample points in the feature space, the 3rd percentile of the distance values is calculated to be 0.5. Therefore, the cutoff distance parameter is set to 0.5. The DBSCAN algorithm performs clustering with a cutoff distance of 0.5, counting the number of neighboring nodes of each feature vector within a radius of 0.5, and obtaining the local density distribution of each vector. A vector representing the impeller rotation feature has a local density of 45. The algorithm uses this vector as the core and aggregates all vectors with a distance less than 0.5 to form an original data cluster containing 120 samples.
[0027] The centroid of this cluster was calculated, and the average values for each dimension were obtained as follows: vibration frequency 128 Hz, vibration amplitude 0.32 mm / s, and pressure fluctuation amplitude 0.15 MPa. The Euclidean distance from each point within the cluster to the centroid was calculated, resulting in a distance set with an average value of 0.15. Since 0.15 is less than the preset convergence threshold of 0.3, it indicates that the data points within this cluster are highly clustered and exhibit consistent fluctuation patterns. Therefore, this cluster was marked as a stable operating condition cluster, corresponding to the stable operation phase of the pump station at its rated speed. Another cluster contained 80 samples, and the average Euclidean distance from each point within the cluster to the centroid was 0.8, which is greater than the threshold of 0.3. This indicates that the data points in this cluster are widely distributed and exhibit large fluctuations. This cluster was marked as a transient disturbance cluster, corresponding to the interference phase where air is mixed into the suction end. The marked stable operating condition cluster and transient disturbance cluster were output as grouped data clusters.
[0028] In step S13, the grouped data clusters are probability-mapped using a pre-trained classifier to obtain anomaly probability scores, including: Assign status labels to the grouped data clusters to construct a labeled training set; The labeled training set is input into a preset support vector machine classifier for training to obtain the decision boundary; Input the data cluster to be tested into the decision boundary, calculate the geometric distance between the data cluster to be tested and the decision boundary, and obtain the geometric distance value; The geometric distance value is input into a preset S-shaped probability mapping function and converted into anomaly probability scores.
[0029] It should be noted that during the construction of the labeled training set, semantic assignment is required based on the different state clusters generated by clustering in the previous steps. Specifically, the system automatically labels each data cluster based on its intra-cluster convergence value and cluster centroid location characteristics. For clusters with an intra-cluster convergence value less than a preset convergence threshold of 0.3 and whose centroid location highly matches the distribution of historical normal operation data centroids, the system automatically labels them as "normal pumping state." For clusters with an intra-cluster convergence value greater than 0.3 and whose centroid location significantly deviates from the normal operation area, the system further subdivides them based on the physical meaning of the feature space region where their centroids are located. For example, when the cluster centroid exhibits high-frequency characteristics in the vibration frequency dimension and large fluctuations in the pressure fluctuation dimension, it is labeled as "cavitation interference state." When the cluster centroid exhibits low-frequency large vibrations in the vibration amplitude dimension and slow changes in the pressure fluctuation dimension, it is labeled as "intake disturbance state." This automatic labeling mechanism is based on statistical analysis of a historical fault case library, and determines the accuracy and consistency of the labels by matching the clustering results with known fault modes.
[0030] Support Vector Machine (SVM) classifiers map low-dimensional features to a high-dimensional space using kernel functions. Within this high-dimensional space, they search for an optimal classification surface that maximizes the margin between two classes of samples; this surface is the decision boundary. The decision boundary is determined through iterative optimization calculations on the labeled training set. During training, the classifier continuously adjusts the position and orientation of the classification surface until the margin between the closest samples from different classes is maximized. This yields an optimal decision boundary that can tolerate a certain level of noise interference, ensuring that minor fluctuations in pump station operating data are not incorrectly identified as faults.
[0031] Once the decision boundary is established, the system enters the real-time monitoring phase. After the data cluster to be tested is input into the decision boundary, the system calculates the vertical distance from the centroid of the feature vector cluster to the optimal decision boundary. The magnitude of this distance reflects the degree to which the sample deviates from the normal state; a larger distance indicates a more typical anomaly. Directly using geometric distance presents two problems: first, the physical meaning of geometric distance is not intuitive and difficult for operators to understand; second, the distance scale is inconsistent under different data distributions, making it difficult to set a uniform threshold. Therefore, a sigmoid probability mapping function is introduced to normalize the distance. This function can smoothly map geometric distance values of any range to the interval between 0 and 1, transforming them into intuitive anomaly probability scores. The sigmoid probability mapping function determines its mapping parameters through cross-validation with historical data, ensuring that the mapping result approaches 0 when the distance value is small, approaches 1 when the distance value is large, and presents a smooth transition in the intermediate region, thus transforming the abstract geometric distance into an intuitive anomaly probability.
[0032] For example, assuming that a stable operating condition cluster and a transient disturbance cluster are obtained from the clustering results, the intra-cluster convergence value of the stable operating condition cluster is 0.12, and the cluster centroid position highly matches the centroid of historical normal operation data, the system automatically marks it as "normal pumping state"; the intra-cluster convergence value of the transient disturbance cluster is 0.85, and the cluster centroid is located at a vibration frequency of 128Hz and a pressure fluctuation of 0.35MPa, with a feature pattern matching degree of 92% with "cavitation disturbance state" in the historical cavitation case library, the system automatically marks it as "cavitation disturbance state", and a labeled training set is constructed. This training set is input into a support vector machine classifier, and the classifier finds an optimal classification surface that maximizes the interval between the two classes of samples, "normal pumping state" and "cavitation disturbance state", through iterative optimization calculation. In real-time monitoring, a new vibration and pressure signal is converted into a feature vector cluster, and the calculated vertical distance from its centroid to the optimal classification surface is 2.5. By substituting the distance value of 2.5 into the S-shaped probability mapping function parameters determined through cross-validation, the anomaly probability score was calculated to be 0.88.
[0033] In step S14, if the anomaly probability score exceeds a preset score threshold, the contribution of each feature in the corresponding classification result is extracted, and the features are sorted in descending order of contribution to obtain a highly significant feature sequence. The highly significant feature sequence is then matched with a preset mapping dictionary to obtain preliminary anomaly labels, including: If the anomaly probability score is greater than a preset score threshold, then extract the multidimensional features from the corresponding classification result; The contribution of each dimension of the multidimensional features to the classification result is evaluated using a preset random forest algorithm, and the feature subset and the contribution weight of each dimension feature are obtained. The feature subsets are sorted according to the contribution weights to generate a highly significant feature sequence; Potential sources of anomalies are determined by matching the highly significant feature sequences with a pre-defined mapping dictionary. Assign a text identifier to the potential source of the anomaly and output an initial anomaly label.
[0034] It should be noted that after the system obtains the anomaly probability score of the current pump station's operating status, it compares it with a preset judgment threshold. The preset score threshold is set based on ROC curve analysis of historical abnormal and normal samples. By calculating the true positive rate and false positive rate under different thresholds, the threshold that maximizes the Youden index is selected as the optimal classification threshold. Statistical analysis of historical fault data shows that this optimal threshold is typically distributed between 0.80 and 0.90. The specific value can be adjusted according to the historical data distribution characteristics of different pump station models; in this embodiment, 0.85 is used. If the currently obtained score is greater than 0.85, the system determines that the current state has a significant anomaly and immediately extracts the multidimensional feature classification result corresponding to that moment.
[0035] The Random Forest algorithm consists of numerous decision trees. Each decision tree is trained using randomly selected subsets of samples and features, improving the stability and accuracy of the evaluation through ensemble learning. When dealing with high-dimensional features such as vibration frequency, bearing temperature, and inlet / outlet pressure difference in pump stations, the algorithm measures the contribution of each feature dimension to the classification result by calculating the reduction in node impurity. Specifically, for each feature dimension, the algorithm calculates the total reduction in impurity when splitting nodes using that feature across all decision trees, divides this total by the sum of the impurity reductions for all features, and obtains the feature importance score. After evaluation, the system selects features with the highest importance scores to form a feature subset, and the score corresponding to each feature is the contribution weight. The selection of feature subsets is usually based on a cumulative contribution of over 90%, ensuring that the core features with the greatest impact on the classification result are retained.
[0036] The system sorts feature subsets in descending order based on contribution weights to generate highly significant feature sequences, which intuitively reflect the core driving factors of the current abnormal state. An internal pre-built mapping dictionary is used to match highly significant feature sequences to specific anomaly sources. The dictionary is constructed by first collecting a historical fault case database. Each case includes a multi-dimensional feature contribution sequence at the time of the fault and a fault cause label confirmed on-site. Data sources include historical pump station operation records, maintenance work orders, and fault analysis reports. Then, the Apriori association rule mining algorithm is used to mine frequent patterns in the historical case data. The Apriori algorithm iteratively generates candidate feature sets and prunes them to find association rules between feature combinations with support and confidence exceeding preset thresholds and fault types. Specifically, the minimum support is set to 0.1 and the minimum confidence to 0.85. For example, when the radial vibration amplitude contribution is greater than 0.4 and the differential pressure fluctuation contribution is greater than 0.3, there is a 90% probability of corresponding to cavitation. These rules are stored in a pre-defined mapping dictionary. Each record in the dictionary contains the judgment conditions for feature combinations and the corresponding fault label. The dictionary is updated as follows: whenever a new fault case is confirmed on-site, the system adds the feature contribution sequence and fault label of that case to the historical case database. Every 30 days, the Apriori algorithm is run again to mine the entire case database and update the association rules in the dictionary. When an uncovered feature combination occurs, i.e., a highly significant feature sequence cannot match the judgment conditions of any rule in the dictionary, the system uses a nearest neighbor matching strategy. It calculates the cosine similarity between the current sequence and each rule feature combination in the dictionary, and selects the rule with the highest similarity exceeding 0.7 as the matching result. If the highest similarity is still below 0.7, the anomaly is marked as an "unknown anomaly" and a manual review process is triggered. After expert confirmation, the anomaly is added to the case database.
[0037] For example, the system obtains an anomaly probability score of 0.91, which is greater than the preset score threshold of 0.85, thus determining that the current state has a significant anomaly and extracting the corresponding classification result. The system uses a random forest algorithm to evaluate the feature contribution. This algorithm consists of 200 decision trees and evaluates the importance of 12-dimensional features, including vibration frequency, vibration amplitude, bearing temperature, inlet and outlet pressure difference, and flow rate change rate. The calculated weights are: impeller radial vibration amplitude 0.45, pump cavity internal fluid pressure difference fluctuation 0.35, motor casing temperature 0.08, flow rate change rate 0.07, and bearing temperature 0.05. The feature subsets are sorted in descending order according to their contribution weights to generate a highly significant feature sequence, with impeller radial vibration amplitude as the first value and fluid pressure difference fluctuation as the second. The sequence is input to an attribute mapping dictionary for matching. The dictionary contains a strong association rule mined using the Apriori algorithm: when the radial vibration amplitude contribution is greater than 0.4 and the differential pressure fluctuation contribution is greater than 0.3, the confidence level is 0.92, corresponding to "cavitation phenomenon". In the current sequence, both the radial vibration amplitude contribution of 0.45 and the differential pressure fluctuation contribution of 0.35 meet the conditions, resulting in a successful match. The system associates this sequence with cavitation phenomena inside the pump body, assigns the text label "suspected cavitation" to the potential anomaly source, and outputs a preliminary anomaly label. If a highly significant feature sequence does not match any rules in a match, the system calculates the cosine similarity. Assuming the highest similarity is only 0.5, below the threshold of 0.7, the anomaly is marked as "unknown anomaly," and the relevant data is recorded for subsequent manual analysis.
[0038] In step S15, the confidence level of the preliminary anomaly label is calculated. If the confidence level exceeds a preset confidence threshold, the preliminary anomaly label is refined into anomaly cause labels according to a preset level, serving as refined anomaly source labels, including: Historical environmental factor data are extracted based on the preliminary anomaly labels, and a time-series feature vector is constructed. The time-series feature vectors are cross-validated to obtain a data subset, and the data subset is fitted by a preset support vector machine model to generate a fusion matrix; The initial anomaly labels are correlated with the fusion matrix to calculate the confidence level of the initial anomaly labels and obtain a confidence score. If the confidence score is greater than the preset confidence threshold, the preliminary anomaly labels are hierarchically divided using a preset classification mapping table to obtain anomaly cause labels, which are then output as refined anomaly source labels.
[0039] It should be noted that after obtaining the initial anomaly tag of "suspected cavitation," the system will trace and extract historical environmental factor data related to that tag. The selection of historical environmental factor data is based on in-depth analysis of the cavitation mechanism. The occurrence of cavitation is closely related to three factors: the liquid level at the pump suction end, atmospheric pressure, and ambient temperature. The liquid level at the suction end directly affects the pump's suction pressure; a low liquid level easily triggers cavitation. Changes in atmospheric pressure affect the vaporization pressure of the liquid. Ambient temperature indirectly affects the critical conditions for cavitation by influencing the liquid's saturated vapor pressure. Therefore, the external environmental data of the pump station over the past 72 hours, including the water tank level, atmospheric pressure, and ambient temperature, are retrieved. The 72-hour time span is chosen based on statistical analysis of typical meteorological change cycles. By analyzing the autocorrelation function of atmospheric pressure and ambient temperature in historical data, it was found that their effective influence duration is approximately 48 to 96 hours. Taking the median value of 72 hours can cover a complete meteorological change process while ensuring sufficient data volume. The environmental data arranged in chronological order were aligned to construct a multi-dimensional time series feature vector. This vector has an hourly time step and contains 72 time points. Each time point contains three dimensions: liquid level, air pressure, and temperature, for a total of 216 feature components, reflecting the dynamic change trajectory of the external environment during the operation of the pumping station.
[0040] The system employs a 5-fold cross-validation method to process the aforementioned time-series feature vectors. The 72-hour data sequence is divided into five non-overlapping subsets, each containing approximately 14.4 hours of data. Four subsets are used alternately as training samples, with the remaining subset serving as the validation sample. The choice of 5-fold cross-validation is based on a comprehensive consideration of model evaluation stability and computational efficiency. It ensures a sufficient training sample size (approximately 57.6 hours of data) while obtaining reliable model performance estimates through five independent validations, avoiding random biases caused by a single partition. Subsequently, a support vector machine (SVM) model is introduced to fit these data subsets. The SVM maps low-dimensional environmental features to a high-dimensional space using a radial basis function kernel, searching for the optimal hyperplane in this space that can distinguish between normal environmental fluctuations and induced abnormal environmental conditions. After multiple iterative fitting iterations, the system extracts the support vectors and their corresponding weights from each fold of validation, constructing a multi-dimensional fusion matrix. Each row of this matrix corresponds to a support vector, and each column corresponds to an environmental feature dimension. The matrix element values reflect the weight contribution of the support vector in the corresponding feature dimension.
[0041] The confidence score for the initial anomaly label "suspected cavitation" is calculated using this fusion matrix. Specifically, the cosine similarity is calculated between the current environmental factor time-series feature vector and each support vector in the fusion matrix. Then, the similarities are weighted and summed using the weight coefficients of each support vector to obtain the confidence score. The fusion matrix quantifies the degree of agreement between the current environmental factor change trajectory and typical cavitation occurrence conditions. The preset confidence threshold is based on statistical analysis of historically verified cavitation cases. Fifty field-confirmed cavitation cases are collected, and the confidence score for each case is calculated. The lower quartile of the confidence scores for all cavitation cases is taken as the threshold; in this embodiment, this value is 0.80. If the confidence score is greater than 0.80, the initial label is hierarchically divided using a preset classification mapping table.
[0042] The system has a pre-built multi-dimensional classification mapping table, organized into three levels: physical system, component location, and triggering mechanism. The physical system level is divided into hydraulic system, mechanical transmission system, electrical control system, etc.; the component location level is further refined according to the pump station structure, such as suction end, impeller cavity, discharge end, bearing housing, etc.; the triggering mechanism level is subdivided into low liquid level triggering, temperature change triggering, pressure fluctuation triggering, and gas-containing medium triggering. The classification mapping table is constructed based on the knowledge system of domain experts, categorizing historical fault cases according to the above three-level structure to form a complete fault classification system. The system inputs high-confidence labels into this mapping table for step-by-step comparison: first matching the physical system level, then the component location level, and finally the triggering mechanism level. After three layers of filtering, the abnormal cause label is obtained and output as a refined abnormal source label.
[0043] For example, after obtaining the initial anomaly label of "suspected cavitation," the system retrieves data on the water level, atmospheric pressure, and ambient temperature of the pump station over the past 72 hours to construct a time-series feature vector. This vector, with an hourly step size, contains 72 time points, each with three dimensions, totaling 216 feature components. Five-fold cross-validation is used to divide the 72-hour data sequence, and a fusion matrix is generated through a support vector machine model. This fusion matrix contains approximately 150 support vectors, each corresponding to a set of environmental feature combinations. The confidence score of the "suspected cavitation" label is calculated using this fusion matrix. The cosine similarity between the current environmental time-series feature vector and each of the 150 support vectors is calculated and then weighted and summed, resulting in a confidence score of 0.89, which is greater than the preset threshold of 0.80. This confirms the label's high confidence and upgrades it to a high-confidence label. The label is input into a classification mapping table. First, it is matched at the physical system level, locating the hydraulic system based on the characteristics of cavitation. Then, it is matched at the component location level, locating the suction end based on the source of the abnormal signal. Finally, it is matched at the inducing mechanism level, matching low liquid level-induced cavitation based on the characteristic of the liquid level being continuously below the critical value of 3.2 meters in historical environmental data. After a step-by-step comparison at these three levels, the final abnormal cause label "low liquid level-induced cavitation at the suction end" is obtained and output as the refining abnormality source label.
[0044] In step S16, the feature distribution density of the refined anomaly source tag is calculated, a real-time data stream is acquired and compared with the feature distribution density to obtain a feedback deviation value. If the feedback deviation value exceeds a preset deviation threshold, the grouping parameters in the clustering process are adjusted, and the initial feature vector is re-clustered to obtain optimized grouped data clusters, including: Extract operational features from the refining anomaly source tags and calculate the feature distribution density of the operational features; Acquire real-time data stream, compare the real-time data stream with the feature distribution density, and calculate the feedback deviation value; If the feedback deviation value is greater than the preset deviation threshold, the weight parameters of the distance metric function in the clustering process are adjusted, and the adjusted distance metric function is used to re-cluster the initial feature vector to generate optimized grouped data clusters.
[0045] It should be noted that after acquiring the source label of the refining anomaly, the system extracts the multi-dimensional operational features corresponding to that label, such as rotor vibration frequency and bearing temperature. The system maps these features into a multi-dimensional feature space and uses kernel density estimation to calculate their feature distribution density. Kernel density estimation uses the feature vector corresponding to the high-confidence label as its core, employing a Gaussian kernel function to statistically analyze the clustering degree of data points within a specific radius neighborhood, thereby quantifying the density of this type of anomaly distribution in the feature space. The neighborhood radius is set based on a global statistical analysis of historical normal operation data during system initialization. Historical feature data of the pump station under various normal operating conditions are collected, and the Euclidean distances between all pairwise sample points are calculated. The median of these distances is taken as a fixed neighborhood radius parameter. This radius remains unchanged during system operation, ensuring the stability of the feature distribution density baseline and avoiding density baseline drift caused by real-time feedback.
[0046] The system introduces a real-time feedback mechanism to verify the accuracy of the aforementioned feature distribution density. It continuously receives real-time data streams from the pump station's on-site monitoring equipment and compares them with the calculated feature distribution density. The feedback deviation value is calculated using the relative entropy method, which quantifies model deviation by measuring the degree of difference between the real-time data distribution and the historical feature distribution. Specifically, the system divides the real-time data stream into multiple data segments according to time windows, statistically analyzes the probability distribution of each data segment across each feature dimension, calculates the relative entropy between these distributions and the historical feature distribution, and takes the average of the relative entropies of all windows as the feedback deviation value. The preset deviation threshold is set based on statistical analysis of the pump station's normal state drift. Real-time data from the pump station under normal operating conditions for one week is collected, and its feedback deviation value is calculated. The 95th percentile of these values is taken as the threshold; in this embodiment, this value is 0.15. If the feedback deviation value is greater than 0.15, the system determines that the current data grouping model cannot accurately reflect the real-time operating status of the pump station, and parameter adjustments are required.
[0047] To address excessive feedback bias, the system triggers an adjustment mechanism for the distance metric function. The standard Euclidean distance function, originally used to measure data point similarity, assigns equal weights to all dimensions in the feature space, failing to reflect the differentiated contributions of different features to anomaly detection. The system replaces this with a weighted Euclidean distance function that dynamically adjusts weights based on feature importance. The weight values are determined based on the feature contribution weights calculated using the random forest algorithm in step S14; features with higher contributions are assigned greater weights. For example, when the contribution of vibration frequency is significantly higher than other features, the system assigns a larger weight coefficient to the vibration frequency dimension in the weighted distance function.
[0048] Based on the adjusted distance metric function, the system re-clusters the initial feature vectors. Specifically, the system first recalculates the weighted distances between all pairs of initial feature vectors using the weighted Euclidean distance function, replacing the original Euclidean distances. Then, using the same DBSCAN algorithm as the original clustering, with the same cutoff distance and minimum sample number parameters, it re-performs density clustering based on the newly calculated weighted distance matrix, aggregating vectors whose weighted distance is less than the cutoff distance and whose density is connected into new clusters. Finally, the regrouped data clusters are obtained as the output of the optimized grouped data clusters.
[0049] It is worth noting that after the system completes the adjustment of clustering parameters and obtains optimized grouped data clusters, the original clustering results, decision boundaries, and feature contribution models may no longer be suitable. To ensure the accuracy and consistency of subsequent steps, the system automatically triggers a model update process. First, based on the optimized grouped data clusters, the intra-cluster convergence value of each new cluster is recalculated, and the stable operating condition clusters and transient disturbance clusters are relabeled. Then, step S13 is re-executed using the newly labeled grouped data clusters, i.e., the support vector machine classifier is retrained to obtain the updated decision boundaries. Next, based on the new decision boundaries and classification results, the random forest feature contribution evaluation in step S14 is re-executed to update the feature subsets and contribution weights. Finally, the updated decision boundaries and contribution weights are used for subsequent real-time monitoring. The mapping dictionary is constructed based on the fault mechanism and has no direct dependency on the clustering results, so it does not need to be updated. However, when new feature combinations appear, the dictionary supplementation mechanism is executed as described in step S14. This closed-loop update mechanism ensures the consistency of the entire model chain from clustering to classification to feature evaluation.
[0050] For example, the system obtains high-confidence labels from the refining anomaly source classification "cavitation induced by low liquid level at the intake end," extracts operating features such as rotor vibration frequency and bearing temperature. During system initialization, a fixed neighborhood radius of 0.3 is determined based on historical normal operating data, and the feature distribution density is calculated using kernel density estimation. The real-time monitoring data stream is divided into 10-minute windows, continuously collecting data from 72 windows. After statistically analyzing the probability distribution of each window, the relative entropy with the historical feature distribution is calculated, and the average value is used to obtain a feedback deviation value of 0.25. This value is greater than the preset deviation threshold of 0.15, indicating that the current model cannot accurately reflect the real-time state. The system triggers an adjustment mechanism. Based on the feature contribution weights calculated by the random forest algorithm in step S14, the vibration frequency contribution is 0.7, and the bearing temperature contribution is 0.3. The standard Euclidean distance function is replaced with a weighted Euclidean distance function, where the weight coefficient for the vibration frequency dimension is set to 0.7, and the weight coefficient for the bearing temperature dimension is set to 0.3. Subsequently, the system recalculates the weighted distances between all pairwise features using the weighted Euclidean distance function, and employs the DBSCAN algorithm to cluster based on the new distance matrix. The cutoff distance remains at 0.5, and the minimum sample size is 10. Points with a weighted distance less than 0.5 and density contiguous are grouped into new clusters. After re-clustering, some boundary samples that originally belonged to the stable operating condition cluster separate due to the increased weighted distance, forming a more compact cluster structure, thus generating optimized grouped data clusters. The system then recalculates the convergence, retrains the support vector machine classifier and updates the decision boundary based on the new clustering results, and re-evaluates the feature contribution, providing a consistent foundation for subsequent real-time monitoring.
[0051] In step S17, the optimized grouped data clusters are used to perform anomaly scoring on the new real-time data stream to obtain the attribution probability. If the attribution probability exceeds a preset warning threshold, a fault warning signal is generated, including: Extract key metrics from the new real-time data stream to form real-time data points; Calculate the Euclidean distance between the real-time data point and the cluster center vector of each cluster in the optimized grouped data cluster; Anomaly scores are obtained using a preset local density calculation formula based on the Euclidean distance. The probability of attribution is calculated using a preset logistic regression function for the abnormal scores. If the attribution probability is greater than the preset warning threshold, a fault warning signal is generated.
[0052] It should be noted that after acquiring the optimized grouped data clusters, the system will simultaneously receive a new round of real-time data streams from the pump station monitoring nodes. These data streams contain key operational indicators such as the pump station's current outlet pressure and instantaneous flow rate. The system extracts these indicators to form real-time data points and calculates the Euclidean distance between these data points and the cluster center vectors in the data grouping model. The cluster center vectors represent the typical operational characteristics of the pump station under different conditions. By calculating the Euclidean distance, the system can intuitively quantify the physical spatial distance of the current real-time data point from the known typical state.
[0053] After obtaining the Euclidean distances between the current data point and each cluster center, the system also needs to evaluate the local density of the data point within its neighborhood to determine whether it is an outlier in a sparse region. To this end, the system employs a nearest-neighbor-based local density calculation method. First, it finds the K nearest neighbor sample points of the current data point in the feature space, calculates the Euclidean distances between the data point and these K nearest neighbors, and takes the average value. Then, it generates an anomaly score using the reciprocal of this average distance. Specifically, the smaller the average distance, the denser the samples around the data point, and the lower the anomaly score; the larger the average distance, the more sparse the data point, and the higher the anomaly score. This anomaly score directly reflects the degree of deviation of the pump station's current operating state from normal operating conditions.
[0054] To further clarify the probability of a fault occurring, the system inputs the anomaly score into a pre-configured logistic regression function. The logistic regression function, through non-linear mapping, transforms continuous anomaly scores into attribution probability values between 0 and 1. This probability value characterizes the tendency of the current pump station state to belong to a specific severe mechanical fault. The parameters of the logistic regression function are trained using historical fault sample data, ensuring that the anomaly score for the normal state approaches 0 after mapping, and the anomaly score for the severe fault state approaches 1 after mapping, with a smooth transition in the intermediate region. The preset warning threshold is set based on statistical analysis of historical fault samples. Fault cases confirmed on-site are collected, and the attribution probability of each case is calculated. The threshold that achieves a fault detection rate of 95% while minimizing the false alarm rate is taken as the warning threshold; in this embodiment, this value is 0.75. If the attribution probability is greater than 0.75, the system determines that the current pump station has an extremely high fault risk, immediately generates a fault warning signal, and immediately triggers the underlying blocking logic, issuing an emergency stop command to the pump station control cabinet to forcibly cut off the power supply to the motor's main circuit, thereby maintaining the operational stability of the entire pump station network.
[0055] For example, after acquiring the optimized data grouping model, the system receives a new round of real-time data stream and extracts the outlet pressure of 5.2 MPa and the instantaneous flow rate of 1.8 cubic meters per second to form real-time data points. The Euclidean distance between this data point and the cluster center vectors in the model is calculated: the distance to the normal state cluster center is 5.2, and the distance to the cavitation state cluster center is 1.8. The system finds the five nearest neighbor samples of this data point in the feature space, and the calculated Euclidean distances are 5.2, 5.5, 5.3, 5.4, and 5.1, with an average of 5.3. The anomaly score is calculated as 78 points using the reciprocal relationship. The anomaly score of 78 points is input into the logistic regression function. The regression coefficients of this function are determined to be 0.05 and -2.5 through training with historical data. The attribution probability is calculated as 0.80 through nonlinear mapping. The preset warning threshold is 0.75. If the probability of failure is 0.80, which is greater than the warning threshold, the system determines that there is an extremely high risk of failure, generates a fault warning signal, and triggers the blocking logic to issue an emergency shutdown command to the pump station control cabinet to maintain the stability of the pump station pipeline operation.
[0056] In summary, this invention discloses an abnormal data detection method for vehicle-mounted mobile pumping stations. It establishes a state recognition database through signal preprocessing and cluster analysis, quantifies the anomaly probability using support vector machines, locates the anomaly source by combining random forests, refines the classification by integrating historical environmental data, and optimizes the detection model through real-time feedback. Ultimately, it achieves anomaly quantification analysis and early warning blocking, solving the problems of low anomaly detection accuracy and poor fault early warning reliability in complex environments caused by insufficient industrial data processing capabilities in existing technologies.
[0057] Reference Figure 2 The second embodiment of the present invention provides an abnormal data detection system for vehicle-mounted mobile pumping stations, comprising: The feature extraction module is used to acquire the real-time vibration sequence and pressure fluctuation sequence of the vehicle-mounted mobile pump station, and to preprocess the real-time vibration sequence and pressure fluctuation sequence to obtain a preliminary feature vector. The clustering and grouping module is used to perform clustering processing based on the preliminary feature vectors to generate grouped data clusters; An anomaly detection module is used to perform probability mapping on the grouped data clusters using a pre-trained classifier to obtain anomaly probability scores; The source initial judgment module is used to extract the contribution of each feature in the corresponding classification result if the anomaly probability score exceeds a preset score threshold, sort them in descending order of contribution to obtain a high significance feature sequence, and match the high significance feature sequence with a preset mapping dictionary to obtain a preliminary anomaly label. The source refinement module is used to calculate the confidence level of the preliminary anomaly label. If the confidence level exceeds a preset confidence threshold, the preliminary anomaly label is refined into anomaly cause labels according to a preset level, which are used as refined anomaly source labels. The model optimization module is used to calculate the feature distribution density of the refined anomaly source label, obtain the real-time data stream and compare it with the feature distribution density to obtain the feedback deviation value. If the feedback deviation value exceeds the preset deviation threshold, the grouping parameters in the clustering process are adjusted and the initial feature vector is re-clustered to obtain the optimized grouped data cluster. The early warning triggering module is used to score the new real-time data stream for anomalies using the optimized grouped data clusters to obtain the attribution probability. If the attribution probability exceeds the preset early warning threshold, a fault early warning signal is generated.
[0058] It should be noted that the vehicle-mounted mobile pumping station abnormal data detection system provided in this embodiment of the invention is used to execute all the process steps of the vehicle-mounted mobile pumping station abnormal data detection method in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.
[0059] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0060] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for detecting abnormal data in a vehicle-mounted mobile pumping station, characterized in that, include: The real-time vibration sequence and pressure fluctuation sequence of the vehicle-mounted mobile pump station are acquired, and the real-time vibration sequence and pressure fluctuation sequence are preprocessed to obtain a preliminary feature vector. Clustering is performed based on the preliminary feature vectors to generate grouped data clusters; Anomaly probability scores are obtained by performing probability mapping on the grouped data clusters using a pre-trained classifier. If the anomaly probability score exceeds a preset score threshold, the contribution of each feature in the corresponding classification result is extracted, and the features are sorted in descending order of contribution to obtain a high-significance feature sequence. The high-significance feature sequence is then matched with a preset mapping dictionary to obtain a preliminary anomaly label. Calculate the confidence level of the preliminary anomaly label. If the confidence level exceeds a preset confidence threshold, refine the preliminary anomaly label into anomaly cause labels according to a preset level, and use them as refined anomaly source labels. Calculate the feature distribution density of the refined anomaly source tag, obtain the real-time data stream and compare it with the feature distribution density to obtain the feedback deviation value. If the feedback deviation value exceeds the preset deviation threshold, adjust the grouping parameters in the clustering process and re-cluster the initial feature vector to obtain optimized grouped data clusters. The optimized grouped data clusters are used to score the new real-time data stream for anomalies and obtain the attribution probability. If the attribution probability exceeds a preset warning threshold, a fault warning signal is generated.
2. The method for detecting abnormal data in a vehicle-mounted mobile pumping station according to claim 1, characterized in that, The process involves acquiring the real-time vibration and pressure fluctuation sequences of the vehicle-mounted mobile pump station, preprocessing the real-time vibration and pressure fluctuation sequences to obtain preliminary feature vectors, including: Real-time vibration and pressure fluctuation sequences of the vehicle-mounted mobile pump station are acquired through a pre-deployed sensor network. The real-time vibration sequence and the pressure fluctuation sequence are filtered to obtain a clean vibration sequence and a clean pressure sequence. The pure vibration sequence and the pure pressure sequence are truncated by time windows to obtain a piecewise matrix; If the local variance of the segmented matrix is greater than a preset variance threshold, then a fitting process is performed on the segmented matrix to determine the time series feature set; The time-series feature set is subjected to dimensionality reduction processing to obtain preliminary feature vectors.
3. The method for detecting abnormal data in a vehicle-mounted mobile pumping station according to claim 1, characterized in that, The step of clustering based on the preliminary feature vectors to generate grouped data clusters includes: The initial feature vectors are mapped to a high-dimensional feature space, and the local density value of each initial feature vector is calculated in the high-dimensional feature space. The adjacent preliminary feature vectors are aggregated based on the local density values to generate the original data clusters; Calculate the centroid of each of the original data clusters, and calculate the Euclidean distance from each of the preliminary feature vectors within the cluster to the corresponding centroid, to obtain a set of Euclidean distances. Determine the intra-cluster convergence value based on the set of Euclidean distances. If the convergence value within the cluster is less than a preset convergence threshold, the original data cluster is marked as a stable operating condition cluster; otherwise, the original data cluster is marked as a transient disturbance cluster. The labeled stable condition cluster and transient disturbance cluster are output as grouped data clusters.
4. The method for detecting abnormal data in a vehicle-mounted mobile pumping station according to claim 1, characterized in that, The step of performing probability mapping on the grouped data clusters using a pre-trained classifier to obtain anomaly probability scores includes: Assign status labels to the grouped data clusters to construct a labeled training set; The labeled training set is input into a preset support vector machine classifier for training to obtain the decision boundary; Input the data cluster to be tested into the decision boundary, calculate the geometric distance between the data cluster to be tested and the decision boundary, and obtain the geometric distance value; The geometric distance value is input into a preset S-shaped probability mapping function and converted into anomaly probability scores.
5. The method for detecting abnormal data in a vehicle-mounted mobile pumping station according to claim 1, characterized in that, If the anomaly probability score exceeds a preset score threshold, the contribution of each feature in the corresponding classification result is extracted, and the features are sorted in descending order of contribution to obtain a highly significant feature sequence. The highly significant feature sequence is then matched with a preset mapping dictionary to obtain preliminary anomaly labels, including: If the anomaly probability score is greater than a preset score threshold, then extract the multidimensional features from the corresponding classification result; The contribution of each dimension of the multidimensional features to the classification result is evaluated using a preset random forest algorithm, and the feature subset and the contribution weight of each dimension feature are obtained. The feature subsets are sorted according to the contribution weights to generate a highly significant feature sequence; Potential sources of anomalies are determined by matching the highly significant feature sequences with a pre-defined mapping dictionary. Assign a text identifier to the potential source of the anomaly and output an initial anomaly label.
6. The method for detecting abnormal data in a vehicle-mounted mobile pumping station according to claim 1, characterized in that, The calculation of the confidence level of the preliminary anomaly label, if the confidence level exceeds a preset confidence threshold, then the preliminary anomaly label is refined into anomaly cause labels according to a preset level, as refined anomaly source labels, including: Historical environmental factor data are extracted based on the preliminary anomaly labels, and a time-series feature vector is constructed. The time-series feature vectors are cross-validated to obtain a data subset, and the data subset is fitted by a preset support vector machine model to generate a fusion matrix; The initial anomaly labels are correlated with the fusion matrix to calculate the confidence level of the initial anomaly labels and obtain a confidence score. If the confidence score is greater than the preset confidence threshold, the preliminary anomaly labels are hierarchically divided using a preset classification mapping table to obtain anomaly cause labels, which are then output as refined anomaly source labels.
7. The method for detecting abnormal data in a vehicle-mounted mobile pumping station according to claim 1, characterized in that, The process involves calculating the feature distribution density of the refined anomaly source tags, acquiring a real-time data stream, comparing it with the feature distribution density to obtain a feedback deviation value, and if the feedback deviation value exceeds a preset deviation threshold, adjusting the grouping parameters in the clustering process and re-clustering the initial feature vectors to obtain optimized grouped data clusters, including: Extract operational features from the refining anomaly source tags and calculate the feature distribution density of the operational features; Acquire real-time data stream, compare the real-time data stream with the feature distribution density, and calculate the feedback deviation value; If the feedback deviation value is greater than the preset deviation threshold, the weight parameters of the distance metric function in the clustering process are adjusted, and the adjusted distance metric function is used to re-cluster the initial feature vector to generate optimized grouped data clusters.
8. The method for detecting abnormal data in a vehicle-mounted mobile pumping station according to claim 1, characterized in that, The process involves using the optimized grouped data clusters to perform anomaly scoring on the new real-time data stream to obtain an attribution probability. If the attribution probability exceeds a preset warning threshold, a fault warning signal is generated, including: Extract key metrics from the new real-time data stream to form real-time data points; Calculate the Euclidean distance between the real-time data point and the cluster center vector of each cluster in the optimized grouped data cluster; Anomaly scores are obtained using a preset local density calculation formula based on the Euclidean distance. The probability of attribution is calculated using a preset logistic regression function for the abnormal scores. If the attribution probability is greater than the preset warning threshold, a fault warning signal is generated.
9. A vehicle-mounted mobile pumping station abnormal data detection system, characterized in that, include: The feature extraction module is used to acquire the real-time vibration sequence and pressure fluctuation sequence of the vehicle-mounted mobile pump station, and to preprocess the real-time vibration sequence and pressure fluctuation sequence to obtain a preliminary feature vector. The clustering and grouping module is used to perform clustering processing based on the preliminary feature vectors to generate grouped data clusters; An anomaly detection module is used to perform probability mapping on the grouped data clusters using a pre-trained classifier to obtain anomaly probability scores; The source initial judgment module is used to extract the contribution of each feature in the corresponding classification result if the anomaly probability score exceeds a preset score threshold, sort them in descending order of contribution to obtain a high significance feature sequence, and match the high significance feature sequence with a preset mapping dictionary to obtain a preliminary anomaly label. The source refinement module is used to calculate the confidence level of the preliminary anomaly label. If the confidence level exceeds a preset confidence threshold, the preliminary anomaly label is refined into anomaly cause labels according to a preset level, which are used as refined anomaly source labels. The model optimization module is used to calculate the feature distribution density of the refined anomaly source label, obtain the real-time data stream and compare it with the feature distribution density to obtain the feedback deviation value. If the feedback deviation value exceeds the preset deviation threshold, the grouping parameters in the clustering process are adjusted and the initial feature vector is re-clustered to obtain the optimized grouped data cluster. The early warning triggering module is used to score the new real-time data stream for anomalies using the optimized grouped data clusters to obtain the attribution probability. If the attribution probability exceeds the preset early warning threshold, a fault early warning signal is generated.