Park equipment abnormity monitoring method and system based on big data
By using a big data-based equipment anomaly monitoring method, a baseline mode for equipment operation is generated and combined with historical data and external feedback data. This solves the problems of high false alarm rate and inaccurate monitoring in existing technologies, and achieves accurate perception and highly reliable monitoring of equipment status.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU YUNWU INTELLIGENT TECH CO LTD
- Filing Date
- 2025-12-16
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies for monitoring equipment anomalies have a high false alarm rate and are difficult to accurately identify potential hazards. They are also unable to adapt to equipment performance drift and changes in operating conditions, resulting in low monitoring sensitivity and frequent false alarms.
The big data-based method for monitoring abnormal equipment in industrial parks acquires multi-source equipment operation data, performs time alignment and feature transformation to generate equipment operation baseline patterns, combines historical operation record databases for real-time comparison and dynamic threshold determination, calculates the probability distribution of anomalies and generates structured abnormal event logs, and introduces external feedback data for integrated analysis and report generation.
It enables adaptive and accurate perception of the dynamic operating status of equipment, improves the reliability and anti-interference ability of anomaly monitoring results, and ensures the accuracy of monitoring reports and business availability.
Smart Images

Figure CN121997164A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial internet and big data analytics, and in particular to a method and system for monitoring abnormal equipment in industrial parks based on big data. Background Technology
[0002] Currently, in the intelligent management system of modern industrial parks, real-time monitoring of equipment operating status and fault early warning are core links to ensure production continuity and safety. With the advancement of Industry 4.0, the scale of equipment in the park is growing exponentially, generating a huge amount of operational data from complex sources. How to accurately identify abnormal signals from this massive amount of data is directly related to the operation and maintenance of enterprises.
[0003] In existing technologies, equipment monitoring solutions typically rely on microcontrollers (MCUs) or traditional PLC systems deployed at the edge for basic data acquisition and judgment. These systems mostly employ monitoring logic based on fixed rules or preset static thresholds; that is, an alarm is triggered when the acquired parameters (such as temperature or vibration frequency) exceed pre-set upper and lower limits. This one-size-fits-all static strategy ignores the performance drift caused by aging and wear during long-term operation, as well as the dynamic impact of different operating conditions (such as load changes and fluctuations in ambient temperature and humidity) on the normal operating parameters of the equipment. Lacking the ability to deeply analyze the periodic patterns and long-term trends behind massive amounts of time-series data, the system cannot accurately distinguish between normal parameter fluctuations and potential fault symptoms. Consequently, when facing complex and ever-changing operating environments, it often misses key abnormal signals due to the inability to effectively integrate information, or frequently generates false alarms due to rigid threshold settings.
[0004] Therefore, existing technologies suffer from high false alarm rates in anomaly monitoring and difficulty in accurately identifying potential hazards. Summary of the Invention
[0005] This invention provides a method and system for monitoring abnormal equipment in a park based on big data, in order to solve the technical problems of high false alarm rate and difficulty in accurately identifying potential hazards in the existing technology.
[0006] Firstly, in order to solve the above-mentioned technical problems, the present invention provides a method for monitoring abnormal equipment in a park based on big data, comprising: The system acquires multi-source equipment operation data from the industrial park and performs time alignment and feature transformation processing to obtain time-series feature segments. Based on the time-series feature segments, periodic pattern recognition and trend fitting are performed to generate a device operation baseline mode, and the time-series feature segments are stored in a preset historical operation record library. Obtain the device operation data at the current moment and construct a feature vector. Compare the feature vector with the device operation baseline mode in real time and make dynamic threshold determination to obtain a set of potential mutation points. Based on the set of potential mutation points and the historical operation record library, historical backtracking and pattern analysis are performed to calculate the anomaly probability distribution and obtain the anomaly confidence assessment value. Based on the anomaly confidence assessment value, risk classification is determined in conjunction with the preset risk level standard, a structured anomaly event log is generated, and an early warning is triggered when the risk is determined to be high. Obtain external feedback data, integrate and analyze the structured anomaly logs and the external feedback data, generate a report, and output the final monitoring report.
[0007] Secondly, the present invention provides a big data-based system for monitoring abnormal equipment in a park, comprising: The data preprocessing module is used to acquire multi-source equipment operation data from the industrial park and perform time alignment and feature transformation processing to obtain time-series feature segments. The benchmark construction module is used to perform periodic pattern recognition and trend fitting processing based on the time-series feature segments, generate a device operation benchmark mode, and store the time-series feature segments in a preset historical operation record library. The real-time monitoring module is used to acquire the current equipment operation data and construct a feature vector. The feature vector is then compared with the equipment operation baseline mode in real time and a dynamic threshold is determined to obtain a set of potential mutation points. The evaluation and analysis module is used to perform historical backtracking and pattern analysis based on the set of potential mutation points and the historical operation record library, calculate the anomaly probability distribution, and obtain the anomaly confidence evaluation value. The early warning decision module is used to determine the risk level based on the anomaly confidence assessment value and the preset risk level standard, generate a structured abnormal event log, and trigger an early warning when the risk level is determined to be high. The report generation module is used to acquire external feedback data, integrate and analyze the structured abnormal event logs and the external feedback data, and generate a final monitoring report.
[0008] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention generates a benchmark mode for equipment operation by performing periodic pattern recognition and trend fitting on time-series feature segments, and compares the current feature vector with the benchmark mode in real time and makes dynamic threshold determination. This method abandons the rigid static threshold determination logic in traditional monitoring and constructs a dynamic reference system that can simultaneously characterize the long-term aging trend and short-term periodic fluctuation law of equipment. It solves the technical problems of high false alarm rate and low monitoring sensitivity caused by the inability of existing technology to adapt to equipment performance drift or cyclical changes in operating conditions, and realizes adaptive and accurate perception of the dynamic operating status of equipment.
[0009] (2) This invention combines historical operation record database to conduct historical retrospective analysis of potential mutation point set, calculates the probability distribution of anomalies and obtains anomaly confidence assessment value; this "initial screening-retrospective-precision judgment" secondary verification mechanism can distinguish between occasional random noise and real anomalies with systemic risks from the perspective of probability statistics; it solves the technical problem that the existing technology is easily affected by instantaneous interference and frequently triggers invalid alarms, and improves the credibility and anti-interference ability of anomaly monitoring results.
[0010] (3) This invention acquires external feedback data and integrates it with structured abnormal event logs for analysis and report generation. This step introduces a closed-loop feedback verification logic of "human-machine collaboration" and uses feedback information from actual operation and maintenance to clean and confirm the logs generated by the algorithm. This solves the problem that the pure data-driven model in the prior art lacks realistic verification methods, resulting in the accumulation of invalid information and difficulty in continuously optimizing the model accuracy, and ensures that the final output monitoring report has high business availability and accuracy. Attached Figure Description
[0011] Figure 1 This is a schematic diagram of the process of the big data-based abnormal monitoring method for park equipment provided in the first embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a big data-based campus equipment anomaly monitoring system provided in the second embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] Reference Figure 1 The first embodiment of the present invention provides a method for monitoring abnormal equipment in a park based on big data, including the following steps: S11: Obtain multi-source equipment operation data from the industrial park, and perform time alignment and feature transformation processing to obtain time-series feature segments; S12, perform periodic pattern recognition and trend fitting processing based on the time-series feature segments to generate a device operation baseline mode, and store the time-series feature segments in a preset historical operation record library; S13, acquire the device operation data at the current moment and construct a feature vector, compare the feature vector with the device operation baseline mode in real time and determine the dynamic threshold to obtain a set of potential mutation points; S14. Based on the set of potential mutation points and the historical operation record library, perform historical backtracking and pattern analysis, calculate the anomaly probability distribution, and obtain the anomaly confidence assessment value. S15. Based on the anomaly confidence assessment value and combined with the preset risk level standard, risk classification is determined, a structured anomaly event log is generated, and an early warning is triggered when the risk is determined to be high. S16, Obtain external feedback data, integrate and analyze the structured abnormal event log and the external feedback data, and generate a report to output the final monitoring report.
[0014] In step S11, multi-source equipment operation data from the industrial park is acquired and processed for time alignment and feature transformation to obtain time-series feature segments, including: The system collects multi-source sensor signals from the industrial park and performs timestamp alignment processing to obtain the raw operational data stream. According to a preset time step, the original running data stream is divided into multiple continuous discrete data subsets by a sliding window. For each of the discrete data subsets, multidimensional statistical features are calculated to construct a high-dimensional statistical feature vector; Based on the high-dimensional statistical feature vector, feature projection and compression are performed to obtain the time-series feature segment.
[0015] In one implementation, this embodiment acquires the multi-source sensor signals through a sensor array deployed at key locations of equipment in an industrial park. To address the issue of significant differences in sampling frequencies among different sensors—for example, vibration signals with frequencies above 2kHz while temperature signals have frequencies below 1Hz—this embodiment employs a feature-first alignment strategy. Specifically, for high-frequency sensors, edge computing is performed using an IoT chip integrated into the sensor to calculate time-domain metrics within a short time window.
[0016] It should be noted that the length of this short window is set to 100ms. This value is chosen because its corresponding 10Hz frequency can meet the low-frequency envelope requirements of most mechanical fault characteristics of park equipment (such as imbalance and misalignment), and can also effectively reduce data transmission bandwidth.
[0017] The conversion process specifically involves collecting all high-frequency sampling points within the 100ms window, calculating their root mean square or peak values, and using the calculated single value as the low-frequency feature value corresponding to that window, thereby forming a low-frequency feature stream that matches the sampling rate of the low-frequency sensor. For low-frequency sensors, the original values are directly collected. Subsequently, this embodiment sets a globally unified master sampling clock frequency, such as 10Hz, and resamples the aforementioned low-frequency feature stream and low-frequency sensor data to a unified time point. During this process, differentiated resampling strategies are adopted for different types of sensor data. For continuously changing analog data such as temperature and pressure, a linear interpolation algorithm is used to calculate intermediate values; for discretely changing digital data such as switch states and alarm bits, a zero-order hold strategy is adopted, that is, the value at the previous sampling time is held until the change at the next sampling time. The linear interpolation calculation of analog quantities follows the following formula: in, Indicates the target time point The interpolation results, and These represent the points in the original data stream located at the target time points. Two adjacent sampling time points, and These represent the points in time. and The corresponding original collected values.
[0018] It is worth noting that, during the above linear interpolation calculations, although the input is already a low-frequency stream, data transmission jitter or clock synchronization errors may still cause variations in the denominator. The value approaches zero, potentially causing division by zero anomalies or numerical instability. This embodiment incorporates a division-by-zero protection logic, which verifies the time difference before calculation. Is it greater than the preset minute amount? ,For example If it is less than this value, then take it directly. As the interpolation result, all aligned data sequences constitute the original running data stream.
[0019] In one implementation, this embodiment divides the original running data stream into sliding window segments according to a preset time step. This segmentation operation uses a fixed-length time window. and a sliding step It should be noted that the window parameters are determined through offline autocorrelation function analysis. This determination process includes selecting a historical data sequence under stable operating conditions during the initial operation phase of the equipment or obtaining it from a preset database, calculating the correlation coefficient between this sequence and the equipment itself at different lag times, then plotting the correlation coefficient as a function of lag time, identifying the first maximum point on the curve, and determining the lag time corresponding to this maximum point as the inherent operating cycle of the equipment. Finally, the time window will be... Set as this period The value is twice that of the window, and this multiple is set based on the minimum requirement of Shannon's sampling theorem for complete coverage of periodic signals, thereby ensuring that the window contains complete periodic information.
[0020] It is worth noting that for equipment with variable frequency operation characteristics, its operating cycle is not a fixed value. This embodiment dynamically adjusts the cycle by real-time acquisition of the equipment's speed or frequency control signal. The adjustment process involves scaling the reference period inversely based on the ratio of the current speed to the rated speed, i.e.: The current dynamic period is obtained. Subsequently, this embodiment uses a dynamic time warping algorithm to perform nonlinear phase alignment. Specifically, the alignment process involves calculating the Euclidean distance matrix between the data sequence within the current dynamic period and the reference period template sequence, using dynamic programming to find a warped path with the minimum cumulative distance from the upper left corner to the lower right corner of the matrix, and mapping the time axis of the current data sequence onto the reference time axis according to this path to achieve phase unification.
[0021] For each segmented discrete data subset, this embodiment calculates its time-domain statistics for each dimension. These time-domain statistics include the mean reflecting central tendency, the standard deviation reflecting dispersion, and the peak-to-peak value reflecting the range of fluctuation. This embodiment concatenates these statistics to construct a single, undimension-reduced numerical vector, which is the high-dimensional statistical feature vector.
[0022] It is worth noting that feature projection and compression are performed based on the high-dimensional statistical eigenvectors. This processing employs principal component analysis (PCA) algorithm. First, this embodiment collects historical data to construct the original feature matrix. And calculate the mean and standard deviation of each column to construct a mean vector. and standard deviation vector This embodiment is for Z-score standardization is performed dimension-wise, which involves subtracting the mean from each column of the matrix and dividing by the standard deviation using vector operations. Subsequently, the covariance matrix is calculated and eigenvalue decomposition is performed. In this embodiment, the cumulative contribution rate threshold is determined based on the scree map. For example, 95%. This threshold is set based on the signal-to-noise ratio characteristics of industrial sensor data. Typically, the first 95% of the variance contains the main physical laws governing equipment operation, while the remaining 5% usually corresponds to high-frequency measurement noise or random environmental interference. Discarding this portion helps improve the robustness of the model. The specific decision-making logic involves calculating the difference between adjacent feature values. And set a smooth threshold, such as the first feature value. 5%. When When the value first falls below the gradual threshold, the location is identified as the elbow point. A projection matrix is constructed by selecting this point and all feature vectors preceding it. This embodiment multiplies the current high-dimensional statistical feature vector by... The dimensionality-reduced vectors are obtained. In order to form temporal samples for subsequent modeling, this embodiment arranges the dimensionality-reduced vectors from multiple consecutive time points in chronological order to form the temporal feature segments.
[0023] It should be noted that, to address the cold start issue during the initial system deployment, this embodiment employs a general projection matrix pre-trained on similar devices during the period lacking historical data. This general projection matrix is obtained from a pre-established cloud-based device knowledge base, which aggregates historical operating data from multiple devices of the same model and specifications under standard operating conditions, and is trained offline through the aforementioned principal component analysis process. The system uses this general matrix to initiate operation and data accumulation. When the accumulated historical data reaches a preset minimum sample size, for example, covering at least one complete monthly operating cycle, the system triggers a localized principal component analysis process, recalculating the covariance matrix and eigenvectors using the locally accumulated data to generate a new projection matrix. It also replaces the original general matrix, completing the localization update of the model.
[0024] It should be noted that the system has a model drift monitoring mechanism. When the mean of the reconstructed residuals shows a slow, monotonous trend change that lasts for more than the aging confirmation window, it is determined to be aging drift and the aforementioned PCA recalculation process is triggered. The aging confirmation window is preset based on the thermal time constant of the equipment, specifically calculated to be 3 to 5 times the time required for the equipment to reach thermal equilibrium (for example, if the equipment's thermal equilibrium time is 2 days, the window is set to 7 days). This filters out reversible parameter drift caused by diurnal variations or seasonal fluctuations in ambient temperature, ensuring that the monitored trend reflects irreversible physical aging. If the residuals show a step-like abrupt change, it is determined to be a potential fault, and the model update is not triggered.
[0025] In step S12, based on the time-series feature segments, periodic pattern recognition and trend fitting processing are performed to generate a device operation baseline mode, and the time-series feature segments are stored in a preset historical operation record library, including: The time series feature segments are subjected to time series component decomposition calculation to separate the long-term trend component, periodic fluctuation component and random residual component; Numerical fitting and smoothing are performed on the long-term trend components to generate a long-term trend baseline curve. Waveform clustering and feature extraction are performed on the periodic fluctuation components to determine typical periodic fluctuation templates; The long-term trend benchmark curve is superimposed and reconstructed with the typical periodic fluctuation template to generate the equipment operation benchmark mode.
[0026] In one implementation, this embodiment performs time series component decomposition calculation on the time series feature segment. Given that the time series feature segment is generated by S11 and includes... For multivariate time series with multiple independent feature dimensions (i.e., principal components), this embodiment employs a dimension-by-dimensional decomposition strategy to accurately capture the independent evolution of each feature dimension. Specifically... For each independent dimension sequence in the given dimensions, this embodiment applies the STL (Seasonal-Trend decomposition using Loess) decomposition algorithm. This algorithm is based on an additive model and assumes a single-dimensional sequence. By trend item Seasonal terms (periodic terms) and residuals The core of this algorithm lies in its use of LOESS (Locally Weighted Regression) for iterative smoothing.
[0027] Specifically, this embodiment sets the seasonal smoothing window length. The inherent operating cycle determined for S11 The period length is an odd multiple (e.g., 7 times). Choosing an odd number ensures the filter's symmetry and avoids introducing phase shift; choosing 7 times the period length allows for the slow evolution of the seasonal component over time while smoothing out short-term random fluctuations, thus achieving a balance between stability and adaptability. After this step, the system obtains... The group corresponds to the long-term trend component, the periodic fluctuation component, and the random residual component.
[0028] In another implementation, the long-term trend components are numerically fitted and smoothed. Considering that multinomial models are prone to numerical divergence (Runge phenomenon) during long-term extrapolation, causing the baseline value to deviate from physical reality, this embodiment preferentially uses the Holt-Winters double exponential smoothing algorithm to model the trend components. This algorithm includes a level term. and trend items Its iterative formula is: It should be noted that, in order to avoid the subjectivity of manually setting parameters, the smoothing coefficient (horizontal smoothing coefficient) is... and trend smoothing coefficient The value is determined through an automated optimization process. This embodiment utilizes the L-BFGS-B (Limited-memory Broyden–Fletcher–Goldfarb–Shanno with Box constraints) numerical optimization algorithm to construct an optimization problem with the sum of squared single-step prediction errors (SSE) of historical data of the long-term trend components stored in the historical running record library as the objective function. It searches for the optimal coefficient combination that minimizes the SSE within the constraint interval [0,1]. Based on this optimal coefficient, the algorithm dynamically updates the level term and trend term, thereby generating the long-term trend baseline curve that can robustly predict the future.
[0029] It is worth noting that waveform clustering and feature extraction are performed on the periodic fluctuation components. To preserve the phase correlation and coupling between different feature dimensions, this embodiment employs a multi-dimensional state slicing approach. First, based on the operating period determined in S11... , will contain The periodic fluctuation component of each dimension is divided into multiple lengths. Multidimensional waveform segments. For each segment, it is derived from... The matrix structure is flattened to a length of The flattened vectors are then grouped using the K-Means clustering algorithm. It should be noted that the number of clusters in the K-Means algorithm... It is determined through silhouette coefficient analysis. This embodiment traverses a set of preset candidate... Values (e.g.) ), for each The average silhouette coefficient of all samples is calculated. The average silhouette coefficient combines intra-cluster cohesion and inter-cluster separation; a value closer to 1 indicates better clustering. This embodiment selects the sample with the highest average silhouette coefficient. The value is used as the optimal cluster number. In this embodiment, the cluster containing the most waveform segments (i.e., the main cluster) is counted, the arithmetic mean of all vectors within this cluster is calculated, and then restored to its optimal value. The matrix form is determined as the typical periodic fluctuation template.
[0030] In one implementation, the long-term trend benchmark curve is superimposed and reconstructed with the typical cyclical fluctuation template. This process is a time-dimensional vector addition operation. For any future prediction time point... This embodiment first predicts its value based on the Holt-Winters model. Trend vectors in each dimension .in The formula for predicting the step size is as follows: That is, predict the target time point With the last update time of the model The time difference between them divided by the system sampling interval At the same time, calculate the time point. Relative to the running cycle phase : And the corresponding period vector is obtained by indexing from the typical periodic fluctuation template. This embodiment calculates the synthesized reference vector. : This synthesized vector sequence constitutes the device's operating baseline mode. Simultaneously, in this embodiment, the time-series feature segments generated in S11 and their corresponding timestamps are appended to the preset historical operation record database via the time-series database's write interface.
[0031] In step S13, the device operation data at the current moment is acquired and a feature vector is constructed. The feature vector is then compared in real time with the device operation baseline mode and a dynamic threshold is determined to obtain a set of potential mutation points, including: Obtain the device operation data at the current moment, and perform feature extraction and numerical vectorization processing to construct the current state feature vector; Based on the current state feature vector, a search and matching process is performed in the device operation baseline mode to determine the corresponding associated baseline feature interval; Calculate the numerical difference between the current state feature vector and the associated baseline feature interval to obtain the real-time deviation index; The real-time deviation index is compared and analyzed with a preset dynamic safety threshold, and sampling points that exceed the dynamic safety threshold are extracted to obtain the set of potential mutation points.
[0032] In one implementation, this embodiment acquires the device's current operating data, performs feature extraction, and constructs a feature vector. To meet the requirements of high-frequency real-time monitoring, this process employs a streaming computing architecture, maintaining a sliding window in memory with the same length as defined in S11. The same real-time cache queue is used. When a new sensor data point arrives, this queue is updated using a first-in, first-out (FIFO) method. To avoid repeatedly traversing and calculating all data within the window, this embodiment employs an incremental statistical algorithm. Specifically, the system persistently maintains two intermediate state variables in memory: the sum (Sum) and the sum of squares (SumSq) of the data within the window; when a new data point arrives... Moved in and old data points When removing, these two variables are updated only through addition and subtraction operations, thus achieving... The mean and standard deviation are calculated in real time with constant time complexity, and the original high-dimensional statistical feature vector is constructed by combining indicators such as peak-to-peak value. It should be noted that, in order to ensure the consistency between the real-time vectors and the historical benchmark model in the feature space, this embodiment must load the projection matrix generated during training in S11. and the standardized parameter vector (i.e., the length is...) mean vector and standard deviation vector This embodiment first describes... Element-wise Z-score normalization is performed, and the calculation formula is as follows: in, This represents element-wise division of vectors; subsequently, a dimension reduction mapping is performed using the projection matrix: The dimension-reduced vector This is the current state feature vector.
[0033] In another implementation, a search and match is performed within the device operating baseline mode based on the current state feature vector. Since the device operating baseline mode is composed of... The multi-dimensional time-series model is constructed by superimposing a "long-term trend benchmark curve" and a "typical cyclical fluctuation template" in multiple dimensions. Therefore, this retrieval and matching is essentially based on phase alignment in the time dimension. This embodiment extracts the sampling timestamp of the current state feature vector. And perform specific benchmark synthesis calculations. First, using the data generated by S12... Group long-term trend prediction models (such as the Holt-Winters model) calculate the trend baseline vector at the current absolute time point. Secondly, calculate the current time relative to the running cycle. phase : And index the periodic reference vector corresponding to this phase in the typical periodic fluctuation template. Finally, the composite reference vector is calculated using vector addition. : The composite reference vector This is the center reference value of the associated benchmark feature interval.
[0034] It is worth noting that this embodiment calculates the numerical difference between the current state feature vector and the associated baseline feature interval. This embodiment uses Euclidean distance to quantify this difference in the multidimensional feature space. The calculation formula is: This distance value This refers to the real-time deviation index. For example, if the current dimensionality-reduced feature vector is [0.5, 1.2], and the synthesized baseline vector calculated based on the time phase is [0.4, 1.0], then the real-time deviation index... .
[0035] In one implementation, the real-time deviation index is compared and analyzed with a preset dynamic safety threshold. It should be noted that the preset dynamic safety threshold is not a fixed value, but is dynamically determined based on the statistical characteristics of the random residual components separated in S12. In this embodiment, the dynamic safety threshold is calculated in S12, including... Random residual sequence of each dimension To set the scalar threshold, this embodiment first calculates the Euclidean norm (i.e., modulus) of the residual vector at each time point. This transforms the multidimensional residual sequence into a one-dimensional residual modulus sequence. Subsequently, this embodiment performs statistical analysis on this residual modulus sequence. Although the 3-sigma principle can be used to set the threshold, considering the potentially non-normal long-tailed characteristics of industrial data residuals, this embodiment employs percentile statistics as a more robust preferred implementation. Specifically, the 99.7th percentile of the residual modulus sequence is calculated and defined as the dynamic safety threshold. .
[0036] It should be noted that the 99.7% threshold was chosen because this percentage corresponds to the confidence interval covered by three standard deviations (positive and negative) of a standard normal distribution. This range can encompass most random fluctuations under normal operating conditions, eliminating only extremely rare and abnormal events, thus significantly reducing the false alarm rate while maintaining the fault detection rate. This embodiment uses the real-time deviation index... With threshold Perform a comparison. If... If the device status within the current time window deviates significantly from the normal statistical pattern, this embodiment marks the time window feature vector corresponding to that time point as an abnormal candidate and stores it in the potential mutation point set.
[0037] In step S14, based on the set of potential mutation points and the historical operation record database, historical backtracking and pattern analysis are performed to calculate the anomaly probability distribution and obtain the anomaly confidence assessment value, including: Based on the set of potential mutation points, scene backtracking and retrieval matching are performed in the historical operation record database to extract related historical scene data; The fluctuation characteristics of the associated historical scene data are statistically analyzed and reproduced to determine the historical reproduction patterns. Based on the characteristics of the historical recurrence patterns, anomaly probability estimation is performed to generate anomaly probability distribution; By combining the aforementioned anomaly probability distribution and the set of potential mutation points, a multidimensional weighted quantization calculation is performed to obtain the anomaly confidence assessment value.
[0038] In one implementation, this embodiment performs scene backtracking and retrieval matching in the historical operation record database based on the set of potential mutation points. The historical operation record database is constructed by accumulating historical time-series feature fragments with timestamps and status tags that are continuously added in S12. This embodiment uses this record database to determine whether the currently detected mutation state has frequently occurred as a security fluctuation in the past. To address the cold start problem in the early stages of system operation, if the total amount of data in the historical operation record database is less than a preset minimum retrieval base, for example, less than... If a record is found, the system will automatically skip the search and matching, directly determine the historical recurrence rate as 0, assign the highest anomaly confidence level, and adopt the most conservative strategy to prevent false negatives. When the data volume meets the requirements, given that the historical operation record database may contain massive time window feature vectors, and the feature dimensions are... The dimensionality may be high. In order to ensure the real-time performance of the retrieval and avoid the computational delay caused by linear scanning, this embodiment pre-constructs a high-dimensional feature index structure.
[0039] As a preferred implementation, this embodiment employs HNSW (Hierarchical Navigable Small World) graph indexing or Product Quantization (PQ) technology. For the specific implementation of the high-dimensional index structure, those skilled in the art can refer to open-source libraries (such as FAISS or Annoy) to construct the HNSW graph index. For each mutation point (i.e., the time window feature vector to be verified) in the set of potential mutation points, this embodiment utilizes the index structure to quickly locate the nearest Euclidean distance point with logarithmic time complexity. The historical data points constitute the associated historical scene data. It should be noted that the number of neighbors... This was determined through cross-validation. In this embodiment, historical labeled data is divided into training and validation sets to test different... The value is used to calculate the category purity of the search results. Category purity refers to the proportion of the dominant category label (normal or faulty) among the retrieved neighbor samples. This embodiment plots the category purity as a function of... For the changing curve, calculate its first derivative, define the point where the derivative first approaches zero as the stationary inflection point, and use the value corresponding to that point as a preset value. This ensures the stability of the sample statistics.
[0040] In addition, the system has a dynamic index update mechanism that maintains an in-memory incremental buffer. When the amount of newly written data in the incremental buffer reaches a preset threshold (e.g., 1% of the full index size or a fixed number of 1000 records), the index merging and reconstruction operation is automatically triggered. This threshold is set to balance the computational overhead of index reconstruction with the timeliness of retrieval results, avoiding system lag caused by frequent reconstructions, and preventing new data from remaining unsearchable for too long; it can be dynamically adjusted according to system memory to ensure a balance between real-time performance and stability.
[0041] In another implementation, fluctuation characteristic statistics and recurrence analysis are performed on the associated historical scene data. This embodiment calculates this... The local density of historical neighboring points is used as a reproduction index. Specifically, this embodiment statistically analyzes these... Among the neighboring points, the number of samples that are historically marked as "normal operation" or "known interference" and whose deviation index is less than the deviation index of the current mutation point. The historical recurrence pattern is defined as the recurrence rate. : If the recurrence rate is high, it indicates that similar fluctuations have occurred frequently in history without causing failures, and are considered inherent background noise of the system; if the recurrence rate is low, it indicates that the fluctuations are extremely rare in history and are highly suspected of being abnormal.
[0042] It is worth noting that this embodiment performs anomaly probability estimation based on the aforementioned historical recurrence patterns, generating an anomaly probability distribution. To quantify the degree of anomaly from a statistical distribution perspective, this embodiment introduces Gaussian kernel density estimation (KDE). This embodiment uses the associated historical scene data to fit a local probability density function. The fitting process involves constructing a Gaussian kernel function centered on each historical neighbor. : Then, all kernel functions are summed and averaged. The bandwidth parameter is... The Scott's Rule is used for adaptive determination, that is: in, The number of neighbor samples determined above. The feature dimension is used. Based on this fitting function, the probability density value of the current mutation point under this distribution is calculated. To address the issue that probability density values might be greater than 1, leading to out-of-bounds calculations or unclear physical meaning, this embodiment employs an exponential decay function. Mapped to normalized density anomaly scores : in, The scaling factor, given that the data has been standardized in the previous steps, is set to 1 in this embodiment as the standard decay rate, used to map the density domain to the probability domain of [0,1]. This mapping ensures that higher density results in lower scores (approaching 0), and lower density results in higher scores (approaching 1). The final anomaly probability... The weighted average of recurrence rate and density anomaly score is calculated using the following formula: in, To balance the weights, they are determined through a grid search on the validation set. The search space is set to a closed interval [0,1] with a step size of 0.05. All candidate values within this space are traversed, and the value that maximizes the F1 score for anomaly detection is selected. The value is used as the final parameter. The value is the quantification result of the abnormal probability distribution.
[0043] In one implementation, a multidimensional weighted quantization calculation is performed by combining the anomaly probability distribution and the set of potential mutation points. This embodiment constructs a multidimensional evaluation feature vector, whose components include the anomaly probabilities calculated above. The real-time deviation index calculated in S13 and the number of time steps during which the mutation state lasts. This embodiment uses the Sigmoid scoring function to calculate the anomaly confidence assessment value. The calculation formula is: in, These are the model weight parameters. It should be noted that these model weight parameters are obtained through offline training using a logistic regression algorithm. In this embodiment, historical fault cases are collected as positive samples (Label=1), and historical false alarm cases are collected as negative samples (Label=0). The aforementioned features are extracted and iteratively solved using maximum likelihood estimation combined with an L-BFGS (Limited-memory Broyden–Fletcher–Goldfarb–Shanno) optimizer. The iteration terminates when the gradient norm of the objective function is less than a preset convergence threshold, for example... Or, it may reach the maximum number of iterations, such as 100. The convergence threshold is set based on the single-precision floating-point machine precision (Machine Epsilon, approximately...) defined by the IEEE 754 standard. The maximum number of iterations is determined in this embodiment to be at least two orders of magnitude higher than the machine precision, ensuring the stability of the parameter solution in numerical computation and filtering out the influence of truncation errors. The maximum number of iterations is determined based on the theoretical convergence order of the L-BFGS algorithm in low-dimensional convex optimization problems. This value covers the theoretical upper limit of the number of iterations required for the second-order quasi-Newton method to converge when the feature dimension is limited (3 dimensions in this embodiment). This yields the optimal weight combination. This method ensures that the evaluation value objectively reflects the contribution of each indicator to the final fault confirmation.
[0044] In step S15, based on the anomaly confidence assessment value and a preset risk level standard, a risk classification judgment is made, a structured anomaly event log is generated, and an early warning is triggered when the risk is determined to be high, including: Based on the anomaly confidence assessment value and the preset risk level standard, interval mapping and level matching are performed to determine the risk level judgment result; Based on the risk level determination result and the preset storage template, data serialization and template filling processes are performed to generate a structured abnormal event log; The risk level determination result is subjected to threshold verification. If the risk level determination result falls into the preset high-risk range, an immediate warning instruction containing the anomaly type and the time of occurrence is generated.
[0045] In one implementation, this embodiment performs interval mapping and level matching based on the anomaly confidence assessment value and the preset risk level standard. The anomaly confidence assessment value is a continuous probability value between 0 and 1, while the preset risk level standard is a lookup table containing several intervals defined by numerical boundaries and discrete level labels (such as low risk, medium risk, and high risk) corresponding to each interval. This embodiment outputs the corresponding risk level determination result by determining which specific numerical interval the anomaly confidence assessment value falls within.
[0046] It should be noted that the numerical boundaries of the preset risk level standards are determined through offline statistical analysis of the performance of the logistic regression model trained in S14 on the validation dataset. Specifically, this embodiment uses the validation set data to plot the model's precision-recall curve (PR Curve). A working point on this curve that meets high precision (e.g., Precision > 99%) is selected, and its corresponding probability threshold is set as the high-risk boundary, ensuring an extremely low false alarm rate when triggering a high-risk alarm. Simultaneously, a working point that meets high recall (e.g., Recall > 98%) is selected, and its corresponding probability threshold is set as the low-risk boundary, ensuring that no suspicious minor fluctuations are missed. The interval between these two is defined as medium risk. This threshold setting method based on statistical performance indicators effectively ensures the objectivity and business applicability of the grading standards.
[0047] In another implementation, this embodiment performs data serialization and template filling processing based on the risk level determination result and a preset storage template. This processing solidifies the instantaneous calculation result into a traceable chain of evidence. This embodiment uses JSON Schema as the preset storage template, which defines the fields that must be included, specifically including a unique event identifier generated based on a Universally Unique Identifier (UUID), the precise time of anomaly detection conforming to ISO 8601 format, the risk level determination result output in S15, a Base64 encoded snapshot of the current state feature vector (for subsequent reproduction), the anomaly confidence assessment value calculated in S14, and the associated environmental parameters collected in S11 (such as device ID and operating mode). This embodiment fills the above data into the corresponding fields of the template and serializes it into a binary stream or compact string to generate the structured anomaly event log, which is then written to a document database (such as MongoDB) or a time-series database (such as InfluxDB) for persistent storage.
[0048] It is worth noting that this embodiment performs threshold verification on the risk level determination result. If the risk level determination result falls into a preset high-risk range, an immediate warning instruction containing the anomaly type and the time of occurrence is generated. The preset high-risk range is defined as a set of risk levels equal to "high risk" or "severe". When the determination result hits this set, the system immediately assembles a lightweight message data packet as the immediate warning instruction. This instruction contains a fault code determined based on the deviation direction of the S12 baseline mode (e.g., trend exceeding limits or period loss), the current system timestamp, a priority flag set to the highest level (e.g., P0 level), and the physical location label of the device's campus grid coordinates. This embodiment pushes the instruction to the monitoring screen or the handheld terminal of the maintenance personnel through a low-latency message queue (e.g., MQTT or Kafka), thereby triggering an audible and visual alarm or work order dispatch process. For medium-risk or low-risk events, the system only logs the event without triggering an active, disruptive warning, thus achieving a tiered response.
[0049] In step S16, external feedback data is acquired, and based on the structured anomaly event log and the external feedback data, integrated analysis and report generation processing are performed to output the final monitoring report, including: The structured anomaly log and the external feedback data are subjected to time-series correlation and content matching processing to eliminate false alarm records and determine the set of anomaly events after verification. Based on the verified abnormal event set, retrieve the associated real-time operating parameters of the device, and perform multi-dimensional state fusion and fault propagation correlation analysis to generate state coupling anomaly details; Based on the details of the state coupling anomaly, and in accordance with the preset report generation rules, data aggregation and visualization rendering are performed to output the final monitoring report.
[0050] In one implementation, this embodiment obtains external feedback data from the park's maintenance work order system or manual inspection terminal through a standardized application programming interface (API). The external feedback data includes the maintenance personnel's confirmation status of specific equipment anomalies (e.g., confirmed fault, false alarm, repaired), handling results, and manually entered fault cause tags. This embodiment performs time-series correlation and content matching processing between the structured anomaly event log and the external feedback data.
[0051] It should be noted that, considering the possibility that manual data entry may lag behind system detection time, this embodiment employs a fuzzy time window matching strategy. Specifically, it uses the machine-generated timestamp of each exception log record. Based on this, a time tolerance window is set. The specific value of this time window, such as 15 minutes before and after, is determined based on the typical average operation delay of park maintenance personnel from receiving the alarm to completing on-site verification and entering the data into the system, to ensure that it covers the time offset of the vast majority of manual feedback. The system searches for data in external feedback that simultaneously meets the requirements of consistent device unique identifier and feedback entry time. In Records of these two conditions within the closed interval. If a match is successful and the feedback status is marked as a false alarm or normal fluctuation, this embodiment marks the corresponding abnormal log as invalid and removes it from the analysis set; otherwise, if it is marked as confirmed or pending processing, the record is retained and a manual label is added. The set of all cleaned and labeled log records is determined as the verified abnormal event set.
[0052] In another implementation, this embodiment retrieves the associated real-time operating parameters of the equipment based on the verified abnormal event set and performs multi-dimensional state fusion and fault propagation correlation analysis. This embodiment utilizes the timestamps recorded in the verified abnormal event set to perform a backtracking retrieval in the original operating data stream described in S11, extracting all sensor reading sequences (such as motor speed, bearing temperature, and input voltage) within a preset time period (e.g., one hour before and after) before and after the time of the abnormality occurrence as real-time operating parameters. To analyze the physical propagation impact of the fault, this embodiment introduces a pre-configured physical topology diagram of the equipment in the database. This topology diagram is constructed during the initial system deployment phase by manually importing a standard topology configuration file, such as JSON or XML format, pre-entered based on the actual physical connections and logical dependencies of the equipment in the park. This topology diagram is stored in an adjacency list or JSON format, defining the upstream power supply, downstream load, or fluid pipeline connection relationships between equipment within the park.
[0053] The specific implementation of fault propagation correlation analysis involves first locating the currently malfunctioning device node in the topology graph and traversing all its directly adjacent nodes, such as upstream power sources or downstream pumps. Then, the key parameter sequence of the currently malfunctioning device is extracted. Key parameter sequences within the same time window as adjacent devices Next, the Pearson correlation coefficient between the two sets of sequences was calculated. The calculation formula is: in, These are the means of the sequences. Finally, the absolute value of the correlation coefficient is calculated. The correlation exceeds a preset strong correlation threshold, such as 0.85. This threshold is set based on historical coupling statistics under normal operating conditions between devices. Specifically, the system calculates the baseline correlation coefficient distribution of adjacent device parameters during historical fault-free periods and uses the 95th percentile of this distribution as the baseline. The 95th percentile is chosen based on the significance level in statistics. This quantile represents a reasonable statistical upper bound for fluctuations in background correlation between devices under normal process control, covering most linkage situations under non-fault conditions. A preset redundancy, such as 0.05, is added to determine the final threshold. The 0.05 threshold is set as a safety margin to filter out critical situations where the correlation coefficient slightly exceeds the statistical boundary due to sensor measurement noise or transient disturbances, effectively distinguishing between normal process linkage and abnormal fault impacts. If the threshold is exceeded, it is determined that the fault has a physical propagation or linkage effect across devices, and the IDs of relevant neighboring devices and their corresponding parameter curves are associated and stored together. In this embodiment, the above-mentioned real-time operating parameters, topology propagation characteristics, and original anomaly logs are structurally merged to generate the state coupling anomaly details containing rich contextual information.
[0054] It is worth noting that this embodiment performs data aggregation and visualization rendering based on the state coupling anomaly details and according to preset report generation rules. The preset report generation rules include multidimensional statistics and Pareto analysis. Specifically, the system first counts and statistically analyzes anomalies by risk level, then sorts them by frequency according to fault type code and calculates the cumulative percentage, identifying the top 20% of critical fault types that cause 80% of fault frequencies, i.e., following the Pareto principle. This embodiment uses an industrial-grade template engine (such as Jinja2) to populate the above statistical data, parameter waveforms of critical faults (including anomaly point annotations), and fault propagation chain diagrams generated based on the topology graph into a predefined HTML or PDF report template. The final rendered electronic document containing a complete data evidence chain and statistical conclusions is the final monitoring report.
[0055] In summary, this invention, by constructing a dynamic benchmark model that includes long-term trends and periodic patterns, overcomes the limitations of traditional static threshold monitoring. Combined with historical backtracking verification and an external feedback closed-loop mechanism, it effectively distinguishes between normal equipment fluctuations and real fault symptoms, achieving adaptive perception and high-precision anomaly early warning for the operating status of massive amounts of equipment in the park.
[0056] Reference Figure 2 The second embodiment of the present invention provides a campus equipment anomaly monitoring system based on big data, comprising: The data preprocessing module is used to acquire multi-source equipment operation data from the industrial park and perform time alignment and feature transformation processing to obtain time-series feature segments. The benchmark construction module is used to perform periodic pattern recognition and trend fitting processing based on the time-series feature segments, generate a device operation benchmark mode, and store the time-series feature segments in a preset historical operation record library. The real-time monitoring module is used to acquire the current equipment operation data and construct a feature vector. The feature vector is then compared with the equipment operation baseline mode in real time and a dynamic threshold is determined to obtain a set of potential mutation points. The evaluation and analysis module is used to perform historical backtracking and pattern analysis based on the set of potential mutation points and the historical operation record library, calculate the anomaly probability distribution, and obtain the anomaly confidence evaluation value. The early warning decision module is used to determine the risk level based on the anomaly confidence assessment value and the preset risk level standard, generate a structured abnormal event log, and trigger an early warning when the risk level is determined to be high. The report generation module is used to acquire external feedback data, integrate and analyze the structured anomaly logs and the external feedback data, and generate a final monitoring report. It should be noted that the big data-based park equipment anomaly monitoring system provided in this embodiment executes all the process steps of the big data-based park equipment anomaly monitoring method described in the above embodiment. The working principles and beneficial effects of both correspond one-to-one, and therefore will not be elaborated further.
[0057] This invention also provides an electronic device. The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a big data-based campus equipment anomaly monitoring program. When the processor executes the computer program, it implements the steps described in the various big data-based campus equipment anomaly monitoring method embodiments above, for example... Figure 1 The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above system embodiments, such as the data preprocessing module.
[0058] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.
[0059] The electronic device may be a desktop computer, laptop, handheld computer, or smart tablet, etc. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the above components are merely examples of electronic devices and do not constitute a limitation on the electronic device. It may include more or fewer components than described above, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.
[0060] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the electronic device, connecting all parts of the electronic device via various interfaces and lines.
[0061] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0062] Wherein, if the modules / units integrated in the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0063] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the system embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0064] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for monitoring abnormal equipment in a park based on big data, characterized in that, include: The system acquires multi-source equipment operation data from the industrial park and performs time alignment and feature transformation processing to obtain time-series feature segments. Based on the time-series feature segments, periodic pattern recognition and trend fitting are performed to generate a device operation baseline mode, and the time-series feature segments are stored in a preset historical operation record library. Obtain the device operation data at the current moment and construct a feature vector. Compare the feature vector with the device operation baseline mode in real time and make dynamic threshold determination to obtain a set of potential mutation points. Based on the set of potential mutation points and the historical operation record library, historical backtracking and pattern analysis are performed to calculate the anomaly probability distribution and obtain the anomaly confidence assessment value. Based on the anomaly confidence assessment value, risk classification is determined in conjunction with the preset risk level standard, a structured anomaly event log is generated, and an early warning is triggered when the risk is determined to be high. Obtain external feedback data, integrate and analyze the structured anomaly logs and the external feedback data, generate a report, and output the final monitoring report.
2. The method for monitoring abnormal equipment in a park based on big data according to claim 1, characterized in that, The process involves acquiring multi-source equipment operation data from the industrial park, performing time alignment and feature transformation processing to obtain time-series feature segments, including: The system collects multi-source sensor signals from the industrial park and performs timestamp alignment processing to obtain the raw operational data stream. According to a preset time step, the original running data stream is divided into multiple continuous discrete data subsets by a sliding window. For each of the discrete data subsets, multidimensional statistical features are calculated to construct a high-dimensional statistical feature vector; Based on the high-dimensional statistical feature vector, feature projection and compression are performed to obtain the time-series feature segment.
3. The method for monitoring abnormal equipment in a park based on big data according to claim 1, characterized in that, The step of performing periodic pattern recognition and trend fitting processing based on the time-series feature segments to generate a device operation baseline mode, and storing the time-series feature segments in a preset historical operation record library, includes: The time series feature segments are subjected to time series component decomposition calculation to separate the long-term trend component, periodic fluctuation component and random residual component; Numerical fitting and smoothing are performed on the long-term trend components to generate a long-term trend baseline curve. Waveform clustering and feature extraction are performed on the periodic fluctuation components to determine typical periodic fluctuation templates; The long-term trend benchmark curve is superimposed and reconstructed with the typical periodic fluctuation template to generate the equipment operation benchmark mode.
4. The method for monitoring abnormal equipment in a park based on big data according to claim 1, characterized in that, The process involves acquiring the device's current operating data and constructing a feature vector, then comparing the feature vector with the device's operating baseline mode in real time and performing dynamic threshold determination to obtain a set of potential mutation points, including: Obtain the device operation data at the current moment, and perform feature extraction and numerical vectorization processing to construct the current state feature vector; Based on the current state feature vector, a search and matching process is performed in the device operation baseline mode to determine the corresponding associated baseline feature interval; Calculate the numerical difference between the current state feature vector and the associated baseline feature interval to obtain the real-time deviation index; The real-time deviation index is compared and analyzed with a preset dynamic safety threshold, and sampling points that exceed the dynamic safety threshold are extracted to obtain the set of potential mutation points.
5. The method for monitoring abnormal equipment in a park based on big data according to claim 1, characterized in that, The step involves performing historical backtracking and pattern analysis based on the set of potential mutation points and the historical operation record database to calculate the anomaly probability distribution and obtain the anomaly confidence assessment value, including: Based on the set of potential mutation points, scene backtracking and retrieval matching are performed in the historical operation record database to extract related historical scene data; The fluctuation characteristics of the associated historical scene data are statistically analyzed and reproduced to determine the historical reproduction patterns. Based on the characteristics of the historical recurrence patterns, anomaly probability estimation is performed to generate anomaly probability distribution; By combining the aforementioned anomaly probability distribution and the set of potential mutation points, a multidimensional weighted quantization calculation is performed to obtain the anomaly confidence assessment value.
6. The method for monitoring abnormal equipment in a park based on big data according to claim 1, characterized in that, The step of determining risk classification based on the anomaly confidence assessment value and a preset risk level standard, generating a structured anomaly event log, and triggering an alert when the risk is determined to be high includes: Based on the anomaly confidence assessment value and the preset risk level standard, interval mapping and level matching are performed to determine the risk level judgment result; Based on the risk level determination result and the preset storage template, data serialization and template filling processes are performed to generate a structured abnormal event log; The risk level determination result is subjected to threshold verification. If the risk level determination result falls into the preset high-risk range, an immediate warning instruction containing the anomaly type and the time of occurrence is generated.
7. The method for monitoring abnormal equipment in a park based on big data according to claim 1, characterized in that, The process involves acquiring external feedback data, integrating and analyzing the structured anomaly logs with the external feedback data, generating a report, and outputting a final monitoring report, including: The structured anomaly log and the external feedback data are subjected to time-series correlation and content matching processing to eliminate false alarm records and determine the set of anomaly events after verification. Based on the verified abnormal event set, retrieve the associated real-time operating parameters of the device, and perform multi-dimensional state fusion and fault propagation correlation analysis to generate state coupling anomaly details; Based on the details of the state coupling anomaly, and in accordance with the preset report generation rules, data aggregation and visualization rendering are performed to output the final monitoring report.
8. A big data-based system for monitoring abnormal equipment in a park, characterized in that, include: The data preprocessing module is used to acquire multi-source equipment operation data from the industrial park and perform time alignment and feature transformation processing to obtain time-series feature segments. The benchmark construction module is used to perform periodic pattern recognition and trend fitting processing based on the time-series feature segments, generate a device operation benchmark mode, and store the time-series feature segments in a preset historical operation record library. The real-time monitoring module is used to acquire the current equipment operation data and construct a feature vector. The feature vector is then compared with the equipment operation baseline mode in real time and a dynamic threshold is determined to obtain a set of potential mutation points. The evaluation and analysis module is used to perform historical backtracking and pattern analysis based on the set of potential mutation points and the historical operation record library, calculate the anomaly probability distribution, and obtain the anomaly confidence evaluation value. The early warning decision module is used to determine the risk level based on the anomaly confidence assessment value and the preset risk level standard, generate a structured abnormal event log, and trigger an early warning when the risk level is determined to be high. The report generation module is used to acquire external feedback data, integrate and analyze the structured abnormal event logs and the external feedback data, and generate a final monitoring report.