Industrial environment monitoring and accident prediction method fusing multi-modal data

Through dynamic knowledge graphs and deep learning models combined with digital twin-driven reinforcement learning, the problems of strong subjectivity, false alarms and omissions in traditional industrial environment monitoring are solved, and the theoretical optimality and practical feasibility of efficient identification of early weak anomalies and preventive interventions are achieved.

CN120493531AActive Publication Date: 2025-08-15SHANGHAI YUNLIN COMM TECH CO LTD

Patent Information

Application Number
CN202510594398.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-15
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

Traditional industrial environmental monitoring and accident prevention methods have problems such as strong subjectivity, lagging response, false alarms or missed alarms, and difficulty in identifying early weak abnormalities and complex faults.

Method used

Dynamic knowledge graphs are used to semantic collection and causal correlation preprocessing of multimodal heterogeneous data, combined with deep learning models to extract deep features and perform cognitive fusion, build a hybrid intelligent prediction engine, use digital twin-driven reinforcement learning to find preventive intervention strategies, and realize continuous system evolution through edge-end-cloud collaboration closed-loop feedback.

Benefits of technology

It significantly improves the perceived sensitivity and identification accuracy of early, weak and complex abnormal states in the industrial environment, ensures the theoretical optimization and practical feasibility of preventive interventions, and forms an intelligent monitoring system that is dynamically adaptable and continuous learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493531A_ABST
    Figure CN120493531A_ABST
Patent Text Reader

Abstract

The invention provides an industrial environment monitoring and accident prediction method fusing multi-modal data, and relates to the technical field of data processing, and the method comprises the steps: carrying out the semantic collection and causal association preprocessing of multi-modal heterogeneous data collected in real time through constructing a dynamic industrial knowledge graph; a customized deep learning model is adopted to extract deep abstract features of each mode, and weak signals and potential risks are accurately represented and uncertainty is quantified; a high-fidelity digital twin model is utilized to drive a deep reinforcement learning algorithm, and dynamic optimization and verification are performed to generate a multi-level and multi-target preventive intervention strategy combination; an intervention strategy is executed through an edge-end-cloud three-layer collaborative intelligent architecture, and online learning and system sustainable evolution are realized by using a closed-loop data feedback mechanism. According to the method, the sensing and early warning capability of the early weak and complex abnormal state of the industrial environment can be remarkably improved, the accident evolution path is accurately predicted, and credible explanation is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to an industrial environment monitoring and accident prediction method integrating multimodal data. Background Art

[0002] Traditional industrial environment monitoring and accident prevention methods mainly rely on the following aspects: First, manual inspection and judgment based on operator experience. This method is highly subjective, labor-intensive, and has a delayed response, making it difficult to cope with the massive amount of information and rapid changes in complex industrial systems; Second, automated monitoring and alarms based on distributed control systems and safety instrument systems. Such systems mainly monitor single or a small number of key process parameters by setting fixed thresholds. Although they can achieve basic process control and safety interlocks, they have significant deficiencies in predictability and comprehensiveness, are prone to a large number of false alarms or missed alarms, and are difficult to identify The gradual process of early weak anomalies and complex failures; third, rule-based expert systems, which solidify the experiential knowledge of domain experts into an IF-THEN rule base for reasoning and judgment, but such systems are difficult to acquire knowledge, have high updating and maintenance costs, and have limited rule coverage, making them difficult to adapt to changing working conditions and unknown risk scenarios; fourth, traditional statistical process control and single physical quantity monitoring technologies, these methods have limited capabilities in processing multi-source heterogeneous data fusion, capturing complex nonlinear correlations, and understanding deep causal mechanisms, making it difficult to form a comprehensive, dynamic, and accurate understanding of the overall operating status and potential risks of industrial systems. Summary of the Invention

[0003] In response to the deficiencies of the prior art, the present invention provides an industrial environment monitoring and accident prediction method that integrates multimodal data, which solves the problems of the prior art.

[0004] To achieve the above objectives, the present invention is implemented through the following technical solutions: an industrial environment monitoring and accident prediction method integrating multimodal data, the industrial environment monitoring and accident prediction method comprising the following:

[0005] Sp1. Semantic Collection and Causal Association Preprocessing of Multimodal Heterogeneous Data Based on Dynamic Knowledge Graph: Real-time collection of multimodal heterogeneous data in industrial environments to obtain a dynamic knowledge graph. Utilizing the dynamically updated industrial domain knowledge graph, the data is semantically annotated, spatially aligned, and preliminarily mined and labeled for potential causal relationships, generating a structured dataset rich in semantic information and causal hypotheses.

[0006] Sp2. Deep feature extraction and multimodal cognitive fusion for weak signals and potential risks: Deep learning models are used to extract deep abstract features that characterize early weak equipment failures, subtle environmental changes, and atypical human behavior from different modal data. These multimodal deep abstract features are then cognitively fused through fusion strategies to generate a unified feature representation that characterizes the current operating status of the industrial system, potential risk factors, and their interactions. The uncertainty of the fusion results is then quantitatively assessed.

[0007] Sp3. Probabilistic prediction and explainability analysis of accident chain evolution paths based on hybrid intelligence and evolutionary learning: Build a hybrid intelligent prediction engine that integrates physical mechanism models, statistical models, and deep learning models to obtain probabilistic prediction results of accident chain evolution paths. Based on the unified feature representation and historical accident data, predict the probability of occurrence, triggering conditions, spatiotemporal evolution paths, and secondary / derivative accident chains of multiple potential accidents within a specific time window in the future. Use explainable artificial intelligence technology to provide key influencing factor analysis and explanation of the accident evolution logic for the prediction results.

[0008] Sp4. Dynamic optimization and verification of preventive intervention strategies driven by digital twins: Build a digital twin model that maps the physical industrial environment in real time with high fidelity. Input the probability prediction results of the accident chain evolution path into the digital twin model. Use a deep reinforcement learning algorithm to simulate in the digital twin environment. With the goal of minimizing expected risks and intervention costs, dynamically optimize and generate a multi-level, multi-objective preventive intervention strategy combination. Verify and iteratively optimize the generated intervention strategy combination in the digital twin environment.

[0009] Sp5, edge-end-cloud collaborative adaptive monitoring, closed-loop feedback and continuous system evolution: Collaboratively perform monitoring and prediction tasks on the device side, edge side and cloud side, collect execution effect data and new data of the intervention strategy in the physical environment, and use them as feedback information to update the knowledge graph, feature extraction and fusion model, prediction engine and reinforcement learning strategy network online, so as to realize closed-loop feedback and adaptive evolution of the entire monitoring and prediction system.

[0010] Preferably, the construction and update of the dynamic knowledge graph in Sp1 further includes: using natural language processing technology to automatically identify and extract entity concepts, attribute information, failure modes, causal relationships and operational constraints related to industrial safety, equipment health, and process flows from unstructured and semi-structured text data including maintenance records, operating manuals, safety regulations and accident investigation reports, and after verification and standardization of these extracted information, structured integration and dynamic update of the dynamic knowledge graph, thereby significantly enhancing the breadth of knowledge coverage, richness of details, real-time nature of information and accuracy of subsequent reasoning decisions of the knowledge graph.

[0011] Preferably, the deep feature extraction for weak signals in Sp2 further includes the following steps:

[0012] Sp2.1. For signals containing high-frequency noise and non-stationary components in the collected raw multimodal data (including vibration signals, acoustic signals, or high-speed visual data), use advanced signal processing techniques such as adaptive wavelet transform, empirical mode decomposition, or Hilbert-Huang transform to perform time-frequency analysis, feature enhancement, and extraction of target frequency components to initially separate and highlight sensitive features associated with early-stage faults or minor anomalies;

[0013] Sp2.2. Input the features or raw data processed by Sp2.1 into a pre-trained or task-specific fine-tuned deep learning network (including a residual network ResNet with an attention mechanism, a densely connected network DenseNet, or a gated recurrent unit GRU network for time series data), and learn and capture deep, high-dimensional, and more discriminative abstract feature representations that indicate the incipient state of early faults, accumulation of microscopic damage to materials, or slight deviations in operating procedures through multi-layer nonlinear transformations, thereby effectively improving the sensitivity, specificity, and robustness of detecting weak signals of potential risks.

[0014] Preferably, when the hybrid intelligent prediction engine in Sp3 faces challenges such as sparse historical data, frequent changes in operating conditions, or the presence of unknown new accident modes, its evolutionary learning capability is further implemented in at least one of the following ways:

[0015] Launch a meta-learning-based rapid adaptation module. This module learns "what to learn" from a large number of historically relevant tasks or simulation tasks, enabling the prediction model to quickly adjust parameters to adapt to new operating conditions or accident types using a very small number of new target domain samples.

[0016] A knowledge transfer module based on transfer learning is applied. The knowledge transfer module selectively transfers and adapts the accident pattern knowledge, fault feature representation or model parameters learned in the source domain to the prediction model of the target domain by measuring and reducing the data distribution differences or feature space differences between the source domain (including similar industrial scenarios with sufficient data and high-fidelity simulation environments) and the target domain (the current specific industrial environment). This effectively improves the cross-domain generalization ability of the prediction engine and the accuracy of early warning for low-probability and high-risk events when there is insufficient data in the target domain or there are unseen accident patterns.

[0017] Preferably, the digital twin model in the Sp4 further includes: an integrated multi-physics field coupling simulation engine based on first principles, which can simulate with high precision the nonlinear interactions between multiple physical quantities involved in thermodynamics, fluid mechanics, electromagnetism, structural mechanics and chemical reaction kinetics in industrial processes and their dynamic evolution process in fault or accident scenarios; providing a configurable parameterized scenario construction interface, allowing users to flexibly define complex accident scenarios including equipment initial state, material properties, environmental parameters, operation sequences, potential fault injection points and propagation paths according to actual needs or historical cases; supporting users to interactively manually review, parameter fine-tune, logic correct or reject the preventive intervention strategies automatically generated by the system through a graphical interface or programming interface, and re-incorporate the manually adjusted strategies into the digital twin environment for rapid verification, thereby realizing the effective combination of human and machine intelligence and ensuring the practical feasibility and acceptability of the final intervention strategy.

[0018] Preferably, the adaptive monitoring, closed-loop feedback and continuous system evolution of the edge-end-cloud collaboration in the Sp5 are specifically implemented as follows: deploying lightweight edge AI models on the device side or embedded system close to the data source to perform data quality verification, basic feature extraction and extremely low-latency simple abnormal event screening and rapid response; deploying edge computing nodes at the industrial site or workshop level, responsible for more complex multimodal data fusion, real-time risk assessment, short-term accident prediction and execution of localized preventive control instructions for the aggregated end device data and local sensor data; running large-scale, highly complex global knowledge graph management, deep learning model training and updating, long-term trend analysis, complex accident chain deduction, global resource optimization scheduling and cross-factory / cross-enterprise federated learning coordination on the cloud computing platform; the closed-loop feedback mechanism ensures that the operating data, intervention effects, and newly discovered risk patterns of the physical world can be captured in a timely manner and used to continuously optimize the models, algorithms and knowledge bases deployed at all levels, forming an intelligent monitoring and prediction system with dynamic adaptation, continuous learning and continuously improved performance.

[0019] Preferably, the industrial environment monitoring and accident prediction system integrating multimodal data includes:

[0020] Multimodal data semantic collection and preprocessing unit, used to execute Sp1;

[0021] Deep feature extraction and cognitive fusion unit for executing Sp2;

[0022] Accident evolution prediction and explainability analysis unit, used to execute Sp3;

[0023] Digital twin driven strategy optimization and verification unit, used to execute Sp4;

[0024] Cooperative control and system evolution unit, used to execute Sp5;

[0025] As well as sensor networks, industrial control system interfaces, distributed edge computing devices, cloud computing platforms, and human-computer interaction and visualization interfaces connected to the above units.

[0026] Preferably, the multimodal data semantic collection and preprocessing unit further includes:

[0027] Multi-source heterogeneous data access module, used to collect multi-source heterogeneous data in real time and batch from sensor networks, PLC / DCS systems, MES systems, video surveillance systems, and manual record databases through standard industrial protocols or customized interfaces;

[0028] A natural language processing and knowledge extraction module is configured to use machine learning models and semantic analysis algorithms to automatically extract key entities, attributes, relationships, events, and rules from maintenance logs, operating procedures, safety standards, and accident report text data to assist in building and updating industrial domain knowledge graphs;

[0029] The knowledge graph management and reasoning engine module is used to store and manage dynamically updated knowledge graphs, and perform semantic queries, logical reasoning, and data association analysis based on the graphs, providing rich background knowledge and contextual information for subsequent data fusion and prediction.

[0030] Preferably, the collaborative control and system evolution unit further includes:

[0031] Distributed model deployment and management module, used to safely and efficiently deploy trained or continuously updated monitoring models, prediction models, and control strategy models to corresponding devices, edge computing nodes, and cloud platforms, and manage the model versions, configurations, and operating status;

[0032] Federated Learning Coordination and Security Aggregation Module: When the system adopts the federated learning mechanism, this module is responsible for coordinating the local model training process of each participant (edge node or facility), securely collecting and aggregating model updates contributed by all parties, updating the global shared model, and distributing the updated model back to each participant, while ensuring data privacy and communication security throughout the process;

[0033] The closed-loop feedback and continuous learning engine is used to collect system operation data, model prediction results, manual intervention records, and strategy execution effects. Through online learning, incremental learning, or reinforcement learning mechanisms, it continuously optimizes model parameters, knowledge base content, and control strategies at all levels to improve the system's overall performance and adaptability to dynamic environments.

[0034] The present invention provides an industrial environment monitoring and accident prediction method that integrates multimodal data. It has the following beneficial effects:

[0035] 1. This invention achieves a deep understanding of massive, heterogeneous industrial data and the precise construction of contextual scenarios by constructing a dynamically updated industrial domain knowledge graph and combining targeted natural language processing technology with multimodal data semantic collection, high-precision spatiotemporal alignment, and preliminary causal association mining. This significantly improves the system's sensitivity, recognition accuracy, and comprehensive identification capabilities for early, subtle, and complex anomalies in the industrial environment. This allows potential risks that were previously difficult to detect or easily overlooked to be accurately captured and characterized at an early stage, providing the data foundation and feature support for subsequent precise predictions and proactive interventions.

[0036] 2. This invention uses deep reinforcement learning driven by high-fidelity digital twins to dynamically optimize and conduct closed-loop verification of preventive intervention strategies. In a digital twin environment that maps the physical world in real time and integrates multi-physics field coupling simulation capabilities, it is possible to automatically explore and learn a combination of multi-level, multi-objective preventive intervention strategies that minimize expected risks and intervention costs. The closed-loop decision-making mechanism of "prediction-deduction-optimization-verification-execution" ensures that the preventive intervention measures taken are not only theoretically optimal, but also robust, safe, and acceptable in actual operating environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a system composition diagram of the present invention;

[0038] Figure 2 It is an operation flow chart of the present invention;

[0039] Figure 3 This is a diagram illustrating the core interactions of multimodal data fusion and prediction driven by the knowledge graph of the present invention. DETAILED DESCRIPTION

[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention. Specific embodiment one:

[0042] like Figures 1 to 3 As shown, a method for industrial environment monitoring and accident prediction that integrates multimodal data includes the following:

[0043] Sp1. Semantic collection and causal association preprocessing of multimodal heterogeneous data based on dynamic knowledge graph: This step first collects all-round, high-fidelity multimodal heterogeneous data in the industrial environment in real time through the multi-heterogeneous sensor network and data interface deployed at the industrial site, including: Visual data: Use industrial cameras with a resolution of not less than 2K and a frame rate of not less than 60fps (including RGB, infrared thermal imaging, and hyperspectral camera arrays) to capture equipment surface status, personnel behavior, and environmental smoke and flame image sequences; Acoustic data: Use a high-precision microphone array covering the 20Hz-100kHz frequency band (with beamforming capabilities) to synchronously collect equipment operation voiceprint information and environmental abnormal sounds; Vibration data: By sticking to the core components of key equipment (including bearing seats, motor housings, gears A triaxial accelerometer (range ±50g, frequency response 0.2Hz-20kHz) on the wheel box acquires equipment vibration signals. Equipment operating parameters: Key process parameters (including temperature, pressure, flow, liquid level, speed, current, voltage, and valve opening) are collected in real time from DCS, PLC, and SCADA systems using standard industrial protocols such as OPC UA, Modbus TCP, and EtherNet / IP, with a sampling period of no more than 100ms. Environmental parameters include temperature and humidity, flammable / toxic gas concentrations (using electrochemical or infrared sensors), and dust concentration (using a laser scattering sensor). Personnel data: Worker location, posture, movement sequence, and important operation records are captured through badge positioning and video behavior analysis. The collected raw data undergoes rigorous quality verification and cleaning, including outlier removal, gap filling (using an algorithm based on multivariate time series interpolation), and signal noise reduction (using adaptive Wiener filtering and wavelet threshold denoising for vibration and acoustic signals). Subsequently, the dynamically updated industrial domain knowledge graph is used to perform high-precision semantic annotation of the data, sub-second spatiotemporal alignment (based on NTP network time synchronization and unified spatial coordinate system conversion), and preliminary mining and labeling of potential causal relationships (using association rule mining based on the Apriori algorithm and causal structure learning based on the PC algorithm to initially screen candidate causal pairs). Ultimately, a structured dataset rich in precise spatiotemporal stamps and deep semantic information is generated, and potential causal links are preliminarily calibrated, providing a high-quality data foundation for subsequent analysis.

[0044] The construction and updating of the dynamic knowledge graph in Sp1 further includes: The knowledge graph is constructed using a combination of top-down and bottom-up approaches. The basic ontology framework is based on ISO15926 and ISA-95 international standards, combined with industry-specific ontologies (OntoCAPE for the chemical industry and the CIM model for the power industry) and the company's internal expert knowledge base (organized through structured interviews and knowledge acquisition tools). Protégé 5.5.0 is used for ontology modeling, defining no fewer than 500 core entity categories (pumps, valves, compressors, reactors, sensors, alarm events, failure modes, operating procedures, safety measures), 1,000 entity attributes (equipment ID, rated power, material, installation location, operating status, fault description, maintenance history), and 200 semantic relationships ("is-a," "part-of," "connected-to," "causes," "prevents," and "requires"). Natural language processing technology is used to automatically identify and extract structured information from massive amounts of unstructured and semi-structured text data (including but not limited to equipment maintenance records in PDF format, operating manuals and process cards in Word documents, safety checklists in Excel spreadsheets, and accident investigation reports and industry standard specifications in HTML or plain text format). Specifically, a domain pre-training model based on BERT (Bidirectional Encoder Representations from Transformers) is adopted (secondary pre-training of Masked Language Model and Next Sentence Prediction tasks is performed on more than 1TB of industrial text corpus, with a model parameter volume of no less than 340M, BERT-Large), combined with a downstream BiLSTM-CRF model for named entity recognition (NER), with an entity recognition F1 value of no less than 95%; a relation extraction model (RE) based on R-BERT or SpanBERT is used to identify predefined semantic relations between entities, with a relation extraction accuracy of no less than 90%; and an event extraction (EE) technology based on dependency syntactic analysis and semantic role labeling is used to extract key events ("equipment A failed X", "operator B performed operation Y") and their arguments.All extracted information undergoes multiple verifications (including consistency verification with existing knowledge graphs, rule base verification, and manual sampling when necessary) and strict standardization (physical units are uniformly based on the International System of Units, professional terms are standardized against standard dictionaries, and time formats are unified to ISO8601). Finally, it is structured and integrated into the Neo4j4.x enterprise graph database in the form of RDF triples or property graphs and dynamically updated (cluster deployment to ensure high availability and scalability, and the graph scale design capacity supports tens of billions of nodes and relationships). This dynamic knowledge graph ensures the breadth of its knowledge coverage (from a single component to the complete hierarchical structure of the entire factory), richness of details (multi-dimensional attributes and complex relationship networks of each entity), real-time information (through the API interface with the real-time data stream, the device status and parameter information in the knowledge graph are synchronized with the physical world in near real time, with a delay of no more than 5 seconds), and the accuracy and reliability of subsequent graph-based complex reasoning (using the Cypher query language for multi-hop association queries, pattern matching, or combined with the rule engine Drools for logical reasoning) and intelligent decision-making.

[0045] Sp2. Deep feature extraction and multimodal cognitive fusion for weak signals and potential risks: Customized deep learning models are used for deep abstract feature extraction for different modal data: Vibration signal: A multi-scale residual convolutional time series network (MSRCTN) is adopted. The network first extracts the multi-scale local patterns of the original vibration time domain signal through parallel one-dimensional convolution kernel groups with different receptive field sizes (3, 5, 7, 9) (each group contains 64 convolution kernels, ReLU activation, BatchNormalization), and then learns its long-term temporal dependency through a residual-connected stacked temporal convolutional network (TCN) module (including dilated convolution layers with a dilation factor sequence of 1, 2, 4, 8 to ensure the capture of dependencies over a large time span). Finally, the output is a high-dimensional feature vector (with a dimension of 512) containing early weak fault features of the equipment (bearing pitting, gear tooth breakage, periodic shock pulses caused by rotor imbalance, or specific frequency modulation phenomena). Acoustic signals: First, a log-Mel spectrogram (with a frame length of 25ms, a frame shift of 10ms, and 128 Mel filters) is extracted. This is then fed into a 2D convolutional neural network based on the EfficientNet-B3 architecture and combined with a channel-wise and spatial attention mechanism (CBAM). This network efficiently extracts the fine-scale temporal-spectral structural features (512 dimensions) of acoustic events (gas leaks, mechanical friction noises, and discharge sounds) through depthwise separable convolutions and squeeze-and-excitation modules. Visual data: For equipment condition monitoring, an image encoder based on SwinTransformerV2-Large is used to extract visual representations of key areas (such as cracks on the equipment surface, abnormal oil color, and instrument readings). For human behavior analysis, a human pose estimation algorithm based on HRNet is used to obtain joint coordinate sequences. A spatiotemporal graph convolutional network (ST-GCN) is then used to identify atypical or illegal behaviors (such as not wearing a helmet, entering a restricted area, and falling down unexpectedly). The network then outputs behavior classification probabilities and associated visual features (512 dimensions). The deep, abstract features extracted from each modality are cognitively fused through a Multi-modal Cross-attention Transformer Fusion Network (MCTFN), based on a multi-head cross-modal attention mechanism. This network comprises multiple encoder layers, each equipped with a self-attention module for each modality and a cross-attention module across all modality pairs. This network dynamically learns complex nonlinear dependencies within and between modalities, and adaptively assigns different fusion weights to features from different modalities and time steps.The output of the MCTFN is a unified feature representation vector with a dimension of 1024. This vector comprehensively characterizes the overall operating state of the current industrial system at a specific moment, the combined effects of potential risk factors (equipment degradation, environmental hazard levels, and unsafe conditions for personnel), and their interactions. To quantitatively assess the uncertainty of the fusion results, this solution employs a deep ensemble learning approach. This involves independently training multiple (five) MCTFN models with different random initialization seeds or slightly different hyperparameters. During inference, their outputs (feature vectors or predicted probabilities for downstream tasks) are averaged or weighted averaged, and the variance or standard deviation of the predictions is calculated as a measure of uncertainty. This uncertainty measure is used to guide subsequent risk decision-making and active learning.

[0046] The deep feature extraction for weak signals in Sp2 further includes the following steps:

[0047] Sp2.1. For signals containing high-frequency noise and non-stationary components in the collected raw multimodal data, particularly vibration signals from high-speed rotating machinery (steam turbines, compressors, and large fans) (which often contain abundant fault harmonics and sideband information but are easily drowned out by background noise), acoustic signals generated by high-pressure fluid (steam, natural gas) leaks or cavitation (which have characteristic frequencies as high as tens of kHz and are bursty and non-stationary), and industrial visual data captured on high-speed production lines (steel rolling, papermaking, and printing) (which suffer from motion blur and uneven lighting), we first employ advanced signal processing techniques based on the Synchrosqueezed Wavelet Transform (SSWT). SSWT redistributes and concentrates signal energy on the time-frequency plane, improving time-frequency resolution. It is particularly suitable for modal separation and instantaneous frequency estimation of multi-component non-stationary signals. By performing time-frequency analysis of the signal using SSWT and combining it with an adaptive threshold denoising algorithm (SUREshrink or Bayes Shri nk based on wavelet coefficients) to enhance the signal-to-noise ratio, target frequency components or transient features with a signal-to-noise ratio improvement of at least 10 dB are extracted. This allows for the preliminary separation and highlighting of sensitive features associated with early-stage faults (tiny spalling less than 0.1 mm in diameter on the bearing outer ring, initial pitting on the gear tooth surface, and tiny cracks in the rotor blades) or minor anomalies (loosening of equipment anchor bolts by less than 0.5 mm, slight changes in oil film stiffness due to trace metal wear debris in the lubricating oil).

[0048] In Sp2.2, the sensitive features extracted after preprocessing in Sp2.1 (such as the instantaneous energy value and kurtosis of a specific fault frequency band in the time-spectrogram obtained by SSWT, or the waveform data of a single IMF component after modal reconstruction) or, in an end-to-end learning scenario, the raw data are directly fed into a deep learning network designed and rigorously fine-tuned for a specific industrial monitoring task. This network employs a Hybrid Attention Residual Temporal Convolutional Network (HARTCN). The input layer of the HARTCN receives the preprocessed feature sequence. Its core structure consists of stacked residual temporal convolutional modules, each of which contains multiple layers of convolutions with dilation factors (the dilation factors are 1, 2, 4, 8, and 16, ensuring an exponential increase in the receptive field). A dual mechanism of channel attention and temporal attention is introduced between the convolutional layers: the channel attention module (SE module) is used to adaptively learn the importance of different feature channels and enhance the weight of key fault-indicating features; the temporal attention (the self-attention mechanism in the Transformer) is used to capture long-term dependencies and key time points within the temporal features. The activation function uniformly uses GeLU, the optimizer uses AdamW, and the learning rate adopts a cosine annealing strategy with preheating. Through multiple layers of nonlinear transformations, the network learns and captures deep, high-dimensional (output feature vector dimension is 512), and highly discriminative abstract feature representations from the input that indicate early failure initiation (the initial 10% of the lifespan of fatigue cracks), microscopic damage accumulation (the initial stages of creep or corrosion in equipment under high-temperature and high-pressure environments), or minor deviations from the operating process (persistent small fluctuations in process parameters within 0.5% of the set value). The goal is to minimize the intra-class distance and maximize the inter-class distance in the feature space between samples of different health states or risk levels in multi-class classification or regression tasks. This process effectively improves the sensitivity (targeting an early detection rate of over 98% for known failure modes), specificity (targeting a false alarm rate of less than 0.5% under normal conditions), and robustness (detection performance degradation of no more than 5% within a ±20% operating condition fluctuation) of detecting weak signals of potential risks.

[0049] Sp3. Probabilistic prediction and explainability analysis of accident chain evolution paths based on hybrid intelligence and evolutionary learning: The core architecture of the hybrid intelligent prediction engine constructed in this solution is a three-layer decision fusion system: the bottom layer is a cluster of physical mechanism models based on first principles, including finite element models for key equipment (using ANSYS Mechanical for rotor dynamics analysis and fatigue life prediction), computational fluid dynamics models (using Fluent to simulate multiphase flow and heat transfer processes in reactors), and system dynamic models based on state-space equations. These models can provide physical-level predictions of theoretical behavior baselines, stress distribution, and failure probabilities for specific equipment or subsystems based on real-time input operating parameters. The middle layer is a library of multivariate statistical analysis and machine learning models, including a multivariate time series prediction model based on Vector Autoregres sion (VAR) (used to capture dynamic correlations and short-term trends between key parameters), a remaining life prediction model for equipment based on SurvivalAnalysis (using Weibull distribution or Cox proportional hazard model, combined with historical failure data and real-time status parameters), and a state transition probability estimation model based on Hidden Markov Model (HMM) or Dynamic Bayesian Network (DBN) (used to predict the characteristics of the system migrating from the current healthy state to different fault states). The top layer is a complex accident pattern recognition and evolution prediction module based on deep learning, the core of which is a spatio-temporal dynamic graph attention network (STDGAT). STDGAT uses time series data of a unified feature representation vector generated by Sp2 and a knowledge graph constructed by Sp1 (representing the physical connections, logical dependencies, and causal relationships between devices and systems) as input. The nodes of the graph represent the various monitored objects (devices, sensors, and regions) in the industrial system, and their features are the outputs of Sp2. The edges of the graph are dynamically constructed based on the knowledge graph and assigned time-varying weights. STDGAT learns and captures the propagation patterns and dependencies of accidents in space (between devices and regions) and time by alternating multiple layers of graph attention layers (GATv2) and temporal convolutional layers (TCN or GRU).The engine is based on the unified feature representation generated by Sp2 and the historical accident database (containing detailed accident records for at least the past five years, each record containing accident ID, occurrence time, duration, equipment involved, triggering factors, direct causes, indirect causes, development process, casualties, economic losses, and structured information of measures taken). It accurately predicts the probability of occurrence (outputting a continuous value between 0 and 1) of various potential accidents (type A compressor surge, overpressure leakage of tank B, electrical fire in area C, and unplanned shutdown of production line D due to key equipment failure) within a specific future time window (which can be set to the next 15 minutes, 1 hour, 8 hours, or 24 hours), the maximum trigger condition combination (identified by back-tracing important paths and node features in STDGAT), the spatiotemporal evolution path (displaying the new sequence and expected time of the accident starting from the source node and spreading to other nodes in the time and space dimensions in the form of a probability graph), and the secondary / derivative accident chain (predicting fire and explosion caused by equipment failure, which in turn leads to structural damage and the spread of toxic substances, forming a domino effect). The target accuracy under the accident probability prediction (AUC) is set at no less than 0.90, and the target mean average precision (MAP) for evolutionary path prediction is set at no less than 0.85. Furthermore, Explainable Artificial Intelligence (XAI) technology is employed, combining integrated gradients (IGGs) and graph-based counterfactual explanations (GCE). IGGs are used to quantify the contribution of input features (the dimensions in the unified feature representation of Sp2) to the STDGAT prediction results (accident probability) and identify key influencing factors. GCE generates counterfactual explanations in the form of "If X had not occurred, the probability of Y accident would have decreased by Z%" by searching the knowledge graph for the smallest subgraph structure or edge perturbation that could change the prediction result. This provides deep insight into the accident evolution logic and provides a basis for decision-making.

[0050] When faced with challenges such as sparse historical data, frequently changing operating conditions, or unknown new accident patterns, the hybrid intelligent prediction engine in Sp3 further implements its evolutionary learning capabilities through at least one of the following methods:

[0051] A rapid adaptation module based on the Model-Agnostic Meta-Learning (MAML) algorithm is activated. MAML learns a universal model initialization parameter set or a meta-network capable of rapidly generating task-specific parameters by meta-training on a large number (hundreds) of historically related tasks or high-fidelity simulation tasks with similar structures but varying parameters (fault prediction of centrifugal pumps of different models, reactor runaway simulation under different catalyst ratios). This meta-knowledge enables the overall prediction model to adapt to new operating conditions (increasing production load from 70% to 95%), new equipment models (replacing sensors or actuators with new models), or historically undocumented accident types (failure modes caused by rare material corrosion) using only a small number (5-10) of new target domain samples. Through a few steps of gradient descent, the model rapidly adjusts its internal parameters to adapt to the new data distribution and task requirements, achieving rapid learning and high-precision prediction for new scenarios. Its learning convergence speed on new tasks is at least 30% faster than traditional transfer learning.

[0052] A knowledge transfer module based on adversarial domain adaptation (ADA) is applied. The core idea of this module is to introduce a domain discriminator (Domain Discriminator) and a feature extractor (typically the underlying graph convolution and temporal convolution layers of the STDGAT) for adversarial training between the source domain (a benchmark factory A with sufficient labeled data and complete operating records, or a large-scale simulation dataset containing multiple failure modes generated in a digital twin environment) and the target domain (a specific industrial environment B with sparse data or unknown patterns). The feature extractor's goal is to learn domain-independent feature representations that can both achieve the primary prediction task (accident probability prediction) and confuse the domain discriminator. The domain discriminator's goal is to distinguish whether features originate from the source or target domain. Through this adversarial game, the feature extractor is forced to learn invariant feature representations that are transferable between the two domains. A partial transfer strategy is used to transfer model parameters: first, the entire STDGAT model is pretrained in the source domain. Then, the parameters of its feature extraction layer are fixed or fine-tuned with a small learning rate. Simultaneously, the top prediction layer and domain discriminator are trained on data from the target domain. This method selectively transfers and adapts the accident pattern knowledge (expressed as feature combinations sensitive to specific faults), fault feature representations (i.e., feature extractor parameters), or model parameters (attention weights) learned in the source domain to the prediction model in the target domain. This effectively improves the cross-domain generalization capability of the prediction engine and the early warning accuracy of low-probability, high-risk events (such as cascading failures caused by extreme weather and raw material quality issues caused by supply chain mutations) when the target domain lacks data (such as a lack of fault samples in the early stages of production of new equipment) or when there are unseen accident patterns. The goal is to increase the prediction accuracy of the target domain by at least 15% (compared to not using transfer learning).

[0053] Sp4. Dynamic optimization and verification of preventive intervention strategies driven by digital twins: The digital twin model constructed in this solution uses the NVIDIA Omniverse platform as the core visualization and collaboration foundation, integrating multi-source data and multi-physics field simulation capabilities. The three-dimensional geometric model is constructed by importing CAD / BIM (Solid Works, Revit, Plant3D) design data and performing lightweight processing, with an accuracy of millimeter level. The physical property database includes equipment materials (mechanical, thermal, and electrical properties corresponding to ASTM standard grades), fluid media (viscosity, density, specific heat capacity), and environmental parameters (atmospheric pressure, soil properties). The behavioral logic is imported into the control system logic (FMU compiled from the Simulink model) and operating procedures (represented in the form of a state machine or behavior tree) through the Functional Simulation Unit (FMU) / Functional Prototype Interface (FMI) standard. This digital twin model uses MQTT and Kafka message queues to access various sensor data collected by Sp1 in real time (with latency less than 100ms) and the unified feature representation generated by Sp2. Using a built-in Kalman filter and data assimilation algorithm, its internal state is synchronized with the physical entity with high fidelity (with state deviation less than 2%). Sp3's predicted probabilistic predictions of the accident chain evolution path are displayed as a dynamic risk heat map and accident evolution sequence superimposed on the digital twin model's 3D scene. The dynamic optimization of preventive intervention strategies utilizes a deep reinforcement learning algorithm based on proximal policy optimization (PPO). The state space of the PPO agent consists of the full system state vector provided by the digital twin model (including key equipment health indicators, process parameters, and environmental parameters, with a dimension of approximately 2000), the unified feature representation of Sp2 (with a dimension of 1024), and the predicted risk vector of Sp3 (including the probability of each potential accident and the probability of key nodes in the evolution path, with a dimension of approximately 500). The action space is a hybrid of discrete and continuous, including: discrete actions: selecting the intervention object (specific equipment, specific area), selecting the intervention type (adjusting parameters, starting backup, performing maintenance, issuing alarms, triggering ESD); continuous actions: adjusting the specific values of parameters (setting target values for temperature, pressure, and flow, floating within the allowable range).The reward function R is designed as: R = w1(ΔSafety)-w2(Cost_Intervention)-w3(Cost_Downtime)-w4(Cost_Energy)+w5(ΔEfficiency), where ΔSafety is the accident probability reduction value or the risk level reduction value, Cost_Intervention is the direct cost of the intervention measure, Cost_Downtime is the production downtime loss caused by intervention or accident, Cost_Energy is the energy consumption change caused by the intervention measure, ΔSafety is the setting value of the safety level, which is defined as (0-10) according to the safety level, where 0 is the lowest safety requirement and 10 is the highest safety requirement. The initial conventional setting is 6, ΔEfficiency is the improvement in production efficiency, and w1-w5 are adjustable weight coefficients, which are set according to the company's risk preference and operating goals. The PPO agent undergoes simulated training for at least 10^7 time steps in a digital twin environment. Through interaction with the environment, it learns a policy network (actor network) and a value network (critic network) that maximize the cumulative expected reward. Both networks employ a neural network architecture with multiple fully connected layers and residual connections, using the Tanh activation function. After training, the policy network dynamically optimizes and generates a multi-level (device-level parameter fine-tuning, system-level process reengineering, plant-level emergency response initiation) and multi-objective preventive intervention strategy combination based on real-time input. The generated strategy combination is first validated in the digital twin environment through at least 1000 Monte Carlo simulations under different initial conditions and random perturbations. The effectiveness (target 99% accident avoidance success rate), robustness (no more than 5% drop in effectiveness for a ±10% perturbation of key parameters), and potential negative impacts (perturbations to other connected systems) are evaluated. Based on the validation results, the policy network parameters and reward function are iteratively optimized until the preset performance indicators are achieved.

[0054] The digital twin model in Sp4 further includes the integration of a first-principles-based multi-physics coupled simulation engine. This is achieved by integrating COMSOL Multiphysics (for electromagnetic-thermal-fluid-structural coupled analysis), ANSYS Fluent (for complex fluid dynamics and combustion simulation), and Siemens Simcenter Amesim (for one-dimensional system-level multi-domain dynamic simulation) as callable simulation solvers into the digital twin platform via the FMI / FMU interface. These solvers enable high-precision simulation of the nonlinear interactions between multiple physical quantities involved in industrial processes, including thermodynamics (heat exchanger efficiency, reaction thermal runaway), fluid mechanics (pipeline pressure loss, pump cavitation), electromagnetics (motor electromagnetic torque, transformer core loss), structural mechanics (pressure vessel stress concentration, pipeline vibration fatigue), and chemical reaction kinetics (catalyst activity changes, product yield), as well as their dynamic evolution under fault (equipment overheating, pipeline rupture) or accident (fire spread, explosion shock wave propagation). The simulation results are consistent with experimental or historical data at least 90%. A web-based, configurable, parameterized scenario-building graphical user interface (GUI) is provided, allowing authorized users (process engineers, safety engineers) to flexibly define complex accident scenarios, including equipment initial state (setting equipment operating time, initial wear coefficient, material defect size), material properties (inputting chemical composition, viscosity, and flash point of different batches of raw materials), environmental parameters (setting extreme temperature, humidity, wind speed, and external vibration sources), operation sequences (simulating misoperation, emergency shutdown, and maintenance processes), potential fault injection points (virtually imposing crack propagation, valve jamming, and sensor failure on specific components in the digital twin model), and accident propagation path constraints (simulating firewall collapse, fire protection facility failure, and adverse conditions of ventilation system reverse operation) through dragging and dropping, parameter input, and logical arrangement, based on actual needs, historical accident cases (imported from the accident database), or HAZOP analysis results. This allows for targeted simulation and emergency plan verification.It supports users to interactively review the preventive intervention strategies automatically generated by the PPO agent through a graphical interface (integrated in the cockpit of the digital twin platform) or Python API, fine-tune parameters (adjust the recommended alarm threshold from 80% to 75%, and limit the opening adjustment rate of the control valve to within 5% / second), correct logic (add a manual confirmation link before automatic shutdown, or adjust the execution priority of multiple parallel intervention measures), or completely reject them (if experts judge based on experience that the strategy will cause more serious consequences or not comply with safety regulations). The manually adjusted strategy (or the choice of executing the manual plan) is then re-incorporated into the digital twin environment for rapid verification (single simulation evaluation completed within 1 minute) and effect comparison, thereby achieving deep collaboration and optimization of human-machine intelligence, ensuring the actual feasibility, operational compliance, regulatory compliance and acceptability of the final intervention strategy to on-site personnel, and making up for the ethical risks and experience blind spots of pure AI decision-making.

[0055] Sp5, edge-end-cloud collaborative adaptive monitoring, closed-loop feedback, and continuous system evolution: This solution adopts a strict three-layer intelligent collaborative architecture. On the device side: ultra-lightweight AI models are deployed at key sensor nodes (smart vibration sensors, smart cameras), actuators (smart valve positioners), or small embedded controllers. These models have been extremely optimized (8-bit integer quantization, weight pruning rate of not less than 70%), and mainly perform: online data quality verification (sensor self-diagnosis, data packet CRC verification, physical range rationality judgment); basic feature extraction (RMS, peak-to-peak value, kurtosis, zero-crossing rate time domain statistics, or simple frequency domain energy ratio); and simple abnormal event screening and instinctive rapid response with extremely low latency (response time less than 10ms) (immediately triggering local sound and light alarms when vibration intensity exceeds the second threshold, and immediately cutting off the heating power when temperature exceeds the limit). Edge ("Edge" Intelligence): Edge computing nodes (equipped with at least 32GB of RAM, a 256GB NVMe SSD, and GPU acceleration support) are deployed at the industrial site or workshop level, connecting to end devices and upper-layer cloud platforms via 5G RLC or TSN industrial Ethernet. Edge nodes are responsible for: performing more complex multimodal data fusion (a lightweight version of Sp2's MCTFN model) on aggregated end-device data and local high-value sensor data (high-frequency vibration, multispectral vision); real-time risk assessment (calculating device health index and regional safety score based on fused features, with a refresh rate of at least 1Hz); short-term accident prediction (using GRU or small Transformer models to predict the probability of specific equipment failure or small-scale accident risk within the next 1-60 minutes); and executing localized, somewhat autonomous preventive control commands (preemptively fine-tuning compressor guide vane opening based on predicted surge risk; automatically isolating relevant pipe sections and initiating exhaust ventilation based on predicted leakage risk). Cloud (“Cloud” Intelligence): Deploy a central intelligent platform with large-scale parallel computing (using MPI, Horovod for distributed training), massive storage (HDFS, Ceph) and complex analytical capabilities on a private cloud (built on OpenStack and Kubernetes) or hybrid cloud architecture.Cloud platform operation: construction, management, reasoning and continuous updating of the global knowledge graph (Sp1); periodic (daily or weekly) retraining and fine-tuning of the ultra-large-scale deep learning models involved in Sp2, Sp3 and Sp4 (Transformer models with more than 1 billion parameters, DRL models that require thousands of GPU hours of training); long-term trend analysis based on historical data and real-time status of the entire plant (equipment group life prediction, in-depth mining of accident occurrence patterns, and global optimization of maintenance strategies); complex accident chain deduction and domino effect analysis (Sp3, simulating major accident scenarios across systems and regions); coordinated optimization and scheduling of global production operations and safety risks (dynamically adjusting production plans to maximize overall benefits while ensuring safety); and serving as a central coordination server and model aggregation center for federated learning (if adopted). A closed-loop feedback mechanism is central to the system's continuous evolution. Real-time operational data from the physical world (including baseline data under normal operating conditions, early warning event data, and detailed process data from existing failures), the effectiveness of preventive intervention strategies (quantified by comparing changes in risk levels, accident rates, and production indicators before and after intervention), manually input feedback (operation and maintenance personnel's confirmation and handling records of alarms, and expert evaluations and scores of prediction accuracy and intervention recommendations), and newly discovered risk patterns or failure mechanisms discovered through the continuous learning module are all structured and securely transmitted to the corresponding model update process. Specifically, the online update module uses a streaming learning algorithm based on HoeffdingTree or OnlineRandomForest to process real-time small batches of data, rapidly fine-tuning the parameters of edge and some cloud models. Periodic model retraining utilizes all accumulated historical data and feedback to comprehensively optimize the core complex model in the cloud. The policy network of the reinforcement learning agent is iteratively improved through continuous online interaction or offline policy evaluation. The entire system forms a virtuous cycle of data-driven, model iteration, and knowledge enhancement, enabling it to dynamically adapt to the ever-changing industrial environment, equipment aging, process adjustments, and the emergence of new risks, and achieve continuous improvement in monitoring accuracy, prediction accuracy, and intervention effectiveness (with the goal of improving key performance indicators by no less than 5% each year).

[0056] The adaptive monitoring, closed-loop feedback and continuous evolution of the edge-end-cloud collaboration in Sp5 are specifically implemented as follows: deploying ultra-lightweight edge AI models that have undergone deep pruning (removing more than 80% of redundant connections), extreme quantization (INT8 or INT4) and dedicated instruction set optimization at the device end or embedded system close to the data source (using RISC-V architecture-based and integrated NPU SoC chips with power consumption less than 5W) to perform data quality verification, basic feature extraction and simple abnormal event screening and rapid response with extremely low latency; deploying edge computing gateways or servers with at least 128TOPSAI computing power and support for multi-channel high-speed data concurrent processing at the industrial site or workshop level, responsible for more complex multimodal data fusion, real-time risk assessment, short-term accident prediction and execution of localized preventive control instructions for the aggregated end device data and local sensor data; The cloud computing platform (which adopts a microservice architecture, performs container orchestration and resource scheduling based on Kubernetes, supports GPU / TPU heterogeneous computing clusters, and integrates the MLOps tool chains Kubeflow and MLflow) runs large-scale, highly complex global knowledge graph management, periodic training and updating of deep learning models (especially ultra-large-scale models that use distributed parameter server architecture or AllReduce algorithm for efficient parallel training), long-term trend analysis, complex accident chain deduction, global resource optimization scheduling, and serves as a coordination server for federated learning; the closed-loop feedback mechanism ensures that the operating data, intervention effects, and newly discovered risk patterns of the physical world can be captured in a timely manner and used to continuously optimize the models, algorithms, and knowledge bases deployed at all levels, forming an intelligent monitoring and prediction system that dynamically adapts, continuously learns, and continuously improves performance. Specific embodiment two:

[0058] like Figures 1 to 3 As shown, an industrial environment monitoring and accident prediction system integrating multimodal data includes:

[0059] The core of the multimodal data semantic collection and preprocessing unit is one or more high-performance data servers (CPU: Intel Xeon Scalable series, RAM: at least 512GB, storage: NVMe SSD RAID10 array: at least 20TB) running specially developed distributed data collection and processing software. This software integrates a multi-heterogeneous data access module for Sp1 (with built-in OPC UA Client / Server SDK, Modbus Master / Slave Library, MQTT Broker / Client, Kafka Producer / Consumer, as well as JDBC / ODBC drivers and file parsing engines for mainstream databases); a data cleaning and validation module (using an outlier detection and repair algorithm based on statistical rules and machine learning); a high-precision spatiotemporal alignment module (with an integrated NTP client and coordinate conversion service); a knowledge graph construction and management module (with a Neo4j Enterprise Edition cluster on the backend and knowledge editing and visualization tools on the front end); and a natural language processing module based on a domain-pretrained Transformer model (for extracting entities, relationships, and events from text). This unit strictly performs data collection, semantic annotation, spatiotemporal alignment, causal preprocessing, and knowledge graph construction and dynamic update tasks in accordance with the technical plan set by Sp1.

[0060] The deep feature extraction and cognitive fusion unit is deployed on edge computing nodes and cloud platforms. The edge nodes (the aforementioned NVIDIA Jetson AGX Orin or equivalent platforms) are responsible for real-time or quasi-real-time feature extraction and preliminary fusion, while the cloud platform (equipped with a large-scale GPU cluster, NVIDIA A100 / H100) is responsible for the training of complex models and deeper fusion. The unit integrates customized deep learning feature extractors for each modality of Sp2 (MSRCTN, EfficientNet-B3+CBAM, Swin Transformer V2-Large+ST-GCN, with model parameters strictly tuned and verified), a Transformer fusion network (MCTFN) module based on a multi-head cross-modal attention mechanism, and an uncertainty quantification module based on deep ensemble learning. This unit strictly follows the technical plan set by Sp2 to perform deep feature extraction for weak signals and potential risks, multimodal cognitive fusion, and uncertainty assessment of fusion results.

[0061] The accident evolution prediction and explainability analysis unit is mainly deployed on the cloud platform, and some lightweight prediction models can be deployed on edge nodes. The unit integrates Sp3's hybrid intelligent prediction engine (including a callable physical mechanism model library (FMU encapsulation), a statistical model library (R / Python script implementation), and the core STDGAT deep learning model), an accident chain deduction module (graph traversal and probabilistic reasoning based on STDGAT), a probabilistic prediction output module (outputs structured accident risk reports), and an XAI tool set that integrates gradients and graph counterfactual interpretations. The unit strictly follows the technical solution set by Sp3 to perform probabilistic predictions of accident chain evolution paths and provide explainable analysis results.

[0062] The digital twin-driven policy optimization and verification unit is deployed on a cloud platform, with a visual interactive interface accessible via the web. This unit integrates NVIDIA Omniverse-based digital twin construction, rendering, and synchronization modules (which enable real-time, two-way communication with the physical world data interface, with latency less than 200ms), a multiphysics simulation engine (the aforementioned COMSOL, Fluent, and Amesim are integrated via the FMI / FMU standard), a PPO-based deep reinforcement learning training and inference framework (using the RayRLlib or Acme distributed RL library), and a human-computer interface that supports parameterized scenario construction and manual policy intervention. This unit dynamically optimizes and verifies preventive intervention strategies in strict accordance with the technical plan established by Sp4.

[0063] The collaborative control and system evolution unit is a distributed unit that spans the three layers of end, edge, and cloud. Its core logic and coordination management functions are deployed on the cloud platform, and the specific execution modules are distributed at all levels. The unit integrates the resource monitoring and task scheduling module of the edge-end-cloud three-layer computing architecture, the distributed model deployment and version management module based on containerization (Docker) and orchestration (Kubernetes) technology (supporting blue-green deployment and canary release), the structured closed-loop data feedback collection and preprocessing module, the online learning and continuous evolution engine (including a streaming learning algorithm library and a regularization method for catastrophic forgetting), and (if enabled) a federated learning coordination and model aggregation module based on secure multi-party computing and differential privacy. The unit strictly implements adaptive monitoring, closed-loop feedback, and continuous system evolution in accordance with the technical solution set by Sp5.

[0064] And the following connected with the above units to form a complete physical information system:

[0065] (1) Sensor network: contains at least 10,000 smart and traditional sensors of various types, connected to the data acquisition gateway through wired (industrial Ethernet, fieldbus) and wireless (5G, LoRaWAN, Wi-SUN) methods;

[0066] (2) Industrial control system interface: ensure seamless and secure (using one-way network gates or secure data channels) two-way data interaction with existing DCS, PLC, and SIS systems;

[0067] (3) Distributed edge computing equipment: deploy at least one edge computing node (specifications as described above) in each major production unit or key area;

[0068] (4) Cloud computing platform: with FP32 computing power of no less than 1000 TFLOPS and storage capacity of 10PB;

[0069] (5) Human-computer interaction and visualization interface: A comprehensive monitoring, early warning and decision support platform developed based on Web technology (React / Vue+Three.js / Babylon.js) provides multi-dimensional data visualization, risk heat map, digital twin interaction, accident evolution animation, early warning information push, intervention strategy recommendation and execution confirmation functions, and supports desktop and mobile access.

[0070] The multimodal data semantic collection and preprocessing unit further includes:

[0071] The core of the multi-heterogeneous data access module is a configurable data acquisition engine with built-in drivers and parsing libraries for mainstream industrial protocols (OPCDA / UA, Modbus TCP / IP&RTU, EtherNet / IP, ProfinetIO, IEC61850, MQTT, CoAP, and DNP3). It can establish stable and efficient data connections with field devices and systems in master-slave or publish-subscribe mode. For database systems (Oracle, SQL Server, MySQL, PostgreSQL, and time series databases InfluxDB and Prometheus), data is extracted through a high-performance JDBC / ODBC connection pool and optimized SQL query statements. For video streams, it supports RTSP, RTMP, and GB / T28181 protocol access and performs H.264 / H.265 stream parsing. For file data, it provides batch import and incremental synchronization functions for CSV, Excel (XLSX), JSON, XML, Parquet, and HDF5 formats. All data is timestamp aligned (based on the central NTP server time) and source identity authenticated upon access.

[0072] The natural language processing and knowledge extraction module is configured to adopt a cascade model architecture based on the "pre-training-fine-tuning-distillation" paradigm: first, a Transformer model (BERT-Large or RoBERTa-Large, with approximately 340 million parameters) pre-trained on more than 10TB of industrial text corpus (including patents, standards, journal articles, technical manuals, and internal corporate documents) is used as a general semantic understanding encoder; second, for the named entity recognition (NER) task, a BiLSTM-CRF layer is built on top of the pre-trained model and fine-tuned on domain-specific annotated data (at least 200,000 entity annotations) to achieve accurate recognition of key entities of equipment, faults, processes, and safety measures (F1 value not less than 0.96); for the relation extraction (RE) task, an attention-based The mechanism-based graph convolutional network (GCN) or path-dependent LSTM model is trained on entity annotated data (at least 100,000 relationship annotations) to extract hierarchical, causal, and association relationships between entities (F1 value not less than 0.92). For the event extraction (EE) task, a model based on the dynamic multi-hop graph attention network (DMHAN) is used to identify key event trigger words and their argument roles (F1 value not less than 0.90). All extracted knowledge triples or event structures are used to assist in the construction and dynamic updating of industrial knowledge graphs.

[0073] The knowledge graph management and reasoning engine module uses a distributed graph database system (Janus Graph combined with HBase / Cassandra backend, or TigerGraph MPP architecture) to store and manage dynamically updated industrial knowledge graphs at the petabyte level; it provides high-performance graph traversal, indexing (including full-text indexing, geospatial indexing, and time series indexing), and query interfaces (supporting Gremlin and SPAR QL1.1); it has a built-in description logic reasoning engine based on RDFS and OWL2DL (an integrated version of Pellet or FaCT++), as well as a customizable rule engine based on SWRL or Drools. It can perform automated logical reasoning (attribute inheritance, relationship transfer, consistency checking, and new knowledge discovery) based on ontology definitions and user rules, and supports representation learning based on graph embeddings (RotatE, ComplEx, GraphSAGE), mapping entities and relationships to low-dimensional vector spaces for advanced graph analysis tasks such as similarity calculation, link prediction, node classification, and community discovery, providing deep, multi-dimensional background knowledge and contextual information for subsequent data fusion and prediction.

[0074] The collaborative control and system evolution unit further includes:

[0075] The distributed model deployment and management module is built on Kubernetes and Kubeflow Pipel ine, and realizes the automation and strategy (selection of deployment nodes based on resource utilization, latency requirements, and data limitations) of AI model deployment from training, verification, packaging (Docker containerization, including model files, dependent libraries, and runtime environment), registration (version control and metadata management in the model warehouse MLflowModelRegistry), to multi-target environments (device side, edge nodes, and cloud). It supports blue-green deployment, canary release, and A / B testing of models to enable model updates and effect evaluation without interrupting services. It provides a unified API gateway and monitoring dashboard for real-time monitoring and alarming of the running status (QPS, inference latency, error rate, resource consumption), input and output data distribution, and prediction drift of deployed models.

[0076] The Federated Learning Coordination and Security Aggregation Module, when the system uses a federated learning mechanism to collaboratively train a global model while protecting the data privacy of multiple parties (different subsidiaries, different production lines, or different suppliers), is responsible for:

[0077] Client selection and task distribution: Based on the availability, data volume, and computing power of the clients (edge nodes or facilities), a subset of clients is selected to participate in this round of training, and the global model and training tasks are distributed.

[0078] Secure parameter aggregation: Receive model updates (gradients, weights, or parameters processed with local differential privacy) uploaded by each client after training on local data through a secure communication channel (TLS1.3). Use a robust aggregation algorithm (FedAvg combined with Krum or Median's anti-poisoning mechanism, FedPro x to address client heterogeneity, or SCAFFOLD to address client drift) to perform weighted averaging or other forms of aggregation on the model updates to generate a new global model. Homomorphic encryption or secure multi-party computation (SMC) can be optionally applied during the aggregation process to enhance privacy protection.

[0079] Global model update and verification: The aggregated global model parameters are distributed to each client for the next round of local training or inference. At the same time, the performance of the updated global model is evaluated in a reserved global validation set or through a simulated environment. The entire process adheres to strict data minimization and purpose limitation principles to ensure that the original data does not leave the local server and only transmits necessary model information.

[0080] Closed-loop feedback and continuous learning engine. The core of this module is an event-driven, configurable automated workflow engine used to:

[0081] Multi-source feedback data collection and alignment: Real-time collection of system operation data (performance indicators, resource utilization), model prediction results (and their confidence and explanatory reports), manual intervention records (operator confirmation of warnings, disposal measures, and effect evaluation), and new true value labels or error samples annotated by the digital twin environment or domain experts. This feedback data is accurately aligned with the original input data, model version, and timestamp.

[0082] Performance drift detection and attribution analysis: Continuously monitors key models' prediction accuracy, recall, and F1 score performance metrics, as well as covariate drift and concept drift in data distribution. When significant performance degradation or data distribution changes are detected, the attribution analysis process is automatically triggered to attempt to locate the root cause of the problem (sensor aging, changed operating conditions, or model obsolescence).

[0083] Automated model retraining and validation: The model retraining process (including data preparation, feature engineering, model training, and hyperparameter optimization) is automatically initiated based on preset rules (performance below a threshold, new data accumulation reaching a certain amount) or manual instructions. After retraining, the model is rigorously evaluated on a new validation set and subjected to A / B testing or shadow deployment with the current online model. The new model is only replaced when it significantly outperforms the old one.

[0084] Continuous learning and knowledge increment: Using continuous learning technologies based on experience replay, elastic weight consolidation, knowledge distillation, or progressive networks, the model can effectively alleviate the problem of catastrophic forgetting when learning new tasks or adapting to new data, retain and accumulate historically learned knowledge and capabilities, and thus continuously optimize model parameters, knowledge base content, and control strategies at all levels to enhance the overall intelligence level of the system and its long-term adaptability to dynamically changing industrial environments. Specific embodiment three:

[0086] Based on the technical solutions of the first and second specific embodiments, further explanations are given in combination with actual conditions:

[0087] In large-scale integrated petrochemical complexes, the core cluster of equipment—a million-ton-capacity ethylene cracker and dozens of downstream large-scale continuous production units for polyolefins and aromatics—forms a massive, highly complex industrial system characterized by interconnected materials and coupled risks. A leak, fire, or explosion in this cluster of equipment could easily trigger a domino effect, resulting in catastrophic consequences. Traditional models based on single-parameter threshold alarms and decentralized emergency response plans are no longer sufficient to meet the holistic safety management requirements of such a complex, massive system. The application of this solution in this scenario first comprehensively deploys the multimodal data acquisition system described in Sp1 in key areas and equipment of cracking furnaces, compressor units, separation towers, storage tank areas, and pipeline corridors to obtain real-time data including high-definition / infrared video (covering key flanges, pump bodies, valve groups, and pipeline-dense areas for leak and flame identification), open optical path gas detector arrays (FTIR / TDLAS technology, used for hydrocarbon and toxic gas trace leak detection), distributed fiber optic vibration / temperature sensors (DVS / DTS, laid along pipeline corridors and key pipelines for abnormal vibration and temperature mutation monitoring), equipment operating parameters (such as cracking furnace outlet temperature and pressure, compressor vibration and shaft displacement, tower bottom liquid level and temperature), and dynamic distribution data of emergency resources (fire stations, rescue teams, material depots, evacuation routes, and medical points) combined with a three-dimensional geographic information system (3DGIS). Sp1's dynamic knowledge graph module combines the material balance relationship between devices, process flow dependencies, the protection layer logic of the safety instrumented system (SIS), the command hierarchy and resource scheduling rules in the emergency plan, and the potential accident scenarios and risk levels obtained from HAZOP / LOPA analysis into a dynamic safety knowledge network covering the entire device group. When Sp2's weak signal feature extraction and fusion module detects local overheating in the radiation section of a cracking furnace (indicating impending creep or coking of the furnace tubes), or weak energy of a specific fault frequency appears in the vibration spectrum of a large centrifugal compressor bearing (indicating early wear), or the open optical path gas detector in the pipeline corridor detects a micro-hydrocarbon concentration that is below the alarm threshold but persists, Sp3's hybrid intelligent prediction engine will combine the device association information in the knowledge graph and the accident propagation model (based on the dynamic risk propagation model of STDGAT) to rapidly predict the probability of the initial abnormal state evolving into a major accident (such as furnace tube rupture, compressor damage, or pipeline leakage causing fire and explosion), the final evolution path (furnace tube rupture causing high-temperature hydrocarbon material leakage, forming a combustible gas cloud upon contact with air, exploding upon encountering an ignition source, and the shock wave causing damage to adjacent pipelines or equipment, triggering secondary accidents), and the estimated critical time window (predicting that the risk of furnace tube rupture will reach an unacceptable level within 30 minutes).At the same time, Sp4's digital twin drive module will dynamically simulate the entire predicted accident evolution process in real time in a high-fidelity three-dimensional digital twin environment that is mapped 1:1 with the petrochemical base, and automatically optimize and generate a multi-level collaborative emergency response strategy covering the entire affected area based on deep reinforcement learning (PPO algorithm). This strategy not only includes emergency response measures for the source of the fault (such as automatically triggering the emergency shutdown procedure of the cracking furnace and starting the compressor standby unit), but more importantly, it can intelligently generate a coordinated emergency command plan for the entire area based on real-time wind direction, wind speed, leakage diffusion simulation results, and the dynamic availability of emergency resources (such as the coverage and effective range of fire monitors, the arrival time of rescue teams, and the safety of different evacuation routes), including: (1) accurately demarcating the warning area and evacuation range, and sending personalized evacuation instructions to personnel in the area through emergency broadcasts and personnel positioning systems; (2) automatically dispatching the fire rescue force with the closest distance and the best matching capabilities, and planning the optimal route; (3) intelligently controlling the emergency shut-off valves, drain valves, and steam curtain safety facilities in the accident area and upstream and downstream related devices to minimize the spread of the accident and reduce material leakage; (4) dynamically adjusting the production load of surrounding unaffected devices to ensure the supply security of important intermediate products or achieve orderly shutdown. Sp5's edge-end-cloud collaborative architecture ensures efficient closed-loop operation of the entire process, from weak signal detection to regional coordinated emergency command: edge nodes are responsible for real-time processing and rapid response of on-site data, while the cloud platform performs complex accident evolution prediction, digital twin simulation, and global emergency strategy optimization, and sends instructions to on-site execution units and mobile terminals of commanders at all levels through the industrial 5G network. The implementation of this solution will enhance the accident prevention capabilities of the petrochemical base from "post-event response" and "local disposal" to "pre-event prediction" and "whole-region coordination". It is expected to reduce the incidence of major accidents by at least 80%, shorten emergency response time by more than 50%, significantly reduce direct economic losses and indirect impacts (such as environmental pollution and social panic) caused by accidents, and comprehensively improve the base's inherent safety level and emergency management efficiency.

[0088] It should be noted that, in this article, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprising a reference structure" do not exclude the presence of other identical elements in the process, method, article or device comprising the elements.

[0089] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for industrial environment monitoring and accident prediction integrating multimodal data, characterized by: The industrial environment monitoring and accident prediction method includes the following: Sp1. Semantic Collection and Causal Association Preprocessing of Multimodal Heterogeneous Data Based on Dynamic Knowledge Graph: Real-time collection of multimodal heterogeneous data in industrial environments to obtain dynamic knowledge graphs; Sp2. Deep feature extraction and multimodal cognitive fusion for weak signals and potential risks: Deep learning models are used to extract deep abstract features that characterize early weak equipment failures, subtle environmental changes, and atypical human behavior based on data from different modalities. Sp3. Probabilistic prediction and explainability analysis of accident chain evolution paths based on hybrid intelligence and evolutionary learning: Build a hybrid intelligent prediction engine that integrates physical mechanism models, statistical models, and deep learning models to obtain probabilistic prediction results of accident chain evolution paths; Sp4. Dynamic optimization and verification of preventive intervention strategies driven by digital twins: Build a digital twin model that maps the physical industrial environment in real time with high fidelity, and input the probability prediction results of the accident chain evolution path into the digital twin model; Sp5, edge-end-cloud collaborative adaptive monitoring, closed-loop feedback and continuous system evolution: Collaboratively perform monitoring and prediction tasks on the device side, edge side and cloud side, collect execution effect data and new data of intervention strategies in the physical environment, and use them as feedback information for online updating of knowledge graphs, feature extraction and fusion models, prediction engines and reinforcement learning strategy networks.

2. The method for industrial environment monitoring and accident prediction integrating multimodal data according to claim 1 is characterized in that: The construction and update of the dynamic knowledge graph in Sp1 further includes: using natural language processing technology to automatically identify and extract entity concepts, attribute information, failure modes, causal relationships and operational constraints related to industrial safety, equipment health, and process flows from unstructured and semi-structured text data including maintenance records, operating manuals, safety regulations and accident investigation reports, and after verification and standardization of these extracted information, structured integration and dynamic update of the dynamic knowledge graph, thereby significantly enhancing the knowledge coverage breadth, detail richness, information real-timeness and accuracy of subsequent reasoning decisions of the knowledge graph.

3. The method for industrial environment monitoring and accident prediction integrating multimodal data according to claim 1, characterized in that: The deep feature extraction for weak signals in Sp2 further includes the following steps: Sp2.1, for signals containing high-frequency noise and non-stationary components in the collected original multimodal data; Sp2.

2. Input the features or raw data processed by Sp2.1 into a pre-trained or task-specific fine-tuned deep learning network, and learn and capture deep, high-dimensional, and more discriminative abstract feature representations that indicate the incipient state of early faults, accumulation of microscopic damage to materials, or slight deviations from operating procedures through multi-layer nonlinear transformations.

4. The method for industrial environment monitoring and accident prediction integrating multimodal data according to claim 1, characterized in that: When faced with challenges such as sparse historical data, frequently changing operating conditions, or the presence of unknown new accident patterns, the hybrid intelligent prediction engine in Sp3 can further achieve its evolutionary learning capabilities through at least one of the following methods: Launch a rapid adaptation module based on meta-learning, where the prediction model uses new target domain samples to quickly adjust parameters to adapt to new working conditions or accident types; A knowledge transfer module based on transfer learning is applied. The knowledge transfer module selectively transfers and adapts the accident pattern knowledge, fault feature representation or model parameters learned in the source domain to the prediction model in the target domain by measuring and reducing the data distribution differences or feature space differences between the source domain and the target domain.

5. The method for industrial environment monitoring and accident prediction integrating multimodal data according to claim 1, characterized in that: The digital twin model in Sp4 further includes: integrating a multi-physics field coupling simulation engine based on first principles, and providing a configurable parameterized scenario construction interface.

6. The method for industrial environment monitoring and accident prediction integrating multimodal data according to claim 1, characterized in that: The adaptive monitoring, closed-loop feedback and continuous system evolution of edge-end-cloud collaboration in Sp5 are specifically implemented as follows: deploy lightweight edge AI models on the device side or embedded system close to the data source to perform data quality verification, basic feature extraction and simple abnormal event screening and rapid response with extremely low latency; deploy edge computing nodes at the industrial site or workshop level, responsible for more complex multimodal data fusion, real-time risk assessment, short-term accident prediction and execution of localized preventive control instructions for the aggregated end device data and local sensor data; and run large-scale, highly complex global knowledge graph management, deep learning model training and updating, long-term trend analysis, complex accident chain deduction, global resource optimization scheduling and cross-factory / cross-enterprise federated learning coordination on the cloud computing platform.

7. A system corresponding to the method for industrial environment monitoring and accident prediction based on the multimodal data integration method according to any one of claims 1 to 6, characterized in that: The industrial environment monitoring and accident prediction system integrating multimodal data includes: Multimodal data semantic collection and preprocessing unit, used to execute Sp1; Deep feature extraction and cognitive fusion unit for executing Sp2; Accident evolution prediction and explainability analysis unit, used to execute Sp3; Digital twin driven strategy optimization and verification unit, used to execute Sp4; Cooperative control and system evolution unit, used to execute Sp5; As well as sensor networks, industrial control system interfaces, distributed edge computing devices, cloud computing platforms, and human-computer interaction and visualization interfaces connected to the above units.

8. The industrial environment monitoring and accident prediction system integrating multimodal data according to claim 7 is characterized in that: The multimodal data semantic collection and preprocessing unit further includes: Multi-source heterogeneous data access module, used to collect multi-source heterogeneous data in real time and batch from sensor networks, PLC / DCS systems, MES systems, video surveillance systems, and manual record databases through standard industrial protocols or customized interfaces; A natural language processing and knowledge extraction module is configured to use machine learning models and semantic analysis algorithms to automatically extract key entities, attributes, relationships, events, and rules from maintenance logs, operating procedures, safety standards, and accident report text data to assist in building and updating industrial domain knowledge graphs; The knowledge graph management and reasoning engine module is used to store and manage dynamically updated knowledge graphs, and perform semantic queries, logical reasoning, and data association analysis based on the graphs, providing rich background knowledge and contextual information for subsequent data fusion and prediction.

9. The industrial environment monitoring and accident prediction system integrating multimodal data according to claim 7, characterized in that: The collaborative control and system evolution unit further includes: Distributed model deployment and management module, used to safely and efficiently deploy trained or continuously updated monitoring models, prediction models, and control strategy models to corresponding devices, edge computing nodes, and cloud platforms, and manage the model versions, configurations, and operating status; Federated Learning Coordination and Security Aggregation Module: When the system adopts the federated learning mechanism, this module is responsible for coordinating the local model training process of each participant, securely collecting and aggregating the model updates contributed by all parties, updating the global shared model, and distributing the updated model back to each participant; Closed-loop feedback and continuous learning engine, used to collect system operation data, model prediction results, manual intervention records and strategy execution effects through online learning, incremental learning or reinforcement learning mechanisms.

Citation Information

Patent Citations

  • Intelligent vibratory digital twinning system and method for industrial environments

    CN115039045A

  • Method for constructing power plant power generation equipment fault knowledge graph

    CN116521898A

  • Operation safety risk identification method based on multi-modal knowledge graph

    CN117408507A

  • Knowledge graph driven semantic governance scheme dynamic generation method and system

    CN118114758A

  • Intelligent fault diagnosis method and system for explosion-proof motor in natural gas industry

    CN119004265A

Cited By

  • Digital training method and system based on safety production

    CN120725835A

  • A digital training method and system based on safety production

    CN120725835B

  • Wheeled inspection robot and camera motion understanding method and system thereof

    CN120766376A

  • Potential disaster intelligent sensing and emergency data engineering system driven by multi-modal data

    CN120769245A

  • Tracking method based on full-automatic flow monitoring of lithium battery positive electrode material production process

    CN120851571A