Industrial risk processing method and device based on multi-source data fusion, equipment and medium

By using a multi-source data fusion method, spatiotemporal alignment and anomaly identification are performed on sensor and video stream data in key industrial areas. This solves the problem of low reliability in risk handling in existing technologies, enables accurate identification and rapid response to industrial risks, and improves safety and management efficiency.

CN121860404APending Publication Date: 2026-04-14CHINA NATIONAL OFFSHORE OIL (CHINA) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies for handling industrial risks have low reliability, relying on manual inspections and single-point sensor alarms. These methods suffer from monitoring blind spots, delayed responses, and underutilization of multi-source data, making it difficult to achieve efficient spatiotemporal alignment and feature correlation, resulting in a high rate of missed detections of safety incidents.

Method used

By using a multi-source data fusion method, environmental data and video stream data collected by sensors in key industrial areas are spatiotemporally aligned to perform target detection and behavioral anomaly identification. Data anomaly identification is performed by combining environmental parameter thresholds, and risk levels are determined through cross-validation, thereby achieving accurate risk identification and response.

Benefits of technology

It improves the reliability of industrial risk handling, enables accurate detection of hidden equipment failures, personnel violations and high-risk environmental events, builds a comprehensive risk proactive defense system, and significantly improves the safety level of industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121860404A_ABST
    Figure CN121860404A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial risk processing, and discloses an industrial risk processing method, device and equipment based on multi-source data fusion and a medium, which can perform space-time alignment on environment data and video stream data acquired by a plurality of sensors at a plurality of continuous timestamps in an industrial key area to obtain multi-source alignment data. Performing target detection and behavior anomaly recognition on the video stream data in the multi-source alignment data to obtain a behavior recognition result; and according to a set environment parameter threshold value, performing data exception identification on the environment data in the multi-source alignment data to obtain a data identification result. And performing cross verification on the behavior identification result and the data identification result to determine the risk level of the industrial key area at the first timestamp. And determining whether to execute early warning and response actions or not according to the risk level of the industrial key area at the first timestamp. According to the invention, the reliability of industrial risk processing can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial risk management, and in particular to an industrial risk management method, apparatus, equipment and medium based on multi-source data fusion. Background Technology

[0002] With the improvement of industrial automation and intelligence, industrial risk management technology is constantly being improved.

[0003] In high-risk industrial sectors such as petrochemicals and power energy, safety accidents often lead to catastrophic consequences. Identifying potential industrial risks and taking corresponding risk management measures are effective ways to deal with industrial risks and help improve industrial safety.

[0004] Related technologies generally rely on manual inspections and single-point sensor alarms to identify potential industrial risks and issue risk warnings, resulting in low reliability in industrial risk handling. Summary of the Invention

[0005] This invention provides an industrial risk processing method, apparatus, equipment, and medium based on multi-source data fusion, which addresses the shortcomings of low reliability in related technologies and improves the reliability of industrial risk processing.

[0006] In a first aspect, the present invention provides an industrial risk processing method based on multi-source data fusion, comprising: Spatiotemporal alignment of environmental data and video stream data collected by multiple sensors in key industrial areas at multiple consecutive time stamps is performed to obtain multi-source aligned data. Target detection and behavior anomaly identification are performed on the video stream data in the multi-source aligned data to obtain behavior identification results; and, based on a set environmental parameter threshold, data anomaly identification is performed on the environmental data in the multi-source aligned data to obtain data identification results. Cross-validate the behavior recognition results and the data recognition results to determine the risk level of the key industrial area at the first time stamp; Based on the risk level of the key industrial areas at the first time stamp, determine whether to implement early warning and response actions.

[0007] Optionally, the plurality of sensors include a plurality of camera devices and environmental sensors disposed at different locations in the industrial critical area; The process involves spatiotemporally aligning environmental data and video stream data collected by multiple sensors in key industrial areas at consecutive time stamps to obtain multi-source aligned data, including: Acquire the location information of each of the camera devices and the video stream data collected by each of the camera devices at each timestamp; and acquire the location information of each of the environmental sensors and the environmental data collected by each of the environmental sensors at each timestamp; Based on the location information of each camera device and each environmental sensor, determine the camera devices and environmental sensors whose location information matches, and group the camera devices and environmental sensors whose location information matches as a sensor group to obtain multiple sensor groups; For any of the sensor groups, the position information of the camera device and the environmental sensor in the sensor group, the video stream data collected by the camera device at the timestamp, and the environmental data collected by the environmental sensor at the same timestamp are taken as the multi-source aligned sub-data corresponding to the timestamp. The multi-source aligned sub-data corresponding to each timestamp is taken as the whole as the multi-source aligned data.

[0008] Optionally, the step of performing target detection and behavior anomaly recognition on the video stream data in the multi-source aligned data to obtain behavior recognition results includes: Target detection and behavior anomaly recognition are performed on the video stream data collected by each camera device at each timestamp in the multi-source aligned data to obtain the behavior recognition result corresponding to each video stream data. The step of identifying data anomalies in the multi-source aligned data based on a set environmental parameter threshold, and obtaining data identification results, includes: Based on the set environmental parameter thresholds, anomaly identification is performed on the environmental data collected by each environmental sensor at each timestamp in the multi-source aligned data to obtain the data identification result corresponding to each environmental data.

[0009] Optionally, the environmental sensor is used to collect data on at least one of temperature, smoke, and gas concentration in the industrial critical area.

[0010] Optionally, the cross-validation of the behavior recognition result and the data recognition result to determine the risk level of the industrial critical area at the first time stamp includes: For the video stream data and environmental data collected by the camera device and the environmental sensor in any of the sensor groups at the first time stamp, if the behavior recognition result corresponding to the video stream data and the data recognition result corresponding to the environmental data are both abnormal, then the risk level of the industrial critical area at the first time stamp is determined to be high. If both the behavior recognition result corresponding to the video stream data and the data recognition result corresponding to the environmental data are normal, then the risk level of the key industrial area at the first time stamp is determined to be low. If one of the behavior recognition results corresponding to the video stream data and the data recognition results corresponding to the environmental data is normal and the other is abnormal, then the risk level of the industrial critical area at the first time stamp is determined to be medium.

[0011] Optionally, determining whether to execute early warning and response actions based on the risk level of the industrial critical area at the first time stamp includes: If the risk level of the industrial critical area is low at the first time stamp, then early warning and response actions are prohibited. If the risk level of the industrial critical area is medium or high at the first time stamp, then a warning mode corresponding to the risk level is determined, and a warning and corresponding response action are executed according to the warning mode.

[0012] Optionally, the response action includes at least one of pushing alarm information to the management platform, controlling the start and stop of field equipment, and activating an audible and visual alarm.

[0013] Secondly, the present invention provides an industrial risk processing device based on multi-source data fusion, comprising: The alignment unit is used to perform spatiotemporal alignment of environmental data and video stream data collected by multiple sensors in critical industrial areas at multiple consecutive time stamps to obtain multi-source aligned data. The first identification unit is used to perform target detection and behavior anomaly identification on the video stream data in the multi-source aligned data to obtain behavior identification results; The second identification unit is used to identify data anomalies in the environmental data in the multi-source aligned data according to the set environmental parameter threshold, and obtain the data identification result. A verification unit is used to cross-verify the behavior recognition result and the data recognition result to determine the risk level of the industrial critical area at the first time stamp. The determining unit is used to determine whether to execute early warning and response actions based on the risk level of the industrial critical area at the first time stamp.

[0014] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the industrial risk processing method based on multi-source data fusion described in the first aspect or any corresponding embodiment thereof.

[0015] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the industrial risk processing method based on multi-source data fusion described in the first aspect or any corresponding embodiment thereof.

[0016] This invention provides an industrial risk management method, apparatus, equipment, and medium based on multi-source data fusion. It can spatiotemporally align environmental data and video stream data collected by multiple sensors in a critical industrial area at consecutive time stamps to obtain multi-source aligned data. Target detection and behavioral anomaly identification are performed on the video stream data within the multi-source aligned data to obtain behavioral identification results. Furthermore, based on set environmental parameter thresholds, data anomaly identification is performed on the environmental data within the multi-source aligned data to obtain data identification results. The behavioral identification results and data identification results are cross-validated to determine the risk level of the critical industrial area at a first time stamp. Based on the risk level of the critical industrial area at the first time stamp, it is determined whether to execute early warning and response actions. This invention can spatiotemporally align environmental data and video stream data collected by sensors to obtain multi-source aligned data, determine the risk level of a critical industrial area at different time stamps based on the multi-source aligned data, and determine whether to execute early warning and response actions, effectively improving the reliability of industrial risk management. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating an industrial risk processing method based on multi-source data fusion, provided as an embodiment of the present invention; Figure 2 A flowchart illustrating another industrial risk processing method based on multi-source data fusion provided in this embodiment of the invention; Figure 3 A flowchart illustrating another industrial risk management method based on multi-source data fusion provided in this embodiment of the invention. Figure 4 A schematic diagram of the structure of an industrial risk processing device based on multi-source data fusion provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0020] The following is combined Figures 1-3 This invention describes an industrial risk management method based on multi-source data fusion.

[0021] like Figure 1 As shown, this embodiment proposes a first industrial risk processing method based on multi-source data fusion, which may include the following steps: S101. Spatiotemporally align environmental data and video stream data collected by multiple sensors in key industrial areas at multiple consecutive time stamps to obtain multi-source aligned data.

[0022] Among them, key industrial areas can be the risk areas within the factory that require special attention.

[0023] The sensors may include environmental sensors and camera devices.

[0024] Specifically, a timestamp can be a time period of a specific duration, such as 1 minute, 5 minutes, or 20 seconds.

[0025] Optionally, the multiple sensors include multiple camera devices and environmental sensors positioned at different locations in a critical industrial area. Step S101 includes: Acquire the location information of each camera device and the video stream data collected by each camera device at each timestamp; and acquire the location information of each environmental sensor and the environmental data collected by each environmental sensor at each timestamp; Based on the location information of each camera device and each environmental sensor, determine the camera devices and environmental sensors that match the location information, and group the camera devices and environmental sensors that match the location information into a sensor group to obtain multiple sensor groups; For any sensor group, the location information of the camera device and the environmental sensor in the sensor group, the video stream data collected by the camera device at the timestamp, and the environmental data collected by the environmental sensor at the same timestamp are taken as the multi-source aligned sub-data corresponding to the timestamp. The multi-source aligned sub-data corresponding to each timestamp is treated as a whole as multi-source aligned data.

[0026] Specifically, in this embodiment, the location information of the camera device and the environmental sensor in the same sensor group, and the video stream data and environmental data collected by the camera device and the environmental sensor at the same time stamp are used as the multi-source aligned sub-data corresponding to that time stamp.

[0027] S102. Perform target detection and behavior anomaly recognition on the video stream data in the multi-source aligned data to obtain behavior recognition results.

[0028] Specifically, this embodiment can perform target detection and behavior anomaly recognition on each video stream data in the multi-source aligned data to obtain the behavior recognition result corresponding to each video stream data.

[0029] S103. Based on the set environmental parameter thresholds, perform data anomaly identification on the environmental data in the multi-source aligned data to obtain the data identification results.

[0030] Optionally, environmental sensors are used to collect data on at least one of temperature, smoke, and gas concentration in critical industrial areas.

[0031] Specifically, when environmental sensors are used to collect multiple environmental parameters, the environmental parameter thresholds can include the data threshold for each environmental parameter.

[0032] Optionally, in other industrial risk processing methods based on multi-source data fusion proposed in this embodiment, step S102 includes: Target detection and behavior anomaly recognition are performed on the video stream data collected by each camera device at each time stamp in the multi-source aligned data to obtain the behavior recognition result corresponding to each video stream data.

[0033] Optionally, step S103 includes: Based on the set environmental parameter thresholds, anomaly identification is performed on the environmental data collected by each environmental sensor at each time stamp in the multi-source aligned data, and the data identification result corresponding to each environmental data is obtained.

[0034] Specifically, this embodiment can perform target detection and behavior anomaly identification on each video stream data in the multi-source aligned data, and perform anomaly identification on each environmental data in the multi-source aligned data, to obtain the data identification results corresponding to each video stream data and each environmental data.

[0035] S104. Cross-validate the behavior recognition results and data recognition results to determine the risk level of key industrial areas at the first time stamp.

[0036] The first timestamp can be one of the timestamps mentioned above.

[0037] Specifically, this embodiment can perform target detection and behavior anomaly identification on video stream data in multi-source aligned data to obtain behavior identification results, and perform data anomaly identification on environmental data in multi-source aligned data to obtain data identification results. After obtaining data identification results, the behavior identification results and data identification results are cross-validated to determine the risk level of key industrial areas at the first time stamp.

[0038] Optionally, step S104 includes: For video stream data and environmental data collected by camera devices and environmental sensors in any sensor group at the first time stamp, if the behavior recognition result corresponding to the video stream data and the data recognition result corresponding to the environmental data are both abnormal, then the risk level of the industrial critical area at the first time stamp is determined to be high. If both the behavior recognition result corresponding to the video stream data and the data recognition result corresponding to the environmental data are normal, then the risk level of the key industrial area at the first time stamp is determined to be low. If the behavior recognition result corresponding to the video stream data and the data recognition result corresponding to the environmental data are both normal and abnormal, then the risk level of the key industrial area at the first time stamp is determined to be medium.

[0039] Understandably, risk levels can also be classified into different levels in other ways, such as high level, higher level, lower level and no risk.

[0040] S105. Based on the risk level of key industrial areas at the first time stamp, determine whether to implement early warning and response actions.

[0041] Specifically, this embodiment can determine whether to execute early warning and response actions based on the risk level of key industrial areas at the first time stamp.

[0042] Optionally, step S105 includes: If the risk level of a critical industrial area is low at the first time stamp, then early warning and response actions are prohibited. If the risk level of a critical industrial area is medium or high at the first time stamp, then the warning method corresponding to the risk level is determined, and the warning and corresponding response actions are executed according to the set warning method.

[0043] Optionally, the response action includes at least one of pushing alarm information to the management platform, controlling the start and stop of field equipment, and activating an audible and visual alarm.

[0044] The industrial risk management method based on multi-source data fusion proposed in this embodiment can spatiotemporally align environmental data and video stream data collected by multiple sensors in critical industrial areas at consecutive time stamps to obtain multi-source aligned data. Target detection and behavioral anomaly identification are performed on the video stream data in the multi-source aligned data to obtain behavioral identification results; and data anomaly identification is performed on the environmental data in the multi-source aligned data based on set environmental parameter thresholds to obtain data identification results. The behavioral identification results and data identification results are cross-validated to determine the risk level of the critical industrial area at the first time stamp. Based on the risk level of the critical industrial area at the first time stamp, it is determined whether to execute early warning and response actions. This embodiment can spatiotemporally align environmental data and video stream data collected by sensors to obtain multi-source aligned data, determine the risk level of the critical industrial area at different time stamps based on the multi-source aligned data, and determine whether to execute early warning and response actions, which can effectively improve the reliability of industrial risk management.

[0045] The current model of manual inspection combined with single-point sensor alarms faces three major challenges: First, the dynamic complexity of high-risk scenarios is rapidly increasing. Modern industrial sites have high equipment density and strong process chain coupling, meaning a single equipment failure can trigger a chain reaction. Point sensors can only capture local parameters, failing to adequately perceive spatial dynamic risks such as minor equipment damage, micro-leaks in pipelines, and personnel trespassing into restricted areas, creating monitoring blind spots. Second, safety standards and cost pressures are driving technological upgrades. Manual inspections suffer from fatigue blind spots and response delays, and the cost of skilled workers is rising, making it urgent for enterprises to replace inefficient manual labor with intelligent methods. Third, the value of multi-source data is not being fully utilized. Although factories have deployed video surveillance, IoT sensors, and positioning systems, the problem of data silos is prominent: video monitoring rooms require dedicated personnel to monitor screens, and effective images are often overlooked; after a sensor alarm, video footage needs to be retrieved across departments for verification, delaying crucial decision-making time. The industry urgently needs a technology that can link equipment status, environmental parameters, and personnel behavior in real time to build a proactive, preventative safety system.

[0046] With the improvement of industrial automation and intelligence, monitoring and early warning systems in industrial scenarios are gradually evolving from single data sources to multi-source data fusion. However, existing technologies still face significant bottlenecks in multi-source data collaborative processing, real-time performance, and accurate identification of abnormal events. Traditional monitoring systems typically rely on a single data source (such as sensors or video streams) or use simple overlay methods to process multi-source data, resulting in low data utilization efficiency. For example, in industrial scenarios, sensor data (temperature, pressure, gas concentration, etc.) and video stream data exhibit heterogeneity in sampling frequency, spatiotemporal resolution, and format, making it difficult to achieve efficient spatiotemporal alignment and feature correlation. Some technologies attempt to fuse data through fixed thresholds or static rules, but lack dynamic adaptive capabilities and cannot cope with interference in complex industrial environments (such as rainy / foggy weather, changes in lighting).

[0047] Deep learning models in the related art (such as object detection algorithms based on convolutional neural networks) face the problems of small samples and long-tailed distribution in industrial scenarios. For example, the data of major safety accidents is scarce, resulting in insufficient recognition ability of the model for rare events. At the same time, complex models (such as three-dimensional convolutional networks) have high computing power requirements for edge devices and are difficult to achieve real-time inference under low-power conditions. Therefore, there is an urgent need for an intelligent monitoring and early warning method based on deep fusion of multi-source data, supporting real-time linkage analysis and dynamic optimization, so as to improve the safety and management efficiency of the industrial environment.

[0048] Industrial safety monitoring is crucial in high-risk scenarios such as petrochemical and power energy industries. The core is to achieve accurate risk perception and rapid response through the coordination of multi-source data. In this context, this embodiment can adopt an intelligent monitoring method based on multi-source data fusion and video linkage. By connecting sensor thresholds, video dynamic analysis, and device positioning information, it realizes a closed-loop management of instant anomaly perception - automatic video review - accurate risk positioning, and essentially solves the rigid requirements of high-risk industrial scenarios for real-time performance, accuracy, and cost controllability.

[0049] The technical problems to be solved in this embodiment are as follows: The general intelligent detection and early warning methods have weak video analysis capabilities. The related technical algorithms are sensitive to dynamic lighting, smoke occlusion, and small target anomalies (such as initial fire points). The general detection models are not adapted to the characteristics of industrial scenarios, and have a high miss detection rate for tasks such as equipment displacement and personnel violations. Moreover, the computing power load restricts edge deployment. The early warning mechanism is lagging. After the sensor gives an alarm, it is necessary to manually retrieve the video for review, and the response delay time is long. In addition, there is a lack of multi-modal cross-verification for low-frequency and high-loss events such as slow pipeline leaks, and the risk of misjudgment is prominent. Although the related art attempts to introduce multi-source fusion, the effect is limited due to the lack of spatio-temporal dynamic calibration, the video model not being optimized for industrial events, the broken decision chain, and the rough edge-cloud collaboration. Especially under the constraints of mandatory standards in high-risk industries, it is urgent to break through deep cross-modal fusion and lightweight edge intelligence technologies.

[0050] The intelligent monitoring and early warning method for industrial scenarios based on multi-source data fusion proposed in this embodiment can include a multi-modal intelligent perception system, intelligent video detection for industrial scenarios, intelligent response and alarm functions, and on-site real-time alarm video linkage. By automatically identifying and analyzing videos and linkage early warning, it can effectively improve the inspection efficiency, reduce the inspection workload of on-site personnel, lower the risk of personnel inspection, and focus on solving the deficiencies of traditional manual inspection modes and the risks of video monitoring, and construct a "anomaly perception - video review - decision execution" closed-loop mechanism to fill the technical gap.

[0051] The multimodal intelligent sensing system achieves precise spatiotemporal alignment of sensor data streams, video frame sequences, and device positioning coordinates by constructing a spatiotemporal collaborative fusion framework. In the temporal dimension, a millisecond-level synchronization network is established using the IEEE 1588v2 precision clock protocol to unify the clock source and embed microsecond-level timestamps into industrial sensor data streams such as temperature, vibration, and gas sensors. A hardware-level video encoder writes frame-level timecodes into each frame. Combined with the TDoA technology of the Ultra Wideband (UWB) positioning system, nanosecond-level spatiotemporal markers are bound to the mobile device coordinates. In the spatial dimension, a multimodal calibration matrix engine is deployed: solving the camera's intrinsic and extrinsic parameter matrices and establishing a pixel coordinate system to an affine transformation model; and constructing the sensor's physical coordinate mapping relationship through UWB anchor point triangulation. Addressing the unique challenges of industrial scenarios, a spatiotemporal alignment execution kernel is designed: employing a sliding time window mechanism to dynamically aggregate multi-source data packets, adaptively downsampling high-frequency sensor data, and simultaneously performing motion trajectory prediction interpolation for occluded targets based on Kalman filtering.

[0052] This intelligent video detection module for industrial scenarios deeply integrates a lightweight AI engine and edge computing architecture to achieve real-time intelligent analysis of industrial video streams. Based on a deeply optimized YOLOv8 framework, it compresses the model size to 1 / 5 of the native architecture through deep separable convolution and structured channel pruning techniques. Combined with the INT8 quantization engine, it achieves ultra-low latency inference of 15ms per frame at edge nodes. Dedicated detection capabilities are built for high-risk industrial scenarios, developing a sub-pixel-level optical flow analysis algorithm to accurately capture subtle equipment angular shifts. It integrates visible light and thermal imaging dual-stream data to improve the initial flame detection distance to 50 meters. An integrated Long Short-Term Memory (LSTM) temporal behavior model identifies violations where the user lingers in prohibited areas for more than 3 seconds. An innovative multi-scale enhancement mechanism is designed: shallow feature feedback paths enhance the preservation of small target features, and attention-guided upsampling technology stably identifies 5×5 pixel-level targets (such as welding slag fire points) in 4K resolution video. A dynamic resolution switching strategy balances computing power and accuracy.

[0053] like Figure 2As shown, the intelligent response and alarm module is used to construct a multimodal cross-validation decision chain: it receives sensor threshold alarm signals and video detection results, and performs bidirectional cross-validation. If the sensor is abnormal (such as a sudden high temperature) and the video synchronously captures a related event (such as a pipeline leak or fire), the system triggers a level one alarm within 0.5 seconds, linking the fire protection system to start. If only a single source of data is abnormal, the intelligent verification channel is activated—dynamically retrieving the equipment's operating pressure curve for historical comparison, activating real-time footage from multi-view cameras within a 15-meter radius, generating a verification report based on a Bayesian model that includes sensor time sequence diagrams, video keyframes, and risk probability values, and pushing it to the console before starting a 15-second countdown for manual confirmation to avoid false alarms. It supports tiered alarm strategies (level one, level two, and level three alarms) and initiates emergency response according to risk level. The on-site real-time alarm video linkage module is used to achieve a visualized closed loop for alarm events: When an alarm is triggered, the system completes a three-level linkage response within 0.3 seconds: automatically retrieves real-time footage and historical video clips from associated cameras (e.g., 30 seconds before the alarm), retrieves the optimal three-view combination (the main view is a fixed bullet camera capturing the leak point coordinates ±0.1m, the secondary view is a pan-tilt-zoom camera tracking the diffusion path, and the third view transmits thermal imaging or other data); simultaneously, it accurately extracts historical video from the 30 seconds before the alarm using the inverted index of edge node timestamps to locate the key frame event trigger. A panoramic event report is generated; the alarm location, video status, and handling plan are dynamically displayed through the positioning management platform, automatically merging multi-source data to generate a dynamic event report, integrating sensor sudden change curves, video annotations of leak morphology, equipment 3D stress cloud maps, and fluid diffusion model predictions to provide decision support on the digital twin platform: the leak point is highlighted and flashed on the factory's 3D map to support rapid decision-making in the command center.

[0054] The target objects are categorized into equipment units, personnel units, and environmental anomaly units. When the target object is categorized as an equipment unit, it represents key equipment in an industrial setting. When the target object is categorized as a personnel unit, it represents staff or unauthorized personnel, including compliant personnel (wearing safety helmets and work uniforms) and non-compliant personnel (not wearing protective equipment or entering restricted areas). Operational risks are determined by behavioral patterns (duration of stay and movement path). When the target object is categorized as an environmental anomaly unit, it represents potential hazards, including hazardous ignition sources (visible flames, infrared hot spots), smoke leaks (visible light detection), gas diffusion (infrared detection), and crude oil leaks (infrared thermal imaging and optical cameras). For different categories of monitoring objects, accurate identification is achieved through multi-source data fusion.

[0055] The beneficial effects of this embodiment are as follows: Based on multi-source data fusion, it achieves accurate capture of hidden equipment faults, personnel violations, and high-risk environmental events; Based on the improved lightweight YOLOv8 model and multimodal cross-validation, it can identify abnormal equipment conditions, initial flames and smoke, and unauthorized personnel intrusion in real time, constructing a comprehensive risk proactive defense system and significantly improving the inherent safety level of industrial scenarios; It studies the automatic risk identification technology for general industrial scenarios and offshore platform-specific industrial scenarios based on various advanced machine vision technologies; Based on the design of a 3D model and network mapping layout interface, it features a real-scene mapping function that is linked with central control abnormal alarm data, fire alarm data, and equipment fault alarm data, facilitating quick retrieval of videos in designated areas for real-time on-site production monitoring; Based on video intelligent analysis technology and video linkage technology, it builds and develops a remote automatic video one-click inspection model, realizing a remote automatic inspection process for unmanned platforms. It supports managers in quickly tracing the source of risk events, intelligently scheduling resources, and accurately issuing emergency commands, optimizing the decision-making efficiency of the entire production and maintenance process, while providing high-confidence data chain support for accident investigation and safety standard iteration.

[0056] Based on the above technical solution, this embodiment may also include the following:

[0057] Furthermore, it also includes an incremental learning unit, which is connected to the intelligent video detection module for industrial scenarios. The incremental learning unit is used to receive the features of false detection samples after manual review and correction, generate a lightweight model update package through knowledge distillation technology, and synchronize it to the edge computing node. The beneficial effects of adopting the above-mentioned further solution are: to achieve continuous model optimization without full retraining, significantly improve the recognition accuracy of complex scenarios such as flames and smoke, and micro-displacement of equipment, while avoiding system interruptions caused by frequent large-scale updates at the edge.

[0058] Furthermore, it also includes a multi-link communication module, through which edge computing nodes connect to the cloud server; the multi-link communication module supports automatic switching between the main channel and the backup channel; the beneficial effect of adopting the above-mentioned further solution is that when the 5G signal is interfered with, the long-distance low-power backup link is automatically activated to ensure the reliability of data transmission, which is especially suitable for signal shielding scenarios such as metal-dense factory areas and underground pipe corridors.

[0059] Furthermore, it also includes a fault self-healing controller, with the edge computing node connected to the fault self-healing controller; the fault self-healing controller automatically switches to the neighboring node for proxy analysis when the node is offline, and triggers a retransmission mechanism when the data packet loss rate is >5%; the beneficial effects of adopting the above further solution are: single point of failure does not affect the global monitoring function, and at the same time, the analysis continuity is ensured through data interpolation compensation, avoiding the missed detection of critical events due to equipment downtime.

[0060] Furthermore, it also includes a multimodal feature fusion unit, which connects the intelligent video detection module for industrial scenarios and the platform's sensor network. The unit adopts a feature fusion algorithm based on an attention mechanism to fuse visible light video, infrared thermal imaging, and sound spectrum features in real time. The beneficial effect of adopting the above-mentioned further solution is that it significantly improves the accuracy and robustness of identifying hidden risks such as leakage, early overheating, and abnormal mechanical friction in complex lighting, occlusion, and high noise environments (such as strong wind and wave areas on offshore platforms and high salt spray environments).

[0061] Furthermore, it also includes an intelligent alarm convergence engine, which connects to the central control system's abnormal alarm data stream, the gas sensor data stream, and equipment fault diagnosis signals. The engine constructs a multi-source alarm association model based on a causal reasoning graph, automatically filters duplicate alarms, and infers the root cause fault location. The beneficial effects of adopting the above-mentioned further solutions are: reducing the number of invalid alarms, accurately locating core risk sources such as pump and valve leaks or bearing overheating, improving emergency response efficiency, and reducing the rate of misoperation.

[0062] like Figure 3 As shown, the system can also detect abnormal equipment operation (such as smoke or leaks) or personnel in hazardous areas / behaviors (such as trespassing into restricted areas or not wearing protective equipment) through real-time video analysis. Simultaneously, it utilizes a sensor network (such as temperature, pressure, gas concentration, and infrared proximity sensors) to monitor whether key parameters exceed preset safety thresholds. When the video recognition result and the sensor threshold trigger event are highly correlated in time and space and meet preset cross-validation conditions—that is, the video event and the sensor event occur almost simultaneously, and the events occur in the same or closely related physical areas (e.g., the video detects smoke from equipment and the temperature sensor simultaneously triggers an over-temperature alarm)—the system automatically determines it as a valid alarm event. For valid alarm events, the system automatically triggers a multi-level response mechanism: on the one hand, it automatically retrieves and highlights real-time and historical monitoring videos of the relevant areas. The purpose is to allow safety personnel to quickly and intuitively confirm the on-site situation and understand the background and development of the event; on the other hand, it immediately activates the alarm device, entering a tiered alarm mechanism. Multi-level alarm mechanisms can be set up to more clearly express the severity and urgency of risks. When video detection and sensor thresholds do not reach a unified judgment on a certain matter, the alarm level can be specifically indicated and reduced to reduce resource waste and false judgment rate. Ultimately, a complete closed loop of "perception-verification-response-handling" is formed, which effectively improves the accuracy of safety warnings and response efficiency.

[0063] Multimodal intelligent sensing systems can also achieve precise spatiotemporal alignment of multi-source data by synchronizing sensor data streams, video frame sequences, and device positioning coordinates with spatial coordinate systems through timestamp synchronization.

[0064] The intelligent video detection module for industrial scenarios can also be used for real-time intelligent analysis of video streams: it integrates an improved lightweight YOLOv8 model, which is customized for training in high-risk industrial events, to detect abnormal equipment displacement, initial flames and smoke, and personnel violations (such as not wearing safety helmets or entering restricted areas) in real time; it adopts a multi-scale feature fusion mechanism to enhance the recognition accuracy of small target anomalies; and it deploys the model based on edge computing nodes to achieve millisecond-level localized analysis of video streams.

[0065] The intelligent response and alarm module can also be used to build a multimodal cross-validation decision chain: it receives sensor threshold alarm signals and video detection results, and performs bidirectional cross-validation: if the sensor is abnormal (such as a sudden high temperature) and the video synchronously captures a related event (such as a pipeline leak or fire), an alarm is triggered; if only a single source of data is abnormal, the video manual review channel is activated to avoid false alarms; it supports hierarchical alarm strategies (level 1 alarm, level 2 alarm, level 3 alarm) and activates emergency response according to risk level.

[0066] The on-site real-time alarm video linkage module can also be used to realize the visual closed loop of alarm events: when an alarm is triggered, it automatically retrieves real-time images and historical video clips from associated cameras (e.g., 30 seconds before the alarm) and generates a panoramic event report; the alarm location, video status and handling plan are dynamically displayed through the positioning management platform to support the command center in making rapid decisions.

[0067] The target objects are categorized into equipment units, personnel units, and environmental anomaly units. When the target object is categorized as an equipment unit, it represents key equipment in an industrial setting. When the target object is categorized as a personnel unit, it represents staff or unauthorized personnel, including compliant personnel (wearing safety helmets and work uniforms) and non-compliant personnel (not wearing protective equipment or entering restricted areas). Operational risks are determined by behavioral patterns (duration of stay and movement path). When the target object is categorized as an environmental anomaly unit, it represents potential hazards, including hazardous ignition sources (visible flames, infrared hot spots), smoke leaks (visible light detection), gas diffusion (infrared detection), and crude oil leaks (infrared thermal imaging and optical cameras). For different categories of monitoring objects, accurate identification is achieved through multi-source data fusion.

[0068] Optionally, the system also includes a fault self-healing controller, with the edge computing node connected to the fault self-healing controller; the fault self-healing controller automatically switches to the neighboring node for proxy analysis when the node is offline, and triggers a retransmission mechanism when the packet loss rate is >5%.

[0069] Optionally, the system also includes a 3D visualization engine. The positioning management platform integrates the 3D visualization engine and overlays sensor data curves, alarm points, and risk diffusion models.

[0070] Optionally, the system also includes a multimodal feature fusion unit, which is connected to the intelligent video detection module for industrial scenarios and the platform's sensor network. The unit adopts a feature fusion algorithm based on an attention mechanism to fuse visible light video, infrared thermal imaging, and sound spectrum features in real time.

[0071] Optionally, the above system also includes an intelligent alarm convergence engine, which connects to the central control abnormal alarm data stream, the fire sensor data stream, and the equipment fault diagnosis signal; the engine constructs a multi-source alarm association model based on a causal reasoning graph, automatically filters duplicate alarms, and infers the root cause fault location.

[0072] Optionally, the system also includes an explosion-proof edge computing unit. The edge computing nodes are deployed inside the explosion-proof edge computing unit, which is equipped with a positive pressure ventilation system and an intrinsically safe heat dissipation module. The industrial scene intelligent video detection module can operate safely in flammable and explosive environments (such as chemical plant areas and oil and gas storage tank areas) through this unit.

[0073] Optionally, the system is configured with an adaptive analysis strategy engine: automatically switching between visible light and infrared modes based on ambient light intensity: enabling visible light mode to identify device displacement when the light intensity is >1000 Lux, and enabling infrared mode to detect thermal anomalies when the light intensity is <100 Lux. The smoke recognition threshold is dynamically adjusted based on meteorological sensor data to eliminate haze interference.

[0074] Optionally, the positioning management platform integrates a digital twin engine: it establishes a 3D factory model and maps the physical entity status in real time; it dynamically presents the equipment temperature gradient through color coding, and supports clicking on the equipment model to retrieve real-time video, operating parameters, and maintenance history.

[0075] It should be noted that the working principle of the above system is as follows: First, a multi-source heterogeneous data fusion framework is constructed, integrating sensor networks. Millisecond-level alignment of flame and gas alarms, temperature and pressure sensors, dual-mode cameras, and infrared thermal imaging is achieved via the IEEE 1588v2 protocol. A unified spatial coordinate system is established, and a sliding window algorithm is used to solve the frequency domain coordination between sensors and video streams. An edge-optimized target detection engine is deployed: through structured channel pruning and a lightweight compression-improved YOLOv8 model, real-time identification of abnormal equipment displacement, pipeline leaks, flames and smoke, and personnel violations is achieved. Furthermore, sensor thresholds are combined for multimodal cross-validation to develop a multi-source data intelligent linkage strategy. This enables cross-modal collaborative decision-making, integrating sensor data (such as temperature, gas concentration, and fire alarms), video streams, and equipment positioning information from industrial scenarios. Dynamic alignment and feature association of multi-source data are achieved, resolving collaboration barriers caused by differences in data sampling frequency and format. By associating sensor threshold features with video target detection features, the robustness of data collaborative analysis is improved, effectively addressing the complexities of industrial environments and unpredictable weather conditions.

[0076] Based on a lightweight, improved YOLOv8 model, this system enables real-time multi-target detection in video streams, covering equipment status (e.g., pipe leaks), personnel behavior (e.g., not wearing safety equipment, intrusion into dangerous areas), and environmental anomalies (e.g., flames, smoke, gas leaks). Through a multimodal cross-validation mechanism, video detection results are dynamically correlated with sensor data. This allows for real-time detection of equipment status, personnel behavior, and environmental anomalies (e.g., flames, smoke, mechanical failures) in the video stream, and cross-validation with sensor thresholds (e.g., pressure surges, gas leaks) to improve the accuracy of anomaly identification. For example, when a sensor detects a gas leak, the system automatically calls upon the associated video stream, uses target detection algorithms to locate the leak source and assess the risk range, and simultaneously generates an emergency response strategy based on personnel location information.

[0077] A multi-source data-triggered early warning process is designed, triggering primary alarms through multi-source data threshold linkage (such as gas concentration exceeding limits and video flame detection), avoiding misjudgments based on single data. In safety supervision, this process achieves closed-loop management of "anomaly perception - video verification - decision execution": the perception layer integrates video behavior recognition and sensor data; the verification layer automatically extracts historical video before the alarm for dynamic analysis; the execution layer triggers equipment control commands in stages, and the execution results are verified by both sensor status feedback and video behavior recognition. In industrial safety supervision, video behavior recognition (such as personnel crossing boundaries or not wearing safety equipment) and sensor data are integrated to trigger alarms in real time and link with equipment control modules, forming a closed-loop risk management mechanism.

[0078] like Figure 4 As shown, this embodiment proposes an industrial risk processing device based on multi-source data fusion, comprising: Alignment unit 101 is used to perform spatiotemporal alignment of environmental data and video stream data collected by multiple sensors in a critical industrial area at multiple consecutive time stamps to obtain multi-source aligned data. The first identification unit 102 is used to perform target detection and behavior anomaly identification on the video stream data in the multi-source aligned data to obtain behavior identification results; The second identification unit 103 is used to identify data anomalies in the environmental data in the multi-source aligned data according to the set environmental parameter threshold, and obtain the data identification result. Verification unit 104 is used to cross-validate the behavior recognition results and data recognition results to determine the risk level of the industrial critical area at the first time stamp. The determination unit 105 is used to determine whether to execute early warning and response actions based on the risk level of the industrial critical area at the first time stamp.

[0079] It should be noted that the processing procedures of the alignment unit 101, the first identification unit 102, the second identification unit 103, the verification unit 104, and the determination unit 105, as well as their beneficial effects, can be referred to respectively. Figure 1 Steps S101 to S105 are not described in detail here.

[0080] Optionally, the multiple sensors include multiple camera devices and environmental sensors located at different locations in critical industrial areas; Alignment unit 101 is also used for: Acquire the location information of each camera device and the video stream data collected by each camera device at each timestamp; and acquire the location information of each environmental sensor and the environmental data collected by each environmental sensor at each timestamp; Based on the location information of each camera device and each environmental sensor, determine the camera devices and environmental sensors that match the location information, and group the camera devices and environmental sensors that match the location information into a sensor group to obtain multiple sensor groups; For any sensor group, the location information of the camera device and the environmental sensor in the sensor group, the video stream data collected by the camera device at the timestamp, and the environmental data collected by the environmental sensor at the same timestamp are taken as the multi-source aligned sub-data corresponding to the timestamp. The multi-source aligned sub-data corresponding to each timestamp is treated as a whole as multi-source aligned data.

[0081] Optionally, the first identification unit 102 is also used for: Target detection and behavior anomaly recognition are performed on the video stream data collected by each camera device at each time stamp in the multi-source aligned data to obtain the behavior recognition result corresponding to each video stream data. Optionally, the second identification unit 103 is also used for: Based on the set environmental parameter thresholds, anomaly identification is performed on the environmental data collected by each environmental sensor at each time stamp in the multi-source aligned data, and the data identification result corresponding to each environmental data is obtained.

[0082] Optionally, environmental sensors are used to collect data on at least one of temperature, smoke, and gas concentration in critical industrial areas.

[0083] Optionally, verification unit 104 is also used for: For video stream data and environmental data collected by camera devices and environmental sensors in any sensor group at the first time stamp, if the behavior recognition result corresponding to the video stream data and the data recognition result corresponding to the environmental data are both abnormal, then the risk level of the industrial critical area at the first time stamp is determined to be high. If both the behavior recognition result corresponding to the video stream data and the data recognition result corresponding to the environmental data are normal, then the risk level of the key industrial area at the first time stamp is determined to be low. If the behavior recognition result corresponding to the video stream data and the data recognition result corresponding to the environmental data are both normal and abnormal, then the risk level of the key industrial area at the first time stamp is determined to be medium.

[0084] Optionally, the determining unit 105 is also used for: If the risk level of a critical industrial area is low at the first time stamp, then early warning and response actions are prohibited. If the risk level of a critical industrial area is medium or high at the first time stamp, then the warning method corresponding to the risk level is determined, and the warning and corresponding response actions are executed according to the set warning method.

[0085] Optionally, the response action includes at least one of pushing alarm information to the management platform, controlling the start and stop of field equipment, and activating an audible and visual alarm.

[0086] The industrial risk processing device based on multi-source data fusion proposed in this embodiment can spatiotemporally align environmental data and video stream data collected by multiple sensors in a critical industrial area at consecutive time stamps to obtain multi-source aligned data. Target detection and behavioral anomaly identification are performed on the video stream data in the multi-source aligned data to obtain behavioral identification results; and data anomaly identification is performed on the environmental data in the multi-source aligned data based on set environmental parameter thresholds to obtain data identification results. The behavioral identification results and data identification results are cross-validated to determine the risk level of the critical industrial area at a first time stamp. Based on the risk level of the critical industrial area at the first time stamp, it is determined whether to execute early warning and response actions. This embodiment can spatiotemporally align environmental data and video stream data collected by sensors to obtain multi-source aligned data, determine the risk level of the critical industrial area at different time stamps based on the multi-source aligned data, and determine whether to execute early warning and response actions, which can effectively improve the reliability of industrial risk processing.

[0087] In this embodiment, the industrial risk processing device based on multi-source data fusion is presented in the form of functional units. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0088] This invention also provides a computer device having the above-described features. Figure 4 The industrial risk processing device shown is based on multi-source data fusion.

[0089] Please see Figure 5 The present invention provides a schematic diagram of the structure of a computer device according to an optional embodiment. The computer device includes one or more processors 10, a memory 20, and interfaces for connecting the various components, including high-speed interfaces and low-speed interfaces. The various components are interconnected via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, multiple processors and / or multiple buses can be used with multiple memories, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 5 Take a processor 10 as an example.

[0090] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0091] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0092] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function. The data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0093] Memory 20 may include volatile memory, such as random access memory. Memory may also include non-volatile memory, such as flash memory, hard disk, or solid-state drive. Memory 20 may also include combinations of the above types of memory.

[0094] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.

[0095] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An industrial risk management method based on multi-source data fusion, characterized in that, include: Spatiotemporal alignment of environmental data and video stream data collected by multiple sensors in key industrial areas at multiple consecutive time stamps is performed to obtain multi-source aligned data. Target detection and behavior anomaly identification are performed on the video stream data in the multi-source aligned data to obtain behavior identification results; and, based on a set environmental parameter threshold, data anomaly identification is performed on the environmental data in the multi-source aligned data to obtain data identification results. Cross-validate the behavior recognition results and the data recognition results to determine the risk level of the key industrial area at the first time stamp; Based on the risk level of the key industrial areas at the first time stamp, determine whether to implement early warning and response actions.

2. The method according to claim 1, characterized in that, The plurality of sensors include multiple camera devices and environmental sensors installed at different locations in the critical industrial area; The process involves spatiotemporally aligning environmental data and video stream data collected by multiple sensors in key industrial areas at consecutive time stamps to obtain multi-source aligned data, including: Acquire the location information of each of the camera devices and the video stream data collected by each of the camera devices at each timestamp; and acquire the location information of each of the environmental sensors and the environmental data collected by each of the environmental sensors at each timestamp; Based on the location information of each camera device and each environmental sensor, determine the camera devices and environmental sensors whose location information matches, and group the camera devices and environmental sensors whose location information matches as a sensor group to obtain multiple sensor groups; For any of the sensor groups, the position information of the camera device and the environmental sensor in the sensor group, the video stream data collected by the camera device at the timestamp, and the environmental data collected by the environmental sensor at the same timestamp are taken as the multi-source aligned sub-data corresponding to the timestamp. The multi-source aligned sub-data corresponding to each timestamp is taken as the whole as the multi-source aligned data.

3. The method according to claim 2, characterized in that, The step of performing target detection and behavior anomaly recognition on the video stream data in the multi-source aligned data to obtain behavior recognition results includes: Target detection and behavior anomaly recognition are performed on the video stream data collected by each camera device at each timestamp in the multi-source aligned data to obtain the behavior recognition result corresponding to each video stream data. The step of identifying data anomalies in the multi-source aligned data based on a set environmental parameter threshold, and obtaining data identification results, includes: Based on the set environmental parameter thresholds, anomaly identification is performed on the environmental data collected by each environmental sensor at each timestamp in the multi-source aligned data to obtain the data identification result corresponding to each environmental data.

4. The method according to claim 2, characterized in that, The environmental sensor is used to collect data on at least one of temperature, smoke, and gas concentration in the key industrial area.

5. The method according to claim 3, characterized in that, The cross-validation of the behavior recognition results and the data recognition results to determine the risk level of the key industrial area at the first time stamp includes: For the video stream data and environmental data collected by the camera device and the environmental sensor in any of the sensor groups at the first time stamp, if the behavior recognition result corresponding to the video stream data and the data recognition result corresponding to the environmental data are both abnormal, then the risk level of the industrial critical area at the first time stamp is determined to be high. If both the behavior recognition result corresponding to the video stream data and the data recognition result corresponding to the environmental data are normal, then the risk level of the key industrial area at the first time stamp is determined to be low. If one of the behavior recognition results corresponding to the video stream data and the data recognition results corresponding to the environmental data is normal and the other is abnormal, then the risk level of the industrial critical area at the first time stamp is determined to be medium.

6. The method according to claim 5, characterized in that, The process of determining whether to implement early warning and response actions based on the risk level of the key industrial area at the first time stamp includes: If the risk level of the industrial critical area is low at the first time stamp, then early warning and response actions are prohibited. If the risk level of the industrial critical area is medium or high at the first time stamp, then a warning mode corresponding to the risk level is determined, and a warning and corresponding response action are executed according to the warning mode.

7. The method according to claim 4, characterized in that, The response actions include at least one of pushing alarm information to the management platform, controlling the start and stop of field equipment, and activating an audible and visual alarm.

8. An industrial risk processing device based on multi-source data fusion, characterized in that, include: The alignment unit is used to perform spatiotemporal alignment of environmental data and video stream data collected by multiple sensors in critical industrial areas at multiple consecutive time stamps to obtain multi-source aligned data. The first identification unit is used to perform target detection and behavior anomaly identification on the video stream data in the multi-source aligned data to obtain behavior identification results; The second identification unit is used to identify data anomalies in the environmental data in the multi-source aligned data according to the set environmental parameter threshold, and obtain the data identification result. A verification unit is used to cross-verify the behavior recognition result and the data recognition result to determine the risk level of the industrial critical area at the first time stamp. The determining unit is used to determine whether to execute early warning and response actions based on the risk level of the industrial critical area at the first time stamp.

9. A computer device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the industrial risk processing method based on multi-source data fusion as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the industrial risk processing method based on multi-source data fusion as described in any one of claims 1 to 7.