Industrial production safety risk real-time early warning method and system based on AI visual detection
By adopting a cloud-edge-device collaborative architecture and multimodal AI visual fusion detection technology, the real-time and accuracy problems of industrial safety monitoring in existing technologies have been solved, enabling full-domain, real-time, and precise safety risk management and control, adapting to complex working conditions and reducing deployment costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-30
- Publication Date
- 2026-06-26
Smart Images

Figure CN122288403A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial safety management and control, and in particular to a method and system for real-time early warning of industrial production safety risks based on AI visual inspection. Background Technology
[0002] Industrial production safety is the core guarantee for the stable operation of industries such as manufacturing, chemical, metallurgy, and mining. Industrial sites face multiple safety risks, including personnel violations, equipment malfunctions and aging, and sudden environmental hazards. Inadequate management can easily lead to safety accidents, causing casualties and property damage. Traditional industrial production safety management mainly relies on manual inspections, fixed sensor monitoring, and manual review of video surveillance. With the upgrading of industrial intelligence, some enterprises have introduced conventional video surveillance systems and single-sensor monitoring equipment to assist in safety management, but overall, it is still in a passive management stage.
[0003] Most existing industrial safety monitoring technologies use general machine vision models without being customized for industrial conditions. Conventional methods only perform simple target detection, lack anti-interference filtering and multi-source verification, resulting in extremely high false alarm and false negative rates. Furthermore, they rely on centralized inference in the cloud, lack edge offline adaptation, and become ineffective when the network is disconnected. They also consume a lot of computing power and have high deployment costs, making it difficult to balance real-time performance and accuracy, and thus difficult to scale up.
[0004] Therefore, there is a need to provide a real-time early warning method and system for industrial production safety risks based on AI visual inspection to solve the above problems. Summary of the Invention
[0005] The purpose of this invention is to provide a real-time early warning method and system for industrial production safety risks based on AI visual inspection. Through a three-level collaborative architecture of cloud-edge-device and multimodal AI visual fusion inspection technology, it can achieve full-domain, real-time, accurate and closed-loop control of industrial safety risks.
[0006] To achieve the above objectives, this invention provides a real-time early warning method for industrial production safety risks based on AI visual inspection, comprising the following steps: S1: Collect multimodal visual data, deploy dual-light fusion visual acquisition equipment in industrial sites, simultaneously acquire visible light video streams and infrared thermal imaging video streams, and collect environmental parameters with sensors to complete real-time acquisition of all-dimensional data in industrial sites; S2: The acquired video stream and environmental parameters are transmitted to the edge computing unit in real time. Preprocessing is performed using an adaptive industrial image enhancement algorithm to remove noise interference from the industrial site. An improved YOLOv8 lightweight target detection model, which integrates GhostNet and CBAM attention mechanisms, is used in conjunction with HRNet human pose estimation and DeepLabV3+ semantic segmentation models to simultaneously detect personnel protection compliance, dangerous behaviors, restricted area crossings, equipment temperature changes and defects, environmental fires and leaks, and passage blockages. Continuous target tracking is achieved through the ByteTrack multi-target spatiotemporal tracking algorithm. Instantaneous interference is removed by combining a spatiotemporal context confidence fusion formula, and then a vision-sensor multi-source data decision fusion formula is used. Complete the verification of the authenticity of the risk; in the formula, Indicates the confidence weight of visual detection; For visual detection confidence; As an auxiliary sensor confidence weight; Confidence level of sensor data; The final risk confidence level is used to determine whether an alert will ultimately be triggered. S3: The three-level early warning system is divided according to the degree of danger of the hidden danger. All early warning information is simultaneously captured with on-site images and video clips for evidence preservation. S4: The edge computing unit uploads effective early warning data, hazard handling records and on-site abnormal samples to the cloud management platform in real time. The cloud management platform visualizes the risk distribution and equipment status across the entire domain through digital twins. Based on the accumulated on-site sample data, it fine-tunes and iterates the edge AI model to optimize detection accuracy and anti-interference capabilities. S5: In offline mode, the edge computing unit automatically switches to offline working mode to continuously complete local video acquisition, AI inference and hierarchical early warning operations, cache early warning data and on-site samples, and automatically synchronize them to the cloud management platform after the network is restored to ensure that the early warning function is not interrupted in the case of network outage.
[0007] Preferably, in S1, a dual-light pixel-level fusion formula is used. Real-time data alignment and fusion are achieved, among which, Raw image data captured by a visible light camera. Raw image data acquired by an infrared thermal imaging camera. The fusion coefficient; A fused image of two-pixel sets; It is equipped with temperature, humidity, gas concentration and vibration auxiliary sensors to collect environmental parameters, with a frame rate of no less than 30fps and a resolution of no less than 1080P.
[0008] Preferably, in S3, the three-level warning levels are specifically divided as follows: Level 1 warning is not wearing protective equipment, minor blockage of safety passages, and normal equipment temperature is too high; Level 2 warning is crossing the restricted area, minor equipment leakage, and personnel violation of regulations; Level 3 warning is open flame and thick smoke, large-scale leakage of hazardous chemicals, personnel falling and losing consciousness, or equipment overheating and overload failure.
[0009] The system of real-time early warning method for industrial production safety risks based on AI visual inspection includes a terminal multimodal perception layer, an edge intelligent computing layer, a cloud management and control platform layer, and a hierarchical early warning linkage layer. The hierarchical early warning linkage layer receives early warning instructions from the edge intelligent computing layer in one direction. The terminal multimodal sensing layer includes an explosion-proof wide dynamic range visible light camera, an infrared thermal imaging camera, a supplementary lighting module, and an auxiliary sensor module; The edge intelligent computing layer includes an edge AI computing box equipped with an NPU computing chip, a data preprocessing module, an AI inference module, and an offline caching module; The cloud-based management platform layer includes a digital twin visualization module, an early warning management module, a work order dispatch module, a model iteration module, and a data storage module; The tiered early warning and linkage layer includes a local audible and visual alarm module, a voice broadcast module, a mobile push module, and an industrial PLC linkage module.
[0010] Preferably, the edge AI computing box supports concurrent processing of 16-32 video streams, and can complete video preprocessing, AI model inference, risk assessment, local early warning control and offline data caching; the edge AI computing box adopts a heterogeneous architecture of ARM and NPU.
[0011] Preferably, the data storage module provides underlying data support for the digital twin visualization module, early warning management module, work order dispatch module, and model iteration module; One end of the early warning management module connects to the data storage module in real time through a data interface to pull the original early warning information and risk level judgment results uploaded from the edge terminal in real time; the other end is unidirectionally connected to the digital twin visualization module and the work order dispatch module respectively, pushing the processed standardized early warning data to the digital twin visualization module for situation display, and at the same time pushing the early warning work orders to be handled to the work order dispatch module to start the closed-loop handling process. The digital twin visualization module can unidirectionally access the full-domain field data, early warning data, and equipment data from the data storage module, while simultaneously receiving real-time early warning information pushed by the early warning management module. The model iteration module is bidirectionally connected to the data storage module, retrieving historical anomaly samples, false positive and false negative samples, and on-site real-world labeled data from the data storage module to complete the fine-tuning, optimization, and performance verification of the edge AI model.
[0012] Preferably, the local sound and light alarm module and the voice broadcast module are arranged in parallel and directly connected to the edge AI computing box, the mobile terminal push module is bidirectionally connected to the edge AI computing box and the early warning management module respectively, and the industrial PLC linkage module is directly connected to the edge AI computing box.
[0013] Preferably, the edge AI computing box has a built-in vision-sensor data cross-validation module that calls multi-source data decision fusion formulas to complete risk verification.
[0014] Therefore, the present invention employs the above-mentioned AI-based visual inspection-based real-time early warning method and system for industrial production safety risks, and the technical effects are as follows: (1) This invention uses dual-light fusion vision equipment combined with auxiliary sensors to cover three major categories of core safety risks: personnel behavior, equipment status, and environmental hazards. Visible light and infrared thermal imaging complement each other to adapt to various complex working conditions. It can be deployed in the whole area without blind spots to monitor and breaks through the limitations of traditional single-point monitoring of sensors and video surveillance without active detection, so as to realize full-dimensional and full-coverage safety perception in industrial sites.
[0015] (2) This invention improves the lightweight AI model, spatiotemporal context continuous frame filtering, multi-sensor cross-validation, and dual-model verification screening four noise reduction mechanisms, combined with dedicated decision fusion and filtering formulas, to specifically optimize the anti-interference capability of complex industrial working conditions and effectively eliminate various instantaneous interferences.
[0016] (3) This invention divides the warning levels into three levels and performs differentiated warnings and linkage operations for different risks to avoid the proliferation of alarm information. At the same time, it forms a closed-loop system of monitoring-analysis-early warning-disposal-iteration. The cloud management platform realizes the full-process management of work order dispatch, review and traceability, and model iteration, transforming safety management from passive post-event traceability to proactive pre-event prevention, and greatly reducing the cost of manual inspection.
[0017] (4) The edge computing unit of this invention has offline caching and local early warning functions. It can still complete the core early warning operation normally under network outage conditions. After the network is restored, the data is automatically synchronized, which is suitable for unstable network scenarios in industrial sites. The hardware adopts explosion-proof, waterproof and dustproof design. The model is optimized for special working conditions such as dust, low light, high and low temperature. It can be widely adapted to various industrial and engineering scenarios such as mechanical manufacturing, chemical industry, metallurgy, mining, and construction, and has strong versatility. Attached Figure Description
[0018] Figure 1 This is a flowchart of the real-time early warning method for industrial production safety risks based on AI visual inspection, as described in this invention. Detailed Implementation
[0019] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0020] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.
[0021] Example 1 like Figure 1 As shown, this invention provides a real-time early warning method for industrial production safety risks based on AI visual inspection, including the following steps: S1: Acquires multimodal visual data. A dual-light fusion visual acquisition device is deployed in the industrial site to simultaneously acquire visible light video streams and infrared thermal imaging video streams. Combined with sensors to collect environmental parameters, it completes real-time acquisition of all-dimensional data from the industrial site. S1 utilizes a dual-light pixel-level fusion formula... Real-time data alignment and fusion are achieved, among which, A fused image of two-pixel sets; Raw image data captured by a visible light camera. Raw image data acquired by an infrared thermal imaging camera. Fusion coefficient; Fusion coefficient under normal operating conditions Automatically switches to low visibility conditions It is equipped with temperature, humidity, gas concentration and vibration auxiliary sensors to collect environmental parameters, and complete the real-time acquisition of all-dimensional data in the industrial site, with a frame rate of no less than 30fps and a resolution of no less than 1080P.
[0022] S2: The acquired video stream and environmental parameters are transmitted to the edge computing unit in real time. Preprocessing is performed using an adaptive industrial image enhancement algorithm, and industrial noise interference is removed using a spatiotemporal context filtering formula. In the formula, For continuous frame windows, +5; For the first i Risk confidence level of the frame; The average risk confidence score of consecutive frames is used to filter out instantaneous interference from single frames. The filtering threshold is set to 0.6 to exclude single-frame interference and eliminate instantaneous interference such as flying insects and light shadows, thereby reducing the false alarm rate. An improved lightweight YOLOv8 target detection model integrating GhostNet and CBAM attention mechanisms is adopted. The standard convolutional layer of the original C2f backbone network is replaced by introducing a lightweight GhostNet module, and GhostConv is used to replace ordinary Conv to achieve feature map reuse. At the same time, the CBAM convolutional attention module is embedded to enhance the extraction of high-risk target features. The loss function and anchor box scale are optimized for complex working conditions such as industrial dust, water mist, and low light. The model's computing power compression ratio reaches more than 45%, the inference latency of a single frame image is less than 30ms, the target detection mAP@0.5 is not less than 95%, and the inference latency of a single frame image is less than 30ms. The core improved formulas and parameters are as follows: By replacing redundant computations in conventional convolutions with feature reuse, Ghost significantly reduces model parameters and computational cost while preserving core feature information, making it suitable for low-computing edge hardware and achieving model lightweighting. Ghost convolution feature generation formula: ,in For the input feature map, The original convolutional kernel is responsible for basic feature extraction; It is a linear low-rank convolution kernel used to generate redundant features; This is an identity mapping to enable feature reuse. s =4 represents the feature reuse factor; it is the optimal parameter measured in industrial scenarios. By focusing channel attention on the core feature channels of high-risk targets and suppressing interference from invalid background channels, and by locating the precise spatial position of the target through spatial attention, the background noise from industrial dust, light, and shadows is mitigated. This dual weighting improves the accuracy of target detection under complex conditions and enhances the detection capability for small and occluded targets. The CBAM attention weighting formula is as follows: ; ; In the formula, This is the channel attention weight matrix, used to filter target core feature channels and suppress invalid background channels; σ is the spatial attention weight matrix, used to locate the target's spatial position and mitigate environmental noise interference; X is the model input feature map; σ is the Hard-Swish activation function, which realizes non-linear feature activation; MLP is a multilayer perceptron, used to compress and reconstruct channel features; AvgPool is global average pooling, which extracts the global mean information of the feature map; MaxPool is global max pooling, which extracts the key extreme value information of the feature map. The kernel size is 7 7; Used to extract spatial features and generate spatial weights; The channel attention dimension is set to 16 to balance feature extraction accuracy and computational power consumption; Function: Focus on the features of potential hazards, suppress industrial background noise, and improve the detection accuracy in complex working conditions.
[0023] Optimize the loss function: ; In the formula, for The loss value represents the degree of difference between the predicted bounding box and the ground truth bounding box; This is the intersection-union ratio (IoU) between the predicted bounding box and the ground truth bounding box; the value range is set to 0-1. Center point of the prediction box Center point of the real frame Squared Euclidean distance; The minimum diagonal distance of the enclosing region that contains the two frames; This is the aspect ratio coefficient; This is a parameter for maintaining the aspect ratio of the frame; ; ; In the formula, This is the actual width of the annotation box. The actual height of the annotation box; The width of the model's predicted bounding box. The height of the predicted bounding box for the model; In this embodiment, the target confidence threshold is set to 0.65, the NMS threshold is set to 0.45, and the anchor box is adapted to small industrial targets (48×48, 96×96, 192×192). The target box regression accuracy is optimized to accelerate model convergence and reduce the false negative rate.
[0024] Combining HRNet human pose estimation and DeepLabV3+ semantic segmentation models, the system simultaneously detects personnel protective compliance, dangerous behaviors, restricted area incursions, equipment temperature changes and defects, environmental fires and leaks, and passageway blockages. Continuous target tracking is achieved through the ByteTrack multi-target spatiotemporal tracking algorithm, and instantaneous interference is eliminated using a spatiotemporal context confidence fusion formula. Finally, a vision-sensor multi-source data decision fusion formula is employed. Complete the verification of the authenticity of the risk; in the formula, The final risk confidence level is used to determine whether an alert will ultimately be triggered. Indicates the confidence weight of visual detection; For visual detection confidence; As an auxiliary sensor confidence weight; Confidence level of sensor data; =0.75; =0.25; the final risk threshold is set to 0.7 to achieve cross-validation of visual and sensor data and verify the authenticity of the risk.
[0025] S3: Based on the degree of danger of the hidden danger, three levels of early warning are divided. All early warning information is simultaneously captured with on-site images and video clips for evidence preservation. In S3, the three levels of early warning are specifically divided as follows: Level 1 warning is not wearing protective equipment, slight blockage of safety passages, and normal equipment temperature is higher than normal; Level 2 warning is crossing the restricted area, slight equipment leakage, and personnel violation of regulations; Level 3 warning is open flame and thick smoke, large-scale leakage of hazardous chemicals, personnel falling and losing consciousness, or equipment overheating and overload failure.
[0026] S4: The edge computing unit uploads effective early warning data, hazard handling records and on-site abnormal samples to the cloud management platform in real time. The cloud management platform visualizes the risk distribution and equipment status across the entire domain through digital twins. Based on the accumulated on-site sample data, it fine-tunes and iterates the edge AI model to optimize detection accuracy and anti-interference capabilities. S5: In offline mode, the edge computing unit automatically switches to offline working mode to continuously complete local video acquisition, AI inference and hierarchical early warning operations, cache early warning data and on-site samples, and automatically synchronize them to the cloud management platform after the network is restored to ensure that the early warning function is not interrupted in the case of network outage.
[0027] The system of real-time early warning method for industrial production safety risks based on AI visual inspection includes a terminal multimodal perception layer, an edge intelligent computing layer, a cloud management and control platform layer, and a hierarchical early warning linkage layer. The hierarchical early warning linkage layer receives early warning instructions from the edge intelligent computing layer in one direction. The terminal multimodal perception layer includes an explosion-proof wide dynamic range visible light camera, an infrared thermal imaging camera, a supplementary lighting module, and an auxiliary sensor module. The camera adopts an explosion-proof, waterproof, and dustproof design, adapting to harsh industrial conditions and providing full coverage without monitoring blind spots. The terminal multimodal perception layer establishes a point-to-point low-latency connection with the edge intelligent computing layer through wired industrial Ethernet / wireless 5G, synchronously transmitting the real-time collected visible light video stream, infrared thermal imaging video stream, and environmental parameter data to the edge AI computing box without intermediate forwarding links, ensuring the real-time performance and stability of data transmission.
[0028] The edge intelligent computing layer communicates bidirectionally with the cloud management platform layer via industrial Ethernet, sending down early warning data and anomaly samples, and receiving cloud model iteration parameters and remote management commands. It connects in real-time with the terminal multimodal perception layer to receive data collected from the front end. Simultaneously, it uses a hardwired or wireless dual-mode direct connection with the hierarchical early warning linkage layer to ensure zero-latency delivery of early warning commands. The edge intelligent computing layer includes an edge AI computing box equipped with an NPU computing chip, a data preprocessing module, an AI inference module, and an offline caching module. The core is responsible for video preprocessing, AI model inference, risk level determination, local early warning control, and offline data caching. The edge AI computing box supports concurrent processing of 16-32 video streams, enabling video preprocessing, AI model inference, risk determination, local early warning control, and offline data caching. It can independently achieve core early warning and linkage functions without relying on a cloud network. The edge AI computing box adopts an ARM and NPU heterogeneous architecture.
[0029] The cloud-based management platform layer comprises a digital twin visualization module, an early warning management module, a work order dispatching module, a model iteration module, and a data storage module. These five modules, based on the platform's internal MQTT protocol and RESTful universal interfaces, employ a bus-based interconnectivity architecture to achieve end-to-end data interaction and business collaboration. One end of the early warning management module connects in real-time with the data storage module via a data interface, retrieving raw early warning information and risk level assessment results uploaded from the edge. The other end connects unidirectionally with both the digital twin visualization module and the work order dispatching module, pushing processed standardized early warning data to the digital twin visualization module for situational display, while simultaneously pushing pending early warning work orders to the work order dispatching module to initiate a closed-loop processing procedure. The data storage module provides underlying data support for the digital twin visualization module, early warning management module, work order dispatch module, and model iteration module. The data storage module establishes a long connection with the edge intelligent computing layer through industrial Ethernet. It is responsible for receiving, encrypting, and persistently storing the original video clips, early warning records, risk assessment data, hidden danger handling results, abnormal training samples, and equipment operation logs uploaded from the edge. At the same time, it provides standardized data read, write, query, retrieval, and download permissions for the digital twin visualization module, early warning management module, work order dispatch module, and model iteration module. It is the core node for data flow across the entire platform.
[0030] The digital twin visualization module can unidirectionally access the full-domain field data, early warning data, and equipment data of the data storage module, while simultaneously receiving real-time early warning information pushed by the early warning management module. It can construct a 1:1 industrial field digital twin scene to realize risk heat map rendering, real-time early warning point marking, equipment operation status visualization, and work order processing progress display. Management personnel operation instructions are transmitted back to the data storage module through the module and simultaneously sent to the edge for execution.
[0031] The model iteration module is bidirectionally connected to the data storage module, retrieving historical anomaly samples, false positive and false negative samples, and on-site real-world labeled data from the data storage module to complete the fine-tuning, optimization, and performance verification of the edge AI model.
[0032] The work order dispatch module connects unidirectionally to the early warning management module, receiving standardized pending warnings and automatically generating work order numbers, matching responsible persons, setting processing time limits, and dispatching and pushing the work order. Downstream, it connects bidirectionally to the data storage module, transmitting work order dispatch records, processing progress feedback, rectification results, and acceptance confirmation information back to the data storage module for storage in real time, forming a full-process data chain of "early warning-dispatch-processing-closed loop". At the same time, it can also send overdue unprocessed work orders back to the early warning management module to activate a secondary reminder mechanism. The tiered early warning and linkage layer includes a local audible and visual alarm module, a voice broadcast module, a mobile push module, and an industrial PLC linkage module. The local audible and visual alarm module and the voice broadcast module are deployed in parallel and directly connected to the edge AI computing box. The local audible and visual alarm module only receives warning commands for all three levels (Level 1, Level 2, and Level 3). When a warning is triggered, it simultaneously activates audible and visual flashing alerts, installed at the corresponding monitoring points to immediately alert personnel in the surrounding area. The warning status is transmitted back to the edge terminal and archived in the cloud in real time. The voice broadcast module supports customized tiered voice content. Upon receiving a warning command of the corresponding level, it plays differentiated warning voices (Level 1 hazard reminder, Level 2 hazard warning, Level 3 hazard emergency evacuation), which are executed synchronously with the local audible and visual alarm module to enhance the on-site warning effect.
[0033] The mobile push module is bidirectionally connected to both the edge AI computing box and the early warning management module, while the industrial PLC linkage module is directly connected to the edge AI computing box. The edge AI computing box has a built-in vision-sensor data cross-validation module that uses multi-source data decision fusion formulas to verify the authenticity of risks. The cloud-based management platform layer is bidirectionally connected, receiving level 2 and 3 early warning commands and pushing early warning information, hazard locations, and on-site captured images to the mobile devices of corresponding managers and safety officers in real time. Simultaneously, it receives feedback commands from managers and transmits them back to the edge and cloud work order modules, forming an early warning-feedback closed loop.
[0034] The industrial PLC linkage module is directly connected to the edge AI computing box via hard wiring. As the highest priority emergency linkage module, it only receives level three major hazard warning commands and directly connects to the industrial site PLC control system and fire protection system. After being triggered, it immediately executes risk blocking operations such as equipment shutdown, power cut-off, fire pre-start, and valve closure. The linkage execution status is synchronously transmitted back to the edge terminal and the cloud to ensure rapid closed-loop blocking of major risks.
[0035] This embodiment has been running continuously for three months. The end-to-end latency of the system is between 300-450ms, the accuracy rate of hidden danger detection is 96.2%, the average daily false alarm of a single camera is 0.21, and there are no missed alarms. The frequency of workshop personnel's violations has decreased by 82%, the early detection rate of equipment abnormalities and hidden dangers is 100%, no safety accidents have occurred, the manpower for manual inspection has been reduced by 75%, the efficiency of safety management has been greatly improved, and it fully meets the actual control needs of industrial sites.
[0036] Therefore, this invention adopts the aforementioned AI-based visual inspection-based real-time early warning method and system for industrial production safety risks. It employs a decoupled collaborative architecture combining local real-time inference at the edge with iterative management from the cloud backend. This architecture pushes all core functions requiring high real-time performance—risk detection, early warning, and linkage—down to the edge computing unit. The cloud is only responsible for non-real-time data analysis, model iteration, work order closure, and global visualization. This invention completely avoids the pain points of high cloud transmission bandwidth consumption, high network latency, and failure due to network outages. It achieves independent offline early warning even under network outage conditions, compressing end-to-end response latency to within 500ms. Simultaneously, the edge unit adopts a physically isolated heterogeneous architecture of ARM+NPU, preventing AI inference from interfering with industrial production control timing and solving the industry challenge of ensuring compatibility between industrial control safety and intelligent detection.
[0037] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for real-time early warning of industrial production safety risks based on AI visual inspection, characterized in that, Includes the following steps: S1: Collect multimodal visual data, deploy dual-light fusion visual acquisition equipment in industrial sites, simultaneously acquire visible light video streams and infrared thermal imaging video streams, and collect environmental parameters with sensors to complete real-time acquisition of all-dimensional data in industrial sites; S2: The acquired video stream and environmental parameters are transmitted to the edge computing unit in real time, and preprocessed by an adaptive industrial image enhancement algorithm to remove noise interference from the industrial site. An improved lightweight YOLOv8 target detection model, integrating GhostNet and CBAM attention mechanisms, is employed. Combined with HRNet human pose estimation and DeepLabV3+ semantic segmentation, it simultaneously detects personnel protective compliance, dangerous behaviors, restricted area incursions, equipment temperature changes and defects, environmental fires and leaks, and passageway blockages. Continuous target tracking is achieved through the ByteTrack multi-target spatiotemporal tracking algorithm. Instantaneous interference is eliminated using a spatiotemporal context confidence fusion formula, and a vision-sensor multi-source data decision fusion formula is then used. Complete the verification of the authenticity of the risk; in the formula, Indicates the confidence weight of visual detection; For visual detection confidence; As an auxiliary sensor confidence weight; Confidence level of sensor data; For the final risk confidence level; S3: The three-level early warning system is divided according to the degree of danger of the hidden danger. All early warning information is simultaneously captured with on-site images and video clips for evidence preservation. S4: The edge computing unit uploads effective early warning data, hazard handling records and on-site abnormal samples to the cloud management platform in real time. The cloud management platform visualizes the risk distribution and equipment status across the entire domain through digital twins. Based on the accumulated on-site sample data, it fine-tunes and iterates the edge AI model to optimize detection accuracy and anti-interference capabilities. S5: In offline mode, the edge computing unit automatically switches to offline working mode to continuously complete local video acquisition, AI inference and hierarchical early warning operations, cache early warning data and on-site samples, and automatically synchronize them to the cloud management platform after the network is restored to ensure that the early warning function is not interrupted in the case of network outage.
2. The method for real-time early warning of industrial production safety risks based on AI visual inspection according to claim 1, characterized in that, S1 uses a dual-light pixel-level fusion formula Real-time data alignment and fusion are achieved, among which, Raw image data captured by a visible light camera. Raw image data acquired by an infrared thermal imaging camera. The fusion coefficient; A fused image of two-pixel sets; It is equipped with temperature, humidity, gas concentration and vibration auxiliary sensors to collect environmental parameters, with a frame rate of no less than 30fps and a resolution of no less than 1080P.
3. The method for real-time early warning of industrial production safety risks based on AI visual inspection according to claim 1, characterized in that, In S3, the three warning levels are specifically divided as follows: Level 1 warning is for not wearing protective equipment, minor blockage of safety passages, and normal equipment temperature being too high; Level 2 warning is for crossing the restricted area, minor equipment leakage, and personnel operating in violation of regulations; Level 3 warning is for open flames and thick smoke, large-scale leakage of hazardous chemicals, personnel falling and losing consciousness, or equipment overheating and overload failure.
4. The system of the real-time early warning method for industrial production safety risks based on AI visual inspection according to any one of claims 1-3, characterized in that: It includes a terminal multimodal perception layer, an edge intelligent computing layer, a cloud management and control platform layer, and a hierarchical early warning and linkage layer. The hierarchical early warning and linkage layer receives early warning instructions from the edge intelligent computing layer in one direction. The terminal multimodal sensing layer includes an explosion-proof wide dynamic range visible light camera, an infrared thermal imaging camera, a supplementary lighting module, and an auxiliary sensor module; The edge intelligent computing layer includes an edge AI computing box equipped with an NPU computing chip, a data preprocessing module, an AI inference module, and an offline caching module; The cloud-based management platform layer includes a digital twin visualization module, an early warning management module, a work order dispatch module, a model iteration module, and a data storage module; The tiered early warning and linkage layer includes a local audible and visual alarm module, a voice broadcast module, a mobile push module, and an industrial PLC linkage module.
5. The system of the real-time early warning method for industrial production safety risks based on AI visual inspection as described in claim 4, characterized in that, The edge AI computing box supports concurrent processing of 16-32 video streams and can complete video preprocessing, AI model inference, risk assessment, local early warning control, and offline data caching. The edge AI computing box adopts a heterogeneous architecture of ARM and NPU.
6. The system for real-time early warning of industrial production safety risks based on AI visual inspection as described in claim 4, characterized in that, The data storage module provides underlying data support for the digital twin visualization module, early warning management module, work order dispatch module, and model iteration module; One end of the early warning management module connects with the data storage module in real time through a data interface to retrieve the original early warning information and risk level judgment results uploaded from the edge terminal in real time; The other end is unidirectionally connected to the digital twin visualization module and the work order dispatch module, respectively. The processed standardized early warning data is pushed to the digital twin visualization module for situation display, and at the same time, the early warning work orders to be handled are pushed to the work order dispatch module to start the closed-loop handling process. The digital twin visualization module can unidirectionally access the full-domain field data, early warning data, and equipment data from the data storage module, while simultaneously receiving real-time early warning information pushed by the early warning management module. The model iteration module is bidirectionally connected to the data storage module, retrieving historical anomaly samples, false positive and false negative samples, and on-site real-world labeled data from the data storage module to complete the fine-tuning, optimization, and performance verification of the edge AI model.
7. The system of the real-time early warning method for industrial production safety risks based on AI visual inspection as described in claim 6, characterized in that, The local sound and light alarm module and the voice broadcast module are deployed in parallel and directly connected to the edge AI computing box. The mobile terminal push module is bidirectionally connected to the edge AI computing box and the early warning management module. The industrial PLC linkage module is directly connected to the edge AI computing box.
8. The system of the real-time early warning method for industrial production safety risks based on AI visual inspection as described in claim 4, characterized in that, The edge AI computing box has a built-in vision-sensor data cross-validation module that calls multi-source data decision fusion formulas to complete risk authenticity verification.