Industrial safety production monitoring method and system based on AI

By building an AI-based industrial safety production monitoring system, the problem of low monitoring accuracy of traditional monitoring methods in high-precision assembly and human-machine collaboration scenarios has been solved. It has achieved full-coverage, real-time safety risk assessment and intelligent decision-making, thereby reducing the accident rate.

CN121564604APending Publication Date: 2026-02-24SHENZHEN OKRA MUTUAL ENTERTAINMENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511624415.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Traditional industrial safety monitoring methods suffer from low monitoring accuracy, poor real-time performance, and limited coverage in scenarios such as high-precision assembly, heavy-duty machinery collaboration, and human-machine hybrid operations. They are unable to achieve multi-dimensional dynamic monitoring and intelligent risk prediction, leading to risks and hidden dangers such as delayed safety warnings and untimely accident response.

Method used

An AI-based industrial safety production monitoring method is adopted. By collecting and synchronously calibrating comprehensive monitoring video streams, a topological risk field of the production environment is constructed. Combined with personnel area matching detection and action safety logic verification, a location safety early warning signal is generated, and intelligent production monitoring decisions are made to achieve real-time identification and intelligent early warning of operator behavior and equipment status.

Benefits of technology

It enables multi-angle, full-coverage visual perception of the production site, improves the integrity and accuracy of data collection, dynamically reflects the status of safety risks, can monitor personnel location in real time and identify abnormal work behaviors, dynamically assess the level of production risks, realize proactive protection and intelligent decision-making, and reduce the incidence of safety accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564604A_ABST
    Figure CN121564604A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of production monitoring, in particular to an industrial safety production monitoring method and system based on AI, and the method comprises the following steps: collecting an omnibearing monitoring video stream of a production environment, carrying out the synchronous calibration, and constructing a synchronous monitoring video stream; performing risk labeling on the synchronous monitoring video stream, and constructing a production environment topology risk field; personnel area matching detection is carried out based on the production environment topology risk field, and a position safety early warning signal is generated; performing action safety logic verification based on the synchronous monitoring video stream, and marking an abnormal operation behavior event; and performing intelligent production monitoring decision based on the position safety early warning signal and the abnormal operation behavior event. According to the invention, the safety early warning efficiency and accuracy of precision industrial production are improved, the production safety is improved, and the generation risk is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of production monitoring technology, and in particular to an AI-based industrial safety production monitoring method and system. Background Technology

[0002] With the rapid development of industrial automation, intelligent manufacturing, and the Industrial Internet, modern industrial production is gradually moving towards a transformation towards high precision, high efficiency, and intelligence. In this process, safety monitoring of the production process has become a core element in ensuring personnel safety, stable equipment operation, and product quality control. Especially in high-risk work scenarios, precision assembly processes, and complex collaborative operations, achieving accurate perception and intelligent early warning of key elements such as worker behavior, equipment status, and the production environment has become a significant challenge for modern industrial safety management systems.

[0003] Traditional industrial safety monitoring methods primarily rely on limited physical sensors, fixed-point surveillance cameras, and manual inspections for anomaly identification and accident prevention. While these methods offer basic monitoring capabilities to some extent, they generally suffer from significant drawbacks such as low monitoring accuracy, poor real-time performance, limited coverage, and reliance on human judgment. They are ill-suited to meet the technological demands of modern precision manufacturing scenarios for "multi-dimensional dynamic monitoring" and "intelligent risk prediction." Particularly in scenarios involving high-precision assembly, heavy-duty machinery collaboration, and human-machine hybrid operations, traditional methods are severely inadequate in detecting fine-grained safety hazards such as micro-motion deviations, unauthorized path intrusions, and tool misoperations. This can lead to delayed safety warnings and untimely accident responses, posing significant risks. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention proposes an AI-based industrial safety production monitoring method and system, thereby resolving at least one of the aforementioned technical issues.

[0005] To achieve the above objectives, this invention provides an AI-based industrial safety production monitoring method, comprising the following steps: Step S1: Collect comprehensive monitoring video streams of the production environment, perform synchronous calibration, and construct a synchronous monitoring video stream; Step S2: Perform risk labeling on the synchronous monitoring video stream and construct a topological risk field for the production environment; Step S3: Perform personnel area matching and detection based on the production environment topological risk field to generate location safety early warning signals; Step S4: Perform motion safety logic verification based on synchronous monitoring video stream and mark abnormal operation behavior events; Step S5: Make intelligent production monitoring decisions based on location safety early warning signals and abnormal operation behavior events.

[0006] This specification provides an AI-based industrial safety production monitoring system for executing the AI-based industrial safety production monitoring method described above, including: The video calibration module is used to acquire comprehensive monitoring video streams of the production environment, perform synchronous calibration, and construct a synchronous monitoring video stream. The risk labeling module is used to label the synchronous monitoring video stream with risks and construct a topological risk field for the production environment. The personnel detection module is used to perform personnel area matching and detection based on the topological risk field of the production environment and generate location safety warning signals. The work behavior analysis module is used to perform action safety logic verification based on synchronously monitored video streams and to mark abnormal work behavior events. The monitoring and decision-making module is used to make intelligent production monitoring decisions based on location safety early warning signals and abnormal operation behavior events.

[0007] The beneficial effects of this invention are as follows: By collecting and synchronously calibrating comprehensive monitoring video streams of the production environment, multi-angle, full-coverage visual perception of the production site can be achieved, avoiding blind spots and improving the integrity of data collection. After time and space synchronous calibration, multiple video sources are fused on the same timeline, ensuring the consistency and comparability of image information, thereby providing high-precision, low-error visual input data for subsequent AI recognition and analysis. By risk labeling the synchronous monitoring video streams, a topological risk field of the production environment is constructed, realizing structured modeling and semantic understanding of the production scene. This topological risk field spatially maps and labels the risk levels of production elements (equipment, personnel passages, hazard sources, etc.), dynamically reflecting the safety risk status of each area. Based on the production environment topological risk field, personnel area matching detection can monitor the location distribution of personnel in the production scene in real time and determine whether they have entered or approached high-risk areas. Through human detection and target tracking algorithms, dynamic matching between people and the spatial risk field is achieved. When personnel are detected approaching a dangerous area, the system can automatically generate a location safety warning signal for early protection. By performing motion safety logic verification on synchronously monitored video streams, the system enables real-time identification and logical judgment of operator behavior. Utilizing AI visual analysis technology, the system identifies personnel movement characteristics, determines whether they conform to standard operating procedures, and automatically marks abnormal work behavior events, such as violations, boundary crossings, and failure to wear protective equipment. Based on location safety warning signals and abnormal work behavior events, the system makes intelligent production monitoring decisions, achieving multi-dimensional information fusion and intelligent decision control. The system comprehensively analyzes personnel location, safety status, and work behavior to dynamically assess production risk levels and can link with safety systems to execute automatic intervention measures, such as alarms, shutdowns, or personnel evacuation commands, thereby achieving proactive protection. Attached Figure Description

[0008] Figure 1 This is a flowchart illustrating the steps of an AI-based industrial safety production monitoring method according to the present invention. Figure 2 This is a detailed flowchart illustrating the implementation steps of step S1. Figure 3 This is a flowchart illustrating the detailed implementation steps of step S2. Detailed Implementation

[0009] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0010] This application provides an AI-based industrial safety production monitoring method and system. The executing entities of the AI-based industrial safety production monitoring method and system include, but are not limited to, the following: mechanical equipment, data processing platform, cloud server node, network upload device, etc., which can be considered as general computing nodes of this application. The data processing platform includes, but is not limited to, at least one of: an audio-visual management system, an information management system, and a cloud data management system.

[0011] Please see Figures 1 to 3 This invention provides an AI-based industrial safety production monitoring method, comprising the following steps: Step S1: Collect comprehensive monitoring video streams of the production environment, perform synchronous calibration, and construct a synchronous monitoring video stream; Step S2: Perform risk labeling on the synchronous monitoring video stream and construct a topological risk field for the production environment; Step S3: Perform personnel area matching and detection based on the production environment topological risk field to generate location safety early warning signals; Step S4: Perform motion safety logic verification based on synchronous monitoring video stream and mark abnormal operation behavior events; Step S5: Make intelligent production monitoring decisions based on location safety early warning signals and abnormal operation behavior events.

[0012] In the embodiments of the present invention, see Figure 1 The diagram below illustrates the steps of an AI-based industrial safety production monitoring method according to the present invention. In this example, the steps of the AI-based industrial safety production monitoring method include: Step S1: Collect comprehensive monitoring video streams of the production environment, perform synchronous calibration, and construct a synchronous monitoring video stream; In this embodiment, a distributed high-definition visual acquisition matrix is ​​deployed within a precision industrial production workshop. Industrial-grade network cameras are used to collect omnidirectional video streams from different perspectives in key work areas. The camera deployment density reaches at least three monitoring nodes per 50 square meters, ensuring a monitoring coverage rate of over 98% and eliminating blind spots. Each camera uses a 1920×1080 resolution and 30 frames per second acquisition standard, configured with a variable zoom lens to adapt to different monitoring distance requirements. Wide-angle lenses with a focal length of 2.8-12 mm are used in close-range work areas, while telephoto lenses with a focal length of 8-32 mm are used in long-range equipment monitoring areas. During video acquisition, the precise timestamp of each frame is recorded synchronously, with timestamp accuracy down to the millisecond level. A Network Time Protocol (NTP) server is used to perform clock synchronization calibration on all camera nodes, ensuring that the time deviation between nodes is less than 5 milliseconds. To address the inter-frame asynchrony problem caused by differences in network transmission latency and processing speed between different cameras, a timing alignment algorithm based on feature point matching is used for synchronization calibration. The specific method involves identifying feature points of the same moving target in multi-view video streams, calculating the time difference of these feature points appearing in different viewpoints, and using this as the time offset for compensation and adjustment of each video stream. Frame interpolation technology is used to resample the time axis, unifying all video streams to the same time reference, with time synchronization accuracy controlled within 10 milliseconds. The acquired raw video streams undergo preprocessing optimization, including lens distortion correction, illumination normalization, and motion blur compensation. Zhang Zhengyou's camera calibration method is used to obtain the intrinsic and extrinsic parameter matrices for each camera, and radial and tangential distortions are mathematically modeled and corrected to ensure that straight objects maintain their straight-line characteristics in the image. Finally, a synchronized monitoring video stream dataset containing spatial location information, time synchronization information, and image quality optimization is constructed, providing a reliable visual data foundation for subsequent scene analysis and target tracking.

[0013] Step S2: Perform risk labeling on the synchronous monitoring video stream and construct a topological risk field for the production environment; In this embodiment, deep learning-driven scene semantic segmentation processing is performed on the synchronous monitoring video stream. An improved DeepLabV3+ semantic segmentation network model is used to perform pixel-by-pixel classification and recognition of video frames, dividing the production environment into different functional areas. By training a semantic segmentation model containing 23 types of industrial scene labels, the recognition accuracy reaches 92.5%, accurately distinguishing different spatial areas such as dangerous boundary areas, restricted operation areas, high-frequency activity areas, equipment concentration areas, material storage areas, personnel passage areas, and emergency escape routes. Spatial boundaries of each identified functional area are extracted, morphological closing operations are used to eliminate boundary fragmentation noise, and clear regional contour lines are obtained through edge detection algorithms. The two-dimensional image coordinates are converted into spatial positions in the three-dimensional world coordinate system. Utilizing multi-view geometric constraints, stereo vision reconstruction technology is used to model the production workshop in three-dimensional space. Through steps such as feature point matching, epipolar correction, disparity calculation, and triangulation, a three-dimensional point cloud model containing depth information is generated, with a point cloud density of 5000 sampling points per square meter. Based on a 3D spatial model, equipment object identification and localization were performed. A Faster R-CNN object detection network was used to identify various production equipment such as lathes, milling machines, stamping machines, welding robots, and overhead cranes, achieving a detection accuracy of 96.8%. Simultaneously, equipment type attributes, spatial coordinates, dimensional parameters, and operational status information were extracted. Fixed facilities in the production environment, such as load-bearing columns, guardrails, electrical cabinets, and fire-fighting facilities, were labeled, constructing a scene knowledge graph containing 537 spatial entities. Risk labeling was performed based on a historical safety accident database. The spatial locations of 283 safety accidents that occurred in the past three years were traced back, and high-incidence areas were marked in the 3D scene model. A kernel density estimation method was used to generate an accident risk heatmap, with heatmap values ​​ranging from 0 to 1 representing risk intensity. Key spatial topological characteristic parameters were calculated, including a minimum safe distance of 1.2 meters between equipment, a minimum width of 1.5 meters for personnel passages, a minimum width of 2 meters for emergency escape routes, and a 3-meter radius isolation warning zone below high-altitude work areas. By analyzing 30 days of personnel movement trajectory data, 12 high-incidence areas of traffic congestion and 8 blind spots were identified, requiring additional safety protection measures. By integrating spatial geometric constraints, equipment hazard level classification, historical accident thermal distribution, and regional functional attributes, a comprehensive 3D visualization model of the "production environment topological risk field" was generated. This model divides the workshop space into 0.5m × 0.5m grid cells using a rasterized approach, assigning a risk score to each grid cell, ranging from 1 to 10 levels, providing a spatial reference benchmark for subsequent personnel location matching and risk early warning.

[0014] Step S3: Perform personnel area matching and detection based on the production environment topological risk field to generate location safety early warning signals; In this embodiment, real-time personnel detection and tracking are performed based on a production environment topological risk field model. The YOLO-v8 target detection algorithm is used to identify workers in the synchronously monitored video stream, achieving a detection speed of 45 frames per second, meeting the requirements of real-time monitoring. Image bounding boxes containing identity codes are generated for detected personnel targets, with pixel-level coordinate accuracy. The DeepSORT multi-target tracking algorithm continuously tracks each person, maintaining tracking continuity even in complex scenarios such as occlusion and intersections, achieving a tracking success rate of 94.3%. Personnel identity authentication is performed by combining facial recognition and employee badge recognition technologies. Facial recognition accuracy reaches 99.2%, and employee badge recognition accuracy reaches 97.5%, ensuring the accuracy of personnel identity through dual verification. Based on the identity information, the authorized work area information of the person is retrieved from the access management database. The authorized area is defined in the topological risk field using polygonal geometry, and each person can have 1-5 authorized work areas. Multi-view visual positioning technology is used to calculate the real-time 3D spatial coordinates of personnel. Triangulation is used to fuse observation data from at least three camera perspectives, achieving a positioning accuracy of ±0.15 meters horizontally and ±0.20 meters vertically. Spatial matching analysis is performed between the real-time personnel location coordinates and the topological risk field of the production environment. A point-within-a-polygon algorithm is used to determine whether a person is within their authorized work area, calculated 10 times per second to ensure timely detection of boundary violations. For detected boundary violations, the minimum Euclidean distance between the person and the area boundary is calculated. A distance less than 0.5 meters is considered a boundary approach warning, and a complete entry into an unauthorized area is considered a boundary violation alarm. For high-risk areas such as high-voltage power distribution areas, toxic chemical storage areas, and high-altitude work areas, stricter monitoring standards are adopted, with a three-level warning mechanism: a level one warning is triggered when a person is 5 meters from the boundary of a high-risk area, a level two warning is triggered when they are 3 meters away, and a level three emergency alarm is triggered when they enter a high-risk area. Based on historical data of personnel movement trajectories, motion intention prediction is performed. A Long Short-Term Memory (LSTM) model is used to predict the movement trajectory of personnel within the next 3 seconds, achieving a prediction accuracy of 87.6%. This model can identify the trend of personnel about to enter dangerous areas in advance. When the prediction model determines that there is a greater than 70% probability that a person will enter a high-risk area within 5 seconds, a location safety warning signal is generated in advance. The warning signal includes information such as personnel identity, current location coordinates, predicted target area, risk level, and suggested handling measures. The warning signal is pushed in real time via the Industrial Internet of Things (IIoT) to on-site audible and visual alarm devices, smart bracelets worn by personnel, and the safety management platform. The audible and visual alarm response time is less than 200 milliseconds, ensuring timely reminders to personnel and management personnel to take measures and effectively prevent personnel location-related safety accidents.

[0015] Step S4: Perform motion safety logic verification based on synchronous monitoring video stream and mark abnormal operation behavior events; In this embodiment, worker behavior recognition and motion analysis are performed based on synchronously monitored video streams. A dual-stream convolutional neural network architecture is used to simultaneously extract spatial and temporal motion features from the video. The spatial feature stream processes single-frame RGB images, capturing the worker's body posture and tool usage status. The temporal feature stream processes optical flow information between consecutive frames, capturing motion patterns and speed features. The two feature streams are fused in later stages to improve the accuracy of behavior recognition. On a test set containing 38 types of industrial operation actions, the recognition accuracy reaches 91.7%. The recognized actions are processed by temporal slicing, decomposing the continuous operation process into multiple action atomic units. Each action unit lasts between 0.5 and 5 seconds, forming a complete sequence of production operation behavior trajectories. A standard operating procedure (SOP) action library is established, defining standard operating procedures for 23 common processes such as lathe machining, welding, assembly, and equipment maintenance. Each procedure contains 5-15 ordered action steps, clearly specifying the necessary actions, prohibited actions, and safety protection requirements for each step. A Hidden Markov Model (HMM) is used to model the sequence of work behaviors. The standard operating procedure is transformed into a state transition probability matrix, with the normal action transition probability set between 0.75 and 0.95, and the abnormal transition probability below 0.1. The real-time identified work behavior sequences are input into the HMM model for sequence matching and anomaly detection. The similarity score between the actual behavior sequence and the standard procedure is calculated. When the similarity is below 0.7, it is judged as non-standard operation; when prohibited actions occur or critical safety steps are skipped, it is judged as abnormal work behavior. The specific detection includes the following eight abnormal behavior patterns: entering the work area without wearing a safety helmet (accuracy rate 96.5%); performing dangerous operations without wearing protective gloves (accuracy rate 94.2%); working at height without a safety belt (accuracy rate 93.8%); unauthorized use of mobile phones or smoking (accuracy rate 97.1%); unauthorized approach to dangerous parts while equipment is running (accuracy rate 89.6%); failure to perform equipment shutdown and power testing procedures (accuracy rate 88.3%); excessive material stacking in violation of safety regulations (accuracy rate 91.4%); and coordination errors during two-person collaborative work (accuracy rate 85.7%). Each detected abnormal work behavior event is recorded in detail, including the event timestamp accurate to milliseconds, spatial location coordinates, personnel identification information, abnormal behavior type classification, duration statistics, automatic extraction and saving of video evidence clips, and event severity rating. The abnormal behavior is classified into four levels based on its potential danger: minor violations, such as not wearing a work hat, are classified as Level 1; general violations, such as improper operating posture, are classified as Level 2; serious violations, such as not wearing protective equipment, are classified as Level 3; and extremely dangerous violations, such as unauthorized contact with high-voltage equipment, are classified as Level 4.A database of abnormal behavior events was established, and statistical analysis of abnormal events over the past six months revealed that violations related to not wearing protective equipment accounted for 38.7%, violations of operating procedures accounted for 27.4%, unauthorized stays in hazardous areas accounted for 18.9%, and unauthorized equipment operation accounted for 15.0%. Continuous monitoring and statistical analysis of abnormal behaviors provide objective data support for optimizing safety training content, revising operating procedures, and conducting individual safety assessments, thereby enabling a shift in safety management from a reactive accident response model to a proactive behavior intervention model.

[0016] Step S5: Make intelligent production monitoring decisions based on location safety early warning signals and abnormal operation behavior events.

[0017] In this embodiment, a multi-source safety risk information fusion processing platform is established to perform spatiotemporal alignment and correlation analysis on location safety early warning signals, abnormal work behavior events, equipment operation status monitoring data, and environmental parameter monitoring data. A timestamp matching algorithm is used to temporally correlate safety events from different sources, with a time window set to ±2 seconds, identifying composite risk scenarios with causal relationships in time and space. A risk correlation knowledge graph is constructed, containing 146 risk coupling patterns across three categories: personnel-equipment, personnel-environment, and equipment-environment. For example, the composite risk of personnel crossing boundaries into a hazardous area of ​​equipment and the equipment exhibiting operational abnormalities has a risk amplification factor of 2.3 times. A multi-dimensional risk assessment is performed on each detected safety event, with assessment dimensions including event occurrence probability, potential consequence severity, impact scope, response and handling difficulty, and frequency of similar historical incidents. Each dimension is quantitatively scored from 1 to 10 points, and the weights of each dimension are determined using the Analytic Hierarchy Process (AHP) to be 0.25, 0.30, 0.15, 0.15, and 0.15, respectively, to calculate the comprehensive risk score. Dynamic risk levels are assigned based on a comprehensive risk score: 0-3 points indicate low risk (green), 3-5 points indicate medium risk (yellow), 5-7 points indicate high risk (orange), and 7-10 points indicate extremely high risk (red). Differentiated intelligent monitoring and decision-making strategies are developed for each risk level: low-risk levels involve event recording and trend analysis without immediate intervention; medium-risk levels trigger on-site audio-visual alerts and push information to the team leader's mobile terminal, requiring on-site verification within 5 minutes; high-risk levels activate automatic voice warning broadcasts, freeze the equipment operation permissions of involved personnel, and push an emergency notification to the workshop safety supervisor, requiring on-site response within 2 minutes; extremely high-risk levels immediately trigger the equipment interlock protection system to execute an emergency shutdown, activate the area emergency lighting and evacuation guidance system, automatically dial the emergency command center, and push alarm information to the plant-level safety management level. An intelligent decision support system is established, trained using deep learning based on 4327 historical safety incident handling cases. Deep reinforcement learning algorithms are employed to optimize risk response strategies and learn the optimal handling solution selection in different scenarios. The system can automatically generate personalized risk management solutions based on constraints such as current production status, personnel distribution, equipment status, and emergency resource availability, with a solution generation time of less than 1.5 seconds. The solutions include personnel evacuation route planning, considering the shortest path, corridor congestion, and secondary risk avoidance, achieving an evacuation route planning accuracy of 92.1%; equipment shutdown sequence optimization, adhering to process dependencies and safety priority principles to avoid equipment damage caused by sudden shutdowns; and emergency resource scheduling plans, including emergency team personnel deployment, protective equipment allocation, and medical rescue preparation.A risk management effectiveness evaluation and feedback mechanism was established to quantitatively assess the accuracy of early warnings, timeliness of response, and effectiveness of handling each risk event. The evaluation results were fed back to the decision-making model for parameter optimization and adjustment, enabling the system to have autonomous learning and continuous evolution capabilities. Through six months of actual operation data verification, the intelligent production monitoring system reduced the accident rate by 67.3%, achieved an 88.5% success rate in intervening in abnormal behaviors, and shortened the emergency response time to 42% of its original length. This represents a fundamental shift from traditional passive monitoring to intelligent proactive protection, providing a comprehensive and intelligent technical support system for the safe production of precision industries.

[0018] In this embodiment, see Figure 2 The diagram below illustrates the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include: Deploy multi-angle high-definition vision cameras in industrial sites to collect comprehensive monitoring video streams of the production environment; The illumination intensity of the work area is calculated based on the all-round monitoring video stream to obtain the natural lighting conditions. Local contrast is enhanced based on natural lighting conditions to generate a brightness-optimized video stream; Calculate the inter-frame interval parameters of the brightness-optimized video stream; perform global delay analysis based on the inter-frame interval parameters to obtain global delay characteristics; Frame rate adaptive adjustment is performed based on global latency characteristics to obtain a frame rate latency optimized video stream; The frame rate delay optimized video stream is subjected to time stamp calculation for video acquisition from each angle, and synchronous calibration is performed to construct a synchronous monitoring video stream.

[0019] In this embodiment, when deploying multi-angle high-definition vision cameras in an industrial production site, the number and location of cameras need to be determined based on the spatial layout of the production area, equipment distribution, and the visibility requirements of key operational processes. Generally, within a production workshop of approximately 2000 square meters, 8 to 16 cameras are used, with a resolution of at least 8 megapixels (4K) and a frame rate maintained at 60 frames per second to ensure the clarity and continuity of video data. Each camera is connected to a centralized acquisition port via industrial Ethernet or fiber optic transmission, using H.265 high-efficiency video encoding for real-time transmission. There is a 10% to 15% overlap between the viewing angles of each camera to ensure the integrity of the field of view and the accuracy of subsequent video fusion. Lenses all possess wide dynamic range and automatic exposure capabilities to adapt to the workshop environment with drastic lighting changes. The acquired multi-angle video streams are uniformly time-stamped for subsequent time synchronization and multi-angle data calibration, thus forming the raw video foundation data available for AI visual analysis. Continuous frame images are extracted from the video stream, converted to grayscale, and the mean and variance of pixel brightness are calculated. Then, based on the camera's exposure time and gain parameters, the brightness value is physically corrected to convert it into an approximate illuminance value (lux). To improve accuracy, multiple representative sampling areas (such as workstation surfaces, passageways, and equipment surfaces) can be selected within the image to calculate the average brightness distribution, reflecting the overall lighting level. By smoothing the time-series data, the trend of natural light changes can also be reflected, such as fluctuations caused by daytime light decay or lighting changes. The final natural lighting condition data serves as an important reference for subsequent brightness adaptive processing, determining whether contrast enhancement and illumination compensation are needed for local areas, ensuring that the visual recognition model maintains consistent detection performance under different illumination conditions.

[0020] The system determines whether the illumination level is below a set threshold (e.g., 300 lux). If it is below this value, the video frame is divided into fixed-size regions (typically 64×64 pixels). The brightness distribution of each region is calculated, and local histogram equalization is performed to enhance details in dark areas. Next, a multi-scale Retinex model is used to separate the illumination and reflection components, and the reflection portion is smoothly enhanced to maintain a natural brightness transition between bright and dark areas. To prevent over-enhancement from introducing noise or color cast, a maximum gain limit (e.g., no more than 1.8 times) is set during contrast enhancement. The brightness-optimized video image exhibits clearer details in both bright and dark areas, with more prominent edge features, making it suitable for feature extraction and classification analysis in subsequent AI visual recognition tasks, ensuring high recognition accuracy even under complex lighting conditions. Temporal continuity analysis is performed on the brightness-optimized video stream to calculate the time interval parameter between frames, reflecting the latency characteristics during acquisition and transmission. The time signature of each frame is recorded through a precise time synchronization mechanism, and the time difference between adjacent frames is the inter-frame interval parameter. At an ideal acquisition rate (60 frames / second), the inter-frame interval is approximately 16.67 milliseconds. By statistically analyzing the time difference between consecutive frames, characteristic values ​​such as average interval, fluctuation amplitude, and time jitter can be obtained, reflecting the overall latency stability. If the fluctuation amplitude of the inter-frame interval is too large, it indicates that there is latency or frame dropping during video stream transmission. Further analysis of the latency distribution trend and fluctuation pattern can be performed using time series fitting. The obtained global latency characteristics include parameters such as average latency, fluctuation variance, and peak offset, serving as a quantitative basis for subsequent video frame rate adjustment. This analysis process ensures that multi-angle video maintains a stable temporal structure during the acquisition and transmission phases, providing a foundation for synchronous processing.

[0021] When the average latency and fluctuation are both within the set range (e.g., latency less than 20 milliseconds, jitter less than 1 millisecond), the original frame rate output is maintained. When a continuous increase in latency or jitter is detected, the output frame rate is gradually reduced (e.g., from 60 frames / second to 45 frames / second) to alleviate the data processing and transmission load. During frame rate adjustment, the video bitrate and image clarity are kept in harmony, and a dynamic bitrate adjustment strategy is used to avoid image quality degradation. To prevent instability caused by frequent switching, a hysteresis interval is set for frame rate changes, and the frame rate is only gradually increased when the latency recovers to below the threshold. The video stream after frame rate optimization can maintain stable output even under bandwidth fluctuations or high load conditions, thus ensuring the continuity and stability of the AI ​​visual analysis module. In the parallel acquisition of multi-angle video, time synchronization is a key factor in ensuring the accuracy of visual recognition. Each video frame is accompanied by a high-precision time stamp during acquisition, and global calibration is performed using a unified time reference signal. The time synchronization error Δt_sync is obtained by calculating the time deviation between corresponding frames from different cameras. When the deviation exceeds a predetermined range (usually no more than 3 milliseconds), time alignment can be achieved through frame interpolation or frame dropping, ensuring that video frames from different angles correspond to the same production state at a unified moment. After calibration, the video stream is reassembled in chronological order to form a multi-view monitoring screen with consistent time and corresponding angles. The synchronized multi-angle video enables cross-view object tracking, workstation behavior recognition, and 3D spatial positioning, ensuring that the information captured by each camera is completely corresponding in the temporal domain. The synchronized monitoring video stream obtained in this way provides a stable and time-consistent data foundation for the application of AI visual recognition in precision industrial safety monitoring.

[0022] In this embodiment, see Figure 3 The diagram below illustrates the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include: Multi-region semantic hierarchical recognition is performed on the synchronous monitoring video stream to extract regions of different operation types, including dangerous boundary areas, restricted operation areas, and high-frequency activity areas; Analyze the building structure of different work areas, visualize the results, and construct a multi-level work area map; Based on the analysis of equipment spacing, tool displacement trajectory and logistics path between workstations using multi-level work area map, the spatial topological characteristics of different areas are obtained. Based on the spatial topological characteristics, a temporal change analysis was performed, and potential spatial collision risks and traffic congestion were analyzed to extract collision risk areas and traffic congestion areas. Risks are marked on multi-level work area maps based on collision risk areas and traffic congestion areas to construct a production environment topology risk field.

[0023] In this embodiment, a semantic segmentation model is used to classify scenes in video frames at the pixel level. A commonly used network structure is a deep convolutional network based on multi-scale feature fusion (such as DeepLab or SegFormer). The input is a multi-angle fused video frame, and the output is a semantic label map of the same size as the original image. Different colors or indices in the label map represent different work types, such as equipment surfaces, passageways, personnel activity areas, and hazardous facility areas. By superimposing the recognition results from multiple angles and using Conditional Random Field (CRF) optimization, the region boundaries can be made more accurate and smooth. Based on the identified semantic region features, further classification is defined according to production rules: areas with high-temperature, rotating parts, or high-pressure equipment are classified as hazardous boundary areas; areas where only specific types of personnel are allowed to enter are marked as restricted operation areas; and areas where personnel frequently appear and work is intensive are defined as high-frequency activity areas. After semantic hierarchical recognition is completed, the video image is structured into a spatial image with semantic labels, laying the foundation for subsequent building structure analysis and spatial relationship analysis. The geometric boundaries and area ratios of each region are extracted, and the pixel coordinates in the two-dimensional image are mapped to actual three-dimensional spatial coordinates according to camera calibration parameters to achieve spatial reconstruction. By fusing matching points from multi-angle videos, the position matrix and height information of the region in three-dimensional space are obtained. For large equipment, partition walls, or supporting structures, their geometric shapes are extracted using edge detection and plane fitting algorithms, and relative spacing is calculated. Building structure analysis includes not only static geometric dimensions but also the identification of passable paths, obstacle distribution, and visually obstructed areas. Subsequently, the identification results are visualized: different types of spatial elements are overlaid in layers, with the base layer representing fixed building structures, the middle layer representing mobile equipment areas, and the top layer representing the distribution of work activities. This layered visualization method constructs a multi-level work area map, allowing the spatial organization of the production site to be presented intuitively in a graphical way, providing spatial semantic support for subsequent topology and risk analysis.

[0024] The 3D coordinate center points of the equipment in the field map are extracted, and the Euclidean distance and relative orientation angle between any two equipment are calculated to obtain the equipment spacing matrix. For the activity trajectories of personnel, tools, or vehicles, continuous position sequences are extracted using time-series tracking algorithms (such as target tracking based on optical flow or cross-frame tracking using depth detection networks) to form tool displacement trajectories and personnel movement paths. Based on this, the logistics routes and movement frequencies between workstations are analyzed to identify high-frequency passage paths and bottleneck areas. The spatial topology features are generated and expressed in the form of a graph model: nodes represent equipment or workstations, edges represent passage paths or interaction relationships, and edge weights reflect distance, frequency, or passage intensity. In this way, the spatial connectivity, operational coupling, and potential interference relationships between various areas can be obtained. Spatial topology features not only characterize static structural relationships but also reflect dynamic operating patterns, providing a quantitative reference for subsequent time-series analysis and risk area identification. The activity intensity of each node (workstation, equipment, personnel area) within a continuous time period is statistically analyzed to form a time-space distribution curve. If multiple movement trajectories intersect in the same area within a certain time period, and the trajectories have opposite velocity vector directions or a small angle, a potential collision risk is identified. By setting time thresholds (e.g., trajectory intersection events within 0.5 seconds) and distance thresholds (e.g., spatial overlap less than 1 meter), potential risk points can be accurately located. For traffic congestion analysis, the rate of change in the number of objects on the path per unit time is used as an indicator. When the density continuously exceeds a preset threshold (e.g., 2 people / square meter) and the average speed decreases significantly, it is marked as a congested area. Through continuous time-series analysis, a dynamic risk heat map distribution can be formed, reflecting the spatial risk change trend over different time periods. This time-series topology-based identification method enables safety hazards in the production site to be exposed in advance in a data-driven and traceable manner.

[0025] Based on the spatial coordinates and temporal weights of collision and congestion risks, multi-layered annotations are applied to the field map. Risk levels are graded according to probability and impact, for example, using levels 1 to 5 to identify risk areas of different severity. High-risk areas (such as frequent trajectory intersections) are displayed using color weighting, with red indicating high risk, orange indicating medium risk, and yellow indicating potential risk. Subsequently, spatial clustering analysis is performed on adjacent risk areas to identify risk-dense zones and risk propagation paths, thereby constructing a complete risk topology. This topological risk field not only displays the current risk distribution in the production environment but also predicts future risk evolution trends based on time-series data. The resulting topological risk field provides intuitive risk visualization results for industrial safety monitoring systems, providing a basis for safety scheduling, personnel diversion, equipment layout optimization, and intelligent early warning, enabling proactive prevention and continuous optimization of production safety.

[0026] In this embodiment, step S3 includes the following steps: Perform deep visual recognition and image segmentation on the synchronous monitoring video stream, and mark the image frames of the operators; Real-time spatial position calculation and temporal tracking are performed on the image frame of the operator to obtain the position coordinate sequence; Perform identity recognition on the image frames of the workers and mark the type of workers; Based on the type of workers and their location coordinate sequence, a personnel area matching analysis is performed on the topological risk field of the production environment to identify the authorized location matching degree of each worker. Based on the authorized location matching degree, dangerous area entry prediction is performed. When it is detected that a worker has entered a high-risk area, a restricted area, or an unauthorized area, a location safety warning signal is generated.

[0027] In this embodiment, a target detection model based on a convolutional neural network (CNN) structure is used to analyze video frames one by one. This model typically employs a multi-scale detection structure, such as the YOLO series or Mask R-CNN framework, to simultaneously achieve detection and segmentation functions. Semantic features at different scales are extracted through a feature pyramid network, and then combined with an anchor box matching mechanism to accurately locate personnel targets. The segmentation network further performs pixel-level segmentation within the detection box, distinguishing between people and background areas, ensuring the integrity of edge contours. Each detected worker is marked with a rectangular box in the video frame and assigned a unique number for subsequent tracking. To adapt to the complex environment of industrial sites (such as uneven lighting, occlusion, reflective metallic backgrounds, etc.), the detection algorithm improves stability through multi-angle input and brightness normalization enhancement. After depth recognition and segmentation processing, each worker in the video is accurately marked as a trackable target object, forming the basic input data for subsequent location and identity calculations. Combining camera calibration parameters and multi-view geometric relationships, a spatial projection model is used to convert image coordinates into spatial positions in the world coordinate system. The center point of each worker's image frame from different camera perspectives is triangulated to obtain spatial coordinates (X, Y, Z). To maintain continuity, the worker's position in each frame is recorded in a time series, forming a position coordinate sequence. The tracking phase employs a multi-target temporal matching algorithm based on target association, such as SORT or DeepSORT. Kalman filtering is used to predict the worker's possible position in the next frame, and matching correction is performed based on appearance features (such as color and posture contour). For cases of occlusion or temporary disappearance, temporal interpolation is used to maintain trajectory continuity. As workers move within the production area, the coordinate sequence records their displacement path and velocity information at fixed time intervals (e.g., 10 samples per second). This process enables precise positioning and dynamic trajectory tracking of workers within the industrial space, providing fundamental spatial data support for subsequent area matching and risk identification.

[0028] The system extracts physical features of personnel within the image frame, such as clothing color, reflective vest markings, helmet style, and type of tools worn. These features are compared with a pre-established occupation feature database, and the occupation category is determined by deep feature similarity calculation (usually using cosine distance or feature embedding vector distance). For example, a person wearing a specific colored safety helmet is identified as an electrician, a person carrying a tool bag and wearing dark blue overalls is identified as a maintenance worker, and a person wearing yellow reflective clothing is identified as an operator. If there is overlap in occupation identification, it is corrected by temporal consistency constraints and context-related features, that is, by combining the type of activity of the person in a specific work area to help confirm the occupation identity. After identification, each worker's image frame is labeled, including personnel number, occupation type, and identity confidence, providing semantic identification for subsequent authorized location matching analysis. The personnel coordinates are projected onto a three-dimensional spatial model of the topological risk field, and the matching degree is calculated based on the relationship between location and area boundaries. The authorized areas corresponding to different occupation types are predefined in the topological risk field, such as the power distribution area allowed for electricians, the assembly line area for assemblers, and the passage range for logistics personnel. By calculating the distance and positional relationship between personnel coordinates and the boundaries of each authorized area, it is determined whether personnel are within the compliant range. The matching degree can be defined as a weighted value of the proportion of time spent within the authorized area and the degree of spatial overlap. When the matching degree is higher than a set threshold (e.g., 0.85), it indicates normal personnel activity; below the threshold, it represents a risk of boundary crossing or unauthorized entry. Through continuous frame analysis, the trend of personnel approaching high-risk areas can be identified, enabling early detection of unauthorized movement. Personnel area matching analysis integrates personnel behavior data with environmental risk models to achieve dynamic monitoring of spatial authorization status.

[0029] By performing time-series analysis on the matching degree curve over a continuous time period, an entry risk is identified when the personnel trajectory gradually approaches the boundary of a high-risk area and the matching degree shows a continuous downward trend. The movement path within several future frames can be predicted by combining the velocity vector direction, and the minimum distance to the high-risk area boundary can be calculated using trajectory extrapolation. When this distance is lower than a preset threshold (e.g., 0.5 meters) and the personnel's current job does not have entry authorization, a location safety warning signal is immediately generated. The warning information includes the personnel number, current location, risk type, and risk level, used to prompt the monitoring platform or safety control unit to intervene. For restricted areas, if personnel enter the area continuously for more than a set time (e.g., 2 seconds), the warning signal is upgraded to a high-risk status. In this way, early warnings can be provided before personnel actually come into contact with dangerous equipment or perform boundary violations, enabling AI visual monitoring to not only have recognition capabilities but also behavior prediction and proactive safety intervention capabilities, providing an intelligent safety protection mechanism for precision industrial production.

[0030] In this embodiment, step S4 includes the following steps: Perform time-series analysis of the worker image frame to generate a sequence of worker production actions; The sequence of production actions of operators is divided into temporal segments and action units, and multiple action behavior units are extracted. Based on the preset standard operation action library, perform action safety logic verification on multiple action behavior units and mark abnormal operation action units; Based on the abnormal operation action unit, the time, location, personnel information and action type of the abnormality are calculated to generate abnormal operation behavior events.

[0031] In this embodiment, key pose points of each worker's image frame are detected and located. Key skeletal nodes, including 17 key points such as the head, shoulders, elbows, wrists, waist, knees, and ankles, are identified using deep neural network-based human pose estimation algorithms (such as OpenPose, HRNet, or ViTPose architectures). Subsequently, the trajectory of pose change over time is calculated based on the relative position, angle changes, and movement speed between these key points. Action feature vectors are generated by modeling the skeletal angle curves and movement directions in consecutive frames. A time window sliding mechanism is then used to combine consecutive frames within a short period into a time segment to describe a complete action unit process. Finally, all time segments are arranged sequentially to form a sequence of worker's production actions in the time dimension, covering the complete action cycle from preparation and execution to completion. This action sequence is the core input data for subsequent behavior decomposition and safety logic analysis, providing a temporal feature basis for action recognition and safety determination. The continuous action sequence is segmented based on the changes in the time curve of the action features. First, the velocity and acceleration curves of key points are calculated. Points showing significant changes (such as velocity increasing from 0 to a peak or abrupt angle changes) are considered action boundary points. Then, the Dynamic Time Warping (DTW) algorithm is used to analyze the similarity between action sequences, identifying the periodic structure of repetitive actions, such as basic units like grasping, carrying, installing, twisting, and inspecting. For complex action sequences, Hidden Markov Models (HMMs) or Temporal Convolutional Networks (TCNs) are used for sequence decomposition, automatically extracting multiple action units from long-term behavior. Each action unit contains feature parameters, such as duration, joint angle range, direction of movement, and body center of gravity displacement. These features constitute a semantic description of the action, used for subsequent matching with standard action templates. After temporal segmentation, the originally continuous work behavior is decomposed into multiple structured and quantifiable action units, facilitating standardized comparison and safety logic verification.

[0032] After extracting multiple action units, logical verification is required based on a pre-defined standard operating procedure (SOP) action library to determine whether the behavior complies with safety regulations. The SOP action library consists of standardized action samples for each job type. Each action template includes the key skeletal angle range, action duration, spatial displacement direction, and permissible operational range. For example, in assembly work, the standard bolt tightening action includes a standard curve with an arm angle between 30° and 90° and a duration not exceeding 3 seconds; while in welding operations, body tilt angle, hand stability, and working posture are strictly regulated. The action verification process uses a feature vector comparison algorithm to match the real-time extracted action units with the standard action templates and calculate a similarity score. If the similarity is below a set threshold (e.g., 0.75) or abnormal parameters are detected (e.g., excessive action range, reversed posture, abnormal operation frequency), it is determined to be an abnormal action unit. Verification can also be combined with job type and work area to form behavioral logical constraints; for example, non-maintenance workers performing equipment disassembly and assembly is considered a violation. Through this verification process, accurate identification and automatic labeling of violations, fatigue, and dangerous postures can be achieved. The start and end timestamps of the abnormal unit are read from the corresponding video frame sequence to determine the specific time of the anomaly. Combined with worker tracking data, the corresponding spatial coordinates are obtained, and the work area is determined through coordinate matching, such as the vicinity of high-risk equipment, logistics channels, or assembly stations. Subsequently, the worker's ID and job type information are associated to fully describe the anomaly. Based on action recognition tags, the type of abnormal behavior is determined, such as "out-of-bounds operation," "non-standard posture," "tool misuse," "work interruption," or "approaching dangerous action." These elements are integrated into an attribute set for the abnormal event, forming an event record. If the same person exhibits multiple abnormal actions within a short period, they are grouped into consecutive abnormal events using a temporal clustering algorithm to reflect fatigue or violation trends. The final generated abnormal work behavior event possesses traceability, spatial location, and behavior type identification, and can be directly used for safety monitoring records, risk assessment, and automatic warnings, providing a data-driven and interpretable behavioral evidence basis for AI-based safety supervision in industrial sites.

[0033] In this embodiment, the specific steps of step S5 are as follows: Real-time monitoring of equipment safety status parameters, analysis of life status trend evolution, and generation of equipment status trend characteristics; Spatiotemporal correlation analysis of equipment status trend characteristics is performed based on abnormal operation behavior events to extract the causal relationship between action errors and equipment status; Based on the aforementioned causal relationships, the chain reaction of risky actions is predicted, and the degree of risk is evolved to obtain a risk score. Adaptive motion risk warning analysis is performed based on risk scores to obtain motion risk warning signals; Intelligent production monitoring decisions are made based on motion risk warning signals and location safety warning signals, and intelligent risk handling strategies are constructed.

[0034] In this embodiment, the changing trends of the "life state" of production equipment can be obtained by continuously monitoring key operating parameters. The monitored parameters include various physical quantities such as temperature, vibration, pressure, current, voltage, rotational speed, and liquid level. Each equipment's acquisition port outputs data at a fixed sampling frequency (e.g., 10Hz or higher) with a time stamp, synchronized with video frame timestamps. Subsequently, a sliding window analysis is performed on the multidimensional parameter sequence to extract dynamic features such as short-term mean, variance, gradient, and abnormal offset values. Time series modeling methods (such as ARIMA or LSTM networks) are used to predict the equipment's state trend, determining whether the current operating state exhibits deterioration, fluctuations, or abnormal drift. For example, when the vibration frequency shift exceeds three standard deviations of the normal range or the temperature rise rate consistently exceeds a preset threshold, a potential deterioration trend is identified. The changing patterns of each monitored parameter are fused to form equipment state trend features, reflecting the evolution of the equipment from normal to abnormal, providing a continuous state description for subsequent correlation analysis with personnel behavior events. Equipment status trend data and abnormal behavior events are aligned along a unified timeline, and their relative positions in physical space are determined using a spatial mapping matrix. For example, if equipment experiences a temperature increase and a sudden increase in vibration within a specific time period, and simultaneously, personnel in the video are performing non-standard operating actions (such as violent impacts, incorrect assembly, or accidental touches to the control panel), a causal relationship can be determined based on temporal overlap and spatial proximity. Correlation analysis employs methods based on Granger causality tests or time-delay correlation analysis. By calculating the temporal correlation coefficient and response delay time between the behavioral event sequence and the equipment status sequence, the degree of influence of behavior on equipment status changes is assessed. If the behavioral signal is temporally leading and significantly correlated, it is considered a potential "action leading to equipment abnormality" relationship. This process achieves cross-domain correlation from video behavioral data to equipment physical status, enabling production safety monitoring to possess behavior-status linkage perception capabilities.

[0035] A behavior trigger-equipment response chain is established based on a causal model, with key actions as input variables and equipment state changes as output responses. A multivariate time-series prediction model (such as an LSTM-CausalNet structure) is used to predict equipment state trends over future time periods, analyzing potential chain reactions of anomalies caused by action errors, such as overheating, shutdown, sudden energy consumption increases, or structural fatigue. To reflect the degree of risk evolution, time-related factors and response amplitude are weighted to form a risk propagation function. The risk level is expressed as a comprehensive score, with the scoring model considering parameters such as the severity of the action type (e.g., impact, accidental contact, electrical contact), equipment importance level, environmental sensitivity, and potential loss range. The scoring range is typically set from 0 to 100 points, with scores exceeding 70 points defined as high-risk events, 50 to 70 points as medium-risk, and below 50 points as controllable risks. This process enables predictive extrapolation from a single abnormal action to a potential accident chain, laying a quantitative foundation for intelligent early warning. A risk grading threshold model is established, with different response strategies corresponding to different risk levels. When the risk score rises rapidly within a short period or exceeds the classification threshold, a risk warning signal is automatically triggered. To enhance stability, a time smoothing and hysteresis control mechanism is introduced to avoid false alarms caused by short-term fluctuations. The warning analysis process considers not only the current risk value but also the rate of change and duration of the score for comprehensive judgment. For example, a continuous increase of more than 20% for 5 seconds is considered a critical risk. The risk signal contains core information such as action type, associated equipment number, risk level, predicted duration, and spatial location. The signal can be presented visually on the monitoring interface, such as through color, flashing, or indicator lights to distinguish risk levels. The adaptive analysis mechanism enables the warning to automatically adjust the threshold and response intensity according to the on-site conditions, thereby achieving agile response and dynamic perception of sudden risks and ensuring the real-time and accuracy of safety control decisions.

[0036] When both action risk warning signals and location safety warning signals are generated simultaneously, the intelligent decision-making phase begins to construct a comprehensive risk management strategy. The decision-making logic centers on multi-signal fusion, unifying the modeling of multi-dimensional signals such as personnel behavior risk, spatial location risk, and equipment status risk. First, the two types of warning signals are synchronized temporally and spatially to determine if they belong to the same event trigger domain, such as the same worker performing abnormal operations in a high-risk area and associated equipment experiencing abnormal status. Subsequently, a decision-making model combining a rule engine and reinforcement learning algorithms classifies and responds to different risk scenarios. For low-level risks, visual cues or voice alerts are triggered; for medium-level risks, safety interlocks are triggered, such as restricting equipment movement or reducing operating power; for high-level risks, emergency shutdown or safety isolation is implemented. The decision-making strategy dynamically adjusts the scope of action based on risk scores, risk duration, and the number of people in the vicinity. In this way, risk management shifts from passive alarm to proactive prevention, achieving intelligent coordination and adaptive safety control among people, machines, and the environment, forming a closed-loop intelligent production monitoring and decision-making system.

[0037] In this embodiment, the specific steps for real-time detection of equipment safety status parameters, life state trend evolution analysis, and generation of equipment status trend characteristics are as follows: Real-time monitoring of device safety status parameters via IoT sensor networks; Based on the safety status parameters of the equipment, analyze the vibration spectrum signal, temperature distribution field and acoustic characteristic data; Fast Fourier Transform analysis was performed on the vibration spectrum signal to extract abnormal peak drift, harmonic distortion degree and fundamental frequency energy attenuation characteristics, thus obtaining vibration spectrum feature fingerprint; Identify local overheated areas and abnormal temperature gradients based on the temperature distribution field; The degree of equipment degradation is quantitatively assessed using acoustic characteristic data to obtain a quantitative acoustic degradation curve; Based on the acoustic degradation quantization curve, vibration spectrum characteristic fingerprint, local overheating areas, and temperature gradient anomalies, the life state trend evolution analysis is performed to generate equipment state trend characteristics.

[0038] In this embodiment, to achieve real-time awareness of equipment operating status during precision industrial safety monitoring, a high-density IoT sensor network is needed to continuously collect multi-dimensional parameters during equipment operation. The monitoring content covers key indicators such as mechanical vibration, temperature, acoustic signals, current, voltage, torque, and lubricating oil status. Sensor nodes are installed in a distributed layout at key parts of the equipment, such as bearing housings, drive shafts, housings, and heat dissipation areas, to achieve full structural coverage. Sensor sampling frequencies are typically set between 1kHz and 20kHz to meet the requirements for capturing high-frequency mechanical vibrations and acoustic signals. Data is transmitted back to the monitoring center in real time via wireless transmission modules (such as LoRa, Wi-Fi 6, or Industrial Ethernet). To ensure timing accuracy, each sensor node maintains data time consistency through a high-precision clock synchronization protocol (such as PTP precision time protocol). The collected data is recorded in time-series format, including basic information such as sensor intensity, timestamp, and equipment number, forming a multi-dimensional real-time equipment operation dataset. This provides fundamental support for subsequent spectrum analysis, temperature field modeling, and acoustic degradation assessment, enabling continuous monitoring of the equipment's safety status throughout the entire process. The vibration signal is decomposed and mapped in the frequency domain to extract the vibration spectrum signal. The signal sampling rate is typically maintained above 10kHz to preserve the high-frequency resonance information of the equipment. Subsequently, a temperature distribution field model is generated based on the data from the thermal sensor array. Through interpolation and spatial fitting methods, the temperature data from different monitoring points are mapped into continuous two-dimensional or three-dimensional temperature distribution maps for observing local thermal characteristic changes in the equipment. Simultaneously, short-time Fourier transform (STFT) and Mel-frequency cepstral coefficient (MFCC) analyses are performed on the acoustic data to extract acoustic characteristic parameters reflecting mechanical wear, impact, or loosening. Vibration, temperature, and acoustic data are synchronized in the time dimension to form a multi-source physical signal fusion matrix. Through this multi-dimensional analysis process, the dynamic signals during equipment operation can be separated from the original noise, identifying potential abnormal trends and providing high-precision input for subsequent feature extraction and trend modeling.

[0039] The collected vibration signals are analyzed using Fast Fourier Transform (FFT) to convert the time-domain signals into frequency-domain features, revealing the structural resonance characteristics and dynamic response patterns during equipment operation. In the spectrum obtained after FFT transformation, each frequency component corresponds to the inherent vibration mode of the equipment's mechanical components. By comparing with the normal operating spectrum template, abnormal peak drift phenomena can be detected, i.e., the deviation of the equipment's key resonant frequency from the standard frequency. A deviation exceeding ±5% is typically considered an indication of structural loosening or component wear. Simultaneously, the energy distribution of harmonic components is analyzed. Significant enhancement of second or third harmonic energy indicates nonlinear vibration or mechanical coupling problems. The fundamental frequency energy attenuation rate reflects the declining trend of the equipment's mechanical efficiency; its continuous decrease is usually related to bearing wear or alignment misalignment. These three characteristic parameters together constitute a vibration spectrum fingerprint, used to identify the unique spectral morphology of the equipment's operating state. By continuously tracking fingerprint changes, a dynamic archive of the equipment's structural health can be formed, providing fundamental characteristics for the evolution of its life state trends. A temperature distribution field model constructed based on infrared or thermosensitive array data enables spatial characterization of temperature changes on the equipment's surface and internal structure. By performing differential analysis and gradient calculation on continuous temperature distribution frames, regions where the local temperature rise rate exceeds twice the average value are identified and defined as "local overheating regions." These regions are typically located around friction points, contact points, or high-load components. Further calculation of temperature gradient anomalies—pixels with abrupt changes in temperature change rate—reflects abnormal heat conduction or uneven heat dissipation. Gaussian filtering and spatial interpolation algorithms are used to smooth the temperature field data, reducing noise interference and improving the accuracy of anomaly region location. Time-series comparison allows tracking of the persistence and spread trends of overheating regions, thereby assessing the energy transfer efficiency and structural health of the equipment. When local overheating is accompanied by abnormal vibration spectra, it usually indicates mechanical friction or energy imbalance problems, providing key input features for subsequent life cycle trend modeling.

[0040] Acoustic signals can directly reflect the energy loss and wear status of mechanical equipment during operation. Audio signals acquired by high-sensitivity acoustic sensors, after short-time energy analysis and spectral envelope extraction, can identify characteristic acoustic signatures generated by mechanical collisions, friction, and vibrations. First, Mel-frequency cepstral coefficients (MFCC) features are extracted to capture spectral morphology changes. Then, spectral centroid and bandwidth are used to analyze the concentration of acoustic energy and the extent of frequency distribution expansion. Combined with the sound pressure level decay curve over time, an acoustic degradation quantification model is established. The degradation quantification curve has time on the horizontal axis and degradation degree on the vertical axis, with a value range typically from 0 to 1; the closer the value is to 1, the more severe the degradation. When acoustic features show high-frequency enhancement, low-frequency energy decay, or frequent fluctuations, it indicates increased internal friction or bearing wear. This curve can be continuously updated, reflecting the gradual process of equipment acoustic performance from stable to abnormal, providing a dynamic assessment basis for equipment health monitoring. Multi-source features are normalized and time-series aligned to construct an equipment state vector sequence. Subsequently, a multi-feature weighted fusion model was employed to comprehensively analyze the mechanical structural changes reflected by the vibration spectrum fingerprint, the thermal conduction imbalance revealed by the temperature gradient anomaly, and the frictional loss represented by the acoustic degradation curve. The trend direction and rate of change of each feature were calculated using a sliding time window to construct a life state evolution curve. This curve can be divided into four stages: "stable period," "sub-healthy period," "deterioration period," and "abnormal period." Through trend fitting and inflection point identification, the turning points of the equipment's health status were determined. When vibration anomalies, overheating, and acoustic degradation characteristics rise simultaneously and the duration exceeds a set threshold (e.g., 120 seconds), the equipment is judged to have entered the early deterioration stage. The final generated equipment status trend features possess temporal continuity and cross-domain consistency, providing data support for risk prediction and maintenance decision-making in industrial AI monitoring.

[0041] In this embodiment, an AI-based industrial safety production monitoring system is provided to execute the AI-based industrial safety production monitoring method described above, including: The video calibration module is used to acquire comprehensive monitoring video streams of the production environment, perform synchronous calibration, and construct a synchronous monitoring video stream. The risk labeling module is used to label the synchronous monitoring video stream with risks and construct a topological risk field for the production environment. The personnel detection module is used to perform personnel area matching and detection based on the topological risk field of the production environment and generate location safety warning signals. The work behavior analysis module is used to perform action safety logic verification based on synchronously monitored video streams and to mark abnormal work behavior events. The monitoring and decision-making module is used to make intelligent production monitoring decisions based on location safety early warning signals and abnormal operation behavior events.

[0042] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0043] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein are implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. An AI-based method for monitoring industrial safety production, characterized in that, Includes the following steps: Step S1: Collect comprehensive monitoring video streams of the production environment, perform synchronous calibration, and construct a synchronous monitoring video stream; Step S2: Perform risk labeling on the synchronous monitoring video stream and construct a topological risk field for the production environment; Step S3: Perform personnel area matching and detection based on the production environment topological risk field to generate location safety early warning signals; Step S4: Perform motion safety logic verification based on synchronous monitoring video stream and mark abnormal operation behavior events; Step S5: Make intelligent production monitoring decisions based on location safety early warning signals and abnormal operation behavior events.

2. The AI-based industrial safety production monitoring method according to claim 1, characterized in that, The specific steps of step S1 are as follows: Deploy multi-angle high-definition vision cameras in industrial sites to collect comprehensive monitoring video streams of the production environment; The illumination intensity of the work area is calculated based on the all-round monitoring video stream to obtain the natural lighting conditions. Local contrast is enhanced based on natural lighting conditions to generate a brightness-optimized video stream; Calculate the inter-frame interval parameters of the brightness-optimized video stream; perform global delay analysis based on the inter-frame interval parameters to obtain global delay characteristics; Frame rate adaptive adjustment is performed based on global latency characteristics to obtain a frame rate latency optimized video stream; The frame rate delay optimized video stream is subjected to time stamp calculation for video acquisition from each angle, and synchronous calibration is performed to construct a synchronous monitoring video stream.

3. The AI-based industrial safety production monitoring method according to claim 1, characterized in that, The specific steps of step S2 are as follows: Multi-region semantic hierarchical recognition is performed on synchronous monitoring video streams to extract regions of different operation types, including dangerous boundary areas, restricted operation areas, and high-frequency activity areas; Analyze the building structure of different work areas, visualize the results, and construct a multi-level work area map; Based on the analysis of equipment spacing, tool displacement trajectory and logistics path between workstations using multi-level work area map, the spatial topological characteristics of different areas are obtained. Based on the spatial topological characteristics, a temporal change analysis was performed, and potential spatial collision risks and traffic congestion were analyzed to extract collision risk areas and traffic congestion areas. Risks are marked on multi-level work area maps based on collision risk areas and traffic congestion areas to construct a production environment topology risk field.

4. The AI-based industrial safety production monitoring method according to claim 1, characterized in that, Step S3 is as follows: Perform deep visual recognition and image segmentation on the synchronous monitoring video stream, and mark the image frames of the operators; Real-time spatial position calculation and temporal tracking are performed on the image frame of the operator to obtain the position coordinate sequence; Perform identity recognition on the image frames of the workers and mark the type of workers; Based on the type of workers and their location coordinate sequence, a personnel area matching analysis is performed on the topological risk field of the production environment to identify the authorized location matching degree of each worker. Based on the authorized location matching degree, dangerous area entry prediction is performed. When it is detected that a worker has entered a high-risk area, a restricted area, or an unauthorized area, a location safety warning signal is generated.

5. The AI-based industrial safety production monitoring method according to claim 1, characterized in that, The specific steps of step S4 are as follows: Perform time-series analysis of the worker image frame to generate a sequence of worker production actions; The sequence of production actions of operators is divided into temporal segments and action units, and multiple action behavior units are extracted. Based on the preset standard operation action library, perform action safety logic verification on multiple action behavior units and mark abnormal operation action units; Based on the abnormal operation action unit, the time, location, personnel information and action type of the abnormality are calculated to generate abnormal operation behavior events.

6. The AI-based industrial safety production monitoring method according to claim 1, characterized in that, The specific steps of step S5 are as follows: Real-time monitoring of equipment safety status parameters, analysis of life status trend evolution, and generation of equipment status trend characteristics; Spatiotemporal correlation analysis of equipment status trend characteristics is performed based on abnormal operation behavior events to extract the causal relationship between action errors and equipment status; Based on the aforementioned causal relationships, the chain reaction of risky actions is predicted, and the degree of risk is evolved to obtain a risk score. Adaptive motion risk warning analysis is performed based on risk scores to obtain motion risk warning signals; Intelligent production monitoring decisions are made based on motion risk warning signals and location safety warning signals, and intelligent risk handling strategies are constructed.

7. The AI-based industrial safety production monitoring method according to claim 6, characterized in that, The specific steps for analyzing the life state trend evolution of the real-time detection equipment safety status parameters and generating equipment status trend characteristics are as follows: Real-time monitoring of device safety status parameters via IoT sensor networks; Based on the safety status parameters of the equipment, analyze the vibration spectrum signal, temperature distribution field and acoustic characteristic data; Fast Fourier Transform analysis was performed on the vibration spectrum signal to extract abnormal peak drift, harmonic distortion degree and fundamental frequency energy attenuation characteristics, thus obtaining vibration spectrum feature fingerprint; Identify local overheated areas and abnormal temperature gradients based on the temperature distribution field; The degree of equipment degradation is quantitatively assessed using acoustic characteristic data to obtain a quantitative acoustic degradation curve; Based on the acoustic degradation quantization curve, vibration spectrum characteristic fingerprint, local overheating areas, and temperature gradient anomalies, the life state trend evolution analysis is performed to generate equipment state trend characteristics.

8. The AI-based industrial safety production monitoring method according to claim 6, characterized in that, The intelligent risk management strategy specifically includes personnel evacuation instructions, emergency equipment shutdown, temporary area closure, and emergency resource allocation.

9. An AI-based industrial safety production monitoring system, characterized in that, The method for performing the AI-based industrial safety production monitoring method as described in claim 1 includes: The video calibration module is used to acquire comprehensive monitoring video streams of the production environment, perform synchronous calibration, and construct a synchronous monitoring video stream. The risk labeling module is used to label the synchronous monitoring video stream with risks and construct a topological risk field for the production environment. The personnel detection module is used to perform personnel area matching and detection based on the topological risk field of the production environment and generate location safety warning signals. The work behavior analysis module is used to perform action safety logic verification based on synchronously monitored video streams and to mark abnormal work behavior events. The monitoring and decision-making module is used to make intelligent production monitoring decisions based on location safety early warning signals and abnormal operation behavior events.