Factory safety monitoring and emergency linkage system based on AI vision
The factory safety monitoring system, which integrates multimodal sensor fusion and AI, solves the feature extraction problem of traditional systems in low-light and obstructed environments, enabling all-weather monitoring and personalized emergency response, and improving the level of intelligence in factory safety management.
Patent Information
- Application Number
- CN202511170492.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-28
AI Technical Summary
Traditional factory safety monitoring systems fail to extract features when ambient light is poor or there are obstructions, making it difficult to assess individual interactions and their potential risks, and emergency response procedures cannot dynamically adapt to sudden situations.
Employing multimodal sensor fusion technology, data is acquired using visible light, infrared thermal imaging, and millimeter-wave radar. The feature completion module reconstructs human posture in low-light environments, while the behavior analysis module constructs a human interaction map. The risk assessment module identifies abnormal behaviors, and the deductive planning module dynamically predicts accident evolution. The linkage execution module controls the response of emergency equipment.
It improves the robustness of human behavior recognition and the intelligence of emergency response in complex environments, accurately identifies high-risk behaviors, dynamically plans escape routes, and enables collaborative response from multiple terminal devices.
Smart Images

Figure CN121033764A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of factory safety monitoring technology, and in particular to a factory safety monitoring and emergency response system based on AI vision. Background Technology
[0002] Factory safety monitoring technology refers to a series of technologies and measures used to ensure the safety of personnel and assets within a factory or industrial area. It typically includes, but is not limited to, video surveillance, intrusion detection, fire alarms, and gas leak monitoring. With the development of AI technology, modern factory safety monitoring systems are becoming increasingly intelligent. For example, they use computer vision technology to automatically identify abnormal behavior and utilize IoT devices to monitor changes in environmental parameters in real time. These technologies work together to prevent accidents, promptly detect safety hazards, and quickly take corresponding measures to reduce losses. Therefore, how to utilize advanced technologies to improve the intelligence and security of factory safety monitoring has become one of the most pressing issues to be addressed.
[0003] In the field of file encryption, traditional monitoring systems fail when ambient light is poor or there are obstructions. The feature extraction method based on visible light will fail. Furthermore, traditional personnel monitoring systems can only provide location information and it is difficult to assess the interaction between individuals and their potential risks. At the same time, conventional emergency response procedures are pre-set and cannot dynamically adapt to changes in sudden situations. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides an AI vision-based factory safety monitoring and emergency response system to solve the problems that traditional monitoring systems fail when the ambient light is poor or there are obstructions. Traditional personnel monitoring systems often only provide location information and are difficult to assess the interaction between individuals and their potential risks. At the same time, conventional emergency response procedures are pre-set and cannot dynamically adapt to changes in sudden situations.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] In a first aspect, the present invention provides a factory safety monitoring and emergency response system based on AI vision, comprising:
[0008] The module includes a perception and acquisition module, a feature completion module, a behavior analysis module, a risk assessment module, a deduction and planning module, a linkage execution module, and a model update module.
[0009] The sensing and acquisition module is used to deploy multimodal sensing units in the factory area to acquire visible light images, infrared thermal imaging data and millimeter-wave radar point clouds, and to perform time stamp alignment and spatial coordinate mapping on the three types of data to output the original sensing dataset.
[0010] The feature completion module is used to judge the environmental quality based on the brightness and contrast indicators of the visible light image. When the indicators are lower than the preset threshold, it extracts human point cloud clusters based on millimeter-wave radar point cloud, generates three-dimensional pose key points through clustering and skeleton regression algorithms, and performs fusion verification by combining the human thermal contour in the infrared image to output human pose feature map.
[0011] The behavior analysis module is used to perform target detection on the human posture feature map, extract the position and posture vector of the person, construct the person interaction graph with the person as the node and the angle between the spatial distance and the line of sight as the edge weight, input the graph attention network to calculate the attention score of each node, and output the interaction risk vector representing the degree of abnormality in group cooperation.
[0012] The risk assessment module is used to match the interactive risk vector with the preset group violation pattern. When the correlation between the risk vector and the combination of unlicensed operation, lack of supervision or unauthorized operation exceeds the threshold, a group safety risk event is generated.
[0013] The simulation and planning module is used to call the digital twin model of the factory area after a mass safety risk event is generated, map the location information of the affected personnel to the BIM space, start the smoke diffusion simulation in combination with the ventilation system layout and combustible material distribution, predict the range of the dangerous area in the future time step, generate a dynamic obstacle avoidance grid based on the dangerous area and personnel location, and use a path search algorithm to calculate the optimal escape path in the obstacle avoidance grid.
[0014] The linkage execution module is used to control the emergency lighting system in the factory area to illuminate the guide route according to the optimal escape route, send path navigation instructions to the AR terminals worn by the affected personnel, send rescue task assignment information to the mobile terminal of the nearest inspection personnel, and start the broadcast system to play directional voice prompts.
[0015] The model update module is used to retain the original data locally in each factory area when the same system architecture is deployed in multiple factory areas, and periodically upload the model parameter gradients to the central server. The central server performs weighted aggregation of the gradients to generate a globally updated model.
[0016] As a preferred embodiment of the AI vision-based factory safety monitoring and emergency response system of the present invention, wherein: in the sensing and acquisition module, the multimodal sensing unit is installed on a fixed bracket above the key operating areas and passageways of the factory area, and the specific steps are as follows:
[0017] An integrated sensing unit is deployed at each monitoring point. This sensing unit includes a visible light camera, an infrared thermal imager, and a millimeter-wave radar. The three types of sensors share the same mounting base.
[0018] The sensing and acquisition module also includes a time synchronization unit, used to attach a unified time reference to the data acquired by each type of sensor, specifically including:
[0019] The time synchronization unit receives external time signals and generates high-precision time pulses;
[0020] When each sensor acquires a data frame, it appends the current time pulse value as a timestamp to the header of the data packet, so that the visible light image, infrared image and radar point cloud have the same time reference.
[0021] After time alignment is completed, spatial coordinate mapping is performed, a calibration target is placed at a preset distance in front of the sensing unit, and data is collected by the three types of sensors at the same time.
[0022] Extracting the feature point pixel coordinates u of the calibration target from a visible light image. i ,v i , where i represents the i-th feature point;
[0023] Extracting the three-dimensional spatial coordinates x of corresponding feature points from millimeter-wave radar point clouds. j ,y j ,z j , where j corresponds one-to-one with i;
[0024] The projection relationship is established by the expression:
[0025]
[0026] Where s is the scaling factor, K is the intrinsic parameter matrix of the visible light camera, which includes the focal length and principal point parameters, R is the rotation matrix, and T is the translation vector;
[0027] Spatial calibration is completed by minimizing the reprojection error of all feature points and solving for the optimal R and T.
[0028] The radar point cloud acquired at any time can be projected onto the visible light image plane through this transformation matrix to achieve cross-modal space alignment;
[0029] The output raw perception data set contains time-synchronized and spatially aligned trimodal data, with each set having the same timestamp, and is used for subsequent feature fusion and behavior analysis.
[0030] As a preferred embodiment of the AI vision-based factory safety monitoring and emergency response system of the present invention, wherein: the feature completion module receives the original sensing data set, and the specific steps are as follows:
[0031] Perform a quality assessment on the current frame of the visible light image to determine whether it meets the conditions for normal visual analysis.
[0032] Convert the image to grayscale to obtain a pixel matrix;
[0033] The overall average brightness L of the image is calculated, which is defined as the arithmetic mean of the gray values of all pixels, and the expression is:
[0034]
[0035] Where M is the image height and N is the image width;
[0036] The image is divided into several regular sub-blocks. The gray-level range within each sub-block is calculated as the local contrast, and the average contrast of all sub-blocks is calculated as the global contrast.
[0037] The L and global contrast are compared with preset quality judgment conditions. If L is lower than the first threshold and the global contrast is lower than the second threshold, the current environment is judged to be a low-quality visual environment.
[0038] Initiate a cross-modal feature completion process by acquiring a point cloud dataset from a millimeter-wave radar, where each point has three-dimensional coordinates;
[0039] Density clustering algorithm is used to process the point cloud dataset. The neighborhood search radius and minimum number of points are set, and points that meet the density connectivity are divided into the same cluster, resulting in multiple personnel point cloud clusters.
[0040] Apply a pre-trained 3D pose estimation model to each point cloud cluster to output the 3D coordinates of key human joints;
[0041] Meanwhile, the thermal radiation area of the human body is extracted from the infrared image, the high temperature area is obtained by temperature threshold segmentation, and then the complete thermal outline of the human body is obtained by edge detection and morphological closing operation.
[0042] The radar-generated skeleton structure is projected onto the infrared image plane, and its spatial overlap with the thermal profile is calculated. It is defined as the ratio of the length of the overlap between the skeleton line segment and the thermal profile to the total length of the skeleton.
[0043] If the spatial overlap is less than the preset matching threshold, the clustering parameters are adjusted and point cloud clusters are extracted again until the overlap meets the requirements, and a completed human pose feature map is generated.
[0044] As a preferred embodiment of the AI vision-based factory safety monitoring and emergency response system of the present invention, wherein: the behavior analysis module receives human posture feature maps, and the specific steps are as follows:
[0045] The geometric center of each person's joint coordinates is calculated to obtain their center position in the monitoring screen.
[0046] Extract the pose vector, which includes the head orientation angle, torso tilt angle, and limb extension state;
[0047] Construct a personnel interaction diagram with each person as a node;
[0048] For any two nodes j and i, calculate their Euclidean distance d. ij The expression is:
[0049]
[0050] If d ij If the distance is less than the preset threshold, then a connection edge is established between node j and i;
[0051] The weight of an edge is determined by both spatial proximity and line-of-sight consistency;
[0052] The constructed personnel interaction graph is input into a graph attention network, which contains multiple layers of attention mechanisms. Each layer calculates the attention coefficient of each node i to its neighbor node j, expressed as:
[0053]
[0054] Among them, h i Let be the feature vector of node i, W be the trainable weight matrix, a be the attention parameter vector, || denote vector concatenation, and N be the feature vector of node i. i Let i be the set of neighbors of node i;
[0055] Using α ij Weighted aggregation of neighbor features is used to update the node representation;
[0056] After multiple layers of propagation, the attention score of each node is output. All attention scores constitute the interaction risk vector S, which is used to characterize the salience of an individual's behavior in the group.
[0057] As a preferred embodiment of the AI vision-based factory safety monitoring and emergency response system of the present invention, the risk assessment module receives an interactive risk vector, and the specific steps are as follows:
[0058] Retrieve preset group violation pattern templates from the local rule base, including templates for unlicensed operation, lack of supervision, and unauthorized combination templates;
[0059] The relevance between the interaction risk vector and each template is calculated using the following expression:
[0060]
[0061] Where · represents the vector dot product;
[0062] Each R m Matching threshold R th Comparison, if R m >R th If so, then the current scenario is determined to match the violation pattern of the m-th group;
[0063] Generate a group safety risk event. The event record includes the event type identifier m, the occurrence time t, the set of personnel numbers involved and their center location coordinates;
[0064] The event is transmitted as a trigger signal to the simulation planning module to initiate subsequent emergency simulation procedures.
[0065] As a preferred embodiment of the AI vision-based factory safety monitoring and emergency response system of the present invention, the simulation and planning module receives group safety risk events, and the specific steps are as follows:
[0066] The digital twin model of the factory area is invoked. This model is built based on the building information model and includes the three-dimensional geometry and attribute data of walls, doors and windows, stairs, ventilation ducts and fire protection facilities.
[0067] Map the location coordinates of the people involved in the incident to the three-dimensional spatial coordinate system of the digital twin model to determine their floor, room number, and relative position;
[0068] Initiate the physical field simulation process to simulate the accident evolution path, including smoke diffusion or harmful gas spread; divide the three-dimensional factory space into cubic cells with a side length of δ to form a three-dimensional mesh;
[0069] Set the initial state: set the initial concentration and temperature for the cells at the accident source location, and set the background value for other cells;
[0070] The state of each cell is updated iteratively according to time steps, and the concentration update formula is:
[0071]
[0072] Among them, C t (c) represents the concentration of cell c at time t, N(c) is its neighborhood cell set, β is the diffusion coefficient, and γ is the sedimentation coefficient;
[0073] Temperature updates take into account the effects of thermal buoyancy and the ventilation system; the vertical airflow component is determined by the ventilation pressure difference.
[0074] Run the simulation to the preset time end point and predict the expansion range of the danger zone at multiple future time steps;
[0075] Generate a dynamic obstacle avoidance mesh and mark cells with a concentration higher than the safety threshold as impassable nodes;
[0076] Starting from each person's current location and ending at the nearest safe exit, a pathfinding algorithm is run on the obstacle avoidance grid to calculate the optimal escape route that avoids dangerous areas, outputting a path point sequence {p1, p2, ..., p...}. L}
[0077] As a preferred embodiment of the AI vision-based factory safety monitoring and emergency response system of the present invention, the linkage execution module receives the path point sequence of the optimal escape route, and the specific steps are as follows:
[0078] Analyze the path point sequence, extract the emergency lighting device numbers that need to be lit, control them to be activated in sequence, and gradually change the light color from green at the beginning to red at the end to indicate the direction of travel;
[0079] The route navigation instructions are packaged and sent to the augmented reality terminals worn by the affected personnel. The instructions include a sequence of path point coordinates and forward direction arrows.
[0080] After receiving the data, the terminal overlays a virtual guide graphic onto its display interface, which updates in real time as the user moves.
[0081] Generate rescue mission assignment information, which includes the target location coordinates, the number of personnel involved, and the estimated response time, and send it to the mobile terminal of the nearest inspection personnel;
[0082] Activate the broadcasting system and play voice prompts to the risk area through directional audio equipment. The voice content is dynamically generated based on the location of the personnel, and the pointing angle of the playback device is adjusted according to the personnel's position.
[0083] All the linked operations are initiated after the path point sequence is received.
[0084] As a preferred embodiment of the AI vision-based factory safety monitoring and emergency response system of the present invention, the model update module operates when the same system architecture is deployed in multiple factories, and the specific steps are as follows:
[0085] Each factory terminal retains the original image, video, and sensor data locally and does not transmit them across domains.
[0086] Within a set period, the parameters of the graph attention network and risk assessment model are fine-tuned using newly added monitoring samples to obtain the model parameter update amount;
[0087] The model parameter updates are encrypted and uploaded to the central server, which then receives updates from N factory areas.
[0088] Based on the number of risk events occurring in each factory area during the past period, E i The aggregate weight is calculated using the following expression:
[0089]
[0090] The global update direction is calculated using the following expression:
[0091]
[0092] Using Δθ global Update global model parameters θ global ;
[0093] The updated global model was encrypted and distributed to terminals in each factory area;
[0094] Each terminal receives the data and replaces its local model, completing one model update cycle.
[0095] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the computer program, when executed by the processor, implements any step of the AI vision-based factory safety monitoring and emergency response system as described in the first aspect of the present invention.
[0096] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the AI vision-based factory safety monitoring and emergency response system as described in the first aspect of the present invention.
[0097] The beneficial effects of this invention are as follows: By constructing a closed-loop safety management system that integrates multimodal perception fusion and AI, the robustness of personnel behavior recognition and the level of intelligent emergency response in complex industrial environments are significantly improved. In visually limited scenarios such as low light and occlusion, millimeter-wave radar and infrared thermal imaging are used to achieve non-visual modal attitude completion, ensuring all-weather monitoring capabilities. By constructing a personnel interaction graph and introducing a graph attention network, abnormal behavior patterns in group collaboration are effectively captured, improving the accuracy of identifying high-risk violations such as unlicensed operations and lack of supervision. Combined with digital twin and physical simulation technologies, dynamic prediction of accident evolution trends and personalized escape route planning are achieved, and a collaborative response mechanism is formed by linking multiple terminal devices such as lighting, AR navigation, and broadcasting. Attached Figure Description
[0098] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0099] Figure 1 This is a schematic diagram of the AI vision-based factory safety monitoring and emergency response system in Example 1. Detailed Implementation
[0100] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0101] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0102] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0103] Example, refer to Figure 1 This embodiment of the invention provides a factory safety monitoring and emergency response system based on AI vision, comprising:
[0104] The module includes a perception and acquisition module, a feature completion module, a behavior analysis module, a risk assessment module, a deduction and planning module, a linkage execution module, and a model update module.
[0105] The sensing and acquisition module is used to deploy multimodal sensing units in the factory area to acquire visible light images, infrared thermal imaging data and millimeter-wave radar point clouds, and to perform time stamp alignment and spatial coordinate mapping on the three types of data to output the raw sensing dataset.
[0106] Furthermore, in the sensing and acquisition module, the multimodal sensing unit is installed on a fixed bracket above the key operating areas and passageways in the factory area. The specific steps are as follows:
[0107] An integrated sensing unit is deployed at each monitoring point. This sensing unit includes a visible light camera, an infrared thermal imager, and a millimeter-wave radar. The three types of sensors share the same mounting base.
[0108] The sensing and acquisition module also includes a time synchronization unit, used to attach a unified time reference to the data acquired by each type of sensor, specifically including:
[0109] The time synchronization unit receives external time signals and generates high-precision time pulses;
[0110] When each sensor acquires a data frame, it appends the current time pulse value as a timestamp to the header of the data packet, so that the visible light image, infrared image and radar point cloud have the same time reference.
[0111] After time alignment is completed, spatial coordinate mapping is performed, a calibration target is placed at a preset distance in front of the sensing unit, and data is collected by the three types of sensors at the same time.
[0112] Extracting the feature point pixel coordinates u of the calibration target from a visible light image. i ,v i , where i represents the i-th feature point;
[0113] Extracting the three-dimensional spatial coordinates x of corresponding feature points from millimeter-wave radar point clouds. j ,y j ,z j , where j corresponds one-to-one with i;
[0114] The projection relationship is established by the expression:
[0115]
[0116] Where s is the scaling factor, K is the intrinsic parameter matrix of the visible light camera, which includes the focal length and principal point parameters, R is the rotation matrix, and T is the translation vector;
[0117] Spatial calibration is completed by minimizing the reprojection error of all feature points and solving for the optimal R and T.
[0118] The radar point cloud acquired at any time can be projected onto the visible light image plane through this transformation matrix to achieve cross-modal space alignment;
[0119] The output raw perception data set contains time-synchronized and spatially aligned trimodal data, with each set having the same timestamp, and is used for subsequent feature fusion and behavior analysis;
[0120] It should be noted that the integrated design of the multimodal sensing unit not only ensures the rigid connection of the three types of sensors in space, avoiding calibration failure due to equipment displacement, but also achieves microsecond-level time synchronization through a unified time synchronization mechanism, providing a high-precision spatiotemporal reference for subsequent cross-modal data fusion. This structure is suitable for complex factory environments such as day-night cycles and rain / fog obstruction, ensuring the continuity and consistency of the sensed data.
[0121] The feature completion module is used to judge the environmental quality based on the brightness and contrast indicators of the visible light image. When the indicators are lower than the preset threshold, it extracts human point cloud clusters based on millimeter-wave radar point cloud, generates three-dimensional pose key points through clustering and skeleton regression algorithms, and performs fusion verification by combining the human thermal contour in the infrared image to output human pose feature map.
[0122] Furthermore, the feature completion module receives the raw perceptual data set, and the specific steps are as follows:
[0123] Perform a quality assessment on the current frame of the visible light image to determine whether it meets the conditions for normal visual analysis.
[0124] Convert the image to grayscale to obtain a pixel matrix;
[0125] The overall average brightness L of the image is calculated, which is defined as the arithmetic mean of the gray values of all pixels, and the expression is:
[0126]
[0127] Where M is the image height and N is the image width;
[0128] The image is divided into several regular sub-blocks. The gray-level range within each sub-block is calculated as the local contrast, and the average contrast of all sub-blocks is calculated as the global contrast.
[0129] The L and global contrast are compared with preset quality judgment conditions. If L is lower than the first threshold and the global contrast is lower than the second threshold, the current environment is judged to be a low-quality visual environment.
[0130] Initiate a cross-modal feature completion process by acquiring a point cloud dataset from a millimeter-wave radar, where each point has three-dimensional coordinates;
[0131] Density clustering algorithm is used to process the point cloud dataset. The neighborhood search radius and minimum number of points are set, and points that meet the density connectivity are divided into the same cluster, resulting in multiple personnel point cloud clusters.
[0132] Apply a pre-trained 3D pose estimation model to each point cloud cluster to output the 3D coordinates of key human joints;
[0133] Meanwhile, the thermal radiation area of the human body is extracted from the infrared image, the high temperature area is obtained by temperature threshold segmentation, and then the complete thermal outline of the human body is obtained by edge detection and morphological closing operation.
[0134] The radar-generated skeleton structure is projected onto the infrared image plane, and its spatial overlap with the thermal profile is calculated. It is defined as the ratio of the length of the overlap between the skeleton line segment and the thermal profile to the total length of the skeleton.
[0135] If the spatial overlap is less than the preset matching threshold, the clustering parameters are adjusted and the point cloud clusters are extracted again until the overlap meets the requirements, and the completed human pose feature map is generated.
[0136] It should be noted that the feature completion module can accurately reconstruct human posture even when the visible light image quality degrades by fusing the 3D point cloud of millimeter-wave radar with the temperature distribution information of infrared thermal imaging. This effectively overcomes the limitations of a single visual modality in low-light and smoke-dust interference scenarios. The introduced fusion verification mechanism improves the reliability of posture estimation and prevents subsequent analysis deviations caused by false detections or missed detections.
[0137] The behavior analysis module is used to perform target detection on human posture feature maps, extract personnel position and posture vectors, construct personnel interaction graph with personnel as nodes and the angle between spatial distance and line of sight as edge weights, input graph attention network to calculate the attention score of each node, and output interaction risk vector representing the degree of abnormality in group cooperation;
[0138] Furthermore, the behavior analysis module receives human posture feature maps, and the specific steps are as follows:
[0139] The geometric center of each person's joint coordinates is calculated to obtain their center position in the monitoring screen.
[0140] Extract the pose vector, which includes the head orientation angle, torso tilt angle, and limb extension state;
[0141] Construct a personnel interaction diagram with each person as a node;
[0142] For any two nodes j and i, calculate their Euclidean distance d. ij The expression is:
[0143]
[0144] If d ij If the distance is less than the preset threshold, then a connection edge is established between node j and i;
[0145] The weight of an edge is determined by both spatial proximity and line-of-sight consistency;
[0146] The constructed personnel interaction graph is input into a graph attention network, which contains multiple layers of attention mechanisms. Each layer calculates the attention coefficient of each node i to its neighbor node j, expressed as:
[0147]
[0148] Among them, h i Let be the feature vector of node i, W be the trainable weight matrix, a be the attention parameter vector, || denote vector concatenation, and N be the feature vector of node i. iLet i be the set of neighbors of node i;
[0149] Using α ij Weighted aggregation of neighbor features is used to update the node representation;
[0150] After multiple layers of propagation, the attention score of each node is output. All attention scores constitute the interaction risk vector S, which is used to characterize the salience of an individual's behavior in the group.
[0151] It should be noted that the interaction graph constructed by the behavior analysis module not only reflects the spatial proximity between individuals, but also introduces the consistency of gaze direction as an edge weight element, which can more accurately depict the intentional interaction and collaborative state between people. The graph attention network can identify abnormal attention distributions that deviate from the normal collaborative mode by learning the dynamic influence weights between nodes, providing interpretable evidence for group behavior risk assessment.
[0152] The risk assessment module is used to match interactive risk vectors with preset group violation patterns. When the correlation between the risk vector and the combination of unlicensed operation, lack of supervision, or unauthorized operation exceeds a threshold, a group safety risk event is generated.
[0153] Furthermore, the risk assessment module receives interactive risk vectors, and the specific steps are as follows:
[0154] Retrieve preset group violation pattern templates from the local rule base, including templates for unlicensed operation, lack of supervision, and unauthorized combination templates;
[0155] The relevance between the interaction risk vector and each template is calculated using the following expression:
[0156]
[0157] Where · represents the vector dot product;
[0158] Each R m Matching threshold R th Comparison, if R m >R th If so, then the current scenario is determined to match the violation pattern of the m-th group;
[0159] Generate a group safety risk event. The event record includes the event type identifier m, the occurrence time t, the set of personnel numbers involved and their center location coordinates;
[0160] The event is transmitted as a trigger signal to the simulation planning module to initiate subsequent emergency simulation procedures;
[0161] It should be noted that the risk assessment module adopts a vector similarity matching mechanism, which compares the real-time calculated interactive risk vector with the preset violation pattern template. This avoids the strong dependence of traditional rule engines on explicit logic. By setting a matching threshold, it achieves flexible judgment, taking into account both recognition sensitivity and false alarm suppression. It is suitable for risk identification in a variety of typical high-risk operation scenarios.
[0162] The simulation and planning module is used to call the digital twin model of the factory area after a mass safety risk event is generated, map the location information of the affected personnel to the BIM space, start the smoke diffusion simulation in combination with the ventilation system layout and combustible material distribution, predict the range of the danger zone in the future time step, generate a dynamic obstacle avoidance grid based on the danger zone and personnel location, and use a path search algorithm to calculate the optimal escape path in the obstacle avoidance grid.
[0163] Furthermore, the simulation and planning module receives information on mass security risk events, and the specific steps are as follows:
[0164] The digital twin model of the factory area is invoked. This model is built based on the building information model and includes the three-dimensional geometry and attribute data of walls, doors and windows, stairs, ventilation ducts and fire protection facilities.
[0165] Map the location coordinates of the people involved in the incident to the three-dimensional spatial coordinate system of the digital twin model to determine their floor, room number, and relative position;
[0166] Initiate the physical field simulation process to simulate the accident evolution path, including smoke diffusion or harmful gas spread; divide the three-dimensional factory space into cubic cells with a side length of δ to form a three-dimensional mesh;
[0167] Set the initial state: set the initial concentration and temperature for the cells at the accident source location, and set the background value for other cells;
[0168] The state of each cell is updated iteratively according to time steps, and the concentration update formula is:
[0169]
[0170] Among them, C t (c) represents the concentration of cell c at time t, N(c) is its neighborhood cell set, β is the diffusion coefficient, and γ is the sedimentation coefficient;
[0171] Temperature updates take into account the effects of thermal buoyancy and the ventilation system; the vertical airflow component is determined by the ventilation pressure difference.
[0172] Run the simulation to the preset time end point and predict the expansion range of the danger zone at multiple future time steps;
[0173] Generate a dynamic obstacle avoidance mesh and mark cells with a concentration higher than the safety threshold as impassable nodes;
[0174] Starting from each person's current location and ending at the nearest safe exit, a pathfinding algorithm is run on the obstacle avoidance grid to calculate the optimal escape route that avoids dangerous areas, outputting a path point sequence {p1, p2, ..., p...}. L};
[0175] It should be noted that the simulation planning module combines the factory area BIM model with physical field simulation algorithms, which can quickly predict the evolution trend of dangerous areas in the early stage of an accident. It breaks through the limitations of the static response of traditional emergency plans. The combination of dynamic obstacle avoidance grid and optimal path search provides personalized escape guidance for each trapped person, significantly improving the scientific nature and timeliness of emergency evacuation.
[0176] The linkage execution module is used to control the emergency lighting system in the factory area to illuminate the guide route according to the optimal escape route, send path navigation instructions to the AR terminals worn by affected personnel, send rescue task assignment information to the mobile terminals of the nearest inspection personnel, and start the broadcast system to play directional voice prompts.
[0177] Furthermore, the coordinated execution module receives the path point sequence of the optimal escape route, and the specific steps are as follows:
[0178] Analyze the path point sequence, extract the emergency lighting device numbers that need to be lit, control them to be activated in sequence, and gradually change the light color from green at the beginning to red at the end to indicate the direction of travel;
[0179] The route navigation instructions are packaged and sent to the augmented reality terminals worn by the affected personnel. The instructions include a sequence of path point coordinates and forward direction arrows.
[0180] After receiving the data, the terminal overlays a virtual guide graphic onto its display interface, which updates in real time as the user moves.
[0181] Generate rescue mission assignment information, which includes the target location coordinates, the number of personnel involved, and the estimated response time, and send it to the mobile terminal of the nearest inspection personnel;
[0182] Activate the broadcasting system and play voice prompts to the risk area through directional audio equipment. The voice content is dynamically generated based on the location of the personnel, and the pointing angle of the playback device is adjusted according to the personnel's position.
[0183] All linked operations are initiated upon receiving the path point sequence;
[0184] It should be noted that the linkage execution module enables the coordinated control of multiple types of terminal devices, including emergency lighting, AR navigation, mobile terminals and broadcasting systems, forming a three-dimensional emergency response system covering visual, auditory and spatial guidance. All instructions are triggered synchronously based on a unified spatiotemporal reference, ensuring the consistency and authority of information transmission and reducing cognitive confusion among personnel in emergency situations.
[0185] The model update module is used to retain the original data locally in each plant when the same system architecture is deployed in multiple plants. The model parameter gradients are periodically uploaded to the central server, and the central server performs weighted aggregation of the gradients to generate a globally updated model.
[0186] Furthermore, when the model update module is deployed with the same system architecture in multiple factory areas, the specific steps are as follows:
[0187] Each factory terminal retains the original image, video, and sensor data locally and does not transmit them across domains.
[0188] Within a set period, the parameters of the graph attention network and risk assessment model are fine-tuned using newly added monitoring samples to obtain the model parameter update amount;
[0189] The model parameter updates are encrypted and uploaded to the central server, which then receives updates from N factory areas.
[0190] Based on the number of risk events occurring in each factory area during the past period, E i The aggregate weight is calculated using the following expression:
[0191]
[0192] The global update direction is calculated using the following expression:
[0193]
[0194] Using Δθ global Update global model parameters θ global ;
[0195] The updated global model was encrypted and distributed to terminals in each factory area;
[0196] Each terminal receives the data and replaces its local model, completing one model update cycle.
[0197] It should be noted that the model update module adopts a federated learning architecture, which enables cross-plant knowledge sharing without centralized transmission of raw data. This protects the data privacy and security of each plant, and through a weighted aggregation mechanism, the global model continuously absorbs abnormal behavior features under new scenarios, thus possessing good scalability and long-term evolution capabilities.
[0198] This embodiment also provides a computer device suitable for an AI vision-based factory safety monitoring and emergency response system, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the AI vision-based factory safety monitoring and emergency response system proposed in the above embodiment.
[0199] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0200] This embodiment also provides a storage medium storing a computer program. When executed by a processor, the program implements the AI vision-based factory safety monitoring and emergency response system as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0201] In summary, this invention significantly improves the robustness of personnel behavior recognition and the intelligence level of emergency response in complex industrial environments by constructing a closed-loop safety management system based on multimodal perception fusion and AI-driven technology. In visually limited scenarios such as low light and occlusion, millimeter-wave radar and infrared thermal imaging are used to achieve non-visual modal attitude completion, ensuring all-weather monitoring capabilities. By constructing a personnel interaction graph and introducing a graph attention network, abnormal behavior patterns in group collaboration are effectively captured, improving the accuracy of identifying high-risk violations such as unlicensed operations and lack of supervision. Combining digital twin and physical simulation technologies, dynamic prediction of accident evolution trends and personalized escape route planning are achieved, and a collaborative response mechanism is formed by linking multiple terminal devices such as lighting, AR navigation, and broadcasting.
[0202] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A factory safety monitoring and emergency response system based on AI vision, characterized in that: include: The module includes a perception and acquisition module, a feature completion module, a behavior analysis module, a risk assessment module, a deduction and planning module, a linkage execution module, and a model update module. The sensing and acquisition module is used to deploy multimodal sensing units in the factory area to acquire visible light images, infrared thermal imaging data and millimeter-wave radar point clouds, and to perform time stamp alignment and spatial coordinate mapping on the three types of data to output the original sensing dataset. The feature completion module is used to judge the environmental quality based on the brightness and contrast indicators of the visible light image. When the indicators are lower than the preset threshold, it extracts human point cloud clusters based on millimeter-wave radar point cloud, generates three-dimensional pose key points through clustering and skeleton regression algorithms, and performs fusion verification by combining the human thermal contour in the infrared image to output human pose feature map. The behavior analysis module is used to perform target detection on the human posture feature map, extract the position and posture vector of the person, construct the person interaction graph with the person as the node and the angle between the spatial distance and the line of sight as the edge weight, input the graph attention network to calculate the attention score of each node, and output the interaction risk vector representing the degree of abnormality in group cooperation. The risk assessment module is used to match the interactive risk vector with the preset group violation pattern. When the correlation between the risk vector and the combination of unlicensed operation, lack of supervision or unauthorized operation exceeds the threshold, a group safety risk event is generated. The simulation and planning module is used to call the digital twin model of the factory area after a mass safety risk event is generated, map the location information of the affected personnel to the BIM space, start the smoke diffusion simulation in combination with the ventilation system layout and combustible material distribution, predict the range of the dangerous area in the future time step, generate a dynamic obstacle avoidance grid based on the dangerous area and personnel location, and use a path search algorithm to calculate the optimal escape path in the obstacle avoidance grid. The linkage execution module is used to control the emergency lighting system in the factory area to illuminate the guide route according to the optimal escape route, send path navigation instructions to the AR terminals worn by the affected personnel, send rescue task assignment information to the mobile terminal of the nearest inspection personnel, and start the broadcast system to play directional voice prompts. The model update module is used to retain the original data locally in each factory area when the same system architecture is deployed in multiple factory areas, and periodically upload the model parameter gradients to the central server. The central server performs weighted aggregation of the gradients to generate a globally updated model.
2. The factory safety monitoring and emergency response system based on AI vision as described in claim 1, characterized in that: In the sensing and acquisition module, the multimodal sensing unit is installed on a fixed bracket above the key operating areas and passageways in the factory area. The specific steps are as follows: An integrated sensing unit is deployed at each monitoring point. This sensing unit includes a visible light camera, an infrared thermal imager, and a millimeter-wave radar. The three types of sensors share the same mounting base. The sensing and acquisition module also includes a time synchronization unit, used to attach a unified time reference to the data acquired by each type of sensor, specifically including: The time synchronization unit receives external time signals and generates high-precision time pulses; When each sensor acquires a data frame, it appends the current time pulse value as a timestamp to the header of the data packet, so that the visible light image, infrared image and radar point cloud have the same time reference. After time alignment is completed, spatial coordinate mapping is performed, a calibration target is placed at a preset distance in front of the sensing unit, and data is collected by the three types of sensors at the same time. Extracting the feature point pixel coordinates u of the calibration target from a visible light image. i ,v i , where i represents the i-th feature point; Extracting the three-dimensional spatial coordinates x of corresponding feature points from millimeter-wave radar point clouds. j ,y j ,z j , where j corresponds one-to-one with i; The projection relationship is established by the expression: Where s is the scaling factor, K is the intrinsic parameter matrix of the visible light camera, which includes the focal length and principal point parameters, R is the rotation matrix, and T is the translation vector; Spatial calibration is completed by minimizing the reprojection error of all feature points and solving for the optimal R and T. The radar point cloud acquired at any time can be projected onto the visible light image plane through this transformation matrix to achieve cross-modal space alignment; The output raw perception data set contains time-synchronized and spatially aligned trimodal data, with each set having the same timestamp, and is used for subsequent feature fusion and behavior analysis.
3. The factory safety monitoring and emergency response system based on AI vision as described in claim 2, characterized in that: The feature completion module receives the original sensing data set, and the specific steps are as follows: Perform a quality assessment on the current frame of the visible light image to determine whether it meets the conditions for normal visual analysis. Convert the image to grayscale to obtain a pixel matrix; The overall average brightness L of the image is calculated, which is defined as the arithmetic mean of the gray values of all pixels, and the expression is: Where M is the image height and N is the image width; The image is divided into several regular sub-blocks. The gray-level range within each sub-block is calculated as the local contrast, and the average contrast of all sub-blocks is calculated as the global contrast. The L and global contrast are compared with preset quality judgment conditions. If L is lower than the first threshold and the global contrast is lower than the second threshold, the current environment is judged to be a low-quality visual environment. Initiate a cross-modal feature completion process by acquiring a point cloud dataset from a millimeter-wave radar, where each point has three-dimensional coordinates; Density clustering algorithm is used to process the point cloud dataset. The neighborhood search radius and minimum number of points are set, and points that meet the density connectivity are divided into the same cluster, resulting in multiple personnel point cloud clusters. Apply a pre-trained 3D pose estimation model to each point cloud cluster to output the 3D coordinates of key human joints; Meanwhile, the thermal radiation area of the human body is extracted from the infrared image, the high temperature area is obtained by temperature threshold segmentation, and then the complete thermal outline of the human body is obtained by edge detection and morphological closing operation. The radar-generated skeleton structure is projected onto the infrared image plane, and its spatial overlap with the thermal profile is calculated. It is defined as the ratio of the length of the overlap between the skeleton line segment and the thermal profile to the total length of the skeleton. If the spatial overlap is less than the preset matching threshold, the clustering parameters are adjusted and point cloud clusters are extracted again until the overlap meets the requirements, and a completed human pose feature map is generated.
4. The factory safety monitoring and emergency response system based on AI vision as described in claim 3, characterized in that: The behavior analysis module receives the human posture feature map, and the specific steps are as follows: The geometric center of each person's joint coordinates is calculated to obtain their center position in the monitoring screen. Extract the pose vector, which includes the head orientation angle, torso tilt angle, and limb extension state; Construct a personnel interaction diagram with each person as a node; For any two nodes j and i, calculate their Euclidean distance d. ij The expression is: If d ij If the distance is less than the preset threshold, then a connection edge is established between node j and i; The weight of an edge is determined by both spatial proximity and line-of-sight consistency; The constructed personnel interaction graph is input into a graph attention network, which contains multiple layers of attention mechanisms. Each layer calculates the attention coefficient of each node i to its neighbor node j, expressed as: Among them, h i Let be the feature vector of node i, W be the trainable weight matrix, a be the attention parameter vector, || denote vector concatenation, and N be the feature vector of node i. i Let i be the set of neighbors of node i; Using α ij Weighted aggregation of neighbor features is used to update the node representation; After multiple layers of propagation, the attention score of each node is output. All attention scores constitute the interaction risk vector S, which is used to characterize the salience of an individual's behavior in the group.
5. The factory safety monitoring and emergency response system based on AI vision as described in claim 4, characterized in that: The risk assessment module receives the interactive risk vector, and the specific steps are as follows: Retrieve preset group violation pattern templates from the local rule base, including templates for unlicensed operation, lack of supervision, and unauthorized combination templates; The relevance between the interaction risk vector and each template is calculated using the following expression: Where · represents the vector dot product; Each R m Matching threshold R th Comparison, if R m >R th If so, then the current scenario is determined to match the violation pattern of the m-th group; Generate a group safety risk event. The event record includes the event type identifier m, the occurrence time t, the set of personnel numbers involved and their center location coordinates; The event is transmitted as a trigger signal to the simulation planning module to initiate subsequent emergency simulation procedures.
6. The factory safety monitoring and emergency response system based on AI vision as described in claim 5, characterized in that: The simulation and planning module receives mass security risk events, and the specific steps are as follows: The digital twin model of the factory area is invoked. This model is built based on the building information model and includes the three-dimensional geometry and attribute data of walls, doors and windows, stairs, ventilation ducts and fire protection facilities. Map the location coordinates of the people involved in the incident to the three-dimensional spatial coordinate system of the digital twin model to determine their floor, room number, and relative position; Initiate the physical field simulation process to simulate the accident evolution path, including smoke diffusion or harmful gas spread; divide the three-dimensional factory space into cubic cells with a side length of δ to form a three-dimensional mesh; Set the initial state: set the initial concentration and temperature for the cells at the accident source location, and set the background value for other cells; The state of each cell is updated iteratively according to time steps, and the concentration update formula is: Among them, C t (c) represents the concentration of cell c at time t, N(c) is its neighborhood cell set, β is the diffusion coefficient, and γ is the sedimentation coefficient; Temperature updates take into account the effects of thermal buoyancy and the ventilation system; the vertical airflow component is determined by the ventilation pressure difference. Run the simulation to the preset time end point and predict the expansion range of the danger zone at multiple future time steps; Generate a dynamic obstacle avoidance mesh and mark cells with a concentration higher than the safety threshold as impassable nodes; Starting from each person's current location and ending at the nearest safe exit, a pathfinding algorithm is run on the obstacle avoidance grid to calculate the optimal escape route that avoids dangerous areas, outputting a path point sequence {p1, p2, ..., p...}. L } 7. The factory safety monitoring and emergency response system based on AI vision as described in claim 6, characterized in that: The linkage execution module receives the path point sequence of the optimal escape path, and the specific steps are as follows: Analyze the path point sequence, extract the emergency lighting device numbers that need to be lit, control them to be activated in sequence, and gradually change the light color from green at the beginning to red at the end to indicate the direction of travel; The route navigation instructions are packaged and sent to the augmented reality terminals worn by the affected personnel. The instructions include a sequence of path point coordinates and forward direction arrows. After receiving the data, the terminal overlays a virtual guide graphic onto its display interface, which updates in real time as the user moves. Generate rescue mission assignment information, which includes the target location coordinates, the number of personnel involved, and the estimated response time, and send it to the mobile terminal of the nearest patrol personnel; Activate the broadcasting system and play voice prompts to the risk area through directional audio equipment. The voice content is dynamically generated based on the location of the personnel, and the pointing angle of the playback device is adjusted according to the personnel's position. All the linked operations are initiated after the path point sequence is received.
8. The factory safety monitoring and emergency response system based on AI vision as described in claim 7, characterized in that: The model update module operates when the same system architecture is deployed in multiple factory areas. The specific steps are as follows: Each factory terminal retains the original image, video, and sensor data locally and does not transmit them across domains. Within a set period, the parameters of the graph attention network and risk assessment model are fine-tuned using newly added monitoring samples to obtain the model parameter update amount; The model parameter updates are encrypted and uploaded to the central server, which then receives updates from N factory areas. Based on the number of risk events E occurring in each factory area over the past period i The aggregate weight is calculated using the following expression: The global update direction is calculated using the following expression: Using Δθ global Update global model parameters θ global ; The updated global model was encrypted and distributed to terminals in each factory area; Each terminal receives the data and replaces its local model, completing one model update cycle.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the AI vision-based factory safety monitoring and emergency response system as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the AI vision-based factory safety monitoring and emergency response system as described in any one of claims 1 to 8.
Citation Information
Cited By
Water moving object detection method and device based on unmanned aerial vehicle, and electronic equipment
CN121414786A
Unmanned aerial vehicle cluster control method, device, equipment and medium
CN121657709A
Unmanned aerial vehicle swarm control method, apparatus, device, and medium
CN121657709B