Safety production monitoring method and system based on AI video analysis

By utilizing target detection models and image matching technology in industrial settings, key images with high confidence and completeness are selected, solving the problem of low efficiency in monitoring systems caused by too many views and achieving efficient and accurate safety production monitoring.

CN120823566BActive Publication Date: 2025-11-21ZHONGNAN INFORMATION TECH (SHENZHEN) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511333888.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-11-21
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

Existing technologies cannot effectively select key images when faced with too many views, resulting in low efficiency in alarm analysis of monitoring systems and problems of false detection and missed detection.

Method used

By acquiring current-time images from multiple acquisition devices, the target detection model outputs confidence and integrity factors, calculates image matching degree, and filters out key images for early warning analysis. Combining persistence score and feature change score, images with high confidence and integrity are prioritized for monitoring.

Benefits of technology

It significantly reduces computing resource consumption, avoids false detections, improves the accuracy and efficiency of early warning analysis, and enhances the robustness of the monitoring system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823566B_ABST
    Figure CN120823566B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing, in particular to a safety production monitoring method and system based on AI video analysis. The method comprises: acquiring current time images of a plurality of acquisition devices in each target grid area; inputting each current time image into a target detection model to output target detection results of each frame image; calculating the matching degree of any two current time images, sorting the matching degrees in descending order, and selecting the top m matching degrees corresponding to the image pairs; calculating the screening criteria of the two current time images in any image pair respectively, and taking the current time image corresponding to the larger one of the two screening criteria of any image pair as the key image; and performing early warning analysis on the key images of each target grid area to output alarm information. The scheme of the present application can improve the efficiency of monitoring and alarm analysis when facing excessive views.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology. More specifically, this invention relates to a safety production monitoring method and system based on AI video analysis. Background Technology

[0002] In the fields of intelligent manufacturing and industrial automation, real-time monitoring of production areas is a crucial link in ensuring production efficiency, safety, and quality. With the rapid development of computer vision and artificial intelligence technologies, image-based target detection and analysis have been widely applied in production environment monitoring. These systems deploy cameras on the production site and utilize deep learning models (such as target detection and behavior recognition) to analyze real-time video streams. This allows them to automatically identify unsafe behaviors or states, such as not wearing safety helmets, smoking in violation of regulations, and personnel entering dangerous areas, thus achieving the purpose of monitoring. Once a pre-set violation is detected, the system generates alarm information and pushes it to safety management personnel, thereby achieving automated, 24 / 7 supervision of the production site.

[0003] In existing technologies, target detection is typically based on a single-frame image, outputting the bounding box coordinates and confidence score of the target in the image. However, in practice, it has been found that due to the complex industrial environment (such as changes in lighting, angular occlusion, and object reflection) and the inherent sensitivity of the algorithm itself, single-frame detection is easily affected by the aforementioned factors such as changes in lighting, occlusion, or noise, leading to false detections or missed detections. A significant portion of these false alarms are caused by algorithm misjudgment or are instantaneous alarms triggered by brief, unintentional, and non-continuous human actions.

[0004] Moreover, with the continuous increase in the number of video surveillance cameras, key control points have achieved full coverage of video surveillance. At this point, not only is there the problem of excessive resource consumption and inability to select views due to the large number of monitored views being displayed in excess, but there is also the problem of constantly occurring instantaneous alarms, making safety production monitoring quite difficult.

[0005] Therefore, existing technologies face the problem of inefficient alarm analysis in monitoring systems due to the inability to effectively select key images caused by too many views. Summary of the Invention

[0006] The purpose of this invention is to propose a safety production monitoring method and system based on AI video analysis, in order to solve the problem that the efficiency of monitoring and alarm analysis cannot be guaranteed when faced with excessive views in the prior art; to this end, this invention provides solutions in the following two aspects.

[0007] In a first aspect, the present invention provides a safety production monitoring method based on AI video analysis, comprising:

[0008] Acquire current-time images from multiple acquisition devices within each target grid area;

[0009] The target detection model inputs the images at each current time moment and outputs the target detection results for each frame; the target detection results include the confidence scores of multiple targets;

[0010] Calculate the matching degree between any two images at the current time, sort the matching degrees in descending order, and select the image pairs corresponding to the first m matching degrees, where m≥1;

[0011] Calculate the selection criteria for the two current time images in any image pair, and take the current time image corresponding to the larger of the two selection criteria for any image pair as the key image; the selection criteria are positively correlated with the confidence score and the completeness factor of each target in the corresponding current time image; the completeness factor characterizes the completeness of the target within the bounding box;

[0012] Early warning analysis is performed on key images of each target grid area to output alarm information.

[0013] The above scheme analyzes the target detection of images captured by multiple acquisition devices within the same target grid area at the current moment, as well as the matching of images at different angles at the current moment. Combined with the target integrity factor, key images are selected, enabling the monitoring system to focus more on important views for subsequent monitoring and analysis. This significantly reduces the consumption of computing and storage resources, while also avoiding instantaneous alarms caused by false detections, thus improving the accuracy of early warning analysis.

[0014] Optionally, the screening criterion is the sum of the products of the confidence scores of all targets in the image at the corresponding current time and the completeness factor.

[0015] The above scheme combines confidence level and target completeness factor, prioritizing images that include high confidence and display more complete target locations, making it suitable for monitoring needs.

[0016] Optionally, the integrity factor is the product of the ratio of the number of detected keypoints to the total number of keypoints and the penalty coefficient; the penalty coefficient is the ratio of the number of detected valid keypoints to the total number of valid keypoints; and the valid keypoints are keypoints that have an interactive relationship between the target and the dangerous area.

[0017] The aforementioned completeness factor can prioritize views that show the entire target or key parts.

[0018] Optionally, the confidence level is the product of the correction coefficient and the initial confidence level; the correction coefficient is positively correlated with the persistence score and feature change score of the corresponding target;

[0019] The process of obtaining the persistence score is as follows:

[0020] Acquire a sequence of grayscale images containing the current time frame from any acquisition device, use a tracking algorithm to track the target, and use the percentage of frames in which the target appears as the corresponding persistence score.

[0021] The feature change score is negatively correlated with the Euclidean distance between the feature vector of the target in the current moment image and the corresponding target in the previous moment image.

[0022] The above approach, by leveraging the temporal change characteristics of targets (persistence score and feature change score), can assess the temporal stability of targets, providing data support for the subsequent selection of key images.

[0023] Optionally, the feature vector is the coordinates of the bounding box of the target in the target detection result, specifically including the horizontal and vertical coordinates of the center point of the bounding box, and the length and width of the bounding box.

[0024] The aforementioned feature vectors can quantify the changes in the appearance of the target at different times, providing data support for subsequent screening of key images.

[0025] Optionally, the matching degree is the degree of similarity between any two images at the current time obtained by using an image registration algorithm.

[0026] Optionally, the target grid region is:

[0027] Before acquiring the current-time images from each acquisition device, the production area is first divided into grids to obtain the acquisition devices in multiple grid areas. Each grid area is then marked, and the grid areas marked as dangerous are designated as target grid areas.

[0028] Optionally, the object detection model is a YOLOv8-L model or a Faster R-CNN model.

[0029] Optionally, the step of performing early warning analysis on key images of each target grid region to output alarm information includes:

[0030] The cross-union ratio (CURRR) of the bounding box of the target in the target detection result and the corresponding danger area in the grid region is determined. When the CURRR is greater than the threshold, an alarm message is output.

[0031] In the second aspect, the safety production monitoring system based on AI video analysis includes:

[0032] processor;

[0033] The memory stores computer instructions for safety production monitoring based on AI video analysis, which, when executed by the processor, cause the system to perform the aforementioned safety production monitoring method based on AI video analysis.

[0034] The beneficial effects of this invention are as follows:

[0035] The solution of this invention integrates multi-device image acquisition, target detection, matching degree calculation and early warning analysis to form a complete closed-loop system that can adapt to complex and ever-changing environments and handle scenes with different perspectives, lighting conditions or target complexity, thus enhancing robustness. Attached Figure Description

[0036] Figure 1 This illustration schematically shows a flowchart of the steps of the safety production monitoring method based on AI video analysis in this embodiment;

[0037] Figure 2 The diagram illustrates the structural block diagram of the safety production monitoring system based on AI video analysis in this embodiment. Detailed Implementation

[0038] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0039] In safe production scenarios, video data is massive, but most of it consists of normal production footage, with only a small portion involving safety incidents. Therefore, it is crucial to filter out important and non-repeating key images from the video data to reduce the resource consumption during viewing. Furthermore, it is essential to analyze these key images to achieve safety monitoring and improve the monitoring efficiency of the surveillance system.

[0040] Therefore, this invention proposes a safety production monitoring method and system based on AI video analysis, which aims to filter out video clips containing safety hazards (such as workers entering dangerous areas) from the massive video data generated by factory or workshop monitoring systems, so as to achieve real-time early warning and key display.

[0041] Taking monitoring workers entering hazardous areas as an example, this embodiment introduces the safety production monitoring method based on AI video analysis.

[0042] like Figure 1 As shown, the safety production monitoring method based on AI video analysis in this embodiment includes the following steps:

[0043] Step S1: Obtain the current time images of multiple acquisition devices within each target grid area.

[0044] For production areas (such as workshop production lines), to ensure coverage of critical areas (such as workbenches, conveyor belts, and hazardous boundaries), cameras are arranged in a grid pattern, with adjacent cameras partially overlapping their fields of view to ensure that the same target (such as workers or equipment) is captured from multiple perspectives. The video stream is transmitted to a local server via Real-Time Transport Protocol (RTSP). The server decodes the video and extracts keyframes at 0.5-second intervals to capture rapidly changing safety events (such as worker misoperation or workers entering hazardous areas).

[0045] In this embodiment, the production area is first divided into multiple grid regions, and each grid region is marked. Grid regions classified as hazardous are designated as target grid regions. It should be noted that a target grid region may contain some hazardous areas, or it may consist entirely of hazardous areas.

[0046] The aforementioned hazard markings can be done manually or automatically based on equipment location, such as marking production lines, areas near high-temperature equipment, and areas of mechanical movement as hazardous.

[0047] Each target grid area includes multiple acquisition devices, which can acquire multiple images of the target grid area at the same time.

[0048] For example, the current moment image is acquired by multiple high-definition cameras deployed within the target grid area, each camera capturing a 1080p (1920×1080 pixels) video stream at a frame rate of 30 frames per second.

[0049] In this embodiment, multiple current-time images within each grid area are also preprocessed, such as denoising, image enhancement, grayscale processing, etc.

[0050] Step S2: Select key images from multiple images at the current time.

[0051] The process of acquiring key images is as follows:

[0052] Step S21: Input the images at each current time into the target detection model and output the target detection results of each frame image; the target detection results include the bounding boxes of multiple targets and their corresponding confidence scores.

[0053] The target detection results include the detection results of multiple targets and their corresponding camera numbers; each current moment image corresponds to one target detection result, and the target detection result includes the initial confidence scores of multiple targets. The initial confidence scores are directly provided by the target detection model, which is a confidence assessment based on appearance features.

[0054] The object detection models mentioned above can be YOLOv8-L or Faster R-CNN.

[0055] Taking the YOLOv8-L model as an example, the YOLOv8-L model is first trained, and then the trained YOLOv8-L model is used to perform forward inference calculations on the input image at the current time step, outputting a set of target detection results for the image at the current time step. Here, the detection result of the k-th target in the t-th image at the current time step is denoted as... .

[0056] Furthermore, each test result may also include:

[0057] The bounding box of the target, such as the coordinates of the bounding box of the i-th target in the image at time t. It precisely defines the position and range of the target in the image. x and y represent the horizontal and vertical coordinates of the center point of the bounding box, respectively, and w and h represent the length and width of the bounding box, respectively. For example, there are bounding boxes of 5 targets in the image at the t-th time.

[0058] Before obtaining the target detection results of the images at each current time, the YOLOv8-L model is also trained. The specific training process is as follows:

[0059] Obtain the training set and category labels; the training set can be historical images under surveillance. The historical images are manually labeled to determine the category labels in the images, such as labeling workers in dangerous areas as 1 and workers not in dangerous areas as 0.

[0060] The training set is input into the YOLOv8-L model for training, and the loss value is calculated using the loss function. The parameters of the network prediction model are adjusted using the gradient descent algorithm until the loss value between the output value and the label is less than the threshold or the number of training iterations reaches the set number. At this point, training stops, and the trained YOLOv8-L model is obtained.

[0061] Since the training process of the YOLOv8-L model is based on existing technology, it will not be described in detail here.

[0062] In this embodiment, to better reflect the true credibility of the target, the initial confidence level needs to be corrected to obtain the final confidence level. That is, the confidence level is the product of the initial confidence level and the correction coefficient.

[0063] Specifically, the process of obtaining the correction coefficient is as follows:

[0064] First, obtain a persistence score.

[0065] Specifically, the process of obtaining the persistence score is as follows:

[0066] Obtain a sequence of grayscale images containing the current time frame from any angle. Use a tracking algorithm to track the target and use the percentage of frames in which the target appears as the corresponding persistence score, which is the ratio of the number of frames in which the target appears to the total number of frames in the grayscale image sequence.

[0067] The number of frames in which the target appears can be determined by setting rules. If the set rules are met, it is considered that the target in the current frame also appears in other frames.

[0068] The above setting rule is that the intersection-union ratio of the bounding boxes of any target in the current image with the bounding boxes of targets in any other frame is greater than a preset threshold.

[0069] The tracking algorithm mentioned above can be IoU-based tracking or Kalman filtering; since IoU-based tracking or Kalman filtering are existing technologies, they will not be described in detail here.

[0070] The preset threshold can be set to 0.5.

[0071] The aforementioned persistence score reflects the spatiotemporal stability of a target by analyzing its continued presence in multiple consecutive frames of images.

[0072] Secondly, obtain feature change scores.

[0073] Specifically, the process of obtaining the feature change score is as follows:

[0074] Obtain the bounding box of any target in the current image and its corresponding region in the previous frame; extract the feature vectors of the bounding box of any target in the current image and its corresponding region in the previous frame; calculate the feature change score based on the Euclidean distance between the two feature vectors.

[0075] Here, the feature vector represents the coordinates of the bounding box of the corresponding target. For example, the coordinates of the bounding box of the i-th target in the image at time t are... x and y represent the horizontal and vertical coordinates of the center point of the bounding box, respectively, and w and h represent the length and width of the bounding box, respectively.

[0076] Specifically, feature change scoring for:

[0077] ;

[0078] in, Let be the feature vector of the bounding box of the i-th target in the t-th frame of the image. Let be the feature vector of the region corresponding to the bounding box of the i-th target in the (t-1)-th frame image, and dist() is the function for calculating the Euclidean distance; The function to find the maximum value.

[0079] The above This is the maximum Euclidean distance between the feature vectors of each target in all frames of the grayscale image sequence.

[0080] A higher feature change score indicates that the target exhibits minor changes and is considered a stable target; while a lower feature change score indicates significant changes and may indicate a dynamic target. In other words, the feature change score measures the stability of a target's appearance features; a higher score indicates a stable appearance (potentially a static or slowly changing target), while a lower score indicates significant changes in appearance (potentially a dynamic target).

[0081] Then, the persistence score and feature change score are weighted and fused to obtain the correction coefficient.

[0082] In one embodiment, the correction coefficient can be obtained by weighted fusion, wherein the weight settings can be adjusted according to the persistence and the importance of feature changes.

[0083] In another embodiment, the correction factor can be the average of the feature change score and the persistence score.

[0084] The combination of persistence scoring and feature change scoring can effectively filter out spurious targets (such as noise or false detections) that appear only briefly or have unstable features. This allows for a better determination of whether a target is a real, persistent target or a spurious target caused by environmental changes or detection errors. This provides data support for reducing the false alarm rate (filtering out spurious targets that appear only briefly) and the false alarm rate (enhancing confidence in stable targets) in subsequent monitoring.

[0085] Step S22: Calculate the matching degree between any two images at the current time, sort the matching degrees in descending order, and select the image pairs corresponding to the first m matching degrees.

[0086] In addition, the images captured by different cameras in the target grid area at the current time need to be matched to determine whether the images captured by different cameras at the current time contain a large amount of duplicate information.

[0087] When performing matching, an image registration algorithm can be used to determine the similarity between the two images obtained at the current time as the matching degree.

[0088] In this embodiment, based on the matching degree, the top m image pairs with the highest matching degree are selected to avoid missing the current key information.

[0089] Step S23: Calculate the selection criteria for the two current time images in any image pair, and take the current time image corresponding to the larger of the two selection criteria for any image pair as the key image.

[0090] Among them, the selection criteria are positively correlated with the confidence level and the completeness factor of each target in the image at the current time.

[0091] The aforementioned target completeness factor takes into account situations where the target is occluded and its information is incomplete. For example, when monitoring whether workers have illegally entered a dangerous area (e.g., their hands are near dangerous machinery), more attention needs to be paid to the workers' hands. If the hands are occluded, the image may not provide useful monitoring information, meaning it cannot be used as a key image for subsequent monitoring analysis. Therefore, the target completeness factor is incorporated into the screening criteria for key images.

[0092] In one embodiment, the completeness factor is the ratio of the number of detected keypoints to the total number of keypoints. The more keypoints detected, the higher the completeness factor.

[0093] For example, the human body has 17 key points. When 15 key points are detected, the integrity of the target is relatively high.

[0094] In another embodiment, the integrity factor can also be the product of the ratio of the number of detected keypoints to the total number of keypoints and the penalty coefficient, wherein the penalty coefficient is the ratio of the number of detected valid keypoints to the total number of valid keypoints.

[0095] It is important to note that whether the detected keypoints include valid keypoints is also crucial. For example, when monitoring whether a worker has illegally entered a hazardous area (e.g., their hands are near dangerous machinery), there may be an interaction between the hand and body postures and the hazardous area; therefore, the visibility of hand and body postures is particularly important. In this case, the valid keypoints to prioritize are: left wrist, right wrist (for determining hand movements, such as whether a tool is being held), left elbow, right elbow (to assist in determining arm posture), left shoulder, and right shoulder (to determine body orientation). However, if the hand area is occluded, i.e., the hand keypoints are not detected, the integrity factor should be reduced, and the image corresponding to this target may not be a key image.

[0096] The above key points were detected by the human pose estimation model. Since the human pose estimation model is an existing technology, it will not be described in detail here.

[0097] The above-mentioned correction using a penalty coefficient allows for more accurate selection of views with higher integrity factors as the basis for analysis, ensuring that key parts are clearly visible. This provides data support for subsequent accurate judgment of whether workers have violated regulations (such as placing their hands near dangerous areas).

[0098] Among them, the screening criteria for:

[0099] ;

[0100] in, Let be the confidence score of the i-th target in the image at the k-th current time. Let I be the integrity factor of the i-th target in the image at the k-th current time, and let I be the total number of targets.

[0101] Step S3: Perform early warning analysis on the key images of each target grid area to output alarm information.

[0102] In this embodiment, each target in the key image of each target grid region is analyzed according to preset rules to determine whether it is abnormal.

[0103] The preset rules include at least the following: if at least one target appears in a dangerous area in the key image, an alarm message will be output.

[0104] Specifically, the criterion for determining if at least one target appears in a dangerous area in a key image is: the intersection-union ratio (IUU) of the target's bounding box with the dangerous area of ​​the corresponding grid region. If the IUU is greater than a threshold, an alarm message is output.

[0105] The threshold value mentioned above can be 0.4; of course, it can also be determined according to the actual situation. For example, the threshold can also be the average value of the intersection-union ratio of the worker's bounding box and the danger zone in multiple different historical images under the critical safety state.

[0106] In order to select key images from current-time images from multiple angles, the present invention first matches current-time images from different angles to find image pairs with a high degree of matching. Then, for each current-time image in the image pair, by using the confidence level of the target and the target integrity factor in the current-time images from each angle, the view that can see the whole picture of the target or key parts can be selected first, avoiding the problem of excessive views for alarm processing and improving alarm analysis efficiency.

[0107] This invention also provides a safety production monitoring system based on AI video analysis. For example... Figure 2 As shown, the system includes a processor and a memory, the memory storing computer program instructions, which, when executed by the processor, implement the safety production monitoring method based on AI video analysis according to the present invention.

[0108] The system also includes other components well known to those skilled in the art, such as communication buses and communication interfaces, the settings and functions of which are known in the art and therefore will not be described in detail here.

[0109] In this invention, the aforementioned memory can be any tangible medium containing or storing a program that can be used or combined with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc., or any other medium that can be used to store desired information and can be accessed by an application, module, or both. Any such computer storage medium can be part of a device or accessible to or connected to a device. Any application or module described in this invention can be implemented by computer-readable / executable instructions stored or otherwise maintained on such a computer-readable medium.

[0110] In the description of this specification, "multiple" means at least two, such as two, three or more, etc., unless otherwise expressly and specifically defined.

[0111] While various embodiments of the invention have been shown and described in this specification, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention.

Claims

1. A method for safety production monitoring based on AI video analysis, characterized in that, The method comprises: acquiring current time images of a plurality of collection devices in each target grid area; inputting each current time image into a target detection model to output a target detection result of each frame image; the target detection result comprises initial confidence of a plurality of targets; correcting the initial confidence to obtain confidence; the confidence is a product of a correction coefficient and the initial confidence; the correction coefficient is positively correlated with a persistence score and a feature change score of the corresponding target; the acquisition process of the persistence score is: acquiring a gray image sequence of a plurality of continuous frames comprising the current time image under any collection device, tracking the target by using a tracking algorithm, and taking a frame number ratio of the target as the corresponding persistence score; the feature change score is negatively correlated with a Euclidean distance of a feature vector of the target of the current time image and a corresponding target of a previous time image of the current time image; the feature vector is a coordinate of a bounding box of the target in the target detection result, and specifically comprises horizontal and vertical coordinates of a center point of the bounding box, length and width of the bounding box; calculating a matching degree of any two current time images, sorting the matching degrees in descending order, and selecting a front m matching degrees corresponding to image pairs, m≥1; respectively calculating a screening standard of the two current time images in any image pair, and taking a current time image corresponding to a larger one of the two screening standards of any image pair as a key image; the screening standard is positively correlated with confidence of each target in the corresponding current time image and a completeness factor of the target; the completeness factor represents completeness of the target in the bounding box; performing early warning analysis on the key images of each target grid area to output alarm information. 2.The safety production monitoring method based on AI video analysis according to claim 1, characterized in that, The screening standard is a sum of products of confidence of all targets in the corresponding current time image and the completeness factor. 3.The safety production monitoring method based on AI video analysis according to claim 1, characterized in that, The completeness factor is a product of a ratio of a number of detected key points to a total number of key points and a penalty coefficient; The penalty coefficient is a ratio of a number of detected effective key points to a total number of effective key points; The effective key point is a key point of the target having an interaction relationship with a dangerous area. 4.The safety production monitoring method based on AI video analysis according to claim 1, characterized in that, The matching degree is a similarity degree of any two current time images obtained by using an image registration algorithm. 5.The safety production monitoring method based on AI video analysis according to claim 1, characterized in that, The target grid area is: Before acquiring the current time images of the collection devices, the production area is divided into a plurality of grid areas to obtain collection devices in the grid areas, and each grid area is marked; a grid area marked as dangerous is recorded as a target grid area. 6.The safety production monitoring method based on AI video analysis according to claim 1, characterized in that, The target detection model is a YOLOv8-L model or a Faster R-CNN model.

7. The safety production monitoring method based on AI video analysis according to claim 1, characterized in that, The early warning analysis on the key images of each target grid area to output alarm information comprises: judging an intersection-over-union of a bounding box of the target in the target detection result and a dangerous area of the corresponding grid area; when the intersection-over-union is greater than a threshold value, the alarm information is output.

8. A safety production monitoring system based on AI video analysis, characterized in that, The system comprises: a processor; a memory storing computer instructions for AI video analysis-based safety production monitoring, when the computer instructions are run by the processor, the system executes the AI video analysis-based safety production monitoring method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Image processing method and device based on distributed monitoring system and terminal equipment

    CN118799814A

  • Early warning method and device, electronic equipment and storage medium

    CN119863903A