Industrial safety monitoring method and device based on target identification, equipment and storage medium

By extracting and preprocessing images frame by frame from real-time monitoring video streams, and combining deep learning algorithms with double-threshold nonmaximum suppression, the problem of low target recognition accuracy in complex industrial scenarios is solved, and high-precision industrial safety monitoring is achieved.

CN120853100APending Publication Date: 2025-10-28SHANDONG IND INTERNET DEV RES CENT
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510915747.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In complex industrial scenarios, existing target recognition methods suffer from low accuracy due to factors such as changes in lighting, target occlusion, and background complexity.

Method used

By extracting and preprocessing images frame by frame from real-time monitoring video streams, target recognition is performed using deep learning algorithms, and the detection box set is optimized through double threshold nonmaximum suppression to remove redundant detection boxes and generate labeled video frames and alarm signals.

Benefits of technology

It improves the accuracy and reliability of target recognition in complex industrial scenarios, reduces the false detection rate and missed detection rate, and enables real-time safety monitoring and timely early warning of industrial sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853100A_ABST
    Figure CN120853100A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial safety monitoring method and device based on target recognition, equipment and a storage medium, relates to the technical field of safety monitoring, and discloses an industrial safety monitoring method based on target recognition, which comprises the following steps: extracting a frame-by-frame image based on a real-time monitoring video stream, and preprocessing the frame-by-frame image to obtain a preprocessed image; performing target identification on the preprocessed image to obtain coordinate information and confidence information of an initial detection frame set; performing dual-threshold non-maximum suppression processing on the initial detection frame set based on the coordinate information and the confidence information to obtain an optimized detection frame set; and obtaining a labeled video frame and an alarm signal based on the optimized detection frame set so as to monitor industrial safety. And the target identification precision in the security monitoring process in a complex industrial scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of security monitoring technology, and in particular to industrial security monitoring methods, devices, equipment and storage media based on target recognition. Background Technology

[0002] In the process of safety supervision in industrial production, it is necessary to monitor the complex industrial site environment in real time and identify various potential safety hazards, including personnel violations, abnormal equipment conditions, and hazardous factors in the environment, in order to ensure the improvement of industrial production safety and reduce the accident rate.

[0003] Currently, target recognition methods in industrial safety monitoring are mainly based on convolutional neural network technology in deep learning. These methods train on large amounts of labeled data to learn the feature representations of targets, thereby enabling the detection and recognition of different targets in monitoring videos. However, due to factors such as varying lighting conditions, target occlusion, diverse target shapes, and complex backgrounds in complex industrial environments, the accuracy of target recognition is relatively low. Improving the accuracy of target recognition in safety monitoring processes under complex industrial scenarios remains an unresolved issue.

[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide an industrial safety monitoring method, device, equipment and storage medium based on target recognition, aiming to solve the technical problem of how to improve the target recognition accuracy in the safety monitoring process in complex industrial scenarios.

[0006] To achieve the above objectives, this application proposes an industrial safety monitoring method based on target recognition, the method comprising:

[0007] Frame-by-frame images are extracted from real-time monitoring video streams, and the frame-by-frame images are preprocessed to obtain preprocessed images;

[0008] The preprocessed image is subjected to target recognition to obtain the coordinate information and confidence information of the initial detection box set;

[0009] Based on the coordinate information and the confidence information, the initial detection box set is subjected to double threshold nonmaximum suppression processing to obtain an optimized detection box set.

[0010] Based on the optimized detection box set, labeled video frames and alarm signals are obtained to monitor industrial safety.

[0011] In one embodiment, the step of performing double-threshold non-maximum suppression processing on the initial detection box set based on the coordinate information and the confidence information to obtain an optimized detection box set includes:

[0012] Obtain the detection boxes to be processed from the initial set of detection boxes;

[0013] Based on the confidence information, a reference detection box is obtained from the initial detection box set;

[0014] Based on the coordinate information, double threshold non-maximum suppression processing is performed on the detection box to be processed and the reference detection box to obtain the non-maximum suppression result of the detection box to be processed.

[0015] The optimized set of detection boxes is determined based on the nonmaximum suppression results of the detection boxes to be processed.

[0016] In one embodiment, the step of performing dual-threshold non-maximum suppression processing on the detection box to be processed and the reference detection box based on the coordinate information to obtain the non-maximum suppression result of the detection box to be processed includes:

[0017] Obtain a first cross-union ratio (CUP) threshold and a second CUP threshold, wherein the first CUP threshold is greater than the second CUP threshold;

[0018] The intersection-union ratio (IUU) of the detection box to be processed and the reference detection box is calculated based on the coordinate information to obtain the target IUU.

[0019] The non-maximum suppression result of the detection box to be processed is determined based on the first cross-union ratio threshold, the second cross-union ratio threshold, and the target cross-union ratio.

[0020] In one embodiment, the step of determining the non-maximum suppression result of the detection box to be processed based on the first cross-union ratio (CUP) threshold, the second CUP threshold, and the target CUP includes:

[0021] When the target cross-union ratio is greater than the first cross-union ratio threshold, the non-maximum suppression result of the detection box to be processed is determined as the confidence level of suppressing the detection box to be processed;

[0022] When the target cross-union ratio is less than or equal to the first cross-union ratio threshold and greater than or equal to the second cross-union ratio threshold, the non-maximum suppression result of the detection box to be processed is determined to be a reduction in the confidence of the detection box to be processed according to a linear decay function;

[0023] When the target cross-union ratio is less than the first cross-union ratio threshold, the non-maximum suppression result of the detection box to be processed is determined as the confidence level for retaining the detection box to be processed.

[0024] In one embodiment, the step of extracting frame-by-frame images based on a real-time monitoring video stream includes:

[0025] Extract video frame data from the real-time monitoring video stream based on the start code and end identifier;

[0026] Add frame sequence number, timestamp, monitoring device number and metadata to the video frame data to obtain message data;

[0027] The message data is converted to a video format according to a preset encoding format, and the converted message data is compressed to obtain the transmitted video stream data.

[0028] Frame-by-frame images are extracted from the transmitted video stream data according to a preset frame rate.

[0029] In one embodiment, the step of preprocessing the frame-by-frame images to obtain a preprocessed image includes:

[0030] Calculate the global brightness mean of the frame-by-frame image, and perform linear brightness adjustment processing on the frame-by-frame image based on the brightness mean to obtain a brightness-normalized image;

[0031] The brightness-normalized image is subjected to histogram equalization to obtain a contrast-enhanced image;

[0032] The contrast-enhanced image is subjected to white balance correction to obtain a color-corrected image;

[0033] The color-corrected image is resized and edge-filled to obtain a pre-processed image.

[0034] In one embodiment, the step of obtaining labeled video frames and alarm signals based on the optimized detection box set to monitor industrial safety includes:

[0035] Obtain the coordinate information, confidence information, and target category information of each optimized detection box in the optimized detection box set, and obtain the color mapping relationship table;

[0036] The target color is determined from the color mapping table based on the target category information;

[0037] Based on the coordinate information, the confidence information, and the target color, a labeled video frame is generated, and a corresponding alarm signal is generated based on the labeled video frame to monitor industrial safety.

[0038] Furthermore, to achieve the above objectives, this application also proposes an industrial safety monitoring device based on target recognition, the industrial safety monitoring device based on target recognition comprising:

[0039] The data processing module is used to extract frame-by-frame images based on real-time monitoring video streams and preprocess the frame-by-frame images to obtain preprocessed images.

[0040] The target recognition module is used to perform target recognition on the preprocessed image to obtain the coordinate information and confidence information of the initial detection box set;

[0041] The detection optimization module is used to perform double threshold non-maximum suppression processing on the initial detection box set based on the coordinate information and the confidence information to obtain an optimized detection box set.

[0042] The safety monitoring module is used to obtain labeled video frames and alarm signals based on the optimized detection frame set in order to monitor industrial safety.

[0043] Furthermore, to achieve the above objectives, this application also proposes an industrial safety monitoring device based on target recognition, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the industrial safety monitoring method based on target recognition as described above.

[0044] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the industrial safety monitoring method based on target recognition as described above.

[0045] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the target recognition-based industrial safety monitoring method described above.

[0046] One or more technical solutions proposed in this application have at least the following technical effects:

[0047] By extracting and preprocessing images frame by frame from real-time monitoring video streams, image quality is improved, providing a clear and accurate image foundation for subsequent target recognition. Deep learning algorithms are then used to perform target recognition on the preprocessed images, enabling accurate detection of various targets in industrial scenarios. Furthermore, dual-threshold non-maximum suppression is applied to optimize the initial detection box set, removing redundant detection boxes and reducing false positive and false negative rates. Finally, labeled video frames and alarm signals are generated based on the optimized detection box set, achieving real-time safety monitoring of industrial sites, timely detection and early warning of safety risks, and improving monitoring efficiency and industrial production safety. This application's solution integrates image preprocessing, target recognition, and dual-threshold non-maximum suppression, effectively improving the accuracy and reliability of target recognition during safety monitoring in complex industrial scenarios. Attached Figure Description

[0048] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0049] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is a flowchart illustrating an embodiment of the industrial safety monitoring method based on target recognition provided in this application.

[0051] Figure 2 This is a flowchart illustrating Embodiment 2 of the industrial safety monitoring method based on target recognition in this application;

[0052] Figure 3 This is a diagram of the intelligent safety monitoring system architecture provided in Embodiment 2 of the industrial safety monitoring method based on target recognition in this application;

[0053] Figure 4 This is a schematic diagram of the main monitoring interface of the intelligent safety monitoring system provided in Embodiment 2 of the industrial safety monitoring method based on target recognition in this application;

[0054] Figure 5 This is a schematic diagram of the monitoring center interface of the intelligent safety monitoring system provided in Embodiment 2 of the industrial safety monitoring method based on target recognition in this application;

[0055] Figure 6 This is a schematic diagram of the user management interface of the intelligent safety monitoring system provided in Embodiment 2 of the industrial safety monitoring method based on target recognition in this application;

[0056] Figure 7 A simplified flowchart illustrating the target recognition-based industrial safety monitoring method provided in Embodiment 2 of this application;

[0057] Figure 8 This is a schematic diagram of the module structure of an industrial safety monitoring device based on target recognition, as described in an embodiment of this application.

[0058] Figure 9 This is a schematic diagram of the equipment structure of the hardware operating environment involved in the industrial safety monitoring method based on target recognition in the embodiments of this application.

[0059] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0060] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0061] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0062] The main solution of this application embodiment is as follows: extract frame-by-frame images based on real-time monitoring video streams, and preprocess the frame-by-frame images to obtain preprocessed images; perform target recognition on the preprocessed images to obtain coordinate information and confidence information of an initial detection box set; perform double-threshold non-maximum suppression processing on the initial detection box set based on the coordinate information and the confidence information to obtain an optimized detection box set; and obtain labeled video frames and alarm signals based on the optimized detection box set to monitor industrial safety.

[0063] In this embodiment, for ease of description, the following description uses the intelligent safety monitoring system as the executing entity.

[0064] Currently, target recognition methods in industrial safety monitoring are mainly based on convolutional neural network technology in deep learning. These methods train on large amounts of labeled data to learn the feature representations of targets, thereby enabling the detection and recognition of different targets in surveillance videos. However, due to factors such as varying lighting conditions, target occlusion, diverse target shapes, and complex backgrounds in complex industrial environments, the accuracy of target recognition is relatively low.

[0065] This application provides a solution that improves image quality by extracting and preprocessing images frame by frame from real-time monitoring video streams, providing a clear and accurate image foundation for subsequent target recognition. Deep learning algorithms are then used to perform target recognition on the preprocessed images, enabling accurate detection of various targets in industrial scenarios. Furthermore, dual-threshold non-maximum suppression is applied to optimize the initial detection box set, removing redundant detection boxes and reducing false positive and false negative rates. Finally, labeled video frames and alarm signals are generated based on the optimized detection box set, achieving real-time safety monitoring of industrial sites, timely detection and early warning of safety risks, and improving monitoring efficiency and industrial production safety. This application's solution integrates image preprocessing, target recognition, and dual-threshold non-maximum suppression, effectively improving the accuracy and reliability of target recognition during safety monitoring in complex industrial scenarios.

[0066] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or intelligent safety monitoring system capable of performing the above functions. The following description uses an intelligent safety monitoring system as an example to illustrate this embodiment and the subsequent embodiments.

[0067] Based on this, embodiments of this application provide an industrial safety monitoring method based on target recognition, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the industrial safety monitoring method based on target recognition in this application.

[0068] In this embodiment, the industrial safety monitoring method based on target recognition includes steps S10 to S40:

[0069] Step S10: Extract frame-by-frame images from the real-time monitoring video stream and preprocess the frame-by-frame images to obtain preprocessed images;

[0070] It should be noted that the solution in this application is applied to an intelligent safety monitoring system based on target recognition. The intelligent safety monitoring system includes a data acquisition module, a central processing module, and an application service module. The intelligent safety monitoring system acquires real-time monitoring video images from industrial sites and uses a pre-trained target detection model to identify and locate abnormal targets. It can monitor and issue early warnings for abnormal targets, providing data for backend analysis or user viewing.

[0071] Specifically, in the intelligent safety monitoring system, the data acquisition module allows the system to acquire and format real-time video at the industrial site, the central processing module allows the algorithm model to identify and predict video frames, and the application service module can process, display, and issue warnings on the model's calculation results.

[0072] It should be understood that the architecture of the intelligent safety monitoring system is divided into an application service layer, a data storage layer, a model layer, and a perception and acquisition layer. The data acquisition module is deployed in the perception and acquisition layer. In the perception and acquisition layer, multiple industrial cameras or other monitoring devices are used to collect monitoring video data of industrial scenes. Based on the device information of the industrial cameras, the monitoring video data is stamped with timestamps, location markers, and other information to obtain real-time monitoring video streams.

[0073] Specifically, real-time monitoring video streams can be obtained through transport layer protocols, a reliable connection can be established through TCP / IP protocols, and the video stream data can be transmitted in blocks, numbered, acknowledged, and routed and addressed.

[0074] It should be noted that video decoding can extract individual images from a real-time monitoring video stream based on information such as encoding format, frame rate, and resolution, resulting in frame-by-frame images. Performing a series of image preprocessing operations on each frame (such as brightness normalization, contrast enhancement, color correction, size scaling, and edge filling) yields a preprocessed image.

[0075] In one feasible implementation, step S10, which involves extracting frame-by-frame images based on a real-time monitoring video stream, may include steps S111 to S114:

[0076] Step S111: Extract video frame data from the real-time monitoring video stream based on the start code and end identifier;

[0077] It's important to note that start codes and end markers are special byte sequences defined in the video encoding format, used to mark the start and end positions of a video frame. In real-time surveillance video streams, by scanning the video stream data and finding the byte sequences corresponding to the start codes and end markers, the complete video frame data can be accurately extracted. The start codes and end markers are inserted by the video encoder when generating the video stream to help the decoder correctly parse the video content. Video frame data is the raw data of a single video frame extracted from the real-time surveillance video stream, containing all the pixel information that constitutes the video frame, as well as related encoding information.

[0078] Step S112: Add frame sequence number, timestamp, monitoring device number and metadata to the video frame data to obtain message data;

[0079] It's important to note that the frame sequence number is a unique number assigned to each video frame after the video frame data is extracted, according to the extraction order. Starting from 1, the frame sequence number increments by 1 for each extracted video frame. The frame sequence number is used to identify the order of video frames during subsequent transmission and processing, ensuring that video frames are not out of order during transmission and processing, and facilitating the reconstruction of the video stream in the correct order at the receiving end.

[0080] In addition, the timestamp is precise time information generated by the intelligent safety monitoring system, indicating the moment when the video frame was extracted. It is used to record the acquisition time of the video frame, which is convenient for determining the time point of the event in subsequent monitoring and analysis. It also helps to synchronize and recover data in the event of video stream interruption or packet loss.

[0081] Additionally, the monitoring device number is used to uniquely identify the industrial camera or other monitoring device that captured the video frame. Each monitoring device has a pre-configured unique number in the intelligent security monitoring system. By adding monitoring device numbers, video frames captured by different devices can be distinguished in multi-device monitoring scenarios, facilitating the management and analysis of monitoring videos from different areas or devices.

[0082] Additionally, metadata is auxiliary information used to describe the characteristics of video frames, including the video encoding format, resolution, frame rate, etc., to help the receiving end correctly decode and display video frames.

[0083] It should be understood that encapsulating information such as video frame data, frame sequence number, timestamp, monitoring device number, and metadata yields message data. During transmission, message data serves as the basic unit of transmission, ensuring that the receiving end can correctly parse and reconstruct the video frames.

[0084] Step S113: Convert the message data into a video format according to a preset encoding format, and compress the converted message data to obtain the transmitted video stream data.

[0085] It should be noted that the preset encoding format is a predefined video encoding format used by the intelligent safety monitoring system, such as H.264 or H.265. When converting the message data to a video format, the original video frame data needs to be converted to message data in the preset encoding format. During the conversion process, the video frame data is re-encoded to conform to the target encoding format specifications. The message data, now converted to the preset encoding format, is then compressed to obtain the transmitted video stream data, which reduces the data volume and improves transmission efficiency.

[0086] Step S114: Extract frame-by-frame images from the transmitted video stream data according to a preset frame rate.

[0087] It should be noted that the preset frame rate is the number of video frames extracted per second from the transmitted video stream data by the intelligent security monitoring system, measured in frames per second. The preset frame rate can be determined based on actual monitoring needs and the performance of the intelligent security monitoring system. Extracting individual image frames from the transmitted video stream data according to the preset frame rate yields still images within a specified time interval, i.e., frame-by-frame images.

[0088] In one feasible implementation, the step of preprocessing the frame-by-frame image in step S10 to obtain a preprocessed image may include steps S121 to S124:

[0089] Step S121: Calculate the global brightness mean of the frame-by-frame image, and perform linear brightness adjustment processing on the frame-by-frame image based on the brightness mean to obtain a brightness-normalized image;

[0090] It's important to note that the global average brightness can be obtained by averaging the brightness values ​​of all pixels in each frame of the image. Brightness values ​​are typically represented by grayscale values. For color frame-by-frame images, they can be first converted to grayscale, then the grayscale value of each pixel can be calculated, and the average of all grayscale values ​​can be obtained. Specifically, the formula is: Global Average Brightness = (Sum of Brightness Values ​​of All Pixels) / (Total Number of Pixels). The global average brightness reflects the overall brightness level of the image and is an important basis for brightness adjustments.

[0091] It should be understood that during linear brightness adjustment, a target average brightness value is set, and the difference between the current global average brightness value and the target average brightness value is calculated. Based on this difference, the brightness value of each pixel in the image is linearly adjusted. Specifically, the formula for linear adjustment is: Adjusted pixel brightness value = Original pixel brightness value + (Target average brightness value - Global average brightness value). After linear brightness adjustment, the average brightness value of the brightness-normalized image is adjusted to within the preset target average brightness value range. Linear brightness adjustment effectively eliminates differences in image brightness under different lighting conditions, ensuring that images at different times and in different environments have relatively consistent brightness levels, reducing false positives and false negatives caused by brightness variations.

[0092] Step S122: Perform histogram equalization on the brightness-normalized image to obtain a contrast-enhanced image;

[0093] It's important to note that histogram equalization is used to enhance image contrast. By transforming the gray-level histogram of a brightness-normalized image, the gray-level value distribution is mapped to a wider range, thus improving image contrast and resulting in a contrast-enhanced image. Specifically, during histogram equalization, the gray-level histogram of the brightness-normalized image is calculated, and the gray-level values ​​are redistributed according to the cumulative distribution function, resulting in a more uniform gray-level distribution in the transformed contrast-enhanced image. Histogram equalization can further highlight details in the image, making the difference between the target and the background more obvious, and providing clearer image features for subsequent target recognition.

[0094] Step S123: Perform white balance correction processing on the contrast-enhanced image to obtain a color-corrected image;

[0095] It's important to note that white balance correction adjusts the color balance of an image to eliminate color casts, resulting in more realistic and natural colors. During white balance correction, the gain coefficients of each color channel in the contrast-enhanced image are calculated. Based on these gain coefficients, the pixel values ​​are adjusted to ensure that white objects appear true white, resulting in a color-corrected image. White balance correction effectively removes color casts caused by different lighting conditions, improves color fidelity, and provides more accurate color information for subsequent target recognition.

[0096] Step S124: The color-corrected image is scaled and edge-filled to obtain a pre-processed image.

[0097] It should be noted that resizing is used to adjust the size of the color-corrected image to fit the input of the object detection model. Edge padding is used during image resizing to fill the edge regions of the image in order to maintain the aspect ratio and prevent distortion of target information. Since different object detection models have different requirements for the size of the input image, methods such as bilinear interpolation or nearest neighbor interpolation are needed to scale the color-corrected image. Based on the scaling ratio between the target size and the original size, the image pixels are resampled and interpolated. Simultaneously, specific pixel values ​​(such as grayscale value 128) are used to fill the image edges, ensuring that the scaled preprocessed image meets the target size requirements while avoiding target distortion and information loss caused by direct cropping or stretching.

[0098] Step S20: Perform target recognition on the preprocessed image to obtain the coordinate information and confidence information of the initial detection box set;

[0099] It should be noted that pre-trained deep learning-based target detection models, such as YOLOv5, can extract video frame features from pre-processed images to determine whether industrial safety risk targets exist in the pre-processed images. Industrial safety risk targets include personal protective equipment (PPE) targets (such as safety helmets and masks), personnel behavior risk targets (such as sleeping on duty, smoking, and using mobile phones), and environmental risk targets (such as smoke and stagnant water). After identifying industrial safety risk targets, bounding boxes corresponding to the targets are generated on the pre-processed images, and the confidence scores of the bounding boxes are obtained. By setting an appropriate confidence threshold, bounding boxes with confidence scores below the threshold and those that may be misclassified can be removed, resulting in an initial set of detection boxes. The position (i.e., coordinate information) of the target in the pre-processed image and the reliability of the algorithm's target recognition results are also obtained, yielding the coordinate information and confidence information of the initial set of detection boxes.

[0100] Step S30: Based on the coordinate information and the confidence information, perform double threshold nonmaximum suppression processing on the initial detection box set to obtain an optimized detection box set;

[0101] It should be noted that Dual-Threshold Non-Maximum Suppression (DNMS) is an improved post-processing technique for target detection, designed to improve upon traditional non-maximum suppression algorithms. By setting two intersection-over-union (IoU) thresholds, it achieves a dynamic confidence decay range (e.g., 0.4-0.7). Through linear interpolation, it preserves potential targets while suppressing redundant detection boxes, thus addressing the issue of missed or false detections in dense target scenes. Based on coordinate and confidence information, DNMS can be applied to the initial detection box set to obtain an optimized set. Compared to the initial set, the optimized set contains fewer detection boxes, and each box has a higher confidence level and more accurate location.

[0102] In one feasible implementation, step S30 may include steps S31 to S34:

[0103] Step S31: Obtain a reference detection box from the initial detection box set based on the confidence information;

[0104] It should be noted that the reference detection box is the detection box with the highest confidence selected from the initial detection box set based on the confidence information during the double threshold nonmaximum suppression process. It is used to compare with other detection boxes to be processed and to calculate the intersection over union (IoU) between the detection boxes to be processed and the reference detection box.

[0105] Step S32: Obtain the detection box to be processed from the initial detection box set;

[0106] It should be understood that the detection boxes to be processed are those selected from the initial set of detection boxes that have not yet undergone double threshold nonmaximum suppression.

[0107] Step S33: Based on the coordinate information, perform double threshold non-maximum suppression processing on the detection box to be processed and the reference detection box to obtain the non-maximum suppression result of the detection box to be processed;

[0108] It should be noted that, based on coordinate information, dual-threshold non-maximum suppression (NMS) can be performed on both the target detection box and the reference detection box to determine the suppression result for the target detection box, thus obtaining the NMS suppression result for the target detection box. The NMS suppression result for the target detection box includes suppressing the confidence level of the target detection box, reducing the confidence level of the target detection box according to a linear decay function, and retaining the confidence level of the target detection box.

[0109] In one feasible implementation, step S33 may include steps S331 to S333:

[0110] Step S331: Obtain a first cross-union ratio (CUP) threshold and a second CUP threshold, wherein the first CUP threshold is greater than the second CUP threshold;

[0111] It should be noted that the first cross-union ratio (CUP) threshold is a high CUP threshold set for the double-threshold nonmaximum suppression (NMS) process. The second CUP threshold is a low CUP threshold set for the double-threshold NMS process.

[0112] Step S332: Calculate the intersection-union ratio (IUU) of the detection box to be processed and the reference detection box based on the coordinate information to obtain the target IUU.

[0113] It should be noted that the coordinate information includes the coordinates of the top-left and bottom-right corners of the detection box, which can be defined using (x1, y1, x2, y2), where (x1, y1) is the top-left corner coordinate and (x2, y2) is the bottom-right corner coordinate. Based on the detection box B to be processed... i The coordinates of the intersection region can be calculated using the coordinates of the reference bounding box M and the coordinates of the reference bounding box M. The coordinates of the top-left corner of the intersection region are the larger of the top-left corner coordinates of the two bounding boxes. The lower right corner coordinate of the intersection region is the smaller value of the lower right corner coordinates of the two bounding boxes, i.e. If the coordinates of the top-left corner of the intersecting region are greater than or equal to the coordinates of the bottom-right corner (i.e., there is no intersection), then the area of ​​the intersecting region is 0. Otherwise, the width of the intersecting region is... Height is The area of ​​the intersection region is the width multiplied by the height. The area of ​​the union region can be obtained by subtracting the area of ​​the intersection region from the sum of the areas of the two boxes, i.e., Area(M) + Area(B). i )-Area(M∩B i The cross-union ratio (IoU) between the target detection box and the reference detection box can be calculated using the cross-union formula, thus obtaining the target IoU(M,B). i The formula for calculating the intersection-union ratio is as follows:

[0114]

[0115] In the formula, Area(M∩B) i Area(M∪B) is the area of ​​the union region of the detection box to be processed and the reference detection box. i ) represents the area of ​​the intersection region between the detection box to be processed and the reference detection box.

[0116] Step S333: Determine the non-maximum suppression result of the detection box to be processed based on the first cross-union ratio threshold, the second cross-union ratio threshold, and the target cross-union ratio.

[0117] It should be understood that by comparing the target cross-union ratio with the first cross-union ratio threshold and the second cross-union ratio threshold, the suppression result of the detection box to be processed can be determined, and the non-maximum suppression result of the detection box to be processed can be obtained.

[0118] In one feasible implementation, step S333 may include: when the target cross-union ratio (CUN) is greater than the first CUN threshold, determining that the non-maximum suppression result of the detection box to be processed is to suppress the confidence of the detection box to be processed; when the target CUN is less than or equal to the first CUN threshold and greater than or equal to the second CUN threshold, determining that the non-maximum suppression result of the detection box to be processed is to reduce the confidence of the detection box to be processed according to a linear decay function; and when the target CUN is less than the first CUN threshold, determining that the non-maximum suppression result of the detection box to be processed is to retain the confidence of the detection box to be processed.

[0119] It should be noted that when the target crossover ratio is greater than the high crossover ratio threshold N, i When the first intersection-union threshold is reached, it indicates that the detection box B to be processed is... i The overlap with the reference detection box M is very high, at which point the detection box B to be processed is considered to be... i It is highly likely that the detection box is redundant and overlaps with M. It is necessary to directly suppress the detection box to be processed, that is, delete the detection box to be processed, so as to remove redundant detection results and retain the detection box with the highest confidence in each local region, thereby improving the accuracy and simplicity of the detection results.

[0120] Additionally, when the target crossover ratio is less than or equal to the high crossover ratio threshold N... i That is, the first cross-union ratio threshold, and greater than or equal to the lower cross-union ratio threshold N. t When the second intersection-union threshold is reached, it indicates that the detection box B to be processed is... i The overlap with the reference detection box M is at a moderate level, between the high and low intersection-union (IU) thresholds. In this case, B cannot be completely determined. i Whether it is an independent target or a target that overlaps with M, the confidence of the detection box to be processed needs to be reduced by using a linear decay function to punish its potential redundancy, reduce the possibility of false detection, and at the same time retain a certain level of confidence to avoid directly deleting potentially correct detection boxes.

[0121] Additionally, when the target cross-union ratio is less than the low cross-union ratio threshold N... i When the first intersection-union threshold is reached, it indicates that the detection box B to be processed is... i The overlap with the reference detection box M is low, and the detection box B to be processed...i It is likely that the detected target is an independent target, rather than a target that repeats with M. In this case, the original confidence of the detection box to be processed should be preserved and no suppression should be performed to avoid missed detection.

[0122] For example, the formula for dual-threshold nonmaximum suppression is as follows:

[0123]

[0124] Where s i It represents the confidence score of the detection boxes to be processed, M is the current highest-scoring box, i.e., the reference detection box, and B is the confidence score of the detection boxes to be processed. i This is the box to be processed. IoU(M,B) i ) represents the target intersection-union ratio (CUI) between the detection box to be adjusted and the reference detection box M. i N is the first crossover ratio threshold. t The second crossover ratio threshold is N, where N is the crossover ratio threshold. i <N t The larger the IoU, the more s i The greater the decay, the more s i (1-IoU(M,B i )) is a linear decay function.

[0125] In this implementation, the confidence of the detection box is adjusted by a decay function based on the IoU value between the detection box and the detection box with the highest score, instead of deleting it directly. This avoids losing potentially correct detection boxes due to hard threshold deletion in scenes with dense targets.

[0126] Step S34: Determine the optimized detection box set based on the non-maximum suppression result of the detection box to be processed.

[0127] It should be noted that after obtaining the non-maximum suppression results of each unprocessed detection box in the initial detection box set, the corresponding unprocessed detection boxes will be retained, have their confidence adjusted, or be deleted based on the non-maximum suppression results to obtain optimized detection boxes, and then the optimized detection box set will be obtained.

[0128] It should be understood that each optimized detection box in the optimized detection box set is either directly retained because its cross-union ratio with the reference detection box is less than the second cross-union ratio threshold, or it is retained because it still has a high confidence after confidence adjustment when its cross-union ratio is between the first cross-union ratio threshold and the second cross-union ratio threshold.

[0129] Step S40: Based on the optimized detection box set, labeled video frames and alarm signals are obtained to monitor industrial safety.

[0130] It should be noted that, based on the optimized detection box set, the coordinate information, confidence information, and target category information of the industrial safety risk targets corresponding to each optimized detection box can be obtained. Based on the coordinate information, confidence information, and target category information, the frame-by-frame images corresponding to the preprocessed image can be annotated to obtain annotated video frames, and corresponding alarm signals can be generated to remind users to view and monitor industrial safety.

[0131] This embodiment improves image quality by extracting and preprocessing images frame by frame from real-time monitoring video streams, providing a clear and accurate image foundation for subsequent target recognition. Deep learning algorithms are then used to perform target recognition on the preprocessed images, accurately detecting various targets in industrial scenarios. Furthermore, dual-threshold non-maximum suppression is applied to optimize the initial detection box set, removing redundant detection boxes and reducing false positive and false negative rates. Finally, labeled video frames and alarm signals are generated based on the optimized detection box set, enabling real-time safety monitoring of industrial sites, timely detection and early warning of safety risks, and improved monitoring efficiency and industrial production safety. This application's solution integrates image preprocessing, target recognition, and dual-threshold non-maximum suppression, effectively improving the accuracy and reliability of target recognition during safety monitoring in complex industrial scenarios.

[0132] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 Step S40 may include steps S41 to S43:

[0133] Step S41: Obtain the coordinate information, confidence information, and target category information of each optimized detection box in the optimized detection box set, and obtain the color mapping relationship table;

[0134] It should be understood that the optimized detection box set is the set of detection boxes obtained after double threshold nonmaximum suppression processing, containing the coordinate information, confidence information, and target category information of each optimized detection box. Specifically, the coordinate information defines the target position of the industrial safety risk target in the optimized detection box within each frame of the image; the confidence information indicates the reliability of the detection result for the industrial safety risk target in the optimized detection box; and the target category information indicates the specific category of the industrial safety risk target corresponding to the optimized detection box, such as safety helmet, mask, sleeping on duty, smoking, using a mobile phone, and smoke.

[0135] It should be noted that the color mapping table stores the mapping relationship between target categories and target colors. Based on the target category information, the color corresponding to the industrial safety risk target represented by each optimized detection box can be determined from the color mapping table.

[0136] Step S42: Determine the target color from the color mapping table based on the target category information;

[0137] It should be understood that the target color is determined from a color mapping table based on the target category information. Different categories correspond to different colors, making it possible to intuitively distinguish different target categories in the labeled video frames.

[0138] Step S43: Generate labeled video frames based on the coordinate information, the confidence information, and the target color, and generate corresponding alarm signals based on the labeled video frames to monitor industrial safety.

[0139] It should be noted that after obtaining the coordinate information, confidence information, and target color of each industrial safety risk target identified from the preprocessed image, a bounding box is drawn on the frame-by-frame image corresponding to the preprocessed image based on the coordinate information, the bounding box is filled with the target color, and a text label containing the target category and confidence information is drawn near the bounding box to obtain the labeled video frame timestamp and monitoring device number.

[0140] Specifically, when generating text labels containing target category and confidence information, the display position of the labeled text on the video frame is calculated based on the coordinates of the upper left corner of the bounding box area. The display position is set above and to the left of the upper left corner of the bounding box to avoid the text overlapping with the bounding box and to ensure that the labeled information is clear and readable.

[0141] In this implementation, the video target detection results are visualized by marking detected targets with bounding boxes of different colors, and annotating the boxes with the target category and confidence score. Color mapping and font adjustment techniques ensure that the annotation information is clear and readable, allowing users to quickly obtain detection results even in complex scenes.

[0142] It should be understood that after obtaining the labeled video frames, the timestamp and monitoring device number of the labeled video frames are acquired. Based on the timestamp, monitoring device number, target category information, and labeled video frames, alarm information is generated and pushed to the main monitoring interface of the intelligent safety monitoring system. Alarm signals are issued via sound alarm or flashing lights to remind users to check the specific warning situation for monitoring industrial safety. Specifically, different risk levels can be set for different target category information, with different alarm signals corresponding to different risk levels.

[0143] For example, if the timestamp of the labeled video frame that identifies an industrial safety risk target is "February 19, 10:09", the device name labeled by the monitoring device number is "factory camera 001", the target category information is "sleeping on duty", "playing on mobile phone", and "absent from duty", and the risk level of the target category information is "general risk", then the alarm information pushed to the monitoring main interface of the intelligent safety supervision system can be "February 19, 10:09; automatic alarm; factory camera 010; sleeping on duty, playing on mobile phone, abstaining from duty; general risk".

[0144] For example, please refer to Figure 3 , Figure 3 This is a diagram of the intelligent safety monitoring system architecture provided in Embodiment 2 of the industrial safety monitoring method based on target recognition in this application. Figure 3 As shown, the architecture of the intelligent safety monitoring system is divided into an application service layer, a data storage layer, a model layer, and a perception and acquisition layer. The perception and acquisition layer collects video streams through industrial cameras. These video streams can contain equipment information, video encoding, image resolution, event type, and frame rate. The collected video streams cover scenarios such as production processes, personnel management, environmental safety, perimeter protection, and remote monitoring. The data storage layer performs persistent storage, retrieval, and backup / recovery of the video streams from the perception and acquisition layer through backup databases and cloud databases. It also performs single-frame image format conversion and normalization processing on the video streams. The cloud database uses object storage. The single-frame images processed in the model layer then pass through the head input, baseline network, neck network, and head output of the algorithm model to achieve target recognition, identifying various targets such as flames, water accumulation, safety helmets, people sleeping on duty, and smoking. The application service layer integrates functional modules such as early warning and alert, algorithm management, permission management and data dashboard. It can provide corresponding functions to users in the form of a client, and divide the permission levels according to different users. It can receive user operation instructions and realize the advanced functions and user interaction of the intelligent safety monitoring system.

[0145] For example, the intelligent safety monitoring system can utilize object storage for data storage, employing Java and JavaScript as the primary programming languages. It uses a front-end / back-end separation model, leveraging Vue and Spring Cloud as frameworks for full-platform development, and highly encapsulates the back-end infrastructure components. Simultaneously, it integrates Sentinel for multi-dimensional protection, including traffic control, circuit breaking and degradation, and system load balancing. The registry and configuration center utilize Nacos to enhance inter-module collaboration. The authentication center is built on Spring Security, creating a multi-terminal authentication system to achieve mutual isolation of token permissions between services. After logging into the intelligent safety monitoring system, enterprise users can monitor connected cameras and various production processes in the factory, receiving real-time alerts. The main monitoring interface is shown below. Figure 3 As shown. Users can also select the camera to view live video, video playback, and alarm images.

[0146] For example, please refer to Figure 4 , Figure 4 This is a schematic diagram of the main interface of the intelligent safety monitoring system provided in Embodiment 2 of the industrial safety monitoring method based on target recognition of this application. Figure 4 As shown, the main monitoring interface displays multiple information modules, including device overview, device alarms, abnormal devices, real-time monitoring video, connected device details, alarm list, and algorithm alarms. The device overview displays the overall device status in a circular dashboard, such as the number of connected devices, the number of alarmed devices, the number of abnormal devices, the alarm resolution rate, and the alarm security index. Device alarms are displayed in a pie chart, showing the number of devices without alarms and the percentage of alarmed devices. Abnormal devices are displayed in a pie chart, showing the number of normal and abnormal devices and their percentages. Connected device details are displayed in a table, showing detailed information about connected devices, including serial number, device name, installation location, device status (e.g., alarm, normal), and risk level (e.g., general risk, low risk). Real-time monitoring can display multiple real-time monitoring screens, each showing a different monitoring area. Alarm details are displayed in a line graph, showing alarm trends over time. The alarm list can record specific alarm events and push alarm information such as "February 19, 10:09; Automatic alarm; Factory camera 010; Sleeping on duty, using mobile phone, leaving the post; General risk". Algorithm alarms can display statistical data of different alarm types (such as safety helmet, smoke, personnel fall, etc.) in the form of bar charts.

[0147] For example, please refer to Figure 5 , Figure 5 This is a schematic diagram of the monitoring center interface of the intelligent safety monitoring system provided in Embodiment 2 of the industrial safety monitoring method based on target recognition of this application. Figure 5 As shown, users can select industrial cameras to view real-time video, video playback, and alarm images. The video playback and alarm image tracing time can be accurate to the second.

[0148] For example, please refer to Figure 6 , Figure 6 This is a schematic diagram of the user management interface of the intelligent safety monitoring system provided in Embodiment 2 of the industrial safety monitoring method based on target recognition of this application. Figure 6 As shown, compared to enterprise employee users, enterprise administrators can retrieve information on all users within the enterprise and perform operations such as role binding and password reset to manage user permissions, thereby further enhancing system security and business process standardization and effectively reducing risks.

[0149] For example, the system allows administrators to configure the mapping relationship between cameras and target recognition algorithms, configure the algorithm task list of the target detection model, turn the alarm status of selected cameras on or off, and adjust the algorithm recognition frame rate to adapt to different application scenarios. Table 1 shows the algorithm task list configured in a certain intelligent security monitoring system. Administrators can edit the algorithm task list as needed to delete, modify, or add corresponding algorithm tasks.

[0150] Table 1

[0151] Selection box Alarm Name Algorithm Tags Equipment Name Alarm status □ Smoke Detection Smoke Detection Factory camera 001 closure □ Smoke Detection Smoke Detection Test 222 by the Federation of Trade Unions closure □ Safety helmet identification Helmet identification Factory camera 002 closure □ Helmet identification Helmet identification Factory camera 001 closure □ Safety helmet identification Safety helmet identification Factory camera 003 closure □ Smoke Detection Smoke Detection Factory camera 012 closure □ Smoke Detection Smoke Detection Factory camera 011 closure

[0152] This embodiment acquires coordinate, confidence, and target category information from an optimized detection frame set, enabling precise location and identification of various targets in complex industrial scenarios. This provides reliable data support for subsequent monitoring and decision-making. By combining this with a color mapping table to determine target colors, the generated labeled video frames can intuitively and clearly display different target categories, improving users' ability to quickly understand and judge the situation on-site. Alarm signals generated based on these labeled video frames can promptly alert users when potential safety risks are detected, enabling real-time and effective monitoring of industrial sites and improving the safety and management efficiency of industrial production.

[0153] For example, to help understand the implementation process of the target recognition-based industrial safety monitoring method obtained by combining this embodiment with the above embodiment one, please refer to... Figure 7 , Figure 7 A simplified flowchart of an industrial safety monitoring method based on target recognition is provided, specifically:

[0154] First, the intelligent safety monitoring system acquires video data from monitoring equipment in real time using the TCP / IP protocol, and performs format conversion and compression on this data. Next, the system performs image normalization on the video frames, extracts features and predicts target bounding boxes using an algorithmic model, and simultaneously predicts confidence levels and class probabilities. Overlapping bounding boxes are removed using double-threshold non-maximum suppression. Finally, the system provides predictive warnings based on the target detection results and visualizes the results.

[0155] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the industrial safety monitoring method based on target recognition in this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0156] This application also provides an industrial safety monitoring device based on target recognition, please refer to... Figure 8 The target recognition-based industrial safety monitoring device includes:

[0157] The data processing module 10 is used to extract frame-by-frame images based on the real-time monitoring video stream and preprocess the frame-by-frame images to obtain preprocessed images.

[0158] The target recognition module 20 is used to perform target recognition on the preprocessed image to obtain the coordinate information and confidence information of the initial detection box set;

[0159] The detection optimization module 30 is used to perform double threshold non-maximum suppression processing on the initial detection box set based on the coordinate information and the confidence information to obtain an optimized detection box set.

[0160] The safety monitoring module 40 is used to obtain labeled video frames and alarm signals based on the optimized detection frame set in order to monitor industrial safety.

[0161] In one embodiment, the detection optimization module 30 is further configured to: obtain a detection box to be processed from the initial detection box set; obtain a reference detection box from the initial detection box set based on the confidence information; perform double threshold non-maximum suppression processing on the detection box to be processed and the reference detection box based on the coordinate information to obtain a non-maximum suppression result for the detection box to be processed; and determine an optimized detection box set based on the non-maximum suppression result for the detection box to be processed.

[0162] In one embodiment, the detection optimization module 30 is further configured to obtain a first cross-union ratio (CUP) threshold and a second CUP threshold, wherein the first CUP threshold is greater than the second CUP threshold; calculate the CUP between the detection box to be processed and the reference detection box based on the coordinate information to obtain a target CUP; and determine the non-maximum suppression result of the detection box to be processed based on the first CUP threshold, the second CUP threshold and the target CUP.

[0163] In one embodiment, the detection optimization module 30 is further configured to: determine that the non-maximum suppression result of the detection box to be processed is to suppress the confidence of the detection box to be processed when the target cross-union ratio is greater than the first cross-union ratio threshold; determine that the non-maximum suppression result of the detection box to be processed is to reduce the confidence of the detection box to be processed according to a linear decay function when the target cross-union ratio is less than or equal to the first cross-union ratio threshold and greater than or equal to the second cross-union ratio threshold; and determine that the non-maximum suppression result of the detection box to be processed is to retain the confidence of the detection box to be processed when the target cross-union ratio is less than the first cross-union ratio threshold.

[0164] In one embodiment, the data processing module 10 is further configured to extract video frame data from the real-time monitoring video stream based on the start code and end identifier; add frame sequence number, timestamp, monitoring device number and metadata to the video frame data to obtain message data; perform video format conversion on the message data according to a preset encoding format, and compress the message data after video format conversion to obtain transmission video stream data; and extract frame-by-frame images from the transmission video stream data according to a preset frame rate.

[0165] In one embodiment, the data processing module 10 is further configured to calculate the global brightness mean of the frame-by-frame image, and perform linear brightness adjustment processing on the frame-by-frame image according to the brightness mean to obtain a brightness-normalized image; perform histogram equalization processing on the brightness-normalized image to obtain a contrast-enhanced image; perform white balance correction processing on the contrast-enhanced image to obtain a color-corrected image; and perform size scaling and edge filling processing on the color-corrected image to obtain a pre-processed image.

[0166] In one embodiment, the safety monitoring module 40 is further configured to acquire coordinate information, confidence information, and target category information of each optimized detection box in the optimized detection box set, and acquire a color mapping relationship table; determine the target color from the color mapping relationship table according to the target category information; generate labeled video frames according to the coordinate information, the confidence information, and the target color, and generate corresponding alarm signals based on the labeled video frames to monitor industrial safety.

[0167] The target recognition-based industrial safety monitoring device provided in this application, employing the target recognition-based industrial safety monitoring method described in the above embodiments, can solve the technical problem of how to improve the target recognition accuracy in the safety monitoring process under complex industrial scenarios. Compared with the prior art, the beneficial effects of the target recognition-based industrial safety monitoring device provided in this application are the same as those of the target recognition-based industrial safety monitoring method provided in the above embodiments, and other technical features in the target recognition-based industrial safety monitoring device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0168] This application provides an industrial safety monitoring device based on target recognition. The industrial safety monitoring device based on target recognition includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the industrial safety monitoring method based on target recognition in the first embodiment described above.

[0169] Reference below Figure 9This document illustrates a structural schematic diagram of an industrial safety monitoring device based on target recognition, suitable for implementing embodiments of this application. The industrial safety monitoring device based on target recognition in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and vehicle-mounted terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 9 The industrial security monitoring device based on target recognition shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0170] like Figure 9 As shown, the target recognition-based industrial safety monitoring device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the target recognition-based industrial safety monitoring device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, LCDs (Liquid Crystal Displays), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the target-recognition-based industrial security monitoring equipment to exchange data wirelessly or via wired communication with other devices. Although the figure shows target-recognition-based industrial security monitoring equipment with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0171] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0172] The target recognition-based industrial safety monitoring device provided in this application, employing the target recognition-based industrial safety monitoring method described in the above embodiments, can solve the technical problem of how to improve the target recognition accuracy in the safety monitoring process under complex industrial scenarios. Compared with the prior art, the beneficial effects of the target recognition-based industrial safety monitoring device provided in this application are the same as those of the target recognition-based industrial safety monitoring method provided in the above embodiments, and other technical features of this target recognition-based industrial safety monitoring device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0173] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0174] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0175] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the target recognition-based industrial safety monitoring method in the above embodiments.

[0176] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory), or flash memory, optical fiber, CD-ROM (CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0177] The aforementioned computer-readable storage medium may be included in an industrial security monitoring device based on target recognition; or it may exist independently and not be assembled into an industrial security monitoring device based on target recognition.

[0178] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the target-based industrial safety monitoring equipment, cause the target-based industrial safety monitoring equipment to: extract frame-by-frame images from a real-time monitoring video stream and preprocess the frame-by-frame images to obtain preprocessed images; perform target recognition on the preprocessed images to obtain coordinate information and confidence information of an initial set of detection boxes; perform double-threshold non-maximum suppression processing on the initial set of detection boxes based on the coordinate information and the confidence information to obtain an optimized set of detection boxes; and obtain labeled video frames and alarm signals based on the optimized set of detection boxes to monitor industrial safety.

[0179] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LAN (Local Area Network) or WAN (Wide Area Network)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0180] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0181] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0182] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described target recognition-based industrial safety monitoring method. This addresses the technical problem of improving target recognition accuracy during safety monitoring in complex industrial scenarios. Compared to existing technologies, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the target recognition-based industrial safety monitoring method provided in the above embodiments, and will not be elaborated upon here.

[0183] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the target recognition-based industrial safety monitoring method described above.

[0184] The computer program product provided in this application can solve the technical problem of how to improve the target recognition accuracy in the security monitoring process in complex industrial scenarios. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the target recognition-based industrial security monitoring method provided in the above embodiments, and will not be repeated here.

[0185] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. An industrial safety monitoring method based on target recognition, characterized in that, The industrial safety monitoring method based on target recognition includes: Frame-by-frame images are extracted from real-time monitoring video streams, and the frame-by-frame images are preprocessed to obtain preprocessed images; The preprocessed image is subjected to target recognition to obtain the coordinate information and confidence information of the initial detection box set; Based on the coordinate information and the confidence information, the initial detection box set is subjected to double threshold nonmaximum suppression processing to obtain an optimized detection box set. Based on the optimized detection box set, labeled video frames and alarm signals are obtained to monitor industrial safety.

2. The method as described in claim 1, characterized in that, The step of performing double-threshold non-maximum suppression processing on the initial detection box set based on the coordinate information and the confidence information to obtain an optimized detection box set includes: Obtain the detection boxes to be processed from the initial set of detection boxes; Based on the confidence information, a reference detection box is obtained from the initial detection box set; Based on the coordinate information, double threshold non-maximum suppression processing is performed on the detection box to be processed and the reference detection box to obtain the non-maximum suppression result of the detection box to be processed. The optimized set of detection boxes is determined based on the nonmaximum suppression results of the detection boxes to be processed.

3. The method as described in claim 2, characterized in that, The step of performing dual-threshold non-maximum suppression processing on the detection box to be processed and the reference detection box based on the coordinate information to obtain the non-maximum suppression result of the detection box to be processed includes: Obtain a first cross-union ratio (CUP) threshold and a second CUP threshold, wherein the first CUP threshold is greater than the second CUP threshold; The intersection-union ratio (IUU) of the detection box to be processed and the reference detection box is calculated based on the coordinate information to obtain the target IUU. The non-maximum suppression result of the detection box to be processed is determined based on the first cross-union ratio threshold, the second cross-union ratio threshold, and the target cross-union ratio.

4. The method as described in claim 3, characterized in that, The step of determining the non-maximum suppression result of the detection box to be processed based on the first cross-union ratio threshold, the second cross-union ratio threshold, and the target cross-union ratio includes: When the target cross-union ratio is greater than the first cross-union ratio threshold, the non-maximum suppression result of the detection box to be processed is determined as the confidence level of suppressing the detection box to be processed; When the target cross-union ratio is less than or equal to the first cross-union ratio threshold and greater than or equal to the second cross-union ratio threshold, the non-maximum suppression result of the detection box to be processed is determined to be a reduction in the confidence of the detection box to be processed according to a linear decay function; When the target cross-union ratio is less than the first cross-union ratio threshold, the non-maximum suppression result of the detection box to be processed is determined as the confidence level for retaining the detection box to be processed.

5. The method as described in claim 1, characterized in that, The steps for extracting frame-by-frame images from a real-time monitoring video stream include: Extract video frame data from the real-time monitoring video stream based on the start code and end identifier; Add frame sequence number, timestamp, monitoring device number and metadata to the video frame data to obtain message data; The message data is converted to a video format according to a preset encoding format, and the converted message data is compressed to obtain the transmitted video stream data. Frame-by-frame images are extracted from the transmitted video stream data according to a preset frame rate.

6. The method as described in claim 1, characterized in that, The step of preprocessing the frame-by-frame images to obtain preprocessed images includes: Calculate the global brightness mean of the frame-by-frame image, and perform linear brightness adjustment processing on the frame-by-frame image based on the brightness mean to obtain a brightness-normalized image; The brightness-normalized image is subjected to histogram equalization to obtain a contrast-enhanced image; The contrast-enhanced image is subjected to white balance correction to obtain a color-corrected image; The color-corrected image is resized and edge-filled to obtain a pre-processed image.

7. The method according to any one of claims 1 to 6, characterized in that, The step of obtaining labeled video frames and alarm signals based on the optimized detection box set to monitor industrial safety includes: Obtain the coordinate information, confidence information, and target category information of each optimized detection box in the optimized detection box set, and obtain the color mapping relationship table; The target color is determined from the color mapping table based on the target category information; Based on the coordinate information, the confidence information, and the target color, a labeled video frame is generated, and a corresponding alarm signal is generated based on the labeled video frame to monitor industrial safety.

8. An industrial safety monitoring device based on target recognition, characterized in that, The device comprises: The data processing module is used to extract frame-by-frame images based on real-time monitoring video streams and preprocess the frame-by-frame images to obtain preprocessed images. The target recognition module is used to perform target recognition on the preprocessed image to obtain the coordinate information and confidence information of the initial detection box set; The detection optimization module is used to perform double threshold non-maximum suppression processing on the initial detection box set based on the coordinate information and the confidence information to obtain an optimized detection box set. The safety monitoring module is used to obtain labeled video frames and alarm signals based on the optimized detection frame set in order to monitor industrial safety.

9. An industrial safety monitoring device based on target recognition, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the industrial security monitoring method based on target recognition as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the industrial safety monitoring method based on target recognition as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Anti-interference industrial visual counting method and device based on spatio-temporal context perception

    CN122199550A