Intelligent operation safety monitoring method and device based on multi-source video and storage medium
By synchronously collecting and analyzing multi-view video data and environmental data, and combining it with the intelligent model of the cloud platform, the problem of insufficient data integration and linkage in existing technologies has been solved, realizing intelligent and real-time management of work safety monitoring and improving the monitoring effect of work safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN STARCAM TECH
- Filing Date
- 2026-03-11
- Publication Date
- 2026-04-24
AI Technical Summary
Existing intelligent monitoring methods for work safety suffer from poor data integration and insufficient linkage between analysis results and subsequent monitoring and handling, resulting in inadequate effectiveness and timeliness of work safety monitoring.
By synchronously collecting multi-view video data of the work area and its associated environmental data, performing time alignment and data association processing, monitoring-related data is formed. The target detection model and safety feature recognition model of the cloud analysis platform are used for analysis to generate violation identification reports. Combined with alarm information and equipment control commands, intelligent and real-time monitoring and management are achieved.
It achieves comprehensive and timely operation safety monitoring, enabling timely detection and rapid handling of safety hazards, and improving the accuracy and efficiency of operation safety monitoring.
Smart Images

Figure CN121924239A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to a method, device, and storage medium for intelligent monitoring of operational safety based on multi-source video. Background Technology
[0002] In various high-risk work scenarios such as construction and mining operations, work safety monitoring is a key link in avoiding production safety accidents and ensuring the personal safety of workers. With the development of video surveillance and intelligent analysis technology, intelligent monitoring methods based on video acquisition have gradually replaced traditional manual inspections and become the main technical means of work safety monitoring.
[0003] Currently, existing video-based intelligent monitoring methods for work safety generally involve collecting video data from the work area using camera equipment and uploading it to an analysis platform. The platform then analyzes the video data and provides feedback on the monitoring results. However, existing technologies have several shortcomings in practical applications. First, the collected data on the work area is of limited variety, and various related data are not effectively integrated, resulting in insufficient basis for the analysis platform to identify safety conditions and affecting the effectiveness of the identification results. Second, after completing video analysis, the analysis platform cannot form effective follow-up action linkages based on the analysis results, and the transmission of monitoring information lacks timeliness, making it difficult for relevant personnel to quickly grasp the safety status of the work area and take countermeasures. Therefore, existing intelligent monitoring methods for work safety suffer from poor data integration and insufficient linkage between analysis results and follow-up monitoring and response, failing to achieve efficient real-time intelligent monitoring of work safety.
[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0005] This application provides a method, device, and storage medium for intelligent monitoring of work safety based on multi-source video, which can realize intelligent and real-time monitoring and management of work safety, and improve the comprehensiveness and timeliness of work safety monitoring.
[0006] In a first aspect, embodiments of this application provide a method for intelligent monitoring of work safety based on multi-source video, applied to a work safety monitoring system. The work safety monitoring system includes a multi-view camera device deployed within a work area, an edge computing node communicatively connected to the multi-view camera device, a cloud analysis platform communicatively connected to the edge computing node, and a user terminal communicatively connected to the cloud analysis platform. The multi-view camera device is used to collect multi-view video data of the work area. The method for intelligent monitoring of work safety based on multi-source video includes: Simultaneously collect multi-view video data of the work area and its associated environmental data, and perform time alignment and data association processing on the multi-view video data and the environmental data to form monitoring associated data; The monitoring-related data is uploaded to the cloud-based analysis platform; The cloud-based analysis platform analyzes multi-view video data from the monitoring and related data using a preset target detection model and safety feature recognition model to identify the safety equipment wearing status and behavior of the workers and generate an analysis report containing the results of violation identification. Based on the analysis report, alarm information and equipment control commands are generated; The alarm information, the device control commands, and the real-time monitoring screen are pushed to the user terminal.
[0007] Furthermore, in some embodiments of this application, the synchronous acquisition of multi-view video data of the work area and its associated environmental data, and the time alignment and data association processing of the multi-view video data and the environmental data to form monitoring-related data, include: The multi-camera device simultaneously acquires video streams covering different angles and depths of field within the work area using at least one fixed-focus camera module and at least two zoom camera modules. Simultaneously collect real-time ambient light intensity data of the work area; The video stream and the ambient light intensity data at the same timestamp are bound together to form the monitoring association data.
[0008] Furthermore, in some embodiments of this application, the multi-camera device includes a protective housing, an image processing unit, and a GPS positioning module; the image processing unit is used to perform distortion correction and preliminary enhancement processing on the acquired raw video stream; the GPS positioning module is used to obtain the real-time geographical location information of the multi-camera device and attach it to the monitoring associated data.
[0009] Furthermore, in some embodiments of this application, the target detection model is a personnel detection model trained based on the YOLOv8 framework; the color and shape feature analysis includes determining whether the head region presents the standard color of a safety helmet and the outline features of an approximate circle or hemisphere.
[0010] Furthermore, in some embodiments of this application, the edge computing node is used to encode and compress the multi-view video data, cache it locally, and route it to the cloud analysis platform.
[0011] Furthermore, in some embodiments of this application, the step of generating alarm information and device control commands based on the analysis report includes: When a high incidence of violations is detected in a specific area, a control command is generated to adjust at least one of the pan-tilt angle, focal length, or fill light parameters of the multi-camera device associated with the specific area.
[0012] Furthermore, in some embodiments of this application, the user terminal also provides a user interface configured to perform at least one of the following operations: View real-time panoramic surveillance footage and footage from specific cameras; Receive and display violation alarm information, wherein the alarm information includes at least one of the following: violation type, time of occurrence, location, and screenshot; Query historical violation records and statistical reports; Remotely control the gimbal, focal length, and supplementary lighting equipment of the multi-camera device.
[0013] Secondly, this application also provides an intelligent monitoring device for work safety based on multi-source video, comprising: The data acquisition module is used to synchronously acquire multi-view video data of the work area and its associated environmental data, and to perform time alignment and data association processing on the multi-view video data and the environmental data to form monitoring associated data; The data upload module is used to upload the monitoring-related data to the cloud analysis platform; The cloud-based analysis module is used to analyze multi-view video data in the monitoring-related data through the cloud-based analysis platform, using a preset target detection model and safety feature recognition model, to identify the safety equipment wearing status and behavior of the workers, and generate an analysis report containing the results of violation identification. The control module is used to generate alarm information and equipment control commands based on the analysis report; The push module is used to push the alarm information, the device control commands, and the real-time monitoring screen to the user terminal.
[0014] Thirdly, this application also provides a storage medium storing a computer program that can be loaded by a processor and executed as described in the first aspect: a multi-source video-based intelligent monitoring method for job safety.
[0015] This application provides a method, device, and storage medium for intelligent monitoring of work safety based on multi-source video. First, multi-view video data and associated environmental data of the work area are collected synchronously. Time alignment and data association processing are performed on the two types of data to form monitoring-related data, providing a complete, correlated, and standardized data source for subsequent analysis on the cloud analysis platform. Then, after uploading the monitoring-related data to the cloud analysis platform, the multi-view video data is analyzed using a preset target detection model and safety feature recognition model. This systematically and accurately identifies the safety equipment wearing status and behavior of workers, generating an analysis report containing the results of violation identification, thus achieving intelligent judgment of work safety status and improving the accuracy of violation identification. Next, based on the analysis report, alarm information and equipment control commands are further generated, allowing the analysis results to be directly translated into specific monitoring and response measures, avoiding the problem of disconnect between analysis results and actual monitoring and handling. Finally, alarm information, equipment control commands, and real-time monitoring images are simultaneously pushed to the user terminal, enabling relevant personnel to promptly and comprehensively grasp the safety status of the work area and corresponding handling commands, thereby improving the timeliness of work safety monitoring and allowing staff to quickly take countermeasures against violations. In summary, this embodiment, by constructing an integrated intelligent monitoring system for operational safety, can promptly detect safety hazards during operations and facilitate rapid response, thereby achieving intelligent and real-time monitoring and management of operational safety, and improving the comprehensiveness, accuracy, and response efficiency of operational safety monitoring. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is an application environment diagram of the intelligent monitoring method for work safety based on multi-source video provided in the embodiments of this application; Figure 2 This is a flowchart illustrating the intelligent monitoring method for work safety based on multi-source video provided in an embodiment of this application; Figure 3 This is a schematic diagram of the data acquisition process provided in an embodiment of this application; Figure 4 This is a schematic diagram of the cloud-based analysis process provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of the intelligent monitoring device for work safety based on multi-source video provided in the embodiments of this application; Figure 6This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0018] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of systems and methods consistent with those detailed in the appended claims or with some aspects of this application.
[0019] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover descriptions such as non-exclusive inclusion, so that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.
[0020] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0021] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustration and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0022] To address the aforementioned technical problems and overcome the shortcomings of existing technologies, this application provides a method, device, and storage medium for intelligent monitoring of work safety based on multi-source video, which can realize intelligent and real-time monitoring and management of work safety, and improve the comprehensiveness and timeliness of work safety monitoring.
[0023] Figure 1 This is an application environment diagram of an intelligent monitoring method for job safety based on multi-source video in one embodiment. (Refer to...) Figure 1This multi-source video-based intelligent monitoring method for work safety is applied to a multi-source video-based intelligent monitoring system for work safety, which includes a terminal 110 and a server 120. The terminal 110 and server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal, specifically a mobile phone, tablet, or laptop. The server 120 can be a standalone server or a server cluster composed of multiple servers. The server 120 is configured to execute the aforementioned multi-source video-based intelligent monitoring method for work safety, including: synchronously collecting multi-view video data of the work area and its associated environmental data; performing time alignment and data association processing on the multi-view video data and environmental data to form monitoring-related data; uploading the monitoring-related data to a cloud analysis platform; the cloud analysis platform analyzing the multi-view video data in the monitoring-related data using a preset target detection model and safety feature recognition model to identify the safety equipment wearing status and behavior of the workers, generating an analysis report containing the results of violation identification; generating alarm information and equipment control commands based on the analysis report; and pushing the alarm information, equipment control commands, and real-time monitoring images to the user terminal.
[0024] This embodiment mainly applies a multi-source video-based intelligent monitoring method for work safety to a work safety monitoring system. The work safety monitoring system includes a multi-view camera device deployed in the work area, an edge computing node that is communicatively connected to the multi-view camera device, a cloud analysis platform that is communicatively connected to the edge computing node, and a user terminal that is communicatively connected to the cloud analysis platform. The multi-view camera device is used to collect multi-view video data of the work area.
[0025] Specifically, the work safety monitoring system provided in this embodiment includes a multi-camera device, an edge computing node, a cloud analysis platform, and a user terminal; the multi-camera device is communicatively connected to the edge computing node, the edge computing node is communicatively connected to the cloud analysis platform, and the cloud analysis platform is communicatively connected to the user terminal.
[0026] Among them, multi-view camera devices are deployed at key monitoring points within the work area to collect multi-view video data of the work area, providing a basic video data source for the entire intelligent monitoring process and serving as a prerequisite for all subsequent safety analysis work. During the overall operation of the system, the multi-view camera devices continuously collect video data from the entire work area from multiple perspectives, capturing real-time information such as the work status of personnel and the on-site environment. The collected multi-view video data is transmitted in real-time to the edge computing nodes connected to it, completing the output and transmission of front-end data.
[0027] For edge computing nodes, acting as intermediary nodes between multi-camera devices and cloud analytics platforms, they function as data routing and forwarding nodes, ensuring that front-end acquired data can be stably and efficiently transmitted to the cloud analytics platform. During the overall system operation, they receive video-related data from the multi-camera devices in real time, schedule the data according to the system's preset transmission protocol, and upload the received complete data to the cloud analytics platform accurately and in real time through the communication link. This achieves seamless data interaction between the front-end acquisition devices and the cloud analytics platform, ensuring that the cloud analytics platform can promptly obtain video data from the work area.
[0028] The cloud-based analytics platform performs professional intelligent analysis and processing on the received front-end data, identifying the safety status of workers, determining violations, and generating various instructions and information based on the analysis results. During the system's overall operation, it receives relevant data uploaded from edge computing nodes in real time, calls upon pre-set target detection and safety feature recognition models within the platform, performs in-depth analysis of multi-view video data, accurately identifies the safety equipment wearing status and work behavior of workers, determines whether violations exist, and generates an analysis report. Based on the analysis report, it further generates corresponding alarm information and equipment control instructions, and finally integrates the alarm information, equipment control instructions, and real-time monitoring images of the work area, pushing them to the user terminals connected to the platform.
[0029] For user terminals, the system receives various information and images pushed by the cloud analytics platform, providing real-time monitoring information viewing channels for operational safety monitoring and management personnel. Throughout the system's operation, it receives alarm information, equipment control commands, and real-time monitoring images of the work area from the cloud analytics platform, clearly displaying this information and images. This allows monitoring and management personnel to promptly and comprehensively grasp the safety status of the work area, including whether there are any violations and the specific details of those violations. Simultaneously, it obtains corresponding equipment control commands, providing information support for managers to take timely safety measures and execute equipment control operations.
[0030] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the intelligent monitoring method for work safety based on multi-source video provided in this application. Specifically, the intelligent monitoring method for work safety based on multi-source video provided in this application may include the following steps: S1. Synchronously collect multi-view video data of the work area and its associated environmental data, and perform time alignment and data association processing on the multi-view video data and environmental data to form monitoring associated data; Specifically, for step S1, multi-view video data and related environmental data of the work area are collected simultaneously. The multi-view video data is collected by multi-camera devices deployed in the work area, and the environmental data includes various environmental data related to the video acquisition and safety status assessment of the work area. After data collection, the acquired multi-view video data and environmental data are time-aligned to ensure consistency in the time dimension of the two types of data. Then, the time-aligned data are correlated to integrate the related video data and environmental data, ultimately forming integrated monitoring and correlation data.
[0031] For example, in the work area of a construction site, multi-camera devices simultaneously collect construction video data from different perspectives, such as the construction area, material storage area, and high-altitude work area. At the same time, they also collect real-time lighting, ambient temperature and humidity data for each area. The video data collected at 10:00 AM is then bound to the lighting, temperature and humidity data at that time point. In this way, the time alignment and association of data throughout the entire time period are completed, forming the monitoring and association data for construction site operations.
[0032] S2. Upload the monitoring-related data to the cloud analysis platform; Specifically, in step S2, the generated monitoring-related data is transmitted through an edge computing node that is connected to the multi-camera device. Relying on the communication connection between the edge computing node and the cloud analysis platform, the integrated monitoring-related data is uploaded to the cloud analysis platform completely and efficiently, providing a complete and standardized data source for subsequent intelligent analysis work.
[0033] The S3 cloud-based analytics platform analyzes multi-view video data from monitoring and related data using preset target detection and safety feature recognition models to identify the safety equipment wearing status and behavior of workers, and generates an analysis report that includes the results of violation identification. Specifically, in step S3, after receiving the uploaded monitoring data, the cloud-based analysis platform extracts multi-view video data as the core basis for analysis. It then calls upon the platform's preset target detection and safety feature recognition models to conduct intelligent analysis of the multi-view video data. Through model analysis, it accurately identifies the safety equipment wearing status of workers within the work area and their on-site work behavior. Combined with preset work safety regulations, it determines whether workers have violated regulations. Finally, it integrates the identification and judgment results to generate an analysis report containing the violation identification results.
[0034] For example, in mining operations, cloud-based analytics platforms use models to analyze mining operation videos to identify whether workers are wearing safety helmets, face shields, and other safety equipment as required. They also identify whether workers have violated regulations by entering dangerous areas or operating equipment improperly. The platforms record instances of not wearing face shields or entering blasting areas without authorization and generate corresponding violation analysis reports.
[0035] S4. Generate alarm information and equipment control commands based on the analysis report; Specifically, for step S4, based on the generated analysis report and combined with the actual needs of work safety monitoring, corresponding alarm information is generated for the violations identified in the report to provide early warning of potential work safety hazards; at the same time, equipment control instructions related to the control of on-site monitoring equipment are generated to make targeted adjustments to the monitoring equipment in the work area and achieve dynamic optimization of the monitoring process.
[0036] S5. Push alarm information, equipment control commands, and real-time monitoring screens to the user terminal; Specifically, for step S5, the cloud analysis platform integrates the generated alarm information, equipment control commands, and real-time monitoring images of the work area. Through the communication connection with the user terminal, the platform pushes the above information and images to the user terminal in real time, allowing monitoring and management personnel to view the relevant content on the terminal in real time.
[0037] This embodiment utilizes a work safety monitoring system to synchronously collect and integrate multi-view video data and related environmental data from the work area, providing a complete and standardized data source for cloud-based analysis and ensuring comprehensive safety status identification. Leveraging pre-set professional models in the cloud, intelligent analysis accurately identifies the safety equipment wearing status and work behavior of personnel, determining violations and achieving intelligent analysis of work safety monitoring. Based on the analysis report, alarm information and equipment control commands are generated and pushed to the user terminal along with real-time footage. This achieves seamless integration from safety analysis to hazard warning, equipment control, and information synchronization, enabling monitoring and management personnel to have a real-time and comprehensive understanding of the safety status of the work area and promptly address work safety hazards, thereby improving the real-time performance and effectiveness of work safety monitoring.
[0038] Furthermore, such as Figure 3 As shown, in some embodiments, step S1, "synchronously acquiring multi-view video data of the work area and its associated environmental data, and performing time alignment and data association processing on the multi-view video data and environmental data to form monitoring-related data," may specifically include: S11. Simultaneously acquire video streams covering different angles and depths of field of the work area using at least one fixed-focus camera module and at least two zoom camera modules in a multi-camera device; Specifically, in step S11, video stream acquisition is completed through the camera modules mounted on the multi-camera device. The device is equipped with at least one fixed-focus camera module and at least two zoom camera modules. Each camera module works collaboratively and starts the acquisition action synchronously. The fixed-focus camera module is used to acquire stable video streams of fixed key monitoring areas within the work area. The at least two zoom camera modules are used to acquire video streams of areas with different depths of field within the work area by adjusting their focal lengths, thereby achieving comprehensive video coverage of different spatial perspectives and different depths of field within the work area. During the acquisition process, the video stream acquisition time of each module is completely synchronized, forming multi-dimensional video stream data of the work area.
[0039] For example, in the working area of a construction site, one fixed-focus camera module of a multi-camera system continuously collects video streams of a fixed area, namely the working face of the foundation pit. Two zoom camera modules are adjusted to different focal lengths, one to collect video streams of the mid-depth-of-field area around the tower crane, and the other to collect video streams of the far-end material stacking area of the construction site. The three camera modules start collecting simultaneously, synchronously obtaining three video streams covering different perspectives and different depths of field of the construction site.
[0040] S12. Synchronously collect real-time ambient light intensity data of the work area; Specifically, in step S12, during the synchronous acquisition of the video stream, real-time ambient light intensity data of the work area is collected simultaneously. The collected light intensity data is the real-time light value of the work area that matches the video acquisition area. The acquisition action is completely consistent with the video stream acquisition of each camera module in the time dimension, ensuring the synchronization of light intensity data and video stream data in the acquisition time, and providing a time basis for subsequent data association.
[0041] For example, during the video stream acquisition process at the aforementioned construction site, while the fixed-focus module acquires video streams of the foundation pit working face and the zoom module acquires video streams of the area around the tower crane and the material storage area, the real-time light intensity values of these three areas are simultaneously acquired, ensuring that the light data and the video stream data acquisition time are completely correlated.
[0042] S13. Bind video streams and ambient light intensity data at the same timestamp to form monitoring-related data; Specifically, for step S13, a unified timestamp is added to each set of synchronized video stream data and light intensity data as a time identifier for data collection. All video streams with different perspectives and depths of field collected by the multi-view camera device at the same timestamp are bound and integrated one-to-one with the real-time ambient light intensity data of the work area collected synchronously at that time point. The scattered video stream data and environmental data are merged into a whole, and finally structured and correlated monitoring data are formed.
[0043] For example, at the timestamp of 15:00:00, the video stream of the foundation pit collected by the fixed-focus module, the video streams of the tower crane perimeter and material stacking area collected by the two zoom modules, and the light intensity data of the three areas collected synchronously at that time point are bound together to form the monitoring associated data corresponding to that time point; at the new timestamp of 15:00:01, the binding of the corresponding video stream and light data is completed in the same way, thereby realizing the continuous generation of monitoring associated data throughout the entire time period.
[0044] This embodiment uses a combination of fixed-focus and zoom camera modules to simultaneously acquire video streams from different angles and depths of field in the work area. Combined with the synchronously acquired light intensity data, it is precisely bound by timestamps, resulting in more comprehensive and structured monitoring data. This provides a high-quality data source with rich dimensions and accurate correspondence for subsequent cloud-based intelligent analysis.
[0045] Furthermore, in some embodiments, the multi-camera device includes a protective housing, an image processing unit, and a GPS positioning module; the image processing unit is used to perform distortion correction and preliminary enhancement processing on the acquired raw video stream; the GPS positioning module is used to obtain the real-time geographical location information of the multi-camera device and attach it to the monitoring-related data.
[0046] Specifically, the multi-view camera acquisition device provided in this embodiment integrates a protective shell, an image processing unit, and a GPS positioning module. The image processing unit and the GPS positioning module are both integrated inside the protective shell. The components form a data interaction and collaborative relationship, together constituting a multi-view video acquisition device that is adapted to the deployment of the work area.
[0047] The protective housing, as the external physical structure of the multi-camera device, serves as the fundamental load-bearing component of the entire device. It employs an integrated encapsulation design, completely enclosing all internal core components, including the image processing unit, GPS positioning module, and camera acquisition modules. It acts as the protective carrier for the entire device, with no additional external connection structures. Only communication and power supply interfaces with external devices are reserved, and these interfaces are sealed. The protective housing provides comprehensive physical protection for all internal core functional components, resisting the harsh environmental influences of the work area and preventing damage to internal components due to external factors. This ensures that the multi-camera device can stably and continuously conduct video acquisition in complex working environments.
[0048] The image processing unit, as the core data processing component of the multi-camera device, is integrated on a core circuit board inside the protective housing. It establishes a direct electrical signal and data connection with the device's camera acquisition module, and simultaneously maintains data interaction with the GPS positioning module within the housing. It can receive raw data from the camera acquisition module and also receive location information from the GPS positioning module and integrate the data. The image processing unit specifically performs targeted image processing on the raw video stream acquired in real-time by the camera acquisition module. Its core functions include distortion correction and preliminary enhancement processing. This corrects image distortion caused by lens optical characteristics and shooting angles in the raw video stream, while also performing basic image quality optimization to improve video clarity, providing a high-quality basic video source for subsequent video data processing and analysis.
[0049] The GPS positioning module, as an internal positioning component of the multi-camera device, is an independent satellite positioning module integrated within the protective housing. It establishes a data connection with the image processing unit, independently receiving satellite positioning signals and calculating its location. Simultaneously, it transmits the calculated geographic location information to the image processing unit in real time. The GPS positioning module acquires the multi-camera device's own geographic location information in real time and accurately, then transmits this information to the image processing unit. The image processing unit then appends this real-time geographic location information to the video-related data, which has undergone distortion correction and preliminary enhancement processing. Ultimately, this data is integrated into the monitoring-related data, giving it a clear identification of the device's acquisition location.
[0050] The protective housing provided in this embodiment provides physical protection for the multi-camera device in complex operating environments, ensuring stable operation of the equipment; the image processing unit optimizes the quality of the original video stream, and the GPS positioning module adds accurate geographical location information to the monitoring-related data, comprehensively improving the overall quality of the monitoring-related data from the device end, while enriching the spatial dimension information of the data.
[0051] Furthermore, such as Figure 4 As shown, in some embodiments, step S3, "The cloud analysis platform analyzes multi-view video data from the monitoring-related data using a preset target detection model and safety feature recognition model to identify the safety equipment wearing status and behavior of the workers," may specifically include: S31. Use an object detection model to detect people in video frame images and obtain the bounding box coordinates of each person target; Specifically, in step S31, the cloud analysis platform retrieves a preset target detection model, decomposes the multi-view video data in the monitoring-related data into single-frame video images, uses the model to perform full-domain personnel target detection on each video frame, identifies all the workers in the image through model analysis, and generates a corresponding bounding box for each identified worker target, accurately obtaining the planar coordinate information of the bounding box of each worker target. This coordinate information can clearly define the position and range of each worker in the video frame image, providing accurate location basis for subsequent area interception.
[0052] For example, in a video frame image of a construction site, the target detection model identifies 5 workers in the image, generates a rectangular bounding box for each of the 5 workers, and obtains the coordinates of the upper left and lower right corners of each bounding box. These coordinates can be used to determine the specific location and outline of each worker in the image.
[0053] S32. For each detected person target, extract the head region image and the torso region image according to the bounding box coordinates of the person target; Specifically, for step S32, for each detected person target, based on the bounding box coordinates corresponding to the person target, according to the preset region division rules, the head region image and torso region image of the person target are respectively cropped in the video frame image. During the cropping process, the body contour of the person target is strictly followed to ensure the integrity and independence of the head and torso region images, avoid the overlap of region images of different person targets, and prevent the cropped area from containing too much irrelevant background, so as to ensure the accuracy of the objects of subsequent feature analysis.
[0054] For example, based on the bounding box coordinates of a worker in the above construction site video frame image, the head region image of the worker is cropped with the upper half of the bounding box as the reference, and the torso region image is cropped with the middle main body part of the bounding box as the reference. The head and torso region images of each worker are extracted and stored separately to form 5 sets of corresponding head and torso region images.
[0055] S33. Perform color and shape feature analysis on the head region image to determine whether a preset style of safety helmet is being worn, and obtain the first judgment result; Specifically, for step S33, a special color and shape feature analysis is performed on the head area image of each captured personnel target. Combined with the style requirements for safety helmets in the work safety specifications, the feature analysis is used to determine whether the personnel are wearing a safety helmet that meets the requirements. The judgment result is recorded as the first judgment result. The first judgment result only includes two situations: "wearing a safety helmet that meets the requirements" and "not wearing a safety helmet that meets the requirements", so as to make the judgment result clear and unambiguous.
[0056] For example, when analyzing images of the head area of construction workers, we can determine whether the main color of the head area meets the color requirements of the safety helmet, and at the same time, we can determine whether the outer contour of the head area is the typical contour of the safety helmet. If a person's head area image does not have corresponding color and contour features, it is determined that the person is not wearing a safety helmet that meets the requirements, thus forming the corresponding first judgment result.
[0057] S34. Perform high reflectivity feature analysis on the torso region image to determine whether the wearer is wearing a pre-defined standard reflective vest, and obtain a second judgment result; Specifically, for step S34, a special high reflectivity feature analysis is performed on the torso area image of each captured personnel target. Combined with the standard requirements for reflective clothing in the work safety regulations, the system identifies whether the torso area has high reflectivity features that meet the requirements, and determines whether the personnel are wearing reflective clothing that meets the requirements. The determination result is recorded as the second determination result. The second determination result only includes two situations: "wearing reflective clothing that meets the requirements" and "not wearing reflective clothing that meets the requirements", ensuring the uniqueness of the determination result.
[0058] For example, when examining images of the torso of construction workers, the system detects whether there are continuous and standardized highly reflective stripes on the torso. If no such highly reflective features are found in the torso image, it is determined that the worker is not wearing a compliant reflective vest, thus forming a corresponding second judgment result.
[0059] S35. Based on the first and second judgment results, determine the safety equipment wearing status of the personnel target. If any equipment is detected not being worn as required, it shall be marked as a violation. Specifically, for step S35, the first judgment result and the second judgment result corresponding to each personnel target are integrated, and the overall safety equipment wearing status of the personnel target is determined based on the two results. According to the requirements of the work safety specifications, if it is detected that either the safety helmet or the reflective vest is not worn as required, the personnel target is directly marked as having a violation. If both pieces of equipment are worn as required, the personnel target's safety equipment wearing status is judged to be compliant and there is no violation.
[0060] For example, among the five construction workers mentioned above, if a worker's first judgment result is "not wearing a safety helmet" and the second judgment result is "wearing a reflective vest", then that worker is marked as having violated regulations; if both judgment results are compliant, then it is determined that there is no problem with the wearing status; if both judgment results are non-compliant, then it is also marked as having violated regulations.
[0061] This embodiment uses a target detection model to accurately locate personnel targets, selectively extracts the head and torso areas and conducts exclusive feature analysis, and combines the dual results to determine the safety equipment wearing status and mark violations, forming a standardized and procedural judgment logic, making the identification of workers' safety equipment wearing status and violations more targeted and accurate.
[0062] Furthermore, in some embodiments, the target detection model is a personnel detection model trained based on the YOLOv8 framework; color and shape feature analysis includes determining whether the head region presents the standard color of a safety helmet and the outline features of an approximate circle or hemisphere.
[0063] Specifically, in the process of detecting personnel targets in video frame images, this embodiment uses a dedicated personnel detection model trained on the YOLOv8 framework as the core detection model. The video frame images in the monitoring-related data are input into the model. The model uses its own network structure to extract features, regress targets, and filter them, accurately identifying all the personnel targets in the image. At the same time, it generates a corresponding bounding box for each identified personnel target and outputs the precise coordinate information of the bounding box, thereby completing the detection and location definition of personnel targets.
[0064] For the captured image of the head region of the target personnel, the color features of the whole domain are extracted and analyzed. First, the main color tone and main color distribution features of the head region image are obtained. Then, they are compared with the standard color of the safety helmet specified in the work safety regulations one by one to determine whether the color features of the head region match the standard color of the safety helmet. If the head region shows the standard color features, the color dimension is determined to meet the requirements; otherwise, the color dimension is determined to not meet the requirements.
[0065] While performing color feature analysis of the head region, the shape and contour of the head region image are extracted and analyzed. Through image contour recognition algorithm, the overall external contour of the head region is outlined. The extracted contour features are then compared with the contour features of approximately circular or hemispherical shapes to determine whether the overall contour of the head region conforms to the typical contour features of a safety helmet. If it presents an approximately circular or hemispherical contour, the shape dimension is determined to meet the requirements; otherwise, the shape dimension is determined to not meet the requirements.
[0066] The color feature determination results and shape contour feature determination results are integrated. Only when the head region image simultaneously meets the two requirements of presenting the standard color of a safety helmet and having an approximately circular or hemispherical contour feature can the head region of the target person be determined to meet the feature requirements of a safety helmet. If either feature does not meet the requirements, the head region of the target person is directly determined not to meet the feature requirements of a safety helmet.
[0067] This embodiment employs a personnel detection model trained based on the YOLOv8 framework to improve the accuracy and efficiency of personnel target detection; it clarifies the dual feature analysis dimensions of the standard color of the safety helmet and the approximately circular / hemispherical outline, establishing a unified and objective standard for determining whether a safety helmet is worn, thereby improving the accuracy and standardization of safety helmet wearing status determination.
[0068] Furthermore, in some embodiments, the intelligent monitoring method for job safety based on multi-source video further includes: S36. Perform fusion calculation on the security equipment identification results of the same person target in multiple consecutive video images; Specifically, for step S36, after completing the identification of the safety equipment wearing status of the personnel target in a single frame video image, for the same personnel target in the video stream, extract the safety equipment identification results for each of the multiple consecutive video images. The identification results include the determination results of wearing safety helmets and reflective vests. These scattered single-frame identification results are summarized and integrated, and the multi-frame results are processed by data integration according to the preset fusion calculation rules to form a fusion calculation dataset of the personnel target's safety equipment identification results, thereby realizing the centralized collection and quantification processing of multi-frame identification results.
[0069] For example, in a video stream from a construction site, a construction worker is identified as the target. The results of determining whether the worker is wearing a safety helmet or reflective vest are extracted from each of 10 consecutive video frames. These 10 sets of recognition results are then aggregated and fused according to rules to form a safety equipment recognition fusion calculation dataset corresponding to the construction worker.
[0070] S37. Based on the fusion calculation results, determine whether the consistency of the recognition results of multiple frames exceeds a preset threshold; Specifically, for step S37, based on the obtained fusion calculation dataset, for the wearing determination results of safety helmets and reflective vests, the proportion or frequency of the same determination results in multiple consecutive frames is counted, and the statistical results are compared with the system's preset consistency threshold. It is then determined whether the consistency of the multi-frame recognition results of safety helmets and reflective vests reaches or exceeds the preset threshold. The preset threshold is a value pre-set by the system according to the accuracy requirements of work safety monitoring, and it is the core standard for determining whether the multi-frame recognition results are stable.
[0071] For example, the system's preset consistency threshold for multi-frame recognition results is 80%. For the 10-frame fusion calculation dataset of the construction workers mentioned above, it was found that in the results of their safety helmet wearing judgment, 9 frames were "not wearing", with a consistency rate of 90%, which exceeds the preset threshold of 80%; in the results of their reflective vest wearing judgment, 10 frames were "wearing", with a consistency rate of 100%, which also exceeds the preset threshold.
[0072] S38. When the consistency exceeds a preset threshold, determine the final safety equipment wearing status of the personnel target; Specifically, for step S38, if the consistency of the multi-frame recognition results of a certain type of safety equipment (helmet / reflective vest) exceeds a preset threshold, the same judgment result that dominates the type of equipment in the multiple frames is taken as the final judgment result of the person's wearing of that type of equipment; after completing the final judgment of the helmet and reflective vest respectively, the final results of the two types of equipment are integrated to comprehensively determine the overall safety equipment wearing status of the person.
[0073] For example, if the consistency of the safety helmet identification results for the construction workers exceeds the threshold by 90%, "not wearing" will be the final determination of the safety helmet; if the consistency of the reflective vest identification results exceeds the threshold by 100%, "wearing" will be the final determination of the reflective vest; after integration, the final safety equipment wearing status of the personnel target will be determined as "not wearing a safety helmet, wearing a reflective vest".
[0074] This embodiment effectively avoids recognition deviations caused by accidental factors such as image blurring and partial occlusion in single-frame images by fusing and calculating the identification results of the same person's safety equipment in multiple consecutive frames, and determines the final wearing status through consistency threshold verification. This reduces the probability of misjudgment and improves the reliability of safety equipment wearing status determination.
[0075] Furthermore, in some embodiments, step S36, "fusing and calculating the security equipment identification results of the same person target in multiple consecutive video images," may specifically include: S361. Track the same person target across frames in a continuous sequence of video frames; Specifically, in step S361, for the continuous video frame sequence in the monitoring and associated data, dynamic matching and continuous tracking of personnel targets detected frame by frame are performed through temporal correlation analysis and target tracking logic. Combining the position, outline, and motion trajectory features of the personnel targets in the video frames, a unique target identifier is assigned to each independent personnel target. This identifier enables accurate association of the same personnel target in different video frames, ensuring that regardless of positional movement or posture changes within the work area, the personnel target can be continuously tracked and identified as the same analysis object, avoiding confusion in the identification results of different personnel targets.
[0076] For example, in a continuous video frame sequence at a construction site, if a construction worker is detected moving around the foundation pit, a unique target identifier is assigned to him / her. As the worker moves from the east side of the foundation pit to the west side, the worker will be matched with the same identifier in every frame of the video, thus achieving continuous cross-frame tracking of the worker and preventing him / her from being identified as a new person target due to position movement.
[0077] S362. For each tracked target ID, record the recognition results and confidence level of the safety helmet and reflective vest for each target ID in N consecutive frames; Specifically, for step S362, an independent information record file is established for each unique target identifier determined through cross-frame tracking. First, the number of video frames N to be continuously tracked and analyzed is set. Then, for each target identifier, the recognition result of its helmet-wearing state and the corresponding confidence value are recorded sequentially across N consecutive video frames, according to the temporal order of the video frames. Similarly, the recognition result of its reflective vest-wearing state and the corresponding confidence value are also recorded. All recorded information corresponds one-to-one with the target identifier and the video frame order, forming a structured multi-frame recognition information ledger, completely preserving all the data required for fusion calculation.
[0078] For example, a record file is established for the target identification of the aforementioned pit workers. The number of continuous tracking frames N is set to 10 frames, and the frames are recorded sequentially: Frame 1: Safety helmet not worn (confidence 0.98), reflective vest worn (confidence 0.99); Frame 2: Safety helmet not worn (confidence 0.97), reflective vest worn (confidence 0.99)... up to Frame 10. The identification results and corresponding confidence scores of the two types of equipment in the 10 frames are completely recorded in the dedicated file of the target identification.
[0079] This embodiment assigns a unique ID to the same person target through cross-frame tracking, ensuring the uniqueness of the analysis object. At the same time, it standardizes the recording of the identification results and confidence levels of safety helmets and reflective vests in N consecutive frames, providing complete, detailed and traceable basic data for fusion computing, and ensuring the effectiveness and accuracy of fusion computing from the data source.
[0080] Furthermore, in some embodiments, step S38, "when consistency exceeds a preset threshold, determine the final safety equipment wearing status of the personnel target," may specifically include: S381. Perform statistical analysis on the N identification results of the safety helmet and reflective vest corresponding to each target ID; Specifically, for step S381, for each unique target ID determined by cross-frame tracking, the recorded helmet and reflective vest wearing status identification results in N consecutive frames are retrieved. Frequency statistical analysis is performed according to the two independent dimensions of helmet and reflective vest. The specific frequency of the two identification results of "wearing" and "not wearing" under each dimension is counted one by one to form frequency statistics data for the two dimensions. During the statistical process, it is ensured that the data accurately corresponds to the target ID and equipment type, and there is no dimension confusion or data mismatch.
[0081] For example, in a construction scenario, the N consecutive frames of target ID001 are 10 frames. After statistical analysis of its recognition results, we find that: in the safety helmet dimension, "not worn" appears 9 times and "worn" appears 1 time; in the reflective vest dimension, "worn" appears 10 times and "not worn" appears 0 times, forming independent frequency statistics for the two equipment dimensions of this target ID.
[0082] S382. If any equipment is identified as being worn more frequently than a preset first frequency threshold, then it is determined that the equipment has been worn. S383. If the frequency of any equipment being identified as not being worn exceeds a preset second frequency threshold, then it is determined that the equipment is not being worn. Specifically, for steps S382 and S383, based on preset judgment rules, a first frequency threshold (wearing threshold) and a second frequency threshold (not wearing threshold) are set for determining the wearing status of safety helmets and reflective vests, respectively. The frequency statistics for both equipment dimensions are compared with their corresponding thresholds to independently determine the wearing status of each type of equipment: if the frequency of a certain type of equipment being identified as "wearing" exceeds the first frequency threshold, the equipment is directly determined to be in a worn state; if the frequency of a certain type of equipment being identified as "not wearing" exceeds the second frequency threshold, the equipment is directly determined to be in a not-wearing state. The determinations for the two types of equipment are independent and do not affect each other.
[0083] For example, in the above scenario, the first frequency threshold is preset to 7 times and the second frequency threshold is preset to 6 times. The statistical data of target ID001 is compared with the threshold: the safety helmet is "not worn" 9 times > the second frequency threshold of 6 times, so the safety helmet is determined to be in an unworn state; the reflective vest is "worn" 10 times > the first frequency threshold of 7 times, so the reflective vest is determined to be in a worn state, thus completing the independent determination of the two types of equipment.
[0084] S384. Based on the independent assessment results of safety helmets and reflective vests, determine the final safety equipment wearing status of the personnel target; Specifically, for step S384, the results of the safety helmet wearing status determination and the reflective vest wearing status determination are integrated and combined with the determination conclusions of the two types of equipment to form the overall and final safety equipment wearing status of the personnel target corresponding to the target ID. The determination results fully reflect the actual wearing status of the two types of equipment and clearly reflect whether there is a situation where the equipment is not worn as required.
[0085] For example, if the determination result of target ID001 is that the safety helmet is not worn and the reflective vest is worn, the final safety equipment wearing status of the target person is determined to be "not wearing a safety helmet, wearing a reflective vest"; if the determination result of another target ID002 is that the safety helmet is worn but the reflective vest is not worn, its final status is "wearing a safety helmet, not wearing a reflective vest".
[0086] This embodiment performs multi-dimensional statistical analysis on the N recognition results of safety helmets and reflective vests, independently determines the wearing status of various equipment according to preset frequency thresholds, and then integrates them, making the final wearing status determination process quantifiable and objective, eliminating the arbitrariness of subjective judgment, and further improving the objectivity and persuasiveness of the determination results.
[0087] Furthermore, in some embodiments, the intelligent monitoring method for job safety based on multi-source video further includes: S301. Real-time monitoring of light intensity and light direction in associated environmental data; S302. When the scene is determined to be backlit based on the light intensity and light direction, the exposure parameters of the video capture are automatically adjusted, and local image enhancement processing is performed on the head area image and the torso area image.
[0088] Specifically, this embodiment also provides a method for accurately identifying backlit scenes through real-time monitoring of the lighting environment, and optimizes the image quality of the core analysis area under backlight through parameter adjustment and local image enhancement. The specific process is as follows: During the continuous acquisition of video data in the work area, environmental data associated with the video acquisition is monitored in real time. The monitoring targets are two core parameters in the environmental data: light intensity and light direction. The monitoring process maintains synchronization with the acquisition of video frames in the time dimension. The light parameter data is dynamically updated and fed back in real time with the video acquisition, forming continuous light environment monitoring data. This provides objective and real-time environmental basis for subsequent scene judgment, ensuring that changes in the light environment can be captured in a timely manner. For example, in the work area of an open-air construction site, while the video acquisition equipment continuously captures the foundation pit work surface, it will monitor the real-time light intensity (in lux) of the work surface and the direction of natural light (e.g., from the northeast in the morning and from the northwest in the evening). The light data is updated synchronously with each frame of video acquisition, reflecting the changes in the light environment of the work surface in real time.
[0089] Using real-time monitored light intensity and direction data as the core criteria, combined with the actual shooting direction of the video capture device, a scene determination logic is established to determine whether the current video capture is in a backlight scene. During the determination process, the relative angle between the light direction and the camera's shooting direction, as well as the magnitude of the light intensity, are considered to comprehensively determine whether lighting factors have caused typical backlight shooting characteristics such as a dark foreground and blurred subject features in the captured image. If the conditions for a backlight scene are met, subsequent targeted processing actions are triggered; if it is a normal lighting scene, the original video capture and image processing methods remain unchanged. For example, if the construction site video capture device is facing south to capture the work surface, and the real-time monitoring shows direct sunlight from south to north, and the light intensity exceeds the preset backlight determination threshold, then the outline of the workers in the captured image will be dark and the details will be blurred, meeting the characteristics of backlight, thus determining that the current scene is backlight. If the light direction is from the east, forming a 90° angle with the camera's shooting direction, even if the light intensity is high, it is not determined to be a backlight scene.
[0090] When the current video capture is determined to be in a backlit scene, two targeted processing actions are immediately initiated and executed simultaneously: First, the exposure-related parameters of the video capture are automatically adjusted to optimize the overall image quality and alleviate the problem of a dark foreground caused by backlighting. Second, for the head and torso images of personnel used for subsequent safety equipment identification and analysis, separate local image enhancement processing is performed, focusing on the core analysis area and specifically improving the image clarity and feature recognition of that area, without performing global image enhancement to avoid image quality interference from irrelevant background areas. For example, when the video capture of a mining operation area is determined to be in a backlit scene, the system will automatically increase the exposure compensation value of the video capture and adjust the appropriate exposure time to improve the overall brightness of the captured image; at the same time, local brightness and contrast enhancement processing is performed on the head and torso areas of the workers in the captured video frames to make the outline of the head area and the reflective features of the torso area clearer, while the background in the image, such as walls and equipment areas, is not subject to additional enhancement processing.
[0091] This embodiment monitors the light intensity and direction of the work area in real time, automatically adjusts the video acquisition exposure parameters when a backlight scene is determined, and performs local image enhancement on the core areas of the head and torso, effectively solving the problem of blurry images under backlight and ensuring the stability and accuracy of the identification of safety equipment features of workers in backlight scenes.
[0092] Furthermore, in some embodiments, step S302, "when the scene is determined to be backlit based on the light intensity and light direction, automatically adjust the exposure parameters of the video capture and perform local image enhancement processing on the head region image and the torso region image," includes: Determine whether it is a severe backlighting scene based on the angle between the direction of the light and the direction of the camera, as well as the brightness contrast between the foreground target and the background. Specifically, combining the collected ambient lighting data with the visual characteristics of the video footage, the determination of severe backlighting scenes is carried out from two core dimensions: first, calculating the angle between the direction of the light and the direction of the camera's shooting to determine the degree to which the light directly hits the camera lens; second, analyzing the brightness contrast between the foreground target (workers) and the background area in the video footage to quantify the difference in the degree of darkness in the foreground and brightness in the background. Based on the combined results of these two dimensions, and according to preset criteria, it is determined whether the current video capture scene is a severe backlighting scene. If both the degree of direct light and the difference in brightness contrast between the foreground and background are met, it is determined to be a severe backlighting scene, triggering subsequent parameter adjustments and image enhancement actions; if the criteria are not met, video capture and processing are performed in the normal mode.
[0093] For example, in an open-air road construction scene, the camera is facing due west to film the work area. In the afternoon, the sunlight is shining directly from west to east. The angle between the direction of the sunlight and the direction of the camera is close to 0°. Moreover, the brightness value of the construction workers (foreground) in the video is much lower than the brightness value of the road surface and the sky (background). The contrast difference between the two exceeds the preset standard. At this time, the scene is determined to be a severe backlight scene.
[0094] If the scene is determined to be a severe backlighting scene, the multi-camera device is controlled to increase the exposure compensation value and reduce the sensitivity of the image sensor in order to suppress background overexposure and improve the details of the foreground target. Specifically, upon determining a severe backlighting scene, the exposure-related parameters of the multi-camera system are immediately and automatically adjusted. Two key parameter adjustments are executed: first, the camera system increases exposure compensation to compensate for insufficient brightness of the foreground target due to backlighting; second, the sensitivity of the camera's image sensor is reduced to suppress overexposure of the background area caused by strong light. These two parameter adjustments are performed simultaneously and in tandem, effectively improving the detail of the foreground personnel target while preventing the background from appearing washed out and losing detail, thus optimizing the basic image quality from the source of video capture.
[0095] For example, in the severe backlighting scenario of road construction mentioned above, the system automatically adjusts the exposure compensation value of the camera device to an appropriate value, while lowering the sensitivity of the image sensor from a high value. After adjustment, the overexposed sky and road surface areas in the video are restored to normal visual effect, and the main body of the construction workers under backlight is no longer excessively dark, and their body outlines and basic features can be clearly distinguished.
[0096] In the image processing unit, for the adjusted acquired image, local contrast enhancement algorithm and edge sharpening algorithm are applied to the head and torso areas of the personnel target to enhance the edge contour of the safety helmet and the high reflective strip features of the reflective vest. Specifically, in the image processing unit, instead of performing global image enhancement on the video images acquired after exposure parameter adjustments, the focus is on the head and torso areas of the workers—two core areas for safety equipment identification—and targeted algorithm enhancement processing is applied. Specifically, a local contrast enhancement algorithm is used to improve the image depth in the core areas, and an edge sharpening algorithm is used to strengthen the contours and details of the core areas. Ultimately, this achieves precise enhancement of the edge contours of the safety helmet and the highly reflective stripes of the reflective vest, allowing the key identification features of the safety equipment to stand out from the backlit image.
[0097] For example, in the adjusted video images of the road construction scene mentioned above, the algorithm only processes the head and torso areas of the construction workers. By enhancing the local contrast, the boundary between the safety helmet and the hair and clothing becomes clearer. By sharpening the edges, the highly reflective strips of the reflective vest form obvious visual features. Even in backlight conditions, the outline of the safety helmet and the position of the reflective strips of the reflective vest can be clearly distinguished.
[0098] This embodiment combines the lighting angle and the brightness contrast of the foreground and background to accurately determine severe backlighting scenes, and adjusts the exposure parameters to achieve image quality balance between the foreground and background. Then, it uses local algorithms to enhance the key identification features of safety helmets and reflective vests, fundamentally solving the problem of difficult identification of core features under severe backlighting and ensuring the accuracy of safety equipment identification in this scene.
[0099] Furthermore, in some embodiments, the intelligent monitoring method for job safety based on multi-source video also includes: S61. Obtain the GPS location information of at least two multi-view camera devices in the work safety monitoring system; Specifically, for step S61, within the work area where the multi-camera devices have been deployed, the GPS geographic location information of at least two multi-camera devices is extracted from the work safety monitoring system. This information is acquired in real time by the GPS positioning module of each device and uploaded to the system. It includes core spatial data such as the precise latitude, longitude, and elevation of the device deployment points. Moreover, the location information of each device is valid data under the same geographic coordinate system, providing a real and accurate reference for the equipment location to be established in the future.
[0100] For example, in large open-pit mining areas, the system deploys multiple multi-camera devices to extract GPS information from the multi-camera devices at two locations: the mining face and the entrance to the transportation channel, thereby obtaining precise spatial location data for these two locations.
[0101] S62. Establish a global spatial coordinate system for the work area based on GPS geographic location information; Specifically, for step S62, based on the actual geographical area of the work area, the GPS geographical location information of at least two multi-camera devices obtained in step 1 is used as the core spatial reference point. Following the rules for constructing a three-dimensional spatial coordinate system, a fixed landmark within the work area (such as the work area gate or the base of core equipment) is selected as the origin of the coordinate system. The east, north, and vertical directions are set as the three-dimensional coordinate axes. The location information of all extracted multi-camera devices is mapped to this coordinate system, completing the calibration and establishment of the coordinate system and forming a global spatial coordinate system covering the entire work area. This coordinate system provides a unified spatial positioning benchmark for all deployed multi-camera devices, allowing the shooting perspectives of each device to be spatially correlated.
[0102] For example, taking the west gate of the aforementioned mining operation area as the origin of the global spatial coordinate system, and setting east as the X-axis, north as the Y-axis, and vertical upward as the Z-axis, the GPS information of the multi-camera devices at the mining face and the transportation channel entrance is converted into three-dimensional coordinates (X1, Y1, Z1) and (X2, Y2, Z2) in this coordinate system. At the same time, the positions of other multi-camera devices in the operation area are also mapped to this coordinate system, realizing unified spatial calibration of the positions of all equipment.
[0103] S63. Perform feature point matching and image registration on video streams acquired from different devices, and generate panoramic monitoring images based on registration relationships and weighted fusion algorithms; Specifically, for step S63, for the video streams of the work area synchronously acquired by each multi-camera device based on the global spatial coordinate system, feature point matching is first performed on the video streams of different devices to extract significant spatial feature points (such as building corners, equipment outlines, road markings, etc.) in each video frame and perform precise spatial matching to establish spatial correspondence between different video streams; then, based on this matching relationship, image registration of the video streams of different devices is completed, so that the video images from different perspectives can achieve precise spatial alignment in the global spatial coordinate system, eliminating image misalignment and offset problems; finally, based on the registration relationship, a weighted fusion algorithm is used to fuse the registered video images, and then the fused images are combined as a whole to finally generate a panoramic monitoring image covering the entire work area.
[0104] For example, in the global spatial coordinate system of a construction site, the synchronous video streams from three multi-camera devices in the foundation pit area, tower crane area, and material storage area are analyzed. Feature points such as the tower crane body, foundation pit edge, and material silo corners in each video frame are extracted and matched. Image registration is performed on the video images of the three devices to achieve spatial alignment. Then, a weighted fusion algorithm is used to process overlapping areas of the images (such as the junction between the foundation pit and tower crane areas). Finally, the fused images are combined to generate a panoramic monitoring image covering all work areas of the construction site, with no blind spots in the shooting angle.
[0105] This embodiment establishes a global spatial coordinate system for the work area based on the GPS geographic location information of at least two multi-camera devices. By using feature point matching, image registration, and weighted fusion, a panoramic monitoring image is generated, effectively eliminating blind spots in the monitoring perspective of a single camera device. This achieves full-area visual monitoring of the work area and improves the comprehensiveness and intuitiveness of work safety monitoring.
[0106] Furthermore, in some embodiments, step S63, "performing feature point matching and image registration on video streams acquired by different devices, and generating a panoramic monitoring image based on the registration relationship and a weighted fusion algorithm," may specifically include: S631. The scale-invariant feature transform algorithm is used to extract and match feature points from synchronized video frames acquired from different devices; Specifically, in step S631, for synchronous video frames acquired simultaneously by different multi-camera devices within the work area, a scale-invariant feature transform algorithm is used to extract and match feature points. This algorithm can accurately extract unique and stable salient feature points from video frames of different scales and perspectives, such as building corners, equipment outlines, road markings, and pit edges within the work area—features that are not easily changed by the shooting angle. After extraction, the algorithm compares the similarity of feature points in different video frames, establishing a one-to-one correspondence between feature points in different video frames, forming matched feature point pairs, and laying a precise feature foundation for subsequent image spatial alignment.
[0107] For example, in a construction site, the scale-invariant feature transform algorithm is used to extract feature points such as pit corners and edge protection frame nodes from the video frames of two multi-camera monitoring devices in the pit area and tower crane area, and feature points such as tower crane connection points and foundation support corners from the video frames of the pit area. After comparison by the algorithm, the corresponding feature points of adjacent protective frames and tower crane supports in the two video frames are matched to form multiple sets of feature point pairs.
[0108] S632. Based on the matched feature point pairs, solve the homography transformation matrix between images to complete image registration; Specifically, in step S632, based on the obtained matching feature point pairs, the planar coordinate information of each feature point pair in different video frames is extracted. A spatial geometric transformation algorithm is used to solve for the homography transformation matrix that describes the spatial mapping relationship between different video frames. This matrix accurately reflects the differences in shooting angle, scale, and rotation angle between different video frames. The solved homography transformation matrix is applied to the spatial coordinate transformation of non-reference video frames to remap the pixel positions of the video frames, enabling precise spatial alignment of video frames acquired by different devices in the global spatial coordinate system. This eliminates problems such as image misalignment, stretching, and rotation caused by different shooting device positions, completing image registration for all video frames.
[0109] For example, to address the registration requirements of the foundation pit area and tower crane area at the construction site, the video frame of the foundation pit area is used as the reference frame. The homography transformation matrix is calculated based on the coordinates of the matching feature point pairs. This matrix is then applied to the video frame of the tower crane area to perform a spatial transformation on its pixel positions. This ensures that the feature points of the tower crane support in the transformed tower crane area video frame are precisely aligned with the feature points of the protective frame in the foundation pit area video frame in terms of spatial position. The two video frames achieve spatial alignment without misalignment, thus completing the image registration.
[0110] S633. For the overlapping areas of the registered images, assign fusion weights according to the pixel positions and the distances to the centers of each source image, perform weighted average fusion, and obtain the fused image; Specifically, in step S633, for the multiple video images after registration, the overlapping areas between different images are first identified. These areas are the intersections of shooting perspectives from different devices and are the key part of panoramic image stitching. For each pixel in the overlapping area, a fusion weight is assigned according to the distance between the pixel position and the center of each source image, following the allocation rule that "the closer a pixel is to the center of a source image, the higher the fusion weight of that source image for that pixel; the farther away from the center of the source image, the lower the fusion weight." Based on the assigned weights, the pixel values in the overlapping area are calculated by weighted averaging to obtain the fused pixel value. The non-overlapping areas directly retain the pixel information of the original image, ultimately forming a single fused image, effectively avoiding problems such as image ghosting, uneven brightness, and obvious stitching marks in the overlapping areas.
[0111] For example, in the registered images of the foundation pit area and the tower crane area of the construction site, there is an overlapping area at the junction of the protective frame and the tower crane support. Pixels in this area that are closer to the center of the foundation pit area image are assigned a higher weight, and pixels that are closer to the center of the tower crane area image are assigned a higher weight. After weighted average fusion, the transition of the overlapping area is natural, which not only preserves the details of the foundation pit protective frame, but also clearly presents the characteristics of the tower crane support, without obvious splicing marks.
[0112] S634. Combine the fused multi-frame images to generate a panoramic monitoring image; Specifically, in step S634, all the fused images are combined with the global spatial coordinate system of the work area, and the fused images are systematically stitched and combined according to the actual deployment position and shooting angle of each multi-view camera device. This fills in the blank areas between different shooting angles, so that each fused image forms a continuous and complete picture in the global spatial coordinate system. The edge transition of the combined picture is optimized to ensure the visual consistency of the whole picture. Finally, a panoramic monitoring picture covering the entire work area is generated, realizing the full-area visualization of the work area without blind spots.
[0113] For example, in a construction site, the images of the foundation pit area, tower crane area, material storage area, and construction access area are stitched together sequentially according to the three-dimensional position of each device in the global spatial coordinate system. This fills in the visual gaps between the areas, optimizes the brightness and color transition at the edges of the image, and ultimately generates a panoramic monitoring image of the construction site covering all work areas such as the foundation pit, tower crane, material storage, and construction access. The image is continuous, complete, and without any stitching flaws.
[0114] This embodiment uses a scale-invariant feature transform algorithm to accurately match feature points of video frames from different devices. It achieves precise image registration through a homography transformation matrix, and assigns weights according to pixel positions to smoothly fuse overlapping areas and combine them into a panoramic image. This results in a panoramic image without stitching marks and with strong visual consistency, significantly improving the quality of panoramic monitoring images and the accuracy of full-area coverage.
[0115] Furthermore, in some embodiments, the intelligent monitoring method for job safety based on multi-source video also includes: S71. When a new multi-camera device is detected to be accessing the system, obtain the GPS location information of the new multi-camera device; Specifically, for step S71, the work safety monitoring system maintains real-time monitoring of the equipment access status. When a new multi-camera device is detected to be connected to the system's communication link, the system will automatically trigger the equipment information collection process, retrieve the GPS positioning module of the new device, and obtain its accurate real-time geographical location information. This information includes core spatial data such as the latitude, longitude, and elevation of the new device's deployment location, and the data format is consistent with the geographic coordinate system already adopted by the system, providing accurate and suitable basic data for subsequent spatial location calculations.
[0116] For example, in large open-pit mining areas, the system originally deployed multiple multi-camera devices to cover the mining face and main transportation channels. When a new multi-camera device is connected to the auxiliary material storage area of the mine, the system detects the connection signal in real time and immediately obtains the accurate GPS geographical location information of the new device's deployment point.
[0117] S72. Calculate the relative positions of the new multi-camera device and the existing equipment based on the global spatial coordinate system; Specifically, for step S72, using the established global spatial coordinate system of the work area as a unified benchmark, the acquired GPS geographical location information of the new multi-camera device is converted into three-dimensional coordinates in the global spatial coordinate system according to the mapping rules of the coordinate system, thus clarifying the specific location of the new device in the global space of the work area; then, through spatial geometric calculation algorithms, with the three-dimensional coordinates of the coordinate system of the existing multi-camera devices deployed in the surrounding area as a reference, the relative position parameters such as the three-dimensional spatial distance and shooting angle between the new device and the existing devices are accurately calculated, establishing the spatial relationship between the new device and the existing devices, and providing spatial location basis for the subsequent fusion and connection of video streams.
[0118] For example, the GPS information of the new equipment in the mining area's auxiliary material storage area is converted into three-dimensional coordinates (X8, Y8, Z8) in the mining area's global spatial coordinate system. With the coordinates (X5, Y5, Z5) of the existing transportation channel equipment in the surrounding area as a reference, the spatial distance between the two, the horizontal and vertical angles of the shooting angle are calculated, and the specific location of the new equipment in the mining area's global monitoring layout is determined, as well as the connection relationship between its shooting angle and the shooting angle of the existing equipment.
[0119] S73. Automatically integrate the video stream captured by the new multi-camera device into the existing panoramic monitoring screen; Specifically, for step S73, the system determines the spatial connection position of the video stream acquired by the new device in the existing panoramic monitoring screen based on the calculated relative position parameters. Then, it automatically performs feature point matching, image registration and other processing on the video stream acquired by the new device. Then, through a weighted fusion algorithm, it seamlessly connects the video screen of the new device to the corresponding position in the existing panoramic monitoring screen, completing the automatic integration of the video stream of the new device. The whole process does not require manual re-stitching of the panoramic screen of the entire work area. The original panoramic monitoring screen will be naturally expanded with the integration of the new device.
[0120] For example, after calculating the relative position of the new equipment in the auxiliary material storage area of the aforementioned mining area, it was determined that its video feed should be connected to the south side of the existing transportation channel entrance monitoring screen. The system automatically performs feature point matching and image registration on the video stream of the new equipment, and performs weighted fusion of its video feed with the existing transportation channel entrance screen. Finally, the video stream of the new equipment is seamlessly integrated into the existing panoramic monitoring screen of the mining area, and the panoramic screen naturally extends to the auxiliary material storage area, covering the original monitoring blind spot of the area.
[0121] In this embodiment, when a new multi-camera device is connected, its GPS information is automatically acquired and its relative position to existing equipment is calculated. Its video stream is seamlessly integrated into the existing panoramic monitoring screen without the need for manual re-stitching. This achieves dynamic expansion of the panoramic monitoring screen, improves the flexibility and scalability of the monitoring system, and can continuously fill the monitoring blind spots in the work area.
[0122] Furthermore, in some embodiments, the intelligent monitoring method for job safety based on multi-source video also includes: S81. Activate the supplementary lighting equipment to illuminate the monitored area; Specifically, in step S81, when the work area enters a low-light environment at night, or when the ambient light intensity is lower than a preset lighting threshold, the system automatically activates the supplementary lighting equipment deployed in the work area. The supplementary lighting equipment works in conjunction with the multi-camera device to provide uniform and soft supplementary lighting to the monitored area. During the supplementary lighting process, the equipment automatically adjusts the brightness and illumination range of the supplementary light according to the actual lighting conditions of the monitored area, ensuring that the supplementary lighting area covers the entire video acquisition range. This avoids both excessive supplementary lighting that leads to overexposure and glare, and insufficient supplementary lighting that fails to improve the low-light environment, thus providing stable and suitable lighting conditions for nighttime video acquisition.
[0123] For example, when working at a construction site at night, when the ambient light intensity drops below a preset threshold, the system automatically activates the supplementary lighting equipment deployed around the foundation pit and under the tower crane. The supplementary lighting equipment is adjusted to a suitable brightness to evenly illuminate the entire construction area, avoiding blurry images of the faces of construction workers and safety equipment due to insufficient light, and ensuring that video captures clear outlines of personnel and equipment features.
[0124] S82. Perform multi-frame noise reduction and high dynamic range image synthesis on the acquired nighttime video stream to obtain an enhanced image; Specifically, for step S82, for the video stream acquired after nighttime supplementary lighting, due to the low light environment and uneven supplementary lighting, the video image is prone to problems such as noise, imbalance of brightness and darkness, and loss of details. Therefore, the video stream needs to be optimized in two ways: First, multi-frame noise reduction processing, extracting multiple consecutive video frame images from the video stream, and using algorithms to superimpose and average the multiple frame images to suppress random noise in the image (such as image graininess and color spots) and restore image details; Second, high dynamic range image synthesis, acquiring video frames under different exposure parameters, and synthesizing multiple frames that are too dark, normal, and overexposed to balance the bright and dark areas of the image, improve the clarity of details in dark areas (such as clothing and safety equipment of workers), and suppress overexposure in bright areas (such as highlights in the supplementary lighting area), ultimately obtaining an enhanced image with clear image quality, balanced brightness and darkness, and effective noise suppression.
[0125] For example, in nighttime operation scenarios in industrial and mining areas, the captured nighttime video streams exhibit significant grainy noise, with overly bright supplementary lighting areas and overly dark corners far from these areas. By employing multi-frame noise reduction processing and superimposing 10 consecutive video frames, the graininess of the image is eliminated. Then, through high dynamic range image synthesis, three video frames with different exposures (dark exposure, normal exposure, and bright exposure) are combined to suppress highlights in the supplementary lighting areas and clearly reveal details in the dark corners, resulting in an enhanced image with improved image quality.
[0126] S83. Sharpen and suppress the background of the enhanced image to obtain an optimized image that highlights the outline of the human target; Specifically, for step S83, based on the obtained enhanced image, further image optimization processing is carried out to focus on the clear presentation of the human target. This is divided into two steps: First, image sharpening processing, which uses edge sharpening algorithms to enhance the edge contours of objects in the image (such as the body contours of workers, the edges of safety helmets, and the contours of reflective clothing), improving the clarity and recognizability of the contours and solving the problem of blurred contours in nighttime images; Second, background suppression processing, which uses image segmentation algorithms to identify and weaken irrelevant backgrounds in the enhanced image (such as walls, equipment, and the ground in the work area), reducing the brightness and contrast of the background area, while preserving and enhancing the image features of the human target area, allowing the human contour of the worker to stand out from the background, reducing the impact of background interference on subsequent recognition work, and finally obtaining an optimized image that highlights the human target contour.
[0127] For example, by processing the enhanced images of the mining area at night, the head and torso outlines of the workers are enhanced by sharpening algorithms, making the hemispherical outline of the safety helmet and the edge of the reflective vest clearer; by background suppression processing, the details of mining equipment and the ground in the image are weakened and their brightness is reduced, making the human outline of the workers more prominent in the image, and human targets can be quickly locked even in complex nighttime working environments.
[0128] S84. Target detection and security feature recognition analysis based on optimized images; Specifically, in step S84, the optimized image is used as the core analysis basis to carry out target detection and safety feature recognition analysis. The optimized image has prominent human target outlines, less noise, and clear details, which can provide accurate identification objects for target detection, making it easy to quickly identify all worker targets in the image; at the same time, the clear image details can make safety feature recognition more accurate, effectively identifying safety features such as the wearing status of safety helmets and reflective clothing of workers, ensuring that safety identification work can be carried out stably and accurately in nighttime environments, and obtaining reliable identification results.
[0129] For example, based on the optimized image highlighting human-shaped targets, the system quickly identified three mining workers in the image, accurately cropped the head and torso areas of each person, and clearly identified that two of them were wearing safety helmets and reflective vests, while one person was not wearing a safety helmet. The system completed target detection and safety feature recognition analysis, and the recognition results were accurate and reliable.
[0130] This embodiment improves the low-light environment at night by activating supplementary lighting equipment. After multi-frame noise reduction, high dynamic range synthesis and sharpening, and background suppression to optimize the nighttime video stream, target detection and safety feature recognition are carried out based on the optimized image. This effectively solves the problem of difficult recognition under low light conditions at night and achieves all-weather coverage of work safety monitoring.
[0131] Furthermore, in some embodiments, step S82, "performing multi-frame noise reduction and high dynamic range image synthesis on the acquired nighttime video stream," may specifically include: S821. Perform multi-frame noise reduction on the nighttime video stream to obtain a noise-reduced image sequence; Specifically, in step S821, for video streams acquired in low-light conditions at night, which are prone to image noise due to insufficient ambient light (such as graininess, random color spots, and bright / dark noise), multi-frame noise reduction processing is performed on the nighttime video stream. The algorithm extracts consecutive frames from the video stream, performs inter-frame matching, superposition, and averaging calculations on the pixels of each frame to suppress random noise while preserving the true details of the video image. This avoids problems such as image blurring and detail loss during the noise reduction process, ultimately resulting in a noise-reduced image sequence with effectively suppressed noise and a clean image. This sequence completely preserves the temporal and image information of the nighttime video stream.
[0132] For example, in a nighttime video stream from an open-pit mine, there is obvious black grain noise in the image. By performing multi-frame noise reduction on 20 consecutive frames of the video stream and averaging the inter-frame pixel stacking, grain noise is eliminated, resulting in a noise-reduced image sequence with no obvious noise and clear outlines of underground workers and equipment details.
[0133] S822. Select image frames with different exposure parameters from the denoised image sequence; Specifically, in step S822, from the image sequence obtained after multi-frame noise reduction, image frames with different exposure parameters are selected based on the exposure characteristics of the nighttime video capture. Due to these differences in exposure parameters, these image frames each present high-quality details in different areas of the nighttime scene: image frames with low exposure parameters retain complete details in dark areas without overexposure in bright areas; image frames with normal exposure parameters have balanced brightness and clear outlines of the main target; and image frames with high exposure parameters highlight bright features and clearly present details of reflective and luminous objects. The selection process ensures that image frames with each exposure parameter cover the key details of the nighttime scene, providing diverse image detail support for subsequent fusion processing.
[0134] For example, from the noise-reduced image sequence of the above-mentioned mining operation area, three image frames with different exposure parameters were selected: one low-exposure frame, which clearly preserves the details of the equipment deep in the underground roadway; one normal-exposure frame, which fully presents the body outline of the workers; and one high-exposure frame, which clearly shows the highly reflective stripe features of the workers' reflective clothing. Each of the three images has its own advantages in detail.
[0135] S823. Merge image frames with different exposure parameters to generate an enhanced image; Specifically, in step S823, a high dynamic range image fusion algorithm is used to perform fusion processing on the selected image frames with different exposure parameters. The algorithm accurately extracts high-quality details from each frame, integrates the dark details of low-exposure frames, the main outline details of normal-exposure frames, and the bright feature details of high-exposure frames, and simultaneously performs brightness and contrast equalization processing on each frame to suppress image defects in a single exposure frame, such as overall darkness in low-exposure frames and local overexposure in high-exposure frames. Finally, all high-quality details are fused into the same frame image to generate an enhanced image that is clean, has balanced brightness and darkness, and is rich in detail.
[0136] For example, by fusing three frames of low, medium, and high exposure images of a selected mining area, the details of the equipment in the tunnel in the low exposure frame, the outlines of the personnel in the normal exposure frame, and the features of the reflective clothing in the high exposure frame are extracted. After brightness equalization processing, the enhanced image generated clearly shows the outlines of the equipment and workers in the depths of the tunnel, as well as the features of the reflective clothing. The image is free of noise, overexposure, and dark areas, resulting in a significant improvement in overall image quality.
[0137] This embodiment first performs multi-frame noise reduction on the nighttime video stream to suppress image noise, and then selects image frames with different exposure parameters for fusion. The resulting enhanced image has low noise, balanced brightness and darkness and rich details, effectively solving the problems of dark nighttime video images, high noise and loss of details, and providing a high-quality base image for subsequent image processing.
[0138] Furthermore, in some embodiments, step S81, "activating the supplementary lighting device to illuminate the monitored area," further includes: S811. Based on the distance between the multi-camera device and the monitored target, dynamically adjust the illumination brightness and beam focusing area of the supplementary lighting device.
[0139] Furthermore, in some embodiments, step S811, "dynamically adjusting the illumination brightness and beam focusing area of the supplementary lighting device according to the distance between the multi-view camera device and the monitored target," may specifically include: S8111. Based on the bounding box coordinates of the personnel target and the preset camera parameters, estimate the approximate distance between the personnel target and the multi-view camera device in the current frame; Specifically, in step S8111, after obtaining the bounding box coordinates of the person target in the video frame image, distance estimation is performed based on these coordinates and the preset camera parameters of the multi-view camera device. The bounding box coordinates reflect the pixel size of the person target in the video frame. The preset camera parameters include fixed parameters such as focal length, imaging sensor size, and lens angle of view. Through a spatial geometric conversion algorithm, the pixel size of the person target, the camera parameters, and the conventional dimensions of the actual human body are correlated and calculated to accurately estimate the approximate actual distance between the person target and the multi-view camera device in the current frame. This distance provides a core quantitative basis for subsequent adjustment of the supplementary lighting parameters.
[0140] For example, in a video frame of a construction site, the bounding box of a worker is identified, with a pixel height of 200 pixels. Combined with the preset camera parameters of the multi-view camera device, such as a focal length of 8mm and an imaging sensor size of 1 / 2.8 inches, as well as the reference value of a normal human height, the approximate distance between the worker and the multi-view camera device is estimated to be 6 meters through geometric conversion.
[0141] S8112. Based on the approximate distance, consult the preset distance and brightness correspondence table to calculate the target illumination brightness value of the supplementary lighting device; Specifically, for step S8112, the system pre-establishes and stores a distance-brightness correspondence table based on the video acquisition requirements of different work scenarios and the lighting characteristics of the supplementary lighting equipment. This table clearly defines the optimal lighting brightness reference value for the supplementary lighting equipment at different monitoring distances, and the reference value has been verified through actual scenario testing to adapt to the video acquisition lighting requirements at the corresponding distance. The estimated approximate distance is used as the query basis, and the basic brightness value at that distance is matched in the correspondence table. Then, a small correction calculation is performed based on the actual lighting conditions of the current environment to finally determine the target lighting brightness value of the supplementary lighting equipment.
[0142] For example, in the table of distance and brightness correspondence in the above construction site scenario, the basic brightness value corresponding to a distance of 6 meters is 400 lux. In the current nighttime environment, there is no other auxiliary lighting, so no additional correction is needed. The target lighting brightness value of the supplementary lighting equipment is directly calculated to be 400 lux. If there is weak site lighting in the environment, the basic brightness value can be corrected to 350 lux as the target value.
[0143] S8113. Control the drive circuit of the supplementary lighting device to adjust the output brightness to the target illumination brightness value; Specifically, in step S8113, the drive circuit of the supplementary lighting device is the core execution component for regulating its illumination brightness. It can achieve precise brightness adjustment by changing parameters such as power supply current and power. The system converts the calculated target illumination brightness value into a corresponding electrical signal control command and sends this command to the drive circuit of the supplementary lighting device. After receiving the command, the drive circuit automatically adjusts its own power supply parameters, changing the output power of the light-emitting element of the supplementary lighting device, thereby precisely adjusting the actual illumination brightness of the supplementary lighting device to the target illumination brightness value, achieving quantification and precise control of the illumination brightness.
[0144] For example, for the target lighting brightness value of 400 lux, the system sends an adjustment command to the drive circuit of the LED supplementary lighting device. The drive circuit adjusts the power supply current from the initial 0.8A to 1.2A, so that the output power of the light-emitting element of the supplementary lighting device matches the target brightness, and the actual lighting brightness is accurately reached to 400 lux.
[0145] S8114. Synchronously control the optical lens or reflector mechanism of the supplementary lighting device to match the focusing area of the beam with the estimated target distance; Specifically, in step S8114, while adjusting the illumination brightness of the supplementary lighting equipment, the system simultaneously sends a control command to the optical lens or reflector mechanism of the supplementary lighting equipment. By controlling the mechanical movement of this mechanism, the divergence angle of the beam of the supplementary lighting equipment is changed, thereby adjusting the focusing area of the beam so that the focusing area accurately matches the estimated target distance. For personnel targets at close range, the control mechanism increases the beam divergence angle, making the beam form a wide focusing area that covers the working movement range of the personnel target; for personnel targets at long distances, the control mechanism decreases the beam divergence angle, making the beam form a narrow focusing area, improving the concentration of illumination, and allowing the limited illumination to be accurately projected onto the personnel targets at long distances.
[0146] For example, for the target distance of 6 meters, the system synchronously controls the optical lens mechanism of the supplementary lighting equipment to adjust the beam divergence angle to 30°, forming a focusing area adapted to the 6-meter distance, accurately covering the working range of the operator; if the distance of a certain person target is estimated to be 10 meters, the beam divergence angle is adjusted to 15° to form a narrow focusing area, allowing the light to be accurately focused on the distant person target.
[0147] This embodiment combines the bounding box coordinates of the personnel target with preset camera parameters to accurately estimate the target distance, calculates the target brightness through a distance-brightness correspondence table, and then synchronously controls the drive circuit and optical mechanism of the supplementary lighting device to adjust the brightness and beam focusing area, thereby achieving accurate and dynamic adaptation of supplementary lighting parameters, making supplementary lighting in low-light environments more targeted and efficient.
[0148] Furthermore, in some embodiments, edge computing nodes are used to encode and compress multi-view video data, cache it locally, and route it to a cloud analytics platform.
[0149] Specifically, in this embodiment, the edge computing node serves as the core data processing and communication hub between the multi-view camera device and the cloud analytics platform. It is responsible for encoding and compressing the multi-view video data collected by the multi-view camera device, caching it locally, and routing it to the cloud analytics platform. The edge computing node is equipped with a high-efficiency video encoding algorithm. While ensuring the video image clarity and detail integrity meet the subsequent cloud analysis requirements, it removes redundant information and optimizes the data format of the original video data, significantly reducing its size and creating lightweight, compressed video data, thus laying the foundation for efficient subsequent transmission.
[0150] During video data transmission, in the event of network fluctuations, network interruptions, or unforeseen circumstances such as the cloud analytics platform being temporarily unable to receive data due to server maintenance or high load, edge computing nodes will automatically store the encoded and compressed multi-view video data temporarily in their local cache space, preventing data loss due to transmission link anomalies. Once the network returns to normal and the cloud analytics platform resumes its data receiving capability, the edge computing nodes will immediately resume forwarding the cached video data to the cloud, ensuring continuous data transmission. Simultaneously, because the cached data is lightweight and encoded, it effectively saves local storage resources on the edge computing nodes, preventing storage overflow.
[0151] Edge computing nodes serve as dedicated data transmission hubs between front-end acquisition devices and cloud analytics platforms. They incorporate pre-defined routing and communication protocols, ensuring precise routing and forwarding of video data. Edge computing nodes identify and schedule encoded and compressed multi-view video data suitable for transmission. Following pre-defined communication protocols and fixed routing paths, they stably, efficiently, and accurately forward the data to the corresponding cloud analytics platforms. Simultaneously, they can uniformly schedule and orderly forward video data from multiple multi-camera devices, preventing confusion and mistransmission of video data from different perspectives and areas during transmission. This ensures that the cloud analytics platforms receive the corresponding video data according to region and time sequence.
[0152] In this embodiment, edge computing nodes encode and compress multi-view video data, significantly reducing bandwidth consumption for data transmission; local caching prevents data loss due to network or cloud failures; and routing and forwarding enables orderly and accurate transmission of video data. The three functions work together to ensure efficient, stable, and complete transmission of video data from the front end to the cloud, providing reliable data support for cloud analysis.
[0153] Furthermore, in some embodiments, step S4, "generating alarm information and equipment control instructions based on the analysis report," may specifically include: When a high incidence of violations is detected in a specific area, control commands are generated to adjust at least one of the following: the gimbal angle, focal length, or lighting parameters of the multi-camera device associated with that specific area.
[0154] Specifically, the analysis report containing the results of violation identification is first subjected to in-depth data analysis. Core data such as the location, type, frequency, and time distribution of all violations recorded in the report are extracted. Through location-based violation frequency statistics and trend analysis, specific areas with high incidence of violations are identified, clarifying the types of violations and peak time periods within those areas. This process uses the geographical zoning of the work area and monitoring points as the basis for accurately locating areas where violations are concentrated, providing clear target areas for subsequent equipment control.
[0155] Based on the identified high-incidence areas of violations, and combined with information such as the deployment locations and monitoring coverage of multi-camera devices in the work safety monitoring system, the multi-camera devices directly associated with the high-incidence areas are accurately located, i.e., the multi-camera devices currently responsible for video acquisition and image monitoring in the area. The device number, deployment location, and other information are clearly identified to ensure that subsequent control commands can be accurately issued to the corresponding devices and avoid device mismatch.
[0156] Targeting the types of violations and monitoring shortcomings in specific areas with high rates of misconduct, and combining the adjustable parameters of multi-camera devices, control commands are generated to adjust at least one of the following parameters: pan-tilt angle, focal length, or supplementary lighting parameters of the associated multi-camera devices in that area. The adjustment direction is perfectly aligned with the actual monitoring needs of the area: if there are blind spots in the high-incidence area, the pan-tilt angle is adjusted to cover the blind spots; if the field of view is too wide, making it difficult to capture details of violations, the focal length is adjusted to magnify the target area; if insufficient lighting makes it difficult to identify violations, the supplementary lighting parameters are adjusted to optimize the lighting effect.
[0157] The generated targeted equipment control commands are accurately sent to the multi-camera devices associated with high-incidence violation areas through the communication link of the operation safety monitoring system. After receiving the commands, the devices automatically perform parameter adjustment operations, dynamically changing the pan-tilt angle, focal length, or supplementary lighting parameters. The adjusted parameters will continue to adapt to the monitoring needs of the area until the frequency of violations in the area drops back to normal levels.
[0158] After identifying high-incidence areas of violations in the work area, this embodiment generates targeted control commands to adjust the pan-tilt angle, focal length, or supplementary lighting parameters of the associated multi-camera device. This effectively compensates for the monitoring shortcomings in high-incidence areas, improves the comprehensiveness and clarity of video acquisition in the area, strengthens the precise supervision of high-incidence areas of violations, and helps reduce the occurrence of violations in the area.
[0159] Furthermore, in some embodiments, the user terminal also provides a user interface configured to perform at least one of the following operations: View real-time panoramic surveillance footage and footage from specific cameras; Receive and display violation alerts, which include at least one of the following: violation type, time of occurrence, location, and screenshot. Query historical violation records and statistical reports; Remotely control the pan-tilt head, focal length, and lighting equipment of a multi-camera system.
[0160] Specifically, the user terminal in this embodiment also provides a user interface, which integrates the core operation functions of the entire process of work safety monitoring and supports at least one operation such as viewing monitoring screens, receiving alarm information, querying historical data, and remotely controlling equipment.
[0161] The user interface features multi-dimensional video viewing capabilities, with core support for switching between real-time panoramic monitoring footage and specific camera feeds. The main view displays a complete real-time panoramic view of the work area, allowing managers to intuitively grasp the overall operational status without blind spots. Simultaneously, the interface includes a map showing the distribution of multi-camera locations within the work area. Managers can easily access real-time feeds from any multi-camera by clicking on locations or entering device numbers. It also supports auxiliary operations such as zooming in, zooming out, and freezing the frame, meeting both overall monitoring needs and detailed local viewing requirements.
[0162] The user interface is linked in real time with the cloud analytics platform, enabling real-time reception and display of violation alarm information. When the cloud analytics platform identifies a violation in the work area and generates an alarm, it immediately pushes the alarm information to the user interface. The interface will alert management personnel through pop-ups, red highlighting, and alarm sounds. The pushed alarm information includes rich core dimensions, covering at least the type of violation (such as not wearing a safety helmet, unauthorized entry into a high-risk restricted area, not wearing reflective clothing, etc.), the precise time of occurrence of the violation, the specific location of the violation (such as the east side of the foundation pit, the north entrance of the transportation channel, 3 meters below the tower crane, etc.), and real-time screenshots of the violation. In some scenarios, short video clips of the violation can also be attached, allowing management personnel to quickly and comprehensively grasp the key information of the violation without having to view the monitoring screen.
[0163] The user interface incorporates historical violation data query and statistical report generation functions, providing data support for safety management analysis. The interface supports multi-dimensional filtering of historical violation records, with selectable filtering conditions including time (day / week / month / year / custom time period), violation type, location, and responsible work team. After managers set the filtering conditions according to their needs, the interface quickly displays all historical violation records that meet the criteria. Each record contains complete violation information and subsequent handling results, and records can be exported and printed. Simultaneously, the system automatically performs statistical analysis on historical violation data, generating visual statistical reports. These reports can use bar charts, line charts, pie charts, etc., to intuitively display data such as violation frequency, violation trends, and the proportion of various violations from different dimensions, allowing managers to clearly understand the shortcomings in safety management of the work area.
[0164] The user interface provides a remote control entry point for monitoring equipment, enabling convenient remote control of multi-camera devices and supplementary lighting equipment without requiring on-site personnel. The interface features built-in visual control buttons and parameter adjustment sliders. Managers can send control commands to the equipment via simple clicks and drags to remotely adjust the pan / tilt-zoom angle and optical / digital focal length of the multi-camera devices. They can also remotely control the on / off state, illumination brightness, and beam focus area of the supplementary lighting equipment. Equipment parameters can be adjusted in real time according to monitoring needs, ensuring the monitored image remains clear and effective at all times.
[0165] The user interface of the user terminal provided in this embodiment integrates functions such as viewing multi-dimensional monitoring screens, receiving multi-information dimension violation alarms, querying historical violation data and statistical reports, and remote control of monitoring equipment. It realizes the integration of operation safety monitoring, making the operation of monitoring personnel more convenient and the information acquisition more comprehensive, and greatly improving the operation efficiency and management efficiency of operation safety monitoring.
[0166] To better understand the intelligent monitoring method for work safety based on multi-source video provided in this embodiment, this embodiment also provides a construction site violation identification and monitoring system based on multi-view cameras. This system includes multi-view cameras, a cloud server, and terminal devices, wherein: A multi-view camera system comprises at least one bullet camera and at least two PTZ cameras, used for comprehensive monitoring of the construction site environment. Each multi-view camera integrates multiple independent image acquisition modules, each equipped with an independent lens and image sensor to acquire image data from different perspectives. The multi-view camera also includes an image processing unit for preliminary processing of the acquired image data, such as distortion correction and image enhancement.
[0167] The cloud server includes an image analysis module, a human detection module, a feature extraction module, a classification and recognition module, and a database module. The image analysis module analyzes image data uploaded from multiple cameras. The human detection module uses the YOLOv8 algorithm to detect people in the images and outputs their coordinates and number. The feature extraction module extracts features from the head and torso areas of each person. The classification and recognition module identifies safety helmets and reflective vests based on the extracted features. The database module stores standard feature data for safety helmets and reflective vests.
[0168] The terminal device includes a display screen, a network interface, a storage unit, and a user interface. The terminal device communicates with the cloud server via the network interface to receive classification and recognition results, which are then displayed in real time on the screen. The user interface provides basic setup functions, such as device adjustment and data backup.
[0169] The system also includes an intelligent analysis module, which comprises an environmental assessment submodule and a behavior analysis submodule. The environmental assessment submodule dynamically adjusts the camera's exposure parameters and focal length based on real-time ambient light intensity, temperature, and other factors to adapt to different lighting conditions. The behavior analysis submodule analyzes the image sequences captured by the multi-camera system to identify the safety status of construction workers and generates timely alarm information when violations are detected.
[0170] The system also supports multi-device clustering, using a GPS module to acquire the location information of each device in real time, establishing a global coordinate system for the construction site, and calculating the relative positions and angles between devices. The system employs the SIFT algorithm for feature point matching and weighted fusion, enabling image fusion between multiple devices, and supports dynamic calibration and integration into the panorama when devices are added.
[0171] In night mode, the system employs multi-frame noise reduction technology to denoise the image, combined with HDR algorithms to expand the dynamic range and enhance image details. The system uses an intelligent control module to adjust the brightness and focus area of the white light, achieving adaptive lighting for people at different distances. It also supports nighttime human figure optimization, using sharpening and background suppression techniques to highlight human targets.
[0172] In a specific embodiment, this embodiment also provides a construction site violation identification and monitoring system based on multi-view cameras. The system includes multi-view cameras, a cloud server, and terminal devices.
[0173] The multi-view camera system includes one bullet camera and two PTZ cameras for comprehensive monitoring of the construction site environment. Internally, the multi-view camera integrates three independent image acquisition modules, each equipped with an independent 1080P resolution CMOS sensor and a wide-angle lens to acquire image data from different perspectives. The multi-view camera also includes an image processing unit based on an ARM Cortex-A72 processor for preliminary processing of the acquired image data, such as distortion correction and image enhancement.
[0174] The cloud server includes a GPU-accelerated image analysis module, a human detection module using the YOLOv8 algorithm, a feature extraction module, a classification and recognition module, and a database module. The image analysis module decodes and normalizes image data uploaded from the multi-camera system. The human detection module uses the YOLOv8 algorithm to detect people in the images and outputs their coordinates and number. The feature extraction module extracts features from the head region (top 1 / 5 of the bounding box) and torso region (middle 2 / 5) of each person. The classification and recognition module identifies helmets based on the extracted head color features (RGB values R>180, G>180, or B>180) and shape features (circular or hemispherical), and identifies reflective vests based on the reflective strip features of the torso region (grayscale value>200). The database module stores standard feature data for helmets and reflective vests.
[0175] The terminal device includes a 10.1-inch IPS display, a gigabit Ethernet port, a 128GB solid-state drive, and a touchscreen interface. The terminal device communicates with the cloud server via the gigabit Ethernet port, receives the classification and recognition results, and displays them on the screen in real time. The interface provides basic functions such as device adjustment, data backup, and system settings.
[0176] The system also includes a deep learning-based intelligent analysis module, comprising an environmental assessment submodule and a behavior analysis submodule. The environmental assessment submodule dynamically adjusts the camera's exposure time (30~1 / 1000 seconds) and sensitivity (ISO100~ISO51200) parameters based on real-time ambient light intensity (300~1000 lux), temperature (-20~60℃), and other factors to adapt to different ambient lighting conditions. The behavior analysis submodule analyzes the image sequences captured by the multi-camera system to identify the safety status of construction workers and promptly generates alarm information when violations such as not wearing safety helmets or reflective vests are detected.
[0177] The system also supports multi-device clustering, using the built-in GPS module of each device to acquire longitude, latitude, and altitude information in real time, establishing a global coordinate system for the construction site, and calculating the relative positions and angles between devices. The system uses the SIFT algorithm for feature point matching and weighted fusion to achieve image fusion between multiple devices, and supports dynamic calibration and integration into the panorama when devices are added.
[0178] In night mode, the system employs a multi-frame noise reduction technique of 5 frames of mean filtering + 3 frames of median filtering to denoise the image, combined with HDR algorithms to expand the dynamic range and enhance image details. The system uses an intelligent control module to adjust the brightness (0~1000 lumens) and focus area of the white light to achieve adaptive lighting for people at different distances, and supports a night-time human figure optimization function. The Sobel operator is used to sharpen the edges of human figures, improving the clarity of the outline, while reducing the brightness of non-target areas to highlight human figures.
[0179] In a specific embodiment, this embodiment also provides a construction site violation identification and monitoring system based on multi-view cameras. The system includes multi-view cameras, a cloud server, and terminal devices.
[0180] The multi-camera system includes one 4K resolution bullet camera and two 1080P resolution PTZ cameras for comprehensive monitoring of the construction site environment. Internally, the multi-camera system integrates three independent image acquisition modules, each equipped with an independent 1 / 2.8-inch CMOS sensor and an 8-24mm wide-angle lens to acquire image data from different perspectives. The system also includes an image processing unit based on an NVIDIA Jetson AGX Xavier processor for preliminary processing of the acquired image data, such as distortion correction and image enhancement.
[0181] The cloud server includes a GPU-accelerated image analysis module, a human detection module using the YOLOv8 algorithm, a feature extraction module, a classification and recognition module, and a database module. The image analysis module decodes and normalizes image data uploaded from the multi-camera system. The human detection module uses the YOLOv8 algorithm to detect people in the images and outputs their coordinates and number. The feature extraction module extracts features from the head region (top 1 / 5 of the bounding box) and torso region (middle 2 / 5) of each person. The classification and recognition module identifies helmets based on the extracted head color features (RGB values R>180, G>180, or B>180) and shape features (circular or hemispherical), and identifies reflective vests based on the reflective strip features of the torso region (grayscale value>200). The database module stores standard feature data for helmets and reflective vests.
[0182] The terminal device includes a 15-inch OLED display, a 10Gbps fiber optic interface, a 512GB NVMe solid-state drive, and a touchscreen user interface. The terminal device communicates with the cloud server via the 10Gbps fiber optic interface to receive classification and identification results, which are then displayed in real-time on the screen. The user interface provides basic functions such as device adjustment, data backup, and system settings.
[0183] The system also includes a deep learning-based intelligent analysis module, which comprises an environmental assessment submodule and a behavior analysis submodule. The environmental assessment submodule dynamically adjusts the camera's exposure time and sensitivity parameters based on real-time ambient light intensity, temperature, and other factors to adapt to different lighting conditions. The behavior analysis submodule analyzes the image sequences captured by the multi-camera system to identify the safety status of construction workers and promptly generates alarm information when violations such as not wearing safety helmets or reflective vests are detected.
[0184] The system also supports multi-device clustering, using the built-in GPS module of each device to acquire longitude, latitude, and altitude information in real time, establishing a global coordinate system for the construction site, and calculating the relative positions and angles between devices. The system uses the SIFT algorithm for feature point matching and weighted fusion to achieve image fusion between multiple devices, and supports dynamic calibration and integration into the panorama when devices are added.
[0185] In night mode, the system employs a multi-frame noise reduction technique of 3 frames median filtering + 2 frames mean filtering to denoise the image, combined with HDR algorithms to expand the dynamic range and enhance image details. The system uses an intelligent control module to adjust the brightness (0~2000 lumens) and focus area of the white light to achieve adaptive lighting for people at different distances, and supports a nighttime human figure optimization function, highlighting human targets through sharpening and background suppression techniques.
[0186] It should be understood that, although Figure 2 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 2 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0187] To facilitate better implementation of the intelligent monitoring method for work safety based on multi-source video in the embodiments of this application, this application also provides an intelligent monitoring device for work safety based on multi-source video, which is based on the aforementioned intelligent monitoring method for work safety. The meanings of the terms used are the same as in the aforementioned intelligent monitoring method for work safety based on multi-source video, and specific implementation details can be found in the descriptions in the method embodiments.
[0188] Please see Figure 5 , Figure 5 The structural diagram of the intelligent monitoring device for work safety based on multi-source video provided in this application embodiment can specifically include: The data acquisition module 201 is used to synchronously acquire multi-view video data of the work area and its associated environmental data, and to perform time alignment and data association processing on the multi-view video data and environmental data to form monitoring associated data; Data upload module 202 is used to upload monitoring-related data to the cloud analysis platform; The cloud analysis module 203 is used to analyze multi-view video data in the monitoring-related data through the cloud analysis platform, and to identify the safety equipment wearing status and behavior of the workers by using a preset target detection model and safety feature recognition model, and to generate an analysis report containing the results of violation identification. Control module 204 is used to generate alarm information and equipment control commands based on the analysis report; The push module 205 is used to push alarm information, equipment control commands and real-time monitoring screens to the user terminal.
[0189] The multi-source video-based intelligent monitoring device for work safety provided in this embodiment can promptly detect safety hazards in operations and promote rapid response by constructing an integrated intelligent monitoring system for work safety. This enables intelligent and real-time monitoring and management of work safety, thereby improving the comprehensiveness, accuracy, and response efficiency of work safety monitoring.
[0190] Specific limitations regarding the intelligent monitoring device for work safety based on multi-source video can be found in the limitations of the intelligent monitoring method for work safety based on multi-source video mentioned above, and will not be repeated here. Each module in the aforementioned intelligent monitoring device for work safety based on multi-source video can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0191] Furthermore, embodiments of this application also provide an electronic device, such as... Figure 6 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically: The electronic device may include components such as a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, a power supply 303, and an input unit 304. Those skilled in the art will understand that... Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 301 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 302, and by calling data stored in the memory 302, thereby providing overall monitoring of the electronic device. Optionally, the processor 301 may include one or more processing cores; preferably, the processor 301 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 301.
[0192] The memory 302 can be used to store software programs and modules. The processor 301 executes various functional applications and a multi-source video-based intelligent monitoring method for job safety by running the software programs and modules stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 302 may also include a memory controller to provide the processor 301 with access to the memory 302.
[0193] The electronic device also includes a power supply 303 that supplies power to various components. Preferably, the power supply 303 can be logically connected to the processor 301 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 303 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0194] The electronic device may also include an input unit 304, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0195] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 301 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 302 according to the following instructions, and the processor 301 runs the applications stored in the memory 302 to realize various functions, as follows: The system synchronously collects multi-view video data and associated environmental data of the work area, performs time alignment and data association processing on the multi-view video data and environmental data to form monitoring-related data; the monitoring-related data is uploaded to the cloud analysis platform; the cloud analysis platform analyzes the multi-view video data in the monitoring-related data through preset target detection models and safety feature recognition models to identify the safety equipment wearing status and behavior of the workers, and generates an analysis report containing the results of violation identification; alarm information and equipment control commands are generated based on the analysis report; the alarm information, equipment control commands, and real-time monitoring screens are pushed to the user terminal.
[0196] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0197] This application embodiment constructs an integrated intelligent monitoring system for operational safety, which can promptly detect safety hazards in operations and promote rapid response, thereby achieving intelligent and real-time monitoring and management of operational safety, and improving the comprehensiveness, accuracy, and response efficiency of operational safety monitoring.
[0198] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0199] Therefore, embodiments of this application provide a storage medium storing multiple instructions that can be loaded by a processor to execute steps in any of the multi-source video-based intelligent monitoring methods for job safety provided in embodiments of this application. For example, the instructions can execute the following steps: The system synchronously collects multi-view video data and associated environmental data of the work area, performs time alignment and data association processing on the multi-view video data and environmental data to form monitoring-related data; the monitoring-related data is uploaded to the cloud analysis platform; the cloud analysis platform analyzes the multi-view video data in the monitoring-related data through preset target detection models and safety feature recognition models to identify the safety equipment wearing status and behavior of the workers, and generates an analysis report containing the results of violation identification; alarm information and equipment control commands are generated based on the analysis report; the alarm information, equipment control commands, and real-time monitoring screens are pushed to the user terminal.
[0200] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0201] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0202] Since the instructions stored in the storage medium can execute the steps in any of the intelligent monitoring methods for work safety based on multi-source video provided in the embodiments of this application, the beneficial effects that any of the intelligent monitoring methods for work safety based on multi-source video provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0203] The above provides a detailed description of a method, device, and storage medium for intelligent monitoring of work safety based on multi-source video, as provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for intelligent monitoring of work safety based on multi-source video, applied to a work safety monitoring system, characterized in that, The operation safety monitoring system includes a multi-camera device deployed in the operation area, an edge computing node communicatively connected to the multi-camera device, a cloud analysis platform communicatively connected to the edge computing node, and a user terminal communicatively connected to the cloud analysis platform. The multi-view camera device is used to collect multi-view video data of the work area; The intelligent monitoring method for work safety based on multi-source video includes: Simultaneously collect multi-view video data of the work area and its associated environmental data, and perform time alignment and data association processing on the multi-view video data and the environmental data to form monitoring associated data; The monitoring-related data is uploaded to the cloud-based analysis platform; The cloud-based analysis platform analyzes multi-view video data from the monitoring and related data using a preset target detection model and safety feature recognition model to identify the safety equipment wearing status and behavior of the workers and generate an analysis report containing the results of violation identification. Based on the analysis report, alarm information and equipment control commands are generated; The alarm information, the device control commands, and the real-time monitoring screen are pushed to the user terminal.
2. The intelligent monitoring method for work safety based on multi-source video according to claim 1, characterized in that, The cloud-based analysis platform analyzes multi-view video data from the monitoring and related data using a preset target detection model and safety feature recognition model to identify the safety equipment wearing status and behavior of workers, including: The target detection model is used to detect people in video frame images, and the bounding box coordinates of each person target are obtained; For each detected person target, the head region image and the torso region image are extracted according to the bounding box coordinates corresponding to the person target; The color and shape features of the head region image are analyzed to determine whether a preset style of safety helmet is being worn, and a first determination result is obtained. The image of the torso region is subjected to high reflectivity feature analysis to determine whether the wearer is wearing a reflective vest of a preset standard, and a second judgment result is obtained. Based on the first and second judgment results, the safety equipment wearing status of the personnel target is determined. If any equipment is detected not being worn as required, it is marked as a violation.
3. The intelligent monitoring method for work safety based on multi-source video according to claim 2, characterized in that, The method further includes: The results of security equipment identification of the same person target in multiple consecutive video images are fused and calculated. Based on the fusion calculation results, determine whether the consistency of the recognition results of multiple frames exceeds a preset threshold; When the consistency exceeds a preset threshold, the final safety equipment wearing status of the personnel target is determined.
4. The intelligent monitoring method for work safety based on multi-source video according to claim 3, characterized in that, The fusion calculation of the security equipment identification results of the same person target in multiple consecutive video images includes: Tracking the same person target across frames in a continuous sequence of video frames; For each tracked target ID, record the recognition results and confidence levels of the safety helmet and reflective vest for each target ID in N consecutive frames.
5. The intelligent monitoring method for work safety based on multi-source video according to claim 3, characterized in that, When the consistency exceeds a preset threshold, determining the final safety equipment wearing status of the personnel target includes: Statistical analysis was performed on the N identification results of the safety helmet and reflective vest corresponding to each target ID; If any device is identified as being worn more frequently than a preset first frequency threshold, then it is determined that the device has been worn. If any equipment is identified as not being worn more than a preset second frequency threshold, then it is determined that the equipment is not being worn. Based on the independent assessment results of the safety helmet and reflective vest, the final safety equipment wearing status of the personnel target is determined.
6. The intelligent monitoring method for work safety based on multi-source video according to claim 2, characterized in that, The method further includes: Real-time monitoring of light intensity and direction in associated environmental data; When a backlit scene is determined based on the light intensity and the light direction, the exposure parameters of the video capture are automatically adjusted, and local image enhancement processing is performed on the head area image and the torso area image.
7. The intelligent monitoring method for work safety based on multi-source video according to claim 6, characterized in that, When a backlighting scene is determined based on the light intensity and light direction, the automatic adjustment of the exposure parameters for video capture and the local image enhancement processing of the head region image and the torso region image include: Determine whether it is a severe backlighting scene based on the angle between the direction of the light and the direction of the camera, as well as the brightness contrast between the foreground target and the background. If the scene is determined to be a severe backlighting scene, the multi-camera device is controlled to increase the exposure compensation value and reduce the sensitivity of the image sensor in order to suppress background overexposure and improve the details of the foreground target. In the image processing unit, for the adjusted acquired image, local contrast enhancement algorithm and edge sharpening algorithm are applied to the head region and torso region of the personnel target to enhance the edge contour of the safety helmet and the highly reflective strip features of the reflective vest.
8. The intelligent monitoring method for work safety based on multi-source video according to claim 1, characterized in that, The method further includes: Obtain the GPS location information of at least two of the multi-view camera devices in the operation safety monitoring system; A global spatial coordinate system for the work area is established based on the GPS geographic location information; Feature point matching and image registration are performed on video streams collected from different devices, and panoramic monitoring images are generated based on the registration relationship and weighted fusion algorithm.
9. The intelligent monitoring method for work safety based on multi-source video according to claim 8, characterized in that, The process of performing feature point matching and image registration on video streams acquired from different devices, and generating panoramic monitoring images based on registration relationships and weighted fusion algorithms, includes: The scale-invariant feature transform algorithm is used to extract and match feature points from synchronized video frames acquired from different devices; Based on the matched feature point pairs, the homography transformation matrix between images is solved to complete image registration; For the overlapping areas of the registered images, fusion weights are assigned based on pixel positions and distances from the centers of each source image, and a weighted average fusion is performed to obtain the fused image. The fused multi-frame images are combined to generate the panoramic monitoring screen.
10. The intelligent monitoring method for work safety based on multi-source video according to claim 8, characterized in that, The method further includes: When a new multi-camera device is detected to be connected to the system, its GPS location information is obtained. Based on the global spatial coordinate system, calculate the relative positions of the new multi-view camera device and the existing equipment; The video streams captured by the new multi-camera device are automatically integrated into the existing panoramic surveillance footage.
11. The intelligent monitoring method for work safety based on multi-source video according to claim 1, characterized in that, The method further includes: Activate the supplemental lighting equipment to illuminate the monitored area; Multi-frame noise reduction and high dynamic range image synthesis are performed on the acquired nighttime video stream to obtain an enhanced image; The enhanced image is then sharpened and background suppressed to obtain an optimized image that highlights the outline of the human-shaped target. The target detection and security feature recognition analysis are performed based on the optimized image.
12. The intelligent monitoring method for work safety based on multi-source video according to claim 11, characterized in that, The process of performing multi-frame noise reduction and high dynamic range image synthesis on the acquired nighttime video stream includes: The nighttime video stream is subjected to multi-frame noise reduction processing to obtain a noise-reduced image sequence; From the denoised image sequence, select image frames with different exposure parameters; The image frames with different exposure parameters are fused together to generate an enhanced image.
13. The intelligent monitoring method for work safety based on multi-source video according to claim 11, characterized in that, The activation of the supplementary lighting device to illuminate the monitored area also includes: The illumination brightness and beam focusing area of the supplementary lighting device are dynamically adjusted according to the distance between the multi-camera device and the monitored target.
14. The intelligent monitoring method for work safety based on multi-source video according to claim 13, characterized in that, The step of dynamically adjusting the illumination brightness and beam focusing area of the supplementary lighting device based on the distance between the multi-view camera device and the monitored target includes: Based on the bounding box coordinates of the target and preset camera parameters, estimate the approximate distance between the person target and the multi-view camera device in the current frame; Based on the approximate distance, a preset distance-brightness correspondence table is consulted to calculate the target illumination brightness value of the supplementary lighting device; The drive circuit of the supplementary lighting device is controlled to adjust the output brightness to the target illumination brightness value; The optical lens or reflector mechanism of the supplementary lighting device is synchronously controlled to match the focusing area of the light beam with the estimated target distance.
15. A work safety intelligent monitoring device based on multi-source video, characterized in that, include: The data acquisition module is used to synchronously acquire multi-view video data of the work area and its associated environmental data, and to perform time alignment and data association processing on the multi-view video data and the environmental data to form monitoring associated data; The data upload module is used to upload the monitoring-related data to the cloud analysis platform; The cloud-based analysis module is used to analyze multi-view video data in the monitoring-related data through the cloud-based analysis platform, using a preset target detection model and safety feature recognition model, to identify the safety equipment wearing status and behavior of the workers, and generate an analysis report containing the results of violation identification. The control module is used to generate alarm information and equipment control commands based on the analysis report; The push module is used to push the alarm information, the device control commands, and the real-time monitoring screen to the user terminal.
16. A storage medium, characterized in that, The computer program is stored in which the intelligent monitoring method for job safety based on multi-source video as described in any one of claims 1-14 can be loaded by a processor and executed.
Citation Information
Patent Citations
Operation safety monitoring method and device based on cloud edge collaboration
CN117812079A
Working plane multi-view fusion monitoring method and system based on power transmission line scene
CN118887618A
Intelligent operation monitoring system based on multi-source perception
CN121354310A
Object monitoring method and device based on multi-source data and storage medium
CN121462725A
Video and job data fusion analysis system and method for smart engineering
CN121580303A