Monitoring alarm method and device based on field operation environment dangerous point AI identification

By using edge computing devices and AI technology in high-altitude operation scenarios, real-time identification of workers' seat belt wearing status and dangerous areas of scaffolding, the problem that traditional methods are difficult to capture dangerous points in dynamic operations in real time is solved, and efficient safety prevention and control is achieved.

CN120048067AInactive Publication Date: 2025-05-27STATE GRID ZHEJIANG ELECTRIC POWER CO LTD QUZHOU POWER SUPPLY CO
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510518190.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In high-altitude operation scenarios, traditional hazard point identification methods are difficult to capture abnormal wear of workers' seat belts and damage to scaffolding structures in dynamic operations in real time, resulting in missed detection and lag risks, and lack the ability to analyze and classify early warning in multi-dimensional risk correlation.

Method used

Work images are collected through the on-site camera and transmitted to the edge computing device. AI is used to identify whether workers wear seat belts and scaffolds. There are dangerous areas in which they are used to generate an alarm signal.

Benefits of technology

Real-time identification of abnormal seat belt wear and scaffolding hidden dangers, dynamically evaluate compound risks and classify early warnings, improving the accuracy and initiative of safety prevention and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048067A_ABST
    Figure CN120048067A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent safety management, and provides a monitoring and alarming method and device based on field operation environment dangerous point AI identification, and the method comprises the steps: transmitting a field operation image collected by a field camera to an edge computing device; identifying a worker state and a scaffold state in the field operation image on the edge computing equipment to obtain a first dangerous point identification result for indicating whether a worker wears a safety belt or not and a second dangerous point identification result for indicating whether a dangerous area exists in the scaffold or not; and different levels of emergency alarm prompts are generated by combining the first dangerous point identification result and the second dangerous point identification result. Therefore, the abnormal wearing condition of the safety belt can be accurately identified in real time, meanwhile, the hidden danger of the scaffold can be rapidly detected, the dynamic evaluation of the composite risk and the graded early warning are realized, and the improvement of the accuracy and initiative of safety prevention and control is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent safety management, and more specifically, to a monitoring and alarm method and device based on AI identification of dangerous points in an on-site working environment. Background Art

[0002] In aerial work scenarios, accurate identification of dangerous points in the safe working environment is a core requirement to ensure the safety of personnel and the efficient advancement of the project. Due to the dynamic complexity of the aerial environment, hidden dangers such as the lack of workers' protective equipment (such as improper wearing of safety belts) and structural defects of scaffolding (such as deformation of components and loose connections) can easily lead to major accidents such as falls and collapses. However, traditional methods of identifying dangerous points face significant limitations: first, the status monitoring of workers' safety equipment relies on static image sampling, which makes it difficult to capture in real time the abnormal wearing of safety belts caused by instantaneous violations (such as unbuckled during work) during dynamic operations, and there is a significant risk of missed detection and lag; second, the integrity detection of facilities such as scaffolding mostly adopts periodic manual visual inspection, which cannot timely identify sudden structural damage (such as local tilt, lack of support) caused by load changes, material fatigue or accidental impact during construction, resulting in the response to hidden dangers lagging behind the evolution of risks; third, the existing technology monitors the two types of dangerous points of "personnel protection" and "facility status" in an isolated and separated manner, lacks multi-dimensional risk correlation analysis and graded early warning capabilities, and is difficult to formulate differentiated emergency strategies based on complex risk levels (such as the superposition of "not wearing a safety belt + scaffolding danger"), which restricts the accuracy and initiative of safety prevention and control.

[0003] Therefore, a monitoring and alarm solution based on AI identification of dangerous points in the on-site working environment is expected. Summary of the invention

[0004] In view of the shortcomings of the prior art, the present application provides a monitoring and alarm method and device based on AI identification of dangerous points in the on-site working environment.

[0005] According to one aspect of the present application, a monitoring and alarm method based on AI identification of dangerous points in a field operation environment is provided, which includes: Collecting on-site operation images through an on-site camera and transmitting the on-site operation images to an edge computing device; On the edge computing device, extracting a worker status image from the on-site operation image, and identifying whether the worker wears a safety belt based on the worker status image to obtain a first danger point identification result, including: inputting the on-site operation image into a human target detection network based on a first YOLO model to obtain the worker status image; performing safety belt wearing feature alignment analysis based on a reference image on the worker status image to obtain the first danger point identification result; On the edge computing device, extracting a scaffolding status image from the on-site operation image, and determining whether there is a dangerous area on the scaffolding based on the scaffolding status image to obtain a second dangerous point identification result; On the edge computing device, an alarm signal is generated based on the first danger point identification result and the second danger point identification result.

[0006] According to another aspect of the present application, a monitoring and alarm device based on AI identification of dangerous points in a field operation environment is provided, comprising: An image acquisition module is used to acquire on-site operation images through an on-site camera and transmit the on-site operation images to an edge computing device; A first danger point identification module is used to extract a worker status image from the on-site operation image on the edge computing device, and identify whether the worker is wearing a safety belt based on the worker status image to obtain a first danger point identification result, wherein the first danger point identification module includes: an image extraction unit, used to input the on-site operation image into a human target detection network based on a first YOLO model to obtain the worker status image; a safety belt wearing analysis unit, used to perform safety belt wearing feature alignment analysis on the worker status image based on a reference image to obtain the first danger point identification result; A second danger point identification module is used to extract a scaffolding status image from the field operation image on the edge computing device, and determine whether there is a dangerous area on the scaffolding based on the scaffolding status image to obtain a second danger point identification result; An alarm module is used to generate an alarm signal on the edge computing device based on the first danger point identification result and the second danger point identification result.

[0007] This application has significant technical effects due to the adoption of the above technical solutions: The monitoring and alarm method and device based on AI identification of dangerous points in the on-site working environment provided by the present application transmit the on-site working images collected by the on-site camera to the edge computing device, and identify the worker status and scaffolding status in the on-site working images on the edge computing device to obtain the first dangerous point identification result indicating whether the worker is wearing a safety belt and the second dangerous point identification result indicating whether there is a dangerous area on the scaffolding, and generate emergency warning prompts of different levels in combination with the first dangerous point identification result and the second dangerous point identification result. In this way, it is possible to accurately and real-timely identify abnormal wearing of safety belts, and quickly detect hidden dangers of scaffolding, realize dynamic assessment of compound risks and graded warning, and help improve the accuracy and initiative of safety prevention and control. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0009] Figure 1 It is a flowchart of a monitoring and alarm method based on AI identification of dangerous points in an on-site working environment according to an embodiment of the present application.

[0010] Figure 2 This is a flowchart of step S2 in the monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to an embodiment of the present application.

[0011] Figure 3 This is a flowchart of step S22 in the monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to an embodiment of the present application.

[0012] Figure 4 This is a flowchart of step S223 in the monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to an embodiment of the present application.

[0013] Figure 5 This is a flowchart of step S2232 in the monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to an embodiment of the present application.

[0014] Figure 6 This is a flowchart of step S2232-3 in the monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to an embodiment of the present application.

[0015] Figure 7 It is a block diagram of a monitoring and alarm device based on AI identification of dangerous points in a field working environment according to an embodiment of the present application. DETAILED DESCRIPTION

[0016] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.

[0017] Based on the technical problems raised by the above background technology, the present application provides a monitoring and alarm method based on AI identification of dangerous points in the on-site operation environment, which collects on-site operation images through on-site cameras and transmits the on-site operation images to edge computing devices, and identifies the worker status and scaffolding status in the on-site operation images on the edge computing device to obtain a first dangerous point identification result indicating whether the worker is wearing a safety belt and a second dangerous point identification result indicating whether there is a dangerous area on the scaffolding, and generates emergency warning prompts of different levels in combination with the first dangerous point identification result and the second dangerous point identification result. In this way, the wearing of workers' safety belts can be monitored in real time, no longer limited to static image sampling, and workers' instantaneous violations in dynamic operation processes can be discovered in a timely manner, which greatly improves the accuracy and timeliness of identifying abnormal wearing of workers' safety belts, effectively reduces the risk of falling accidents caused by improper wearing of safety belts, and through automatic analysis of scaffolding status images in on-site operation images, safety hazards of scaffolding can be identified more quickly. In addition, the monitoring of two types of dangerous points, "personnel protection" and "facility status", is linked. Different levels of alarm signals are generated according to different combinations of dangerous point identification results. Differentiated emergency strategies can be formulated based on complex risk levels, which helps to improve the accuracy and initiative of safety prevention and control and make safety management more scientific and effective.

[0018] Figure 1 FIG. 1 is a flow chart of a monitoring and alarm method based on AI identification of dangerous points in a field operation environment according to an embodiment of the present application. Figure 1 As shown, according to the monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to the embodiment of the present application, it includes: S1, collecting on-site working images through an on-site camera, and transmitting the on-site working images to an edge computing device; S2, at the edge computing device, extracting a worker status image from the on-site working image, and identifying whether the worker is wearing a safety belt based on the worker status image to obtain a first dangerous point identification result; S3, at the edge computing device, extracting a scaffolding status image from the on-site working image, and determining whether there is a dangerous area on the scaffolding based on the scaffolding status image to obtain a second dangerous point identification result; S4, at the edge computing device, generating an alarm signal based on the first dangerous point identification result and the second dangerous point identification result.

[0019] In step S1, the on-site operation image is collected by the on-site camera, and the on-site operation image is transmitted to the edge computing device. It should be understood that the on-site operation image contains worker status information, scaffolding structure information, and work environment area information. Specifically, the worker status information includes safety equipment wearing information (such as whether the safety belt is worn, whether the wearing position complies with the specification, whether the safety rope is connected to the anchor point, etc.) and worker dynamic behavior information (such as the worker's action, position, multi-person collaboration status, etc. in high-altitude operations); scaffolding structure information includes scaffolding component integrity information (such as whether the crossbar, vertical pole, and connector are complete, whether there is deformation, rust, breakage or missing), connection stability information (such as whether the connection status of key nodes (such as intersections, support bases) is firm, whether there is looseness or detachment, etc.) and overall morphological abnormality information (such as local or overall tilt, settlement, twisting and other macro deformations of the scaffolding); the work area environment information includes lighting condition information, surrounding obstacle information, etc. Accordingly, considering that in high-altitude operation scenes, it is very important to detect dangers in time and issue alarms. If the image data is transmitted to the cloud for processing, the time delay between network transmission and cloud computing may result in delayed warning of danger. However, the edge computing device is located near the site and can quickly complete image processing and dangerous point identification locally, triggering the alarm signal in seconds, greatly reducing response delays and meeting the real-time requirements of high-risk scenarios.

[0020] In step S2, the edge computing device extracts a worker status image from the on-site work image, and identifies whether the worker is wearing a safety belt based on the worker status image to obtain a first hazard point identification result. It should be understood that in high-altitude work scenarios, the life safety of workers is of vital importance, and correctly wearing a safety belt is a key measure to prevent workers from falling accidents. By processing and analyzing the worker status in the on-site work image on the edge device, the monitoring focus can be focused on the worker, a key safety factor, and the worker's safety belt wearing situation can be understood in real time.

[0021] In particular, the traditional safety belt wearing recognition method has significant application limitations in dynamic working scenes in high-altitude environments, and its recognition accuracy and reliability are difficult to meet the refined requirements of safety protection equipment compliance monitoring. Specifically, it is manifested in the following three dimensions: First, the existing methods mostly use shallow judgment logic such as target detection or color matching, which can only verify the physical existence of the safety belt, but cannot effectively identify the core indicator of "wearing standardization". For example, it is impossible to distinguish key details such as whether the safety belt is fixed across the shoulder in a standardized manner and the buckle locking state, resulting in dangerous behaviors such as loose winding or wrong hanging being misjudged as compliance, posing a major safety hazard. Secondly, the complex environmental factors in high-altitude working scenes pose a severe challenge to the performance of the algorithm - the frequent changes in the body posture of the operators (such as bending, turning), dynamic lighting interference (backlight, shadow alternation) and local occlusion caused by tools and equipment (such as the buckle area is blocked) will significantly affect the feature extraction accuracy of the traditional algorithm, thereby causing missed detection or false detection problems. More importantly, the existing technology lacks a feature-level comparison mechanism with safety standard images, making it difficult to verify from a semantic level whether the wearing status complies with operating specifications (such as position deviation of the seat belt fixing point). This makes it difficult to achieve true compliance verification of seat belt wearing identification.

[0022] Based on this, the technical concept of the present application is to first use a human target detection network to extract a worker status image from a field operation image, then use a seat belt wearing area target detection network to extract a seat belt wearing area ROI image from the worker status image, and then extract the seat belt wearing state image features from the seat belt wearing area ROI image and the seat belt wearing specification reference image extracted from the background database of the edge computing device to obtain the seat belt wearing specification reference state features and the seat belt wearing real state features, and finally intelligently obtain the first danger point identification result of whether the worker wears a seat belt based on the joint perception alignment representation between the seat belt wearing specification reference state features and the seat belt wearing real state features. In this way, by introducing the seat belt wearing specification reference image, and using the image twin extractor and the seat belt wearing feature structure alignment joint perception operation to analyze the two images from the semantic level, a feature level comparison mechanism can be provided for seat belt wearing recognition, which helps to achieve true compliance verification, avoids the limitation of traditional methods that only verify physical existence, greatly reduces the safety hazards caused by improper wearing, and can further reduce the impact of complex environmental factors on the recognition results, and reduce the probability of missed detection and false detection.

[0023] Specifically, Figure 2 FIG. 1 is a flow chart of step S2 in the monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to an embodiment of the present application. Figure 2As shown, the step S2 includes: S21, inputting the on-site work image into a human target detection network based on a first YOLO model to obtain the worker status image; S22, performing a seat belt wearing feature alignment analysis based on a reference image on the worker status image to obtain the first danger point identification result.

[0024] In step S21, the field operation image is input into the human target detection network based on the first YOLO model to obtain the worker status image. It should be understood that the aerial work site image usually contains scaffolding structures, construction equipment, dynamically moving workers and complex background elements. The safety belt detection directly in the panoramic image is susceptible to interference from non-target areas, resulting in feature extraction deviation. In order to effectively exclude irrelevant information such as scaffolding crossbars and tool occlusions, the analysis scope is limited to the worker body area, thereby improving the accuracy of subsequent safety belt wearing status recognition. In this application, it is necessary to input the field operation image into the human target detection network based on the first YOLO model to obtain the worker status image. Those of ordinary skill in the art should know that the YOLO model, as a single-stage target detection framework, has the core feature of reconstructing the target detection task into a single global regression problem, dividing the image area by gridding and synchronously predicting the bounding box and category probability to achieve efficient multi-scale target positioning. Compared with the traditional two-stage detection model, YOLO abandons the candidate region generation step and adopts an end-to-end training and reasoning mechanism. It significantly reduces the computing delay while ensuring high detection accuracy, and can effectively meet the millisecond-level response requirements of high-altitude work scenes. In view of the multi-posture characteristics of workers in high-altitude construction scenes (such as climbing, bending, and turning), the YOLO model integrates semantic information at different levels through a feature pyramid network, and can adaptively capture the scale changes and deformation characteristics of human targets to ensure stable detection in dynamic work scenes. For example, when a worker is in a gap between scaffolding or part of his limbs is blocked, the model can still accurately infer the complete human bounding box through local visible features (such as the head and shoulder contours), avoiding the problem of missed detection caused by target truncation. Based on the generated worker status image for safety belt wearing recognition, the details of the worker's safety belt wearing can be observed more clearly, reducing the impact of background information on the recognition results, which is conducive to improving the accuracy of the recognition results.

[0025] In step S22, the worker status image is subjected to a safety belt wearing feature alignment analysis based on the reference image to obtain the first danger point identification result. Specifically, Figure 3 FIG. 1 is a flow chart of step S22 in the monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to an embodiment of the present application. Figure 3As shown, the step S22 includes: S221, extracting a safety belt wearing specification reference image from the background database of the edge computing device; S222, inputting the worker status image into a safety belt wearing area target detection network based on a second YOLO model to obtain a safety belt wearing area ROI image; S223, performing a safety belt wearing feature alignment analysis on the safety belt wearing specification reference image and the safety belt wearing area ROI image to obtain a safety belt wearing reference-real image semantic level alignment coding feature map; S224, obtaining the first danger point identification result based on the safety belt wearing reference-real image semantic level alignment coding feature map.

[0026] In step S221, a safety belt wearing specification reference image is extracted from the background database of the edge computing device. It should be understood that the safety belt wearing specification reference image can provide a standardized template for safety belt wearing compliance verification. For example, the safety belt should be fixed across the shoulders in a standardized manner, and the two shoulder straps should be evenly distributed on both shoulders. There should be no single shoulder stress or shoulder strap slippage. The waist belt part should fit tightly to the waist, generally above the hip bone, not too loose or too tight, and the whole should be kept horizontal. The presentation of this standard posture can provide a posture reference for subsequent comparison of actual wearing conditions. That is, by obtaining a safety belt wearing specification reference image, the problem of normative misjudgment caused by the lack of standard reference in traditional methods can be solved.

[0027] In step S222, the worker status image is input into the target detection network of the seat belt wearing area based on the second YOLO model to obtain the ROI image of the seat belt wearing area. It should be understood that although the worker status image has been focused on the human target, the specific components of the seat belt (such as buckles, shoulder straps, leg straps, etc.) may be difficult to analyze directly due to posture changes or partial occlusion. In the aerial work scene, the standardized wearing of the seat belt requires that its components strictly follow the fixed position and interaction relationship (such as the shoulder straps need to cross the shoulders and be fixed to the back anchor point, and the buckles need to be completely locked), but if only relying on the rough detection of the overall area of ​​the worker, such detailed features cannot be captured. Through the precise positioning of the seat belt wearing area by the second YOLO model, the focus can be further narrowed to the local area where the key components of the seat belt are located, eliminating the interference of other parts of the human body (such as arms, tools), so as to provide high-resolution fine-grained input data for subsequent normative verification.

[0028] In step S223, the seat belt wearing feature alignment analysis is performed on the seat belt wearing specification reference image and the seat belt wearing area ROI image to obtain a seat belt wearing reference-real image semantic level alignment encoding feature map. Specifically, Figure 4 FIG. 1 is a flow chart of step S223 in the monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to an embodiment of the present application. Figure 4 As shown, the step S223 includes: S2231, inputting the seat belt wearing specification reference image and the seat belt wearing area ROI image into an image twin extractor including a first seat belt wearing state image feature extractor and a second seat belt wearing state image feature extractor to obtain a seat belt wearing specification reference state feature map and a seat belt wearing real state feature map; S2232, performing seat belt wearing feature structure alignment and joint perception on the seat belt wearing specification reference state feature map and the seat belt wearing real state feature map to obtain the seat belt wearing reference-real image semantic level alignment encoding feature map.

[0029] In step S2231, the seat belt wearing specification reference image and the seat belt wearing area ROI image are input into an image twin extractor including a first seat belt wearing state image feature extractor and a second seat belt wearing state image feature extractor to obtain a seat belt wearing specification reference state feature map and a seat belt wearing real state feature map. Accordingly, considering that the original images (seat belt wearing specification reference image and seat belt wearing area ROI image) contain a large amount of redundant pixel information (such as background texture, clothing wrinkles), directly comparing pixel values ​​(such as traditional color matching methods) cannot capture the core semantics of "wearing specification" (such as buckle closure, shoulder strap tension). In order to project the two types of images into a high-dimensional semantic space, eliminate irrelevant noise (such as illumination differences, background interference), and capture functional structural features directly related to compliance (such as buckle locking status, webbing direction, and fixed point position), in this application, it is necessary to input the seat belt wearing specification reference image and the seat belt wearing area ROI image into an image twin extractor including a first seat belt wearing state image feature extractor and a second seat belt wearing state image feature extractor to obtain a seat belt wearing specification reference state feature map and a seat belt wearing real state feature map. In particular, the essence of the image twin extractor is a twin neural network model, and the first seat belt wearing state image feature extractor and the second seat belt wearing state image feature extractor are two sub-networks that share weights and constitute the twin neural network model. They use exactly the same network architecture and parameters (such as convolutional layer, pooling layer, residual module) to ensure that the input image is encoded in the same feature space, so that the feature comparison is interpretable. Specifically, the first seat belt wearing state image feature extractor is used to process the seat belt wearing specification reference image, and the second seat belt wearing state image feature extractor is used to process the seat belt wearing area ROI image. The shared weight twin structure forces the standard reference image and the real ROI image to be processed by the same set of convolution kernels to ensure that the encoding rules of the feature maps of the two are completely consistent. This can ensure the consistency of the generated safety belt wearing standard reference state features and the safety belt wearing real state features in the feature space, avoiding misjudgment caused by different network parameters. And through the attention mechanism and multi-scale feature fusion, the model can focus on the key areas of the functional parts of the safety belt (such as the contact surface between the buckle lock tongue and the lock hole). Even if there is a millimeter-level deviation in actual wearing (such as the lock tongue is not fully inserted), the local response difference of the image can still be accurately captured. In short, through the processing of the image twin extractor, the leap from "seat belt existence detection" to "semantic verification of wearing standardization" is achieved. The feature mapping of the standard reference image and the real ROI image not only provides a quantifiable semantic benchmark for compliance judgment, but also through the shared weight and attention mechanism, it gives the model strong robustness to complex interference (occlusion, posture, illumination) in dynamic working scenes, which can lay the core technical foundation for the precise and intelligent monitoring of high-altitude work safety protection.

[0030] In step S2232, the seat belt wearing standard reference state feature map and the seat belt wearing real state feature map are subjected to seat belt wearing feature structure alignment joint perception to obtain the seat belt wearing reference-real image semantic level alignment encoding feature map. Specifically, Figure 5 FIG. 2 is a flow chart of step S2232 in the monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to an embodiment of the present application. Figure 5 As shown, the step S2232 includes: S2232-1, feature decoupling the seat belt wearing specification reference state feature map and the seat belt wearing actual state feature map to obtain a set of seat belt wearing specification reference state local feature coding matrices and a set of seat belt wearing actual state local feature coding matrices; S2232-2, based on the feature correspondence between any two seat belt wearing specification reference state local feature coding matrices and the seat belt wearing actual state local feature coding matrices in the set of seat belt wearing specification reference state local feature coding matrices and the set of seat belt wearing actual state local feature coding matrices. According to the bit alignment matching degree, the set of local feature coding matrices of the seat belt wearing specification reference state and the set of local feature coding matrices of the seat belt wearing actual state are subjected to dynamic search and alignment of the seat belt wearing feature structure to obtain a set of structurally aligned {seat belt wearing specification reference state local feature coding matrix, seat belt wearing actual state local feature coding matrix} feature pairs; S2232-3, based on the set of structurally aligned {seat belt wearing specification reference state local feature coding matrix, seat belt wearing actual state local feature coding matrix} feature pairs, obtain the seat belt wearing reference-real image semantic level aligned coding feature map.

[0031] It should be understood that in high-altitude work scenarios, the difference in perspective, dynamic posture and occlusion interference between the safety belt wearing specification reference image and the ROI image of the actual work scene leads to the problem of spatial geometric structure misalignment (such as buckle position offset, shoulder strap direction difference) and insufficient semantic information complementarity (such as the isolation of the standard tension feature of the reference image and the relaxation state feature of the actual image) in the feature map extracted. When the traditional method directly performs feature vector comparison, it ignores the spatial structural correlation of the feature map and cannot capture the geometric correspondence of local areas (such as buckle closure, fixed point position), resulting in the failure of semantic-level compliance verification. Based on this, in the present application, the safety belt wearing specification reference state feature map and the safety belt wearing actual state feature map are jointly perceived to align the safety belt wearing feature structure to obtain the safety belt wearing reference-real image semantic-level alignment encoding feature map. Specifically, by decoupling the feature map into a collection of local feature encoding matrices, the structural similarity metric in the feature space is used to dynamically search and align local regions of cross-level features (for example, a soft alignment mapping is established between the buckle region of the reference image and the corresponding position of the real image), and then feature aggregation is performed under the drive of the attention mechanism (the feature contribution of key areas such as buckles and fixed points is strengthened through self-attention weights). This process not only corrects the spatial misalignment caused by posture changes (such as the rotational alignment of shoulder strap features when workers turn around), but also improves the integrity of feature representation through complementary fusion (such as the contrast enhancement of the standard tension texture of the reference image and the relaxation texture of the real image). The semantic-level aligned encoding feature map finally generated organically integrates the prior knowledge of the normative reference (such as the standard geometric form of the buckle closure) with the dynamic features of real wear (such as the buckle contour reasoning under partial occlusion), forming a high-quality feature representation with joint perception capabilities. That is, through spatial structure alignment and semantic complementary enhancement, it is possible to accurately identify wearing compliance details, such as detecting hidden violations such as half-closed buckles and shoulder strap slippage, thereby upgrading the judgment of the seat belt wearing status from "existence detection" to "functional compliance verification", which can provide more discriminative feature inputs for subsequent state identifiers, ultimately ensuring the accuracy and reliability of seat belt wearing monitoring during high-altitude operations.

[0032] Specifically, in the embodiment of the present application, the step S2232-1 performs feature decoupling on the seat belt wearing specification reference state feature graph and the seat belt wearing actual state feature graph to obtain a set of seat belt wearing specification reference state local feature coding matrices and a set of seat belt wearing actual state local feature coding matrices, which can be expressed as the following formula: in, is the reference state characteristic diagram of the safety belt wearing specification, is the characteristic diagram of the real state of wearing the seat belt, To perform feature decoupling operations, , , and are the first, second, and third local feature encoding matrices of the seat belt wearing standard reference state. and A local feature encoding matrix of the seat belt wearing standard reference state, , , and are the first, second, and third local feature encoding matrices of the actual state of seat belt wearing. and A local feature encoding matrix of the actual seat belt wearing state.

[0033] It should be understood that in aerial work scenarios, the compliance verification of the seat belt wearing status is limited by the spatial deformation sensitivity and semantic confusion problems of the global feature map. When traditional methods directly process high-dimensional global features, it is difficult to capture the geometric structure and dynamic changes of local details such as buckles and shoulder straps (such as the deviation of the shoulder strap direction caused by the worker turning around). Feature decoupling can transform complex global representations into more easily operable local semantic units by decomposing global features into a set of local feature encoding matrices. This operation not only reduces the computational complexity, but also enhances the ability to express details by focusing on local areas (such as buckle closure and fixed point texture), providing a fine-grained basis for subsequent dynamic alignment. That is, the set of local feature encoding matrices of the reference state and the real state of the seat belt wearing specification obtained by decoupling respectively encodes standard tension features (such as the ideal geometric shape of buckle closure) and dynamic wearing features (such as the relaxed contour under occlusion). This decomposition overcomes the sensitivity of global features to spatial misalignment, enabling the model to independently analyze the local semantics of key areas such as shoulder strap direction and buckle position, thereby providing an operable input unit for cross-image structural alignment.

[0034] Specifically, in the embodiment of the present application, the step S2232-2, based on the feature phase alignment matching degree between any two local feature coding matrices of the seat belt wearing specification reference state and the local feature coding matrices of the seat belt wearing actual state in the set of the seat belt wearing specification reference state, and the local feature coding matrices of the seat belt wearing actual state, performs a seat belt wearing feature structure dynamic search alignment on the set of local feature coding matrices of the seat belt wearing specification reference state and the set of local feature coding matrices of the seat belt wearing actual state to obtain a set of structural alignment {seat belt wearing specification reference state local feature coding matrix, seat belt wearing actual state local feature coding matrix} feature pairs, which can be expressed as the following formula: in, is the transpose operation, To calculate the Frobenius norm, yes and The feature structure alignment matching degree between To return the maximum value value, It is to find the location of the maximum approximate matching value in the set of local feature encoding matrices of the actual state of seat belt wearing.

[0035] Accordingly, due to the spatial misalignment of local regions caused by perspective differences and dynamic postures (such as the position offset of the buckle in the reference image and the real image), the traditional static alignment method cannot establish an effective cross-level correspondence. By calculating the feature phase alignment matching degree between any two local feature encoding matrices of the seat belt wearing specification reference state and the local feature encoding matrices of the seat belt wearing real state in the set of local feature encoding matrices of the seat belt wearing specification reference state and the set of local feature encoding matrices of the seat belt wearing real state, the dynamic search strategy can find the feature pairs with the highest structural consistency in the feature space (such as the closed buckle feature of the reference image and the semi-closed buckle contour of the real image). This process does not rely on the rigid correspondence of physical coordinates, but uses the semantic similarity of feature vectors to establish a soft alignment mapping to solve the geometric mismatch problem caused by occlusion or posture changes. That is, the set of structural alignment {local feature encoding matrix of the seat belt wearing specification reference state, local feature encoding matrix of the seat belt wearing real state} feature pairs generated by dynamic alignment corrects spatial deformations such as shoulder belt rotation and buckle position offset through soft mapping. For example, the occluded buckle area in the real image can be associated with the closed standard features of the reference image through structural matching, achieving local alignment based on semantic reasoning. This alignment method breaks through the limitations of traditional rigid matching and can provide a geometrically consistent feature interaction basis for the subsequent attention mechanism.

[0036] Specifically, Figure 6 FIG. 2 is a flow chart of step S2232-3 in the monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to an embodiment of the present application. Figure 6As shown, the step S2232-3 includes: S2232-31, performing collaborative perception of seat belt wearing features on each structure alignment {seat belt wearing specification reference state local feature coding matrix, seat belt wearing actual state local feature coding matrix} in the set of the structure alignment {seat belt wearing specification reference state local feature coding matrix, seat belt wearing actual state local feature coding matrix} feature pairs to obtain a set of seat belt wearing reference-actual state collaborative perception attention weights: S2232-32, based on the set of seat belt wearing reference-actual state collaborative perception attention weights, performing attention-driven complementary feature aggregation on the set of the structure alignment {seat belt wearing specification reference state local feature coding matrix, seat belt wearing actual state local feature coding matrix} feature pairs to obtain the seat belt wearing reference-actual image semantic level alignment coding feature map.

[0037] More specifically, in an embodiment of the present application, the step S2232-31 performs collaborative perception of seat belt wearing features on each structure alignment {seat belt wearing specification reference state local feature coding matrix, seat belt wearing actual state local feature coding matrix} in the set of structure alignment {seat belt wearing specification reference state local feature coding matrix, seat belt wearing actual state local feature coding matrix} feature pairs to obtain a set of seat belt wearing reference-actual state collaborative perception attention weights.

[0038] It should be understandable that simple structural alignment cannot distinguish the importance differences of feature pairs (such as the buckle closure feature is more critical to compliance judgment than the shoulder strap color feature). The seat belt wearing feature collaborative perception operation adaptively learns the weight allocation rules by modeling the interactive relationship between feature pairs (such as the contrast enhancement of standard tension features and true relaxation features). Specifically, the operation uses the cross-attention mechanism to analyze the complementarity of feature pairs (such as the fixed point position prior of the reference image and the dynamic occlusion information of the real image) and redundancy (such as repeated background textures), and highlight key violation clues (such as the weak edge difference of the half-closed buckle). That is, the generated set of seat belt wearing reference-real state collaborative perception attention weights realizes the dynamic importance screening of feature pairs. For example, a high weight is given to the buckle closure feature pair, while irrelevant texture differences caused by illumination changes are suppressed. This mechanism enables the model to focus on functional compliance indicators (such as whether the shoulder strap force state meets the standard tension threshold) rather than surface visual similarity, significantly improving the detection sensitivity of hidden violations (such as shoulder straps slipping but not completely occluded).

[0039] More specifically, in the embodiment of the present application, the step S2232-31 includes: calculating the association matrix between each structural alignment {seat belt wearing specification reference state local feature coding matrix, seat belt wearing real state local feature coding matrix} in the set of the structural alignment {seat belt wearing specification reference state local feature coding matrix, seat belt wearing real state local feature coding matrix} feature pairs to obtain a set of seat belt wearing specification reference-real state associated feature coding matrices; performing trace measurement operation on the set of seat belt wearing specification reference-real state associated feature coding matrices to obtain a set of seat belt wearing specification reference-real state associated feature trace measurement values; based on the set of seat belt wearing specification reference-real state associated feature trace measurement values, obtaining the set of seat belt wearing reference-real state collaborative perception attention weights. The above process can be expressed by the formula: in, Calculate the feature co-aware attention weights. It is a structural alignment {seat belt wearing specification reference state local feature encoding matrix, seat belt wearing real state local feature encoding matrix} feature pair, is the trace value of the matrix, yes and The seat belt wearing standard reference-real state correlation characteristic trace measurement value between is the normalization function, for and The seatbelt wearing reference-ground truth state co-perception attention weights.

[0040] Preferably, in another embodiment of the present application, based on the set of the seat belt wearing specification reference-real state associated characteristic trace measurement values, the set of the seat belt wearing reference-real state collaborative perception attention weights is obtained, including: the set of the seat belt wearing specification reference-real state associated characteristic trace measurement values ​​is embedded and compactified into a manifold structure of interlocking connection strength to obtain a set of optimized seat belt wearing specification reference-real state associated characteristic trace measurement values; based on the set of optimized seat belt wearing specification reference-real state associated characteristic trace measurement values, the set of the seat belt wearing reference-real state collaborative perception attention weights is obtained. The above process can be expressed by the formula:

[0041] in, yes The transposed matrix of yes ,Right now The corresponding structural alignment seat belt wearing real state local feature encoding matrix, yes Middle The eigenvalues ​​at the positions, yes Middle The eigenvalues ​​at the positions, and , , For calculation and Between distance, For calculation and Between distance, is the preset threshold, To calculate the number of conditions that meet the condition, yes Distance interlock connection strength parameter, yes Distance interlock connection strength parameter, The natural constant The exponential function value with base , It is the point multiplication by position. yes The optimized matrix is yes The optimized matrix is yes The optimized seat belt wearing standard reference-real state correlation feature trace measurement value after optimization, for and The seatbelt wearing reference-ground truth state co-perception attention weights.

[0042] In particular, by introducing a reconstruction mechanism based on interlocking connection strength and manifold embedding compactification, the core goal is to resolve the feature interaction ambiguity caused by dynamic posture and perspective differences, thereby laying a mathematical foundation for the precise allocation of attention weights. Specifically, the optimization first calculates the local feature encoding matrix of the structural alignment feature pair and The L1 and L2 double distance constraints between them are used to screen out matching pairs that meet the threshold ε, and their number is counted to generate the interlocking connection strength parameter and This process is essentially a dynamic quantification of the quality of feature alignment: Characterizes the phase alignment density of the reference image's standard features and the real image's dynamic features on the local geometric structure (such as the overlap between the standard contour of the snap-fit ​​closure and the actual contour), and It reflects the consistency of their distribution in the manifold space (such as the topological similarity between the standard path and the actual path of the shoulder strap). The design of the dual distance constraint takes into account both spatial tolerance and semantic stability - the L1 distance allows a certain degree of geometric deviation (such as the translation of the buckle position caused by the worker turning around), while the L2 distance ensures the reliability of the matching pairs at the semantic level by suppressing noise interference (such as texture differences caused by lighting changes). For example, when the buckle in the real image is only partially visible due to occlusion, the L2 distance can filter out false matches caused by missing contours, and only retain feature pairs with semantically similar closed morphology to the reference image. Further manifold structure embedding compaction operations apply nonlinear transformations to the original matrix through exponential mapping, and fuse the interlocking connection strength and the safety belt wearing specification reference-real state association trace value into the feature space. This optimization has a dual role: first, - The difference dynamically adjusts the matrix scaling factor, strengthens the compactness of highly aligned feature areas (such as unobstructed buckle closure areas) in the manifold space, and thus improves the contrast of local structures; secondly, through iterative updates, the model can adaptively balance the global energy relationship between standard and real features (such as the complementarity of standard tension features and real relaxation features). This compactification operation not only enhances the robustness of long-range feature associations (such as cross-view fixed point position reasoning), but also amplifies the distinguishability of key violation clues through nonlinear transformations (such as the slight gradient difference of feature vectors when the buckle is half closed). For example, in the real image, when the degree of closure of the buckle is difficult to distinguish by the naked eye due to angle offset, the manifold compactification based on the optimized seat belt wearing specification reference-real state association feature trace value maps such subtle differences into significant feature distance changes, allowing the attention mechanism to keenly capture such hidden dangers. In other words, this optimization step provides a calibrated feature interaction field for the generation of attention weights. Through the reconstruction operation driven by the interlocking connection strength, the model can distinguish the importance level of feature pairs: highly aligned and smoothly distributed feature pairs (such as clearly visible shoulder strap fixing points) are given higher weights, while weakly aligned feature pairs caused by occlusion or deformation (such as blurred buckle edges) are suppressed through manifold compaction. In general, this optimization mechanism enables the generated seat belt wearing reference-real state collaborative perception attention weights to not only have the ability to correct spatial structure, but also strengthen the semantic contrast through manifold embedding, thereby providing a highly discriminative decision basis for hidden violation detection in subsequent compliance verification (such as shoulder straps slipping but not completely occluded).

[0043] More specifically, in the embodiment of the present application, the step S2232-32, based on the set of the seat belt wearing reference-real state collaborative perception attention weights, performs attention-driven complementary feature aggregation on the set of the structural alignment {seat belt wearing specification reference state local feature encoding matrix, seat belt wearing real state local feature encoding matrix} feature pairs to obtain the seat belt wearing reference-real image semantic level alignment encoding feature map, which can be expressed as the following formula: in, and are the first collaborative sensing weight matrix and the second collaborative sensing weight matrix, It is subtracted by position point. It is added by location point. yes and The collaborative seat belt wearing reference-real state collaborative perception feature matrix is ​​the first in the set of seat belt wearing reference-real state collaborative perception feature matrices. A seat belt wearing reference-real state collaborative perception feature matrix, , and are the first, second and third in the set of seat belt wearing reference-real state collaborative perception feature matrices. A seat belt wearing reference-real state collaborative perception feature matrix, It is the seat belt wearing reference-real image semantic level alignment encoding feature map.

[0044] It should be understandable that simple feature concatenation or averaging will dilute key violation signals (such as subtle gradient changes in half-closed buckles). Based on the attention-driven complementary feature aggregation operation, weighted fusion is used to highlight the contrasting semantics between the standard and real features (such as the difference between the standard closed shape and the real relaxed state), while retaining complementary information (such as the reference image provides the geometric prior of the occluded area, and the real image supplements the actual contour). This process constructs a joint perception field in the feature space, nonlinearly fusion of normative knowledge (such as the standard angle range of buckle closure) and dynamic observation (such as the actual buckle angle offset). That is, the final generated seat belt wearing reference-real image semantic level alignment encoding feature map has both spatial structure correction and semantic contrast enhancement characteristics. For example, the aggregation weights highlight the difference characteristics of the closure of the buckle area, while the normative-real contrast texture of the shoulder strap direction is integrated. This feature representation enables the subsequent state recognizer to directly parse the "functional compliance" indicator characteristics, rather than just judging whether the seat belt exists, thereby upgrading the monitoring dimension from "existence verification" to "effectiveness verification", significantly improving the recognition accuracy of high-risk hidden violations such as half-closure and misaligned fixation.

[0045] In step S224, based on the safety belt wearing reference-real image semantic level alignment coding feature map, the first danger point identification result is obtained. Specifically, in an embodiment of the present application, the step S224 includes: inputting the safety belt wearing reference-real image semantic level alignment coding feature map into a classifier-based state identifier to obtain the first danger point identification result, and the first danger point identification result is used to indicate whether the worker wears a safety belt. It should be understood that although the obtained safety belt wearing reference-real image semantic level alignment coding feature map contains rich information, it is only a collection of features and cannot intuitively indicate the worker's safety belt wearing status. The introduction of a classifier-based state identifier aims to convert these complex features into clear and easy-to-understand judgment results, that is, whether the worker wears a safety belt, so as to solve the problem that traditional methods are difficult to accurately determine the wearing standardization. Specifically, the classifier can learn the complex correlation between features, and use the semantic level alignment coding feature map extracted in the early stage to obtain accurate wearing status judgment through complex algorithms and training. For example, when the feature map shows that the buckle feature is slightly different from the standard, and the shoulder strap position is slightly offset, the classifier can comprehensively judge that this situation is not wearing properly based on the experience accumulated in the training data, that is, it is essentially not wearing a safety belt. It should also be noted that the lightweight classifier is adapted to edge computing devices and can complete reasoning in a very short time. Even if the worker wears the safety belt abnormally for a moment during the operation, such as unfastening the buckle within 0.3 seconds, it can be detected in time. This process converts the complex feature processing results in the early stage into a clear first danger point identification result, which can not only determine whether the safety belt exists, but also determine whether the wearing method is standardized. It can provide a reliable basis for subsequent graded alarms, thereby greatly improving the accuracy, timeliness and reliability of judging the wearing status of workers' safety belts in high-altitude work safety monitoring.

[0046] In summary, the step S2 is clearly explained, which uses deep learning-based image processing technology to first extract the worker status image from the on-site work image, then extract the seat belt wearing area ROI image from the worker status image, and then extract the seat belt wearing state image features from the seat belt wearing area ROI image and the seat belt wearing specification reference image extracted from the background database of the edge computing device to obtain the seat belt wearing specification reference state features and the seat belt wearing real state features, and finally intelligently obtain the first danger point identification result of whether the worker wears a seat belt based on the joint perception alignment representation between the seat belt wearing specification reference state features and the seat belt wearing real state features. In this way, by analyzing the seat belt wearing area ROI image and the seat belt wearing specification reference image at the feature granularity, a feature level comparison mechanism can be provided for seat belt wearing identification, which helps to achieve compliance verification in the true sense, avoids the limitation of traditional methods that only verify physical existence, greatly reduces the safety hazards caused by improper wearing, and can further reduce the impact of complex environmental factors on the identification results, and reduce the probability of missed detection and false detection.

[0047] In step S3, the edge computing device extracts a scaffolding state image from the on-site operation image, and determines whether there is a dangerous area on the scaffolding based on the scaffolding state image to obtain a second dangerous point identification result. It should be understood that the on-site operation image usually contains multiple elements such as workers, mechanical equipment, and building materials, and the scaffolding structure may be blocked by dynamic targets (such as moving workers) or complex backgrounds (such as temporarily stacked steel pipes). By extracting the scaffolding state image from the on-site operation image, it is possible to eliminate the interference of irrelevant information, focus on the scaffolding as a key facility, and conduct detailed analysis and evaluation on it in order to timely discover potential dangerous hazards. That is, extracting the scaffolding state image and analyzing it using AI technology can realize real-time monitoring of the scaffolding state, make up for the shortcomings of traditional detection methods, improve the efficiency and accuracy of discovering scaffolding dangers, and ensure the safety of high-altitude operations. Specifically, first, the scaffolding state image is processed using a deep learning-based detection model (a convolutional neural network model specifically for scaffolding structure) pre-trained in the edge computing device. The model has been trained with a large number of scaffolding image samples containing normal and dangerous areas, and can identify various features in scaffolding images, such as the shape, position and connection relationship of components such as uprights, crossbars, and diagonal braces. Then, the model will analyze the scaffolding components in the image to detect whether there are component deformation (such as the bending degree of the rod exceeds the normal range), loose connections (such as bolts not tightened, fasteners loose, etc.), local tilt (by calculating the deviation of the tilt angle of the whole or part of the scaffolding from the normal angle), missing support (compared with the standard scaffolding structure model to determine whether there are missing support components), etc. For example, for the detection of rod deformation, the model will extract the contour features of the rod and compare them with the features of normal rods to calculate the degree of deformation; for the detection of loose connections, the detailed features of the connection parts will be analyzed to determine whether the connection is firm. Finally, based on the detection results of the model for various dangerous features, it is comprehensively judged whether the scaffolding has a dangerous area. If one or more of the above dangerous situations are detected, it is determined that the scaffolding has a dangerous area; if no dangerous features are detected, it is determined that the scaffolding does not have a dangerous area. For safety managers, the second danger point identification result is an important basis for formulating safety management measures. If the identification result shows that there is a dangerous area on the scaffolding, the manager can arrange professionals to inspect, repair or reinforce the scaffolding according to the severity of the danger to ensure the safety of the scaffolding. In addition, the identification result can also be combined with the first danger point identification result (whether the worker wears a safety belt) to conduct multi-dimensional risk association analysis and graded warning. In short, the second danger point identification result obtained by analysis not only provides real-time protection for the safety of the scaffolding itself, but also constructs a multi-dimensional risk prevention and control network through dynamic association with the protection status of personnel, which can significantly improve the overall safety of high-altitude operation scenes.

[0048] In step S4, the edge computing device generates an alarm signal based on the first danger point identification result and the second danger point identification result. Specifically, in an embodiment of the present application, the step S4 includes: in response to the first danger point identification result that the worker is not wearing a safety belt and the second danger point identification result is that there is a dangerous area on the scaffolding, generating a third-level emergency alarm prompt; in response to the first danger point identification result that the worker is wearing a safety belt and the second danger point identification result is that there is a dangerous area on the scaffolding, generating a second-level emergency alarm prompt; in response to the first danger point identification result that the worker is not wearing a safety belt and the second danger point identification result is that there is no dangerous area on the scaffolding, generating a first-level emergency alarm prompt. It should be understood that the safety risks of aerial work scenes are complex and multi-dimensional. Single consideration of the worker's safety belt wearing situation or the scaffolding status cannot comprehensively and accurately assess the overall safety situation on site. By combining the first danger point identification result and the second danger point identification result, the comprehensive risk of the work site can be accurately assessed. For example, when workers do not wear safety belts and there is a dangerous area on the scaffolding, the risk brought by this situation is much higher than the situation where only one of the dangers exists. Through comprehensive analysis, the actual risk situation on site can be more clearly understood. And different risk situations require different levels of emergency response. By generating corresponding levels of alarm prompts (level 1, level 2, level 3) according to different combinations of dangerous points, graded warnings can be achieved. Graded warnings enable safety managers and operators to quickly understand the severity of the risk and take corresponding emergency measures according to different levels. Specifically, for the superimposed risk of "workers not wearing safety belts + dangerous areas on scaffolding", after triggering the third-level emergency alarm, the power supply to the dangerous area is automatically cut off (to prevent electric shock caused by falling), drones are dispatched for real-time monitoring, workers are guided to evacuate through voice calls, and 3D positioning information is pushed to safety officers for rescue assistance; for a single risk scenario, the second-level alarm is used to limit the number of people working in the dangerous area and after automatically generating a maintenance work order, an inspection robot is dispatched to the site for inspection within 5 minutes, or a first-level alarm is used to send a vibration signal to the worker's smart bracelet to remind him to wear a safety belt. In long-term operation, the closed-loop feedback formed by alarm data and disposal results drives strategy iteration. For example, bolt selection and inspection paths are optimized based on the alarm frequency of loose fasteners. At the same time, the tamper-proof alarm records stored in blockchain provide accurate traceability basis for accident responsibility definition. In general, by dynamically associating "personnel-facility" risks and generating graded alarms, not only can accurate quantification of safety hazards be achieved, but also safety prevention and control can be transformed from "post-event disposal" to "pre-event prevention-in-event interception-post-event tracing" through differentiated emergency response mechanisms.

[0049] In summary, the monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to the embodiment of the present application is explained, which transmits the on-site working images collected by the on-site camera to the edge computing device, and identifies the worker status and scaffolding status in the on-site working images on the edge computing device to obtain the first dangerous point identification result indicating whether the worker is wearing a safety belt and the second dangerous point identification result indicating whether there is a dangerous area on the scaffolding, and generates emergency warning prompts of different levels in combination with the first dangerous point identification result and the second dangerous point identification result. In this way, the abnormal wearing of the safety belt can be accurately identified in real time, and the hidden dangers of the scaffolding can be quickly detected, realizing the dynamic assessment of complex risks and graded warning, which helps to improve the accuracy and initiative of safety prevention and control.

[0050] Figure 7 FIG. 1 is a block diagram of a monitoring and alarm device based on AI identification of dangerous points in a field operation environment according to an embodiment of the present application. Figure 7 As shown, according to an embodiment of the present application, a monitoring and alarm device 100 based on AI identification of dangerous points in a field operation environment includes: an image acquisition module 110, which is used to acquire field operation images through a field camera and transmit the field operation images to an edge computing device; a first dangerous point identification module 120, which is used to extract a worker status image from the field operation image on the edge computing device, and identify whether the worker is wearing a safety belt based on the worker status image to obtain a first dangerous point identification result; a second dangerous point identification module 130, which is used to extract a scaffolding status image from the field operation image on the edge computing device, and determine whether there is a dangerous area on the scaffolding based on the scaffolding status image to obtain a second dangerous point identification result; an alarm module 140, which is used to generate an alarm signal on the edge computing device based on the first dangerous point identification result and the second dangerous point identification result.

[0051] Here, those skilled in the art can understand that the specific functions and operations of the various units and modules in the monitoring and alarm device 100 based on AI identification of dangerous points in the on-site working environment have been described in the above reference. Figures 1 to 6 The description of the monitoring and alarm method based on AI identification of dangerous points in the on-site working environment has been introduced in detail, and therefore, its repeated description will be omitted.

[0052] In summary, the monitoring and alarm device 100 based on AI identification of dangerous points in the on-site working environment according to the embodiment of the present application is explained, which transmits the on-site working images collected by the on-site camera to the edge computing device, and identifies the worker status and scaffolding status in the on-site working images on the edge computing device to obtain the first dangerous point identification result indicating whether the worker is wearing a safety belt and the second dangerous point identification result indicating whether there is a dangerous area on the scaffolding, and generates emergency warning prompts of different levels in combination with the first dangerous point identification result and the second dangerous point identification result. In this way, it is possible to accurately and real-timely identify abnormal wearing of safety belts, and quickly detect hidden dangers of scaffolding, realize dynamic assessment of complex risks and graded warning, and help improve the accuracy and initiative of safety prevention and control.

Claims

1. A monitoring and alarm method based on AI identification of dangerous points in the on-site working environment, characterized in that: include: Collecting on-site operation images through an on-site camera and transmitting the on-site operation images to an edge computing device; On the edge computing device, extracting a worker status image from the on-site operation image, and identifying whether the worker wears a safety belt based on the worker status image to obtain a first danger point identification result, including: inputting the on-site operation image into a human target detection network based on a first YOLO model to obtain the worker status image; performing safety belt wearing feature alignment analysis based on a reference image on the worker status image to obtain the first danger point identification result; On the edge computing device, extracting a scaffolding status image from the on-site operation image, and determining whether there is a dangerous area on the scaffolding based on the scaffolding status image to obtain a second dangerous point identification result; On the edge computing device, an alarm signal is generated based on the first danger point identification result and the second danger point identification result.

2. The monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to claim 1 is characterized in that: Performing a safety belt wearing feature alignment analysis based on a reference image on the worker status image to obtain the first danger point identification result includes: Extracting a safety belt wearing specification reference image from a background database of the edge computing device; Inputting the worker status image into a seat belt wearing area target detection network based on a second YOLO model to obtain a seat belt wearing area ROI image; Performing seat belt wearing feature alignment analysis on the seat belt wearing specification reference image and the seat belt wearing area ROI image to obtain a seat belt wearing reference-real image semantic level alignment encoding feature map; Based on the seat belt wearing reference-real image semantic level alignment coding feature map, the first danger point recognition result is obtained.

3. The monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to claim 2 is characterized in that: Performing seat belt wearing feature alignment analysis on the seat belt wearing specification reference image and the seat belt wearing area ROI image to obtain a seat belt wearing reference-real image semantic level alignment encoding feature map, including: Inputting the seat belt wearing specification reference image and the seat belt wearing area ROI image into an image twin extractor including a first seat belt wearing state image feature extractor and a second seat belt wearing state image feature extractor to obtain a seat belt wearing specification reference state feature map and a seat belt wearing real state feature map; The seat belt wearing specification reference state feature map and the seat belt wearing real state feature map are jointly perceived by aligning the seat belt wearing feature structures to obtain the seat belt wearing reference-real image semantic level aligned encoding feature map.

4. The monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to claim 3 is characterized in that: Performing seat belt wearing feature structure alignment and joint perception on the seat belt wearing standard reference state feature map and the seat belt wearing real state feature map to obtain the seat belt wearing reference-real image semantic level alignment encoding feature map, including: Performing feature decoupling on the seat belt wearing specification reference state feature graph and the seat belt wearing actual state feature graph to obtain a set of seat belt wearing specification reference state local feature coding matrices and a set of seat belt wearing actual state local feature coding matrices; Based on the feature phase alignment matching degree between any two local feature coding matrices of the seat belt wearing specification reference state and the local feature coding matrices of the seat belt wearing actual state in the set of the local feature coding matrices of the seat belt wearing specification reference state and the set of the local feature coding matrices of the seat belt wearing actual state, a seat belt wearing feature structure dynamic search alignment is performed on the set of the local feature coding matrices of the seat belt wearing specification reference state and the set of the local feature coding matrices of the seat belt wearing actual state to obtain a set of structurally aligned {seat belt wearing specification reference state local feature coding matrix, seat belt wearing actual state local feature coding matrix} feature pairs; Based on the set of feature pairs of the structural alignment {local feature encoding matrix of the seat belt wearing standard reference state, local feature encoding matrix of the seat belt wearing real state}, the seat belt wearing reference-real image semantic level alignment encoding feature map is obtained.

5. The monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to claim 4 is characterized in that: Based on the set of feature pairs of the structural alignment {local feature encoding matrix of the seat belt wearing specification reference state, local feature encoding matrix of the seat belt wearing real state}, the seat belt wearing reference-real image semantic level alignment encoding feature map is obtained, including: Performing seat belt wearing feature collaborative perception on each structure alignment {seat belt wearing specification reference state local feature encoding matrix, seat belt wearing actual state local feature encoding matrix} in the set of feature pairs of the structure alignment {seat belt wearing specification reference state local feature encoding matrix, seat belt wearing actual state local feature encoding matrix} to obtain a set of seat belt wearing reference-actual state collaborative perception attention weights; Based on the set of seat belt wearing reference-real state collaborative perception attention weights, the set of structurally aligned {seat belt wearing specification reference state local feature encoding matrix, seat belt wearing real state local feature encoding matrix} feature pairs is subjected to attention-driven complementary feature aggregation to obtain the seat belt wearing reference-real image semantic-level aligned encoding feature map.

6. The monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to claim 5 is characterized in that: For each structure alignment {seat belt wearing specification reference state local feature encoding matrix, seat belt wearing actual state local feature encoding matrix} in the set of feature pairs of the structure alignment {seat belt wearing specification reference state local feature encoding matrix, seat belt wearing actual state local feature encoding matrix}, seat belt wearing feature collaborative perception is performed to obtain a set of seat belt wearing reference-actual state collaborative perception attention weights, including: Calculate the association matrix between each structure alignment {seat belt wearing specification reference state local feature coding matrix, seat belt wearing actual state local feature coding matrix} in the set of the structure alignment {seat belt wearing specification reference state local feature coding matrix, seat belt wearing actual state local feature coding matrix} feature pairs to obtain a set of seat belt wearing specification reference-actual state association feature coding matrices; Performing a trace measurement operation on the set of seat belt wearing specification reference-real state associated feature coding matrices to obtain a set of seat belt wearing specification reference-real state associated feature trace measurement values; Based on the set of the seat belt wearing specification reference-real state associated feature trace measurement values, a set of the seat belt wearing reference-real state collaborative perception attention weights is obtained.

7. The monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to claim 6 is characterized in that: Based on the set of the seat belt wearing specification reference-real state associated characteristic trace measurement values, a set of the seat belt wearing reference-real state collaborative perception attention weights is obtained, including: Performing interlocking connection strength manifold structure embedding compactification on the set of seat belt wearing specification reference-real state associated characteristic trace measurement values ​​to obtain an optimized set of seat belt wearing specification reference-real state associated characteristic trace measurement values; Based on the set of optimized seat belt wearing specification reference-real state associated feature trace measurement values, a set of seat belt wearing reference-real state collaborative perception attention weights is obtained.

8. The monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to claim 2 is characterized in that: Based on the safety belt wearing reference-real image semantic level alignment coding feature map, the first danger point identification result is obtained, including: inputting the safety belt wearing reference-real image semantic level alignment coding feature map into a classifier-based state identifier to obtain the first danger point identification result, and the first danger point identification result is used to indicate whether the worker is wearing a safety belt.

9. The monitoring and alarm method based on AI identification of dangerous points in the on-site working environment according to claim 1 is characterized in that: The edge computing device generates an alarm signal based on the first danger point identification result and the second danger point identification result, including: In response to the first danger point identification result being that the worker does not wear a safety belt and the second danger point identification result being that there is a dangerous area on the scaffold, generating a third-level emergency warning prompt; In response to the first danger point identification result being that the worker is wearing a safety belt and the second danger point identification result being that a dangerous area exists on the scaffold, generating a secondary emergency warning prompt; In response to the first danger point identification result being that the worker is not wearing a safety belt and the second danger point identification result being that there is no danger area on the scaffold, a first-level emergency alarm prompt is generated.

10. A monitoring and alarm device based on AI identification of dangerous points in the on-site working environment, characterized in that: include: An image acquisition module is used to acquire on-site operation images through an on-site camera and transmit the on-site operation images to an edge computing device; A first danger point identification module is used to extract a worker status image from the on-site operation image on the edge computing device, and identify whether the worker is wearing a safety belt based on the worker status image to obtain a first danger point identification result, wherein the first danger point identification module includes: an image extraction unit, used to input the on-site operation image into a human target detection network based on a first YOLO model to obtain the worker status image; a safety belt wearing analysis unit, used to perform safety belt wearing feature alignment analysis on the worker status image based on a reference image to obtain the first danger point identification result; A second danger point identification module is used to extract a scaffolding status image from the field operation image on the edge computing device, and determine whether there is a dangerous area on the scaffolding based on the scaffolding status image to obtain a second danger point identification result; An alarm module is used to generate an alarm signal on the edge computing device based on the first danger point identification result and the second danger point identification result.

Citation Information

Cited By

  • Activity monitoring method and system applied to fault zone

    CN120669297A

  • Hazardous chemical substance production scene safety risk prediction method and system based on artificial intelligence

    CN120673539A

  • Engineering safety intelligent management method and system based on AI hidden danger recognition

    CN121095136A

  • Invisible electronic enclosure system of tower crane on construction site and transient closed-loop method of invisible electronic enclosure system

    CN121553851A

  • Live-line work scene dangerous point identification method and system based on combination of large and small models, electronic equipment and readable storage medium

    CN121706958A