Construction site safety dynamic early warning methods, electronic devices and readable storage media
By using edge computing units and machine learning models to identify and provide immediate warnings of abnormal behavior at construction sites, the problem of delayed transmission of safety early warning information at construction sites has been solved, enabling real-time on-site intervention and 24/7 management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIANGXI SCI & TECH NORMAL UNIV
- Filing Date
- 2026-04-16
- Publication Date
- 2026-06-30
AI Technical Summary
The existing construction site safety early warning system suffers from delays in the transmission of warning information, which prevents on-site personnel from being informed of potential dangers in a timely manner, increasing the difficulty of safety management and the risk of accidents.
Edge computing units are used for real-time video stream processing, machine learning models are used to identify abnormal behavior, and warning information, including captured images and warning text, is displayed on the display terminal at the construction site in real time. The system provides real-time warnings by combining visual perception with the on-site display terminal.
It has enabled a shift from "back-end detection" to "on-site intervention," significantly shortening the reaction time from the occurrence of an action to its perception by personnel, reducing the risk of accidents, and achieving 24/7 automated safety management.
Smart Images

Figure CN122049822B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image communication technology, specifically relating to a construction site safety dynamic early warning method, electronic device, and readable storage medium. Background Technology
[0002] Currently, construction sites present complex environments with numerous safety hazards and uncertainties. Real-time monitoring and immediate intervention of abnormal behavior by construction workers has become a core objective of smart construction site development. These abnormal behaviors typically include unauthorized smoking, falls, failure to wear safety protective equipment, and unauthorized entry into high-risk areas.
[0003] However, existing safety warning methods generally suffer from problems of disconnection and delay. Current early warning models are mostly closed-loop systems, meaning that the system only issues an alarm in the monitoring room after detecting abnormal behavior. The drawback of this approach is that personnel involved in the incident and relevant safety management personnel at the construction site cannot be informed of the danger immediately, leading to delays in the transmission of warning information and making it difficult to stop violations or initiate rescue operations in a timely manner. This increases the difficulty of safety management and the risk of accidents. Summary of the Invention
[0004] The purpose of this application is to provide a construction site safety dynamic early warning method, electronic device, and readable storage medium, which can reduce the difficulty of safety management and the risk of accidents.
[0005] To solve the above-mentioned technical problems, this application is implemented as follows:
[0006] In a first aspect, embodiments of this application provide a dynamic early warning method for construction site safety, applied to an edge computing unit, the method comprising:
[0007] Acquire real-time video streams collected by the visual perception unit at the construction site;
[0008] Based on the real-time video stream, a machine learning model is invoked to identify abnormal behavior.
[0009] Upon detecting abnormal behavior, an early warning message is generated; wherein the early warning message is displayed on the display terminal at the construction site to alert personnel in the vicinity of the display terminal.
[0010] Secondly, embodiments of this application also provide a dynamic early warning method for construction site safety, applied to a display terminal at a construction site, the method comprising:
[0011] Based on an asynchronous callback function, the early warning information generated by the edge computing unit is obtained from the message queue maintained by the edge computing unit; wherein, the early warning information is generated when the edge computing unit identifies abnormal behavior, the abnormal behavior is obtained by the edge computing unit calling a machine learning model to identify the real-time video stream, and the real-time video stream is collected by the visual perception unit of the construction site;
[0012] Upon receiving the warning information, a warning insertion mode is executed, switching the transparency of the warning layer corresponding to the warning information from 0 to 1, and simultaneously switching the transparency of the carousel layer corresponding to the display terminal carousel sequence from 1 to 0. The captured image and warning text in the warning information are overlaid and rendered to obtain a rendering result, which is then displayed to warn people around the display terminal. The captured image is an image of abnormal behavior captured by the edge computing unit from the real-time video stream, and the warning text is generated based on the captured image.
[0013] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0014] Fourthly, embodiments of this application provide a computer-readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0015] Fifthly, embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.
[0016] The method for early warning of abnormal behavior at construction sites provided in this application has the following advantages compared to the prior art:
[0017] 1. This solution changes the traditional "closed-loop" early warning model. In existing technologies, abnormal behavior alarms are typically confined to the monitoring center, leaving on-site personnel unaware of potential hazards. This solution pushes early warning information to display terminals at the construction site, delivering alerts directly to the work area. This shifts the focus from "back-end detection" to "on-site intervention," fundamentally solving the problem of disconnected early warning information from the work site and significantly shortening the reaction time from the occurrence of the behavior to personnel awareness.
[0018] 2. By combining visual perception with on-site display terminals, a visual warning is issued instantly upon detecting abnormal behavior (such as smoking in violation of regulations, falling, not wearing a safety helmet, or unauthorized entry). This immediate warning not only reminds the violator to correct their behavior promptly but also alerts nearby workers and safety management personnel to intervene quickly, transforming safety management from reactive post-incident tracking to proactive in-process intervention, effectively reducing the risk of accidents.
[0019] 3. Using display terminals for warnings extends beyond the individual violator to all people in the vicinity. This public warning directly conveys the message that safety regulations are being followed, deterring potential violations. Furthermore, in emergencies such as falls, it quickly draws the attention of those nearby, facilitating immediate mutual assistance and rescue.
[0020] 4. By replacing manual inspections with machine vision, 24 / 7 monitoring of the construction site is achieved. When abnormal behavior occurs, there is no need for dedicated personnel to forward warning information in the background, reducing human communication costs, avoiding missed reports due to human negligence, and realizing the automation upgrade of construction site safety management. Attached Figure Description
[0021] Figure 1 This is one of the flowcharts illustrating the dynamic early warning method for construction site safety provided in some embodiments of this application;
[0022] Figure 2 This is a flowchart illustrating the VB-Block process for processing input features provided in some embodiments of this application;
[0023] Figure 3 This is a flowchart illustrating the AG-Up process for processing input features, provided in some embodiments of this application.
[0024] Figure 4 This is a structural diagram of an improved model provided by some embodiments of this application;
[0025] Figure 5 This is one of the flowcharts illustrating the dynamic early warning method for construction site safety provided in some embodiments of this application;
[0026] Figure 6 These are internal structural diagrams of electronic devices provided in some embodiments of this application. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0028] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0029] An applicable system according to an embodiment of this application includes a visual perception unit, an edge computing unit, and a display terminal.
[0030] The visual perception unit may include, but is not limited to: fixed monitoring equipment set up at the construction site, and drones used for construction site inspection.
[0031] In some embodiments, hardware coupling and data interaction between the visual sensing unit, the edge computing unit, and the display terminal are achieved through industrial-grade communication protocols. Specifically:
[0032] First, the physical link and network architecture are established. In this embodiment, the visual perception unit is connected to the edge computing unit via a switch, and the raw video stream is acquired using a real-time streaming protocol. Simultaneously, the edge computing unit can establish a communication link with the display terminal deployed at the construction site via gigabit Ethernet. To address the challenges of a complex construction site environment and numerous sources of electromagnetic interference, the system can employ double-shielded twisted-pair cables for wired connections, in conjunction with a high-performance switch, to ensure the bandwidth requirements of the video stream transmission and the real-time nature of control commands.
[0033] Secondly, hardware control and logical link initialization are performed. The edge computing unit and the display terminal can be directly connected via a physical interface or establish a command channel based on the Transmission Control Protocol (TCP) over a local area network (LAN). During system startup, the edge computing unit sends a connection request based on the display terminal's unique LAN IP address and dedicated communication port. After successful connection, both parties establish a long-term handshake. The edge computing unit monitors the online status of the display terminal in real time and reserves a high-priority buffer for asynchronous takeover signals, preparing the underlying communication for instantaneous switching warning information.
[0034] The construction site safety dynamic early warning method of this application embodiment can be applied to electronic devices. In specific implementation, the method can be executed by electronic devices, or by components of electronic devices, such as processors, chips, or chip systems, or by logic modules or software that implement all or part of the functions of electronic devices. In practical applications, electronic devices can be edge computing units, display terminals, etc., and further, edge computing units can be servers, service platforms, cloud platforms, distributed systems, etc.
[0035] In an exemplary embodiment, this application proposes a dynamic early warning method for construction site safety, which can be applied to the edge computing unit of the above-mentioned system. The dynamic early warning method for construction site safety provided by the embodiments of this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.
[0036] Reference Figure 1 The method includes steps 102-106. Wherein:
[0037] Step 102: Obtain the real-time video stream collected by the visual perception unit at the construction site.
[0038] Real-time video streams are acquired in real time by visual perception units, such as high-definition surveillance cameras deployed at construction site entrances, living areas, and work sites, or by drones conducting low-altitude inspections at these locations. It can be understood that real-time video streams can contain on-site video data with varying lighting conditions, different work backgrounds, and different shooting angles.
[0039] In some embodiments, high-definition surveillance cameras can be used for real-time data acquisition, while drones can be used for periodic data acquisition.
[0040] Step 104: Based on the real-time video stream, call the machine learning model to identify abnormal behavior.
[0041] In some embodiments, abnormal behavior may include, but is not limited to, smoking in violation of regulations, falling down, not wearing safety protective equipment, and illegally entering high-risk areas.
[0042] In some embodiments, the algorithms used to build machine learning models may include, but are not limited to: traditional object detection methods (such as Haar, HOG + SVM, AdaBoost, etc.) and deep learning-based object detection algorithms (such as YOLO series, convolutional neural networks, etc.).
[0043] It should be noted that illegal smoking, as a typical cause of fires, has become a challenge for construction site safety management due to its small size and the fact that the behavior is concealed.
[0044] To address the aforementioned challenges, specifically the pain points of construction site scenarios such as semi-transparent smoke obscuring the view due to dust, low contrast due to strong backlighting, and the easy submersion of subtle features like cigarette butts, embers, and smoke by noise, some embodiments involve using a machine learning model to identify abnormal behavior based on the real-time video stream, including:
[0045] Based on the real-time video stream, a machine learning model is invoked to generate an identification result indicating whether smoking is in violation of regulations.
[0046] The machine learning model is constructed based on the Real-Time Detection Transformer (RT-DETR); the standard convolutions in the RT-DETR are replaced with Visual State-Conv (VS-Conv) convolutions in the neck network of the machine learning model.
[0047] The visual state space convolution includes a two-dimensional convolutional layer, a batch normalization layer, and a SiLU nonlinear activation function.
[0048] In this embodiment, firstly, the first input feature of the input visual state spatial convolution is encoded into spatial features through a two-dimensional convolutional layer to extract basic texture information; secondly, the feature distribution is standardized and reconstructed through a batch normalization (BN) layer to eliminate the internal covariate shift of the encoded features; finally, the SiLU activation function is used to introduce smooth nonlinear characteristics, thereby enhancing the machine learning model's response to weak smoke features. The forward propagation formula of VS-Conv is:
[0049]
[0050] in, This represents the feature map of the final output of the VS-Conv operator; This represents the original feature map input to the operator, i.e., the first input feature; Indicates the kernel size as Two-dimensional spatial convolution operation; and These represent the channel mean and variance of the input features in the current batch, respectively. This represents a very small constant to prevent the denominator from being zero; and These represent the learnable scaling factor and bias term in the batch normalization layer, respectively. This represents the SiLU nonlinear activation function.
[0051] In some embodiments, the kernel size of the two-dimensional spatial convolution operation Set to 3×3, stride of 1, and padding of 1 to ensure consistent input and output feature map sizes; learnable scaling factor for batch normalization layers. The initial value is 1, and the bias term is... The initial value is 0; a minimal constant to prevent the denominator from being zero. The value is 1×10 -5 This ensures a linear response to weak features with low amplitude and nonlinear suppression of features with high amplitude.
[0052] In this embodiment, the visual state space convolution is a lightweight perception operator specifically designed to meet the needs of extracting weak features under harsh working conditions at construction sites. This operator constructs a cascaded perception architecture of spatial filtering, statistical normalization, and nonlinear activation. It performs three-level cascaded optimization for the semi-transparent textures and small targets in construction site scenes. Utilizing the smooth nonlinear characteristics of the SiLU activation function, it effectively solves the problem of traditional standard convolution being easily overwhelmed by noise when processing weak features. This allows for accurate preservation of the edge and texture information of weak targets under strong noise interference, thereby improving robustness in handling high-frequency noise such as construction site dust and strong light.
[0053] To address the aforementioned challenges, specifically the common issues in construction site environments such as arm obstruction, faces turned away from the camera, obstruction by construction equipment, and incomplete target information due to dense crowds, some embodiments replace the residual blocks in the RT-DETR with a Vision Bottleneck-Block (VB-Block) module in the backbone network of the machine learning model.
[0054] The visual bottleneck module is constructed based on the State Space Model (SSM).
[0055] In this embodiment, refer to Figure 2 First, the second input feature of the SSM is preprocessed through linear projection and depthwise convolution. Second, the preprocessed 2D features are unfolded into a 1D sequence along the four diagonal directions. Then, based on the four-directional recursive scanning mechanism of the Selective Scanning 2D (SS2D) unit, linear recursive modeling is performed for each 1D sequence using state space parameters to aggregate global context information. Finally, the scanning results in the four directions are restored to 2D feature maps and weighted fusion is performed. The formula for VB-Block is:
[0056]
[0057] in, Indicates the second input feature; A linear projection layer representing the input and output; Representation layer normalization; This indicates that the two-dimensional feature is expanded into the first... Transformation operations on one-dimensional sequences in each direction =1,2,3,4; Indicates the first A recursive state-space model in one direction; This indicates a sequence reversal and recombination operation; This represents the gated branch signal used for feature selection; This indicates the enhanced features with global context attributes extracted by VB-Block.
[0058] In some embodiments, the dimensions of the input and output linear projection layers must be consistent with the number of input feature channels. For example, for an RT-DETR model with an input size of 640×640, the number of input feature channels of the VB-Block in the backbone network can be set to 64, 128, 256, and 512 respectively; the kernel size of the depthwise convolution is 3×3, the stride is 1, and the padding is 1; the four diagonal directions of the four-way recursive scan are top left-bottom right, top right-bottom left, bottom left-top right, and bottom right-top left; the state dimension of the SSM is set to 16, and the discretization stride is set to 1 / 16; the gated branch signal is generated by linearly projecting the second input feature through a 1×1 convolution, and the activation function is the Sigmoid function to achieve adaptive selection of effective features.
[0059] This embodiment leverages the global receptive field characteristics of SSM (Self-Signaling Model), specifically its four-way recursive scanning mechanism, to overcome the technical bottleneck of limited receptive field in traditional convolution. A VB-Block module is designed to replace the traditional residual block, specifically addressing the problems of long-distance spatiotemporal dependency capture and feature fragmentation in construction site occlusion scenarios. Specifically, addressing the common issues of arm or face occlusion in smoking behavior, the VB-Block module achieves global image context modeling with linear computational complexity, capturing long-distance spatiotemporal topological dependencies between images. This maintains high-precision recognition even when the target pose is incomplete or severely occluded, effectively reducing the false negative rate.
[0060] To address the aforementioned challenges, specifically in construction site scenarios where cameras are typically installed at high positions, resulting in targets such as cigarette butts and safety helmet clips being far away and small in size, the upsampling process in the Feature Pyramid Network (FPN) is prone to semantic loss and blurred edges for small targets. In some embodiments, the backbone network of the machine learning model uses an attention gating-up (AG-Up) module to replace the nearest neighbor interpolation-based upsampling module in the RT-DETR.
[0061] The attention-gated upsampling module includes a dual-path reconstruction unit and a gating filtering unit.
[0062] This embodiment designs a dual-path reconstruction and gating screening architecture adapted to small target detection, replacing the conventional single interpolation upsampling, i.e., the upsampling module based on nearest neighbor interpolation.
[0063] In some embodiments, refer to Figure 3 The first path of the dual-path reconstruction unit is a semantic reconstruction path, which utilizes the learnable properties of transposed convolution (Conv Transpose) to recover the nonlinear details of the third input feature through reverse mapping. The second path of the dual-path reconstruction unit is an auxiliary interpolation path, which uses bilinear interpolation to provide a basic spatial amplification benchmark for the third input feature. After the two features are fused, they are adaptively weighted through a gating filter unit. This gating filter unit is a channel attention gate generated by the global descriptor, thereby accurately suppressing the aliasing noise generated by upsampling. The third input feature is the feature input to the attention gate upsampling module. The AG-Up formula is:
[0064]
[0065] in, This represents the high-resolution feature map output by the AG-Up module; This represents the low-resolution feature map of the input, i.e., the third input feature; This indicates a global average pooling operation; Represents the interaction between dimensionality reduction and features. convolution; This represents the Hardsigmoid activation function; This represents the generated channel attention weight vector; This indicates the transpose convolution operation; This indicates a bilinear interpolation upsampling operation; This indicates a feature concatenation operation along the channel dimension; This indicates the intermediate fusion features after splicing; This represents element-wise multiplication based on a broadcast mechanism; Represents the output layer Merge convolutions.
[0066] In some embodiments, the transposed convolution kernel size is 4×4, the stride is 2, and the padding is 1, ensuring that the feature map size is magnified by a factor of 2 after upsampling; the magnification factor of bilinear interpolation upsampling is consistent with that of the transposed convolution, which is 2 times; the 1×1 convolution used for dimensionality reduction and feature interaction has an output channel number that is 1 / 4 of the input channel number; the global average pooling operation compresses the two-dimensional features of each channel into a 1×1 global descriptor; the threshold of the Hardsigmoid activation function is set to [-3, 3] to ensure the stable generation of attention weights; the number of output channels of the 1×1 fusion convolution in the output layer is consistent with the number of channels of the input low-resolution feature map.
[0067] In this embodiment, the AG-Up module abandons the conventional linear interpolation method and adopts a dual-path reconstruction and gating filtering strategy. It utilizes the learnable properties of transposed convolution to recover the nonlinear details of features and precisely suppresses aliasing noise generated during upsampling through a channel attention gating mechanism. This design enables the model to retain the geometric features of distant, small targets to the greatest extent when fusing features at multiple scales, solving the problem of semantic ambiguity of small targets caused by traditional interpolation algorithms and significantly improving the detection performance of extremely small targets.
[0068] In some embodiments, taking abnormal behavior as a violation (such as smoking in violation of regulations, not wearing a helmet, etc.) as an example, the training process of the above machine learning model can be specifically carried out in the following steps:
[0069] Step 1: Multi-source data mining and construction of a multi-modal scenario library.
[0070] Step 2: Target feature refinement cleaning and classification strategy.
[0071] Step 3: Behavioral representation and truth labeling based on fine-grained semantics.
[0072] Step 4: Divide the dataset into a validation set, a training set, and a test set, and finally process it into a recognition format for machine learning models.
[0073] Step 5: Use the partitioned dataset to iteratively train the untrained machine learning model.
[0074] For step one:
[0075] The dataset can be constructed using a hybrid strategy combining fixed surveillance data collection with drone patrols. First, high-definition surveillance cameras are deployed at construction site entrances, living areas, and work surfaces. Drones are then used regularly for low-altitude patrols to collect on-site video data under varying lighting conditions, with different work backgrounds, and from different shooting angles. The collected raw images are then manually screened and cleaned, removing severely blurry images, images with obstructed targets, and images with excessive repetition.
[0076] In some embodiments, in order to improve the adaptability of machine learning models to complex environments, image enhancement techniques can be used to process low-contrast images to construct a polymorphic scene library containing abnormal behaviors such as smoking and not wearing a helmet.
[0077] For step two:
[0078] Automated cleaning is performed on the massive data in the polymorphic scene library, and invalid frames are removed based on resolution, completeness, and target visibility. A category balancing strategy is adopted to divide the filtered samples into significant violation samples and hidden violation samples, and time series sorting is performed based on the coherence of abnormal actions such as smoking.
[0079] For step three:
[0080] The cleaned images were annotated with fine-grained tools. The annotation range not only covered physical targets such as cigarettes and mobile phones, but also behavioral representations of the hand approaching the face and the flow of smoke that are unique to smoking.
[0081] For step four:
[0082] To scientifically evaluate model performance, the labeled dataset was randomly divided into training, validation, and test sets in an 8:1:1 ratio. The training set was used for iterative updates of model parameters, the validation set for hyperparameter tuning during training, and the test set for evaluating the model's generalization ability. Finally, the dataset was processed into a recognition format suitable for machine learning models.
[0083] Step 106: Upon detecting abnormal behavior, generate early warning information; wherein the early warning information is used for display on the display terminal at the construction site to alert personnel in the vicinity of the display terminal.
[0084] In some embodiments, the warning information may be a captured image of abnormal behavior obtained by the edge computing unit from a real-time video stream, and a warning text generated based on the abnormal behavior.
[0085] In some embodiments, the next batch of images in the real-time video stream is identified if no abnormal behavior is detected.
[0086] In some embodiments, the surrounding personnel can be people around the location where the abnormal behavior occurs, such as workers smoking in violation of regulations, workers who have fallen, managers, etc.
[0087] The embodiments of this application have the following beneficial effects:
[0088] 1. This solution changes the traditional "closed-loop" early warning model. In existing technologies, abnormal behavior alarms are typically confined to the monitoring center, leaving on-site personnel unaware of potential hazards. This solution pushes early warning information to display terminals at the construction site, delivering alerts directly to the work area. This shifts the focus from "back-end detection" to "on-site intervention," fundamentally solving the problem of disconnected early warning information from the work site and significantly shortening the reaction time from the occurrence of the behavior to personnel awareness.
[0089] 2. By combining visual perception with on-site display terminals, a visual warning is issued instantly upon detecting abnormal behavior (such as smoking in violation of regulations, falling, not wearing a safety helmet, or unauthorized entry). This immediate warning not only reminds the violator to correct their behavior promptly but also alerts nearby workers and safety management personnel to intervene quickly, transforming safety management from reactive post-incident tracking to proactive in-process intervention, effectively reducing the risk of accidents.
[0090] 3. Using display terminals for warnings extends beyond the individual violator to all people in the vicinity. This public warning directly conveys the message that safety regulations are being followed, deterring potential violations. Furthermore, in emergencies such as falls, it quickly draws the attention of those nearby, facilitating immediate mutual assistance and rescue.
[0091] 4. By replacing manual inspections with machine vision, 24 / 7 monitoring of the construction site is achieved. When abnormal behavior occurs, there is no need for dedicated personnel to forward warning information in the background, reducing human communication costs, avoiding missed reports due to human negligence, and realizing the automation upgrade of construction site safety management.
[0092] In some embodiments, to avoid misidentification of abnormal behavior caused by momentary flickering, further judgment can be made when abnormal behavior is identified.
[0093] In this embodiment, further judgment refers to performing a time-series voting anti-interference judgment based on a sliding window. Specifically, to address the potential for momentary flickering in the detection results, the edge computing unit maintains a fixed-length state queue for each tracked target within the entire frame. After each frame of detection is completed, the edge computing unit pushes the current recognition result into the state queue and calculates the cumulative confidence score of the target category within the sliding window. Only when the cumulative confidence score exceeds a preset confirmation threshold is the abnormal behavior considered to have actually occurred, thereby effectively filtering out false alarms caused by sudden changes in light and shadow. Here, the category refers to the aforementioned abnormal behaviors, such as illegal smoking or a person falling, and the target category refers to the same category, for example, both being illegal smoking. Specifically:
[0094] The step of generating early warning information upon detecting abnormal behavior includes:
[0095] If abnormal behavior is detected, it is determined whether the abnormality rate within the sliding window corresponding to the abnormal behavior reaches a preset abnormality rate; wherein, the sliding window is used to select a preset number of images in the real-time video stream.
[0096] If the anomaly rate reaches the preset anomaly rate, the abnormal behavior is determined to be valid, and an early warning message is generated. If the anomaly rate does not reach the preset anomaly rate, the abnormal behavior is determined to be invalid, i.e., no abnormal behavior is determined.
[0097] In this embodiment, the formula for determining the timing-based voting interference immunity based on the sliding window is as follows:
[0098]
[0099] in, This represents the anomaly confidence score within the current time window; Indicates the frame length of the sliding window; Indicates the frame order index within the window; This represents the indicator function, which takes the value of 1 when the class consistency and confidence level meet the criteria, and 0 otherwise. Indicates the first The category label identified in the frame; Indicates the preset category of abnormal behavior; Indicates the first The prediction confidence score of the frame target; This represents the confidence threshold for a valid determination of a single frame. This represents the cumulative score threshold for confirming abnormal behavior, also known as the preset abnormality rate. This represents the final output stable anomaly determination signal.
[0100] In some embodiments, the frame length N of the sliding window can be set to 5, that is, the preset frame number of the image selected by the sliding window in the real-time video stream is 5, which corresponds to a 200ms time window of a 25FPS video stream, taking into account both real-time performance and anti-interference capability. It can be set to 0.5, which means that a single frame is only considered a valid recognition when the confidence level of the single frame detection result is greater than 0.5; It can be set to 0.6, meaning that at least 3 frames within a 5-frame window must have valid recognition results before the abnormal behavior is finally determined to have actually occurred, thus effectively filtering out instantaneous false alarms caused by sudden changes in light and shadow or dust interference.
[0101] In some embodiments, the display terminal shows the warning information to alert people in the vicinity of the display terminal.
[0102] In some embodiments, displaying the warning information includes:
[0103] Upon receiving the aforementioned warning information, the warning insertion mode is executed.
[0104] In the warning insertion mode, the transparency of the warning layer corresponding to the warning information is switched from 0 to 1; wherein, when the transparency of the warning layer is switched from 0 to 1, the transparency of the carousel layer corresponding to the carousel sequence of the display terminal is switched from 1 to 0.
[0105] The captured image and warning text in the warning information are overlaid and rendered to obtain a rendering result; wherein, the captured image is an image corresponding to abnormal behavior captured by the edge computing unit from the real-time video stream, and the warning text is generated based on the captured image.
[0106] The rendering results are displayed.
[0107] In this embodiment, the display terminal employs a dual-layer dashboard rendering engine, achieving takeover by dynamically adjusting the transparency weight coefficients of the layers. Specifically, the edge computing unit automatically captures on-site images and overlays automatically generated warning text to construct an early warning layer; simultaneously, the regular carousel layer is placed at the bottom layer. When takeover is triggered, the transparency weight coefficient of the early warning layer quickly switches to its maximum value, completely covering the display content of the carousel layer, achieving immediate exposure and deterrence of violations. After the set playback duration is reached, the system automatically resumes the regular carousel sequence.
[0108] In some embodiments, addressing the challenges of dispersed deployment of display terminals (such as advertising screens) in construction site scenarios, unstable network environments, difficulties in synchronous management of multiple terminals, and the need for smooth switching between advertising playback and early warning insertion, this embodiment designs an asynchronous callback takeover mechanism based on a high-priority message queue. This mechanism performs asynchronous takeover and layer rendering of display terminals, adapting to the complex deployment environment of construction sites with multiple terminals and strong electromagnetic interference. Compared to traditional push architectures, which cannot support immediate takeover needs (specifically, traditional playback architectures typically use a preset list mode and lack priority scheduling capabilities), this embodiment can achieve a smooth asynchronous switch from a normal display stream to an instantaneous early warning stream when emergency abnormal behavior is detected. This allows early warning materials to occupy the screen immediately, thereby achieving optimal deterrence and reducing rescue time.
[0109] When an abnormal behavior is determined to have actually occurred, the edge computing unit generates an alert message and pushes it into a message queue. The display terminal then retrieves the alert information generated by the edge computing unit from the message queue maintained by the edge computing unit via an asynchronous callback function. Furthermore, it immediately suspends regular carousel threads, such as advertising threads, thereby interrupting the current playback content and forcibly rendering the alert information at the top. The formulas for asynchronous takeover and rendering output are as follows:
[0110]
[0111] in, This indicates the operating mode of the display terminal in the next moment; and These represent the early warning insertion mode and the regular carousel mode, respectively. This represents an abnormal behavior payload in the message queue; This indicates an empty set, representing a warning with no pending processing. Indicates the duration the warning information has been displayed; This indicates the minimum mandatory playback duration for warning messages; Indicates time Display the mixed pixel matrix output by the terminal; This represents the transparency weight coefficient of the warning layer. When takeover is triggered, this value quickly switches to 1, thus changing the transparency of the warning layer from 0 to 1. This refers to captured images, such as images of the moment a violation occurs; This indicates a warning message; This indicates an asynchronous overlay rendering operation between the image layer and the text layer; This indicates that the next frame of the carousel sequence should be retrieved. This indicates a preset carousel sequence, such as an advertising carousel sequence.
[0112] In some embodiments, the minimum mandatory playback duration of the warning information It can be set to 10 seconds to ensure that people in the vicinity fully receive the warning information; the transparency weighting coefficient of the warning layer. The value is 0 when there is no abnormal behavior, and it linearly switches to 1 within 100ms after the warning is triggered, so as to achieve full-screen coverage of the warning information without any lag; the dedicated communication port for TCP long connection is set to 8088, the heartbeat keep-alive packet sending interval is 30 seconds, and the timeout reconnection time is set to 5 seconds; the depth of the message queue of the high priority buffer is set to 10 to ensure that the warning instructions are transmitted without loss or delay.
[0113] This application embodiment displays warning information through a display terminal at the construction site, realizing an instant visual feedback mechanism on site. This allows violators to realize that they have been recorded, and safety hazards can be corrected on-site at the nascent stage, eliminating the huge time blind spot between the discovery of violations and human intervention.
[0114] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0115] For ease of understanding, the effectiveness of the dynamic early warning method for construction site safety proposed in this application will be explained below in conjunction with experimental verification and performance evaluation:
[0116] To objectively verify the effectiveness and robustness of the construction site safety dynamic early warning method proposed in this application embodiment, rigorous experimental tests were conducted in a unified hardware and software environment.
[0117] The experimental environment was built on a high-performance computing platform, with hardware configuration including an Intel Core i9-13900K processor, 64GB DDR5 memory, and an NVIDIA GeForce RTX 4090 (24GB VRAM) graphics accelerator card; the software environment ran on the Ubuntu 22.04 LTS operating system, and the deep learning framework used was PyTorch 2.0.1, along with the CUDA 11.8 acceleration library.
[0118] The experimental dataset consists of 5,000 images of the construction site collected by fixed monitoring and drone inspections, covering various complex working conditions such as strong light, dust, and nighttime, and including typical violations such as smoking and not wearing safety helmets. The dataset is randomly divided into training, validation, and test sets in an 8:1:1 ratio, and Mosaic and Mixup data augmentation strategies are used to expand the diversity of the samples.
[0119] During model training, the input image size was uniformly adjusted to 640×640, the batch size was set to 16, and the total number of iterations was 300. The AdamW optimizer was used, with an initial learning rate of 0.0001 and a weight decay coefficient of 0.05. A cosine annealing strategy was introduced to dynamically adjust the learning rate to ensure that the model converges to the global optimum.
[0120] Experiment 1:
[0121] To verify the effectiveness of the improved RT-DETR core modules proposed in this application, a detailed ablation experiment was conducted. The experiment used the standard RT-DETR model as a benchmark, with an initial detection accuracy (mAP@0.5) of 79.2.
[0122] Table 1: Ablation Experiment Results
[0123]
[0124] As shown in Table 1, after introducing VS-Conv into the backbone network, the model's feature purification ability under dust and strong light interference is significantly enhanced, and the detection accuracy is improved to 80.8%, with a gain of 1.6%. Secondly, VB-Block is further added, utilizing its four-way recursive scanning mechanism to solve the problems of limb occlusion and side-view recognition, capturing long-distance spatiotemporal dependencies, and further improving the accuracy to 82.1%. Finally, AG-Up is integrated into the neck network, which effectively suppresses aliasing noise in the upsampling process through dual-path reconstruction and gating screening mechanisms, significantly optimizing the detection performance of distant small targets. Ultimately, the improved model proposed in this application (corresponding to experiment number 4 in Table 1) achieved an mAP@0.5 of 83.4% on the test set, a cumulative improvement of 4.2% compared to the baseline model, fully demonstrating the positive gain effect of each improved module in complex construction site scenarios.
[0125] The structure of the improved model in experiment number 4 can be referred to Figure 4 .
[0126] Experiment 2:
[0127] Table 2: Performance Comparison of Different Detection Models in Construction Site Scenarios
[0128]
[0129] The experimental results in Table 2 show that the improved model in this application demonstrates significant advantages in all key performance indicators. Its mAP@0.5 is as high as 83.4%, which is significantly better than YOLOv8n and RT-DETR models. In particular, the recall rate reaches 81.2%, which effectively solves the problem of missed detection under complex working conditions. At the same time, thanks to the lightweight design of VS-Conv and the linear complexity of VB-Block, the model inference speed breaks through to 112 FPS.
[0130] Experiment 3:
[0131] To further verify the specific performance of this application embodiment for the core pain points of construction sites, this embodiment supplements the core target specific detection experiment and the occlusion scene robustness experiment. The experimental environment and dataset are completely consistent with the aforementioned main experiment, only the test set subset is changed. The experimental results are shown in Tables 3 and 4.
[0132] Table 3: Comparison of Performance of Specific Detection for Core Targets
[0133]
[0134] Experimental results show that the improved model in this application embodiment significantly improves the detection performance of core control targets on construction sites. Specifically, the AP@0.5 for detecting small cigarette butt targets, which is a core pain point, is improved by 9.5% compared to the benchmark model, and the AP@0.5 for detecting smoking behavior is improved by 6.7%. This fully demonstrates the optimization effect of this application embodiment on the specific pain points of construction site scenarios, rather than a general accuracy improvement.
[0135] Table 4: Comparison of detection performance in occluded scenarios
[0136]
[0137] Mild occlusion refers to a target occlusion ratio of ≤30%, while severe occlusion refers to a target occlusion ratio of 30%-60%. Experimental results show that the improved model in this embodiment achieves a 12.7% improvement in mAP@0.5 compared to the baseline model in severe occlusion scenarios, fully demonstrating the global modeling capability of the VB-Block module for construction site occlusion scenarios and solving the problem of missed detection in occlusion scenarios.
[0138] Meanwhile, this embodiment also tested the real-time performance of the asynchronous takeover mechanism. In 1,000 events triggered by abnormal behavior, the average latency from the completion of the abnormal behavior judgment by the edge computing unit to the completion of the full-screen rendering of the warning information by the display terminal was 127ms, and the maximum latency was 215ms. This fully meets the design requirements of millisecond-level takeover, with a 100% success rate in switching, and no packet loss, stuttering, or content conflict issues.
[0139] In summary, the embodiments of this application have the following advantages compared to the prior art:
[0140] 1. To address the harsh working conditions of construction sites, including pervasive dust and drastic changes in lighting, a Visual State Space Convolution (VS-Conv) operator was designed to replace standard convolution. This operator constructs a cascaded perception architecture of spatial filtering, statistical normalization, and nonlinear activation, effectively solving the problem of feature inactivation in traditional convolution when processing semi-transparent smoke textures or extremely small cigarette butt flames by leveraging the smooth nonlinear characteristics of the SiLU activation function. Through deep purification of visual state space features, it can accurately preserve the edge and texture information of faint targets under strong noise interference, significantly improving the model's robustness in detecting concealed violations.
[0141] 2. The proposed Visual Bottleneck Module (VB-Block) overcomes the technical bottleneck of limited receptive field in traditional convolution by utilizing a four-way recursive scanning mechanism based on a two-dimensional state-space model. Addressing the common issues of arm or face occlusion in smoking behavior, this module achieves global image context modeling with linear computational complexity. It can capture long-distance spatiotemporal topological dependencies between images, thus maintaining high-precision recognition even when the target pose is incomplete or severely occluded, effectively reducing the false negative rate.
[0142] 3. The attention-gated upsampling module (AG-Up) designed in the feature fusion stage solves the problem of semantic ambiguity of small targets caused by traditional interpolation algorithms. This module abandons the conventional linear interpolation method and adopts a dual-path reconstruction and gating filtering strategy. It utilizes the learnable properties of transposed convolution to recover the nonlinear details of features and accurately suppresses aliasing noise generated during upsampling through a channel attention gating mechanism. This design enables the model to retain the geometric features of distant small targets to the greatest extent during multi-scale feature fusion, significantly improving the detection performance of extremely small targets.
[0143] 4. A millisecond-level safety management closed loop based on an asynchronous takeover mechanism was constructed, solving the problem of delayed feedback in traditional monitoring systems. By breaking down the communication barriers between the visual algorithm server and the on-site advertising display terminal, and utilizing TCP long connections and asynchronous callback technology, millisecond-level forced interruption and takeover of regular commercial advertisements by early warning information was achieved. Once the system determines that a violation has occurred, the large screen on-site immediately displays the captured image and warning message, realizing a shift from passive background recording to immediate on-site deterrence, greatly improving the effectiveness of construction site safety management, and simultaneously activating the safety education value of idle advertising screens.
[0144] In one exemplary embodiment, reference is made to Figure 5 This application proposes a dynamic early warning method for construction site safety, which can be applied to the display terminal of the above-mentioned system. The dynamic early warning method for construction site safety includes:
[0145] Step 502: Based on the asynchronous callback function, obtain the early warning information generated by the edge computing unit from the message queue maintained by the edge computing unit; wherein, the early warning information is generated when the edge computing unit identifies abnormal behavior, the abnormal behavior is obtained by the edge computing unit calling a machine learning model to identify the real-time video stream, and the real-time video stream is collected by the visual perception unit of the construction site.
[0146] Step 504: Upon receiving the warning information, execute the warning insertion mode, switch the transparency of the warning layer corresponding to the warning information from 0 to 1, and simultaneously switch the transparency of the carousel layer corresponding to the display terminal carousel sequence from 1 to 0. Overlay and render the captured image and warning text in the warning information to obtain a rendering result, and display the rendering result to warn people around the display terminal; wherein, the captured image is an image corresponding to abnormal behavior captured by the edge computing unit from the real-time video stream, and the warning text is generated based on the captured image.
[0147] It should be noted that the embodiments of the construction site safety dynamic early warning method applied to the display terminal are basically the same as the embodiments of the construction site safety dynamic early warning method applied to the edge computing unit, and will not be repeated here.
[0148] This application also provides an electronic device, such as... Figure 6 As shown, the electronic device 600 includes:
[0149] One or more processors 610;
[0150] The memory 620 stores one or more programs that, when executed by one or more processors 610, enable the one or more processors 610 to implement the site safety dynamic early warning method described in any of the above embodiments.
[0151] The memory 620, as a non-transitory network system, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, the memory 620 may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 620 may optionally include remotely located memories 620 relative to the processor 610, which can be connected to the processor 610 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0152] The memory 620 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 620 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 620 and called by the processor 610 to execute the methods of the embodiments of this application.
[0153] The processor 610 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0154] In some embodiments, the electronic device further includes:
[0155] Input / output interfaces are used to implement information input and output;
[0156] The communication interface is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0157] The bus transmits information between various components of the device, such as processor 610, memory 620, input / output interfaces, and communication interfaces.
[0158] The processor 610, memory 620, input / output interface, and communication interface can communicate with each other within the device via a bus.
[0159] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0160] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0161] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0162] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A dynamic early warning method for construction site safety, characterized in that, The construction site safety dynamic early warning method, applied to edge computing units, includes: Acquire real-time video streams collected by the visual perception unit at the construction site; Based on the real-time video stream, a machine learning model is invoked to identify abnormal behavior. The abnormal behavior includes illegal smoking. The process of identifying abnormal behavior based on the real-time video stream and calling a machine learning model includes: Based on the real-time video stream, a machine learning model is invoked to generate an identification result of whether smoking is in violation of regulations; The machine learning model is constructed based on the real-time detection converter RT-DETR; the standard convolutions in the RT-DETR are replaced with visual state space convolutions in the neck network of the machine learning model. The forward propagation of the visual state space convolution is achieved through the following formula: Among them, the The feature map representing the final output of the visual state space convolution; The first input feature represents the input to the visual state space convolution; Indicates the kernel size as Two-dimensional spatial convolution operation; the With the These represent the channel mean and variance of the current batch of input features, respectively; This represents a minimal constant to prevent the denominator from being zero; the stated With the These represent the learnable scaling factor and bias term in the batch normalization layer, respectively; This represents the SiLU nonlinear activation function; The residual blocks in the RT-DETR are replaced with visual bottleneck modules in the backbone network of the machine learning model. The formula for the visual bottleneck module is as follows: Among them, the The second input feature is represented; The linear projection layer representing the input and output; Representation layer normalization; the This indicates that the two-dimensional feature is expanded into the first... Transformation operations on one-dimensional sequences in each direction =1,2,3,4; the aforementioned Indicates the first A recursive state-space model in each direction; the This indicates a sequence reverse recombination operation; the This represents the gated branch signal used for feature selection; This refers to the enhanced features with global context attributes extracted by the visual bottleneck module; The attention-gated upsampling module in the RT-DETR is replaced with the nearest neighbor interpolation-based upsampling module in the backbone network of the machine learning model. The formula for the attention-gated upsampling module is as follows: Among them, the This represents the high-resolution feature map output by the attention-gated upsampling module; The third input feature represents the input; This represents a global average pooling operation; the... Represents the interaction between dimensionality reduction and features. Convolution; the This represents the Hardsigmoid activation function; the... This represents the generated channel attention weight vector; This represents the transpose convolution operation; the... This indicates a bilinear interpolation upsampling operation; the This represents a feature concatenation operation along the channel dimension; the This indicates the intermediate fusion features after splicing; the This represents element-wise multiplication based on a broadcast mechanism; the... Represents the output layer Merge convolutions; Upon detecting abnormal behavior, an early warning message is generated; wherein the early warning message is displayed on the display terminal at the construction site to alert personnel in the vicinity of the display terminal.
2. The construction site safety dynamic early warning method according to claim 1, characterized in that, The step of generating early warning information upon detecting abnormal behavior includes: If abnormal behavior is detected, it is determined whether the abnormality rate within the sliding window corresponding to the abnormal behavior reaches a preset abnormality rate; wherein, the sliding window is used to select a preset number of frames of images in the real-time video stream; If so, the abnormal behavior is confirmed to be valid, and an early warning message is generated.
3. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein when the program or instructions are executed by the processor, they implement the steps of the construction site safety dynamic early warning method as described in any one of claims 1-2.
4. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions, which, when executed by a processor, implement the steps of the construction site safety dynamic early warning method as described in any one of claims 1-2.