Comprehensive identification method for unsafe behaviors of construction site personnel
By performing edge detection and attitude analysis on construction site images, combined with monocular projection mapping and instance segmentation technology, the unsafe behavior of construction site workers is identified, and the problem of difficulty in identifying "pressure situations" in the existing technology is solved, and high-precision identification and risk warning of the approaching edge and tilting actions are achieved.
Patent Information
- Application Number
- CN202510580200.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing construction site image recognition system is difficult to effectively identify the unsafe behavior of operators in high altitude operations, side-side operations, etc., especially those "pressure situations" that do not exceed the physical boundary but have a center of gravity close to the boundary or have a dynamic outward tendency.
By performing edge detection, Hough transformation and spatial aggregation on the construction site monitoring images, a boundary segment connection diagram is constructed, a monocular projection mapping structure and buffer zone are established, combined with instance segmentation and pose estimation, the operator's action center of gravity points and attitude spatial expression are extracted, the angle between the main attitude axis and the boundary normal is calculated, the direction consistency matrix and the extraverted behavioral mark are generated, and the resident segment is identified in the spatial risk mask area and the oppressive violation proximity behavior label is output.
It significantly improves the accuracy of identification of near-side, tilting actions and compression behaviors on the construction site, and has high real-time, robustness and applicability, and can identify unsafe behaviors in advance and provide risk warnings.
Smart Images

Figure CN120088866A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image analysis and recognition at construction sites. More specifically, the present invention relates to a comprehensive method for identifying unsafe behaviors of construction site personnel. Background Art
[0002] At construction sites, in positions such as high-altitude operations, floor edges, bridge edges, and scaffolding platforms, it is a very common scenario for workers to work close to the structural edge. Although protective facilities such as guardrails and safety ropes are usually set in these areas, in actual operations, actions such as cleaning residual materials, turning around to carry, and standing at the edge often cause personnel to enter the "edge buffer area" or even work close to the edge position, forming a serious risk of falling. Traditional recognition systems mainly rely on physical boundary judgment, such as whether the safety line is crossed or whether the platform edge is exceeded, while ignoring an extremely critical high-risk phenomenon - although the worker has not crossed the boundary, the body center of gravity has approached the boundary, or the posture is dynamically inclined outwards, being in a "pressing posture" at the edge. Such behaviors are often judged as "normal state" by the recognition system because no fall has occurred and no obvious warning line has been triggered, but in fact, they are in extremely high risk.
[0003] In addition, the actions of different work types vary greatly when approaching the edge. For example, the binder needs to bend frequently, and the cleaner may stand sideways, which further increases the recognition difficulty. Currently, image recognition systems generally lack the ability to model the "pressing posture at the edge" and cannot judge whether a person is in a dangerous state of pressing against the boundary based on dimensions such as the person's posture, orientation, action amplitude, and residence time. This technical blind spot directly affects the early recognition ability of unsafe behaviors in high-altitude construction areas, bridge edges, cantilever structures, etc. There is an urgent need to develop an "spatial compressibility" judgment mechanism based on image features to achieve early recognition and risk warning of edge-approaching unsafe behaviors during construction worker operations.
[0004] To solve the above problems, a technical solution is provided as follows. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a comprehensive method for identifying unsafe behaviors of construction site personnel to solve the problems raised in the above background art.
[0006] To achieve the above object, the present invention provides the following technical solutions: S1: Perform edge detection on the monitoring images of the construction site and extract the edge pixel point set; S2: Construct a boundary line segment connection graph through Hough transform and spatial aggregation to delimit the boundary of the candidate edge area; S3: Based on the boundaries of the edge candidate regions, construct a monocular projection mapping structure, establish a buffer zone according to the planar projection trajectories of the boundary line segments, and generate a spatial risk mask area coverage map; S4: Perform instance segmentation and pose estimation on the workers in the monitoring images, extract the coordinate sequence of human joint nodes, and deduce the action center of gravity positions of individual workers; S5: Based on the pre-constructed mapping relationship between operation types and pose structures, screen key points, and output the relative vector sequence of key points with the center of gravity point as the reference origin as the pose space expression; S6: Extract the main pose axis vectors in consecutive frames from the relative vector sequence of key points, calculate the angle between the main pose axis and the boundary normal with the normal vector of the edge contour line as the reference axis, and generate a direction consistency matrix and an outward inclination behavior flag; S7: Identify the resident sections in the spatial risk mask area coverage map and output the labels of oppressive violation approaching behaviors.
[0007] In a preferred embodiment, in S1, performing edge detection on the construction site monitoring images and extracting the edge pixel point set specifically includes: Perform smoothing processing based on a Gaussian kernel on the construction site monitoring images obtained from the construction site and convert them into single-channel luminance maps; Perform Canny edge detection operations using a double-threshold gradient operator on the images, retain the boundary pixel set composed of floor slabs, scaffolding, and guardrail lines, and store all edge pixel point sets in Cartesian coordinate form.
[0008] In a preferred embodiment, in S2, constructing a boundary line segment connection graph through Hough transformation and spatial aggregation and delimiting the boundaries of the edge candidate regions specifically includes: Project the Cartesian edge pixel point set into the parameter space, perform Hough space transformation through an accumulator array, and calculate the angle and displacement parameters of all linear edge distributions; Identify the high-density voting units with continuous responses exceeding the set density threshold, and derive the straight line segment expressions under the corresponding image coordinates; In the set of image straight line segments obtained by projection, perform spatial density aggregation on the geometric center positions and directions of the line segments, construct a line segment connection graph using Euclidean distance constraints and angle similarity limitations, and delimit the subsets with both the average length and overlap degree in the connection graph exceeding the set threshold as the boundaries of the edge candidate regions.
[0009] In a preferred embodiment, in S3, based on the boundaries of the edge candidate regions, construct a monocular projection mapping structure, establish a buffer zone according to the planar projection trajectories of the boundary line segments, and generate a spatial risk mask area coverage map specifically includes: Construct a monocular projection mapping structure between the two-dimensional image coordinates and the physical space of the actual site based on the coordinate data corresponding to the identified boundary line segments of the adjacent edge candidate regions; Use the linear least squares method to regress the planar projection trajectories of the boundary line pixel points in the physical space, map each boundary line to the construction plane, and calibrate the linear segment length and relative azimuth offset at the actual scale; On the basis of the regressed boundary line trajectories, construct a buffer zone extension structure consistent with the boundary normal direction, extend each line segment along the normal direction to the preset buffer width, and form a continuous parallel boundary space strip; Remap the buffer zone area to the image space according to the pixel coordinates, generate a continuous mask structure and use the polygon scanning method to fill and generate a bitmap-level risk area mask layer, and output the spatial risk area coverage map of the corresponding frame.
[0010] In a preferred embodiment, in S4, performing instance segmentation and pose estimation on the workers in the monitoring image, extracting the coordinate sequence of the human joint nodes and calculating the action center of gravity position of the individual worker specifically includes: Obtain the real-time monitoring image of the construction site, calibrate and perform instance segmentation on the bounding boxes of the workers in the image, and output the detection box coordinates and image area masks corresponding to each target individual; Perform pose estimation network inference on the image area within the detection box, extract the coordinate sequence of the human joint nodes of the worker, and represent all coordinates in two-dimensional image pixels and output the node confidence list; In the preset center of gravity node sequence, perform the calculation of the action center of gravity position of the individual worker based on coordinate averaging and weight filtering.
[0011] In a preferred embodiment, in S5, screening key points based on the pre-constructed mapping relationship between the operation type and the pose structure, and outputting the relative vector sequence of the key points with the center of gravity point as the reference origin as the pose space expression specifically includes: Based on the personnel position coordinates in the real-time image frame corresponding to the construction area, retrieve the task type identifier of the operation area corresponding to the current frame through the space mapping table; Call the pre-constructed mapping relationship between the operation type and the pose structure, retrieve the set of action skeleton nodes of the target work type according to the operation type, and output the node combination template for pose vector construction; Screen the key point indexes consistent with the node combination template in the joint node coordinate sequence, and remove the auxiliary nodes irrelevant to the current work type; Based on the screened key point coordinates, construct a relative coordinate structure of the key points with the center of gravity point as the reference origin, and output the relative vector sequence as the pose space expression.
[0012] In a preferred embodiment, in S6, the main attitude axial vectors in consecutive frames are extracted from the relative vector sequence of key points, and the angle between the main attitude axis and the boundary normal is calculated with the normal vector of the adjacent edge contour line as the reference axis, and the direction consistency matrix and the outward inclination behavior flag are generated, which specifically include: Perform time window segmentation on the relative vector sequence of key points, extract the main attitude axial vectors formed by the key points in consecutive frames, and calculate the unit direction vector of each frame; With the normal vector of the adjacent edge contour line as the reference axis, calculate the angle between the main attitude direction of each frame and the normal vector, record the angle offset sequence and construct the offset rate time graph; Introduce the inter-frame offset variance to calculate the attitude stability index of the current action phase, and mark the time period with variance greater than the specified threshold as the attitude transition section; Combine the inter-frame movement direction of the action center of gravity point and the main attitude direction consistency relationship to judge whether the movement trend of the person is towards the risk area, and establish the direction consistency boolean identification matrix; In the period when the continuous angle offset exceeds the set angle and the direction consistency boolean matrix is true, record the continuous frame number and mark the section exceeding the set time threshold with the outward inclination behavior flag.
[0013] In a preferred embodiment, in S7, the residence section is identified in the spatial risk mask area coverage map, and the oppressive violation approaching behavior label is output, which specifically includes: Based on the person segmentation area in the consecutive image frames marked with the outward inclination behavior flag, establish a single-frame pixel-level mask, count the total number of covered pixels within the buffer mask range, and generate the in-frame overlap area ratio sequence; Perform a sliding window average operation on the in-frame overlap ratio sequence, construct a stable residence section index table, and screen the time segments with continuous coverage ratios higher than the set threshold; Calculate the continuous frame number for each high-coverage time segment, discard the short residence segments with a duration not reaching the minimum time threshold, and retain the set of stable oppressive behavior sections; Perform risk score calculation on the oppressive behavior section, construct a level function based on the average overlap ratio, residence duration, and orientation offset amplitude, and output the oppressive violation approaching behavior label.
[0014] The technical effects and advantages of the comprehensive identification method for unsafe behaviors of construction site personnel of the present invention: The boundary of the edge area in the construction site monitoring image is extracted through edge detection and Hough transformation, and a buffer zone and a spatial risk mask area are established by using monocular projection mapping, realizing the precise delineation of the risk area in the construction site. The human joint node sequence of the operator is extracted based on instance segmentation and pose estimation, and the skeleton node set is screened through the pre-constructed mapping relationship between the operation type and the pose structure. An action skeleton structure vector diagram is constructed and the individual center of gravity position is calculated, improving the accuracy of pose expression for different operation types. The calculation of the angle between the main pose axial vector and the boundary normal and the generation mechanism of the direction consistency matrix are introduced to realize the extraction of the out-of-inclination behavior flag and the judgment of the direction trend. Combining the identification of the residence section and the risk score in the spatial risk mask area can accurately output the label of the oppressive violation approaching behavior.
[0015] The method of the present invention can significantly improve the recognition accuracy of the edge approach, inclination action and oppression behavior in the construction site, and has high real-time performance, robustness and applicability. Brief Description of the Drawings
[0016] Figure 1 It is a schematic diagram of a comprehensive method for identifying unsafe behaviors of construction site personnel according to the present invention. Detailed Embodiment
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.
[0018] Embodiment 1 Figure 1 A comprehensive method for identifying unsafe behaviors of construction site personnel according to the present invention is given, which includes the following steps: S1: Perform edge detection on the construction site monitoring image and extract the edge pixel point set; S2: Construct a boundary line segment connection diagram through Hough transformation and spatial aggregation to delimit the boundary of the edge candidate area; S3: Based on the boundary of the edge candidate area, construct a monocular projection mapping structure, establish a buffer zone according to the plane projection trajectory of the boundary line segment, and generate a spatial risk mask area coverage map; S4: Perform instance segmentation and pose estimation on the operator in the monitoring image, extract the human joint node coordinate sequence and calculate the action center of gravity position of the individual operator; S5: Screen key points based on the pre-constructed mapping relationship between the operation type and the pose structure, and output the relative vector sequence of key points with the center of gravity point as the reference origin as the pose space expression; S6: Extract the main attitude axial vectors in consecutive frames from the relative vector sequence of key points, calculate the angle between the main attitude axis and the boundary normal vector with the normal vector of the adjacent edge contour line as the reference axis, and generate the direction consistency matrix and the outward inclination behavior flag. S7: Identify the residence section in the spatial risk mask area coverage map and output the oppressive violation proximity behavior label.
[0019] In S1, perform edge detection on the construction site monitoring image and extract the edge pixel point set.
[0020] The monitoring images of the construction site are collected by fixedly installed camera devices, which are usually set at the edge of the construction floor or erected at designated safety monitoring points to provide effective monitoring of a vast area. The installation height of the camera device is usually between 2.5 meters and 5 meters, and the tilt angle is between 30° and 60°. The collected monitoring images are usually in RGB format, with a resolution of 1920×1080 pixels and a frame rate between 20 and 30 frames. The obtained image data is input into image processing in the form of a three-dimensional matrix.
[0021] In the image preprocessing stage, first perform grayscale processing on the input RGB image to reduce the interference of color information and improve the consistency of edge feature extraction. Use the standard weighted average method to perform weighted combination on the R, G, and B channels to obtain a single-channel luminance image. Then apply Gaussian filtering to the grayscale image to suppress high-frequency noise caused by changes in light, dust interference, or complex backgrounds at the construction site. The size of the filter kernel is generally selected as 5×5, and the standard deviation is set to 1.4 to achieve a reasonable balance between smoothing and edge retention. The filtered output image is used as the input data for subsequent edge detection.
[0022] In the edge detection stage, this embodiment uses the Canny edge detection algorithm to extract the target boundaries such as the floor slab edge, scaffolding structure, and guardrail line from the smoothed image. The algorithm includes four main steps: gradient calculation, non-maximum suppression, double-threshold detection, and hysteresis connection. First, calculate the gradient components in the horizontal and vertical directions of the grayscale image through the Sobel operator, and comprehensively calculate the gradient magnitude and gradient direction angle of each pixel point based on the gradient components in these two directions. Then, perform non-maximum suppression on each pixel point according to the gradient direction information to remove non-edge points to ensure that the output edge lines have high precision and clarity.
[0023] To further distinguish real edges from noise, this embodiment adopts a dual-threshold strategy in edge detection, grading the detection results by setting a high threshold and a low threshold. Pixel points above the high threshold are directly identified as "strong edge points", while pixel points between the high threshold and the low threshold are identified as "weak edge points". In actual operation, the low threshold is generally taken as 0.5 to 0.7 times the high threshold to balance the effects of edge retention and noise suppression. For weak edge points, this embodiment introduces a hysteresis connection strategy to avoid false detection and breaks. Specifically, weak edge points connected to strong edge points are detected, retained, and marked as valid edge points. Otherwise, they are eliminated. This process is particularly important in complex construction environments because subtle features of structures such as scaffolding and guardrails may be overlooked in the standard detection process, and hysteresis connection can improve the continuity and integrity of the edges. All edge pixel points retained by the Canny algorithm are stored in Cartesian coordinate form.
[0024] In S2, a boundary segment connection graph is constructed through Hough transform and spatial aggregation to delimit the boundary of the candidate adjacent edge region.
[0025] After edge detection is completed and an edge pixel point set stored in Cartesian coordinate form is obtained, these pixel points are projected into the parameter space through Hough transform to identify boundary structures that conform to the straight-line feature. Specifically, each edge point is converted into a curve in the parameter space, and by accumulating and counting the intersection conditions of the curves in the parameter space, the boundary segments with linear features in the image are determined. To improve accuracy and reliability, an accumulator array is set in the parameter space, and all units with a voting count exceeding the set threshold are identified. In this embodiment, this threshold is usually set to more than 50 votes (specifically set according to the array size) to ensure that the extracted segments have sufficient validity and significance. For all segments that meet the conditions, they are mapped back from the parameter space to the image coordinate system, and a segment expression with a starting point and an ending point is generated. All identified segments are stored in structured data form, including the starting point, ending point, geometric center, and direction vector of each segment. The direction vector is used to characterize the orientation of the segment to facilitate subsequent aggregation and connection operations. The following is a specific example of the segment structured data: Segment A (floor edge): starting point coordinates (150, 600), ending point coordinates (700, 600), geometric center (425, 600), direction vector (1, 0) (indicating a horizontal line, along the positive direction of the horizontal axis), segment length 550 pixels; Segment B (scaffold crossbar): starting point coordinates: (200, 700), ending point coordinates: (600, 800), geometric center: (400, 750), direction vector: (0.89, 0.45) (indicating a slope to the upper right), segment length: 447 pixels; Line segment C (guardrail line): Starting coordinates: (100, 500), ending coordinates: (100, 300), geometric center: (100, 400), direction vector: (0, -1) (indicating vertically downward), line segment length: 200 pixels.
[0026] To construct a complete boundary of the adjacent edge region, spatial density aggregation and direction similarity screening are performed on the identified set of line segments. First, based on the geometric center positions of each line segment, the relative distances between the line segments are calculated. If the geometric center distance (Euclidean distance) between two line segments is less than the set distance threshold, they are marked as having an adjacency relationship. In this embodiment, the distance threshold is usually set to 20 pixels to ensure good spatial continuity between adjacent line segments. Second, direction similarity analysis is performed on the line segments with an adjacency relationship. Direction similarity is judged by calculating the included angle between two line segments. If the included angle between the direction vectors of two line segments is less than 10 degrees, they are considered to have the same direction. Under the conditions that both adjacency and direction consistency are satisfied, the two line segments are included in the same connection subset to establish a complete line segment connection graph.
[0027] After constructing the connection graph, further screening is performed on all line segment subsets to extract valid boundary structures that conform to the characteristics of the adjacent edge region. The specific judgment criteria are that the average line segment length exceeds 100 pixels, and the overlap ratio between line segments is greater than 80%. If the total length and overlap degree of the line segments in a line segment subset both meet the above requirements, then the subset is marked as a candidate boundary of the adjacent edge region. The identification of such line segment subsets can effectively avoid misjudgments of fine structures or noise points and ensure that the extracted boundary has high continuity and integrity.
[0028] In S3, a monocular projection mapping structure is constructed based on the candidate boundary of the adjacent edge region, a buffer zone is established according to the plane projection trajectory of the boundary line segments, and a spatial risk mask region coverage map is generated.
[0029] A projection mapping structure from two-dimensional image coordinates to construction plane coordinates is constructed based on the installation parameters of the camera device. Since the monitoring camera device is usually a fixed monocular camera, its position, angle, and focal length are all known parameters. Based on the perspective projection principle, by establishing a projection matrix, the conversion from image pixel points to physical space points is achieved. In this process, it is assumed that the construction plane is a flat horizontal plane, and the point directly below the camera device is used as the plane origin. Through this mapping relationship, the boundary line segments in the image can be accurately converted into spatial line segments on the construction plane, laying a foundation for the subsequent construction of the buffer zone. After completing the projection mapping, each boundary line segment mapped to the construction plane is accurately calibrated and described. Specifically, the point set of the boundary line segment is fitted by the linear regression method to obtain the best straight line expression that conforms to the physical structure of the actual site. Each boundary line segment is represented by its starting point, ending point, and direction vector on the plane, and the corresponding line segment length and direction are calibrated.
[0030] After obtaining the spatial representations of all boundary line segments, a method for constructing the buffer zone is further introduced. The generation of the buffer zone is extended based on the normal direction of the boundary line segment. By constructing parallel strip-shaped regions on both sides of the boundary line segment, the coverage of potential dangerous areas is achieved. In specific operations, the normal direction of each boundary line segment is calculated based on its direction vector and extended appropriately in the normal direction. For example, for the outer edge line segments of the construction floor slab, 0.5 meters to 1 meter is extended on both the inner and outer sides of each line segment, thus forming a strip-shaped region with a certain width. The width of this strip-shaped region can be adjusted according to the safety standards of the construction site or specific application scenarios to meet different safety monitoring requirements.
[0031] To ensure the continuity and integrity of the buffer zone area, splicing and fusion processing are performed on all generated strip-shaped regions. By smoothing and compensating the edges of adjacent regions, multiple discrete buffer zone segments can be connected into a complete continuous space. Especially in areas with intersections or overlaps, automatic merging is performed, and the edges are smoothed to eliminate breaks and incoherences caused by differences in the directions of different line segments. After this process, a complete and coherent strip-shaped body of the risk area can be generated.
[0032] After completing the construction and fusion of the buffer zone, the generated risk area is mapped back to the image space to facilitate the generation and visual display of the risk area mask. By filling and rendering the polygon of the area mapped back to the image space, a risk area mask layer represented in bitmap form is generated. The display method of the mask layer is as follows: Form: A binary layer in bitmap form, with the same size as the original image; Mask region value: The pixel value of the risk area is 255 (white), and the pixel value of the safe area is 0 (black). Overlay display: By overlaying the mask layer and the original image in a semi-transparent form, an image with highlighted markings is generated.
[0033] In this way, the high-risk areas at the construction site can be intuitively displayed, providing an accurate spatial reference for subsequent personnel behavior recognition and violation detection.
[0034] In S4, instance segmentation and pose estimation are performed on the workers in the surveillance image, the coordinate sequence of human joint nodes is extracted, and the center of gravity point of the worker's individual movement is calculated.
[0035] In the construction site surveillance image, in order to accurately detect and regionally calibrate multiple workers in a complex scene, a deep learning model is introduced for object detection and instance segmentation. By deploying a convolutional neural network with high precision and efficiency (such as YOLOv8 or Mask R-CNN), real-time inference and processing of the construction site surveillance image are performed.
[0036] In the image, each detected worker is output in the form of a bounding box and an instance mask. The bounding box is used to provide the position and size of the individual in the image, while the instance mask is used to accurately segment the regional contour of each person. Through instance segmentation, workers adjacent or overlapping with each other can be distinguished, thus providing high-quality input data for subsequent pose estimation. The detection results are stored in the form of structured data, including the upper left and lower right coordinates of each detection box, the detection confidence, and the instance mask matrix, and object detection and instance segmentation are continuously performed within the surveillance frame rate range.
[0037] After detecting and segmenting the workers, pose estimation is performed on the image region within each detection box to extract the joint node sequence and characterize the actions and postures of the personnel. A pose estimation model based on a deep neural network (such as OpenPose, HRNet, or PoseNet) is used to extract the corresponding image region from each detection box and perform pose estimation inference. The model can identify multiple key nodes of the human body and output the two-dimensional pixel coordinates and confidence values of each node. Generally speaking, the detected nodes include parts such as the top of the head, neck, shoulders, elbows, wrists, hips, knees, and ankles. The detection result of each node consists of coordinate values and a confidence score, and the confidence is used to measure the accuracy and reliability of the detection. In this embodiment, a standard structure of 17 key nodes (COCO-17 standard) is adopted, and each node is uniformly numbered and semantically calibrated. The pixel coordinates of the nodes are all represented by their positions in the two-dimensional plane of the image, and the detection results of all nodes are output in the form of a data list.
[0038] After extracting the joint nodes of each worker, these nodes are further processed to calculate the center of gravity point of the individual's movement. The calculation of the center of gravity point is the core step for pose analysis and behavior determination. To accurately calculate the center of gravity point, this embodiment introduces a method for calculating the center of gravity based on node weighted average and filtering processing. The scheme presets a set of node sequences related to the calculation of the center of gravity, including the hip, shoulder, and knee positions (which can be increased or decreased according to accuracy requirements). According to the preset weights of different nodes in the human body structure, the coordinates of these nodes are calculated by weighted average to obtain the two-dimensional pixel coordinates of the overall center of gravity point. The calculated center of gravity point information is output in the form of each frame and saved together with the previously extracted joint node sequence. Through the dynamic analysis and trajectory evaluation of the center of gravity point.
[0039] In S5, key points are screened based on the pre-constructed mapping relationship between the operation type and the pose structure, and a sequence of relative vectors of key points with the center of gravity point as the reference origin is output as the pose space expression.
[0040] In the construction site, different areas usually correspond to different types of operation tasks. For example, the high-altitude operation area may be concentrated on the outer edge of the floor slab and the scaffolding area, while the ground operation may involve material handling and equipment installation. To effectively identify the behavior of workers in different scenarios, a matching mechanism of a pre-marked (manually marked) spatial mapping table and the work type task type identifier is introduced. Through the pixel coordinates of the personnel position in the real-time image frame, it is mapped to the physical space coordinates of the construction site. Through the retrieval of the spatial mapping table, the operation area where the current personnel is located is determined, and the corresponding work type task type identifier is extracted from it. For example, when a certain person is detected in the area at the edge of the floor slab, it is marked as "high-altitude operation"; when a person is detected in the material stacking area on the ground, it is marked as "handling operation".
[0041] Different operation types correspond to different movement and pose requirements. To more accurately describe the pose of personnel, through the pre-constructed mapping relationship table, each work type is associated with the most relevant set of action skeleton nodes. The construction of this mapping relationship is based on the collection and analysis of a large number of construction site operation scenarios, and the most representative node combinations are extracted for different task types, such as: High-altitude operations (such as formwork installation, welding, etc.): Pay attention to the stability of the upper limbs and the center of gravity, so the nodes are concentrated on the head, shoulders, elbows, wrists, and hips; Handling operations (such as material handling, equipment transfer, etc.): Focus on the movement and pose changes of the lower limbs and the center of gravity, so the nodes are concentrated on the hips, knees, ankles, and shoulders; Maintenance operations (such as guardrail inspection, equipment maintenance, etc.): Pay attention to the relative position of the hands and the body, so the nodes are concentrated on the wrists, elbows, shoulders, and head.
[0042] After obtaining the node combination template that matches the current task type, screen and reconstruct the previously extracted joint node sequence. Since the joint node sequence usually includes the detection results of multiple joint points (such as 17 points in the COCO-17 standard), but different job types have different requirements for these points. According to the node indexes defined in the node combination template, selectively extract the original joint node sequence. All nodes that meet the template requirements will be retained, while the nodes not in the template will be removed. In this process, not only the joint point information required for the current type of work task is retained, but also the calculation and interference of invalid information are reduced by removing the non-joint points.
[0043] After completing the node screening, perform a unified structured expression on the retained nodes to construct an action skeleton structure vector graph that conforms to the current task type. After constructing the action skeleton structure vector graph, further perform a spatial structured expression on these nodes to construct a relative vector sequence for pose space expression. To ensure the stability and accuracy of pose expression, the strategies of center of gravity point alignment and relative coordinate structure are introduced when constructing the pose space. Specifically, first calculate the position of the center of gravity point within each detection frame, and use the center of gravity point as the reference origin to construct the relative coordinate structure of all key nodes. In this way, the unified expression of the human skeleton structure can be maintained under different poses and actions. All relative vectors are represented in the form of two-dimensional vectors and form a complete node set.
[0044] In S6, extract the main pose axial vector in the relative vector sequence of key points from consecutive frames, calculate the angle between the main pose axis and the boundary normal with the normal vector of the adjacent edge contour line as the reference axis, and generate a direction consistency matrix and an outward inclination behavior flag.
[0045] Based on the continuity of the image frames, segment the pose data of each worker according to time windows. Usually, each time window contains 5 to 10 frames to ensure the sensitivity and stability to action changes. Within each time window, extract the relative vectors of all key points and calculate the main pose axis formed by them in space. The extraction of the main pose axis is achieved by direction aggregation and weighted averaging of the key point sequence. Select the core nodes (such as the head, shoulders, hips, knees, etc.) calibrated in the node combination template, and calculate their average direction vectors within the time window. To ensure the unity and robustness of the direction vectors, each direction vector is normalized to obtain a unit direction vector to represent the main pose direction of the current frame.
[0046] After extracting the main pose direction of each frame, by comparing it with the normal vector of the adjacent edge area, it is determined whether there is an offset of the operator's orientation from the boundary area. Since a complete set of boundary line segments and normal vectors of the adjacent edge candidate areas has been constructed in the previous steps, the pose direction of each person can be compared with the normal vector frame by frame. Taking the normal vector of the adjacent edge contour line as the reference axis, the angle between the main pose direction and the normal vector of each frame is calculated. To ensure the calculation accuracy, the angles are weighted averaged within each time window, and the time series of the angle offset is recorded.
[0047] To determine whether the actions of the operator during construction are highly unstable or suddenly changing, a stability analysis method based on the variance of frame - to - frame offset is introduced. While calculating the angle offset sequence, the variance of the offset data within each time window is calculated to quantify the stability of the actions. The variance of the offset sequence within each time window is calculated and compared with a preset stability threshold. When the variance value exceeds the set threshold (specifically set based on risk identification requirements and the type of construction operation), the current time window is marked as the "pose transition section". This process can effectively identify sudden changes or unstable phenomena in the actions of the operator and monitor and record these high - risk segments in real time.
[0048] On the basis of completing the analysis of pose stability, further combined with the moving direction of the action center of gravity and the main pose direction for consistency judgment to identify whether the moving trend of the operator is towards the risk area. Specifically, the position of the center of gravity of each frame is recorded, and the moving direction vector between adjacent frames is calculated. By comparing it with the main pose direction, the direction consistency relationship between the two is judged. When the moving direction of the center of gravity is consistent with the main pose direction, it is marked as "direction consistent". To improve the accuracy of the judgment, a direction consistency boolean identification matrix is constructed to mark the direction state of each frame. In the direction consistency matrix, if the direction difference between the two is less than the set threshold (such as 15 degrees), it is marked as "true"; otherwise, it is marked as "false". Through this method, it can accurately judge whether there is a continuous moving trend of the operator towards the boundary area.
[0049] After completing the comprehensive analysis of direction consistency and attitude deviation, by judging continuity and time length, an outward inclination behavior flag is generated and a risk assessment is carried out. Specifically, when it is detected that the deviation angle between the main attitude direction of the operator and the normal vector of the edge exceeds a set angle (such as 45 degrees) within a period of time, and the direction consistency matrix is "true", it is considered that the operator has an outward inclination behavior towards the edge area. In this case, the eligible time periods are cumulatively recorded and compared with a preset time threshold. If the continuous outward inclination behavior lasts for more than the set time threshold (such as 3 seconds), an outward inclination behavior flag is generated and relevant warning information is output. This flag can be used as the basis for subsequent risk assessment and detection of violation behaviors, and provides real-time monitoring data for safety management at the construction site.
[0050] In S7, dwelling sections are identified in the spatial risk mask area coverage map, and an oppressive violation approaching behavior label is output.
[0051] After performing person detection and instance segmentation on the real-time monitoring image, a segmentation area mask for each operator is generated. Based on the previous steps, the mask exists in the form of a binary layer, where the pixel value of the target area is 255 (white) and the pixel value of the background area is 0 (black). For each frame of the monitoring image, a person segmentation mask of the same size as the original image is output.
[0052] To identify possible oppressive approaching behaviors, the previously constructed buffer mask is used as the coverage benchmark for the risk area. The buffer mask usually exists in the form of a bitmap, with a pixel value of 255 for the risk area and 0 for the safe area. The shape and size of the mask are determined according to the edge structure of the construction site and the buffer zone construction method. In each frame, the segmentation area mask of the operator is subjected to a pixel-by-pixel intersection operation with the buffer mask. By counting the total number of all pixel points in the intersection area, the total number of covered pixels of the operator's segmentation area in the buffer zone in the current frame can be calculated.
[0053] A sliding window average operation is performed on the coverage ratio sequence to eliminate fluctuations caused by single-frame jitter or noise. The size of the sliding window is dynamically adjusted according to the monitoring frame rate and the movement speed of the operator, usually 5 to 10 frames. Within the sliding window, all coverage ratio values are averaged to obtain a smoother and more stable coverage ratio sequence. After obtaining the smoothed coverage ratio sequence, a stable dwelling section index table is further constructed. For each time segment, if the coverage ratio continuously exceeds a set threshold (specifically set according to risk identification requirements and construction operation types), it is considered that the operator has a stable dwelling behavior in the buffer zone. For example, the coverage ratio threshold is set to 30% or higher to ensure the accuracy of the determination result.
[0054] After the construction of the stable residence section is completed, calculate and filter the duration of each section. Since the operators at the construction site may pass through the buffer zone briefly during movement without constituting a substantial dangerous behavior, irrelevant short-term residence segments are excluded by setting a time threshold. Calculate the number of frames that each high-coverage section lasts and convert it into the actual time length. For example, if the frame rate of the monitoring device is 30 frames per second, a section containing 60 frames means that the person has stayed in the buffer zone for 2 seconds. Compare this time length with the preset minimum time threshold. For example, if the time threshold is set to 1 second, all sections with a residence time less than 1 second will be discarded.
[0055] After identifying the eligible residence sections, calculate the risk scores for these sections. The calculation of the risk scores is based on a comprehensive evaluation of the average coverage ratio, residence duration, and orientation offset amplitude. Specifically, the weighted summation method is used to ensure that the three parameters can accurately represent a risk score value on the same scale.
[0056] Among them, the higher the coverage ratio, the larger the coverage area of the operator in the buffer zone, and the higher the risk score. At the same time, a higher risk score will be given to the behavior of staying for a long time. In addition, combine the angle offset between the main posture direction and the edge normal vector extracted previously to calculate the average value of the offset amplitude. If the offset amplitude is large, it indicates that the operator's posture has a tendency to tilt outwards, and the risk score should be increased. Define the risk level intervals (low, medium, high risks) for the risk score values, output the oppressive violation approaching behavior labels based on the risk levels, and give real-time warning reminders for the behavior risks of construction workers based on the oppressive violation approaching behavior labels.
[0057] The above formulas are all calculated by taking the numerical values without dimensions. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters and threshold selection in the formulas are set by technicians in this field according to the actual situation.
[0058] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0059] Those of ordinary skill in the art can realize that the modules and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0060] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0061] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings, direct couplings, or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be in an electrical, mechanical, or other form.
[0062] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical module, and it may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0063] In addition, in each embodiment of the present application, each functional module can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0064] If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0065] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0066] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A comprehensive identification method for unsafe behaviors of construction site personnel, characterized in that: The steps include: S1: Perform edge detection on the construction site monitoring image and extract edge pixel point sets; S2: Construct a boundary segment connection graph through Hough transform and spatial aggregation to delineate the boundary of the adjacent edge candidate area; S3: Construct a monocular projection mapping structure based on the boundary of the edge candidate area, establish a buffer zone according to the plane projection trajectory of the boundary segment, and generate a spatial risk mask area coverage map; S4: Perform instance segmentation and posture estimation on the operator in the monitoring image, extract the coordinate sequence of the human joint nodes and calculate the center of gravity of the individual action of the operator; S5: Filter key points based on the pre-built mapping relationship between the operation type and the posture structure, and output the relative vector sequence of key points with the center of gravity as the reference origin as the posture space expression; S6: extract the main posture axis vector in the continuous frames from the relative vector sequence of the key points, calculate the angle between the main posture axis and the boundary normal with the normal vector of the edge contour as the reference axis, and generate the direction consistency matrix and the outward tilt behavior mark; S7: Identify the resident segments in the spatial risk mask area coverage map and output the oppressive violation proximity behavior label.
2. A comprehensive identification method for unsafe behaviors of construction site personnel according to claim 1, characterized in that: In S1, edge detection is performed on the construction site monitoring image, and edge pixel point sets are extracted, including: The construction site monitoring images acquired at the construction site are smoothed based on a Gaussian kernel and converted into a single-channel brightness image; The Canny edge detection operation consisting of a double threshold gradient operator is applied to the image, and the boundary pixel set consisting of the floor, scaffolding and guardrail line is retained. All edge pixel point sets are stored in the form of Cartesian coordinates.
3. A comprehensive identification method for unsafe behaviors of construction site personnel according to claim 2, characterized in that: In S2, a boundary segment connection graph is constructed through Hough transform and spatial aggregation, and the boundaries of the adjacent edge candidate regions are delineated, including: Project the Cartesian edge pixel point set into the parameter space, perform Hough space transform through the accumulator array, and calculate the angle and displacement parameters of all linear edge distributions; Identify high-density voting units whose continuous responses exceed a set density threshold, and derive the line segment expression under the corresponding image coordinates; In the set of straight line segments of the projected image, the geometric center positions and directions of the line segments are spatially density aggregated, and a line segment connection graph is constructed using Euclidean distance constraints and angle similarity constraints. The subsets in the connection graph whose average length and overlap exceed the set threshold are defined as the boundaries of the adjacent edge candidate regions.
4. A comprehensive identification method for unsafe behaviors of construction site personnel according to claim 3, characterized in that: In S3, a monocular projection mapping structure is constructed based on the boundary of the adjacent candidate area, a buffer zone is established according to the plane projection trajectory of the boundary line segment, and a spatial risk mask area coverage map is generated, which specifically includes: According to the corresponding coordinate data of the boundary line segments of the identified adjacent candidate regions, a monocular projection mapping structure between the two-dimensional image coordinates and the actual site physical space is constructed; The linear least squares method is used to regress the plane projection trajectory of the boundary line pixel points in the physical space, correspond each boundary line segment to the construction plane, and calibrate the linear segment length and relative orientation offset at the actual scale; Based on the regressed boundary line trajectory, a buffer zone extension structure is constructed in the same direction as the boundary normal line, and each line segment is extended to a preset buffer width along the normal line direction to form a continuous parallel boundary space strip; The buffer zone area is remapped to the image space according to pixel coordinates to generate a continuous mask structure and fill it with a bitmap-level risk area mask layer using polygon scanning method, and the spatial risk area coverage map of the corresponding frame is output.
5. A comprehensive identification method for unsafe behaviors of construction site personnel according to claim 4, characterized in that: In S4, instance segmentation and posture estimation are performed on the operator in the monitoring image, the human joint node coordinate sequence is extracted, and the center of gravity of the individual action of the operator is calculated. Specifically, the following steps are performed: Obtain real-time monitoring images of the construction site, calibrate the bounding boxes of the workers in the image and perform instance segmentation, and output the detection box coordinates and image area masks corresponding to each target individual; Perform posture estimation network reasoning on the image area within the detection frame, extract the operator's body joint node coordinate sequence, all coordinates are represented by two-dimensional image pixels and output a node confidence list; In the preset center of gravity node sequence, the center of gravity position of the individual action of the operator is estimated based on coordinate averaging and weight filtering.
6. A comprehensive identification method for unsafe behaviors of construction site personnel according to claim 5, characterized in that: In S5, key points are selected based on the pre-built mapping relationship between the operation type and the posture structure, and the relative vector sequence of key points with the center of gravity as the reference origin is output as the posture space expression, which specifically includes: Based on the personnel position coordinates in the real-time image frame corresponding to the construction area, the task type identification of the operation area corresponding to the current frame is retrieved through the spatial mapping table; Call the pre-built mapping relationship between job type and posture structure, retrieve the target job action skeleton node set according to the job type, and output the node combination template for posture vector construction; Filter the key point indexes that are consistent with the node combination template in the joint node coordinate sequence, and remove auxiliary nodes that are not related to the current type of work; Based on the filtered key point coordinates, the key point relative coordinate structure is constructed with the center of gravity as the reference origin, and the relative vector sequence is output as the posture space expression.
7. A method for comprehensively identifying unsafe behaviors of construction site personnel according to claim 6, characterized in that: In S6, the main posture axis vector in the continuous frames is extracted from the relative vector sequence of the key points, and the angle between the main posture axis and the boundary normal is calculated with the normal vector of the edge contour as the reference axis, and the direction consistency matrix and the outward behavior mark are generated, which specifically include: Perform time window segmentation on the relative vector sequence of key points, extract the main posture axis vectors formed by the key points in consecutive frames, and calculate the unit direction vector of each frame; Taking the normal vector of the edge contour as the reference axis, calculate the angle between the main posture direction and the normal vector of each frame, record the angle offset sequence and construct the offset rate time graph; The inter-frame offset variance is introduced to calculate the posture stability index of the current action stage, and the period when the variance is greater than the specified threshold is marked as the posture transfer section; Combined with the consistency relationship between the inter-frame moving direction of the action center of gravity and the main posture direction, it is determined whether the movement trend of the personnel is towards the risk area, and a direction consistency Boolean identification matrix is established; During the period when the continuous angle deviation exceeds the set angle and the direction consistency Boolean matrix is true, the number of continuous frames is recorded and the segment exceeding the set time threshold is marked with an outward tilt behavior flag.
8. A comprehensive identification method for unsafe behaviors of construction site personnel according to claim 7, characterized in that: In S7, the resident segment is identified in the spatial risk mask area coverage map, and the oppressive violation approach behavior label is output, specifically including: Based on the segmented areas of people in the continuous image frames marked with extroverted behavior signs, a single-frame pixel-level mask is established, and the total number of covered pixels within the mask range of the buffer is counted to generate a sequence of overlapping area ratios within the frame; Perform sliding window averaging operation on the overlap ratio sequence within the frame, build a stable resident segment index table, and filter the time segments where the continuous coverage ratio is higher than the set threshold; The number of continuous frames is calculated for each high coverage time segment, and the short-term dwelling segments whose duration does not reach the minimum time threshold are discarded, and the set of stable compression behavior segments is retained; A risk score calculation is performed on the oppressive behavior segment, and a grade function is constructed based on the mean overlap ratio, dwell time, and direction deviation amplitude to output the oppressive violation proximity behavior label.
Citation Information
Patent Citations
Power construction personnel behavior identification method based on unmanned aerial vehicle
CN117975310A
Climbing operation risk intelligent identification method and system based on industrial scene
CN118840785A
High-altitude operation lifeline early warning method
CN119068412A
Field monitoring method based on engineering informatization
CN119204668A
Building construction monitoring method
CN119251761A
Cited By
Video monitoring early warning method based on AI analysis
CN121170684A
A video monitoring early warning method based on AI analysis
CN121170684B
Tunnel working face sealing hidden danger identification method and device based on time sequence detection
CN121214344A
A tunnel operation face closure hidden danger identification method and device based on time sequence detection
CN121214344B
Construction site safety risk real-time early warning method and system based on artificial intelligence
CN122454512A