Construction site safety monitoring method based on machine vision

By performing image distortion correction and geometric benchmark modeling under a unified coordinate system at the construction site, the real-time and accuracy problems of construction site safety monitoring in existing technologies have been solved, and high-precision collision early warning has been achieved.

CN121564657APending Publication Date: 2026-02-24NANJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511814489.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing construction site safety monitoring methods cannot identify collision risks between personnel and equipment in a real-time, interpretable and highly accurate manner, especially when vehicles interact with workers, and there are systematic errors and response lag issues.

Method used

By using a machine vision-based method, image distortion correction is performed in a unified coordinate system, and pixel coordinates are mapped to the actual scale on site. Using worker ground contact points, vehicle ground contact points, and excavator key points as geometric references, the working range and danger boundaries are dynamically updated to achieve real-time collision warning.

Benefits of technology

It enables spatial modeling of personnel ground contact points and equipment geometry at a unified scale, reduces systematic errors, and provides real-time, interpretable, and high-precision collision warnings, making it suitable for rapid implementation on construction sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564657A_ABST
    Figure CN121564657A_ABST
Patent Text Reader

Abstract

The invention discloses a construction site safety monitoring method based on machine vision, and belongs to the technical field of safety monitoring, and the method comprises the steps: obtaining a construction site monitoring video, and carrying out the target perception of each frame of image, the targets comprise workers, vehicles and construction machinery at the construction site, and when the target perception of the construction machinery is carried out, the target perception is carried out; capturing positions of a plurality of preset key points of the construction machinery; in the same coordinate system, for each frame of image, each target perceived by the image is subjected to potential danger area demarcation, if the potential danger areas of any two targets have an intersection or inclusion relation, an early warning is given out if a collision risk exists, otherwise, the collision risk does not exist, and continuous monitoring is carried out. According to the method, spatial modeling can be carried out on personnel ground contact points and equipment geometric quantities under a unified scale; in addition, the operation range and the danger boundary are dynamically updated according to the driving state and the operation posture, so that real-time, interpretable and high-precision collision early warning is achieved under the universal monitoring condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of safety monitoring technology, specifically relating to a construction site safety monitoring method based on machine vision. Background Technology

[0002] Construction sites are high-risk environments involving both human and machinery operations. Heavy vehicles and construction machinery frequently interact with workers in confined spaces, and collisions between workers and machinery are a major safety hazard due to factors such as driver blind spots and equipment encroachment on the work area. Current safety management relies heavily on manual inspections, physical barriers, and fixed warning devices, which suffer from limited coverage and delayed response to dynamic working conditions, making continuous risk identification and early warning during operations difficult. Therefore, establishing an effective early warning mechanism for collision risks between construction workers and machinery has clear engineering needs and application value.

[0003] With the widespread adoption of video surveillance and computing resources, computer vision-based automated monitoring is increasingly being applied to on-site security management. Related methods generally fall into two categories: First, area risk modeling and intrusion detection based on monocular video. This involves detecting personnel positions and comparing them with pre-defined danger zones to provide risk warnings. A common practice is to pre-delineate static restricted areas or buffer zones on the camera screen. Second, target detection and tracking are used to obtain the positions of personnel, vehicles, and other targets. The center point of the bounding box is used as a representative point to determine whether someone has entered a restricted area or is too close to a vehicle. However, using the center of the bounding box as a representative point, rather than the actual point of contact with the ground, can lead to discrepancies between the alarm location and the actual threatened location. Dangerous areas are often fixed rings or sectors that do not update with changes in speed, orientation, or turning. Secondly, regarding the identification of personnel behavior or status to reduce false positives and false negatives and improve identification confidence, specifically, the historical trajectories of personnel and equipment are first obtained through detection and tracking, and then the short-term movement trend is predicted using uniform speed or other simple models. If the prediction shows that the two will meet at a certain location, the system calculates the meeting time and compares it with a time threshold determined by distance or speed. However, this type of method is still based on the centroid or bounding box center of the target, and the modeling of equipment geometry and operational envelope is insufficient: vehicles are often simplified to rectangles, and the landing projection and swing edge of the excavator are not considered. The risk of front-to-back asymmetry caused by the vehicle's orientation and speed is not reflected, and the adaptation to reversing and turning sweep is also insufficient. Relying solely on time thresholds is not sensitive enough to close-range stationary occupancy or slow approach. Summary of the Invention

[0004] This invention addresses the shortcomings of existing technologies by providing a machine vision-based construction site safety monitoring method. This method can spatially model personnel ground contact points and equipment geometry, ensuring that personnel and equipment are on a uniform scale. Furthermore, it dynamically updates the work area and danger boundaries based on driving status and working posture, enabling real-time, interpretable, and high-precision collision warnings under general monitoring conditions.

[0005] This invention provides the following technical solution: Firstly, a machine vision-based method for construction site safety monitoring is provided, comprising the following steps: Acquire monitoring video of the construction site and perform target perception on each frame of the image. The targets include workers, vehicles and construction machinery at the construction site. When performing target perception on the construction machinery, capture the positions of several preset key points of the construction machinery. In the same coordinate system, for each frame of image, potential hazard areas are delineated for each perceived target. Before delineating the potential hazard area for workers, the contact point between the worker and the ground is determined by using the midpoint of the bottom edge of the worker's bounding box and applying a preset pixel offset. When delineating the potential hazard area for vehicles, the initial ground contact point is first determined by using the midpoint of the bottom edge of the vehicle's bounding box, and then by applying preset pixel offset correction and oblique driving deviation correction. The potential hazard area is then dynamically determined based on the driving state. When delineating the potential hazard area for construction machinery, one key point is selected as a reference benchmark, and the maximum distance between the remaining key points and the reference benchmark is used as the benchmark radius. After adding a safety margin, the potential hazard area is determined. If any two potential danger zones of any two targets intersect or contain each other in a continuous frame, a collision risk warning is issued; otherwise, there is no collision risk and monitoring continues.

[0006] Optionally, the step of delineating potentially hazardous areas for workers specifically involves: The pixel coordinates of the contact point between the worker and the ground are determined by using the midpoint of the bottom edge of the worker's bounding box and applying a preset pixel offset. : ; in, and These are the x and y coordinates of the top-left corner of the worker's bounding box, respectively. and These are the width and height of the worker's bounding box, respectively. The preset pixel offset is set based on the deviation between the actual and labeled footpoint positions in several pre-selected frame samples. The pixel coordinates of the worker's contact point with the ground After distortion correction, the coordinates are mapped to the BEV coordinate system to obtain the ground coordinates of the worker's contact point with the ground. The center of the potential danger zone is determined by this center; the pixels at both ends of the bottom edge of the worker's bounding box are distorted and mapped to the BEV coordinate system to obtain the ground coordinates at both ends of the bottom edge of the worker's bounding box. and with The potential danger zone for workers is defined by the safety radius of the potential danger zone.

[0007] Optionally, a concentric potential hazard zone for workers is also set outside the designated potential hazard zone for workers, wherein the radius of the potential hazard zone for workers is... for: ,in, This is a preset safety margin.

[0008] Optionally, the step of delineating the potentially hazardous area of ​​the vehicle specifically involves: The initial ground contact point is set by using the midpoint of the bottom edge of the vehicle bounding box as the initial ground contact point and applying a preset pixel offset along the vertical direction of the image to complete the preset pixel offset correction of the initial ground contact point. After distortion correction of the initial ground contact point with preset pixel offset, it is mapped to the BEV coordinate system to obtain the ground coordinates of the vehicle-ground contact point. Simultaneously, the center pixel coordinates of the vehicle bounding box are distorted and mapped to the BEV coordinate system to obtain the detection centroid. ; Ground coordinates of the vehicle's contact point with the ground Given the vehicle's heading angle and nominal dimensions, generate an initial vehicle body rectangle and obtain its geometric centroid. ; The forward and backward direction correction under the centroid constraint is generated according to the following formula. And the amount of correction in the lateral direction : ; in, , , For the vehicle's heading angle, and These are the nominal lengths of the vehicles. and nominal width The limiting coefficient; Correction amount in the forward and backward directions And the amount of correction in the lateral direction To determine the ground coordinates of the vehicle's contact point with the ground Perform oblique driving deviation correction; ; in, These are the ground coordinates of the vehicle's contact point with the ground after correction for oblique driving deviation. The four corner points of the vehicle body are determined based on the corrected vehicle-to-ground contact point and the vehicle heading angle, and the forward buffer zone, rear buffer zone and lateral buffer zone of the potential danger zone of the vehicle are set accordingly. The coordinates of the four corner points of the vehicle body are: ; ; ; ; in, , , and These are the coordinates of the top right, top left, bottom right, and bottom left corners, respectively.

[0009] Optionally, the vehicle's heading angle Determined based on vehicle speed, the vehicle speed for: ; ; When velocity modulus Below the preset threshold At that time, maintain the orientation of the previous moment. ;when At that time, the orientation angle is defined by the direction of velocity. ; in, , , and The first Frames and The ground coordinates of the vehicle's contact point with the ground in the frame image, with an interval of [missing information]. .

[0010] Optionally, the construction machinery includes an excavator, and the preset key points of the excavator include: the rear end of the excavator body, the connection point between the cab and the robotic arm, the intermediate hinge point of the robotic arm, the connection point between the robotic arm and the bucket, the left end point of the bucket, and the right end point of the bucket.

[0011] Optionally, the delineation of potential hazardous areas for construction machinery specifically includes: After distorting the pixel coordinates of all preset key points, transform them to the BEV coordinate system to obtain the ground coordinates of all key points; The connection point between the cockpit and the robotic arm is selected as the reference point, and the maximum distance between the remaining key points and the reference point is used as the reference radius. The radius of the potential hazardous area is determined by adding a safety margin. With the reference point as the center, the radius of the area Delineate the potential danger zone for construction machinery based on the radius; ; in, The configurability factor representing the reference radius. The set radius value of the potential hazard zone. and These are the upper and lower limit radii set respectively.

[0012] Optionally, a peripheral potential hazard zone for the construction machinery is concentrically provided outside the potential hazard zone of the construction machinery, wherein the radius of the peripheral potential hazard zone for the construction machinery is... for: ; in, This is the configurable coefficient representing the radius of the potential hazardous area for construction machinery. The set radius value of the potential danger zone at the outer limit.

[0013] Optionally, if there are peripheral potential danger zones outside the potential danger zone of a target, a tiered warning system is adopted. Specifically, if the peripheral potential danger zones of two targets intersect or contain each other, a low-level warning is issued; if the peripheral potential danger zone of one target intersects or contains the potential danger zone of another target, a medium-level warning is issued; and if the potential danger zones of two targets intersect or contain each other, a high-level warning is issued.

[0014] In a second aspect, a computer-readable storage medium is provided for storing a computer program; when the computer program is executed by a processor, it implements the steps of the machine vision-based construction site safety monitoring method described in any one of the first aspects.

[0015] Compared with the prior art, the beneficial effects of the present invention are: (1) Under a unified coordinate system, this invention maps pixel coordinates to the actual on-site scale through image distortion correction and homography transformation, eliminating distance and area judgment deviations caused by perspective. In addition, this invention uses the worker's ground contact point, vehicle's ground contact point, and excavator's rotation center as geometric references for judgment, avoiding systematic errors caused by approximating the center of the bounding box or the bottom edge center. Among them, the dangerous area of ​​the vehicle adopts a lightweight alignment correction constrained by the centroid of the bounding box, so that the dangerous area under different driving angles is consistent with the vehicle's ground projection, reducing the shape deviation caused by the axis-aligned circumscribed rectangle. The dangerous area of ​​the excavator is determined by the center and radius of the circle based on the ground projection of key points, so that the coverage area matches the actual operation envelope. Under a unified scale, this invention performs spatial modeling of the personnel's ground contact point and equipment geometry, enabling real-time, interpretable, and high-precision collision warning under general monitoring conditions.

[0016] (2) At the judgment level, this invention achieves graded early warning based on the inclusion or intersection relationship between worker hazardous areas and equipment hazardous areas. The rules are clear and the implementation is simple. It can output in real time without the need for three-dimensional reconstruction or multi-sensor fusion. It is implemented by relying on a monocular fixed camera and a general computing platform. The structure is simple, the parameters are configurable, and the deployment and maintenance costs are low, making it suitable for rapid implementation on construction sites. Attached Figure Description

[0017] Figure 1 This is a flowchart of the construction site safety monitoring method based on machine vision of the present invention; Figure 2 This is a schematic diagram illustrating the homography mapping principle and implementation of the camera view to bird's-eye view of the present invention; Figure 3 This is a schematic diagram of the key points of the excavator of this invention; Figure 4 This is a visualization of the effect of the invention when the vehicle is far away from the worker. Figure 5 This is a visualization of the effect of the present invention when a vehicle approaches a worker. Figure 6 This is a visualization of the effect of the present invention when the potential danger area and the trust area overlap in front of the vehicle. Detailed Implementation

[0018] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the scope of protection of the present invention. It should be noted that the term "comprising" and any variations thereof in the specification, claims and the above-mentioned drawings of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or devices.

[0019] Example 1 like Figure 1 As shown, a construction site safety monitoring method based on machine vision includes the following steps: Step 1: Acquire monitoring video of the construction site and perform target perception on each frame of the image. The targets include workers, vehicles and construction machinery at the construction site. When performing target perception on construction machinery, capture the positions of several preset key points of the construction machinery.

[0020] In this embodiment, a monocular camera can be fixedly installed at a designated location; a checkerboard pattern is laid on the ground and calibration images are acquired to obtain the camera's intrinsic parameters, distortion coefficients, and homography transformation matrix from the pixel plane to the BEV coordinate system; based on this, a bird's-eye view coordinate system and canvas range are established on the ground reference plane, and the unit conversion relationship of millimeters / pixels is determined as a unified benchmark for subsequent ground scale determination.

[0021] Target detection is performed on workers, vehicles, and excavators at the construction site, and multi-target temporal correlation is conducted to obtain a temporal trajectory sequence containing target bounding boxes and unique identifiers (IDs). Specific target detection methods can refer to existing technologies.

[0022] More specifically, see Figure 2 In the construction site, fixed monitoring cameras are used to cover the work area. The cameras communicate with the processing unit for video data transmission and online processing; the connection method is not limited. To perform subsequent geometric determinations at a unified physical scale, the images are first distorted. Based on the calibration results, the homography transformation matrix H from the distorted pixels to the ground is obtained, and a ground reference coordinate system (BEV) canvas is established accordingly as a unified benchmark for geometric calculations and display.

[0023] The camera's intrinsic parameters and distortion parameters are calibrated, and distortion correction is performed on the original pixels. The distortion-corrected pixel coordinates satisfy the homography relation with the ground reference plane coordinates. (1) in, This is the homography matrix from pixel to ground. For the distortion-free pixel coordinates; The coordinates are rectangular coordinates of the ground reference plane; This is the scaling factor. Use it if necessary. Ground features are projected back onto the distortion-free image plane for overlay display.

[0024] After completing the pixel-to-ground mapping as described in equation (1), the coverage area of ​​the ground reference coordinate system is selected. , And set the ground resolution ,in The width of the BEV canvas, based on a comprehensive trade-off between pixel and physical scale, is determined by factors such as site coverage, required measurement accuracy, camera installation height, and real-time computing power constraints; no specific limit is imposed. ,high The number of pixels are as follows:

[0025] Ground coordinates Pixel position calculation on the BEV canvas in the ground coordinate system for:

[0026] in, , , , The display boundaries of the ground coordinate system are defined; This represents the resolution from the ground to the pixel.

[0027] After establishing a unified BEV geometric benchmark, this embodiment first completes the dataset construction and model training, and then uses the trained weights to generate temporal outputs for detection and tracking, which serve as inputs for spatial decision-making. The dataset is derived from real-world images taken at construction sites, covering multiple camera angles, near-to-far scale variations, and various occlusion scenarios. The labeled categories include at least three types: workers, vehicles, and excavators, to ensure the model can detect core targets.

[0028] The multi-object detector can be a YOLO series or similar general-purpose detector, and the multi-object tracker can be DeepSORT or a similar multi-object association method based on appearance features and motion models. Specific networks and hyperparameters are not limited. After training, detection accuracy and temporal stability are evaluated to select the optimal model weights. After deploying the selected weights to the processing unit, the system processes the video frame by frame: first, the detector outputs candidate boxes, which are then subjected to non-maximum suppression to remove overlapping redundancy. Then, a multi-object association method assigns a consistent ID to the same entity across adjacent frames, outputting a temporal record conforming to the MOT convention, which includes at least the frame number, object identifier, and bounding box.

[0029] To facilitate subsequent geometric calculations, the bounding boxes are uniformly recorded in the TLWH format, where T and L represent the distortion-free pixel coordinates of the top-left corner of the bounding box. W and H represent the width and height of the bounding box, respectively, both in pixels. If the detection side uses other formats, they are first converted to TLWH. If the inference resolution is inconsistent with the distortion-reduced image resolution, the bounding box position and size are linearly converted to the distortion-reduced scale according to the scaling factors in the horizontal and vertical directions. Only the latest record in time is retained for the same ID within the same frame; all records are organized into continuous trajectories by frame number increment; no interpolation or extrapolation is performed for occasional missed detections to preserve the true occlusion and temporal information. The trajectory data standardized as described above serves as the unified input for subsequent steps 2 and 3.

[0030] Step 2: In the same coordinate system, for each frame of image, delineate the potential danger zone for each perceived target.

[0031] Before delineating potential hazard areas for workers, the contact point between the worker and the ground is determined by using the midpoint of the bottom edge of the worker's bounding box and applying a preset pixel offset. When delineating potential hazard areas for vehicles, the initial ground contact point is first determined by using the midpoint of the bottom edge of the vehicle's bounding box, and then by correcting for preset pixel offsets and oblique driving deviations. The potential hazard area is then dynamically determined based on the vehicle's driving status. When delineating potential hazard areas for construction machinery, one key point is selected as a reference benchmark, and the maximum distance between the remaining key points and the reference benchmark is used as the benchmark radius. A safety margin is then added to determine the potential hazard area.

[0032] In this embodiment, step 2 specifically includes: S2.1: Delineate potentially hazardous areas for workers.

[0033] By leveraging object detection and temporal correlation, the bounding box pixel coordinates of the worker are obtained frame by frame and bound to a unique identity, forming a continuous motion trajectory ordered by time. The trajectory includes at least the frame number, the unique identifier of the target, and the bounding box parameters, which are used to describe the worker's image position in each frame.

[0034] Since the homography mapping from pixel to ground is conditional on the working surface being approximately the same reference plane, the mapping is strictly effective for points on the ground. To ensure positioning accuracy in ground coordinates, this embodiment selects the contact point between the worker and the ground as the geometric representative of the worker's position, and uses this point as a reference to perform safety circle delineation and geometric calculations such as distance and intersection, thereby avoiding deviations introduced by upper body swaying, posture changes, and height differences.

[0035] Step S2.1 specifically includes the following steps: S2.1.1: Determine the pixel coordinates of the contact point between the worker and the ground by using the midpoint of the worker's bounding box bottom edge and applying a preset pixel offset. : (4) in, and These are the x and y coordinates of the top-left corner of the worker's bounding box, respectively. and These are the width and height of the worker's bounding box, respectively. The preset pixel offset is set based on the deviation between the actual and labeled pixel positions in several pre-selected frame samples.

[0036] Introduction The reason is that the bottom edge of the detected bounding box usually systematically deviates from the actual edge of the worker's foot. This is because ground shadows are treated as background by the model, occlusion caused by debris and guardrails, and small distant targets cause the model to shrink at the bottom edge. If the midpoint of the bottom edge of the bounding box is directly used as the foot position, a stable downward deviation will occur after ground mapping, thus affecting the radius of the safety circle and the intersection judgment with the equipment's danger zone. Therefore, the midpoint of the bottom edge of the bounding box is used as the initial value, and then a downward pixel offset is applied. This can compensate for the deviation in one go, making the obtained foot contact point closer to the actual foot landing position. The values ​​were obtained through calibration using a small number of samples. Several representative frames were selected, and the row coordinates of the actual foot position were marked. And read the row coordinates of the midpoint of the bottom edge of the corresponding frame bounding box. The result is calculated and used for subsequent mapping. The expression is:

[0037] Where N is the number of sample frames.

[0038] S2.1.2: Set the pixel coordinates of the point of contact between the worker and the ground. After distortion correction, the coordinates are mapped to the BEV coordinate system to obtain the ground coordinates of the worker's contact point with the ground. The center of the potential danger zone is determined by this center; the pixels at both ends of the bottom edge of the worker's bounding box are distorted and mapped to the BEV coordinate system to obtain the ground coordinates at both ends of the bottom edge of the worker's bounding box. and with The potential danger zone for workers is defined by the safety radius of the potential danger zone.

[0039] In some other embodiments, step S2.1 further includes step S2.1.3: outside the delineated potential danger zone for workers, a peripheral potential danger zone for workers is also set up concentrically.

[0040] Worker outer safety zone radius For: in Add a fixed safety margin to the basics The specific formula is expressed as follows:

[0041] This results in two levels of circular domains: a core safety zone and an outer safety zone, used for high- and medium-level risk assessments, respectively. This definition uses the ground contact point as the sole geometric reference, with all measurements performed in ground coordinates. It allows for consistent distance calculations and intersection checks with ground features such as vehicle body boundaries and operational buffer zones, providing results with clear physical scale and facilitating engineering verification. If slight ground undulations exist, the homography mapping may be locally approximate; errors can be controlled within acceptable limits through on-site parameter verification and necessary fine-tuning.

[0042] S2.2: Delineate potential danger zones for vehicles.

[0043] By utilizing object detection and temporal correlation, the bounding box pixel coordinates of the vehicle are obtained frame by frame and bound to a unique identity, forming a continuous motion trajectory ordered by time. The trajectory includes at least a frame number, a unique target identifier, and bounding box parameters to describe the vehicle's image position in each frame. After obtaining the vehicle's temporal trajectory, the vehicle's ground contact point and direction of motion need to be determined in the ground coordinate system.

[0044] S2.2.1: The midpoint of the bottom edge of the vehicle bounding box is used as the initial ground contact point, and a preset pixel offset is applied along the vertical direction of the image to complete the preset pixel offset correction of the initial ground contact point. (Vehicle ground contact point pixel coordinates) for:

[0045] in, The x-coordinate of the top-left corner of the vehicle's target bounding box; The vertical coordinate of the top left corner of the vehicle target box; The width of the target bounding box; The height of the target bounding box; This is the preset pixel offset.

[0046] S2.2.2: After distortion correction of the initial ground contact point by the preset pixel offset, map it to the BEV coordinate system to obtain the ground coordinates of the vehicle-ground contact point. Simultaneously, the center pixel coordinates of the vehicle bounding box are distorted and mapped to the BEV coordinate system to obtain the detection centroid. .

[0047] Specifically, the center pixel coordinates of the bounding box for:

[0048] Will After distortion removal and substitution into equation (1), the ground centroid observation point is obtained. .

[0049] S2.2.3: Ground coordinates of the vehicle's contact point with the ground Given the vehicle's heading angle and nominal dimensions, generate an initial vehicle body rectangle and obtain its geometric centroid. .

[0050] The vehicle's orientation is given by the displacement of the contact point over time. Let the ground coordinates of the contact point in two adjacent frames be... , The video frame interval is Then the ground velocity vector and the scalar velocity are:

[0051]

[0052] To suppress orientation jumps caused by detection jitter, when the velocity modulus... Below the preset threshold At that time, maintain the orientation of the previous moment. ;when At that time, the orientation angle is defined by the direction of velocity. Regarding the recent The frame orientation is used for moving average to further suppress jitter, as mentioned above. and This is a configurable parameter, and its specific value is not limited.

[0053] When driving at an angle, the bounding box is a rectangle aligned with the image coordinate axes, and its geometric center often does not coincide with the centroid of the actual vehicle's projection on the ground. If directly using... When constructing the vehicle body, systematic deviations in overall position and orientation may occur. Therefore, this embodiment introduces bounded geometric alignment fine-tuning. Specifically, it involves first using ground contact points... Current heading angle and the vehicle's nominal dimensions (vehicle width) Train Length Generate an initial car body rectangle and calculate its geometric centroid. Then compare it with the observed centroid C obtained from the center of the bounding box after distortion correction and equation (1), and use the difference between the two. As a basis for correction, but only within the longitudinal and lateral directions of the vehicle body, and with limited range of motion.

[0054] S2.2.4: Generate the forward and backward direction correction under the centroid constraint according to the following formula. And the amount of correction in the lateral direction :

[0055] in, , , For the vehicle's heading angle, and These are the nominal lengths of the vehicles. and nominal width The limiting coefficient, This is used to suppress geometric divergence and jitter caused by over-correction. This means restricting real numbers to an interval. .

[0056] S2.2.5: Correction amount in the forward and backward directions And the amount of correction in the lateral direction To determine the ground coordinates of the vehicle's contact point with the ground Perform oblique driving deviation correction.

[0057]

[0058] in, The ground coordinates of the vehicle's contact point with the ground after correction for oblique driving deviation; after correction by formula (13), the coordinates of the ground coordinates of the vehicle's contact point with the ground are obtained. The proposed vehicle body rectangle is consistent with the observed vehicle projection centroid, effectively offsetting the axis alignment boundary box deviation caused by oblique driving, thus providing a more reliable benchmark for the subsequent construction of forward, lateral and rearward buffer zones.

[0059] S2.2.6: Determine the four corner points of the vehicle body based on the corrected vehicle-to-ground contact point and vehicle heading angle, and set the forward buffer zone, rearward buffer zone and lateral buffer zone for potential vehicle hazards accordingly; The coordinates of the four corner points of the vehicle body are:

[0060]

[0061]

[0062]

[0063] in, , , and These are the coordinates of the top right, top left, bottom right, and bottom left corners, respectively. The vehicle body polygon is obtained based on these four corner points, effectively offsetting system offsets caused by axis-aligned bounding boxes under conditions such as diagonal driving.

[0064] Forward buffer, rearward buffer, and left and right lateral buffers. The dimensions of the buffers are characterized by simple proportional parameters: the forward length is 0.45 times the vehicle length, the rearward length is 0.20 times the vehicle length, and the left and right lateral widths are respectively taken as a percentage of the vehicle width. The above proportionality coefficient is a configurable constant, which can be given according to field specifications and calibration results, and its specific value is not limited.

[0065] It is worth noting that the forward buffer extends forward along the vehicle's heading, with the proximal end aligned with the front edge of the vehicle. The distal end forms a slightly outward-flaring leading edge to cover the effects of vehicle braking distance, driver visibility obstruction, and frontal sweep. Whether this leading edge flares outward and the angle of flare are determined by a single angular parameter. control: Indicates that it is the same width as the vehicle body; This indicates that the width gradually increases with distance. The rearward buffer extends from the rear of the vehicle body in the opposite direction, and is relatively short, primarily covering the risks of reversing and inertial backlash. The left and right lateral buffers extend symmetrically outward along the vehicle's transverse direction, parallel to the long sides of the vehicle body, and are used to cover lateral risk zones caused by lane changes, steering, and wheel sweep. All three types of buffers rotate in real time with the vehicle's heading and maintain a rigid connection with the vehicle body.

[0066] The vehicle's hazardous area is formed by the union of the vehicle body, forward buffer, rearward buffer, and left and right lateral buffers. To avoid system offset caused by axis-aligned bounding boxes when traveling diagonally, this embodiment calibrates the vehicle body with the observed centroid by light alignment of contact points and heading before generating the buffers. The buffer zone is automatically aligned accordingly, without introducing additional complex calculations.

[0067] S2.3: Delineate potential hazard areas for construction machinery.

[0068] This embodiment primarily considers excavators as the construction machinery. During operation, the area of ​​the excavator's end effector (bucket opening and both ends) that can reach the ground continuously changes with the swing angle and the posture of the boom, stick, and bucket. Therefore, the instantaneous working boundary should be the outer edge of all positions reachable by the end effector on the ground reference plane at that moment, rather than the projection of the vehicle's shape onto the ground. If a fixed outward expansion or circumscribed rectangle is used, posture changes such as lifting and lateral swaying will be ignored, easily leading to the omission or overestimation of the actual reachable area of ​​the end effector, resulting in underestimation or overestimation.

[0069] To ensure that the area delineation is consistent with the actual operation, this invention uses the center of rotation. Based on key structural points such as the working end, the image positions of these key points are corrected for distortion and homography and mapped to ground coordinates. The coordinates of these key points are then calculated in real time. The horizontal distance is used, and the maximum value and its direction are used to determine the current working boundary; this boundary is updated synchronously with attitude changes. A total of 6 key points are used, listed in Table 1, and their locations on the equipment are shown in [Table 1]. Figure 3 .

[0070] Table 1 Key Points of Excavator Operation

[0071] The potential hazard areas for construction machinery are delineated, specifically as follows: S2.3.1: After distorting the pixel coordinates of all preset key points, transform them to the BEV coordinate system to obtain the ground coordinates of all key points.

[0072] Let the homogeneous pixel coordinates of any keypoint in the distortion-free image be... Based on the homography matrix H from the pixel to the ground obtained through calibration, the ground projection of this key point is... Calculate using the following formula:

[0073]

[0074] in, Here are the homogeneous pixel coordinates after distortion correction; H is the homography matrix from the pixel to the ground. The homogeneous coordinates are for the ground reference plane; These are the two-dimensional coordinates of the plane; For homogeneous scale components.

[0075] The aforementioned projection is equivalent to the intersection of the line of sight from the camera's optical center through the keypoint and the ground plane, representing the instantaneous horizontal coverage position of the component on the ground in its current posture. After completing the ground coordinates of the keypoint and reference point, all distance calculations, inclusion and intersection determinations are performed in this ground coordinate system to establish a unified physical scale and eliminate perspective measurement bias, without the need to recover height information.

[0076] S2.3.2: Select the connection point between the cockpit and the robotic arm as the reference reference point, and use the maximum distance between the remaining key points and the reference reference point as the reference radius. The radius of the potential hazardous area is determined by adding a safety margin. With the reference point as the center, the radius of the area Define the potential danger zone for construction machinery by radius.

[0077]

[0078]

[0079] in, The configurability factor representing the reference radius. The set radius value of the potential hazard zone. and These are the upper and lower limit radii set respectively. This is the set of ground coordinates for key points.

[0080] In some other embodiments, step S2.3 further includes step S2.3.3: a potential hazard zone surrounding the construction machinery is also set concentrically outside the potential hazard zone of the construction machinery.

[0081] radius of the potential danger zone around construction machinery for:

[0082] in, This is the configurable coefficient representing the radius of the potential hazardous area for construction machinery. The set radius value of the potential danger zone at the outer limit.

[0083] In this embodiment, the ground projection of the key points and the working reference radius It will continuously change with the rotation angle and operating posture. To avoid false triggers caused by detection jitter, partial occlusion, or short-term missing points, this invention performs robust processing in the time dimension: First, low-pass or short-window median smoothing is used for the ground coordinates of key points and the center C to reduce single-frame noise; second, obviously abnormal jump points are identified and removed, and short-term missing points are interpolated; third, ... By setting constraints on the rate of change and upper and lower limits, sudden increases or decreases in a single frame are restricted to prevent jitter at the region boundaries. Finally, a continuous frame consistency criterion is adopted to set a minimum number of continuous frames for warning triggering and deactivation, in order to merge instantaneous fluctuations. Through the above timing robustness and fault tolerance strategies, the dangerous region maintains stable and verifiable output characteristics even under conditions such as rapid attitude changes, partial occlusion, and illumination disturbances.

[0084] In some other embodiments, two types of views are generated simultaneously while maintaining semantic consistency: one is the BEV whiteboard view, which draws the center C, the potential hazard area, and the surrounding potential hazard areas on a unified ground coordinate system, and optionally displays the direction line projected from C to the farthest key point; the other is a distortion-corrected overlay view, which overlays the area and labels consistent with the BEV onto the corrected image for easy on-site viewing. The shape, radius, and classification labels of the hazard areas in the two views correspond one-to-one, ensuring consistency between review and acceptance. The system also records and outputs metadata related to each calculation, including timestamps, device and frame identifiers, center position, etc. , , Including parameters such as outer edge pointing, which can be used for evidence collection, retrospective analysis and parameter calibration; the above display elements and record fields are all provided in a configurable manner, and can be turned on, off or formatted according to the equipment model and on-site specifications.

[0085] The system synchronously generates a BEV view and a distortion-free overlay view, both of which maintain consistency in the shape, radius, and classification identifier of the danger zone; when a frame is output, it records the timestamp, device / frame identifier, and center point. , , , Metadata such as outer edge pointers are used for verification and backtracking. The above display elements and record fields can be enabled or disabled according to on-site specifications.

[0086] Step 3: If any two potential danger zones of targets intersect or contain each other in consecutive frames, a collision risk warning is issued; otherwise, there is no collision risk and monitoring continues.

[0087] Collision risk warning methods can refer to existing technologies. The number of continuous frames can be set based on experience.

[0088] In some other embodiments, if there is an outer potential danger zone outside the potential danger zone of the target, a graded warning method is adopted. Specifically, when the outer potential danger zones of two targets intersect or contain each other, a low-level warning is issued; if the outer potential danger zone of one target intersects or contains the potential danger zone of another target, a medium-level warning is issued; and if the potential danger zones of two targets intersect or contain each other, a high-level warning is issued.

[0089] If different risk assessment results are issued from multiple mechanical devices within the same frame, a single level will be output according to the principle of "higher is better", and the identifiers of the associated devices will be recorded together.

[0090] To verify the spatial judgment and graded early warning effects of this invention under a unified ground coordinate system, a simplified construction scene was constructed on a flat indoor ground. A camera was fixedly mounted above the work surface to capture a top-down view of the scene. The scene included a scaled-down excavator, a scaled-down dump truck, and four worker markers (labeled A, B, C, and D), and calibration and hazard area modeling were completed according to the aforementioned steps. The system detected worker and equipment targets in the scene, constructed hazard areas, and output three risk levels: SAFE, MEDIUM RISK, and HIGH RISK. The visualization effect is shown below. Figures 4-6 As shown.

[0091] like Figure 4As shown, the truck travels straight along the passage at a low speed (v≈11.4cm / s). The minimum ground distance between worker B and the truck's rearward hazard zone is greater than the warning threshold, indicating a SAFE (Safety, Failure). Worker D, within the excavator's turning radius, remains within its core hazard zone, maintaining a HIGH RISK, demonstrating that the system can assess the impact of multiple devices on multiple workers in parallel.

[0092] like Figure 5 As shown, the truck continues to move forward and accelerates (v≈33.2cm / s). Worker C's minimum ground distance from the truck's forward danger zone falls into the warning zone, and the status is upgraded to MEDIUM RISK. Worker D remains in HIGH RISK because he / she is still within the excavator's core danger zone.

[0093] like Figure 6 As shown, the truck approaches further (v≈31.3cm / s), and the forward danger zone significantly overlaps with worker C's safety zone, with the minimum ground distance falling below the high-risk threshold. The system upgrades its status to HIGH RISK to simulate a strong early warning scenario in actual engineering. Worker D, remaining within the excavator's core danger zone, remains in HIGH RISK.

[0094] Example 2 The present invention provides a computer-readable storage medium for storing a computer program; when the computer program is executed by a processor, it implements the steps of the above-described machine vision-based construction site safety monitoring method.

[0095] For more detailed information on the above methods, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.

[0096] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. Regarding the storage medium disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0097] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present invention.

[0098] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should be considered within the scope of protection of the present invention.

Claims

1. A method for safety monitoring at construction sites based on machine vision, characterized in that, Includes the following steps: Acquire monitoring video of the construction site and perform target perception on each frame of the image. The targets include workers, vehicles and construction machinery at the construction site. When performing target perception on construction machinery, capture the positions of several preset key points of the construction machinery. In the same coordinate system, for each frame of image, potential hazard areas are delineated for each perceived target. Before delineating the potential hazard area for workers, the contact point between the worker and the ground is determined by using the midpoint of the bottom edge of the worker's bounding box and applying a preset pixel offset. When delineating the potential hazard area for vehicles, the initial ground contact point is first determined by using the midpoint of the bottom edge of the vehicle's bounding box, and then by applying preset pixel offset correction and oblique driving deviation correction. The potential hazard area is then dynamically determined based on the driving state. When delineating the potential hazard area for construction machinery, one key point is selected as a reference benchmark, and the maximum distance between the remaining key points and the reference benchmark is used as the benchmark radius. After adding a safety margin, the potential hazard area is determined. If any two potential danger zones of any two targets intersect or contain each other in a continuous frame, a collision risk warning is issued; otherwise, there is no collision risk and monitoring continues.

2. The construction site safety monitoring method based on machine vision according to claim 1, characterized in that, The delineation of potentially hazardous areas for workers specifically involves: The pixel coordinates of the contact point between the worker and the ground are determined by using the midpoint of the bottom edge of the worker's bounding box and applying a preset pixel offset. : ; in, and These are the x and y coordinates of the top-left corner of the worker's bounding box, respectively. and These are the width and height of the worker's bounding box, respectively. The preset pixel offset is set based on the deviation between the actual and labeled foot positions in a few pre-selected frame samples. The pixel coordinates of the worker's contact point with the ground After distortion correction, the coordinates are mapped to the BEV coordinate system to obtain the ground coordinates of the worker's contact point with the ground. The center of the potential danger zone is determined by this center; the pixels at both ends of the bottom edge of the worker's bounding box are distorted and mapped to the BEV coordinate system to obtain the ground coordinates at both ends of the bottom edge of the worker's bounding box. and with The potential danger zone for workers is defined by the safety radius of the potential danger zone.

3. The construction site safety monitoring method based on machine vision according to claim 2, characterized in that, In addition to the designated potential hazard zone for workers, a concentric outer potential hazard zone for workers is also set up, the radius of which is... for: ,in, This is a preset safety margin.

4. The construction site safety monitoring method based on machine vision according to claim 1, characterized in that, The delineation of potential danger zones for vehicles specifically involves: The initial ground contact point is set by using the midpoint of the bottom edge of the vehicle bounding box as the initial ground contact point and applying a preset pixel offset along the vertical direction of the image to complete the preset pixel offset correction of the initial ground contact point. After distortion correction of the initial ground contact point with preset pixel offset, it is mapped to the BEV coordinate system to obtain the ground coordinates of the vehicle-ground contact point. Simultaneously, the center pixel coordinates of the vehicle bounding box are distorted and mapped to the BEV coordinate system to obtain the detection centroid. ; Ground coordinates of the vehicle's contact point with the ground Given the vehicle's heading angle and nominal dimensions, generate an initial vehicle body rectangle and obtain its geometric centroid. ; The forward and backward direction correction under the centroid constraint is generated according to the following formula. And the amount of correction in the lateral direction : ; in, , , For the vehicle's heading angle, and These are the nominal lengths of the vehicles. and nominal width The limiting coefficient; Correction amount in the forward and backward directions And the amount of correction in the lateral direction To determine the ground coordinates of the vehicle's contact point with the ground Perform oblique driving deviation correction; ; in, These are the ground coordinates of the vehicle's contact point with the ground after correction for oblique driving deviation. The four corner points of the vehicle body are determined based on the corrected vehicle-to-ground contact point and the vehicle heading angle, and the forward buffer zone, rear buffer zone and lateral buffer zone of the potential danger zone of the vehicle are set accordingly. The coordinates of the four corner points of the vehicle body are: ; ; ; ; in, , , and These are the coordinates of the top right, top left, bottom right, and bottom left corners, respectively.

5. The construction site safety monitoring method based on machine vision according to claim 4, characterized in that, The vehicle's heading angle Determined based on vehicle speed, the vehicle speed for: ; ; When velocity modulus Below the preset threshold At that time, maintain the orientation of the previous moment. ;when At that time, the orientation angle is defined by the direction of velocity. ; in, , , and The first Frames and The ground coordinates of the vehicle's contact point with the ground in the frame image, with an interval of [missing information]. .

6. The construction site safety monitoring method based on machine vision according to claim 1, characterized in that, The construction machinery includes an excavator, and the key points of the excavator include: the rear end of the excavator body, the connection point between the cab and the robotic arm, the middle hinge point of the robotic arm, the connection point between the robotic arm and the bucket, the left end point of the bucket, and the right end point of the bucket.

7. The construction site safety monitoring method based on machine vision according to claim 6, characterized in that, The delineation of potential hazardous areas for construction machinery specifically includes: After distorting the pixel coordinates of all preset key points, transform them to the BEV coordinate system to obtain the ground coordinates of all key points; The connection point between the cockpit and the robotic arm is selected as the reference point, and the maximum distance between the remaining key points and the reference point is used as the reference radius. The radius of the potential hazardous area is determined by adding a safety margin. With the reference point as the center, the radius of the area Delineate the potential danger zone for construction machinery based on the radius; ; in, The configurability factor representing the reference radius. The set radius value of the potential hazard zone. and These are the upper and lower limit radii set respectively.

8. The construction site safety monitoring method based on machine vision according to claim 7, characterized in that, A concentric potential hazard zone for the construction machinery is also provided outside the potential hazard zone for the construction machinery. The radius of the potential hazard zone for the construction machinery is... for: ; in, This is the configurable coefficient representing the radius of the potential hazardous area for construction machinery. The set radius value of the potential danger zone at the outer limit.

9. The construction site safety monitoring method based on machine vision according to claim 1, characterized in that, If there are peripheral potential danger zones outside the potential danger zone of a target, a tiered warning system is adopted. Specifically, if the peripheral potential danger zones of two targets intersect or contain each other, a low-level warning is issued; if the peripheral potential danger zone of one target intersects or contains the potential danger zone of another target, a medium-level warning is issued; and if the potential danger zones of two targets intersect or contain each other, a high-level warning is issued.

10. A computer-readable storage medium, characterized in that, Used to store computer programs; when executed by a processor, the computer programs implement the steps of the machine vision-based construction site safety monitoring method according to any one of claims 1-9.