Edge end image acquisition and AI hidden danger identification linkage construction site safety management and control system
Patent Information
- Application Number
- CN202611291187.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-25
- Publication Date
- 2026-09-22
AI Technical Summary
[0007]本发明要解决的技术问题是提供边缘端图像采集与AI隐患识别联动的工地安全管控系统,用于解决吊物遮挡人员后,人员位置观测与吊物实际偏移同时不确定而导致隐患判断失准的问题
[0025]This invention assesses the reasons for personnel disappearance, potential occupancy of hidden personnel, and the future physical hazard range of the suspended load under the same on-site load boundary offset, enabling the risk boundary after personnel are obscured by the suspended load to be updated collaboratively with the actual lifting status. The residual mapping module uses the tower crane's operating volume to transform the actual obstruction boundary of the previous moment into the current rigid prediction boundary, and forms the boundary motion residual by subtracting the corresponding sampling point position on the current rigid prediction boundary from the paired point position on the current actual obstruction boundary; the coherent residual is used to correct the predicted position, and the discrete residual, camera calibration error, and preset safety margin together form the residual envelope. When the corrected obstruction sweeping zone covers the unseen position of the personnel and personnel disappear, the trajectory holding module establishes the trajectory of the undetermined personnel. Thus, the observable event of the suspended load boundary sweeping in is used as the cause gate for personnel loss of observation, reducing indiscriminate tracing for personnel leaving the site, obstruction by fixed facilities, or occasional missed detection. Subsequently, the personnel reachable area is expanded according to the upper limit of the personnel movement speed and the cumulative actual time from the moment the personnel did not see the position, and intersects with the occupancy allowable area expanded by the residual envelope to form the unresolved occupancy area, so that the occupancy of the hidden personnel retains the reachable range of the accumulated time, and is also constrained by the current occupancy status of the suspended object.
Smart Images

Figure CN122799375A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of construction site image or video recognition and tower crane safety control technology, and in particular to a construction site safety management system that links edge image acquisition with AI hazard identification. Background Technology
[0002] When tower cranes on construction sites move precast wall panels, steel components, or material boxes above floors, video surveillance systems typically identify the loads and workers, then determine whether the workers have entered a danger zone based on their positional relationship. Patent CN110733983A utilizes a visual sensing device to collect image information of workers and loads at the work surface, determines the danger zone based on the load size, and outputs a judgment result when workers are in or approaching the danger zone. Patent CN110733981A obtains the distance between the load and the work surface, the load's position, and the workers' positions, and determines whether workers will enter a warning danger zone based on the trend of their relative positions. Such methods can form a lifting risk assessment when workers are continuously visible and their positions and load positions are continuously obtainable.
[0003] Video target tracking offers another type of personnel state-preserving route. Patent CN112560656A employs a joint attention mechanism and end-to-end training to perform multi-target pedestrian tracking, enabling personnel detection results to be correlated across consecutive video frames. Even after a person is briefly occluded, the tracking system can continue position estimation based on their pre-occlusion motion state, or recover their identity using appearance features after they reappear; the person's reachable area can be expanded according to the person's velocity boundaries and elapsed time. These processes can mitigate short-term detection interruptions, but typically rely on the pre-occlusion state remaining suitable for subsequent extrapolation, and personnel trajectory preservation and hazardous areas of suspended objects each employ independent motion assumptions or safety margins.
[0004] Obstructions caused by the suspended object itself during floor hoisting can render the above prerequisites invalid. Taking the observation of the lateral movement of a precast wall panel by a high-position camera as an example, the AI image recognition module initially outputs personnel detection frames and forms personnel trajectories on one side of the wall panel; after the wall panel moves along the front of the camera's line of sight and covers the personnel, subsequent video frames only retain the suspended object mask. The personnel may have already left, or they may still be behind the wall panel holding the material; the current image shows the personnel detection as missing in both states. If the system deletes the personnel trajectory after losing several consecutive frames, personnel still holding the material will no longer participate in the hoisting risk assessment; if the system extends the trajectory for all cases of missing detection, personnel leaving the frame, being obstructed by fixed facilities, or occasional missed detections will also be continuously retained, easily causing frequent triggering of the control area.
[0005] During the same process, the suspended load may deviate from the rigidly predicted position obtained solely from the slewing angle, luffing position, and hook height due to sling swing, attitude changes, flexible deformation, or calibration propagation errors. Simply extending the track continuation time cannot determine whether the disappearance of personnel detection is caused by the suspended load obstructing the path. When using a fixed distance to extend the control zone of the suspended load, an insufficient extension may miss actual offsets, while an excessive extension will cause the personnel-accessible area to frequently overlap with the control zone during stable hoisting. Even if the obstruction track continuation, personnel-accessible area, and fixed extension control zone are simply combined, they are still governed by independent assumptions, making it difficult to guarantee that the personnel-side occupancy boundary and the suspended load-side danger boundary correspond to the same actual hoisting state.
[0006] Therefore, the problem that existing technologies need to solve is: when the observation of personnel position is interrupted due to the obstruction of the suspended object, and there is a field deviation between the actual movement of the suspended object and the rigid prediction, how to maintain the coordination between the hidden personnel's position and the dangerous range of the suspended object, and how to make verifiable hazard judgments, so as to reduce the omissions or false stops caused by independent fixed margins. Summary of the Invention
[0007] The technical problem to be solved by this invention is to provide a construction site safety management system that links edge image acquisition with AI hazard identification, in order to solve the problem of inaccurate hazard judgment caused by the uncertainty of both the personnel position observation and the actual displacement of the hoisted object when the hoisted object obscures the personnel.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] A construction site safety management system that links edge image acquisition with AI hazard identification includes: a synchronous acquisition module, located at the edge and connected to an image acquisition terminal, used to acquire time-stamped construction site images or videos and tower crane operation data, and to obtain the passable area of the work surface;
[0010] The AI image recognition module is used to identify personnel trajectories and suspended object masks from the construction site images or videos;
[0011] The residual mapping module is used to map the suspended object mask onto the working surface to obtain the actual suspended object occlusion area and its actual occlusion boundary. It uses the tower crane operation amount to transform the actual occlusion boundary of the previous moment into the current rigid prediction boundary, and compares the current actual occlusion boundary with the current rigid prediction boundary to obtain the boundary motion residual and residual envelope.
[0012] The trajectory holding module is used to establish the trajectory of the unseen person when the occlusion sweeping band, after being corrected by the boundary motion residual, covers the unseen position of the person and the person detection disappears. The unseen occupancy area is obtained by expanding the upper bound of the person's movement speed, the cumulative actual time from the time corresponding to the unseen position of the person, and the residual envelope.
[0013] The hazard identification module is used to generate a physical sweep envelope of the hoisted object using the tower crane's operating volume, correct the physical sweep envelope of the hoisted object according to the boundary motion residual, and expand the hoisted object's danger zone using the residual envelope. It is also used to output the hazard identification result of the hoisting of hidden personnel when the hoisted object's danger zone overlaps with the unresolved occupancy zone.
[0014] The linkage control module is used to send the identification results of the hidden personnel hoisting hazard to the tower crane control terminal and the image acquisition terminal.
[0015] Optionally, it also includes a model training module; the AI image recognition module includes a shared feature extraction network, a personnel detection head, a suspended object segmentation head, and a personnel re-identification head; the model training module is used to acquire a sequence of construction site images or video frames with personnel detection box labels, suspended object mask labels, and personnel identity labels, input the construction site image or video frame sequence into the shared feature extraction network to obtain shared image features, and input the shared image features into the personnel detection head, the suspended object segmentation head, and the personnel re-identification head respectively to obtain candidate personnel detection boxes, candidate suspended object masks, and candidate personnel identity features; the model training module is also used to calculate the detection loss between the candidate personnel detection boxes and the personnel detection box labels, and calculate the... The segmentation loss between the candidate suspended object mask and the suspended object mask label is calculated using the personnel identity label. The metric learning loss between the candidate personnel identity features is then calculated using the detection loss, the segmentation loss, and the metric learning loss. The parameters of the shared feature extraction network, the personnel detection head, the suspended object segmentation head, and the personnel re-identification head are then updated using the detection loss, the segmentation loss, and the metric learning loss. The AI image recognition module is used to output personnel detection boxes and the suspended object mask using the trained personnel detection head and the suspended object segmentation head. It is also used to output personnel identity features using the trained personnel re-identification head and associate the personnel detection boxes in adjacent frames according to the personnel identity features and the timestamp of the construction site image or video to obtain the personnel trajectory.
[0016] Optionally, the synchronous acquisition module is further used to acquire camera calibration relationships and working surface geometry. The camera calibration relationships include camera intrinsic parameters and the transformation relationship between the camera coordinate system and the tower crane coordinate system. The residual mapping module is used to construct the line of sight from the camera optical center through the boundary pixels of the suspended object mask, calculate the intersection points of each line of sight with the working surface, and connect the intersection points to form the actual occlusion boundary. The residual mapping module is also used to construct the line of sight of personnel from the camera optical center through the midpoint of the bottom edge of the personnel detection frame, and determine the intersection point of the line of sight of personnel with the working surface as the unseen position of the personnel. The residual mapping module is also used to determine the position of the unseen personnel based on the rotation angle, amplitude position, and... The hook height determines the hook rigid displacement. Based on the actual occlusion boundary at the moment before the hook rigid displacement transformation, the current rigid prediction boundary is obtained. The residual mapping module is also used to pair each sampling point of the current rigid prediction boundary with the point with the smallest distance on the current actual occlusion boundary. The displacement vector obtained by subtracting the position of the corresponding sampling point on the current rigid prediction boundary from the position of the paired point on the current actual occlusion boundary is determined as the boundary motion residual. The median value of each component of the boundary motion residual is determined as the coherent residual. The high quantile value of the amplitude of each boundary motion residual relative to the coherent residual, the camera calibration error, and the preset safety margin are combined to form the residual envelope.
[0017] Optionally, the residual mapping module is used to translate the current rigid prediction boundary according to the coherent residual to obtain the corrected prediction boundary; the residual mapping module is also used to obtain the area swept by the corrected prediction boundary from the previous time to the current time, and to expand the swept area using the residual envelope to obtain the occlusion sweep-in band; the trajectory holding module is used to establish the trajectory of the undetermined person when the detection of the corresponding person disappears within the time alignment tolerance after the occlusion sweep-in band covers the unseen position of the person.
[0018] Optionally, the trajectory-keeping module is used to obtain the upper limit of personnel speed specified in the operating procedure, the marked personnel trajectory samples, and the preset speed safety margin. It divides the distance between adjacent personnel positions of the marked personnel trajectory samples by the difference in their corresponding acquisition times to obtain the sample personnel speed. The sum of the high quantile value of the sample personnel speed and the preset speed safety margin is determined as the sample speed boundary. The larger value between the upper limit of personnel speed and the sample speed boundary is determined as the upper limit of personnel movement speed. The trajectory-keeping module is also used to expand along the passable area to obtain a personnel reachable area, using the unseen personnel position as the initial region, according to the product of the upper limit of personnel movement speed and the cumulative actual time from the time corresponding to the unseen personnel position. The trajectory-keeping module is also used to expand the current actual suspended object obstruction area according to the residual envelope to obtain an obstruction allowable area, and determine the intersection of the personnel reachable area and the obstruction allowable area as the unresolved occupancy area.
[0019] Optionally, the synchronous acquisition module is further configured to acquire the current operating target, the size of the suspended object, and the length of the sling from the tower crane control terminal; the hazard identification module is configured to use the current operating target to perform trajectory interpolation on the tower crane's operating volume to obtain the future hook trajectory within a preset time window, project the physical contour of the suspended object onto each predicted position of the future hook trajectory according to the size of the suspended object and the length of the sling, and merge the physical contours of the suspended object corresponding to each predicted position to obtain the physical sweep envelope of the suspended object; the hazard identification module is further configured to extrapolate the predicted coherent residual of each predicted position according to the coherent residual change at continuous time intervals, limit the predicted coherent residual according to a preset swing upper bound, and move each physical contour of the suspended object according to the predicted coherent residual. The physical contours of the suspended object before and after movement are merged and then expanded using the residual envelope. The hazard identification module is also used to project the physical contours of the suspended object corresponding to each predicted position onto the working surface along the gravity direction and expand them according to the preset fall safety distance to obtain the falling projection of the suspended object. The hook position in the future hook trajectory is connected and expanded according to the hook size to obtain the hook running channel. The expanded physical contours of the suspended object, the falling projection of the suspended object, and the hook running channel are merged into the hazard zone of the suspended object. The linkage control module is used to send a deceleration command or a stop command to the tower crane control terminal when outputting the hidden personnel hoisting hazard identification result, and the safety interlock of the tower crane control terminal executes the deceleration command or stop command.
[0020] Optionally, it also includes an exposure prediction module; the exposure prediction module is used to generate multiple rigid prediction boundaries for future acquisition times according to the tower crane operation volume, correct each rigid prediction boundary according to the boundary motion residual, and use the residual envelope to expand to obtain the prediction occlusion area; the exposure prediction module is also used to determine the area where the previous prediction occlusion area moves out of the next prediction occlusion area in two adjacent prediction occlusion areas as an exposure zone, and when the exposure zone overlaps with the unresolved occupancy area, determine the boundary of the exposure zone that first forms the overlap as the priority exposure edge, and maintain the trajectory of the unresolved personnel while the exposure zone and the unresolved occupancy area remain separated.
[0021] Optionally, the linkage control module is used to send a frame-incrementing instruction containing the preferred exposed edge and the corresponding future acquisition time to the image acquisition terminal that forms the trajectory of the unresolved personnel; the image acquisition terminal is used to increase the acquisition frame rate of the image area where the preferred exposed edge is located at the corresponding future acquisition time; the AI image recognition module is used to extract the features of the personnel before the occlusion before the detection disappearance time, and match the features of the personnel re-detected in the image area where the preferred exposed edge is located with the features of the personnel before the occlusion; the trajectory holding module is used to clear the trajectory of the unresolved personnel when the match is successful and the re-detected personnel position is outside the danger zone of the suspended object, and is also used to shrink the unresolved occupancy area according to the personnel position when the match is successful and the re-detected personnel position is inside the danger zone of the suspended object.
[0022] Optionally, a viewpoint relay module is also included. The viewpoint relay module is used to acquire the field of view and camera calibration relationship of multiple candidate image acquisition terminals that are communicatively connected to the synchronous acquisition module. It maps the suspended object mask identified from the construction site images or videos acquired by each candidate image acquisition terminal to the working surface to obtain candidate occlusion areas. It transforms the candidate occlusion areas at the previous moment according to the tower crane operation volume and compares the transformation result with the candidate occlusion areas at the current moment to obtain candidate residual envelopes. The viewpoint relay module is also used to expand each candidate occlusion area using the corresponding candidate residual envelopes, and select the candidate image acquisition terminal with the smallest overlap area between the expanded candidate occlusion area and the unresolved occupancy area from the candidate image acquisition terminals whose field of view covers the unresolved occupancy area as the relay image acquisition terminal.
[0023] Optionally, the linkage control module is used to send a directional acquisition command to the relay image acquisition terminal, which defines the unresolved occupancy area and the acquisition period; the relay image acquisition terminal is used to acquire relay video according to the directional acquisition command; the AI image recognition module is used to select images of people before the detection disappearance time from the trajectory of the unresolved persons and extract the features of the persons before occlusion, and is also used to identify the position of the relay person in the relay video that matches the features of the persons before occlusion; the trajectory keeping module is used to correct or clear the unresolved occupancy area using the position of the relay person; the AI image recognition module is also used to output a pending cloud review status when the feature matching confidence is less than a preset matching threshold; the synchronous acquisition module is used to store the preset period video before the detection disappearance time and the relay video in a local cache when the pending cloud review status is output, and upload them to the cloud for review when communication is available.
[0024] The beneficial effects of this invention are as follows:
[0025] This invention assesses the reasons for personnel disappearance, potential occupancy of hidden personnel, and the future physical hazard range of the suspended load under the same on-site load boundary offset, enabling the risk boundary after personnel are obscured by the suspended load to be updated collaboratively with the actual lifting status. The residual mapping module uses the tower crane's operating volume to transform the actual obstruction boundary of the previous moment into the current rigid prediction boundary, and forms the boundary motion residual by subtracting the corresponding sampling point position on the current rigid prediction boundary from the paired point position on the current actual obstruction boundary; the coherent residual is used to correct the predicted position, and the discrete residual, camera calibration error, and preset safety margin together form the residual envelope. When the corrected obstruction sweeping zone covers the unseen position of the personnel and personnel disappear, the trajectory holding module establishes the trajectory of the undetermined personnel. Thus, the observable event of the suspended load boundary sweeping in is used as the cause gate for personnel loss of observation, reducing indiscriminate tracing for personnel leaving the site, obstruction by fixed facilities, or occasional missed detection. Subsequently, the personnel reachable area is expanded according to the upper limit of the personnel movement speed and the cumulative actual time from the moment the personnel did not see the position, and intersects with the occupancy allowable area expanded by the residual envelope to form the unresolved occupancy area, so that the occupancy of the hidden personnel retains the reachable range of the accumulated time, and is also constrained by the current occupancy status of the suspended object.
[0026] The same boundary motion residual is also transferred to the future risk assessment on the suspended object side, so that the personnel side and the suspended object side no longer rely on each other's independent fixed margins. The hazard identification module forms a physical sweep envelope of the suspended object based on the tower crane's operating volume, corrects the physical sweep envelope of the suspended object according to the boundary motion residual, expands the residual envelope, and then merges the suspended object's fall projection and the hook's running channel to obtain the suspended object's danger zone; when the suspended object's danger zone overlaps with the unresolved occupancy zone, the system outputs the hidden personnel hoisting hazard identification result, and the safety interlock of the tower crane control terminal executes a deceleration command or a stop command. Under the condition that the prediction time window covers the worst end-to-end delay and braking time measured on site, and the calibration error between the camera, tower crane, and work surface is within the on-site verification boundary, the overlap judgment can provide the tower crane linkage control with a risk basis consistent with the on-site boundary offset, without having to definitively attribute the coherent residual to a certain swing, attitude, or calibration error.
[0027] As an auxiliary effect in ensuring input consistency, the shared feature extraction network provides homogeneous image features to the personnel detection head, the suspended object segmentation head, and the personnel re-identification head. The model training module forms three types of losses through personnel detection box labels, suspended object mask labels, and personnel identity labels, and jointly updates the network parameters. The trained AI image recognition module outputs personnel detection boxes, suspended object masks, and personnel identity features, and then combines them with timestamps to associate personnel detection boxes in adjacent frames. This ensures that the personnel trajectory, the actual occlusion boundary of the suspended object, and the personnel features before occlusion all come from a time-aligned video processing chain, providing consistent input for residual mapping and occlusion recovery.
[0028] For evidence collection after occlusion removal, this invention only increases the acquisition frame rate for the image region where the preferentially exposed edge is located and the corresponding future acquisition time when the exposed zone, after boundary motion residual correction, overlaps with the unresolved occupant area; when the two remain separate, the trajectory of the unresolved personnel is maintained. This conditional processing avoids directly equating the continued movement of the suspended object with the exposure of personnel. Compared to continuously increasing the frame rate of the entire video stream, it can increase the density of risk-related evidence collection and reduce the consumption of computing power and cache at the edge by irrelevant images.
[0029] For multi-camera relay, this invention selects the terminal with the smallest overlap area between the candidate occupant area and the unresolved occupant area from candidate image acquisition terminals whose field of view covers the unresolved occupant area, after being expanded by their respective candidate residual envelopes, and performs directional acquisition. When a match is successful, the system uses the relay personnel's position to correct or clear the unresolved occupant area; when the feature matching confidence is less than a preset matching threshold, only the video of the preset time period before the detection disappears and the relay video are cached, and uploaded to the cloud for verification when communication is available. This processing ensures that the relay viewpoint, acquisition time period, and uploaded content are determined around the unresolved occupant evidence, which can reduce the communication burden in weak network environments compared to the routine uploading of multiple video streams to the cloud. Attached Figure Description
[0030] Figure 1 This is a structural diagram illustrating a deployment scenario for the construction site hoisting safety management system provided in an embodiment of the present invention.
[0031] Figure 2 This is a schematic diagram of the module structure of the construction site safety management system provided in an embodiment of the present invention;
[0032] Figure 3 This is a schematic diagram illustrating the mechanism by which a rigid predicted boundary is compared with the actual occlusion boundary to form a boundary motion residual, as provided in an embodiment of the present invention.
[0033] Figure 4 This is a geometrical schematic diagram illustrating the formation process of the obstruction sweeping zone and the unresolved occupancy area provided in an embodiment of the present invention.
[0034] Figure 5 This is a schematic diagram showing the comparison between the physical sweep envelope correction of the suspended object and the formation process of the dangerous zone of the suspended object provided in an embodiment of the present invention;
[0035] Figure 6 This is a data flow diagram illustrating the multi-task model architecture and training process of the AI image recognition module provided in an embodiment of the present invention.
[0036] Figure 7 A schematic diagram showing the comparison between conditional exposure prediction and region augmentation process provided in an embodiment of the present invention;
[0037] Figure 8 This is a schematic diagram of a multi-view relay acquisition and cloud verification process provided in an embodiment of the present invention.
[0038] Figure Label Explanation: 10-Tower crane; 11-Tower crane control terminal; 12-Edge end; 13-Image acquisition terminal; 14-Working face; 15-Lifted object; 16-Personnel; 17-Passable area; 20-Construction site safety management system linking edge end image acquisition and AI hazard identification; 21-Synchronous acquisition module; 22-AI image recognition module; 23-Residual mapping module; 24-Trajectory holding module; 25-Hazard identification module; 26-Linkage control module; 27-Exposure prediction module; 28-Viewpoint relay module; 29-Model training module; 30-Actual lifted object occlusion area; 31-Actual occlusion boundary; 32-Rigid prediction boundary; 33-Boundary motion residual; 34-Residual envelope; 35-Personnel unseen position; 40-Corrected prediction boundary; 41-Occlusion sweep-in zone; 42-Personnel reachable area; 43-Occlusion 44 - Permissible Area; 50 - Undecided Occupied Area; 51 - Future Hook Trajectory; 52 - Nominal Load Physical Sweep Envelope; 53 - Envelope Corrected by Predicted Coherence Residual; 54 - Load Fall Projection; 55 - Hook Running Channel; 66 - Load Danger Zone; 67 - Construction Site Image or Video Frame Sequence; 68 - Shared Feature Extraction Network; 69 - Personnel Detection Head; 60 - Load Segmentation Head; 61 - Personnel Re-identification Head; 62 - Personnel Detection Box Label; 63 - Load Mask Label; 64 - Personnel Identity Label; 75 - Predicted Occlusion Area; 76 - Exposure Zone; 77 - Priority Exposure Edge; 88 - Frame Increment Command; 89 - Multiple Candidate Image Acquisition Terminals; 80 - Candidate Occlusion Area; 81 - Candidate Residual Envelope; 82 - Relay Image Acquisition Terminal; 83 - Directional Acquisition Command; 84 - Relay Video; 85 - Local Cache; 86 - Cloud Verification. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0040] Figure 1This is a structural diagram illustrating a deployment scenario for the construction site hoisting safety management system provided in this embodiment of the invention. The operating mechanism of the tower crane 10 moves the hoisted load 15 relative to the working surface 14, allowing personnel 16 to work within the passable area 17 on the working surface 14. An image acquisition terminal 13 is installed on the tower body, floor-mounted supports, or other locations with a top-down field of view, covering the area the hoisted load 15 may pass through and the activity area of the personnel 16. An edge terminal 12 communicates with both the image acquisition terminal 13 and the tower crane control terminal 11, enabling a closed loop of image data, tower crane operation data, and control results on-site, ensuring continuous availability without relying on cloud links.
[0041] Figure 2 This is a schematic diagram of the module structure of a construction site safety management system provided in an embodiment of the present invention. The construction site safety management system 20, which integrates edge image acquisition and AI hazard identification, includes a synchronous acquisition module 21, an AI image recognition module 22, and a residual mapping module 23. It also includes a trajectory holding module 24, a hazard identification module 25, and a linkage control module 26. In some embodiments, the system further includes an exposure prediction module 27, a viewpoint relay module 28, and a model training module 29. The synchronous acquisition module 21 provides a unified timeline and on-site geometric reference; the AI image recognition module 22 generates visual observations of personnel and suspended objects; and the residual mapping module 23 converts the visual observations to work surface coordinates. The trajectory holding module 24 maintains the status of personnel with evidence of suspended object obstruction; the hazard identification module 25 defines the future physical hazard range; and the linkage control module 26 is responsible for distributing event results to the acquisition and control sides.
[0042] The synchronous acquisition module 21 can receive encoded video frames from the image acquisition terminal 13, or raw images output by the camera. Each frame in the site image or video frame sequence 60 carries an acquisition timestamp obtained by unified timing from the camera hardware clock or the edge terminal 12. The tower crane control terminal 11 provides slewing angle, slewing speed, and luffing position, and can also provide luffing speed, hook height, and hoisting speed, as well as the current operating target. For older controllers without hardware timestamps, the edge terminal 12 can record a monotonic clock when a message arrives and use the sliding median of the round-trip time delay to compensate for communication transmission deviations.
[0043] In one embodiment, the image acquisition frame rate can be 15 to 30 frames per second, and can be increased to 40 frames per second when the load is moving rapidly or when there are many people. The time synchronization deviation can be controlled within 5 to 50 milliseconds. The edge terminal 12 uses the image timestamp as the query time and performs linear interpolation on the operation volume of two adjacent tower cranes. When the query time falls outside the operation volume message, it is only allowed to extrapolate by speed within a short interval of 100 to 300 milliseconds after the most recent message. If it exceeds this interval, the distorted operation volume is not used to generate a rigid prediction boundary 32, and the exception handling described later is entered.
[0044] The synchronous acquisition module 21 also obtains the passable area 17 from the site plan, electronic fence configuration, or on-site annotation results. The passable area 17 can be represented as a set of polygons in the work surface coordinates, and floor slab openings, edge-restricted areas, fixed equipment bases, and material storage areas are deducted from this set. When temporary fencing changes, the on-site management terminal only updates the polygon vertices, without needing to retrain the visual model. This processing ensures that the reachable expansion after personnel lose observation is constrained by the actual passage conditions, avoiding circular expansion that crosses impassable obstacles.
[0045] During system operation, the AI image recognition module 22 first performs personnel detection, suspended object segmentation, and personnel identity feature extraction on images on the same time axis. The residual mapping module 23 then projects the personnel positions and suspended object contours onto the work surface 14 and generates a current rigid prediction boundary 32 based on the tower crane's operating volume. The actual occlusion boundary 31 obtained from the current vision is compared with the current rigid prediction boundary 32 to generate a boundary motion residual 33. This residual is used to determine whether the disappearance of personnel detection is related to the sweeping in of the suspended object, to expand the possible occupancy range of hidden personnel, and to correct the future position of the suspended object. All three stages are driven by the same on-site observation bias, and therefore do not employ independent fixed margins.
[0046] Figure 6 This is a data flow diagram illustrating the multi-task model architecture and training process of the AI image recognition module 22 provided in this embodiment of the invention. The AI multi-task model takes construction site images or video frame sequences 60 as sample inputs and includes a shared feature extraction network 61, a personnel detection head 62, a suspended object segmentation head 63, and a personnel re-identification head 64. The shared feature extraction network 61 can use MobileNetV3, ResNet-50, or a convolutional network with a feature pyramid to output texture and semantic information at different scales as shared image features. The personnel detection head 62 outputs candidate personnel detection boxes and confidence scores, the suspended object segmentation head 63 outputs candidate suspended object masks corresponding to the original image, and the personnel re-identification head 64 outputs normalized personnel identity feature vectors.
[0047] Model training module 29 extracts training samples from anonymized construction site videos. Personnel detection bounding box labels 65 use the bounding rectangle of the visible body of the person as the annotation caliber; for severely occluded but still identifiable persons, the visible bounding box is retained and the occlusion ratio is recorded. Lifting object mask labels 66 are labeled pixel-by-pixel along the visible outer contour of wall panels, steel components, or material boxes; whether slings and hooks are included in the mask remains consistent within the same dataset. Personnel identity labels 67 group the same person according to their identity across adjacent cameras and adjacent time slices; samples whose identity cannot be confirmed only participate in detection and segmentation training, not metric learning.
[0048] The samples are grouped by construction site or construction date, and then further divided into training, validation, and test sets to avoid the same continuous segment being included in both sets simultaneously. Optionally, the training set comprises 70% to 85% of the samples, the validation set 10% to 20%, and the remainder serves as a separate test set. The model input resolution can be selected between 640×384 pixels and 1280×768 pixels, and the training batch can be 8 to 32 frames. Sample augmentation may include brightness perturbation, slight blurring, scale transformation, and horizontal flipping consistent with the camera's imaging direction, but arbitrary rotations that alter the direction of the suspended object's gravity are not permitted.
[0049] The personnel detection head 62 can employ an anchorless detection structure, with its detection loss consisting of classification loss, bounding box regression loss, and intersection-union (IU) loss. The hanging object segmentation head 63 can utilize a lightweight decoder, with its segmentation loss consisting of pixel cross-entropy loss and Dice loss. The personnel re-identification head 64 brings together samples of the same identity and pushes away samples of different identities; its metric learning loss can be obtained by combining triplet loss and identity classification loss. The initial weights of the three types of losses can be set between 0.2 and 2, exemplarily weighted at 1, 1, and 0.5; during training, they can be adjusted according to the gradient magnitude of the validation set to prevent the loss of a particular task from dominating shared features in the long run.
[0050] The AI multi-task model can be trained for 80 to 200 epochs using AdamW or stochastic gradient descent algorithms, with an initial learning rate of 0.00005 to 0.001, and employs cosine annealing or piecewise decay. The model training module 29 backpropagates the three-class weighted loss to the shared feature extraction network 61, while simultaneously updating the personnel detection head 62, the hanging object segmentation head 63, and the personnel re-identification head 64. If, after 5 to 15 consecutive epochs, the validation loss does not decrease and the personnel detection recall, hanging object mask intersection-over-union ratio, and identity retrieval accuracy do not improve, training can be stopped, and the parameters that best validate the overall performance metrics can be retained.
[0051] During the cold start phase when samples are insufficient, the shared feature extraction network 61 can load pre-trained parameters from a publicly available image dataset, and the three task heads are fine-tuned with a small number of field samples. The initial field samples cover at least daytime, nighttime, and backlighting environments, as well as rainy / foggy environments and common types of suspended objects; uncovered conditions are manually reviewed to form incremental annotations. After deployment, offline retraining is performed weekly or based on the cumulative number of newly added samples. The retrained model only replaces the production model if it does not reduce the recall rate for personnel detection, the intersection-over-union ratio for suspended object segmentation, or the accuracy of identity retrieval or cross-camera association on the independent validation set. This update is performed offline, and the edge device 12 does not directly rewrite the model parameters online based on a single false alarm.
[0052] When deployed at the edge, the number of parameters in the AI multi-task model can be controlled between 8 million and 40 million, using 8-bit integer or 16-bit floating-point quantization. The quantized model is first calibrated on a computing platform consistent with the deployment device. If the output differences of personnel detection boxes, suspended object masks, and identity features exceed the acceptance threshold, the original accuracy model is retained. The model version number, input timestamp, and output timestamp of the inference frame are written to the event log, enabling personnel trajectories, suspended object boundaries, and subsequent control results to be traced back to the same model version.
[0053] When associating person detection boxes between adjacent frames, the AI image recognition module 22 first uses motion thresholds to eliminate impossible matches, and then forms an association cost based on the cosine distance of the person's identity features, the box center displacement, and scale changes. The Hungarian algorithm or cascaded matching algorithm can be used to obtain a one-to-one correspondence. Trajectories that do not match temporarily only enter the candidate lost state and are not immediately deleted; only after subsequent residual mapping confirms the inclusion of the object are the trajectory converted to an undetermined person trajectory. This distinguishes between general missed detections and detections that disappear due to physical occlusion evidence.
[0054] The camera, tower crane, and work surface all use a unified right-handed coordinate system. During calibration, the focal length, principal point, and distortion parameters of the image acquisition terminal 13 are first determined using a checkerboard or dot grid. Then, at least six non-collinear control points are measured in the tower crane coordinate system. These control points can be placed at floor corners, tower reference points, and fixed components. The rigid body transformation from the camera coordinate system to the tower crane coordinate system is calculated using the pixel positions and three-dimensional coordinates of the control points in the image. The work surface 14 is represented by local plane equations or a grid surface with elevation.
[0055] The reprojection error is used to verify the calibration relationship. Its root mean square value can be taken as an acceptable range of 1 to 5 pixels, typically not exceeding 2.5 pixels. If the threshold is too small, it will cause frequent recalibration triggered by ordinary lens distortion and manual point selection errors. If the threshold is too large, it will amplify the pixel error into a positional error on the working surface. For the mobile image acquisition terminal 13, the external parameters can be updated by the pan-tilt encoder, visual positioning markers, or synchronous positioning results. After each update, the reprojection error of the independent control points is recalculated. New external parameters that fail verification are not included in the residual calculation.
[0056] In a scenario where work surface 14 consists of multiple height layers, the floor slab can be divided into multiple local planes, or elevations can be represented using triangular meshes. The residual mapping module 23 selects the intersection plane based on the plane partitions where the image rays are expected to fall. Pixels are not used to form boundaries when a ray is approximately parallel to the work surface, the intersection point falls outside the effective area, or the intersection point depth is negative. This avoids incorrectly projecting distant facades or sky areas onto the floor surface.
[0057] The residual mapping module 23 constructs a ray from the camera's optical center through the boundary pixels of the suspended object mask, and connects the effective intersection points in sequence according to the mask boundary to form the actual suspended object occlusion area 30 and its outer actual occlusion boundary 31. The midpoint of the bottom edge of the personnel detection frame approximately corresponds to the contact position between the personnel's feet and the working surface 14, and the intersection of the ray passing through this point and the working surface 14 is taken as the personnel's unseen position 35. When the lower body of the personnel is temporarily occluded, the weighted result of the historical foot point and the current bottom edge of the detection frame can also be used to reduce jumps.
[0058] Figure 3 This is a schematic diagram illustrating the mechanism by which the rigid predicted boundary 32 and the actual occlusion boundary 31 are compared to form the boundary motion residual 33 in an embodiment of the present invention. Figure 3 The actual occlusion boundary 31 is represented by a solid line, and the rigid prediction boundary 32 is represented by a dashed line. The residual mapping module 23 calculates the rigid displacement of the hook based on the rotation angle, amplitude position, and hook height at adjacent time points, and uses this displacement and the corresponding planar rotation to transform the actual occlusion boundary 31 at the previous time point to obtain the current rigid prediction boundary 32. The unseen position 35 is located on the side where the rigid prediction boundary 32 sweeps into the actual occlusion boundary 31; this relative relationship is used for subsequent occlusion cause gate judgment. Using only rigid prediction will ignore the sling swing, the change in the attitude of the suspended object, and the calibration propagation error; therefore, the current visual boundary still needs to be included in the correction.
[0059] The residual mapping module 23 selects sampling points along the current rigid prediction boundary 32 at equal arc lengths or equal angles, and searches for the pairing point with the smallest distance on the current actual occlusion boundary 31. To maintain pairing continuity, the order of adjacent pairing points along the boundary arc length can be constrained, and a maximum pairing distance of 0.2 meters to 1.5 meters can be set. The position of the current actual pairing point minus the position of the current rigid prediction sampling point yields the boundary motion residual 33 pointing from the prediction to the actual boundary. Figure 3 The direction of each arrow indicates the direction of the subtraction; it is not used in reverse.
[0060] The median values of the lateral and longitudinal components of multiple boundary motion residuals 33 are taken to obtain the coherent residuals. The median value is insensitive to local masking burrs and a small number of mismatches, and can represent the overall translational trend of the suspended object boundary relative to the rigid prediction. The correction performed according to the boundary motion residuals 33 mentioned in this invention refers to the overall correction performed according to the coherent residuals obtained statistically from each boundary motion residual 33. Discrete uncertainties that cannot be explained by the overall translation are carried by the residual envelope 34. The amplitude of each residual minus the coherent residual is used to describe the discrete uncertainty. The high quantile of the residual amplitude can be taken from the 85th to the 99th percentile. When the sample points are small, a higher quantile is taken. When the sample points are sufficient and the noise is stable, the 90th to the 95th percentile can be taken.
[0061] The residual envelope 34 is formed by the residual discrete high quantile, the propagation of calibration error on the working surface, and the preset safety margin. The preset safety margin can be 0.05 meters to 0.30 meters, typically 0.15 meters. If this margin is too small, the boundary discreteness and projection error will lack a safety net; if it is too large, the occupancy area and the danger zone will frequently overlap under stable hoisting. Figure 3 The outer dashed line represents the residual envelope 34 in the diagram, used to carry the remaining uncertainty that cannot be explained by the overall translation of the coherent residuals.
[0062] The coherent residual and residual envelope 34 are updated using different methods. The coherent residual can be updated with exponentially smoothed values from the most recent 3 to 15 frames to reduce single-frame jitter; the residual envelope 34 is allowed to increase rapidly when the suspended object accelerates, turns, or when the mask edge fluctuates more, and then slowly shrinks after continuous stable observations. If the number of valid paired points is less than the preset number within an acquisition cycle, the coherent residual is not updated, and only the most recent valid value is used. This process prevents a small number of erroneous boundary points from suddenly pulling the overall correction direction off course.
[0063] Figure 4 This is a geometrical schematic diagram illustrating the formation process of the occlusion sweep-in zone 41 and the unresolved occupancy area 44 provided in an embodiment of the present invention. The residual mapping module 23 first translates the current rigid prediction boundary 32 according to the coherent residual to obtain the corrected prediction boundary 40. The area swept between the previous and current corrected prediction boundaries 40 is expanded outward by the residual envelope 34 to form the occlusion sweep-in zone 41. The occlusion sweep-in zone 41 reflects the newly covered working surface range of the actual occlusion contour of the suspended object between adjacent acquisition times, rather than a simple fixed outward expansion of the current projection of the suspended object.
[0064] The trajectory-keeping module 24 uses the occlusion sweep-in zone 41 covering the unseen personnel position 35 as the reason for the disappearance of personnel detection. The time alignment tolerance can be from 50 milliseconds to 300 milliseconds, typically 150 milliseconds. If the tolerance is too small, the actual occlusion will be missed due to the time difference between visual inference and the crane's movement; if the tolerance is too large, ordinary missed detections after the crane has passed will be misclassified as crane occlusion. Only when the occlusion sweep-in zone 41 covers or simultaneously covers the unseen personnel position 35, and the corresponding personnel disappearance occurs within this tolerance, will the trajectory be converted to an undetermined personnel trajectory.
[0065] If personnel are continuously lost before the arrival of the occlusion sweep inlet band 41, the trajectory holding module 24 treats the event as a normal visual loss and does not establish an unresolved personnel trajectory. If multiple hanging objects exist simultaneously, the sweep inlet bands for each hanging object are calculated separately, and the hanging object that earliest covers the unseen position 35 of the personnel and meets the time relationship condition is taken as the occlusion association object. When multiple sweep inlet bands meet the condition simultaneously, the system retains multiple candidate associations until the subsequent exposure or viewpoint relay result eliminates ambiguity.
[0066] The upper limit of personnel movement speed is determined jointly by the operating procedures and on-site samples. The model training module 29 or the offline statistical program divides the distance between adjacent personnel positions in the labeled trajectory by the time difference of data collection to obtain the sample speed. The high percentile value of the sample speed can be taken from the 90th to the 99th percentile, plus a speed safety margin of 0.1 to 0.5 meters per second. The final upper limit of personnel movement speed is the larger value between the upper limit of the operating procedures and the boundary of the sample speed.
[0067] The upper limit for personnel movement speed can be set between 1.2 m / s and 2.5 m / s, typically 1.8 m / s. If this value is too small, the location that the material handler might reach during periods of invisibility will fall outside the personnel reachable zone 42; if it is too large, the personnel reachable zone 42 will quickly fill the passable area 17. Different job types can be configured with different upper limits, but when the job type is uncertain, the maximum value among all applicable job types should be used to avoid narrowing the occupied area due to misjudgment of personnel category.
[0068] The trajectory-keeping module 24 takes the last visible position 35 as the starting point, multiplies the upper bound of the person's movement speed by the cumulative actual time since the last visible moment, and obtains the maximum travel distance for the current period. This distance is then used to perform obstacle-constrained wavefront expansion along the passable area 17, forming... Figure 4 Personnel can reach zone 42. The cumulative actual time is obtained directly by subtracting the last seen timestamp from the current synchronization time, and is not calculated based on the number of processed frames. Therefore, temporary frame reduction, inference blocking, or video frame loss will not underestimate the distance that personnel may move.
[0069] The current actual obstruction area 30 is expanded outward according to the residual envelope 34 to form the obstruction allowable area 43. The obstruction allowable area 43 represents the location that may still be obstructed by the suspended object, and is not equivalent to all accessible locations. The intersection of the accessible area 42 and the obstruction allowable area 43 is determined as the unresolved occupancy area 44. Figure 4 The intersecting texture area in the text represents the intersection, so the unresolved occupancy area 44 will expand as people are not visible, while being geometrically limited by the obstruction of the hanging object.
[0070] In some embodiments, the unresolved occupancy area 44 is represented by a grid, vector polygon, or signed distance field. The grid resolution can be between 0.05 meters and 0.25 meters, with a coarser resolution used at the edge 12 when computing power is low, and further refined locally when closer to the hazard zone 55 of the suspended object. When updating the region, the main components connected to the previous time step are retained, and isolated fragments with an area smaller than the area occupied by a single person can be merged into the nearest main component or retained according to the safety side.
[0071] The hazard identification module 25 obtains the current operating target, load size, and sling length from the tower crane control terminal 11. The current operating target can be the target slewing angle, target amplitude, target hook height, or a speed curve already generated by the controller. Future hook states are interpolated using speed and acceleration limits consistent with the tower crane controller, avoiding impossible linear instantaneous movement. The interpolated future hook positions are connected in chronological order to form... Figure 5 The future hook trajectory 50 in the middle.
[0072] The preset time window covers the entire process from image acquisition to effective deceleration generated by the safety interlock. The worst-case end-to-end delay can be 0.15 seconds to 0.8 seconds, typically 0.4 seconds; too small a delay will miss the tail delays of encoding, inference, and message transmission, while too large a delay will unnecessarily lengthen the prediction range. The braking time can be 0.5 seconds to 3 seconds, typically 1.5 seconds; its value is determined by the tower crane load, slewing speed, and brake test results. Too small a value will underestimate the stopping distance, while too large a value will expand the impact range of the shutdown.
[0073] The preset time window can be 2 to 8 seconds, typically 4 seconds. Its lower limit is no less than the sum of the worst-case end-to-end delay and braking time, plus a control margin of 0.2 to 1 second. The maximum speed of the hoisted load is measured during on-site no-load and rated-load test runs. The maximum travel distance within the predicted window is equal to the product of the measured maximum hoisted load speed and the preset time window. This distance is also used to check whether the future hook trajectory 50 covers the area it may pass through before the safety interlock completes braking.
[0074] At each predicted location on the future hook trajectory 50, the hazard identification module 25 projects the physical contour of the suspended object according to its size, attitude, and sling length. The contours at each predicted location are then merged to obtain... Figure 5 The dashed line on the left represents the nominal physical sweep envelope of the suspended object 51. The longer the sling, the greater the permissible lateral offset of the swing. For scenarios where the attitude sensor data of the suspended object is available, the physical contour can be directly rotated according to the attitude. For scenarios where the attitude sensor data is lacking, a conservative contour surrounding each permissible attitude of the suspended object is used.
[0075] The hazard identification module 25 extrapolates the predicted coherent residuals at each predicted position based on the rate of change of the coherent residuals over continuous time intervals. The preset upper limit for the swing can be between 0.2 meters and 1.5 meters, typically 0.8 meters. An upper limit that is too small will truncate reasonable offsets during long slings or rapid rotations, while an upper limit that is too large will cause the correction to lose the constraints of on-site observation. When the predicted coherent residual exceeds the upper limit, it is proportionally limited in direction, and then the corresponding outline of the suspended object is translated to form... Figure 5 The envelope 52 on the left is corrected for predicted coherent residuals.
[0076] To prevent the overall translation correction from overlooking local attitude changes, the nominal suspended object's physical sweep envelope 51 and the envelope 52 corrected by the predicted coherent residuals are first joined, and then extended using the residual envelope 34 and a safety distance determined by time delay, braking, and maximum speed. This sequence preserves two possible positions: the rigid trajectory and the residual extrapolated trajectory, without replacing the nominal envelope with the corrected envelope. The residual envelope 34 is used to cover observation discreteness, and the safety distance is used to cover continued motion before the control action takes effect; both have different sources.
[0077] The hazard identification module 25 also projects the physical outline of the suspended object at each predicted location onto the working surface 14 along the direction of gravity, and extends it outward according to the preset fall safety distance to form the suspended object fall projection 53. The lower limit of the preset fall safety distance is obtained by multiplying the horizontal velocity of the suspended object by the free fall time, and then adding the shape margin for the suspended object to overturn or scatter; the free fall time is calculated based on the height of the suspended object from the working surface 14 and the gravitational acceleration. Under low-position, low-speed hoisting conditions, the configuration benchmark can be taken as 0.3 meters to 2 meters, typically 1 meter; when the calculation lower limit exceeds 2 meters, the calculated value is directly used, without being limited by the 2-meter upper limit. If the distance is too small, it is difficult to cover the horizontal velocity retained by the suspended object at a high position and the overturning or scattering after landing; if the distance is too large, it will include areas unrelated to hoisting in the control.
[0078] The hook position in the future hook trajectory 50 is expanded outward according to the hook's external dimensions and swing margin to obtain the hook running channel 54. The expanded physical contour of the suspended object, the object's fall projection 53, and the hook running channel 54 are geometrically joined to form... Figure 5 The irregular outer boundary on the right represents the hazardous area 55 for the suspended object. The hazardous area is not a rectangle that encloses the entire prediction frame, but is obtained by combining three types of actual physical influence areas. Figure 5 The unresolved occupancy area 44 is superimposed on the dangerous area 55 of the suspended object to visually represent the overlap between the two.
[0079] When the hazardous area 55 of the hoisted object and the unresolved occupancy area 44 overlap in an area larger than one grid cell, or when the distance between their boundaries is less than the configured proximity threshold, the hazard identification module 25 outputs the hazard identification result for the hoisting of hidden personnel. The overlap result carries the trajectory identifiers of the associated hoisted object and the unresolved personnel, as well as the overlap area, the estimated arrival time, and the current residual quality level. The linkage control module 26 simultaneously sends the result to the tower crane control terminal 11 and the corresponding image acquisition terminal 13, ensuring that the control action and evidence collection are based on the same event identifier.
[0080] After receiving the result, the tower crane control terminal 11 verifies the equipment status using existing safety interlocks and executes a deceleration or stop command. The edge terminal 12 directly drives the brake without bypassing the tower crane controller. If the estimated arrival time is long and the danger zone is still expanding, speed limiting can be implemented first; if the estimated arrival time is shorter than the braking time or the overlap is rapidly increasing, a stop command is executed. The image acquisition terminal 13 simultaneously improves the encoding quality of the event area and saves images before and after the hazard occurs for on-site verification.
[0081] In one embodiment, the linkage control module 26 uses an event state machine to avoid message jitter. An event is generated when the overlap condition is met for the first time, and the status is elevated to confirmed after 2 to 5 consecutive frames. If the overlap temporarily disappears but the unresolved personnel trajectory has not been cleared, the warning is maintained and the tower crane speed limit is not immediately revoked. The event only enters the deactivated state after the personnel are re-detected and their identity is matched, the relay viewpoint confirms that the personnel have left, or the risk is clearly eliminated through manual verification.
[0082] Figure 7 This is a schematic diagram illustrating the comparison between conditional exposure prediction and regional framing process provided in an embodiment of the present invention. Starting from the current moment, the exposure prediction module 27 generates multiple rigid boundaries for future data acquisition moments based on the tower crane's operating volume. These boundaries are then translated using coherent residuals and expanded using residual envelope 34 to obtain the predicted obstruction area 70. In the predicted obstruction areas 70 at adjacent moments, the portion of the previous area that moves out of the next area forms an exposure zone 71. The exposure zone 71 represents the work surface area that will be released from the obstruction of the suspended load according to the current operating trend; it does not necessarily indicate that personnel will appear.
[0083] The exposure prediction module 27 sorts multiple exposure zones 71 according to future acquisition times and calculates their intersection with the unresolved occupancy area 44. The earliest overlapping boundary among the exposure zones 71 is determined as the priority exposure edge 72. If the exposure zone 71 remains separate from the unresolved occupancy area 44 during the prediction period, the trajectory maintenance module 24 continues to update the unresolved occupancy area 44 according to the cumulative actual time, without prematurely clearing the trajectory of unresolved personnel due to the movement of the suspended object. This process establishes a clear geometric relationship between exposure opportunities and potential personnel occupancy.
[0084] The linkage control module 26 generates a frame increment command 73 that includes the target image region, the future acquisition time, and the duration, based on the back projection position of the preferentially exposed edge 72 in the image. For example... Figure 7 As shown, the frame increment command 73 is sent by the linkage control module 26 to the image acquisition terminal 13 that forms the trajectory of the undetermined person. In the figure, the unlabeled dashed line pointing from the camera to the target area only represents the line of sight. The control command is separated from the line of sight to avoid misinterpreting the direction of the camera's view as the direction of the control message flow.
[0085] The image acquisition terminal 13 can increase the local sampling frequency of the image region where the priority exposure edge 72 is located while keeping the overall frame rate unchanged, or it can increase the overall frame rate within a limited time period. The duration of the regional frame increment can be 0.5 seconds to 3 seconds, starting 0.2 seconds to 1 second before the expected exposure time. If the terminal does not support regional readout, a short-term full-frame increment is used; the edge terminal 12 automatically restores the normal acquisition configuration after the frame increment ends to prevent long-term bandwidth occupation after the event ends.
[0086] The AI image recognition module 22 saves the identity features of the person from 1 to 10 frames before the detection disappeared, and selects representative features based on clarity, occlusion ratio, and detection confidence. After the person is re-detected within the image area where the priority exposure edge 72 is located, the cosine similarity between the current identity features and the representative features is calculated, and verified in conjunction with the time interval and reachability distance. The preset matching threshold can be set between 0.55 and 0.85, typically 0.70. If the threshold is too low, it is easy to mismatch people with similar clothing; if the threshold is too high, the same person will be rejected due to changes in posture and lighting.
[0087] After a successful match, the re-detected personnel position is first projected onto the work surface 14. If the position is outside the hazardous area 55 of the suspended object, the trajectory holding module 24 clears the corresponding unresolved personnel trajectory; if the position is still within the hazardous area 55 of the suspended object, the original probability expansion center is replaced by the measured position, and the unresolved occupancy area 44 is shrunk according to the current personnel detection position. If the match fails or only low-quality detection is performed, the unresolved state continues, and safety control is not lifted based on a single frame of low confidence result.
[0088] Figure 8 This is a schematic diagram of the multi-viewpoint relay acquisition and cloud-based verification process 87 provided in an embodiment of the present invention. The viewpoint relay module 28 maintains the field of view, camera calibration relationship, and availability status of multiple candidate image acquisition terminals 80. Each candidate terminal performs the same working surface projection on the suspended object mask as the main terminal, forming a candidate occlusion area 81; then, based on the tower crane's operating volume, the candidate occlusion area 81 of the previous moment is changed and compared with the current candidate occlusion area 81 to form a candidate residual envelope 82.
[0089] The candidate residual envelope 82 is calculated based on the observation quality of each camera and cannot directly copy the residual envelope 34 of the main image acquisition terminal 13. The viewpoint relay module 28 first excludes terminals whose field of view cannot cover the unresolved occupancy area 44, whose time synchronization is abnormal, or whose reprojection error fails. Then, it expands the candidate occlusion area 81 with the corresponding candidate residual envelope 82. The terminal with the smallest overlap area with the unresolved occupancy area 44 after expansion is determined as the relay image acquisition terminal 83. When the overlap area is the same, the terminal with a larger line-of-sight angle and lower network latency is selected first.
[0090] The linkage control module 26 sends a directional acquisition command 84 to the relay image acquisition terminal 83, specifying the undetermined occupancy area 44 and the acquisition period. Figure 8 The directional acquisition command 84 is derived from the linkage control module 26 and enters the relay image acquisition terminal 83, which then outputs the relay video 85. The acquisition period covers a short interval before and after the expected exposure time, and candidate cameras are not required to upload all videos routinely.
[0091] The AI image recognition module 22 selects images of unaccounted personnel from their trajectories before occlusion, extracts their identity features, and searches for matching targets in the relay video 85. The matching location is projected onto the work surface 14 using the calibration relationship of the relay image acquisition terminal 83. The trajectory maintenance module 24 uses this location to narrow down the unaccounted area 44, or clears the corresponding area when the personnel have left the danger zone and remain visible. Differences in color and focal length between different cameras can be mitigated through feature normalization and camera domain adaptation samples.
[0092] When the feature matching confidence level is lower than the preset matching threshold, the AI image recognition module 22 outputs a pending review status and hands the event over to the cloud for review 87. The preset time period can be 3 to 15 seconds, typically 8 seconds. Too short a time period may lack a complete occlusion process, while too long a time period will increase event caching and upload volume. The synchronous acquisition module 21 writes the preset time period video and relay video 85 before the detection disappearance time into the local cache 86, saving only the segments related to the event identifier.
[0093] When the network is unavailable, the local cache 86 continues to retain event videos, and on-site deceleration or shutdown is unaffected. After communication is restored, the edge device 12 uploads videos in batches according to the severity of the event and the time of occurrence. The cloud review 87 receives the main viewpoint segment, the relay video 85, and the timestamp, as well as the model version and regional geometric summary. The cloud only handles evidence review and model improvement and does not participate in the tower crane's real-time safety interlocking. When the cache capacity is insufficient, priority is given to retaining unresolved events and high-risk events, while low-risk segments that have completed review are overwritten according to the strategy.
[0094] When the suspended object mask fails continuously, the suspended object completely moves out of the field of view, or the image is contaminated and obscured, the residual mapping module 23 cannot form a new actual obstruction boundary 31. At this time, the system freezes the most recent effective coherent residual and residual envelope 34, and does not interpret the empty mask as the suspension object disappearing. The hazard identification module 25 continues to expand the suspended object danger zone 55 outward according to the preset swing upper limit based on the frozen value, and the linkage control module 26 sends a deceleration request to the tower crane control terminal 11 and outputs the manual verification status to the field terminal.
[0095] The aforementioned freeze strategy is also employed when the time synchronization deviation exceeds the allowable range, the tower crane operation is interrupted, or the image timestamp bounces back. If the interruption lasts for more than 0.3 to 1 second, the future position will no longer use operation extrapolation, and the load danger zone 55 will be conservatively extended according to the nearest speed upper limit and the preset swing upper limit. If a safe lead time cannot be guaranteed, the tower crane control terminal 11 will be stopped by the safety interlock. This degradation path does not rely on a data source that remains correct after a fault occurs.
[0096] Normal recovery requires three conditions to be met simultaneously: the suspended object mask is output stably for 3 to 10 consecutive frames; the time deviation between the image and the tower crane's movement falls back into the time synchronization tolerance; and the reprojection error of the independent control points passes verification. During recovery, the coherent residual is first reinitialized with the current actual boundary, and then the conservative envelope is gradually reduced over an observation period of 3 to 15 frames. The deceleration state is not immediately canceled in the first recovery frame.
[0097] When the system is first started and there is no valid residual history, the coherent residual is initialized to zero. The residual envelope 34 is a conservative combination of the calibration error propagation amount, the preset safety margin, and the preset upper limit of oscillation. After accumulating a sufficient number of valid boundary pairs, the system smoothly transitions to the field statistical envelope. This cold start strategy allows the system to still operate when there is no historical data, but the danger zone is relatively conservative.
[0098] In other embodiments, when the boundary of the suspended object mask contains a large area gap but still has a sufficient number of stable segments, the boundary motion residual 33 can be calculated only on the stable segments. The stable segments are determined by the consistency of the boundary curvature and optical flow of consecutive frames, and updates stop when the length is less than 20% to 40% of the complete boundary. The coherent residuals obtained from local updates are still limited by a preset upper bound for swing, and the residual envelope 34 does not shrink due to the reduction of samples.
[0099] The system can verify the effectiveness of the closed-loop residual from the same source through event video playback. Verification data includes normal hoisting, footage of personnel remaining behind the load and personnel leaving, as well as footage of general missed detections and camera time deviations. Comparison schemes include a scheme using fixed independent margins on both the personnel and load sides, and a scheme where boundary motion residual 33 simultaneously drives the obstruction cause gate, the unresolved occupancy area 44, and the load danger zone 55. Both schemes use the same video, the same model output, and the same tower crane operation volume to avoid input differences affecting the conclusions.
[0100] Envelope coverage is defined as the proportion of effective frames that fall within the range corrected by coherent residuals and expanded by the residual envelope 34, after the actual occlusion boundary 31. Hazard lead time is defined as the time difference between the initial output of the hidden personnel hoisting hazard identification result and the arrival of the expected danger zone in the unresolved occupancy area 44. False stop rate is calculated as the number of events that trigger a stop without actual overlap risk out of the total number of stop events; false detection rate is calculated as the number of events with actual overlap risk that do not output results within the braking lead time out of the total number of actual risk events. During the verification phase, the definition and calculation process of the indicators are recorded, and speculative results are not entered when real test records are lacking.
[0101] To verify the three-terminal time closed loop, image acquisition time, AI inference completion time, and result transmission time can also be recorded, as well as the tower crane control terminal 11 receiving time and safety interlock action time. The worst-case end-to-end delay is calculated using the highest quantile value from multiple playback or on-site commissioning samples, rather than the average. Braking time is determined by the safety test interval under different loads and speeds. When either changes, the preset time window and safety distance are recalculated without altering the residual calculation and personnel positioning algorithm.
[0102] Reference Figure 2 The device can be implemented by one or more processors, memory, and communication interfaces. The synchronous acquisition module 21, AI image recognition module 22, and residual mapping module 23 can be functional units within the same edge computing process. The trajectory holding module 24, hazard identification module 25, and linkage control module 26 can also be distributed across multiple containers sharing a common clock. The exposure prediction module 27 and viewpoint relay module 28 can be selectively enabled based on the number of cameras on site. The model training module 29 is typically deployed on an offline training server.
[0103] The synchronous acquisition module 21 performs timestamp alignment, runtime interpolation, and event caching; the AI image recognition module 22 performs multi-task inference and adjacent frame association; the residual mapping module 23 performs coordinate transformation, ray intersection, boundary pairing, and residual statistics; the trajectory holding module 24 performs pending state migration and region update; the hazard identification module 25 performs future trajectory interpolation, hazard zone construction, and overlap determination; and the linkage control module 26 performs instruction orchestration, retransmission, and safety interlock interface adaptation. The modules communicate with each other via message passing with event timestamps and coordinate system version numbers to avoid mixing regions with different calibration versions in calculations.
[0104] The memory can store camera calibration parameters, work surface geometry, and passable area 17. It can also store AI multi-task model parameters, the most recent valid residual, and pending personnel trajectories, as well as event video indexes. After the equipment restarts, the model parameters and calibration parameters can be loaded directly, while pending personnel trajectories are processed according to the restart interval. If the restart interval is shorter than 1 to 5 seconds and the video and operation volume are continuous, the pending state can be restored; if the interval is longer or the time axis is not continuous, the conservative danger zone is re-established according to the cold start strategy, and normal hoisting is resumed only after on-site confirmation.
[0105] Optionally, the edge terminal 12 uses sequence numbers, timeout retransmission, and message authentication codes for control messages. The tower crane control terminal 11 executes the repeating sequence number only once, rejects deceleration or stop messages that have exceeded their validity period, and reports a link anomaly. The image acquisition terminal 13 returns the frame increment configuration effective time, and the relay image acquisition terminal 83 returns the directional acquisition start time, enabling the linkage control module 26 to verify whether the acquisition action falls within the expected exposure period.
[0106] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Those skilled in the art should understand that modifications can be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features, and such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A construction site safety management system that integrates edge image acquisition with AI hazard identification, characterized in that, include: The synchronous acquisition module, located at the edge and connected to the image acquisition terminal, is used to acquire timestamped images or videos of the construction site and tower crane operation data, and to obtain the passable area of the work surface. The AI image recognition module is used to identify personnel trajectories and suspended object masks from the construction site images or videos; The residual mapping module is used to map the suspended object mask onto the working surface to obtain the actual suspended object occlusion area and its actual occlusion boundary. It uses the tower crane operation amount to transform the actual occlusion boundary of the previous moment into the current rigid prediction boundary, and compares the current actual occlusion boundary with the current rigid prediction boundary to obtain the boundary motion residual and residual envelope. The trajectory holding module is used to establish the trajectory of the unseen person when the occlusion sweeping band, after being corrected by the boundary motion residual, covers the unseen position of the person and the person detection disappears. The unseen occupancy area is obtained by expanding the upper bound of the person's movement speed, the cumulative actual time from the time corresponding to the unseen position of the person, and the residual envelope. The hazard identification module is used to generate a physical sweep envelope of the hoisted object using the tower crane's operating volume, correct the physical sweep envelope of the hoisted object according to the boundary motion residual, and expand the hoisted object's danger zone using the residual envelope. It is also used to output the hazard identification result of the hoisting of hidden personnel when the hoisted object's danger zone overlaps with the unresolved occupancy zone. The linkage control module is used to send the identification results of the hidden personnel hoisting hazard to the tower crane control terminal and the image acquisition terminal.
2. The construction site safety management system linking edge image acquisition and AI hazard identification as described in claim 1, characterized in that, It also includes a model training module; The AI image recognition module includes a shared feature extraction network, a personnel detection head, a suspended object segmentation head, and a personnel re-identification head; The model training module is used to acquire construction site image or video frame sequences with personnel detection box labels, suspended object mask labels and personnel identity labels. The construction site image or video frame sequences are input into the shared feature extraction network to obtain shared image features. The shared image features are then input into the personnel detection head, the suspended object segmentation head and the personnel re-identification head to obtain candidate personnel detection boxes, candidate suspended object masks and candidate personnel identity features. The model training module is also used to calculate the detection loss between the candidate person detection box and the person detection box label, calculate the segmentation loss between the candidate hanging object mask and the hanging object mask label, calculate the metric learning loss between the candidate person identity features using the person identity label, and update the parameters of the shared feature extraction network, the person detection head, the hanging object segmentation head, and the person re-identification head using the detection loss, the segmentation loss, and the metric learning loss. The AI image recognition module is used to output personnel detection boxes and the suspended object mask using the trained personnel detection head and the suspended object segmentation head. It is also used to output personnel identity features using the trained personnel re-identification head, and associate the personnel detection boxes in adjacent frames according to the personnel identity features and the timestamp of the construction site image or video to obtain the personnel trajectory.
3. The construction site safety management system linking edge image acquisition and AI hazard identification as described in claim 2, characterized in that, The synchronous acquisition module is also used to acquire camera calibration relationship and working surface geometry. The camera calibration relationship includes camera intrinsic parameters and the transformation relationship between the camera coordinate system and the tower crane coordinate system. The residual mapping module is used to construct the line of sight through the boundary pixel of the suspended object mask from the camera optical center, obtain the intersection point of each line of sight with the working surface, and connect the intersection points to form the actual occlusion boundary; The residual mapping module is also used to construct a line of sight for a person passing through the midpoint of the bottom edge of the person detection frame from the optical center of the camera, and to determine the intersection of the line of sight for the person and the working surface as the position where the person was not seen. The residual mapping module is also used to determine the rigid displacement of the hook based on the rotation angle, amplitude change position and hook height of two adjacent moments, and to obtain the current rigid prediction boundary by changing the actual occlusion boundary of the previous moment according to the rigid displacement of the hook. The residual mapping module is further configured to pair each sampling point of the current rigid prediction boundary with the point with the smallest distance on the current actual occlusion boundary, determine the displacement vector obtained by subtracting the position of the corresponding sampling point on the current rigid prediction boundary from the position of the paired point on the current actual occlusion boundary as the boundary motion residual, determine the median value of each component of the boundary motion residual as the coherent residual, and combine the high quantile value of the amplitude of each boundary motion residual relative to the coherent residual, the camera calibration error and the preset safety margin to form the residual envelope.
4. The construction site safety management system linking edge image acquisition and AI hazard identification as described in claim 3, characterized in that, The residual mapping module is used to translate the current rigid prediction boundary according to the coherent residual to obtain the corrected prediction boundary; The residual mapping module is also used to obtain the area swept from the correction prediction boundary at the previous time to the correction prediction boundary at the current time, and to expand the swept area using the residual envelope to obtain the occlusion sweep-in band. The trajectory maintaining module is used to establish the trajectory of the undetermined person when the detection of the corresponding person disappears within the time alignment tolerance after the occlusion sweeping band covers the unseen position of the person.
5. The construction site safety management system linking edge image acquisition and AI hazard identification as described in claim 4, characterized in that, The trajectory keeping module is used to obtain the upper limit of personnel speed specified in the operation procedure, the marked personnel trajectory samples, and the preset speed safety margin. The distance between adjacent personnel positions of the marked personnel trajectory samples is divided by the difference of the corresponding acquisition time to obtain the sample personnel speed. The sum of the high quantile value of the sample personnel speed and the preset speed safety margin is determined as the sample speed boundary. The larger value between the upper limit of personnel speed and the sample speed boundary is determined as the upper limit of personnel movement speed. The trajectory-keeping module is also used to expand along the passable area to obtain a personnel reachable area, taking the unseen position of the personnel as the initial area and the product of the upper bound of the personnel's movement speed and the cumulative actual time from the time corresponding to the unseen position of the personnel. The trajectory-keeping module is also used to expand the current actual suspended object obstruction area according to the residual envelope to obtain the obstruction allowable area, and to determine the intersection of the personnel reachable area and the obstruction allowable area as the unresolved occupancy area.
6. The construction site safety management system linking edge image acquisition and AI hazard identification as described in claim 3, characterized in that, The synchronous acquisition module is also used to obtain the current operating target, the size of the suspended object, and the length of the sling from the tower crane control terminal; The hazard identification module is used to interpolate the current operating target of the tower crane to obtain the future hook trajectory within a preset time window, project the physical contour of the suspended object to each predicted position of the future hook trajectory according to the size of the suspended object and the length of the sling, and merge the physical contours of the suspended object corresponding to each predicted position to obtain the physical sweep envelope of the suspended object. The hazard identification module is also used to extrapolate the predicted coherent residual of each predicted position according to the coherent residual change at continuous time, limit the predicted coherent residual according to the preset swing upper limit, move each of the suspended objects' physical contours according to the predicted coherent residuals, merge the suspended objects' physical contours before and after the movement, and then expand them using the residual envelope. The hazard identification module is also used to project the physical outline of the suspended object corresponding to each of the predicted positions onto the working surface along the direction of gravity and expand it outward according to the preset fall safety distance to obtain the falling projection of the suspended object, connect the hook position in the future hook trajectory and expand it outward according to the hook size to obtain the hook running channel, and merge the expanded physical outline of the suspended object, the falling projection of the suspended object and the hook running channel into the suspended object danger zone; The linkage control module is used to send a deceleration command or a stop command to the tower crane control terminal when outputting the hidden personnel hoisting hazard identification result, and the safety interlock of the tower crane control terminal executes the deceleration command or stop command.
7. The construction site safety management system linking edge image acquisition and AI hazard identification as described in claim 1, characterized in that, It also includes a prediction module; The exposure prediction module is used to generate multiple rigid prediction boundaries for future acquisition times according to the tower crane operation volume, correct each rigid prediction boundary according to the boundary motion residual, and use the residual envelope to expand to obtain the prediction occlusion area. The exposure prediction module is also used to determine the area where the previous predicted occlusion area moves out of the next predicted occlusion area in two adjacent predicted occlusion areas as an exposure zone, and when the exposure zone overlaps with the unresolved occupancy area, to determine the boundary of the exposure zone that first forms the overlap as the priority exposure edge, and to maintain the trajectory of the unresolved personnel while the exposure zone and the unresolved occupancy area remain separated.
8. The construction site safety management system linking edge image acquisition and AI hazard identification as described in claim 7, characterized in that, The linkage control module is used to send a frame increment command containing the priority exposed edge and the corresponding future acquisition time to the image acquisition terminal that forms the trajectory of the undetermined person; The image acquisition terminal is used to increase the acquisition frame rate of the image region where the priority exposed edge is located at the corresponding future acquisition time; The AI image recognition module is used to extract the features of the person before the occlusion disappeared, and to match the features of the person re-detected in the image area where the priority exposed edge is located with the features of the person before the occlusion. The trajectory keeping module is used to clear the unresolved personnel trajectory when the matching is successful and the re-detected personnel position is outside the danger zone of the suspended object, and is also used to shrink the unresolved occupancy area according to the personnel position when the matching is successful and the re-detected personnel position is inside the danger zone of the suspended object.
9. The construction site safety management system linking edge image acquisition and AI hazard identification as described in claim 1, characterized in that, It also includes a viewpoint relay module; The viewpoint relay module is used to obtain the field of view and camera calibration relationship of multiple candidate image acquisition terminals that are communicatively connected to the synchronous acquisition module. It maps the suspended object mask identified from the construction site images or videos acquired by each candidate image acquisition terminal to the working surface to obtain candidate occlusion areas. It transforms the candidate occlusion areas at the previous moment according to the tower crane operation volume and compares the transformation result with the candidate occlusion areas at the current moment to obtain candidate residual envelopes. The viewpoint relay module is also used to expand each of the candidate occlusion areas using the corresponding candidate residual envelope, and select the candidate image acquisition terminal with the smallest overlap area between the expanded candidate occlusion area and the unresolved occupancy area from the candidate image acquisition terminals whose field of view covers the unresolved occupancy area as the relay image acquisition terminal.
10. The construction site safety management system linking edge image acquisition and AI hazard identification according to claim 9, characterized in that, The linkage control module is used to send a directional acquisition command to the relay image acquisition terminal, which defines the unresolved occupancy area and the acquisition time period; The relay image acquisition terminal is used to acquire relay video according to the directional acquisition command; The AI image recognition module is used to select images of people before the moment of disappearance from the trajectory of the undetermined persons and extract the features of the persons before the occlusion. It is also used to identify the location of the relay personnel from the relay video that matches the features of the persons before the occlusion. The trajectory-keeping module is used to correct or clear the unresolved occupied area using the position of the relay personnel; The AI image recognition module is also used to output a status pending cloud review when the feature matching confidence is less than a preset matching threshold. The synchronous acquisition module is used to store the video of the preset time period before the detection disappearance time and the relay video in the local cache when the status pending cloud review is output, and upload them to the cloud for review when communication is available.
Citation Information
Patent Citations
Safety monitoring and method and system of tower crane
CN110733981A
Safety control system and control method of tower crane
CN110733983A
Pedestrian multi-target tracking method combining attention mechanism end-to-end training
CN112560656A