Man-machine anti-collision holographic safety early warning system based on multiple field cameras

By deploying multiple cameras at the construction site and utilizing computer vision and deep learning technologies, a virtual fence is dynamically generated for real-time early warning, which solves the problems of insufficient accuracy and real-time performance of existing human-machine collision avoidance early warning systems and improves the efficiency of construction safety management.

CN121640685APending Publication Date: 2026-03-10SHANGHAI PUDONG ROAD & BRIDGE CONSTR +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies in human-machine collision avoidance early warning systems at construction sites lack accuracy and real-time performance, making it difficult to provide effective early warnings for construction machinery and personnel, especially in complex environments where they cannot accurately identify targets or provide timely feedback.

Method used

A holographic safety early warning system based on multiple field cameras is adopted. Combined with computer vision technology, deep learning and convolutional neural networks are used to establish a construction machinery posture recognition model and a personnel distribution detection model. A virtual fence is dynamically generated for real-time early warning, and vibration and sound and light alarms are used to remind operators.

Benefits of technology

It enables accurate identification and real-time early warning of machinery and personnel at construction sites, reduces the occurrence of human-machine collision accidents, improves the efficiency of construction safety management, and provides dynamic hazardous area analysis and early warning reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121640685A_ABST
    Figure CN121640685A_ABST
Patent Text Reader

Abstract

The invention provides a man-machine anti-collision holographic safety early warning system based on multiple field cameras, and relates to the field of building construction safety management. The man-machine anti-collision holographic safety early warning system based on the multiple field cameras comprises a field information acquisition module, a matching module, a comparison module, a prediction model construction module, a safety analysis module and an early warning module. Dynamic dangerous operation area judgment and real-time early warning are realized according to the motion state of the excavator; and finally, a mechanical driver and safety management personnel are simultaneously reminded to complete an early warning process through a vibration alarm arranged under a driver seat and an audible and visual alarm arranged at the top of a driving cabin in the early warning module, so that man-machine collision accidents are reduced. According to the method, a man-machine collision accident feedback mechanism is established, and a safety management research report of the engineering machinery and personnel is generated through statistical analysis of potential dangers of the construction machinery in a key scene, so that the construction safety management efficiency of the construction machinery is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of construction safety management, specifically to a human-machine anti-collision holographic safety early warning system based on multiple field-deployed cameras. Background Technology

[0003] Construction sites are generally defined as structured spaces comprised of various resources, including personnel, equipment, and materials, involved in dynamic work tasks. These resources are frequently in motion and can be close to each other. Among them, construction machinery, as a crucial component of construction sites, plays an indispensable role in construction operations. Its large market share ensures the smooth progress of numerous infrastructure projects, but it also brings many safety issues closely related to construction machinery.

[0004] Currently, my country's construction safety management mainly relies on on-site inspections and supervision by safety management personnel. This approach is inefficient and suffers from significant delays. Furthermore, construction safety officers, limited by experience and skill levels, often struggle to make accurate predictions and timely warnings when dangers are imminent, failing to meet the demands of on-site safety management. Therefore, it is imperative to address the urgent need to utilize technological means in complex construction environments with high personnel and machinery density and mobility. This involves replacing "human intervention" with "technical prevention," intervening in, predicting, and warning of construction safety accidents related to engineering machinery, transforming passive warnings into proactive ones, reducing or avoiding human-machine collisions during construction operations, and improving the efficiency and timeliness of on-site safety management.

[0005] Human-machine collision avoidance early warning is a safety management method targeting collisions caused by workers getting too close to machinery. By monitoring the proximity between workers and equipment (or vehicles), potential hazards can be detected in advance when workers and machinery are too close, providing timely feedback to the workers (such as visual, acoustic, and vibration alarms). This proactive intervention allows workers to prepare for evasive action, thereby reducing the probability of collisions.

[0006] Therefore, the human-machine collision avoidance warning method is essentially a human-machine distance warning. Its warning capability depends on the ability to perceive the target and the ability to judge the distance. Depending on the technology used, it is mainly divided into distance sensor-based methods and computer vision-based methods.

[0007] Distance sensor-based early warning methods are further divided into time-of-flight (TOF) sensing and tag-based sensing. The former primarily measures distance to the surrounding environment (such as obstacles, workers, and machinery) by emitting some form of energy and reading its time of flight; this category includes technologies like lidar, millimeter-wave radar, and sonar. The latter utilizes energy fields (such as electromagnetic fields and magnetic fields) and detects proximity between machinery and workers through signal communication between tags installed on them; this includes radio frequency identification (RFID), magnetic field recognition (MF), and Bluetooth Low Energy (BLE). Because these different sensing technologies have their own strengths and weaknesses in terms of application scope, effectiveness, and cost, they exhibit varying early warning effects in research on human-machine collision avoidance early warning methods for construction sites.

[0008] The most common ranging method for lidar is to use high-power, narrow-bandwidth pulsed lasers to determine the distance to an object by calculating the time difference of laser reflection. This method offers advantages such as high accuracy, strong anti-interference capabilities, and high speed in ranging. Furthermore, multi-axis lasers can scan the 3D environment and perform 3D reconstruction, which can be used as a data source for path planning in machinery. However, like other Time-of-Flight (TOF) based sensors, lidar cannot distinguish the detected object itself, requiring the use of other technologies for target identification. It also suffers from drawbacks such as high power consumption, high cost, large size, and difficulty in installation.

[0009] Sonar (or ultrasonic sensors) are commonly used for obstacle detection in autonomous driving and robotic walking. They measure the distance to an object by emitting high-frequency sound waves and measuring the time-of-flight (TOF) of the echo reflected back from the target object. Because the propagation of sound waves is affected by the physical conditions of the medium, this method is typically only suitable for short-range detection of distances less than 3 meters.

[0010] Millimeter-wave radar uses radio signals (300MHz-40GHz) and does not require a medium for propagation. Therefore, this method has strong anti-interference capabilities, is suitable for many complex outdoor conditions, and has an effective range exceeding 30 meters. Furthermore, this method can measure not only the proximity of objects but also their speed. However, when encountering unfavorable reflective surfaces, such as plastic, wood, and large flat objects commonly found at construction sites, the radio signals are easily absorbed and dispersed. Therefore, this method is not suitable for construction scenarios.

[0011] In contrast, tag-based sensing methods are unaffected by the propagation medium and reflective surfaces, but their measurement accuracy and precision are somewhat inferior. According to tests conducted by Park et al., the distance measurement errors of RFID, MF, and BLE sensors reached 5m, 3.4m, and 2.6m, respectively. The accuracy of GPS sensors is easily affected by weather conditions. These tag-based methods require tagging each target and deploying base stations and signal receivers in multiple locations on the construction site to achieve area coverage. This results in significant application costs in projects involving large numbers of personnel and equipment or large areas. Furthermore, the high rates of loss and damage, coupled with the disruption to workers, make them unsuitable for safety monitoring on construction sites.

[0012] Early warning methods based on computer vision. Currently, cameras are a common means of assessing the overall daily safety status of projects, typically used to identify potential hazards and monitor security. Meanwhile, with breakthroughs in hardware and algorithms and the widespread acceptance and application of video surveillance by owners, computer vision-based methods have shown great potential in the research and application of distance-based early warning. This method, leveraging its ability to analyze and process image data, can accurately capture key information about targets in the scene, offering unique advantages in scene understanding and signal feedback. Early warning methods based on computer vision are divided into two types based on the type of vision: stereo vision and monocular vision.

[0013] Stereo vision typically employs binocular cameras or multiple jointly calibrated monocular cameras to capture monitoring scenes. Depth information within the scene can be obtained through parallax for distance measurement. For example, Seokho Chi et al. developed an image-based automated safety assessment method for earthmoving machinery activities. Using a Bumblebee XB3 binocular camera, they monitored the distance and speed of dump trucks within a 75-meter range using frame-by-frame extraction, achieving overspeed and close-range violation detection for mobile construction machinery. However, they did not test or analyze its real-time performance, and the discontinuation of the binocular camera resulted in weak generalization of the research findings. Ishimoto et al. installed a binocular vision device consisting of two monocular cameras at the rear of an excavator, enabling personnel detection and collision warning in the blind spot behind the machinery. However, due to limitations in camera angle and baseline, target detection failed when the distance between personnel and the camera exceeded 16 meters or was less than 1 meter, failing to provide effective warnings. Fang et al. combined computer vision, sensing, and point cloud 3D reconstruction technologies to develop a real-time active safety assistance framework for truck crane operations. This framework uses distributed sensors to capture motion information of key crane components, reconstructs the crane's posture, and automatically reconstructs and updates the site environment and surrounding obstacles using point cloud data. It can provide early warnings to operators through a graphical user interface, but the system has not yet met the real-time requirements for safety warnings. Li Heng et al. calibrated dump trucks using 3D bounding boxes and achieved target ranging in 3D space using stereo vision constructed from multiple cameras. Due to limitations in the annotation method and the objects being annotated, this method is only applicable to machines with shapes approximating cuboids. It cannot acquire spatial information for deformable bodies like excavators that rotate, and it does not consider the buffer time required for mechanical braking and operator reaction.

[0014] Monocular cameras, due to the loss of depth information, require additional auxiliary information for camera calibration to establish a 2D-to-3D mapping. Kim et al. used a fixed monocular camera and wearable warning devices to warn of human-machine collision risks; however, the target localization and ranging parts require pre-setting three points in the scene for camera calibration, which is complex and requires strict camera stillness, thus failing to achieve fully automated warning. Similarly, MingzhuWang et al. used a fixed camera to monitor the proximity of excavators and workers by taking oblique overhead shots, mainly judging by the position of the target detection boxes in the 2D image. This is strict about the camera's installation angle and cannot solve the problem of position misjudgment caused by perspective. To overcome the limitations of the camera's field of view, Kim et al. used drones for vertical overhead shots, solving the distance error caused by the field of view. However, their distance measurement also requires a reference object of known length on site for camera calibration and lacks an effective way to provide warning information feedback. Summary of the Invention

[0015] (a) Technical problems to be solved: To address the shortcomings of existing technologies, this invention provides a human-machine collision avoidance holographic safety early warning system based on multiple field-deployed cameras. Starting from accuracy, real-time performance, and operability, it adopts a computer vision-based early warning method, using a monocular camera as a data sensing device. Based on the movement status of the excavator, it dynamically judges and provides real-time early warnings of dangerous work areas to address the collision risks between personnel and excavators.

[0016] (II) Technical Solution: To achieve the above objectives, the present invention is implemented through the following technical solution: a human-machine anti-collision holographic safety early warning system based on multiple field cameras, including a field information acquisition module, a matching module, a comparison module, a prediction model construction module, a safety analysis module, and an early warning module; The on-site information acquisition module is used to collect holographic image information of the outside of the construction site and obtain a set of on-site external holographic image information, including site-mounted cameras, vibration sensors, AI servers and wireless transmission modules; The matching module is used to match the set of construction machinery operation information and the set of holographic images of construction personnel distribution to obtain the matching information group and send it to the prediction model construction module; The comparison module is used to perform on-body comparison analysis and driving comparison analysis on the construction machinery operation information set and the construction machinery posture image information set, respectively, and to obtain the on-body comparison analysis set and the driving comparison analysis set, which are then sent to the safety analysis module. The prediction model building module is used to build prediction models and then send the completed prediction models to the security analysis module. The prediction model uses computer vision technology to establish a large database for construction machinery posture recognition and a machinery posture recognition and tracking model. It also establishes a real-time multi-level virtual fence and early warning mechanism based on the working status, and uses an algorithm to automatically generate dangerous areas based on the dynamic changes in the working status and moving speed of the machinery to provide boundary crossing warnings. The safety analysis module is used to comprehensively analyze the safety of the posture of construction machinery and the distribution of construction personnel, and to obtain a comprehensive analysis set which is then sent to the early warning module. The early warning module is used to perform early warning analysis based on the comprehensive analysis set, match the early warning analysis results with the preset early warning strategies, and execute the preset corresponding early warning operations.

[0017] Preferably, the prediction model constructs target detection and posture recognition datasets for personnel and construction machinery in a construction scenario, trains and learns them through a deep learning network, and establishes a deep learning-based detection model and posture recognition model for personnel and machinery to detect and locate the distribution of construction machinery and personnel.

[0018] Preferably, the mechanical posture recognition model in the prediction model identifies the mechanical posture by constructing a key point detection dataset for construction machinery, training and fitting the three-dimensional spatial relationship of key points through a convolutional neural network, and determining the working state of the machinery based on the correspondence between posture and working state and the time relationship between frames.

[0019] Preferably, the automatic danger zone generation algorithm and boundary crossing warning algorithm in the prediction model automatically generate dynamic virtual fences around the danger zones of construction machinery, and issue real-time warnings to construction machinery and construction personnel entering the danger zones. That is, within the direction of travel and working range, the system predicts the possibility of safety accidents based on the type of machinery, working status, distance between people and machinery, and speed of machinery movement, and sets three warning levels from high to low. If personnel enter the dangerous working area of ​​construction machinery, the warning system will trigger vehicle-mounted audio and visual warnings to remind the on-site operator and construction worker, and store image information with time, location and machinery type in the database, and send it to the on-site management personnel for viewing and processing in real time.

[0020] Preferably, the loss function in the keypoint detection algorithm model adopts a multi-objective optimization function. Based on the loss function of YOLOv5, a keypoint loss is added, and the improved loss function includes the target detection box regression localization loss. Target regression loss Target category regression classification loss and key point retrieval and positioning loss Four parts; the target regression loss and the target class regression classification loss adopt the original function definition of YOLOv5, and the specific definitions of the other two losses are as follows: Object detection bounding box regression localization loss CIoU is used as the supervisory signal for the detection box, and the formula is defined as follows: ; Where s is the scale factor index, The target pixel position is k, and the anchor point number is k; For the k-th anchor point The prediction result at the s-th scale of the pixel position, with 3 anchor points of different sizes and 4 scales of output results at each pixel position; Keypoint regression localization loss To address the issue of asynchronous loss functions between keypoints and bounding boxes during training, the principle of IoU loss is extended from bounding boxes to keypoints. Target-Keypoint Similarity (OKS) is used as the evaluation criterion for the loss function, representing the similarity between the predicted keypoint location and the ground truth value. The formula is defined as follows: ; in, The Euclidean distance between the predicted keypoints and the actual keypoints is given by s, where s represents the scaling factor. Weighting factors for each key point, This represents the visibility of keypoints, where 0 indicates that they are not labeled and 1 indicates that they are labeled.

[0021] Preferably, in order to associate the detection box and keypoints in each target and avoid the post-processing of clustering in the bottom-up approach, the keypoint regression localization loss uses the distance loss between the keypoints and the detection box of the same target as a weighting factor. After the detection box anchor point is matched, the regression loss of all keypoints is calculated using that anchor point as the center point, and the OKS of each keypoint is multiplied by the weight factor. After summing, the keypoint regression localization loss is defined as follows: ; in, The loss is calculated by comparing the predicted and ground truth values ​​of the i-th keypoint location with the distance from the anchor center of the target detection box, using Smooth L1 Loss. Therefore, when the anchor positions of the target detection boxes match, the above loss function is effective in image location. At point k, the loss function with an output scale of s is valid; The total loss function of the keypoint detection algorithm for a single image is the sum of the four loss functions mentioned above for all locations, all anchor points, and all output scales in the image. The formula is defined as follows: ; in, , , and These are the hyperparameters that balance the loss values ​​at each scale; By combining an SVM classifier and selecting a time span threshold that balances accuracy and real-time performance, the real-time working status of the excavator can be determined.

[0022] Preferably, the warning area of ​​the dynamic virtual fence is the top of the boom when the excavator's three-bar linkage is at its maximum extension. Using the height threshold as the maximum operating radius of the excavator as L, and using L, 1.3L, and 1.6L as length thresholds, a dynamic virtual fence for a three-level cylindrical space warning area is set, represented by red, yellow, and green respectively.

[0023] Preferably, the human-machine spatial distance in the warning area is equivalent to the worker's foot position with the bottom midpoint of the worker detection frame and the center of the rotation axis with the bottom midpoint of the excavator detection frame, and its world coordinates are: The calculation formula is: .

[0024] (III) Beneficial Effects: This invention provides a holographic safety early warning system for human-machine collision avoidance based on multiple field-deployed cameras. It offers the following advantages: This invention applies computer vision technology to the safety management of construction machinery and personnel at construction sites. Multiple field cameras deployed at the construction site acquire a global view of the machinery's operating area. An embedded core algorithm enables target recognition, status recognition, positioning, and distance measurement of machinery and personnel. Based on the excavator's movement status, it dynamically identifies and provides real-time warnings for hazardous operating areas to address collision risks between personnel and excavators. Finally, a vibration alarm located under the driver's seat and an audible and visual alarm on the top of the cab simultaneously alert the driver and safety management personnel to complete the warning process, reducing the occurrence of human-machine collision accidents. By establishing a human-machine collision accident feedback mechanism and conducting statistical analysis of potential hazards in key construction machinery scenarios, a safety management research report on construction machinery and personnel is generated, improving the efficiency of construction safety management. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the system architecture of the early warning system of the present invention; Figure 2 This is a schematic diagram of the target recognition structure of the present invention; Figure 3 This is a schematic diagram of the on-site panoramic mechanical target detection results of the present invention; Figure 4 This is a schematic diagram of the on-site personnel detection results of the present invention; Figure 5 This is a schematic diagram of the excavator identification skeleton of the present invention; Figure 6 This is a schematic diagram of the training and testing process of the detection model of the present invention; Figure 7 This is a schematic diagram of frame-sampling test of the key point detection model of the present invention; Figure 8 This is a schematic diagram of the excavator key point annotation interface of the present invention; Figure 9 This is a schematic diagram illustrating the hazardous area division of the present invention; Figure 10 This is a schematic diagram of the on-site hazardous area identification interface of the present invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] Example: like Figure 1 As shown, this embodiment of the invention provides a human-machine anti-collision holographic safety early warning system based on multiple field cameras, including a field information acquisition module, a matching module, a comparison module, a prediction model construction module, a safety analysis module, and an early warning module; The on-site information acquisition module collects holographic image information from outside the construction site, obtaining an external holographic image information set, including site-mounted cameras, vibration sensors, an AI server, and a wireless transmission module. The matching module matches the construction machinery operation information set with the holographic image information set of construction personnel distribution, obtaining a matching information set which is then sent to the prediction model construction module. The comparison module performs on-site comparison analysis and driving comparison analysis on the construction machinery operation information set and the construction machinery posture image information set, obtaining an on-site comparison analysis set and a driving comparison analysis set, which are then sent to the safety analysis module. The prediction model construction module constructs a prediction model and sends the completed prediction model to the safety analysis module. The safety analysis module performs a comprehensive analysis of the safety of the construction machinery posture and the distribution of construction personnel, obtaining a comprehensive analysis set which is then sent to the early warning module. The early warning module performs early warning analysis based on the comprehensive analysis set, matches the early warning analysis results with preset early warning strategies, and executes the corresponding preset early warning operations. The early warning module sends early warning commands to the vibration alarm installed under the driver's seat and the audible and visual alarm on the top of the driver's cab to complete the early warning reminder. The early warning system is based on four major architectural layers: application layer, algorithm layer, hardware layer, and data layer. The application layer includes the tracking, detection, positioning, and status display of workers and machinery at the construction site, as well as the display of early warning boundaries. It also includes vehicle-mounted sound and light alarms, vibration alarms, and background screenshots and information notifications. The algorithm layer performs target detection and tracking based on monocular camera linkage calibration, and performs monocular depth estimation and then uses wireless sensors for data fusion and transmission. The hardware layer includes data devices for data acquisition, such as field cameras and vibration sensors; devices for data processing, such as AI servers; devices for data transmission, such as wireless transmission modules; and devices for response, such as sound and light alarms. The data layer includes raw data, such as camera video streams; and backend data, such as screenshots of violations and key events.

[0028] The predictive model employs computer vision technology to establish a large database for construction machinery posture recognition and a machinery posture recognition and tracking model. It also establishes a real-time multi-level virtual fence and early warning mechanism based on the working status, and uses an algorithm to automatically generate dangerous areas based on the dynamic changes in the working status and movement speed of the machinery to provide boundary crossing warnings. It can effectively distinguish dangerous working areas of different types of machinery under different states and provide accurate and timely warnings for human-machine collision accidents.

[0029] Specific implementation examples include: 1) Object detection algorithms based on deep learning frameworks and convolutional neural network algorithms: This study focuses on construction machinery and personnel, constructing datasets for target detection and posture recognition in construction scenarios. Deep learning networks are used for training and learning to establish detection and posture recognition models for personnel and machinery, enabling real-time detection and localization of these objects.

[0030] like Figure 2-4 As shown. Human-machine target detection based on YOLO. The YOLO target detection algorithm is applied to identify excavators and construction workers at the construction site and extract the pixel coordinates of the targets.

[0031] 2) Working state recognition algorithm based on key point identification and posture recognition: Taking commonly used construction machinery (excavators) as the research object, a dataset for key point detection of construction machinery is constructed, such as... Figure 5 As shown, the mechanical posture is identified by training a convolutional neural network to fit the three-dimensional spatial relationship of key points, and the working state of the machine is determined based on the correspondence between posture and working state and the temporal relationship between frames.

[0032] First, kinematic analysis of the backhoe excavator's working device was performed to obtain key parameters of its posture changes. A motion state SVM classifier based on feature descriptors was designed to match the excavator's posture changes with its motion state. Then, based on the excavator's posture analysis results, a relevant image dataset was established. Through training a neural network based on improved YOLOv5 for target detection and feature point extraction, a model for excavator target detection and key point detection regression was built. This model identified the extracted lower main body (composed of the traveling and slewing devices) and key points in the working device, forming the excavator's skeleton, such as... Figure 5 As shown.

[0033] The proposed keypoint detection algorithm model employs a multi-objective optimization function in its loss function. Based on the loss function of YOLOv5, a keypoint loss is added, and the improved loss function includes object detection box regression and localization loss. Target regression loss Target category regression classification loss and keypoint regression localization loss Four parts. The target regression loss and target class regression classification loss use the original YOLOv5 function definitions. The other two losses are defined as follows: Object detection bounding box regression localization loss For the target scale transformation problem, common target detection loss functions often employ IoU-based optimization functions, such as GIoU, DIoU, and CIoU, rather than simple distance-based bounding box loss functions. We use CIoU as the supervision signal for the detection box, defined as follows: ; Where s is the scale factor index, is the target pixel position, and k is the anchor point number. For the k-th anchor point The prediction result at the s-th scale of the pixel location, with 3 anchor points of different sizes and 4 scales of output results at each pixel location.

[0034] Key point retrieval and positioning loss To address the issue of asynchronous loss functions between keypoints and bounding boxes during training, the principle of IoU loss is extended from bounding boxes to keypoints. Target-Keypoint Similarity (OKS) is used as the evaluation criterion for the loss function, representing the similarity between the predicted keypoint location and the ground truth value. The formula is defined as follows: ; in, The Euclidean distance between the predicted keypoints and the actual keypoints is given by s, where s represents the scaling factor. Weighting factors for each key point, The visibility of key points (0 indicates that they are not labeled, and 1 indicates that they are labeled).

[0035] To associate bounding boxes and keypoints for each target and avoid the post-processing of clustering in bottom-up methods, the keypoint regression localization loss uses the distance loss between keypoints and bounding boxes of the same target as a weighting factor. After the detection box anchor point is matched, the regression loss of all keypoints is calculated using that anchor point as the center point, and the OKS of each keypoint is multiplied by the weight factor. After summing, the keypoint regression localization loss is defined as follows: ; in, The loss is calculated by comparing the predicted and ground truth values ​​of the i-th keypoint location with the distance from the anchor center of the target detection box, using Smooth L1 Loss. Therefore, when the anchor positions of the target detection boxes match, the above loss function is effective in image location. At point k, the loss function with an output scale of s is valid.

[0036] The keypoint detection algorithm proposed in this invention provides a total loss function for an image that is the sum of the four loss functions mentioned above for all locations, all anchor points, and all output scales in the image. The formula is defined as follows: ; in, , , and These are the hyperparameters that balance the loss values ​​at each scale. Finally, by combining an SVM classifier and selecting a time span threshold that balances accuracy and real-time performance, the real-time working status of the excavator can be determined.

[0037] like Figure 6 As shown, taking the binary classification problem to be solved as an example, assume a training dataset in a feature space is given: ; in, , , , Let i be the i-th eigenvector. For class markers, when The time is represented as a positive example; The time interval is represented as a negative example. The hyperplane is defined as follows: ,in These are sample points on the hyperplane. It is a vector perpendicular to the hyperplane. The geometric margin of the sample points in the hyperplane is: ; The minimum geometric interval of all sample points is: ; Support vectors are the feature vectors at the points where the geometric margin to the hyperplane is minimized. Based on this definition, the maximum hyperplane problem for solving the SVM model can be expressed as the following constrained optimization problem: ; Through normalization processing ; ; ; get, ; Maximize through calculation Equivalent to maximizing This is also equivalent to minimizing Therefore, solving the maximum partitioning hyperplane problem in the SVM model is transformed into solving a convex quadratic programming problem with inequality constraints.

[0038] ; The optimal solution to this constraint can be found by constructing a Lagrange function: ; in It is a Lagrange multiplier, and Based on the duality of the Lagrange function, the positions of the maximum and minimum variables are swapped, forming a new dual problem: ; Its corresponding classification decision function is the step function: ; in, .

[0039] For linearly separable problems, the inner product between samples can be directly calculated following the steps above. In the binary classification problem of excavator motion state recognition to be implemented in this section, the distribution of key variables is not completely linearly separable. Therefore, by using nonlinear transformation, the problem can be transformed into a linear classification problem in a certain feature space, and then the problem can be solved by learning a linear support vector machine in a high-dimensional feature space.

[0040] This is mainly achieved by replacing the inner product in the objective function and classification decision function with a kernel function. The kernel function, or positive definite kernel, is the inner product function or inner definite kernel between two instances after a nonlinear transformation, representing the existence of a mapping from the input space to the feature space. For any input space and ,satisfy: ; Since the MSF value is determined by key variables in two adjacent frames, when the time span between the two selected frames is large, the accuracy of the SVM classifier can be estimated to be high, but the real-time performance of the result recognition is low; conversely, when the time span between the two selected frames is small, the accuracy of the SVM classifier is low, but the real-time performance is high. Given that the collision avoidance warning signal in this project needs to demonstrate sufficient accuracy while ensuring real-time warning performance, an optimal time period needs to be determined as the discrimination interval for the excavator's movement state to balance accuracy and real-time performance.

[0041] Generally, video is processed by extracting frames at a rate of 3 frames per second to reduce computation. This invention, however, tests are conducted by selecting different time periods of 10 frames per second. The principle is as follows: Figure 7 As shown, assuming the time span of the discrimination interval is T, it is defined as follows: ; Where n is the number of units, with 10 frames as the unit. The frames are used as discrimination intervals. Frame 1 and frame 10n are selected as keyframes to calculate the corresponding excavator motion state descriptor (MSF). Frame 10n belongs to both discrimination intervals; it is both the last frame of the previous interval and the first frame of the next interval. Based on different values ​​of n, MSFs for excavator motion state under different intervals are constructed and used as training and testing sets for the SVM classifier for performance comparison. The MSF under the optimal interval is selected to determine whether the excavator's current motion state is working or stationary. The specific training process is as follows: Select multiple video clips containing an excavator, each 2 seconds long. Divide these videos into segments of 10, 20, and 30 frames each, and select the corresponding 10n (n=1,2,3) frames from each segment as keyframes. The MSF between keyframes in the above segment is calculated as the sample value of the video segment. ; The excavator's status in each video segment is marked by human judgment, indicating when it is in working condition. When at rest, .

[0042] The above results and Pairs constitute the training sample set The model was then used to train an SVM binary classification model, thereby obtaining a classifier for recognizing the movement state of the excavator.

[0043] 3) Multi-level dynamic hazard area automatic generation algorithm and boundary crossing warning algorithm based on monocular vision: Based on the above, this invention provides a dynamic virtual fence that automatically generates hazardous areas for construction machinery, issuing real-time warnings to construction machinery and personnel entering these areas. Within the direction of travel and operating range, the system predicts the likelihood of a safety accident based on machinery type, operating status, distance between operator and machinery, and machinery speed, establishing three warning levels from high to low. If personnel recklessly enter the hazardous operating area of ​​the construction machinery, the warning system will trigger onboard audio-visual warnings to the on-site operator and construction worker, and store accompanying image information including time, location, and machinery type in a database, sending it in real-time to on-site management personnel for review and action. like Figure 8 , Figure 9 and Figure 10 As shown, the warning area of ​​the dynamic virtual fence is based on the apex of the boom when the excavator's three-bar linkage is at its maximum extension. Using the height threshold and the maximum operating radius of the excavator as L, and using L, 1.3L, and 1.6L as length thresholds, a dynamic virtual fence is set for a three-level cylindrical space warning zone, represented by red, yellow, and green respectively. The human-machine spatial distance within the warning zone is equivalent to the worker's foot position at the bottom midpoint of the worker's detection frame, with world coordinates [value missing]. The center of the rotation axis is equivalent to the bottom midpoint of the excavator's detection frame, with world coordinates [value missing]. The calculation formula is: ; Based on the calculation formula for human-machine spatial distance, the worker's status is categorized from farthest to closest: safe state, approaching state, dangerous state, and emergency state. Using medium-sized hydraulic excavators commonly found on construction sites as a reference, this early warning system can provide on-site personnel with 3-6 seconds of reaction time upon successful warning without delay, thereby effectively preventing human-machine collision accidents.

[0044] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A human-machine anti-collision holographic safety warning system based on a plurality of field camera lenses, characterized in that: The method comprises a field information collection module, a matching module, a comparison module, a prediction model construction module, a safety analysis module and a warning module. The field information collection module is used for collecting holographic image information outside the construction site to obtain a set of holographic image information outside the site, and comprises a field layout camera, a vibration sensor, an AI server and a wireless transmission module. The matching module is used for matching the construction machinery operation information set and the construction personnel distribution holographic image information set to obtain a matching information group and transmit the matching information group to the prediction model construction module. The comparison module is used for respectively performing ontology comparison analysis and driving comparison analysis on the construction machinery operation information set and the construction machinery posture image information set to obtain an ontology comparison analysis set and a driving comparison analysis set and transmit the ontology comparison analysis set and the driving comparison analysis set to the safety analysis module. The prediction model construction module is used for constructing a prediction model and transmitting the constructed prediction model to the safety analysis module. The prediction model adopts computer vision technology, establishes a large database of construction machinery posture recognition, a machinery posture recognition model and a tracking model, sets up a real-time multi-level virtual fence based on the working state and a warning mechanism, and performs out-of-bound warning based on a dangerous area automatic generation algorithm based on the dynamic changes of the machinery working state and the moving speed. The safety analysis module is used for comprehensively analyzing the safety of the construction machinery posture and the construction personnel distribution position to obtain a comprehensive analysis set and transmit the comprehensive analysis set to the warning module. The warning module is used for performing warning analysis according to the comprehensive analysis set, matching the warning analysis result with a preset warning strategy, and executing a preset corresponding warning operation.

2. The human-machine anti-collision holographic safety warning system based on multiple field camera according to claim 1, characterized in that: In the prediction model, a personnel and engineering machinery target detection and posture recognition dataset under a building construction scene is constructed, a deep learning network is trained and learned, a personnel and machinery detection model and a posture recognition model based on deep learning are established, and the construction machinery and personnel distribution are detected and positioned.

3. The human-machine anti-collision holographic safety warning system based on multiple field camera according to claim 1, characterized in that: In the prediction model, a construction machinery key point detection dataset is constructed, a convolutional neural network is trained and fitted to recognize the posture of the machinery by fitting the three-dimensional spatial relationship of the key points, and the working state of the machinery is determined according to the correspondence between the posture and the working state and the time relationship between the upper and lower frames.

4. The human-machine anti-collision holographic safety warning system based on multiple field camera according to claim 1, characterized in that: In the prediction model, a dangerous area automatic generation algorithm and an out-of-bound warning algorithm generate a dynamic virtual fence of the dangerous area of the construction machinery, issue a real-time warning to the engineering machinery and construction personnel driving into the dangerous area, that is, in the direction of travel and the range of work, the possibility of a safety accident is predicted according to the type of machinery, the working state, the distance between the man and the machine and the moving speed of the machinery, and three warning levels are set from high to low. If a person enters the dangerous working area of the construction machinery, the warning system will trigger the vehicle-mounted sound and light warning to prompt the on-site operator and construction personnel, and store the image information with the time, position and machinery category in the database and send it to the on-site management personnel for viewing and processing.

5. The human-machine anti-collision holographic safety warning system based on multiple field camera according to claim 3, characterized in that: The loss function in the key point detection algorithm model adopts a multi-objective optimization function, and a key point loss is added on the basis of the loss function of YOLOv5, and the improved loss function comprises four parts of target detection frame regression positioning loss , target regression loss , target category regression classification loss and key point regression positioning loss ; wherein the target regression loss and the target category regression classification loss adopt the function definition originally of YOLOv5, and the other two losses are specifically defined as follows: Target bounding box regression positioning loss : CIoU is used as the supervision signal of the bounding box, and the formula is defined as follows: ; Where s is the scale factor index, The target pixel position is k, and the anchor point number is k; For the k-th anchor point The prediction result at the s-th scale of the pixel position, with 3 anchor points of different sizes and 4 scales of output results at each pixel position; Keypoint regression positioning loss : To solve the problem of loss function asynchronization in the training process of key points and detection boxes, the principle of IoU loss is extended from detection boxes to key points, and OKS is used as the evaluation basis of the loss function. It is the similarity between the predicted key point position and the true value. The formula is defined as follows: ; wherein, is the Euclidean distance between the predicted keypoint and the ground truth keypoint, s denotes a scale factor, is a weight factor for each keypoint, is the visibility of the keypoint, wherein 0 means not labeled and 1 means labeled.

6. The human-machine anti-collision holographic safety warning system based on multiple field camera according to claim 5, characterized in that: In order to associate the bounding box and the key points in each target, avoid the post-processing of clustering in the bottom-up method, the key point regression positioning loss takes the distance loss between the key points of the same target and the bounding box as a weight factor After matching the anchor point of the bounding box, the regression loss of all key points is calculated with the anchor point as the center point, and the OKS of each key point is multiplied by the weight factor After addition, the key point regression positioning loss is defined as follows: ; wherein, is the distance loss of the predicted value of the ith key point position and the real value from the anchor center of the target detection box, and Smooth L1 Loss is used for calculation; therefore, when the anchor position of the target detection box matches, the loss function above is effective at the image position , which belongs to the kth anchor, and the loss function with the output scale s is effective. The total loss function of the key point detection algorithm for an image is the sum of the above four loss functions at all positions, all anchor points and all output scales in the image, and the formula is defined as follows: ; wherein, , , and are hyperparameters balancing the respective scale loss values. At the same time, combined with SVM classifier, select the balanced accuracy and real-time time span threshold, realize the discrimination of real-time working state of excavator.

7. The human-machine anti-collision holographic safety warning system based on multiple field-of-view cameras according to claim 4, characterized in that: The early warning area of the dynamic virtual fence is the top point of the large arm under the maximum arm span of the three connecting rods of the excavator The height threshold is L, the length thresholds are L, 1.3L and 1.6L, and the dynamic virtual fence of the three-cylinder space early warning area is set, which is represented by red, yellow and green respectively.

8. The human-machine anti-collision holographic safety warning system based on multiple field camera according to claim 7, characterized in that: The man-machine space distance in the pre-warning area is equivalent to the worker foot position by the midpoint of the bottom of the worker detection frame, and is equivalent to the center of the swing shaft by the midpoint of the bottom of the excavator detection frame, and the world coordinates are , and the calculation formula is: 。