A Construction Safety Management Method Based on Excavator Perspective Image Fusion

By using excavator-perspective image fusion technology to dynamically correct UWB positioning coordinates, the problem of false alarms caused by tag posture interference was solved, thus ensuring the reliability of safety management in construction scenarios and guaranteeing the timely response and efficiency of the early warning system.

CN120853327BActive Publication Date: 2026-01-06福建融茂水利水电工程有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511374925.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-01-06
Estimated Expiration
2045-09-25

AI Technical Summary

Technical Problem

UWB positioning accuracy is easily affected by the tag wearing posture, leading to false alarms from the electronic fence, affecting the construction process and reducing workers' trust in the alarm system.

Method used

By fusing images from the excavator's perspective, the geographical location and operating conditions of equipment within the construction area are obtained. The visual acquisition device is used to identify personnel bounding boxes and identification codes. The positioning coordinates are corrected by combining posture feature vectors and posture-error compensation models, and dangerous areas are dynamically delineated and graded early warnings are triggered.

Benefits of technology

Significantly reduce false alarms, improve positioning accuracy and the reliability of construction safety management, and ensure the timely response and efficiency of the early warning system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853327B_ABST
    Figure CN120853327B_ABST
Patent Text Reader

Abstract

The application discloses a construction safety management method based on image fusion of excavator visual angle, and belongs to the technical field of construction early warning, and specifically comprises the following steps: acquiring the position and working condition of the excavating equipment and setting a dynamic warning fence; triggering the equipment vision device to scan personnel targets in real time, identifying personnel bounding boxes and extracting personnel identity codes; synchronously calling corresponding original coordinates of ultra-wideband tags, capturing posture visual frames and generating posture feature vectors through topology analysis of skeletal joints; inputting the posture vectors into a posture-error compensation model to output a positioning offset prediction value to correct the original coordinates; judging the spatial topological relationship between the corrected coordinates and the boundary of the warning fence, and triggering a hierarchical early warning instruction when the corrected coordinate point falls within the fence; the application realizes real-time compensation of the positioning error of personnel through dynamic and accurate matching of the construction dangerous area, maximally avoids the occurrence of misjudgment, improves the accuracy of active prediction and alarm of safety risks, and realizes intelligent safety management of the construction site.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of construction early warning technology, specifically to a construction safety management method based on excavator-perspective image fusion. Background Technology

[0002] Ultra-wideband (UWB) technology, with its nanosecond-level extremely narrow pulse signal characteristics, is widely used in high-precision positioning both indoors and outdoors, especially in scenarios such as industrial construction, mining, and smart construction sites. By measuring parameters such as time of flight (TOF) and time difference of arrival (TDoA), it can achieve centimeter-level positioning accuracy, providing core technical support for personnel safety management and equipment collaborative scheduling. For example, in construction scenarios, UWB positioning systems, through the deployment of base stations and tag terminals worn by personnel, can track the location of workers in real time. Combined with electronic fence technology, it can provide early warning of intrusion into dangerous areas, significantly reducing the incidence of safety accidents.

[0003] However, in practical applications, UWB positioning accuracy is easily affected by the tag-wearing posture, leading to a significant increase in positioning errors and seriously impacting system reliability. Specifically, the posture of UWB tags worn by personnel (such as terminals integrated into safety helmets and wristbands) changes due to work actions (such as bending over, looking down, turning around, etc.), causing the tag antenna's radiation direction to deviate, thus altering the signal propagation path. This antenna directional shift leads to attenuation of the direct wave signal and an increase in the proportion of non-line-of-sight (NLOS) propagation; consequently, it triggers false alarms from the electronic fence, disrupting normal construction processes and even causing worker fatigue with the alarm system, creating potential safety hazards. Summary of the Invention

[0004] The purpose of this invention is to provide a construction safety management method based on excavator-view image fusion, and to solve the following technical problems:

[0005] UWB positioning accuracy is easily affected by the way the tag is worn, which can easily trigger false alarms from electronic fences, disrupt normal construction processes, and even cause workers to lose trust in the alarm system, thus creating potential safety hazards.

[0006] The objective of this invention can be achieved through the following technical solutions:

[0007] A construction safety management method based on excavator-view image fusion includes the following steps:

[0008] S1. Obtain the geographical location and operating conditions of all excavating equipment in the construction area, and determine the corresponding alarm fence outside any excavating equipment based on the operating conditions.

[0009] S2. Trigger the vision acquisition device mounted on any of the excavation devices and perform real-time scanning of the surrounding area, identify personnel bounding boxes and extract personnel identification codes through target detection algorithms;

[0010] S3. Retrieve the original coordinate data of the corresponding ultra-wideband positioning tag according to the personnel identification code, and simultaneously capture the posture visual frame of the personnel.

[0011] S4. Based on the posture visual frame, generate the corresponding posture feature vector through skeletal joint spatial topology analysis. The posture feature vector includes the torso pitch angle, horizontal yaw angle and limb joint spatial distribution parameters.

[0012] S5. Input the attitude feature vector into the preset attitude-error compensation model to obtain the predicted value of the positioning coordinate offset. Correct the original coordinates of the person according to the predicted value of the positioning coordinate offset to obtain the corrected coordinates.

[0013] S6. Determine the spatial topological relationship between the corrected coordinates and the alarm fence boundary. Trigger a graded early warning command when the corrected coordinates point falls within the fence buffer zone.

[0014] As a further aspect of the present invention: in S1, the specific process for determining the alarm fence is as follows:

[0015] The attitude parameters of the working mechanism are collected in real time by a multi-dimensional sensor group installed on the excavating equipment. The attitude parameters include the boom pitch angle α, the forearm swing angle β, the bucket tilt angle γ, and the three-dimensional coordinates (x, y, z) of the end effector. A kinematic model of the working mechanism is established based on the inverse kinematics algorithm, and the maximum working radius R of the excavating equipment is calculated. The maximum working radius R is superimposed with the preset safety threshold ΔR to obtain the danger radius d. A circle is drawn with the excavating equipment as the center and the danger radius d as the radius to obtain the corresponding alarm fence.

[0016] As a further aspect of the present invention: in step S2, the specific process for obtaining the personnel identification code is as follows:

[0017] Real-time image streams are acquired from a visual acquisition device, face regions are detected through a cascaded neural network, and illumination compensation and filtering preprocessing are performed on the face regions. For the preprocessed face regions, a deep feature extraction network is used to analyze their geometric structure and obtain the corresponding feature vectors.

[0018] The feature vector is compared with a preset facial feature database, and the similarity value with each feature vector in the database is calculated. The feature vector corresponding to the highest similarity value is selected from all similarity values, and the personnel identification code corresponding to the feature vector in the matching feature database is obtained and used as the personnel identification code of the person.

[0019] As a further aspect of the present invention: in step S3, the process of obtaining the original coordinate data is as follows:

[0020] Establish a mapping table that associates personnel identification codes with ultra-wideband positioning tag MAC addresses. The mapping table stores the unique tag MAC address and spatial location signature corresponding to each personnel identification code.

[0021] After obtaining the personnel identification code through visual recognition, the system queries the associated mapping table to retrieve the corresponding ultra-wideband tag MAC address; the ultra-wideband positioning system receives the nanosecond-level pulse signal emitted by the tag through the distributed base station network, calculates the original spatial coordinates of each tag based on the time difference of arrival algorithm, and generates a real-time positioning data stream with timestamp and MAC address;

[0022] Based on the retrieved tag MAC address and combined with the spatial location signature verification mechanism, the original coordinate data of the target tag at the corresponding time is extracted from the real-time location data stream.

[0023] As a further aspect of the present invention: the specific construction process of the preset attitude-error compensation model is as follows:

[0024] S11. Control the baseline test subject to perform multiple sets of preset posture actions at a fixed spatial position, and simultaneously collect the ultrawideband raw coordinates of the positioning tag terminal and the visual images at the corresponding time to generate a spatiotemporally aligned coordinate-posture dataset.

[0025] S12, Predetermine the reference attitude, and extract the ultrawideband raw coordinate data corresponding to the reference attitude as the reference true coordinates;

[0026] S13. Obtain a visual image of any non-reference pose and generate a non-reference pose feature vector through spatial topology analysis of skeletal joints.

[0027] S14. Calculate the three-dimensional spatial offset between the ultra-wideband original coordinates and the reference true coordinates corresponding to the non-reference attitude, and generate the non-reference attitude positioning error vector.

[0028] S15. The non-reference posture feature vector samples and the positioning error vector samples are aggregated to form a training sample set, and the posture-error compensation model is generated by driving the model training through machine learning algorithms.

[0029] As a further aspect of the present invention, it also includes pre-acquiring the same visual images and performing pre-acquisition at different resolutions, labeling each pre-acquisition image as Y1, Y2, ..., Yn; where n is the total number of pre-acquisition images, sequentially generating pose feature vectors for each pre-acquisition image and recording the generation time each time, generating corresponding data points in the coordinate system with resolution as the abscissa and generation time as the ordinate, and using a polynomial fitting method to perform curve fitting on the data points to obtain the fitting formula T(F).

[0030] As a further aspect of the present invention: In S4, if the same person is simultaneously identified by the visual acquisition devices of multiple excavation equipment, the corresponding resolution of the person in the images acquired by each equipment is obtained and substituted into the fitting formula T(F) to obtain the generation time corresponding to the images acquired by each equipment, and the acquisition image with the shortest generation time is selected as the person's posture visual frame.

[0031] As a further aspect of the present invention: in step S6, the specific process of triggering the graded early warning command is as follows:

[0032] Within the danger radius, several distance levels are set at preset unit distance intervals. The smaller the distance, the higher the warning level. At the same time, different warning levels correspond to different warning information.

[0033] The beneficial effects of this invention are:

[0034] 1) This invention constructs a posture-error compensation model. First, it extracts the personnel posture feature vector (including parameters such as torso pitch angle and limb joint distribution) based on the spatial topology analysis of skeletal joints. Then, it uses the model to predict the positioning coordinate offset and correct the original coordinates. This specifically solves the problem of ultra-wideband positioning deviation caused by changes in personnel posture, enabling the corrected coordinates to more accurately reflect the actual position of the personnel. It significantly reduces false alarms from electronic fences caused by posture interference, effectively avoiding interference with normal construction processes. Simultaneously, it enhances the workers' trust in the safety warning system, providing more reliable technical support for construction safety management.

[0035] 2) This invention uses a visual acquisition device mounted on the excavation equipment to scan the surrounding area in real time, and combines it with personnel identification codes to achieve synchronous acquisition of ultra-wideband positioning tags and posture visual frames. At the same time, based on the fitting formula of resolution and generation time, the optimal posture visual frame is selected when identifying multiple people, ensuring the efficiency of posture feature extraction. This improves the real-time performance and accuracy of personnel position monitoring in construction scenarios, and can quickly capture the behavior of personnel entering dangerous areas. It ensures that the safety management system can respond promptly to changes in personnel position in dynamic working environments, and provides precise technical support for the safe interaction between personnel and equipment.

[0036] 3) When determining the alarm fence, this invention dynamically delineates the danger zone by combining the posture parameters of the excavating equipment's operating mechanism (such as boom pitch angle, forearm swing angle, etc.); during the early warning stage, by correcting the topological relationship between the coordinates and the fence boundary, multiple warning levels are set within the danger radius, with higher warning levels for closer distances and corresponding to different warning information, so that the warning instructions can accurately match the actual risk level. This avoids the interference of excessive warnings in low-risk areas on construction efficiency, and can trigger strong warnings in a timely manner in high-risk situations, effectively balancing the continuity of the construction process and the rigor of safety assurance. Attached Figure Description

[0037] The invention will now be further described with reference to the accompanying drawings.

[0038] Figure 1 This is a schematic diagram of a construction safety management method based on excavator perspective image fusion according to the present invention. Detailed Implementation

[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0040] Please see Figure 1 As shown, this invention is a construction safety management method based on excavator-view image fusion, comprising the following steps:

[0041] S1. Obtain the geographical location and operating conditions of all excavating equipment in the construction area, and determine the corresponding alarm fence outside any excavating equipment based on the operating conditions.

[0042] S2. Trigger the vision acquisition device mounted on any of the excavation devices and perform real-time scanning of the surrounding area, identify personnel bounding boxes and extract personnel identification codes through target detection algorithms;

[0043] S3. Retrieve the original coordinate data of the corresponding ultra-wideband positioning tag according to the personnel identification code, and simultaneously capture the posture visual frame of the personnel.

[0044] S4. Based on the posture visual frame, generate the corresponding posture feature vector through skeletal joint spatial topology analysis. The posture feature vector includes the torso pitch angle, horizontal yaw angle and limb joint spatial distribution parameters.

[0045] S5. Input the attitude feature vector into the preset attitude-error compensation model to obtain the predicted value of the positioning coordinate offset. Correct the original coordinates of the person according to the predicted value of the positioning coordinate offset to obtain the corrected coordinates.

[0046] S6. Determine the spatial topological relationship between the corrected coordinates and the alarm fence boundary. Trigger a graded early warning command when the corrected coordinates point falls within the fence buffer zone.

[0047] In a preferred embodiment of the present invention, the specific process for determining the alarm fence in step S1 is as follows:

[0048] The attitude parameters of the working mechanism are collected in real time by a multi-dimensional sensor group installed on the excavating equipment. The attitude parameters include the boom pitch angle α, the forearm swing angle β, the bucket tilt angle γ, and the three-dimensional coordinates (x, y, z) of the end effector. A kinematic model of the working mechanism is established based on the inverse kinematics algorithm, and the maximum working radius R of the excavating equipment is calculated. The maximum working radius R is superimposed with the preset safety threshold ΔR to obtain the danger radius d. A circle is drawn with the excavating equipment as the center and the danger radius d as the radius to obtain the corresponding alarm fence.

[0049] In S1, the alarm fence determination process is as follows: The multi-dimensional sensor group installed on the excavating equipment will collect the attitude parameters of the working mechanism in real time. For example, an angle sensor installed at the connection shaft between the boom and the body can directly detect the angle of the boom's vertical rotation relative to the body, i.e., the boom pitch angle α; an angle sensor installed at the connection shaft between the forearm and the boom can detect the angle of the forearm's left and right swing relative to the boom, i.e., the forearm swing angle β; an angle sensor installed at the connection shaft between the bucket and the forearm can obtain the angle of the bucket's tilt relative to the forearm, i.e., the bucket tilt angle γ; and the three-dimensional coordinates (x, y, z) of the end effector (such as the bucket tip) can be directly measured by the position sensor installed at the end, or calculated by combining the mechanical structural parameters such as the length of the boom and forearm with α, β, and γ. Based on these attitude parameters, a kinematic model of the working mechanism is established using inverse kinematics algorithms. For example, given the structural dimensions such as the boom length and forearm length, the position of the end effector in space can be deduced by combining α, β, and γ, thereby determining the farthest distance that the end effector can reach during equipment operation, i.e., the maximum working radius R. Then, the maximum working radius R is superimposed with a preset safety threshold ΔR. For example, if the equipment can reach a certain range, an additional distance is added as ΔR to reserve a safety buffer space. The two are added together to obtain the danger radius d. Finally, a circle is drawn with the excavator's own position as the center and the danger radius d as the radius. This circle is the alarm fence corresponding to the excavator.

[0050] It is worth noting that the "position of the end effector in space can be deduced by combining α, β, and γ" is based on the inverse kinematics algorithm to establish a kinematic model of the working mechanism to obtain the maximum working radius R of the excavating equipment. This is an extension of existing technology, referring to the relevant content in the master's thesis "Research on End-of-Line Trajectory Planning and Control of Excavators" (Liu Can, Jinan University, 2024). First, the attitude parameters of the working mechanism are collected in real time by a multi-dimensional sensor group installed on the excavating equipment, including the boom pitch angle α, the forearm swing angle β, and the bucket tilt angle γ. Referring to the MDH parameter method used in the thesis, a kinematic model of the working mechanism is constructed. The transformation relationship between joint space (joint angles such as α, β, and γ) and Cartesian space (end effector position) is established through a homogeneous transformation matrix. Then, the inverse kinematics algorithm (such as the inverse solution process based on the geometric method in the thesis) is used, combined with the motion limit range of each joint (such as the joint angle range in the MDH parameter table of the thesis), to obtain the specific calculation process of the end effector coordinates.

[0051] ;

[0052] ;

[0053] ;

[0054] Where θ2 corresponds to α, θ3 corresponds to β, and θ4 corresponds to γ; θ1 refers to the slewing joint angle of the excavator, that is, the horizontal rotation angle of the slewing platform relative to the base coordinate system (usually with the center of the excavator body as the origin); L3 refers to the length of the excavator stick, which is a fixed link structure parameter in the kinematic model and is used to characterize the spatial extension of the stick link in the calculation of the shovel tip position coordinate; L1 is the boom length, that is, the distance from the hinge point between the boom and the slewing platform to the hinge point between the stick and the boom; L2 is the stick length, that is, the distance from the hinge point between the stick and the boom to the hinge point between the bucket and the stick; and L4 is the bucket link length, that is, the distance from the hinge point between the bucket and the stick to the tip of the bucket tooth (end effector). The coordinates of the end effector in the X-axis direction of the base coordinate system correspond to the "forward and backward direction" of the excavator body (forward is the positive direction along the longitudinal axis of the body); The coordinates of the end effector in the Y-axis direction of the base coordinate system correspond to the "left and right directions" of the excavator body (perpendicular to the longitudinal axis of the body, with the right being the positive direction); The coordinates of the end effector in the Z-axis direction of the base coordinate system correspond to the "vertical height direction" (perpendicular to the ground, with upward as the positive direction).

[0055] ;

[0056] The maximum distance between the end effector and the center of the excavating equipment is calculated, which is the maximum operating radius R. R is superimposed with the preset safety threshold ΔR to obtain the danger radius d. A circle is drawn with the excavating equipment as the center and d as the radius to form an alarm fence. Since the attitude parameters are updated in real time with the operation, R changes dynamically, thereby realizing the dynamic delineation of the danger range.

[0057] The danger zone is dynamically and accurately delineated based on the real-time operating status of the excavating equipment. Changes in the posture of the equipment's operating mechanism directly alter its working range; for example, when the boom is raised or the forearm is extended, the working radius increases. The fence, calculated based on real-time posture parameters, can adjust accordingly, preventing missed or incorrect assessments due to the fixed fence's inability to adapt to the equipment's dynamic operation. Simultaneously, a preset safety threshold ΔR is superimposed, providing personnel with safe space for reaction and evacuation. Using a circle as the fence clearly defines the boundary of the danger zone, facilitating rapid subsequent assessment of whether personnel have entered the danger zone, thus providing an accurate regional benchmark for construction safety early warning.

[0058] In another preferred embodiment of the present invention, the specific process of obtaining the personnel identification code in step S2 is as follows:

[0059] Real-time image streams are acquired from a visual acquisition device, face regions are detected through a cascaded neural network, and illumination compensation and filtering preprocessing are performed on the face regions. For the preprocessed face regions, a deep feature extraction network is used to analyze their geometric structure and obtain the corresponding feature vectors.

[0060] The feature vector is compared with a preset facial feature database, and the similarity value with each feature vector in the database is calculated. The feature vector corresponding to the highest similarity value is selected from all similarity values, and the personnel identification code corresponding to the feature vector in the matching feature database is obtained and used as the personnel identification code of the person.

[0061] The excavating equipment acquires continuous image frames of the surrounding area in real time using visual acquisition devices (such as high-definition cameras mounted on the top of the cab), forming a real-time image stream. This is because cameras can capture image information of people in dynamic scenes through continuous shooting. A cascaded neural network is used to detect facial regions. This cascaded neural network consists of multiple network layers working sequentially. First, shallower networks quickly filter out candidate regions that may contain faces. Then, deeper networks make detailed judgments on these regions. Because faces have a fixed layout of facial features (such as the relative positions of the eyes, nose, and mouth), the neural network can accurately distinguish faces from other objects by learning the features of a large number of facial samples, thus locating the facial regions in the image. Illumination compensation and filtering preprocessing are performed on the detected facial regions. For example, when the light is too strong and the face is too bright, the brightness of that area is reduced; when the light is insufficient, the brightness is increased to make the facial details clearer. Filtering preprocessing removes grainy noise from the image, making the facial contours smoother. This is because noise caused by uneven lighting or equipment vibration can interfere with subsequent feature extraction; preprocessing makes the facial image more regular. For the preprocessed facial region, a deep feature extraction network is used to analyze its geometric structure. The deep network extracts deep facial features through multi-layer computation, such as unique geometric parameters like the distance from the corner of the eye to the corner of the mouth and the tilt angle of the bridge of the nose. The feature vector formed by the combination of these parameters is unique and can distinguish different individuals. The generated feature vector is compared with a pre-set personnel facial feature database (which stores the facial feature vectors of all construction workers). Similarity is calculated because the facial feature vectors of the same person are numerically closer; smaller differences indicate higher similarity. The feature vector corresponding to the highest similarity value is selected from all similarity values. This vector corresponds to the matched person, and the personnel identification code associated with this vector in the feature database is then obtained. This is because the highest similarity means that the two belong to the same person.

[0062] Accurate identification of personnel within the construction area lays the foundation for subsequent association with their ultra-wideband positioning tags and acquisition of coordinate data. Through cascaded neural networks and preprocessing, facial regions can be accurately detected and optimized in complex construction environments (such as changing lighting and cluttered backgrounds), improving the reliability of feature extraction. A deep feature extraction network ensures the uniqueness of identification, preventing confusion between different individuals. Comparison with a pre-set feature library and selection of the highest similarity further guarantees the accuracy of identity matching. These steps collectively ensure the accuracy of personnel identification, providing a reliable basis for subsequent identity-based retrieval of positioning data, coordinate correction, and triggering of warnings. Ultimately, this provides the prerequisite for precise monitoring of personnel and hazardous areas in construction safety management, ensuring that warning commands are linked to specific individuals and improving the effectiveness of safety management.

[0063] In another preferred embodiment of the present invention, the process of obtaining the original coordinate data in step S3 is as follows:

[0064] Establish a mapping table that associates personnel identification codes with ultra-wideband positioning tag MAC addresses. The mapping table stores the unique tag MAC address and spatial location signature corresponding to each personnel identification code.

[0065] After obtaining the personnel identification code through visual recognition, the system queries the associated mapping table to retrieve the corresponding ultra-wideband tag MAC address; the ultra-wideband positioning system receives the nanosecond-level pulse signal emitted by the tag through the distributed base station network, calculates the original spatial coordinates of each tag based on the time difference of arrival algorithm, and generates a real-time positioning data stream with timestamp and MAC address;

[0066] Based on the retrieved tag MAC address and combined with the spatial location signature verification mechanism, the original coordinate data of the target tag at the corresponding time is extracted from the real-time location data stream.

[0067] A mapping table is established to associate personnel identification codes with the MAC addresses of ultra-wideband (UWB) positioning tags. For example, each construction worker is assigned a unique identification code (e.g., employee number 001, 002), and each worker's UWB positioning tag has a unique MAC address (similar to a device's ID number). The spatial location signature can be a specific location feature recorded by the tag during initial positioning (e.g., coordinates at the starting point of the construction area). The mapping table records each employee number with its corresponding MAC address and spatial location signature. This is done because the identification code identifies the personnel, and the MAC address identifies the positioning tag; this association allows the system to know which tag corresponds to which person. When a personnel identification code is obtained through visual recognition, for example, if someone's employee number is identified as 001, the mapping table is queried to find the corresponding UWB tag MAC address. Because the mapping table stores the correspondence between the two, the corresponding tag identifier can be directly retrieved. The distributed base station network of an ultra-wideband positioning system (e.g., multiple signal receiving base stations installed around a construction area) receives nanosecond-level pulse signals emitted by tags. It calculates coordinates based on a time-of-arrival (TOA) algorithm. The principle is that the propagation time of the signal from the tag to different base stations varies. Based on the TOA and the signal propagation speed (speed of light), the tag's position in three-dimensional space can be calculated. The generated real-time positioning data stream contains a timestamp (recording the acquisition time) and the corresponding MAC address for each coordinate, thus distinguishing the positions of different tags at different times. Based on the retrieved tag MAC address, combined with a spatial location signature verification mechanism (e.g., comparing the current tag's location characteristics with the initial signature stored in a mapping table), once it is confirmed to be the same tag, the original coordinate data of that MAC address at the corresponding time is extracted from the real-time positioning data stream. This is because the MAC address uniquely identifies the tag, and signature verification avoids tag confusion, ensuring that the extracted coordinates belong to the target person.

[0068] By ensuring precise association between personnel identity and ultra-wideband (UWB) positioning tags, real-time location data of the corresponding personnel can be accurately obtained. A mapping table is established to achieve a one-to-one correspondence between identity and tag, avoiding confusion of location data from different personnel. MAC address lookup via identity code ensures accurate association between personnel identity and positioning tag. The UWB system calculates coordinates based on the time difference of arrival algorithm, generating tagged location data in real time to meet the real-time location information requirements of construction scenarios. Spatial location signature verification further ensures that the extracted coordinates belong to the target personnel, avoiding data errors caused by tag misreading. These steps collectively provide accurate original location data for subsequent attitude-corrected coordinates and identification of hazardous areas, ensuring that safety warnings are accurately linked to specific personnel, improving the accuracy and reliability of construction safety management.

[0069] In another preferred embodiment of the present invention, the specific construction process of the preset attitude-error compensation model is as follows:

[0070] S11. Control the baseline test subject to perform multiple sets of preset posture actions at a fixed spatial position, and simultaneously collect the ultrawideband raw coordinates of the positioning tag terminal and the visual images at the corresponding time to generate a spatiotemporally aligned coordinate-posture dataset.

[0071] S12, Predetermine the reference attitude, and extract the ultrawideband raw coordinate data corresponding to the reference attitude as the reference true coordinates;

[0072] S13. Obtain a visual image of any non-reference pose and generate a non-reference pose feature vector through spatial topology analysis of skeletal joints.

[0073] S14. Calculate the three-dimensional spatial offset between the ultra-wideband original coordinates and the reference true coordinates corresponding to the non-reference attitude, and generate the non-reference attitude positioning error vector.

[0074] S15. The non-reference posture feature vector samples and the positioning error vector samples are aggregated to form a training sample set, and the posture-error compensation model is generated by driving the model training through machine learning algorithms.

[0075] In a fixed spatial location, a baseline subject (either a real person or a bionic model wearing an UWB tag) is controlled to perform multiple sets of preset postures, such as standing, bending over, turning sideways, and raising an arm. Simultaneously, the raw coordinate data of the tag is collected using an UWB positioning system, and visual images of the subject are captured using the visual acquisition device of the excavation equipment. The acquisition times of both are synchronized using timestamps, ensuring a one-to-one temporal correspondence between the coordinates and images for each posture. This generates a coordinate-posture dataset that achieves spatiotemporal alignment, as only data that corresponds in both time and space can accurately reflect the positioning and visual characteristics of the same posture. A baseline posture is pre-determined, such as the subject standing naturally with an upright body, because the tag position is stable in this posture, the UWB signal propagation path is relatively stable, and the positioning error is small. Therefore, the raw UWB coordinate data corresponding to this posture is extracted as the baseline ground truth coordinates to measure the positioning deviation of other postures. After acquiring a visual image of any non-reference pose (such as bending over or turning sideways), the spatial topology of the skeletal joints is analyzed to identify the positions of key joints such as the head, shoulders, hips, and knees. The pitch angle of the torso relative to the vertical direction (the angle of leaning forward or backward), the horizontal yaw angle relative to the front (the angle of left or right rotation), and the relative distances and distribution between joints are calculated. These parameters together constitute the non-reference pose feature vector, which uniquely identifies the pose because the distribution and angles of joints differ across poses. The positional differences between the ultra-wideband original coordinates and the reference ground truth coordinates in three-dimensional space (front / back, left / right, up / down directions) corresponding to the non-reference pose are calculated. These differences form the non-reference pose positioning error vector, as this vector directly reflects the degree of positioning offset caused by the pose. Feature vector samples from multiple different non-reference poses and their corresponding positioning error vector samples are aggregated to form a training sample set. Machine learning algorithms are used to teach the model the correlation between pose features and positioning errors. For example, the model can automatically output the corresponding error offset based on input features such as the torso pitch angle, ultimately generating a pose-error compensation model.

[0076] It is worth noting the specific implementation steps of the model learning association patterns and generation:

[0077] In laboratory or construction simulation scenarios, the baseline test subject (a real person / bionic model wearing a UWB tag) is fixed in a spatial position with known precise coordinates (excluding the influence of position movement on positioning), and only the subject is controlled to perform preset posture actions (such as standing, bending over 15° / 30°, turning to the side 45°, lowering the head, etc., which cover common construction postures).

[0078] The visual acquisition device of the excavation equipment captures visual frames of each posture in real time, and the original coordinates at the corresponding time are collected by the UWB positioning system. All data are timestamped to ensure strict spatiotemporal alignment of "visual frame (posture) - original coordinates".

[0079] The UWB original coordinates corresponding to the baseline posture of "standing naturally with the tag antenna pointing vertically upward" are used as the baseline ground truth coordinates (the antenna radiation is stable and the positioning error is minimal under this posture). For any non-baseline posture, the posture feature vector (including the spatial coordinate distribution parameters of joints such as torso pitch angle α, horizontal yaw angle β, and shoulder, elbow, hip, and knee joints) is extracted through the spatial topology analysis of the skeletal joints. At the same time, the three-dimensional offset (Δx, Δy, Δz) between the UWB original coordinates and the baseline ground truth coordinates under this posture is calculated as the positioning error vector.

[0080] The final training sample set consists of "input (attitude feature vector) - output (positioning error vector)". The sample size needs to cover typical postures in the construction scenario (e.g., 20-30 postures, with each posture collected 10-20 times to enhance data reliability).

[0081] Prioritize algorithms that can fit nonlinear correlations, use gradient boosting trees as the core of model training, divide the sample set into training and validation sets in a 7:3 ratio, and build a training model.

[0082] Mean squared error (MSE) is used as the loss function;

[0083] ;

[0084] Where (Δxi,Δyi,Δzi) is the true offset corresponding to the i-th training sample, (Δxi',Δyi',Δzi') is the model prediction error corresponding to the i-th training sample, and N is the total number of training samples;

[0085] When the MSE of both the training set and the validation set converges to a preset threshold (e.g., ≤0.01m², corresponding to a positioning error ≤10cm, which meets the accuracy requirements for construction early warning), training stops, and the attitude-error compensation model with the correlation law of "attitude features → positioning error" is obtained.

[0086] A model is constructed to accurately predict positioning errors based on personnel posture characteristics, thereby correcting ultra-wideband positioning deviations caused by posture changes. A fixed spatial position ensures that positioning errors are caused solely by posture changes, eliminating interference from positional variations. A stable baseline posture is selected as the ground truth, providing a reliable reference for error calculation. Analyzing skeletal joint features accurately quantifies different postures, allowing the model to identify posture differences. Calculating the error vector establishes a direct correlation between posture and positioning deviation. Training the model enables the prediction capability from posture features to errors. This series of steps ultimately allows the model to effectively compensate for positioning errors caused by posture interference, making the corrected coordinates closer to the actual position of the personnel. This provides precise location data for accurately determining whether personnel have entered dangerous areas and triggering correct warnings, improving the accuracy of construction safety management.

[0087] In another preferred embodiment of the present invention, the method further includes pre-acquiring the same visual images at different resolutions, and labeling each pre-acquired image as Y1, Y2, ..., Yn; where n is the total number of pre-acquired images. The method generates pose feature vectors for each pre-acquired image in sequence and records the generation time each time. The corresponding data points are generated in the coordinate system with the resolution as the abscissa and the generation time as the ordinate. The data points are then fitted to a curve using a polynomial fitting method to obtain the fitting formula T(F).

[0088] First, identical visual images are acquired, such as images of a person maintaining a fixed posture within a construction area, ensuring complete consistency in image content with only resolution differences. Then, these images are pre-acquired at different resolutions. For example, by adjusting the parameters of the visual acquisition device, images of low, medium, and high resolutions are acquired and labeled Y1, Y2, ..., Yn. This labeling distinguishes images of different resolutions for subsequent processing. Next, posture feature vectors are generated for each pre-acquired image. Specifically, for Y1, the trunk pitch angle, horizontal yaw angle, and limb distribution parameters of the skeletal joints are analyzed to generate corresponding feature vectors. This process is repeated for Y2, Y3, and so on. The time spent generating feature vectors each time is recorded; for example, processing Y1 takes t1, and processing Y2 takes t2. This is because images of different resolutions contain different pixel information; higher resolution images contain more pixels, resulting in more complex information to process and longer feature vector generation time. Then, using resolution as the x-axis (e.g., mapping different resolution values ​​to the x-axis) and generation time as the y-axis (mapping t1, t2, etc. to the y-axis), mark the data points for each resolution and corresponding generation time in the coordinate system. Since there is a certain correlation between resolution and generation time, such as the time generally increasing with increasing resolution, a polynomial fitting method can be used to fit these data points to a curve. That is, find a polynomial function that makes the function curve as close as possible to all data points, and finally obtain the fitting formula T(F). This formula can reflect the correspondence between resolution F and generation time T.

[0089] The specific steps for setting the fitting formula T(F);

[0090] First, fix the target in the construction scene (such as a person maintaining a fixed posture). Then, using the visual acquisition device (such as a high-definition camera) on the excavation equipment, under the same lighting, distance, and viewing angle conditions, only adjust the acquisition parameters (such as pixel size: 640×480, 1280×720, 1920×1080, etc.) to pre-acquire n images with different resolutions but completely identical content, and label them as (Y1,Y2,...,Yn).

[0091] For each pre-acquired image (Yc), a posture feature extraction process consistent with actual working conditions is executed (i.e., generating feature vectors containing parameters such as torso pitch angle and horizontal yaw angle through spatial topology analysis of skeletal joints). A timer is used to accurately record the complete time from image input to feature vector output, which is taken as the generation time (Tc) corresponding to that resolution. At the same time, the resolution F is quantized into a calculable value (such as being represented by the number of horizontal pixels, the number of vertical pixels, or the total area of ​​pixels, such as F=1920×1080 which can be simplified to F=2073600 pixels).

[0092] With resolution F as the abscissa and generation time T as the ordinate, the data points are marked as ((Fc,Tc)) in a two-dimensional coordinate system. Since resolution and generation time usually have a non-linear monotonically increasing relationship (the increase in the number of pixels and the computation time are approximately polynomially related), the data points are curve-fitted using a polynomial fitting method commonly used in this field (such as the least squares method), and finally the fitting formula T(F) is obtained.

[0093] When the same person is simultaneously identified by the vision devices of multiple excavation equipment, the resolution (F1, F2, ..., Fk) of the images acquired by each device is extracted, and the corresponding predicted generation time (T1, T2, ..., Tk) is calculated by substituting them into the fitting formula T(F). The image with the smallest T is directly selected as the pose visual frame. This process does not require actual feature extraction of multiple images. It only quickly predicts the time consumption through the formula, avoiding redundant calculation of high-resolution but inefficient images. While ensuring the accuracy of pose feature extraction (since the image content is consistent, the resolution only needs to meet the requirements of skeleton analysis), the calculation delay is shortened to the minimum, and the efficiency of feature extraction is achieved.

[0094] Establishing a correlation between image resolution and pose feature vector generation time enables rapid assessment of processing efficiency for images at different resolutions in real-world construction scenarios. Since the same person may be simultaneously identified by the visual acquisition devices of multiple excavating machines, and the images acquired by different devices have varying resolutions and processing times, the fitting formula T(F) allows direct determination of the pose feature vector generation time based on resolution. This enables the selection of the image with the shortest generation time, reducing the time spent on pose feature extraction, ensuring real-time coordinate correction and early warning assessment, and preventing delayed safety warnings due to processing delays. Ultimately, this provides efficient and timely technical support for construction safety management, ensuring rapid response to personnel entering hazardous areas.

[0095] In another preferred embodiment of the present invention, in step S4, if the same person is simultaneously identified by the visual acquisition devices of multiple excavation equipment, the corresponding resolution of the person in the images acquired by each equipment is obtained and substituted into the fitting formula T(F) to obtain the generation time corresponding to the images acquired by each equipment, and the acquisition image with the shortest generation time is selected as the person's posture visual frame.

[0096] When the same person is identified by multiple devices, the image with the highest processing efficiency is selected for pose feature extraction, ensuring the timeliness of pose feature vector generation. Because personnel and equipment are constantly changing in construction scenarios, real-time performance is crucial for safety warnings. Selecting images with long generation times may delay pose feature extraction, affecting the timeliness of subsequent coordinate correction and warning judgment. By selecting the image with the shortest generation time, both the efficiency of pose feature extraction and timely data support for subsequent coordinate correction based on feature vectors and determining whether a warning has been triggered are ensured. Ultimately, this guarantees the speed and accuracy of warning responses in construction safety management, avoiding safety hazards caused by processing delays.

[0097] In another preferred embodiment of the present invention, the specific process of triggering the graded early warning command in step S6 is as follows:

[0098] Within the danger radius, several distance levels are set at preset unit distance intervals. The smaller the distance, the higher the warning level. At the same time, different warning levels correspond to different warning information.

[0099] Within a defined danger radius, several distance levels are set according to preset unit distance intervals. For example, the danger radius covers the area from the center of the excavating equipment to a certain furthest distance. The preset unit distance interval can be a fixed length. Starting from the boundary of the danger radius, a level is defined every time this length is extended inward, thus dividing the area into first level, second level, and so on, up to a certain number of levels. Because the closer to the excavating equipment, the higher the degree of danger faced by personnel, classifying levels according to distance intervals can clearly distinguish different levels of danger. The smaller the distance, that is, the closer the personnel's corrected coordinates are to the center of the excavating equipment, the higher the corresponding warning level. For example, the level closest to the equipment is set to the highest level, and the level slightly farther away is set to the second highest level, and so on. This is because closer distance means a greater possibility of danger and potentially more serious consequences, requiring a stronger warning. At the same time, different warning levels will correspond to different warning messages. For example, the highest level warning may be a continuous audible and visual alarm sent to the cab of the excavator and the on-site safety officer will be notified at the same time. The next highest level warning may be a flashing light in the cab and a voice prompt. The lower level warning may just be a slight prompt sound. Different levels of danger require different response methods so that relevant personnel can quickly understand the urgency of the danger and take corresponding measures.

[0100] The intensity and content of warnings are dynamically adjusted based on the actual distance between personnel and excavation equipment, ensuring that the warning information more closely reflects the actual hazardous situation. By setting multiple distance levels, the severity of danger can be differentiated in detail, avoiding the ambiguity caused by using the same warning method regardless of distance. For example, the warnings received when personnel are only slightly approaching the danger zone differ from those received when they are very close. This allows relevant personnel to take different countermeasures accordingly, preventing overreactions that could disrupt construction in low-risk situations and insufficient warnings that could lead to accidents in high-risk situations. This series of practices makes construction safety warnings more accurate and effective, ultimately helping to achieve refined management of personnel safety within the construction area and improving the overall reliability of construction safety management.

[0101] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A construction safety management method based on image fusion of a view angle of a excavator, characterized in that, The method comprises the following steps: S1, obtaining the geographic position and working condition of all excavating equipment in the construction area, and determining the corresponding warning fence outside any excavating equipment according to the working condition; S2, triggering the visual acquisition device carried by any excavating equipment and performing real-time scanning on the surrounding area, identifying the personnel boundary box through a target detection algorithm, and extracting a personnel identity code; S3, calling the original coordinate data of the corresponding ultra-wideband positioning tag according to the personnel identity code, and synchronously capturing the posture visual frame of the personnel; Pre-acquire the same visual image and pre-acquire at different resolutions, and mark each pre-acquired image as Y1, Y2,..., Yn; wherein n is the total number of pre-acquired images, and the posture feature vector is generated for each pre-acquired image in turn and the generation time of each generation is recorded, taking the resolution as the horizontal coordinate and the generation time as the vertical coordinate, generating corresponding data points in the coordinate system, and using a polynomial fitting method to perform curve fitting on the data points to obtain a fitting formula T(F); If the same personnel is recognized by the visual acquisition devices of multiple excavating equipment at the same time, the corresponding resolution of the personnel in the image collected by each device is obtained and substituted into the fitting formula T(F) to obtain the generation time corresponding to the image collected by each device, and the image collected with the shortest generation time is selected as the posture visual frame of the personnel; S4, generating a corresponding posture feature vector based on the posture visual frame through spatial topology analysis of the skeleton joints, wherein the posture feature vector includes a trunk pitch angle, a horizontal yaw angle, and a limb joint spatial distribution parameter; S5, inputting the posture feature vector into a pre-set posture-error compensation model to obtain a positioning coordinate offset prediction value, correcting the original coordinates of the personnel according to the positioning coordinate offset prediction value to obtain corrected coordinates; S6, judging the spatial topological relationship between the corrected coordinates and the boundary of the warning fence, and triggering a hierarchical early warning instruction when the corrected coordinate point falls within the fence buffer area; The specific construction process of the pre-set posture-error compensation model is as follows: S11, controlling a reference subject to perform a plurality of pre-set posture actions at a fixed spatial position, synchronously collecting the ultra-wideband original coordinates of the positioning tag terminal and the visual image at the corresponding time, and generating a spatio-temporally aligned coordinate-posture data set; S12, pre-determining a reference posture, and extracting the ultra-wideband original coordinate data corresponding to the reference posture as a reference true value coordinate; S13, obtaining a visual image of any non-reference posture, and generating a non-reference posture feature vector through spatial topology analysis of the skeleton joints; S14, calculating the three-dimensional spatial offset between the ultra-wideband original coordinates corresponding to the non-reference posture and the reference true value coordinates to generate a non-reference posture positioning error vector; S15, aggregating the non-reference posture feature vector sample and the positioning error vector sample to form a training sample set, and generating a posture-error compensation model through model training driven by a machine learning algorithm.

2. The construction safety management method based on the fusion of the view angle image of the excavator according to claim 1, characterized in that, In S1, the specific determination process of the warning fence is as follows: The posture parameters of the working mechanism are collected in real time by a multi-dimensional sensor group installed on the excavating equipment, the posture parameters including a large arm pitch angle α, a small arm swing angle β, a bucket roll angle γ and a three-dimensional coordinate (x, y, z) of an end effector, a kinematics model of the working mechanism is established based on an inverse kinematics algorithm, and a maximum working radius R of the excavating equipment is calculated, the maximum working radius R is superimposed with a preset safety threshold ΔR to obtain a dangerous radius d, and the corresponding warning fence is obtained by taking the excavating equipment as the center and drawing a circle with the dangerous radius d as the radius. 3.The construction safety management method based on the fusion of the view angle video of the excavator according to claim 1, characterized in that, In the S2, the specific process of obtaining the personnel identity code is as follows: An image stream is obtained from a visual acquisition device, a face region is detected by a cascade neural network, and illumination compensation and filtering preprocessing are performed on the face region, a deep feature extraction network is used to analyze the geometric structure of the preprocessed face region to obtain a corresponding feature vector; The feature vector is compared with a preset personnel face feature library, the similarity values of each feature vector in the library are calculated, the feature vector corresponding to the highest value is selected from all similarity values, and the personnel identity code corresponding to the feature vector in the matching feature library is obtained as the personnel identity code of the personnel.

4. The construction safety management method based on the fusion of the view angle image of the excavator according to claim 1, characterized in that, In the S3, the process of obtaining the original coordinate data is as follows: An association mapping table of the personnel identity code and the ultra-wideband positioning tag MAC address is established, the mapping table stores the unique tag MAC address and the spatial position signature corresponding to each personnel identity code; After the personnel identity code is obtained through visual recognition, the corresponding ultra-wideband tag MAC address is searched in the association mapping table, the ultra-wideband positioning system receives the nanosecond-level pulse signals transmitted by the tags through a distributed base station network, calculates the original spatial coordinates of each tag based on a time difference of arrival algorithm, and generates a real-time positioning data stream with a time stamp and a MAC address; According to the retrieved tag MAC address, the original coordinate data of the target tag at the corresponding time are extracted from the real-time positioning data stream in combination with the spatial position signature verification mechanism. 5.The construction safety management method based on the fusion of the view angle video of the excavator according to claim 2, characterized in that, In the S6, the specific process of triggering the hierarchical early warning instruction is as follows: A plurality of distance levels are set at preset unit distance intervals within the dangerous radius, the smaller the distance, the higher the early warning level, and different early warning levels correspond to different early warning information.

Citation Information

Patent Citations

  • Early warning method and system for collision accident between excavator and person based on computer vision

    CN116469037A

  • Substation personnel safety operation early warning system and method fusing UWB and imaging technology

    CN120108104A