Image recognition-based remote identification method and system for highway construction safety behavior

CN122618569BActive Publication Date: 2026-09-29CHENGDU JIAXIN TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202611105097.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-24
Publication Date
2026-09-29
Estimated Expiration
2046-07-24

AI Technical Summary

Technical Problem

[0004]本发明解决的技术问题是:现有技术在进行目标跟踪关联时,主要依赖预测框与检测框的交并比及中心坐标距离,公路施工现场工人常伴随大幅度姿态变化,且极易被施工机械或材料部分遮挡,在这种情况下,工人的视觉轮廓会发生剧烈形变或缺失,单纯依赖质心距离或矩形框重叠率极易导致目标身份频繁跳变或跟踪丢失,无法维持长时间的稳定监测,现有技术缺乏对环境不确定性的量化与自适应机制,现有技术缺乏一种机制来实时评估观测数据的可信度,也无法根据这种不确定性动态调整预测范围,导致在传感器噪声较大时,系统往往给出过于自信的错误判断,现有技术难以利用历史先验信息来抑制突发的异常检测值

Benefits of technology

[0067]本发明的有益效果:本发明通过构建一种融合图像与点云特征的远程识别架构,显著提升了公路施工复杂场景下安全行为监测的鲁棒性与准确性,其核心优势在于引入了基于欧氏距离与单向豪斯多夫距离的双重关联评价机制,并利用变异系数动态分配权重,本发明通过豪斯多夫距离捕捉轮廓边缘的细微形变,解决了仅仅依赖质心距离导致在工人姿态变化或部分遮挡时匹配失效的问题,确保了多源异构数据在时空上的精准对齐与关联,本发明建立了一套包含图像估计、点云估计及关联维度的三元不确定性耦合评估机制,该机制能够实时量化感知数据的可靠程度,并据此驱动自适应的轨迹预测机制,当系统检测到关联置信度降低或存在遮挡时,会自动调整弥散系数以扩大位置预测的概率分布范围,并结合历史遮挡影响系数修正惯性预测权重,本发明通过计算行为先验分布与均匀分布的加权混合,在高度不确定的情况下能够自动抑制激进的违规判定,体现了算法层面的安全冗余,基于后验分布质心与危险区域的耦合关系,能够输出高风险违规、低风险违规与安全的差异化结果,这不仅大幅降低了因环境因素导致的误报率,还为施工现场管理者提供了分级预警能力,显著提升了安全监管的智能化水平与响应效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122618569B_ABST
    Figure CN122618569B_ABST
Patent Text Reader

Abstract

The application discloses a road construction safety behavior remote identification method and system based on image recognition, relates to the technical field of image recognition, and comprises the following steps: preprocessing, aligning and feature extraction are performed on image data and point cloud data to obtain image local estimation results and point cloud local estimation results; a fusion feature value based on Euclidean distance and one-way Hausdorff distance is constructed to obtain the correlation degree of the image local estimation results and the point cloud local estimation results; based on the image local estimation results, the point cloud local estimation results and the correlation degree, three-uncertainty is calculated; based on the three-uncertainty, the correlation degree, the image local estimation results and the point cloud local estimation results, position distribution and behavior distribution are constructed to perform position and behavior prediction, the predicted results, the image local estimation results and the point cloud local estimation results are weighted and fused, and a safety behavior recognition result is output according to the fusion result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition technology, and in particular to a remote identification method and system for highway construction safety behavior based on image recognition. Background Technology

[0002] With the rapid advancement of transportation infrastructure construction, safety management at highway construction sites has become increasingly important. Due to the complex environment of construction sites, the mixing of personnel and machinery, and the presence of interference factors such as dust and changes in lighting, traditional supervision methods relying on manual inspections are inefficient and have blind spots. In recent years, automated monitoring technologies based on computer vision and deep learning have gradually become a research hotspot, aiming to achieve real-time identification and early warning of unsafe behaviors of construction workers.

[0003] Among the existing related technologies, Chinese Patent Publication No. CN120808442A discloses a method for detecting unsafe behaviors of construction workers based on an improved YOLO. This method mainly enhances the model's attention to key information and small target detection capabilities in complex environments by introducing channel attention, spatial attention, and bidirectional feature pyramid networks into the YOLOv3 model. Existing technologies, by improving visual detection models and introducing multi-target tracking, have to some extent solved the problems of false detection and missed detection in construction worker behavior recognition in general scenarios, and can assess potential risks through trajectory analysis. When performing target tracking and association, existing technologies mainly rely on the intersection-union ratio of predicted boxes and detection boxes and the distance between their center coordinates. Workers at highway construction sites often have significant posture changes and are easily partially obscured by construction machinery or materials. In this case, the visual contour of the worker will be drastically deformed or missing. Simply relying on the centroid distance or the overlap rate of the rectangular boxes can easily lead to frequent jumps in target identity or tracking loss, making it impossible to maintain stable monitoring for a long time. Existing technologies lack quantification and adaptive mechanisms for environmental uncertainties. Existing technologies lack a mechanism to evaluate the credibility of observation data in real time and cannot dynamically adjust the prediction range according to this uncertainty. As a result, when sensor noise is high, the system often makes overconfident erroneous judgments. Existing technologies are unable to use historical prior information to suppress sudden abnormal detection values. Summary of the Invention

[0004] The technical problem solved by this invention is that existing technologies mainly rely on the intersection-union ratio of predicted boxes and detection boxes and the distance between their center coordinates when performing target tracking and association. However, workers at highway construction sites often have significant posture changes and are easily partially obscured by construction machinery or materials. In such cases, the visual outline of the worker will be drastically deformed or missing. Simply relying on the centroid distance or the overlap rate of the rectangular boxes can easily lead to frequent changes in the target's identity or loss of tracking, making it impossible to maintain stable monitoring over a long period of time. Existing technologies lack quantification and adaptive mechanisms for environmental uncertainties. They also lack a mechanism to evaluate the reliability of observation data in real time and cannot dynamically adjust the prediction range based on such uncertainties. This results in the system often making overconfident and erroneous judgments when sensor noise is high. Existing technologies also struggle to use historical prior information to suppress sudden abnormal detection values.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a remote identification method for highway construction safety behaviors based on image recognition, comprising the following steps:

[0006] Step S1: Preprocess, align, and extract features from the image data and point cloud data to obtain local image estimation results and local point cloud estimation results;

[0007] Step S2: By constructing fused feature values ​​based on Euclidean distance and one-way Hausdorff distance, the correlation between the local estimation results of the image and the local estimation results of the point cloud is obtained;

[0008] Step S3: Calculate the ternary uncertainty based on the local image estimation results, the local point cloud estimation results, and the correlation degree;

[0009] Step S4: Based on the ternary uncertainty, correlation degree, image local estimation results and point cloud local estimation results, construct position distribution and behavior distribution to predict position and behavior, perform weighted fusion of prediction results, image local estimation results and point cloud local estimation results, and output safety behavior recognition results based on the fusion results.

[0010] Preferably, the preprocessing of the image data includes denoising and normalization of each frame of the consecutive frame images, and the preprocessing of the point cloud data includes outlier removal and density unification;

[0011] The feature extraction includes identifying worker behavior features from the aligned image sequence using a large model and outputting local image estimation results, which include worker ID, behavior type, behavior contour coordinates, occlusion type, and image estimation confidence.

[0012] The feature extraction also includes identifying worker behavior features from the pixel sequence of the mapped point cloud using a density clustering algorithm, and outputting a local point cloud estimation result, which includes worker ID, position coordinates, velocity vector, and position estimation variance.

[0013] Based on the worker ID, the local image estimation results and the local point cloud estimation results of the same worker are initially correlated to form an initial local state set.

[0014] Preferably, obtaining the fusion feature value specifically includes:

[0015] Calculate the Euclidean distance between the center coordinates of the behavioral contour coordinates in the local estimation result of the image and the position coordinates in the local estimation result of the point cloud, wherein the center coordinates of the contour coordinates are the arithmetic mean of the vertex coordinates of the behavioral contour polygon;

[0016] Calculate the one-way Hausdorff distance from all behavioral contour coordinates to the preset pixel neighborhood of the position coordinates;

[0017] The method for calculating the fusion feature value includes:

[0018] Divide the Euclidean distance of the current frame by the mean of the Euclidean distances of the previous N frames, and divide the one-way Hausdorff distance by the mean of the one-way Hausdorff distances of the previous N frames. Then multiply each by the distance feature contribution weight and sum them up. The Euclidean distance contribution weight is the correlation coefficient of the Euclidean distance in the effective associated samples divided by the sum of the correlation coefficients of the Euclidean distance and the one-way Hausdorff distance. The one-way Hausdorff distance contribution weight is 1 minus the Euclidean distance contribution weight.

[0019] Preferably, obtaining the correlation degree specifically includes:

[0020] If the fusion feature value is less than the preset effective association threshold, the association degree is 1 minus the ratio of the fusion feature value to the mean of the fusion feature values ​​of the previous N frames plus 3 times the standard deviation. If the fusion feature value is greater than or equal to the preset effective association threshold, the association degree is 0.3 multiplied by the negative fusion feature value of the natural index minus the ratio of the effective association threshold to the standard deviation of the effective association samples raised to the power of 3.

[0021] Preferably, the ternary uncertainty is obtained by image estimation confidence, location estimation variance, and correlation degree, wherein the ternary uncertainty includes image estimation uncertainty, point cloud estimation uncertainty, and correlation uncertainty;

[0022] Image estimation uncertainty is calculated based on image estimation confidence and correlation. The specific calculation process includes multiplying the difference between 1 and image estimation confidence and the difference between 1 and correlation.

[0023] The uncertainty of point cloud estimation is calculated based on the variance of location estimation and the degree of correlation. The specific calculation process includes multiplying the variance of location by 1 and adding 0.5 times the difference between 1 and the degree of correlation.

[0024] The uncertainty of association is calculated based on the degree of association. The specific calculation process includes the difference between 1 and the degree of association.

[0025] The product of image estimation uncertainty and association uncertainty is taken as image association coupling uncertainty, and the product of point cloud estimation uncertainty and association uncertainty is taken as point cloud association coupling uncertainty.

[0026] Preferably, step S4 includes the following sub-steps:

[0027] S041: Based on the point cloud position coordinates of the previous N frames, calculate the inertial prediction position and the historical weighted average prediction position respectively. Calculate a historical behavior weight based on historical association reliability and historical occlusion influence coefficient. Fuse the inertial prediction position and the historical weighted prediction position to obtain a final predicted position coordinate.

[0028] Calculating the inertial predicted position involves subtracting the position coordinates of the frame before the frame from twice the position coordinates of the previous frame.

[0029] Calculating the historical predicted position involves linearly weighting the position coordinates of each frame within the previous preset number of frames;

[0030] Calculating the historical correlation reliability includes taking the average correlation between the local image estimation results and the local point cloud estimation results of each frame up to a preset number of frames before the current frame;

[0031] Calculating the historical occlusion impact coefficient includes: statistically analyzing the occlusion types of the previous N frames, numerically mapping the occlusion types, and recursively calculating the historical occlusion impact coefficient using an exponential moving average algorithm.

[0032] The historical behavior weight is calculated by multiplying the historical association reliability by 1 minus the historical occlusion impact coefficient;

[0033] Calculating the final predicted position involves multiplying the historical behavior weight by the historical predicted position, adding 1 and subtracting the historical behavior weight multiplied by the inertial predicted position to obtain the final predicted position coordinates.

[0034] S042: Calculate the standard deviation of the position coordinates of the previous N frames, take half of the standard deviation as the base standard deviation, and take the square of the base standard deviation as the base variance;

[0035] The diffusion coefficient is calculated using the following mathematical expression:

[0036] ;

[0037] in, The dispersion coefficient is... For relevance, The mean of the fused feature values ​​of the first N frames. The standard deviation of the fused feature values ​​from the first N frames;

[0038] The initial variance is obtained by adjusting the base variance using the dispersion coefficient. The adjustment includes multiplying the squared value of the dispersion coefficient by the base variance.

[0039] The initial variance is corrected to obtain the corrected variance. The correction includes multiplying the initial variance by 1 and summing the uncertainty of the coupling with the point cloud.

[0040] Centered on the final predicted position, the predicted position range is defined by the area with the radius determined by half the standard deviation of the position coordinates of the previous N frames and the diffusion coefficient. Each pixel within this predicted position range is traversed, and the probability value of each pixel is calculated using a two-dimensional Gaussian probability density function with the final predicted position as the mean and the corrected variance as the covariance. The probability value is then normalized to obtain the prior position distribution.

[0041] Preferably, step S4 further includes the following sub-steps:

[0042] S043: Statistically analyze the number of times safe and illegal behaviors occurred in the historical data of the previous N frames of the current frame, calculate their proportion, and obtain the initial behavior distribution;

[0043] S044: The initial behavior distribution is modified by weighting and mixing the initial behavior distribution with the uniform distribution [0.5, 0.5] to obtain the behavior prior distribution;

[0044] The mathematical expression for weighted mixing is:

[0045] ;

[0046] in, For the prior distribution of behavior, For image correlation coupling uncertainty, This represents the initial behavior distribution.

[0047] Preferably, step S4 further includes:

[0048] The probability density value at the location coordinates is extracted based on the prior distribution of the location, and the membership degree is calculated based on the probability density value, the prior distribution of behavior, and the degree of association.

[0049] Calculating the membership degree involves multiplying the probability density value at the location coordinates, the probability value of the prior distribution of behavior, and the degree of association, and then dividing the result by the total number of targets.

[0050] The target sum is the sum of the products of the prior probability density value of the location, the probability corresponding to the behavior type in the prior distribution of behavior, and the correlation degree of each worker identified in the current frame.

[0051] The common measurement influence coefficient is equal to 1 and the sum of the uncertainty of point cloud association coupling and the uncertainty of image association coupling;

[0052] The product of the common measurement influence coefficient and membership degree is used as the corrected association probability;

[0053] The location prior distribution is taken as the first Gaussian distribution;

[0054] Construct a second Gaussian distribution, where the mean of the second Gaussian distribution is the center coordinate of the behavior contour coordinates, and the covariance is a diagonal matrix with the image association uncertainty as the diagonal element;

[0055] Construct a third Gaussian distribution, where the mean of the third Gaussian distribution is the location coordinates and the covariance is a diagonal matrix with the point cloud association uncertainty as the diagonal element;

[0056] The generalized covariance interaction algorithm is adopted to use the corrected association probability as the weight of the second Gaussian distribution and the third Gaussian distribution, and to perform weighted fusion with the location prior distribution to obtain the location posterior distribution.

[0057] Preferably, the output of security behavior identification results based on the fusion results includes:

[0058] The weighted centroid of the posterior distribution of the location is extracted as the final mean of the location, and the probability value corresponding to the violation in the prior distribution of the behavior is extracted.

[0059] The ray method is used to determine whether the final average position is located in the danger zone. If the average position is located in the danger warning zone and the probability of violation is greater than the preset safety threshold, it is judged as a high-risk violation.

[0060] If the average location is in the danger warning zone or the probability of violation is greater than the preset safety threshold, it is judged as a low-risk violation.

[0061] If the average location is not in a dangerous area and the probability of violation is less than or equal to the preset safety threshold, then it is considered safe.

[0062] A remote identification system for highway construction safety behaviors based on image recognition includes a data acquisition module, an association module, an estimation module, and an identification module.

[0063] The acquisition module is used to preprocess, align, and extract features from image data and point cloud data to obtain local image estimation results and local point cloud estimation results.

[0064] The correlation module is used to obtain the correlation between the local image estimation result and the local point cloud estimation result by constructing a fusion feature value based on Euclidean distance and one-way Hausdorff distance.

[0065] The estimation module is used to calculate the ternary uncertainty based on the local estimation results of the image, the local estimation results of the point cloud, and the correlation degree.

[0066] The recognition module is used to construct position distribution and behavior distribution based on the ternary uncertainty, correlation degree, image local estimation result and point cloud local estimation result, to predict position and behavior, to perform weighted fusion of prediction result, image local estimation result and point cloud local estimation result, and to output safe behavior recognition result based on fusion result.

[0067] The beneficial effects of this invention are as follows: By constructing a remote recognition architecture that integrates image and point cloud features, this invention significantly improves the robustness and accuracy of safety behavior monitoring in complex highway construction scenarios. Its core advantage lies in introducing a dual association evaluation mechanism based on Euclidean distance and one-way Hausdorff distance, and dynamically allocating weights using the coefficient of variation. This invention captures subtle deformations of contour edges through Hausdorff distance, solving the problem of matching failure when relying solely on centroid distance due to changes in worker posture or partial occlusion. This ensures accurate spatiotemporal alignment and association of multi-source heterogeneous data. Furthermore, this invention establishes a ternary uncertainty coupling evaluation mechanism that includes image estimation, point cloud estimation, and association dimensions. This mechanism can quantify the reliability of perceived data in real time. This invention calculates the weighted mixture of prior distribution and uniform distribution of behavior, and can automatically suppress aggressive violation judgments under highly uncertain conditions, demonstrating safety redundancy at the algorithm level. Based on the coupling relationship between the centroid of the posterior distribution and the dangerous area, it can output differentiated results for high-risk violations, low-risk violations, and safety. This not only significantly reduces the false alarm rate caused by environmental factors, but also provides construction site managers with graded early warning capabilities, significantly improving the intelligence level and response efficiency of safety supervision. Attached Figure Description

[0068] Figure 1 This is a schematic diagram of a remote identification method for highway construction safety behavior based on image recognition, provided as an embodiment of the present invention. Detailed Implementation

[0069] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0070] Reference Figure 1 As an embodiment of the present invention, a remote identification method for highway construction safety behavior based on image recognition includes the following steps:

[0071] Step S1: Preprocess, align, and extract features from the image data and point cloud data to obtain local image estimation results and local point cloud estimation results.

[0072] Step S2: By constructing fused feature values ​​based on Euclidean distance and one-way Hausdorff distance, the correlation between the local estimation results of the image and the local estimation results of the point cloud is obtained;

[0073] Step S3: Calculate the ternary uncertainty based on the local estimation results of the image, the local estimation results of the point cloud, and the correlation degree;

[0074] Step S4: Based on ternary uncertainty, correlation degree, local image estimation results and local point cloud estimation results, construct location distribution and behavior distribution to predict location and behavior. Perform weighted fusion on the prediction results, local image estimation results and local point cloud estimation results, and output the safety behavior recognition results according to the fusion results.

[0075] This invention innovatively introduces a ternary uncertainty calculation and weighted fusion mechanism by constructing a fusion architecture of local image estimation and local point cloud estimation. Unlike the traditional method of directly superimposing sensor data, this method can dynamically evaluate the reliability of the image and point cloud according to environmental changes (such as illumination or occlusion) and adjust their weights in the final decision based on this. This effectively solves the problem of single sensor failure in complex highway construction scenarios and significantly improves the robustness and accuracy of remote identification.

[0076] Image data preprocessing includes denoising and normalization of each frame of a continuous image, while point cloud data preprocessing includes outlier removal and density unification.

[0077] Feature extraction includes identifying worker behavior features from aligned image sequences using a large model and outputting local image estimation results, which include worker ID, behavior type, behavior contour coordinates, occlusion type, and image estimation confidence.

[0078] Feature extraction also includes identifying worker behavior features from the pixel sequence of the mapped point cloud using a density clustering algorithm, and outputting local point cloud estimation results, which include worker ID, position coordinates, velocity vector, and position estimation variance.

[0079] The position coordinates are two-dimensional coordinates mapped to the image pixel coordinate system;

[0080] Based on the worker ID, the local image estimation results and the local point cloud estimation results of the same worker are initially correlated to form an initial local state set.

[0081] In one specific embodiment of the present invention, the image data is collected by monitoring equipment deployed in the construction area to collect continuous frame images, covering the worker's work area, the machinery operation area and the danger warning area, including the foundation pit and the high-voltage line area; the point cloud data is ternary point cloud data collected by point cloud equipment that is consistent with the image monitoring range, and the point cloud density is greater than 100 points / ㎡;

[0082] Image data preprocessing involves Gaussian filtering and grayscale normalization of consecutive frame images to obtain a preprocessed image sequence. Point cloud data preprocessing involves statistical filtering of the original point cloud data to remove outliers, eliminating isolated points caused by dust and light interference, and unifying the point cloud density to a preset point cloud threshold of 50 points / m² to reduce computational load and obtain a preprocessed point cloud sequence.

[0083] Time alignment involves aligning the preprocessed image sequence with the preprocessed point cloud sequence at the corresponding time using the timestamp of the data acquisition, ensuring that multiple source data correspond to the same construction scenario state at the same time.

[0084] Spatial alignment involves establishing a mapping relationship from the ternary point cloud coordinate system to the two-dimensional image pixel coordinate system. First, the ternary coordinates of each point in the preprocessed point cloud data are transformed to the camera coordinate system using a rotation matrix R and a preset translation matrix T, resulting in the corresponding point in the camera coordinate system. The coordinates of the corresponding point in the camera coordinate system satisfy the mathematical expression:

[0085] ;

[0086] in, These are the points in the camera coordinate system that correspond to the points in the point cloud data. For points in a ternary point cloud coordinate system, The rotation matrix is ​​preset to the identity matrix;

[0087] T is the translation matrix, which is preset to [a;b;c], where a is the distance between the camera and the ground, b is the distance between the point cloud device and the ground, and c is the horizontal offset between the camera and the point cloud device.

[0088] Then, the points in the camera coordinate system are transformed to the pixel coordinate system using camera intrinsic parameters. The transformation process satisfies the mathematical expression:

[0089]

[0090] in, A point in the pixel coordinate system of a two-dimensional image. The depth value represents the distance from the point corresponding to a point in the point cloud data in the camera coordinate system to the camera plane;

[0091] and This refers to the camera's focal length, which is a factory-calibrated parameter. , The center offset of the camera coordinate system can be directly read from the intrinsic parameter file provided by the camera manufacturer or calculated using the Zhang calibration method. , These are the x and y coordinates of the pixel's physical scale, obtained by calculating based on the camera sensor size and image resolution. This is the ratio of the sensor width to the image width. Given the ratio of sensor height to image height, obtain the aligned image sequence and the mapped point cloud pixel sequence.

[0092] The anchor frame size is optimized for construction scenarios using the YOLOv8 model. The statistical size settings are based on the worker's safety helmet and body outline. An attention mechanism is introduced to focus on the head and the boundary of the danger zone. The image is aligned for each frame, the worker's behavioral characteristics are identified, and the local image estimation results are output.

[0093] The behavior type is either safe or illegal, and is determined based on highway construction safety behaviors. Safe behaviors include wearing a safety helmet and working in a safe area; illegal behaviors include not wearing a safety helmet and entering a dangerous area. The method of obtaining the information is through matching with preset rules. If a safety helmet outline is detected in the head area and the behavior outline coordinates are within the safe area polygon, it is considered safe by ray casting. If no safety helmet outline is detected or the behavior outline coordinates are within the dangerous area, it is considered illegal.

[0094] The behavior contour coordinates are polygon vertices in the pixel coordinate system. The method is to extract the contour of the worker detection box output by YOLOv8, use Canny edge detection to obtain the edge point set, and then simplify it into polygon vertices through the Douglas-Peucker algorithm.

[0095] The image estimation confidence is obtained by fusing the model detection confidence and contour integrity. The fusion calculation process is that the confidence is equal to 0.7 multiplied by the confidence of the model output plus 0.3 multiplied by the contour integrity. The model output confidence comes from the classification branch of YOLOv8. The contour integrity is the ratio of the actual contour pixel count to the standard pixel count of the complete contour. The standard pixel count is based on the average contour of workers in the construction scene, which is 800±50 pixels. When the contour integrity exceeds 1, it is taken as 1. The contour integrity range is controlled between 0 and 1.

[0096] Occlusion types are classified based on contour integrity, including no occlusion, partial occlusion, and heavy occlusion;

[0097] If the contour integrity is greater than or equal to 0.8, it is judged as no occlusion;

[0098] If the contour integrity is less than 0.8 but greater than 0.4, it is judged as partial occlusion;

[0099] If the contour integrity is less than or equal to 0.4, it is judged as severe occlusion;

[0100] For each frame of mapped point cloud pixels, a density clustering algorithm is used to identify the worker's behavioral characteristics from the mapped point cloud pixel sequence, and output the local estimation result of the point cloud. The neighborhood radius of the density clustering algorithm is 1m, the minimum number of samples is 5, the position coordinates are the coordinates in the pixel coordinate system, the velocity vector is calculated based on the position coordinates of the previous 3 frames, and the position estimation variance is used to reflect the position uncertainty. The method is to calculate the mean square of the Euclidean distance from all points in the density clustering result to the center coordinates, which reflects the dispersion of the point cloud distribution. The center coordinates are the coordinates of the geometric center of each worker point cloud pixel set after density clustering in the pixel coordinate system.

[0101] Establish local estimation association initialization: Based on worker ID, the local estimation results of the image with the local estimation results of the point cloud are initially associated to form an initial local state set.

[0102] This invention not only outputs simple coordinates, but also provides fine metadata support for subsequent data usability assessment by quantifying the integrity of the image contour and the dispersion of the point cloud distribution. Combined with spatiotemporal alignment and downsampling preprocessing, it effectively filters out noise caused by dust and light interference at the construction site while ensuring real-time performance, thus improving the quality of the source data.

[0103] Obtaining fusion feature values ​​specifically includes:

[0104] Calculate the Euclidean distance between the center coordinates of the behavioral contour coordinates in the local estimation result of the image and the position coordinates in the local estimation result of the point cloud. The center coordinates of the contour coordinates are the arithmetic mean of the vertex coordinates of the behavioral contour polygon; that is, the pixel coordinates composed of the average of the x-coordinates and the average of the y-coordinates of all vertices.

[0105] Calculate the one-way Hausdorff distance from all behavior contour coordinates to the preset pixel neighborhood of the position coordinates;

[0106] Methods for calculating fusion eigenvalues ​​include:

[0107] Divide the Euclidean distance of the current frame by the mean of the Euclidean distances of the previous N frames, and divide the one-way Hausdorff distance by the mean of the one-way Hausdorff distances of the previous N frames. Then multiply each by the distance feature contribution weight and sum them up. The Euclidean distance contribution weight is the correlation coefficient of the Euclidean distance in the effective associated samples divided by the sum of the correlation coefficients of the Euclidean distance and the one-way Hausdorff distance. The one-way Hausdorff distance contribution weight is 1 minus the Euclidean distance contribution weight.

[0108] In a specific embodiment of the present invention, the preset pixel neighborhood range is a region with a radius of 5 pixels centered on the position coordinates. The one-way Hausdorff distance is calculated for each point p in the set N of all contour coordinates, and the minimum Euclidean distance from p to all points q in the set M of neighboring pixels is calculated. Then, the maximum value among these minimum distances is taken as the one-way Hausdorff distance.

[0109] The previous N frames are the five frames before the current frame. The five frames before the current frame include historical image data and historical point cloud data. The five frames before the current frame are numbered in chronological order as the first frame, the second frame, the third frame, the fourth frame, and the fifth frame. The fifth frame is the frame before the current frame.

[0110] The mathematical expression for the method of calculating fusion eigenvalues ​​is:

[0111] ;

[0112] in, To fuse feature values, For Euclidean distance, The mean Euclidean distance of the first N frames. For one-way Hausdorff distance, The mean of the one-way Hausdorff distance for the first N frames. The contribution weights for the Euclidean distance. Weights for contributions to Hausdorff distance;

[0113] The calculation of the contribution weights for Euclidean distance and Hausdorff distance includes:

[0114] The ratio of the standard deviation of the Euclidean distance to the mean of the first N frames is used as the coefficient of variation of the Euclidean distance. The ratio of the standard deviation of the one-way Hausdorff distance to the mean of the first N frames is used as the coefficient of variation of the one-way Hausdorff distance. The reciprocal of the coefficient of variation of the Euclidean distance and the reciprocal of the coefficient of variation of the one-way Hausdorff distance are normalized to obtain the contribution weight of the Euclidean distance and the contribution weight of the Hausdorff distance.

[0115] The mathematical expression for normalization is:

[0116] ;

[0117] ;

[0118] in, The contribution weights for the Euclidean distance. is the coefficient of variation of the Euclidean distance. The coefficient of variation of the one-way Hausdorff distance. The contribution weight of Hausdorff distance.

[0119] This invention employs a fusion feature value calculation method that combines Euclidean distance and one-way Hausdorff distance, and introduces a dynamic weight allocation mechanism based on the coefficient of variation. Through Hausdorff distance, this invention can capture the matching degree between image contours and point cloud clusters in terms of shape, overcoming the defect that centroid distance alone cannot accurately match deformed targets. At the same time, the dynamic weight mechanism can automatically suppress the influence of large fluctuations in distance indicators, ensuring that the association matching between the image and the point cloud remains stable and reliable even when the worker's posture changes drastically or there is partial occlusion.

[0120] Obtaining relevance specifically includes:

[0121] If the fusion feature value is less than the preset effective association threshold, the association degree is 1 minus the ratio of the fusion feature value to the mean of the fusion feature values ​​of the previous N frames plus 3 times the standard deviation. If the fusion feature value is greater than or equal to the preset effective association threshold, the association degree is 0.3 multiplied by the negative fusion feature value of the natural index minus the ratio of the effective association threshold to the standard deviation of the effective association samples raised to the power of 3.

[0122] In one specific embodiment of the present invention, the effective association threshold is set to 1;

[0123] The mathematical expression for obtaining the correlation degree based on the fusion feature values ​​is:

[0124] ;

[0125] in, For relevance, To fuse feature values, The mean of the fused feature values ​​of the first N frames. Let be the standard deviation of the fused feature values ​​from the first N frames. This is the preset effective association threshold.

[0126] This invention designs a segmented correlation mapping function, which uses different calculation logic for high and low matching degrees. In particular, when the feature value exceeds the threshold, an exponential decay strategy is adopted, which can identify abnormal associations and make the correlation degree of incorrect matches quickly approach zero. This non-linear scoring mechanism greatly enhances the system's ability to distinguish multi-target intersecting scenes and effectively prevents identity jumps or mistracking caused by dense personnel at construction sites.

[0127] The ternary uncertainty is obtained by using image estimation confidence, location estimation variance, and correlation degree. The ternary uncertainty includes image estimation uncertainty, point cloud estimation uncertainty, and correlation uncertainty.

[0128] Image estimation uncertainty is calculated based on image estimation confidence and correlation. The specific calculation process includes multiplying the difference between 1 and image estimation confidence and the difference between 1 and correlation.

[0129] The mathematical expression for the image estimation uncertainty is:

[0130] ;

[0131] in, To estimate uncertainty for an image, Estimate the confidence level for the image. For relevance;

[0132] The uncertainty of point cloud estimation is calculated based on the variance of location estimation and the degree of correlation. The specific calculation process includes multiplying the variance of location by 1 and adding 0.5 times the difference between 1 and the degree of correlation.

[0133] The mathematical expression for the uncertainty in point cloud estimation is:

[0134] ;

[0135] in, To estimate uncertainty in point cloud data, For location variance, For relevance;

[0136] The uncertainty of association is calculated based on the degree of association. The specific calculation process includes the difference between 1 and the degree of association.

[0137] The mathematical expression for the associated uncertainty is:

[0138] ;

[0139] in, To address the uncertainty of the relationship, For relevance;

[0140] The product of image estimation uncertainty and association uncertainty is taken as image association coupling uncertainty, and the product of point cloud estimation uncertainty and association uncertainty is taken as point cloud association coupling uncertainty.

[0141] The lower the confidence level of image estimation, the lower the confidence level of behavior, the less reliable the association, and the higher the uncertainty of behavior; the larger the location variance, the less reliable the association, and the higher the location uncertainty; the lower the confidence level of association, the higher the uncertainty of association.

[0142] This invention constructs a multidimensional assessment that includes image estimation, point cloud estimation, and correlation uncertainty, and proposes a method for calculating coupled uncertainty. This invention mathematically decouples and recouples the detection error of the sensor itself with the cross-modal matching error, which can accurately identify the source of uncertainty. This provides a precise quantitative basis for subsequent probability distribution correction and avoids false alarms caused by blindly trusting a single data source.

[0143] Step S4 includes the following sub-steps:

[0144] S041: Based on the point cloud position coordinates of the previous N frames, calculate the inertial prediction position and the historical weighted average prediction position respectively. Calculate a historical behavior weight based on historical association reliability and historical occlusion influence coefficient. Fuse the inertial prediction position and the historical weighted prediction position to obtain a final predicted position coordinate.

[0145] Calculating the predicted position involves subtracting the position coordinates of the frame before the frame from twice the position coordinates of the previous frame.

[0146] Calculating the historical predicted location involves linearly weighting the location coordinates of each frame from the previous preset number of frames.

[0147] Calculating historical correlation reliability involves taking the average correlation between the local image estimation results and the local point cloud estimation results of each frame up to a preset number of frames prior to the current frame.

[0148] Calculating the historical occlusion impact coefficient includes: statistically analyzing the occlusion types of the previous N frames, numerically mapping the occlusion types, and recursively calculating the historical occlusion impact coefficient using the exponential moving average algorithm.

[0149] The historical behavior weight is calculated by multiplying the historical association reliability by 1 minus the historical occlusion impact coefficient;

[0150] Calculating the final predicted position involves multiplying the historical behavior weight by the historical predicted position, adding the product of the difference between 1 and the historical behavior weight and the inertial predicted position, to obtain the final predicted position coordinates.

[0151] S042: Calculate the standard deviation of the position coordinates of the first N frames, take half of the standard deviation as the base standard deviation, and take the square of the base standard deviation as the base variance.

[0152] The diffusion coefficient is calculated using the following mathematical expression:

[0153] ;

[0154] in, The dispersion coefficient is... For relevance, The mean of the fused feature values ​​of the first N frames. The standard deviation of the fused feature values ​​from the first N frames;

[0155] The initial variance is obtained by adjusting the base variance using the dispersion coefficient. The adjustment includes multiplying the squared value of the dispersion coefficient by the base variance.

[0156] The initial variance is corrected to obtain the corrected variance. The correction includes multiplying the initial variance by 1 and summing the uncertainty of the coupling with the point cloud.

[0157] Centered on the final predicted position, the predicted position range is defined by the area with the radius determined by half the standard deviation of the position coordinates of the previous N frames and the diffusion coefficient. Each pixel within this predicted position range is traversed, and the probability value of each pixel is calculated using a two-dimensional Gaussian probability density function with the final predicted position as the mean and the corrected variance as the covariance. The probability value is then normalized to obtain the prior position distribution.

[0158] In one specific embodiment of the present invention, the coordinates of the inertial prediction position are the position coordinates of the fifth frame plus the difference between the position coordinates of the fifth frame and the fourth frame.

[0159] Historical predicted position involves linearly weighting the position coordinates of each frame in the previous preset number of frames, with the weight of the frame closer to the current frame being greater. The weight of the first frame is 1, the weight of the second frame is 2, the weight of the third frame is 3, the weight of the fourth frame is 4, and the weight of the fifth frame is 5.

[0160] Calculating the historical occlusion impact coefficient involves: statistically analyzing the occlusion types of the previous N frames, mapping the occlusion types numerically (0 for no occlusion, 1 for partial occlusion, and 2 for severe occlusion), and recursively calculating the historical occlusion impact coefficient using an exponential moving average algorithm. The mathematical expression for calculating the historical occlusion impact coefficient is as follows:

[0161] ;

[0162] in, The historical occlusion impact coefficient. The numerical value mapped to the occlusion type of the current frame. The preset smoothing factor, α, ranges from 0.1 to 0.4. This represents the historical occlusion impact coefficient of the previous frame. The frame index of the current frame is a positive integer used to identify frames arranged in chronological order. The historical occlusion impact coefficient of the first frame is a value mapped to the occlusion type of the first frame.

[0163] The fundamental variance reflects the inherent volatility of the historical movement trajectory of workers;

[0164] The physical meaning of the dispersion coefficient is: the lower the current correlation and the more discrete the historical correlation distribution, the more diffuse the predicted probability field should be.

[0165] The physical meaning of initial variance is that if the current correlation is unstable, it further amplifies the uncertainty of the forecast based on historical volatility.

[0166] The probability value of each pixel is calculated using a two-dimensional Gaussian probability density function and then normalized. The normalization method is to add the probability values ​​of all pixels together to get a sum, and then divide the probability value of each pixel by the sum.

[0167] This invention proposes an adaptive location prediction model based on the diffusion coefficient and the historical occlusion influence coefficient. By using inertial prediction and historical weighted smoothing trajectory, when the correlation decreases or historical occlusion is severe, the prediction range can be automatically expanded by increasing the diffusion coefficient, that is, increasing the uncertainty radius of the prediction. This ensures that even in the case of severe occlusion or correlation failure, effective coverage of the target area can still be maintained, preventing target loss.

[0168] Step S4 also includes the following sub-steps:

[0169] S043: Statistically analyze the number of times safe and illegal behaviors occurred in the historical data of the previous N frames of the current frame, calculate their proportion, and obtain the initial behavior distribution;

[0170] S044: The initial behavior distribution is modified by weighting and mixing the initial behavior distribution with the uniform distribution [0.5, 0.5] to obtain the behavior prior distribution;

[0171] The mathematical expression for weighted mixing is:

[0172] ;

[0173] in, For the prior distribution of behavior, For image correlation coupling uncertainty, This represents the initial behavior distribution.

[0174] In a specific embodiment of the present invention, the initial behavior distribution includes safe behaviors and their proportions, i.e., the probability of safe behaviors occurring, and illegal behaviors and their proportions, i.e., the probability of illegal behaviors occurring. The weighted mixing formula is used to correct the probabilities of safe behaviors and illegal behaviors. The purpose of the correction is to make the predicted probability distribution more certain, i.e., closer to a uniform distribution, when the image uncertainty is high.

[0175] This invention introduces a uniform distribution to perform uncertainty-weighted correction on the prior distribution of behavior. When the uncertainty of image correlation coupling is high, it forces the flattening of the probability distribution of safety and violation, making it tend towards an uncertain state. This mechanism effectively prevents the system from arbitrarily determining whether a worker is safe or violating the rules based solely on historical experience when the sensor data quality is extremely poor.

[0176] Step S4 also includes:

[0177] Extract the probability density value at the location coordinates based on the prior distribution of the location, and calculate the membership degree based on the probability density value, the prior distribution of the behavior, and the degree of association.

[0178] Calculating membership involves multiplying the probability density value at the location coordinates, the probability value of the prior distribution of behavior, and the degree of association, and then dividing the result by the total number of targets.

[0179] The target sum is the sum of the products of the prior probability density value of the location, the probability corresponding to the behavior type in the prior distribution of behavior, and the correlation degree for each worker identified in the current frame.

[0180] The common measurement influence coefficient is equal to 1 and the sum of the uncertainty of point cloud association coupling and the uncertainty of image association coupling;

[0181] The product of the common measurement influence coefficient and membership degree is used as the corrected association probability;

[0182] The location prior distribution is taken as the first Gaussian distribution;

[0183] Construct a second Gaussian distribution, where the mean of the second Gaussian distribution is the center coordinate of the behavior contour coordinates, and the covariance is a diagonal matrix with the image association uncertainty as the diagonal element;

[0184] Construct a third Gaussian distribution, where the mean of the third Gaussian distribution is the location coordinates and the covariance is a diagonal matrix with the point cloud association uncertainty as the diagonal element;

[0185] The generalized covariance interaction algorithm is adopted to use the corrected association probability as the weight of the second Gaussian distribution and the third Gaussian distribution, and to perform weighted fusion with the location prior distribution to obtain the location posterior distribution.

[0186] This claim employs a generalized covariance interaction algorithm to fuse the corrected multi-Gaussian distribution, and calculates the fusion weight using membership degree and common measurement influence coefficient. This method breaks through the limitation of traditional Kalman filtering relying solely on covariance, and can perform probabilistic-level fusion of location prior distribution, image observation distribution, and point cloud observation distribution. Through membership degree calculation, it can accurately handle the probability allocation problem when multiple people overlap in the same pixel area, significantly improving the positioning accuracy under crowded work surfaces.

[0187] The security behavior identification results output based on the fusion results include:

[0188] The weighted centroid of the posterior distribution of the location is extracted as the final mean of the location, and the probability value corresponding to the violation in the prior distribution of the behavior is extracted.

[0189] The ray method is used to determine whether the final average position is located in the danger zone. If the average position is located in the danger warning zone and the probability of violation is greater than the preset safety threshold, it is judged as a high-risk violation.

[0190] If the average location is in the danger warning zone or the probability of violation is greater than the preset safety threshold, it is judged as a low-risk violation.

[0191] If the average location is not in a dangerous area and the probability of violation is less than or equal to the preset safety threshold, then it is considered safe.

[0192] In one specific embodiment of the present invention, the preset security threshold is 0.8.

[0193] This invention establishes a graded risk judgment standard based on dual verification of location mean and violation probability. By transforming the complex posterior probability distribution into three intuitive levels: high-risk violation, low-risk violation, and safety, this technical solution provides construction site managers with graded early warning capabilities. For example, it can trigger an emergency shutdown only for high-risk behaviors that enter dangerous areas, thus balancing the rigor of safety supervision with the efficiency of on-site operations.

[0194] A remote identification system for highway construction safety behaviors based on image recognition includes a data acquisition module, an association module, an estimation module, and an identification module.

[0195] The acquisition module is used to preprocess, align, and extract features from image data and point cloud data to obtain local image estimation results and local point cloud estimation results.

[0196] The correlation module is used to obtain the correlation between the local estimation results of the image and the local estimation results of the point cloud by constructing fusion feature values ​​based on Euclidean distance and one-way Hausdorff distance.

[0197] The estimation module is used to calculate the ternary uncertainty based on the local estimation results of the image, the local estimation results of the point cloud, and the correlation degree.

[0198] The recognition module is used to construct location and behavior distributions based on ternary uncertainty, correlation degree, local image estimation results and local point cloud estimation results, and to predict location and behavior. It performs weighted fusion of the prediction results, local image estimation results and local point cloud estimation results, and outputs the safety behavior recognition results based on the fusion results.

[0199] This invention significantly improves the robustness and accuracy of safety behavior monitoring in complex highway construction scenarios by constructing a remote recognition architecture that integrates image and point cloud features. Its core advantage lies in introducing a dual association evaluation mechanism based on Euclidean distance and one-way Hausdorff distance, and dynamically allocating weights using the coefficient of variation. This invention captures subtle deformations of contour edges through Hausdorff distance, solving the problem of matching failure when relying solely on centroid distance due to changes in worker posture or partial occlusion. This ensures accurate spatiotemporal alignment and association of multi-source heterogeneous data. This invention establishes a ternary uncertainty coupling evaluation mechanism including image estimation, point cloud estimation, and association dimensions. This mechanism can quantify the reliability of perceived data in real time and, based on… This adaptive trajectory prediction mechanism automatically adjusts the diffusion coefficient to expand the probability distribution range of location prediction when the system detects a decrease in correlation confidence or the presence of occlusion. It also corrects the inertial prediction weight by combining the historical occlusion influence coefficient. This invention automatically suppresses aggressive violation judgments under highly uncertain conditions by calculating a weighted mixture of the prior distribution of behavior and the uniform distribution, demonstrating safety redundancy at the algorithm level. Based on the coupling relationship between the centroid of the posterior distribution and the dangerous area, it can output differentiated results for high-risk violations, low-risk violations, and safety. This not only significantly reduces the false alarm rate caused by environmental factors but also provides construction site managers with graded early warning capabilities, significantly improving the intelligence level and response efficiency of safety supervision.

[0200] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0201] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A remote identification method for highway construction safety behaviors based on image recognition, characterized in that, Includes the following steps: Step S1: Preprocess, align, and extract features from the image data and point cloud data to obtain local image estimation results and local point cloud estimation results; Step S2: By constructing fused feature values ​​based on Euclidean distance and one-way Hausdorff distance, the correlation between the local estimation results of the image and the local estimation results of the point cloud is obtained; Step S3: Calculate the ternary uncertainty based on the local image estimation results, the local point cloud estimation results, and the correlation degree; The ternary uncertainty includes image estimation uncertainty, point cloud estimation uncertainty, and association uncertainty; Image estimation uncertainty is calculated based on image estimation confidence and correlation. The specific calculation process includes multiplying the difference between 1 and image estimation confidence and the difference between 1 and correlation. The uncertainty of point cloud estimation is calculated based on the variance of location estimation and the degree of correlation. The specific calculation process includes multiplying the variance of location by 1 and adding 0.5 times the difference between 1 and the degree of correlation. The uncertainty of association is calculated based on the degree of association. The specific calculation process includes the difference between 1 and the degree of association. The product of image estimation uncertainty and association uncertainty is taken as image association coupling uncertainty, and the product of point cloud estimation uncertainty and association uncertainty is taken as point cloud association coupling uncertainty; Step S4: Based on the ternary uncertainty, correlation degree, image local estimation results and point cloud local estimation results, construct position distribution and behavior distribution to predict position and behavior, perform weighted fusion of prediction results, image local estimation results and point cloud local estimation results, and output safety behavior recognition results based on the fusion results.

2. The remote identification method for highway construction safety behavior based on image recognition as described in claim 1, characterized in that, The image data preprocessing includes denoising and normalization of each frame of the continuous frame image, and the point cloud data preprocessing includes outlier removal and density unification. The feature extraction includes identifying worker behavior features from the aligned image sequence using a large model and outputting local image estimation results, which include worker ID, behavior type, behavior contour coordinates, occlusion type, and image estimation confidence. The feature extraction also includes identifying worker behavior features from the pixel sequence of the mapped point cloud using a density clustering algorithm, and outputting a local point cloud estimation result, which includes worker ID, position coordinates, velocity vector, and position estimation variance. Based on the worker ID, the local image estimation results and the local point cloud estimation results of the same worker are initially correlated to form an initial local state set.

3. The remote identification method for highway construction safety behavior based on image recognition as described in claim 2, characterized in that, Obtaining the fusion feature values ​​specifically includes: Calculate the Euclidean distance between the center coordinates of the behavioral contour coordinates in the local estimation result of the image and the position coordinates in the local estimation result of the point cloud, wherein the center coordinates of the contour coordinates are the arithmetic mean of the vertex coordinates of the behavioral contour polygon; Calculate the one-way Hausdorff distance from all behavioral contour coordinates to the preset pixel neighborhood of the position coordinates; The method for calculating the fusion feature value includes: Divide the Euclidean distance of the current frame by the mean of the Euclidean distances of the previous N frames, and divide the one-way Hausdorff distance by the mean of the one-way Hausdorff distances of the previous N frames. Then multiply each by the distance feature contribution weight and sum them up. The Euclidean distance contribution weight is the correlation coefficient of the Euclidean distance in the effective associated samples divided by the sum of the correlation coefficients of the Euclidean distance and the one-way Hausdorff distance. The one-way Hausdorff distance contribution weight is 1 minus the Euclidean distance contribution weight.

4. The remote identification method for highway construction safety behavior based on image recognition as described in claim 3, characterized in that, Obtaining the correlation specifically includes: If the fusion feature value is less than the preset effective association threshold, the association degree is 1 minus the ratio of the fusion feature value to the mean of the fusion feature values ​​of the previous N frames plus 3 times the standard deviation. If the fusion feature value is greater than or equal to the preset effective association threshold, the association degree is 0.3 multiplied by the negative fusion feature value of the natural index minus the ratio of the effective association threshold to the standard deviation of the effective association samples raised to the power of 3.

5. The remote identification method for highway construction safety behavior based on image recognition as described in claim 4, characterized in that, Step S4 includes the following sub-steps: S041: Based on the point cloud position coordinates of the previous N frames, calculate the inertial prediction position and the historical weighted average prediction position respectively. Calculate a historical behavior weight based on historical association reliability and historical occlusion influence coefficient. Fuse the inertial prediction position and the historical weighted prediction position to obtain the coordinates of a final prediction position. Calculating the inertial predicted position involves subtracting the position coordinates of the frame before the frame from twice the position coordinates of the previous frame. Calculating the historical weighted predicted position involves linearly weighting the position coordinates of each frame within the previous preset number of frames; Calculating the historical correlation reliability includes taking the average correlation between the local image estimation results and the local point cloud estimation results of each frame up to a preset number of frames before the current frame; Calculating the historical occlusion impact coefficient includes: statistically analyzing the occlusion types of the previous N frames, numerically mapping the occlusion types, and recursively calculating the historical occlusion impact coefficient using an exponential moving average algorithm. The historical behavior weight is calculated by multiplying the historical association reliability by 1 minus the historical occlusion impact coefficient; Calculating the final predicted position involves multiplying the historical behavior weight by the historical predicted position, adding 1 and subtracting the historical behavior weight multiplied by the inertial predicted position to obtain the final predicted position coordinates. S042: Calculate the standard deviation of the position coordinates of the previous N frames, take half of the standard deviation as the base standard deviation, and take the square of the base standard deviation as the base variance; The diffusion coefficient is calculated using the following mathematical expression: ; in, The dispersion coefficient is... For relevance, The mean of the fused feature values ​​of the first N frames. The standard deviation of the fused feature values ​​from the first N frames; The initial variance is obtained by adjusting the base variance using the dispersion coefficient. The adjustment includes multiplying the squared value of the dispersion coefficient by the base variance. The initial variance is corrected to obtain the corrected variance. The correction includes multiplying the initial variance by 1 and summing the uncertainty of the coupling with the point cloud. Centered on the final predicted position, the predicted position range is defined by the area with the radius determined by half the standard deviation of the position coordinates of the previous N frames and the diffusion coefficient. Each pixel within this predicted position range is traversed, and the probability value of each pixel is calculated using a two-dimensional Gaussian probability density function with the final predicted position as the mean and the corrected variance as the covariance. The probability value is then normalized to obtain the prior position distribution.

6. The remote identification method for highway construction safety behavior based on image recognition as described in claim 5, characterized in that, Step S4 further includes the following sub-steps: S043: Statistically analyze the number of times safe and illegal behaviors occurred in the historical data of the previous N frames of the current frame, calculate their proportion, and obtain the initial behavior distribution; S044: The initial behavior distribution is modified by weighting and mixing the initial behavior distribution with the uniform distribution [0.5, 0.5] to obtain the behavior prior distribution; The mathematical expression for weighted mixing is: ; in, For the prior distribution of behavior, For image correlation coupling uncertainty, This represents the initial behavior distribution.

7. The remote identification method for highway construction safety behavior based on image recognition as described in claim 6, characterized in that, Step S4 further includes: The probability density value at the location coordinates is extracted based on the prior distribution of the location, and the membership degree is calculated based on the probability density value, the prior distribution of behavior, and the degree of association. Calculating the membership degree involves multiplying the probability density value at the location coordinates, the probability value of the prior distribution of behavior, and the degree of association, and then dividing the result by the total number of targets. The target sum is the product of the prior probability density values ​​of the positions of all workers identified in the current frame, the probabilities corresponding to the behavior types in the prior distribution of behavior, and the degree of association. The sum of these products is obtained by summing them up. The product of the common measurement influence coefficient and the membership degree is used as the corrected association probability. The common measurement influence coefficient is calculated by adding 1 with the uncertainty of point cloud association coupling and the uncertainty of image association coupling. The location prior distribution is taken as the first Gaussian distribution; Construct a second Gaussian distribution, where the mean of the second Gaussian distribution is the center coordinate of the behavior contour coordinates, and the covariance is a diagonal matrix with the image association uncertainty as the diagonal element; Construct a third Gaussian distribution, where the mean of the third Gaussian distribution is the location coordinates and the covariance is a diagonal matrix with the point cloud association uncertainty as the diagonal element; The generalized covariance interaction algorithm is adopted to use the corrected association probability as the weight of the second Gaussian distribution and the third Gaussian distribution, and to perform weighted fusion with the location prior distribution to obtain the location posterior distribution.

8. The remote identification method for highway construction safety behavior based on image recognition as described in claim 7, characterized in that, The security behavior identification results output based on the fusion results include: The weighted centroid of the posterior distribution of the location is extracted as the final mean of the location, and the probability value corresponding to the violation in the prior distribution of the behavior is extracted. The ray method is used to determine whether the final average position is located in the danger zone. If the average position is located in the danger warning zone and the probability of violation is greater than the preset safety threshold, it is judged as a high-risk violation. If the average location is in the danger warning zone or the probability of violation is greater than the preset safety threshold, it is judged as a low-risk violation. If the average location is not in a dangerous area and the probability of violation is less than or equal to the preset safety threshold, then it is considered safe.

9. A remote identification system for highway construction safety behaviors based on image recognition, characterized in that, It includes a data acquisition module, a correlation module, an estimation module, and a recognition module; The acquisition module is used to preprocess, align, and extract features from image data and point cloud data to obtain local image estimation results and local point cloud estimation results. The correlation module is used to obtain the correlation between the local image estimation result and the local point cloud estimation result by constructing a fusion feature value based on Euclidean distance and one-way Hausdorff distance. The estimation module is used to calculate the ternary uncertainty based on the local estimation results of the image, the local estimation results of the point cloud, and the correlation degree. The ternary uncertainty includes image estimation uncertainty, point cloud estimation uncertainty, and association uncertainty; Image estimation uncertainty is calculated based on image estimation confidence and correlation. The specific calculation process includes multiplying the difference between 1 and image estimation confidence and the difference between 1 and correlation. The uncertainty of point cloud estimation is calculated based on the variance of location estimation and the degree of correlation. The specific calculation process includes multiplying the variance of location by 1 and adding 0.5 times the difference between 1 and the degree of correlation. The uncertainty of association is calculated based on the degree of association. The specific calculation process includes the difference between 1 and the degree of association. The product of image estimation uncertainty and association uncertainty is taken as image association coupling uncertainty, and the product of point cloud estimation uncertainty and association uncertainty is taken as point cloud association coupling uncertainty; The recognition module is used to construct position distribution and behavior distribution based on the ternary uncertainty, correlation degree, image local estimation result and point cloud local estimation result, to predict position and behavior, to perform weighted fusion of prediction result, image local estimation result and point cloud local estimation result, and to output safe behavior recognition result based on fusion result.

Citation Information

Patent Citations

  • Construction personnel unsafe behavior detection method based on improved YOLO

    CN120808442A

  • Steel platform equipment safety monitoring method based on dynamic characteristic monitoring

    CN112153673A

  • Method, system and device for constructing point cloud local coordinate system and medium

    CN114648582A