Intelligent construction site management methods and systems based on photoelectric and video fences

The intelligent construction site management system, which utilizes photoelectric sensors and video fences, enables rapid detection of cable faults and accurate identification of responsible parties. This solves the problems of insufficient accuracy and predictive prevention in existing systems under complex environments, thereby improving the intelligence and safety of construction site management.

CN119942767BActive Publication Date: 2026-01-06CHINA RAILWAY CONSTR ENG GRP FOURTH CONSTR CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510430483.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2026-01-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

Existing construction site management systems lack accuracy in cable fault detection and safety monitoring, lack effective data fusion and collaborative analysis capabilities, struggle to cope with changes in lighting and shading in complex environments, and lack predictive and preventative mechanisms.

Method used

An intelligent construction site management system based on photoelectric and video fences is adopted. Through the collaborative work of camera devices, cable monitoring devices, environmental status parameter sensors and intelligent switches of construction equipment, combined with multimodal data fusion technology, real-time data acquisition and processing are carried out to identify anomalies and track the responsible parties.

Benefits of technology

The system improves the accuracy and timeliness of cable fault detection, accurately identifies the responsible party, enhances the efficiency and accuracy of construction site supervision, and strengthens the level of intelligent safety management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942767B_ABST
    Figure CN119942767B_ABST
Patent Text Reader

Abstract

The application discloses a smart construction site management method and system based on photoelectric and video fences, which is realized based on a smart construction site management system and comprises a control module and an information acquisition device connected with the control module, wherein the information acquisition device comprises a camera device, a cable monitoring device, an environmental state parameter sensor and a construction equipment intelligent switch; the cable monitoring device is arranged at a predetermined length interval along the direction in which the cable extends. The method comprises the following steps: regularly acquiring and preprocessing real-time data of the construction site; analyzing video data to identify dynamic areas; evaluating cable abnormalities and equipment states; generating a motion intensity map and a key frame candidate set; tracking targets and generating trajectories to determine responsible subjects. Through multi-source data fusion and intelligent analysis, the method realizes comprehensive monitoring and abnormal detection of the construction site, and improves management efficiency and safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to construction site management, including cable fault detection and management, construction site safety management, and in particular, intelligent construction site management methods and systems based on photoelectric and video fences. Background Technology

[0002] In complex and ever-changing construction environments, traditional manual monitoring methods are no longer sufficient to meet the growing demands for safety and efficiency. In construction sites and industrial enterprises, the various electrical cables connecting construction machinery and power tools face severe safety challenges. Due to harsh operating environments and frequent movement, these cables are more susceptible to damage from external forces such as pulling, crushing, impact, and frequent bending, leading to core breakage. However, because of the outer sheath, core breakage points are often not visible on the surface, making it difficult to quickly and accurately locate the fault. To address this, Chinese Patent CN102590714 provides a cable fault detector, but it does not disclose the data processing method for the detector. This lack of systematic monitoring may lead to oversights and errors in recording, and there is a possibility of omissions, posing potential risks, and offering limited guidance for subsequent management.

[0003] To this end, Chinese patent CN116243072B provides a maintenance management method for a systematic maintenance management system for electrical equipment on construction sites, which discloses the collection and processing of information on electrical equipment, but does not conduct traceability and accountability.

[0004] In summary, most existing management systems still operate independently, lacking effective data fusion and collaborative analysis capabilities. In complex construction site environments, existing video analytics technologies often struggle to handle issues such as changing lighting and occlusion, resulting in low recognition accuracy. Monitoring of critical facilities like cables and pile foundations, as well as construction and hazardous areas, remains weak, particularly in detecting hidden problems such as broken core faults. Furthermore, existing systems largely focus on post-event analysis, lacking effective predictive and preventative mechanisms.

[0005] Therefore, research and innovation are needed. Summary of the Invention

[0006] The purpose of this invention is to provide an intelligent construction site management method and system based on photoelectric and video fences, in order to solve the aforementioned problems existing in the prior art.

[0007] The technical solution is an intelligent construction site management method based on photoelectric and video fences, implemented based on an intelligent construction site management system. It includes a control module and an information acquisition device connected to the control module. The information acquisition device includes a camera device, a cable monitoring device, an environmental status parameter sensor, and an intelligent switch for construction equipment. The cable monitoring device is set at predetermined intervals along the direction of cable extension.

[0008] The method includes the following steps:

[0009] Step S1: Acquire real-time data from the construction site at predetermined intervals and preprocess it to form a preprocessed dataset, including enhanced video data, cable operation data, environmental status parameters, and equipment switch status.

[0010] Step S2: Read the enhanced video data, identify and obtain the current construction site status, obtain and label the dynamic areas, and form a dynamic area dataset;

[0011] Step S3: Read cable operation data and equipment switch status, extract cable abnormal parameters from them, assess the severity of the abnormality, output the abnormality level index and compare it with the threshold. If the threshold is exceeded, record the timestamp of the time when the abnormality occurred and the corresponding detection area coordinates.

[0012] Step S4: Based on the timestamp of the time when the anomaly occurred, retrieve a video segment of a predetermined length from the dynamic region dataset, determine the range of key frame locations, and form a key frame candidate set.

[0013] Step S5: Identify and track targets based on keyframes and perform cross-frame matching to generate target trajectories. Obtain a list of responsible entities based on the target trajectories.

[0014] According to another aspect of this application, a smart construction site management system based on photoelectric and video fences is also provided, including a control module and an information acquisition device connected to the control module. The information acquisition device includes a camera device, a cable monitoring device, an environmental status parameter sensor, and a smart switch for construction equipment. The cable monitoring device is installed at predetermined lengths along the direction of cable extension. The control module includes:

[0015] At least one processor; and,

[0016] A memory communicatively connected to at least one of the processors; wherein,

[0017] The memory stores instructions that can be executed by the processor to implement the intelligent construction site management method based on photoelectric and video fences as described in any of the above technical solutions.

[0018] Beneficial effects include the ability to effectively detect and locate hidden faults such as broken cores through multimodal data fusion technology combined with cable parameter monitoring and video analysis; improved accuracy and timeliness of fault diagnosis; accurate identification of responsible parties through target tracking and trajectory analysis; and enhanced efficiency and accuracy of construction site supervision. Attached Figure Description

[0019] Figure 1 This is a flowchart of the present invention.

[0020] Figure 2 This is a flowchart of step S1 of the present invention.

[0021] Figure 3 This is a flowchart of step S2 of the present invention.

[0022] Figure 4 This is a flowchart of step S3 of the present invention.

[0023] Figure 5 This is a flowchart of step S4 of the present invention.

[0024] Figure 6 This is a flowchart of step S5 of the present invention. Detailed Implementation

[0025] According to one aspect of this application, a smart construction site management method based on photoelectric and video fences is provided, which is implemented based on a smart construction site management system. The method includes a control module and an information acquisition device connected to the control module. The information acquisition device includes a camera device, a cable monitoring device, an environmental status parameter sensor, and a smart switch for construction equipment. The cable monitoring device is set at predetermined intervals along the direction of cable extension.

[0026] The method includes the following steps:

[0027] Step S1: At predetermined intervals, real-time data of the construction site is acquired through the information acquisition device and preprocessed to form a preprocessed dataset; the real-time data of the construction site includes enhanced video data of the construction site, cable operation data, environmental status parameters and equipment switch status;

[0028] Step S2: Read the construction site video data in the preprocessed dataset, identify and obtain the current construction site status, obtain dynamic areas and label them to form a dynamic area dataset;

[0029] Step S3: Read the cable operation data and equipment switch status from the preprocessed dataset, and call the pre-configured anomaly detection module to extract cable anomaly parameters; fuse the cable anomaly parameters, equipment switch status and environmental status parameters, call the fuzzy inference system to evaluate the severity of the anomaly, output the anomaly level index and compare it with the threshold. If it exceeds the threshold, record the timestamp of the anomaly occurrence and the corresponding detection area coordinates.

[0030] Step S4: Based on the timestamp, retrieve video segments of a predetermined length from the construction site video data, calculate the motion vector of each frame, and generate a motion intensity map; determine the location range of key frames and form a key frame candidate set;

[0031] Step S5: Identify the tracking target based on the keyframe candidate set, perform cross-frame matching on the tracking target, generate the target trajectory, and obtain the list of responsible entities based on the target trajectory; the tracking targets include personnel and equipment.

[0032] In this embodiment, through the coordinated operation of camera devices, cable monitoring devices, environmental condition parameter sensors, and intelligent switches for construction equipment, the system can comprehensively collect real-time data from the construction site. Multi-source data acquisition significantly improves the comprehensiveness and accuracy of monitoring. Environmental condition parameters enhance the accuracy and reliability of anomaly detection. By generating motion intensity maps and keyframe recognition, the system can quickly locate critical events, improving the response speed to anomalies. Through target tracking and trajectory analysis, the system can accurately identify the responsible party. In summary, this improves the efficiency and accuracy of construction site supervision.

[0033] In this embodiment, the hardware system is as follows:

[0034] The control module, as the central processing unit of the system, is responsible for coordinating and managing the operation of the entire system. It establishes direct data connections with various information acquisition devices, receiving and processing information from these devices.

[0035] Multiple high-definition cameras are installed along the site perimeter and in key areas. The cameras are connected to the control module via wired or wireless networks. Real-time video streams are transmitted to the control module for processing and analysis.

[0036] Cable monitoring devices, similar to cable fault detectors, are set up with a monitoring point every predetermined length (e.g., every 50 meters) along the cable's extension direction.

[0037] Each monitoring point contains the following components: a cable fault tester, directly connected to the cable under test; a fault sensing probe, connected to the audio signal input of the tester; a positioning signal sensor and a path signal sensor, connected to the cable under test and transmitting signals to a positioning signal receiver; and a high-energy impact signal generator and a path signal generator, connected to the cable under test and used to generate test signals. Data from all these components is aggregated through a dedicated data acquisition unit and then transmitted to the control module.

[0038] Environmental condition parameter sensors include various environmental sensors installed at key locations on the construction site, such as temperature, humidity, dust, and noise sensors. These sensors transmit data to the control module in real time via wireless networks (such as LoRa or NB-IoT).

[0039] Smart switches are installed on the main construction equipment. These switches can monitor the equipment's on / off status, operating time, energy consumption, and other information. The data is transmitted to the control module in real time via a wireless network.

[0040] In another embodiment of this application, an array of photoelectric sensors for perimeter protection is also included.

[0041] According to one aspect of this application, step S1 specifically comprises:

[0042] Step S11: At predetermined intervals, acquire the raw video data collected by the construction site camera device and perform enhancement processing to obtain enhanced video data;

[0043] Step S12: Obtain cable parameters through the cable monitoring device to form an original cable dataset, including cable normal mode, abnormal mode, and cable parameter values;

[0044] Step S13: Obtain the initial environmental state parameters collected by each environmental state parameter sensor, and perform fusion processing to obtain fused environmental state parameters;

[0045] Step S14: Collect the working parameters of the intelligent switch of the construction equipment, construct the original equipment data, and perform preprocessing to form the equipment switch status;

[0046] Step S15: Based on the output data from steps S11 to S14, construct a multimodal dataset, and perform time alignment using a multidimensional dynamic time warping algorithm to form a preprocessed dataset, which is the construction site basic dataset, and store it.

[0047] Enhancement processing of raw video data improves video quality, aiding in subsequent video analysis. Acquiring cable parameters through cable monitoring devices and distinguishing between normal and abnormal modes helps in the rapid identification of potential cable problems. Environmental condition parameters.

[0048] According to one aspect of this application, step S2 specifically comprises:

[0049] Step S21: Read the enhanced video data, extract video frames, and for each video frame, use superpixel segmentation to divide the region into several superpixel regions and extract feature vectors. Then, use hierarchical clustering to cluster the feature vectors to obtain the initial scene segmentation result. This method is more efficient than traditional pixel-level processing and can preserve the edge information of the image. Through clustering, it can better adapt to complex construction site scenes.

[0050] Step S22: For the initial scene segmentation result, optimize it using a spatiotemporal Markov random field module and minimize the energy function to obtain a spatiotemporally consistent scene segmentation result; improve the consistency and accuracy of scene segmentation.

[0051] Step S23: Based on the spatiotemporally consistent scene segmentation results, an improved optical flow estimation method is used to calculate the motion vector field between consecutive frames. Dynamic regions are identified based on the motion vector field, and significant motion regions are segmented by calculating the magnitude and direction histograms of the motion vectors to obtain motion analysis results. For each motion in a significant motion region, tracking is performed to obtain dynamic region tracking results, i.e., dynamic region dataset.

[0052] Step S24: Annotate and hierarchically represent the dynamic region tracking results, incorporating and integrating cable operation data and environmental data to construct and output a comprehensive dynamic region description dataset. This involves establishing a mapping relationship between images, cables, and environmental status parameters. The spatiotemporal range of the dynamic region is used as a constraint to filter corresponding cable and environmental information, reducing the total amount of information. Data that may have issues is collected, reducing the scope and workload of subsequent filtering.

[0053] It significantly improves the system's ability to perceive the dynamic situation of the construction site, enabling timely detection and tracking of abnormal activities, and enhancing the effectiveness of construction site safety supervision.

[0054] According to one aspect of this application, step S3 specifically comprises:

[0055] Step S31: Obtain the comprehensive dynamic region description dataset, perform multivariate anomaly detection and dimensionality reduction to obtain the processed multi-source synchronous dataset, including cable operation data, environmental status model, and dynamic region dataset; effectively reducing data redundancy and improving the efficiency of subsequent analysis.

[0056] Step S32: Read cable operation data and perform modal decomposition and construct a multi-dimensional phase space. Use clustering algorithm to initially detect the existence of abnormal patterns. Simultaneously, based on the processed multi-source synchronous dataset, call the pre-configured anomaly detection module. The module uses sliding window singular spectrum analysis and local anomaly factor method to calculate and output the cable anomaly detection results, thus obtaining the cable anomaly parameter set. The cable anomaly parameter set includes: anomaly type, anomaly score, and corresponding timestamp; this improves the accuracy and reliability of anomaly detection.

[0057] Step S33: Fuse the cable anomaly parameter set with the equipment switch status and environmental status parameters in the multi-source synchronous dataset, and output the fused anomaly feature dataset, i.e., the anomaly detection result.

[0058] Step S34: Call the pre-configured fuzzy inference system to evaluate each data point in the abnormal feature dataset (abnormal detection results), obtain the abnormal severity evaluation result, and compare it with the abnormal level threshold to obtain the abnormal level determination result.

[0059] Step S35: Record the timestamp and coordinate information of the severe anomaly assessment results, and output the spatiotemporal positioning results.

[0060] Environmental condition parameters improve the system's ability to identify and respond to abnormal situations at the construction site, providing technical support for the timely detection and handling of potential safety hazards, thereby enhancing the overall safety level of the construction site.

[0061] According to one aspect of this application, step S4 specifically comprises:

[0062] Step S41: Read the dynamic region dataset and anomaly detection results, and extract the corresponding video segments from the dynamic region dataset;

[0063] Step S42: Use an adaptive contrast enhancement algorithm to further enhance and stabilize the video clips, and output an optimized set of video clips;

[0064] Step S43: For each optimized set of video segments, obtain and calculate motion saliency based on the corresponding motion analysis results to obtain a motion intensity map, calculate visual saliency, generate a multi-feature fusion saliency map, and output the saliency sequence of each video segment; generating a multi-feature fusion saliency map can capture important information in the video more comprehensively.

[0065] Step S44: Based on the optimized video clip set, calculate temporal saliency to obtain the potential keyframe location range, determine the potential keyframes, and output a keyframe candidate set. Calculating temporal saliency to determine the potential keyframe location range can effectively locate the critical moments when abnormal events occur.

[0066] In some embodiments, if the detection finds that the dynamic region dataset cannot cover all key video segments, then based on temporal and spatial continuity analysis, additional video frames are searched in the video to supplement the key video frames.

[0067] This not only improves the accuracy of keyframe extraction but also reduces the amount of video data requiring manual review. Video analysis and key information extraction improve the response speed and processing efficiency of abnormal events.

[0068] According to one aspect of this application, step S5 specifically comprises:

[0069] Step S51: Read the keyframe dataset, dynamic region dataset, anomaly detection results, and pre-configured registration database; use the YOLOvx target detection algorithm for identification and tracking; obtain the target list; generally, YOLOv5 or a later version is used.

[0070] Step S52: Perform cross-frame matching on the target results one by one; generate target trajectory data to form a target trajectory; effectively track moving targets and generate continuous target trajectory data.

[0071] Step S53: Combining the target trajectory, use a spatiotemporal correlation analysis algorithm to determine the target closest to the detection area at the time of the anomaly occurrence, and output a list of potential responsible parties. This effectively associates abnormal events with specific targets (such as workers or equipment).

[0072] In most cases, the responsible party information can be obtained from key video frames. If information is missing, such as in cases of broken or missing cross-frame matching, a search is then conducted in the dynamic region dataset. If the information still cannot be found, a search is performed in the enhanced video. Overall, the search is fast and efficient.

[0073] This not only improves the efficiency of accident investigations but also provides objective and reliable evidence for construction site safety management. It enhances the accuracy and intelligence of construction site supervision, providing technical solutions for timely detection of safety hazards and rapid identification of responsible parties.

[0074] According to one aspect of this application, in step S32, cable operation data is read and modal decomposition and multidimensional phase space are constructed, and a clustering algorithm is used to preliminarily detect whether there are abnormal patterns, specifically:

[0075] Step S321: Read each time series in the cable operation data, identify all local extreme points in the time series, generate upper and lower envelopes using the cubic spline interpolation method and calculate the mean envelope, calculate the difference between each time series and the mean envelope, and determine whether the difference satisfies the intrinsic mode condition. If it does, output the intrinsic mode function, calculate the residual between the time series and the intrinsic mode function, and use the residual as a new time series until the residual becomes a monotonic function.

[0076] Step S322: Read each intrinsic mode function, perform Hilbert transform on it to obtain the analytic signal; calculate the instantaneous amplitude and instantaneous phase based on the analytic signal; calculate the instantaneous frequency using the instantaneous phase; output the instantaneous frequency and amplitude data of each intrinsic mode function;

[0077] Step S323: Read the instantaneous frequency and amplitude data of each cable parameter, select a specific frequency component as a dimension, and construct a multi-dimensional phase space; map the data at each time point to this space to form a point set; output the constructed multi-dimensional phase space data.

[0078] Step S324: Read the multidimensional phase space data, set the clustering parameter ε and the minimum number of points threshold; for each point in the space, calculate the number of points in its ε-neighborhood; if the number of points is greater than or equal to the minimum number of points threshold, mark the point as a core point; for each core point, recursively add the points in its ε-neighborhood to the same cluster; mark the points that are not assigned to any cluster as outliers; output the clustering results and outlier labels.

[0079] Step S325: Read the clustering results and outlier markers, and analyze the distribution characteristics of outliers; combine the original cable parameter data to extract the specific parameter values ​​corresponding to the outliers; identify and classify the detected outlier patterns according to the preset outlier pattern judgment criteria; and output the outlier pattern description and the corresponding cable parameter values.

[0080] This method improves the accuracy and reliability of cable anomaly detection, enabling early identification of potential cable problems and thus preventing possible safety accidents. Simultaneously, it provides crucial data support for cable maintenance and lifespan prediction, helping to optimize cable management strategies and enhance electrical safety at construction sites.

[0081] According to one aspect of this application, in step S32, based on the processed multi-source synchronous dataset, a pre-configured anomaly detection module is invoked. This module employs a sliding window singular spectrum analysis method and a local anomaly factor method to calculate and output the cable anomaly detection results, specifically:

[0082] Step S326: Read the multi-source synchronous dataset, extract cable-related parameters, and construct a multi-dimensional time series. Each dimension represents a cable parameter, and each point on the time axis corresponds to a sampling time.

[0083] Step S327: Construct a trajectory matrix based on the multidimensional time series and determine the size of the sliding window; perform singular value decomposition on the estimated matrix, obtain and sort the singular values, reconstruct the signal based on the first N singular values ​​and calculate the reconstruction error;

[0084] Step S328: Read the reconstruction error, calculate the moving average and standard deviation of the reconstruction error, calculate the threshold based on the moving average, standard deviation and pre-configured adjustable parameters, and if the reconstruction error is greater than the threshold, mark it as a potential outlier to form a set of potential outliers.

[0085] Step S329: For each potential outlier, calculate the distance and reachability distance from each point to its k-th nearest neighbor. Calculate the local density of the outlier based on the reachability distance. Calculate the LOF score by averaging the ratios of the local densities of each outlier's neighbors to the outlier's own local density. Determine if the LOF score is greater than a threshold. If it is, confirm it as an outlier and form an outlier set. N and k are natural numbers greater than 0.

[0086] This method improves the sensitivity and reliability of anomaly detection, enabling early detection of potential cable problems and effectively preventing possible safety accidents. Furthermore, its adaptability allows it to better handle different types of cables and varying site environments, enhancing the system's versatility and practicality.

[0087] According to one aspect of this application, step S32 further includes:

[0088] Step S3210: Randomly select several outliers as cluster centers, and assign each remaining outlier to the nearest cluster center. Update the cluster center to the mean of all outliers in that class, until convergence.

[0089] Step S3211: Call the pre-configured decision tree module to perform feature importance analysis on each cluster and generate descriptive labels for each cluster based on feature importance;

[0090] Step S3212: Using the dynamic time warping algorithm, calculate the distance between each cluster and the normal case, and give the anomaly analysis results.

[0091] This not only improves the accuracy of anomaly detection but also provides more detailed and meaningful anomaly information. Through automatic classification and severity assessment, the system reduces the workload of manual analysis and improves the efficiency of anomaly handling. By identifying different types of anomalies and their severity, managers can develop more targeted maintenance strategies and security measures.

[0092] According to one aspect of this application, the process of enhancing the original video data of the construction site in step S11 specifically includes:

[0093] Step S111: Obtain the original video data of the construction site, extract the video frame images, and decompose each video frame image using discrete wavelet transform to obtain four frequency sub-bands.

[0094] Step S112: For each frequency sub-band, calculate the local mean and standard deviation of the frequency sub-band; construct an enhancement function based on the local mean, standard deviation, and pre-stored adjustable parameters; apply the enhancement function to each pixel in the frequency sub-band.

[0095] Step S113: Reconstruct the enhanced frequency subband using a weighted summation method to obtain the enhanced image. Stitch the enhanced images together to form an enhanced video stream, i.e., enhanced video data.

[0096] By adjusting adjustable parameters, the system can flexibly balance detail enhancement and noise suppression, thus adapting to different construction site environments and monitoring needs. The weighted summation method is used to reconstruct the enhanced frequency subbands, effectively fusing information from different frequency components to generate enhanced images with better visual effects. This not only improves image clarity and contrast but also effectively suppresses noise, making it particularly suitable for processing video data in complex and variable environments such as construction sites.

[0097] According to one aspect of this application, step S21 specifically comprises:

[0098] Step S211: Read the enhanced video data and extract video frames;

[0099] Step S212: For each frame of the enhanced video data, divide it into several initial regions and set the center of each region. For each pixel, iteratively calculate its distance from the superpixel center, update the pixel assignment and the superpixel center, until convergence or the maximum number of iterations is reached, and obtain the superpixel segmentation result.

[0100] Step S213: For each superpixel region in the superpixel segmentation result, calculate the color histogram, extract texture features using the improved SURF algorithm, calculate the shape descriptor, and combine these features into a comprehensive feature vector.

[0101] Step S214: Use a hierarchical clustering algorithm to cluster all comprehensive feature vectors, calculate the distance matrix between feature vectors, construct a hierarchical tree, and cut the feature hierarchical tree according to a preset threshold to obtain the initial scene segmentation result.

[0102] In this embodiment, a superpixel segmentation method is used to initially divide the video frames. This method is more efficient than traditional pixel-level processing and can effectively preserve the edge information and local consistency of the image. By iteratively optimizing the superpixel centers and pixel assignments, the system can generate preliminary segmentation results that conform to the image structure.

[0103] According to one aspect of this application, step S22 specifically comprises:

[0104] Step S221: Read the initial scene segmentation results of several consecutive frames, construct a spatiotemporal Markov random field model, and construct an energy function containing data terms and smoothing terms;

[0105] Step S222: Use the Gaussian mixture model to estimate the probability that a pixel belongs to each label, and calculate the data terms in the energy function;

[0106] Step S223: Calculate the weights based on the color similarity and spatial distance of adjacent pixels to obtain the smoothing term in the energy function;

[0107] Step S224: Apply the graph cut algorithm to minimize the energy function to obtain the optimized scene segmentation label, and generate and output the spatiotemporally consistent scene segmentation result based on the scene segmentation label.

[0108] In this embodiment, an energy function comprising data terms and a smoothing term is constructed to effectively balance local observations and global consistency. A Gaussian mixture model is used to estimate the probability of pixels belonging to each label, which better handles complex data distributions and improves segmentation accuracy. A weighting method based on color similarity and spatial distance is introduced, allowing the smoothing term to better preserve the structural information of the image. Applying a graph cut algorithm to minimize the energy function efficiently solves large-scale optimization problems, suitable for processing high-resolution construction site monitoring videos. By considering temporal consistency between consecutive frames, the temporal stability of the scene segmentation results is improved, effectively reducing segmentation jitter caused by factors such as illumination changes and occlusion. The optimization method based on spatiotemporal Markov random fields not only improves the accuracy of single-frame segmentation but also ensures consistency across frames, making it particularly suitable for handling dynamic and complex construction site environments. By generating spatiotemporally consistent scene segmentation results, this method provides a reliable foundation for subsequent tasks such as target tracking and behavior analysis, thereby improving the performance and stability of the entire construction site monitoring system. It significantly enhances the system's ability to understand the construction site environment, providing more accurate and coherent visual information for applications such as safety management and anomaly detection, thereby improving the level of intelligence and decision support capabilities of construction site management.

[0109] According to one aspect of this application, step S23 specifically comprises:

[0110] Step S231: Read the spatiotemporally consistent scene segmentation results and apply an improved optical flow estimation algorithm to consecutive frames;

[0111] Step S232: For each pixel, construct an energy function containing a constant brightness term and a motion smoothing term between two adjacent frames;

[0112] Step S233: Solve the optimization problem using variational method and multi-resolution strategy to obtain pixel-level motion vector field;

[0113] Step S234: Calculate the magnitude and direction histogram of the motion vector field, and use the adaptive thresholding method to segment the salient motion region;

[0114] Step S235: Apply morphological operations, including opening and closing operations, to the segmented motion salient regions to refine the region boundaries and obtain candidate dynamic regions;

[0115] Step S236: For each candidate dynamic region, construct cyclic shift samples and train a kernel correlation filter as a target tracker;

[0116] Step S237: For subsequent frame images, use the trained kernel correlation filter to calculate the response map and locate the point with the maximum response as the new target position;

[0117] Step S238: Update the tracker model according to the new location and output the dynamic region tracking result.

[0118] By constructing an energy function that includes constant brightness and motion smoothing terms, and solving it using variational methods and multi-resolution strategies, the system can effectively handle large-scale motion and occlusion problems. Calculating the amplitude and orientation histograms of the motion vector field and using an adaptive thresholding method to segment salient motion regions effectively identifies true dynamic targets and reduces interference from background motion. Employing kernel correlation filters for target tracking allows for efficient computation while adapting to changes in target appearance, making it particularly suitable for construction site environments with frequent target entry and exit and occlusion. By constructing cyclic shift samples and updating the model online, the system can adjust its tracking strategy in real time, improving the stability of long-term tracking. Combining optical flow estimation and kernel correlation filtering not only improves the accuracy of dynamic target recognition and tracking but also enhances processing efficiency, making it particularly suitable for real-time processing of large-scale construction site monitoring video data.

[0119] According to one aspect of this application, step S44 specifically comprises:

[0120] Step S441: Read each enhanced video segment and its corresponding saliency sequence; calculate the cumulative saliency curve for each video segment, which represents the cumulative change in saliency over time; analyze the inflection point of the cumulative saliency curve to identify preliminary keyframe candidate positions;

[0121] Step S442: Using a pre-trained residual network, extract deep features from each frame of the video; based on the extracted deep features, calculate the visual similarity between adjacent frames; construct a frame similarity graph, where each node represents a frame and the weight of the edges represents the degree of similarity between frames; apply a spectral clustering algorithm to the frame similarity graph to segment the video into several scenes; in each scene, select the frame with the highest saliency as the representative frame of that scene.

[0122] Step S443: Merge the candidate frames obtained based on the cumulative saliency curve and the representative frames obtained based on scene segmentation; output the merged integrated initial keyframe candidate set.

[0123] In this embodiment, calculating the cumulative saliency curve and analyzing its inflection point can effectively identify moments in the video where visual information changes significantly, providing an important basis for the initial location of keyframes. Introducing a pre-trained residual network to extract deep features can capture high-level semantic information that is difficult to identify using traditional methods, significantly improving the accuracy of keyframe selection. By calculating the visual similarity between frames and constructing a frame similarity map, the system can comprehensively understand the temporal structure of the video. Applying a spectral clustering algorithm for scene segmentation can adaptively discover semantic boundaries in the video, thereby better organizing and understanding the video content. Selecting the frame with the highest saliency in each scene as a representative ensures that the extracted keyframes comprehensively cover the main content of the video. By merging candidate frames obtained based on cumulative saliency and scene segmentation, the system can generate a comprehensive keyframe set that considers both local saliency and global structure.

[0124] According to one aspect of this application, critical incidents on construction sites often occur in operations that are not visually salient but are high-risk; conventional salience analysis cannot effectively identify violations of construction specifications; critical safety information may be obscured by dust, obstructions, or a cluttered background at the construction site. Step S44 may also be:

[0125] Receive video clips and a safety level map of the construction area, calculate the saliency sequence of the video clips, and perform regional weighting on the saliency sequence based on the safety level map to generate a safety risk weighted saliency sequence;

[0126] The video clips were analyzed using a construction site operation standard recognition model to identify the operation type and its standard compliance score, and to calculate the operation risk weighting factor.

[0127] Based on the weighted significance sequence of safety risks and the weighted factor of operational risks, a safety-oriented cumulative significance curve is generated;

[0128] Based on the safety-oriented cumulative significance curve, a multi-level threshold detection algorithm is used to identify key frames and output a key frame candidate set.

[0129] The steps for filtering the keyframe candidate set based on the construction site knowledge graph are as follows: calculate the construction semantic similarity between candidate frames; remove redundant frames with highly similar construction semantics; retain key frames that are representative in terms of construction semantics to form the final keyframe candidate set.

[0130] Specifically, in step S44a, each enhanced video segment and its corresponding saliency sequence, as well as the construction area safety level map and operational risk database, are read; based on the construction area safety level map, the saliency sequences are weighted by region to form a safety risk weighted saliency sequence S_r(t).

[0131] Step S44b: Identify standard procedures and non-standard operations in the video:

[0132] Extract the deep features f_t of each frame;

[0133] The pre-trained construction site operation specification recognition model M_op is used to classify the features, resulting in operation type O_t and specification compliance score C_t;

[0134] Frames with a specification compliance score below the threshold T_c are marked as potentially risky frames.

[0135] Calculate the operational risk weighting factor R_op(t) = β·(1-C_t)·I(O_t), where I(O_t) is the inherent risk index of the operational type, and β is an adjustable parameter;

[0136] Step S44c: Generate the safety-oriented cumulative significance curve CSC_safety(t) = ∑_{i=1}^t [α·S_r(i) + (1-α)·R_op(i)]; where α is a balance parameter between significance and operational risk, which is dynamically adjusted according to the engineering stage;

[0137] Step S44d: Identify keyframes using a multi-level threshold detection algorithm.

[0138] Apply an adaptive three-layer threshold {T_low, T_med, T_high} to the CSC_safety curve;

[0139] For high-risk operation areas (critical parts of the project), the T_low threshold is used;

[0140] For medium-risk areas, the T_med threshold is used;

[0141] For low-risk areas, the T_high threshold is used;

[0142] Identify inflection points that meet the corresponding thresholds as keyframe candidates;

[0143] Step S44e: Using a redundant frame filtering method based on construction site knowledge graph, calculate the construction semantic similarity between candidate frames, retain key frames that are representative in construction semantics, and output the final key frame candidate set.

[0144] The salience of different areas on the construction site is weighted according to their safety levels to ensure that high-risk areas receive sufficient attention even if their visual salience is not high. A database of standard operating procedures (SOPs) for the construction site is established to assess the compliance of operational behaviors with SOPs in real time and identify potential violations, even minor ones. Thresholds are dynamically adjusted based on the risk level of each area, employing more sensitive detection strategies for critical processes and high-risk areas.

[0145] According to one aspect of this application, this method addresses issues such as low image contrast and loss of detail caused by dust pollution; uneven exposure due to alternating strong light and shadow; and information loss due to insufficient nighttime construction lighting. Traditional image enhancement methods often employ global or fixed-parameter processing strategies, which cannot adapt to the complex changes in the construction site environment, resulting in unsatisfactory enhancement effects. Video enhancement processing can also include:

[0146] Step S11a: Obtain raw video data and simultaneously read environmental state parameter data at the corresponding time, including light intensity L_env, dust concentration D_env, humidity H_env, and weather conditions W_env;

[0147] Step S11b: Construct an adaptive enhancement parameter set based on environmental state parameters:

[0148] Calculate the Environment Complexity Index (ECI) = f(L_env, D_env, H_env, W_env);

[0149] Determine the number of wavelet decomposition levels N_level = min(4, 2 + ⌊ECI / 0.3⌋);

[0150] Calculate the enhancement intensity factor γ = 1.0 + 0.5·D_env + 0.3·(1-L_env / L_max);

[0151] Set the defogging intensity δ = min(0.8, D_env·1.2);

[0152] Step S11c: For each video frame I(x,y), apply environment-adaptive multi-scale enhancement:

[0153] Apply N-level wavelet decomposition: [LL_N, {LH_i, HL_i, HH_i}_{i=1}^N] = WaveletDecomp(I, N_level);

[0154] Apply adaptive enhancement strategies to different frequency subbands:

[0155] Low-frequency subband (LL_N): Apply adaptive contrast stretching, F_LL(x,y) = g(ECI)·LL_N(x,y) +(1-g(ECI))·μ(x,y);

[0156] Where g(ECI) = 1.2 + 0.3·ECI, and μ(x,y) is the local mean;

[0157] Intermediate frequency sub-bands (LH_{N-1}, HL_{N-1}): Adaptive dust enhancement is applied.

[0158] F_MF(x,y) = MF(x,y)·(1 + α(D_env)·(σ(x,y) / σ_max));

[0159] where α(D_env) = 1.5 + D_env·0.8;

[0160] High-frequency subbands (HH_i, i < N - 1): Apply detail-preserving enhancement,

[0161] F_HF(x,y) = HF(x,y)·(1 + β(L_env)·(1 - exp(-σ(x,y)² / k)));

[0162] where β(L_env) is an adaptive parameter based on the lighting condition;

[0163] Step S11d: For the detected shadow region S, apply shadow adaptive brightening:

[0164] Identify the shadow region S = {(x,y) | I(x,y) < T_shadow(L_env)};

[0165] Apply adaptive gamma correction to the shadow region: I_S(x,y) = I(x,y)^(1 / γ_S);

[0166] where γ_S = 1.0 + 0.8·(1 - L_local / L_max), and L_local is the local illumination;

[0167] Step S11e: For a dusty environment, apply the deep learning dehazing network DeHazeNet:

[0168] Estimate the dust transmission map t(x,y) = DeHazeNet(I, D_env);

[0169] Apply adaptive dehazing intensity: I_clear(x,y) = (I(x,y) A·(1 - δ·t(x,y))) / max(t(x,y), t_min); where A is the estimated value of atmospheric light and t_min is the minimum transmittance threshold;

[0170] Step S11f: Fuse the enhancement results of each layer and adopt weighted reconstruction:

[0171] I_enhanced = WaveletRecon([F_LL, {w_i·F_LH_i, w_i·F_HL_i, w_i·F_HH_i}]);

[0172] where the weights w_i are adaptively adjusted based on the environmental state parameters and frequency characteristics;

[0173] Step S11g: Apply global tone mapping to the reconstructed image to ensure visual consistency, and stitch the enhanced frames to form an enhanced video stream.

[0174] Enhancement parameters and strategies are dynamically adjusted based on real-time environmental parameters (light intensity, dust concentration, humidity, etc.). The number of wavelet decomposition levels is automatically determined according to environmental complexity, and targeted enhancements are applied to different frequency sub-bands. Specific enhancements are performed on shaded areas, dusty areas, etc., to address the challenges of construction site environments.

[0175] According to another aspect of this application, a smart construction site management system based on photoelectric and video fences is also provided, including a control module and an information acquisition device connected to the control module. The information acquisition device includes a camera device, a cable monitoring device, an environmental status parameter sensor, and a smart switch for construction equipment. The cable monitoring device is installed at predetermined lengths along the direction of cable extension. The control module includes:

[0176] At least one processor; and,

[0177] A memory communicatively connected to at least one of the processors; wherein,

[0178] The memory stores instructions that can be executed by the processor to implement the intelligent construction site management method based on photoelectric and video fences as described in any of the above technical solutions.

[0179] Implementation Case: The intelligent construction site management system includes a control module and connected information acquisition devices. The control module consists of an industrial-grade server equipped with an Intel Xeon E5-2680 v4 processor, 128GB RAM, and 8TB SSD storage, running a Linux operating system. The information acquisition devices include:

[0180] Sixteen high-definition network cameras (model: Hikvision DS-2CD5126G0-IZS) were installed around the construction site perimeter and key areas. These cameras have a resolution of 2688×1520, a frame rate of 30fps, a wide dynamic range of 120dB, and support H.265 encoding. The cameras are connected to the control module via Gigabit Ethernet for real-time video data transmission.

[0181] Along the cable's extension direction, a set of cable monitoring devices is installed every 50 meters, totaling 32 sets. Each set includes: a cable fault tester (model: ETCR9500C), with an accuracy of ±2% and a measurement range of 0-600V; a fault sensing probe (model: FD-380A), with a sensitivity ≤0.1mA; a positioning signal sensor (model: DS-500), with a positioning accuracy of ±0.5m; a path signal sensor (model: RS-200), with a maximum detection depth of 3m; a high-energy impact signal generator (model: HSG-800), with a maximum output voltage of 8kV; and a path signal generator (model: PSG-300), with an output frequency of 512Hz / 1kHz / 8kHz / 33kHz selectable. These devices are connected to a data acquisition unit (model: DAQ-2000) via an RS-485 bus, which in turn communicates with the control module via a 4G industrial-grade router to achieve real-time data transmission.

[0182] Multiple environmental sensors were deployed at key locations on the construction site: 12 sets of temperature and humidity sensors (model: THD-100), with an accuracy of ±0.5℃ for temperature and ±3%RH for humidity; and 8 sets of dust concentration sensors (model: PM-2000), with a detection range of 0-1000μg / m³. 3 Accuracy ±10μg / m 3 Noise detector (model: SLM-200), 6 groups, detection range 30-130dB, accuracy ±1.5dB; Gas detector (model: MGD-400), 4 groups, can detect CO, H2S, O2 and combustible gases.

[0183] Each sensor transmits data to the LoRa gateway deployed on the construction site via the LoRaWAN network (operating frequency 470MHz), and then transmits the data to the control module via a wired network.

[0184] 24 sets of smart switches (model: IS-800) were installed on the main construction equipment (including tower cranes, concrete pump trucks, excavators, etc.). Each smart switch has the following functions: equipment start-up and shutdown status monitoring; running time statistics; power monitoring (range 0-100kW, accuracy ±0.5%); overload protection; leakage current detection (sensitivity 10mA); the smart switches realize real-time data transmission through the NB-IoT network, with a communication frequency of once every 30 seconds.

[0185] This system performs data acquisition and preprocessing operations every predetermined period T (T is set to 10 seconds), specifically as follows:

[0186] After acquiring the raw video data from the camera device, discrete wavelet transform is used for enhancement processing. The specific steps are as follows:

[0187] Applying a two-dimensional discrete wavelet transform (DWT) to each video frame I(x,y) yields four frequency sub-bands: LL (low frequency), LH (horizontal high frequency), HL (vertical high frequency), and HH (diagonal high frequency): [LL, LH, HL, HH] = DWT(I(x,y));

[0188] For each frequency sub-band, calculate the local statistical properties: use a 5×5 sliding window W to calculate the local mean μ and standard deviation σ: μ(i,j) = (1 / 25)∑(x,y)∈WI(x,y); σ(i,j) = sqrt((1 / 25)∑(x,y)∈W (I(x,y) μ(i,j)) 2 );

[0189] Construct an adaptive enhancement function F, employing different enhancement strategies for different subbands:

[0190] Apply contrast stretching to the LL subband: F_LL(x,y) = g·LL(x,y) + (1-g)·μ(x,y); where g is the contrast gain factor, set to 1.2;

[0191] Nonlinear enhancement is applied to the LH, HL, and HH subbands: F_HF(x,y) = HF(x,y)·(1 + α·(σ(x,y) / σ) M ax));

[0192] Where HF represents the high-frequency subband (LH, HL, or HH), α is the enhancement factor (set to 1.5), and σ Max The maximum standard deviation of the entire subband;

[0193] The enhanced image is reconstructed using a weighted summation method: I_enhanced(x,y) = IDWT(w_LL·F_LL, w_LH·F_LH, w_HL·F_HL, w_HH·F_HH); where IDWT is the inverse discrete wavelet transform, and w_LL=0.6, w_LH=1.2, w_HL=1.2, w_HH=1.0 are the weights of each sub-band.

[0194] Unlike traditional global histogram equalization, this method considers local image characteristics and frequency domain information, enabling it to improve contrast while suppressing noise amplification. This makes it particularly suitable for construction site environments with significant lighting variations and interference. By adaptively adjusting the enhancement level of each frequency band, the system can improve overall clarity while preserving details, providing high-quality visual input for subsequent scene analysis.

[0195] The raw cable data acquired by the cable monitoring device includes: voltage V(t), current I(t), resistance R(t), temperature T(t), and power factor PF(t). The preprocessing steps are as follows: Data smoothing and denoising: A Savitzky-Golay filter is applied to each parameter to remove high-frequency noise: X_smooth(t) = SG_filter(X(t), window_size=11,polynomial_order=3); where X represents any cable parameter.

[0196] Outlier detection and handling: A modified Z-score method is used to identify outliers: Z(t) = |X_smooth(t)median(X_smooth)| / MAD(X_smooth); where MAD is the absolute deviation of the median. If Z(t) > 3.5, X_smooth(t) is replaced with the local median.

[0197] Feature extraction: Calculate the statistical characteristics of each cable parameter, including mean, standard deviation, maximum value, minimum value, peak factor, and skewness.

[0198] Parameter correlation analysis: Calculate the cross-correlation matrix C between the parameters, where C ij This represents the Pearson correlation coefficient between parameters i and j.

[0199] By combining statistical feature extraction and parameter correlation analysis, not only is noise eliminated, but the weak signal characteristics of potential faults are also preserved. Compared with traditional methods that rely solely on threshold detection, this approach can more comprehensively capture the intrinsic relationships between cable parameters.

[0200] The initial environmental state parameters obtained from the environmental state parameter sensors include: temperature Te(t), humidity H(t), PM2.5 concentration P(t), noise intensity N(t), and gas concentration G(t). The preprocessing steps are as follows:

[0201] Standardization: Transforming various environmental state parameters into a unified scale X_norm(t) = (X(t) X M in) / (X M axX M in);

[0202] Weight calculation: Based on the degree of influence of the parameters on construction safety, assign a weight vector W = [w_Te, w_H, w_P, w_N, w_G], where w_Te=0.15, w_H=0.10, w_P=0.20, w_N=0.15, w_G=0.40;

[0203] Environmental Risk Index (ERI) calculation: ERI(t) = ∑(wi × X_norm,i(t));

[0204] Environmental state model construction: Based on the environmental risk index ERI(t), the site environmental state is divided into three categories: safe (S), warning (W), and dangerous (D). The following is achieved using a Bayesian classifier: P(S|ERI) = p(ERI|S)·p(S) / p(ERI); P(W|ERI) = p(ERI|W)·p(W) / p(ERI); P(D|ERI) = p(ERI|D)·p(D) / p(ERI); Environmental state = argmax{P(S|ERI), P(W|ERI), P(D|ERI)}.

[0205] A risk-based environmental status assessment model is introduced, comprehensively considering the impact of multiple environmental factors on construction site safety. Compared with traditional methods that process each environmental status parameter independently, this model provides a more holistic environmental risk assessment and is closely integrated with the subsequent anomaly detection module, significantly improving the system's ability to identify anomalies caused by environmental factors.

[0206] The operating parameters collected from the intelligent switches of the construction equipment include: switch status St(t), running time RT(t), power P(t), load ratio LR(t), and leakage current LC(t). The preprocessing steps are as follows: Outlier detection: Outliers are detected using the Local Outlier Factor (LOF) method; Equipment usage pattern recognition: The k-means clustering algorithm is applied to classify equipment usage patterns into four categories: no-load, normal operation, full load, and overload; Equipment status risk assessment: Combining the switch parameters and equipment usage patterns, the equipment status risk coefficient DSR(t) = α·St(t) + β·f(RT(t)) + γ·g(P(t)) + δ·h(LR(t)) + ε·j(LC(t)) is calculated; where α=0.1, β=0.2, γ=0.25, δ=0.25, and ε=0.2 are weighting coefficients, and f, g, h, and j are risk mapping functions for each parameter.

[0207] By combining equipment parameter monitoring and usage pattern recognition, the system can more accurately assess equipment status risks. By establishing a Data Risk Rating (DSR) coefficient, the system can quantify the degree of equipment anomalies, providing crucial information for subsequent anomaly detection.

[0208] To address the issue of inconsistent sampling frequencies across multiple data sources, a multidimensional dynamic time warping (MD-DTW) algorithm is employed for time alignment: each data source is represented as a time series matrix: X = [x1, x2, ..., x...]. n ], where x nGiven a data vector at time point n; construct a multidimensional distance matrix: D(i,j) = ||X i Y j || 2 X and Y are two data sources to be aligned; calculate the optimal alignment path: use dynamic programming to solve for the minimum cumulative distance matrix; perform time resampling based on the optimal path to align each data source to a unified time standard.

[0209] Compared to traditional linear interpolation methods, this approach better handles the nonlinear temporal distortion of multi-source data, making it particularly suitable for processing multimodal construction site data with varying sampling frequencies and delays. Through precise time alignment, the system can accurately capture cross-modal anomalous patterns, laying the foundation for subsequent multi-source data fusion analysis.

[0210] Superpixel segmentation is performed on the preprocessed enhanced video, as follows:

[0211] The SLIC (Simple Linear Iterative Clustering) algorithm is used for initial superpixel segmentation: the image is divided into a grid of approximately size s×s (s=20 pixels), with the center of each grid serving as the superpixel center; the distance metric D = sqrt((l_p-l_k)) between each pixel p and its surrounding superpixel centers k is calculated. 2 + (a_p-a_k) 2 + (b_p-b_k) 2 ) / m 2 + sqrt((x_p-x_k) 2 + (y_p-y_k) 2 ) / s 2 ; where (l,a,b) are CIELAB color space coordinates, (x,y) are pixel positions, and m is the balance parameter between color distance and spatial distance (set to 10); iteratively update the superpixel center and pixel affiliation until convergence (set the maximum number of iterations to 10).

[0212] Features are extracted for each superpixel region:

[0213] Color characteristics: Calculate the histogram H_color (16-dimensional vector) of the Lab color space;

[0214] Texture features: Employed the SURF (Speeded-Up Robust Features) algorithm;

[0215] Calculate the Haar wavelet responses dx and dy (scales set to 1.2, 2.4, 3.6, 4.8); construct the directional gradient histogram H_grad (8-dimensional vector); calculate the gradient variances σ_dx, σ_dy, and covariance σ_dxdy.

[0216] Shape features: Calculate the shape descriptor F_shape, including area ratio, contour complexity, orientation, compactness, etc. (6-dimensional vector);

[0217] Construct a comprehensive feature vector F = [w_c·H_color, w_t·H_grad, w_t·[σ_dx, σ_dy, σ_dxdy], w_s·F_shape]; where w_c=0.4, w_t=0.4, and w_s=0.2 are the weights of each feature;

[0218] Compared to the original SURF, this study incorporates multi-scale Haar wavelet response analysis, which more effectively captures texture features in construction site environments. The comprehensive feature vector F integrates color, texture, and shape information, providing a more comprehensive region description and richer visual features for subsequent scene segmentation.

[0219] Hierarchical clustering based on superpixel feature vectors enables preliminary scene segmentation:

[0220] Calculate the distance matrix M(i,j) = d(F) between eigenvectors. i , F j = sqrt(∑(w_k·(F i ,k F j ,k) 2 )); where w_k is the weight of the k-th dimension feature;

[0221] The hierarchical tree is constructed using Ward's minimum variance method:

[0222] Initialization: Each superpixel is treated as an independent cluster;

[0223] Iteration: The two clusters with the smallest increase in variance after each merge are merged until a preset number of clusters or a variance threshold is reached;

[0224] The hierarchical tree is cut according to the preset threshold τ_cluster (set to 0.3) to obtain the initial scene segmentation result.

[0225] Compared to traditional k-means, it does not require pre-specifying the number of clusters, can adaptively discover the hierarchical structure of the data, and is more suitable for handling multi-class region segmentation problems in complex construction site scenarios. By adjusting the cutting threshold τ_cluster, the granularity of segmentation can be flexibly controlled to adapt to different analytical needs.

[0226] To improve the spatiotemporal consistency of scene segmentation, a spatiotemporal Markov random field (ST-MRF) model is used for optimization:

[0227] Construct the energy function E(L) = ∑E_data(L) i ) + λ_s·∑E_smooth(L i , L j ) + λ_t·∑E_temp(L i ,t, L i ,t-1); where L is the label configuration, E_data is the data item, E_smooth is the spatial smoothing term, E_temp is the temporal consistency term, and λ_s=0.6 and λ_t=0.4 are the weight coefficients;

[0228] Data item calculation: The probability E_data(L) of a pixel belonging to each label is estimated using a Gaussian mixture model (GMM). i ) =-log(P(I i |L i )); where P(I i |L i The values ​​are provided by GMM, and each label uses 3 Gaussian components;

[0229] Spatial smoothing term calculation: E_smooth(L i , L j ) = exp(-β·||I i I j || 2 )·δ(L i ≠ L j ); where β is the color difference coefficient (set to 0.1), and δ(·) is the indicator function;

[0230] Time consistency term calculation: E_temp(L i ,t, L i ,t-1) = w_t·δ(L i ,t ≠ L i ,t-1); where w_t is the time penalty weight, which is adaptively adjusted according to the inter-frame difference.

[0231] The energy function is minimized using the α-expansion graph cut algorithm to obtain the optimized scene segmentation result.

[0232] By unifying spatial and temporal consistency constraints into an energy optimization framework, this method overcomes the temporal instability of traditional scene segmentation methods. By considering spatiotemporal relationships, it effectively reduces segmentation jitter caused by factors such as illumination variations and camera shake, improving the stability and consistency of segmentation results and providing a reliable foundation for subsequent dynamic region recognition.

[0233] Based on spatiotemporally consistent scene segmentation results, an improved optical flow estimation method is used to identify dynamic regions:

[0234] The Farneback algorithm is used to calculate the pixel-level dense optical flow field: for each pixel (x,y), the displacement vector (u,v) is estimated between adjacent frames; the energy function E(u,v) = ∑(I_1(x,y) I_2(x+u,y+v)) is constructed. 2 + λ·(||grad u|| 2 + ||grad v|| 2 ); where I_1 and I_2 are adjacent frames, and λ is the weight of the smoothing term (set to 0.05);

[0235] The energy function is solved using the variational method and a multi-resolution strategy to obtain the optical flow field (u,v);

[0236] Calculate the optical flow amplitude and direction histogram: Optical flow amplitude M(x,y) = sqrt(u(x,y) 2 + v(x,y) 2 Optical flow direction θ(x,y) = arctan(v(x,y) / u(x,y)); Construct an 8-direction histogram H_flow;

[0237] Adaptive thresholding for motion-salient regions: Calculating the global mean μ of optical flow amplitude. M and standard deviation σ M Determine the adaptive threshold T Motion = μ M + k·σ M Where k is a coefficient (let's say 2.0); if M(x,y)>T Motion If so, then mark it as a significant point of motion;

[0238] Morphological processing optimizes dynamic regions: Opening operations are applied to eliminate noise (structuring element is a 3×3 rectangle); Closing operations are applied to fill holes (structuring element is a 5×5 rectangle); Connectivity analysis is performed to remove regions with an area smaller than the threshold T_area (100 pixels).

[0239] Building upon the traditional Farneback algorithm, a multi-resolution strategy and variational optimization framework are added, enabling more accurate handling of large displacements and deformations. An adaptive thresholding mechanism dynamically adjusts segmentation parameters based on optical flow statistics, adapting to changes in motion patterns across different construction site scenarios. Morphological post-processing further optimizes the boundaries and connectivity of dynamic regions, improving the stability of recognition results and providing high-quality initial regions for subsequent target tracking.

[0240] For the identified dynamic region, the kernel correlation filter (KCF) algorithm is used for target tracking:

[0241] Target model initialization: Extract HOG features x from the dynamic region; construct the ideal response output y (Gaussian shape, with 1 at the center and 0 at the periphery);

[0242] Calculate the filter w = F in the Fourier domain. - ¹(F(y) OF(x)* / (F(x) OF(x)* + λ)); where F is the Fourier transform, O is the element-wise product, * is the conjugate, and λ is the regularization parameter (set to 0.001);

[0243] Target location update: Extract HOG features z from the search region in subsequent frames;

[0244] Calculate the response plot r = F -1 (F(w) OF(z)); Find the point with the maximum response as the new target location;

[0245] Model update: Extract HOG features at new locations x new Update the filter w using the learning rate η. new = (1-η)·w +η·w new Where η is set to 0.02, w new For x new The new filter is calculated;

[0246] Assign a unique ID to each tracked target and record its trajectory information;

[0247] Featuring high computational efficiency and robustness, it is particularly suitable for the real-time tracking of multiple targets in construction site environments. By performing rapid calculations in the Fourier domain, it can effectively handle changes in target appearance and partial occlusion, providing technical assurance for long-term stable tracking. The cyclic shifting sample strategy and online model update mechanism enable the tracker to adapt to gradual changes in target appearance, improving the system's tracking performance in complex construction site environments.

[0248] Empirical Mode Decomposition (EMD) is performed on cable operating data to extract intrinsic mode functions (IMFs); all local extrema in the cable time series X(t) are identified; and the upper envelope e is generated using cubic spline interpolation. Max (t) and lower envelope e Min (t); Calculate the mean envelope m(t) = (e Max (t) + e Min (t)) / 2; Calculate the difference h(t) = X(t) m(t); Check if h(t) satisfies the IMF condition: the difference between the number of extrema and the number of zero crossings does not exceed 1; the local mean is approximately zero; if satisfied, then h(t) is an IMF component; Calculate the residual signal r(t) = X(t) h(t), and repeat steps 1-5 with r(t) as a new signal until r(t) becomes a monotonic function; Empirical mode decomposition can separate different frequency components in X(t) X(t) = ∑c i (t) + r_n(t); where c i (t) represents the IMF component, and r_n(t) represents the residual trend;

[0249] Unlike traditional Fourier analysis, which can handle nonlinear and non-stationary signals, this method is particularly suitable for analyzing complex time-varying signals such as cable parameters. By adaptively decomposing the signal into multiple intrinsic mode functions, this method can effectively extract cable characteristics at different time scales, providing fine-grained frequency information for anomaly pattern recognition.

[0250] Perform a Hilbert transform on each IMF component and analyze its instantaneous characteristics:

[0251] For each IMF component c i (t) Calculate the Hilbert transform H[c i [(t)] = (1 / π)∫(c i (τ) / (t-τ))dτ;

[0252] Constructing the analytic signal z i (t) = c i (t) + j·H[c i [(t)] = a i (t)·exp(j·φ i (t)); where a i (t) represents the instantaneous amplitude, φ i (t) represents the instantaneous phase;

[0253] Calculate the instantaneous frequency ω i (t) = dφ i (t) / dt; Construct the Hilbert spectrum HS(ω,t) = ∑a i (t)·δ(ω-ωi (t)); where δ is the Dirac function.

[0254] By combining EMD decomposition and Hilbert transform, this method can provide a time-frequency distribution map of the signal, intuitively displaying the frequency changes of cable parameters over time. By calculating the instantaneous frequency and amplitude, this method can accurately capture transient changes in cable parameters, making it particularly suitable for detecting instantaneous faults such as broken cores, and significantly improving the system's ability to identify potential faults early.

[0255] Based on the Hilbert spectral analysis results, a multidimensional phase space was constructed and density clustering was applied:

[0256] For each cable parameter (voltage, current, resistance, etc.), select characteristic frequency components, typically the instantaneous frequency and amplitude corresponding to the first three IMF components; construct a multidimensional phase space P, where each dimension represents a characteristic component:

[0257] P = [ω_V1, a_V1, ω_V2, a_V2, ω_V3, a_V3, ω i1 , a i1 , ..., ω_R3, a_R3];

[0258] Where ω_Vn and a_Vn are the instantaneous frequency and amplitude of the nth IMF component of the voltage, respectively.

[0259] Perform DBSCAN density clustering in P-space: Set parameters: ε = 0.15 (neighborhood radius), MinPts = 5 (minimum number of points); For each point p, calculate the number of points in its ε-neighborhood N_ε(p); If |N_ε(p)| ≥ MinPts, then p is a core point; Recursively add all density-reachable points to the same cluster; Unassigned points are marked as noise points (potential outliers);

[0260] Analyze the characteristics of anomalies: calculate the Mahalanobis distance between anomalies and each normal cluster; extract the original cable parameter values ​​corresponding to anomalies; and determine the anomaly type by matching against a pre-set anomaly pattern template library.

[0261] By mapping the time-frequency characteristics of cable parameters into a high-dimensional space, normal operating modes form dense clusters, while abnormal modes manifest as outliers. Compared to traditional univariate threshold detection, this method can capture complex correlations between parameters, significantly improving the sensitivity and accuracy of anomaly detection. Especially for hidden faults such as broken cores, which are difficult to detect using traditional monitoring, this method can issue early warnings by analyzing the abnormal trajectories of cable parameters in phase space, providing crucial protection for construction site safety management.

[0262] To further improve the stability of anomaly detection, the sliding window singular spectral analysis (SSA) method is adopted:

[0263] Construct a multidimensional cable parameter time series matrix X, where each row represents a parameter and each column represents a time point;

[0264] For a sliding window W of length L, construct the trajectory matrix A = [X... t X t+1 , ..., X t+K- 1]; where K = W - L + 1, X t Let be the parameter vector for time t;

[0265] Perform singular value decomposition (SVD) on the trajectory matrix A: A = UΣV^T; where Σ is the singular value diagonal matrix, and U and V are orthogonal matrices; reconstruct the signal by selecting the components corresponding to the first N largest singular values: A* = ∑ i=1 N σ i u i v i T ; where σ i For the i-th singular value, u i and v i Let X be the corresponding left and right singular vectors; calculate the reconstruction error e_t = ||X_t – A*_t||2; calculate the adaptive threshold T_t = μ_e + γ·σ_e; where μ_e and σ_e are the moving average and standard deviation of the reconstruction error, respectively, and γ is the sensitivity parameter (set to 3.0); if e_t>T_t, then mark it as a potential outlier.

[0266] It can capture the main patterns and structural changes in multivariate time series. By extracting the main components of the signal and monitoring reconstruction errors, it can effectively detect abnormal changes in cable parameters, maintaining high sensitivity even in the presence of noise and interference. The adaptive threshold mechanism dynamically adjusts the judgment criteria based on the statistical characteristics of the signal, enabling the system to adapt to the cable operating conditions under different circumstances, reducing false alarm rates while improving detection rates.

[0267] For potential outliers detected by the SSA method, local anomaly factor (LOF) analysis was further applied for validation:

[0268] For each potential outlier point p, calculate its k-nearest neighbor distance d_k(p), where k is set to 5;

[0269] Calculate the reachability distance reach-dist_k(p,o) = max{d_k(o), dist(p,o)}; where o is a neighbor of p, and dist(p,o) is the Euclidean distance between the two points; calculate the local reachability density lrd_k(p) = 1 / (∑_{o∈N_k(p)}reach-dist_k(p,o) / |N_k(p)|); where N_k(p) is the set of k nearest neighbors of p.

[0270] Calculate the local anomaly factor LOF_k(p) = ∑_{o∈N_k(p)} (lrd_k(o) / lrd_k(p)) / |N_k(p)|; if LOF_k(p)>T_LOF (threshold set to 1.5), then it is confirmed as an anomaly.

[0271] By comparing the local density of data points with their neighbors, local outliers in areas with different densities can be effectively detected. Compared with global anomaly detection methods, LOF is particularly suitable for processing cable parameter data with multi-mode distributions, and can identify anomaly patterns that deviate significantly in local areas. By combining LOF with the SSA method, the system implements a two-level verification mechanism, significantly reducing the false alarm rate while maintaining sensitivity to minor anomalies, providing reliable assurance for cable safety monitoring at construction sites.

[0272] To further understand the causes of the anomalies, cluster analysis was performed on the detected outliers:

[0273] An improved k-means++ algorithm is applied to cluster outliers: k initial centroids are selected using a distance-weighted probability selection strategy; standard k-means iterative optimization is applied; and the optimal number of clusters is automatically determined using silhouette coefficients.

[0274] Feature importance was analyzed using a random forest decision tree model: a training set was constructed, consisting of normal samples (label 0) and various abnormal samples (labels 1, 2, ...); a random forest model (100 trees, maximum depth 12) was trained; and feature importance scores based on Gini impurity were calculated.

[0275] Descriptive labels are generated for each anomaly cluster based on feature importance, such as: "Voltage transient anomaly - mainly affects the IMF2 frequency component"; "Current-temperature coupling anomaly - related to ambient temperature";

[0276] The similarity between anomalous and normal patterns is calculated using the Dynamic Time Warping (DTW) algorithm: DTW(X,Y) = min{∑w i ·d(x_{i1},y_{i2})};where w id represents the path weight, and d is the distance function; based on DTW distance, a severity level (1-5) is assigned to each anomaly cluster.

[0277] This system not only detects anomalies but also automatically classifies and explains their causes. Through random forest feature importance analysis, it identifies key parameters and frequency components leading to anomalies, providing site managers with intuitive and understandable descriptions of the anomalies. The dynamic time warping algorithm quantifies the difference between anomalies and normal patterns, providing an objective basis for anomaly severity assessment. This intelligent anomaly interpretation mechanism significantly improves the system's interpretability and usability, enabling even non-professionals to understand complex cable anomaly patterns and providing decision support for timely and targeted maintenance measures.

[0278] Integrate cable anomaly detection results with equipment status and environmental status parameters:

[0279] Construct an anomaly feature vector F_anomaly = [E_cable, DSR, ERI, Loc, T]; where E_cable is the cable anomaly parameter, DSR is the equipment status risk, ERI is the environmental risk index, Loc is the location information, and T is the timestamp;

[0280] Calculate the correlation matrix C for each anomaly source, using the Spearman rank correlation coefficient: C ij = ρ(F i , F j =cov(rg i , rg j ) / (σ_{rg i}·σ_{rg j}); where rg i Let i be the rank sequence of feature i. Anomaly correlation analysis is performed based on the correlation matrix to identify potential causal relationships.

[0281] By combining cable parameters, equipment status, and environmental factors, a comprehensive anomaly feature vector is constructed. By analyzing the correlations and causal relationships between anomaly features, the system can more accurately understand the root causes and propagation paths of anomalies, providing a multi-dimensional information foundation for anomaly severity assessment. For example, the system can distinguish between cable parameter fluctuations caused by changes in ambient temperature and genuine cable faults, significantly reducing the false alarm rate.

[0282] A fuzzy inference system is used to assess the severity of anomalies.

[0283] Define the following input fuzzy variables: Cable anomaly level (E_cable): {low, medium, high}; Equipment status risk (DSR): {safe, warning, dangerous}; Environmental risk index (ERI): {normal, warning, severe}; Duration: {short, medium, long}.

[0284] Define the output fuzzy variable: Severity: {Slight, Moderate, Severe, Critical, Catastrophic};

[0285] Establish a fuzzy rule base, including 25 rules, for example: if (E_cable=high) and (DSR=dangerous) and (ERI=severe) and (Duration=long), then (Severity=catastrophic); if (E_cable=medium) and (DSR=warning) and (ERI=normal) and (Duration=short), then (Severity=moderate).

[0286] Applying Mamdani fuzzy inference: Mapping the input to a fuzzy set; applying fuzzy operators (min for AND, max for OR); merging the outputs of all rules; calculating the output value using the centroid method; comparing the defuzzified result with a threshold to determine the anomaly level.

[0287] It can handle the uncertainty and ambiguity in anomaly assessment, and is more in line with the decision-making thinking of human experts. Unlike traditional hard threshold judgment, fuzzy reasoning, by simulating the human reasoning process, can comprehensively consider the combined impact of multiple factors on the severity of anomalies, and make a more reasonable assessment. Especially in complex and ever-changing environments such as construction sites, fuzzy reasoning systems have shown better adaptability and robustness, providing more accurate prioritization judgments for anomaly handling, helping managers to rationally allocate resources and prioritize the handling of high-risk anomalies.

[0288] For anomalies assessed as severe, perform precise spatiotemporal localization: record the exact timestamp T of the anomaly's occurrence. anomaly Based on the location information of the cable monitoring device, determine the preliminary spatial coordinates (x, y, z) of the anomaly. initial ;

[0289] Precise positioning process: A high-precision scan is performed within ±5 meters of the initial coordinates; the position is optimized by applying the principle of triangulation and combining multi-point signal strength; the final anomaly coordinates (x, y, z) are calculated. final The accuracy is improved to ±0.5 meters; the abnormal coordinates are mapped onto the 3D model of the construction site to generate visualized abnormal points.

[0290] Combining timestamp recording and spatial coordinate positioning, it can accurately determine the time and location of anomalies. High-precision scanning and multi-point triangulation optimization significantly improve positioning accuracy, making it particularly suitable for locating hidden fault points in cables. The visualization function, which maps anomaly points to a 3D site model, intuitively displays the spatial distribution of anomalies, providing strong support for maintenance personnel to quickly locate faults and significantly improving troubleshooting efficiency.

[0291] Extract relevant segments from stored video data based on abnormal timestamps:

[0292] Extraction time range: [T anomaly Δt before , T anomaly + Δt after ]; where Δt before =30 seconds, Δt after =60 seconds;

[0293] Adaptive contrast enhancement is performed on the extracted video segments, including: calculating the luminance histogram of each frame of the video segment;

[0294] The CLAHE (Contrast-Limited Adaptive Histogram Equalization) algorithm is applied to the image, including: dividing the image into an 8×8 grid; performing histogram equalization independently on each grid, with a contrast limit threshold set to 3.0; and merging grid boundaries using bilinear interpolation.

[0295] Video stabilization algorithms are applied to reduce jitter, including: calculating feature point matching between consecutive frames; estimating the global motion model (affine transformation); applying an exponential smoothing filter to smooth the transformation parameters; and re-transforming each frame based on the smoothed parameters.

[0296] Compared to traditional global histogram equalization, this method preserves local details and edge information while suppressing noise amplification, making it particularly suitable for processing videos with uneven lighting in construction site environments. The video stabilization algorithm improves video quality and the accuracy of subsequent analysis by smoothing camera shake. This video enhancement method provides high-quality input data for keyframe extraction and target recognition, significantly improving the system's visual analysis capabilities in harsh environments.

[0297] Based on the enhanced video clips, motion intensity maps are calculated and visual saliency is analyzed:

[0298] Calculate the inter-frame motion vector field: The Farneback optical flow algorithm is used to calculate the pixel-level motion vector (u,v); calculate the motion amplitude matrix M: M(x,y) = sqrt(u(x,y)). 2 + v(x,y) 2 );

[0299] Generate motion intensity map MI: For each pixel position (x,y), calculate the temporal cumulative motion intensity MI(x,y) =∑_t w_t·M_t(x,y); where w_t is the temporal weight, with recent frames having a higher weight;

[0300] Computational visual saliency map VS:

[0301] Calculate multi-feature saliency using the Itti-Koch model: extract brightness, color, and orientation features; construct a 9-scale Gaussian pyramid; calculate center-periphery differences; perform cross-scale combination and normalization;

[0302] By fusing motion intensity and visual saliency, a multi-feature fusion saliency map FS is obtained: FS = α·norm(MI) + (1-α)·norm(VS); where α=0.7 is the motion saliency weight and norm is the normalization function.

[0303] By integrating dynamic information and static visual features, it can comprehensively capture important content in videos. Compared with traditional single-feature saliency analysis, it can more accurately identify key targets and events in the construction site environment, maintaining high performance even under complex backgrounds and varying lighting conditions. By generating high-quality saliency maps, the system provides a reliable basis for subsequent keyframe selection, improving the efficiency and accuracy of video analysis.

[0304] Based on the saliency sequence, calculate temporal saliency and select keyframes:

[0305] Calculate the cumulative significance curve CSC(t) = ∑ i=1 t mean(FS i ); where mean(FS) i ) represents the average value of the saliency map of the i-th frame.

[0306] Analyze the inflection points of the CSC curve and use the second derivative method to detect points of significant change: d 2 CSC(t) / dt 2 = (CSC(t+1)2·CSC(t) + CSC(t-1)); if |d 2 CSC(t) / dt 2 |>T i If nflection (threshold set to 0.05) is used, then t is the inflection point, and the corresponding frame is the preliminary candidate frame.

[0307] Deep Feature Extraction and Scene Segmentation: A pre-trained ResNet-50 network is used to extract deep features f_t for each frame; the inter-frame similarity matrix S is calculated: S(i,j) = cos_sim(f_i,j) i , f jThe video is segmented into K scenes using a spectral clustering algorithm (K is automatically determined by the feature vector gap method); the frame with the highest saliency in each scene is selected as the representative frame; the candidate frames obtained based on inflection points and scene segmentation are merged to output a comprehensive keyframe candidate set.

[0308] By considering the temporal structure and semantic content of the video, the system can select the most representative keyframes. Cumulative saliency curve analysis captures significant changes in the video, while scene segmentation based on deep features identifies semantic boundaries. The combination of these two methods ensures that the keyframes contain important events and cover the main content of the video. In particular, the introduction of a pre-trained ResNet-50 network to extract deep features enables the system to understand the high-level semantics of the construction site scene, improving the quality and representativeness of keyframe selection and providing refined visual data for subsequent target recognition and tracking.

[0309] Based on the keyframe candidate set, an improved YOLOv5 algorithm is used for object detection:

[0310] YOLOv5 model configuration: backbone network is CSPDarknet53, input resolution is 640×640; feature pyramid is PANet structure, multi-scale feature fusion; anchor box size is 9 prior boxes obtained by adaptive clustering; loss function is CIoU loss for bounding box regression and BCE loss for classification.

[0311] Special optimizations for construction site target detection: Adding simulations of common construction site environmental interferences such as low light, rain, fog, and dust; using Focal Loss to mitigate class imbalance; adding additional small-scale detection heads;

[0312] Target recognition process: Apply the YOLOv5 model to each keyframe to obtain detection results (category, confidence, bounding box); apply the non-maximum suppression (NMS) algorithm to merge overlapping boxes, with the IoU threshold set to 0.5; filter low-confidence detection results (threshold of 0.4); output a target list, including common construction site objects such as people, vehicles, and equipment;

[0313] Cross-frame target matching and trajectory generation: The SORT (Simple Online and Realtime Tracking) algorithm is used for target tracking; a Kalman filter is used to predict target positions; a Hungarian algorithm is used for data association based on IoU distance; target appearance and disappearance are handled (maximum number of disappearance frames is set to 30); complete trajectory data for each target is generated;

[0314] Specifically optimized for construction site environments, innovative improvements such as data augmentation, class balancing, and small target enhancement have significantly improved target detection performance in complex construction site conditions. In particular, the addition of data augmentation strategies simulating common environmental disturbances on construction sites makes the model more robust, maintaining good performance even under harsh conditions such as low light and dust. The SORT tracking algorithm, combining Kalman filter prediction and Hungarian algorithm data association, maintains stable tracking performance even under conditions of short-term target occlusion and rapid movement, providing reliable trajectory data for identifying responsible parties.

[0315] Based on the target trajectory and anomaly information, identify potential responsible parties:

[0316] For each tracked target i, calculate its position p at the anomaly time t_anomaly. i (t_anomaly); Calculates the spatial distance d between the target location and the anomaly location. i = ||p i (t_anomaly) p_anomaly||; calculates the minimum distance d between the target and the anomaly region within the time window [t_anomaly-Δt, t_anomaly]. M in,i;

[0317] Behavior analysis: Extract the target's motion characteristics (velocity, acceleration, direction change) before the anomaly; calculate the behavior anomaly score s_behavior,i based on the deviation from the normal behavior pattern; analyze the interaction pattern between the target and the equipment / cable;

[0318] Calculate the responsibility score i = w d ·exp(-d Min,i / σ d ) + w b ·s_behavior,i + w i ·s identity,i ;where w d =0.5, w b =0.3, w i =0.2 is the weighting coefficient, s identity,i The targets are ranked according to their prior scores based on their identities, and a list of potential responsible parties is generated.

[0319] By combining spatiotemporal correlation analysis and behavioral analysis, this system comprehensively considers the spatial proximity, temporal correlation, and behavioral anomalies between the target and the anomaly, achieving more accurate responsibility determination. In particular, the introduction of a behavioral anomaly score allows the system to identify potential abnormal behaviors by analyzing the target's movement patterns and interactions with equipment / cables before the anomaly. The responsibility score calculation formula integrates multiple factors and uses an exponential decay function to reasonably model the impact of spatial distance, making responsibility determination more objective and reasonable. This intelligent responsibility tracking mechanism provides site managers with objective evidence for accident responsibility analysis, helping to improve the targeting and effectiveness of safety management.

[0320] After deploying this system at a high-rise building construction site, a three-month comparative test was conducted with traditional monitoring and cable fault detection systems. The system performance indicators are as follows:

[0321] Cable core breakage detection: Detection rate 95.8% (traditional system: 72.3%); false alarm rate 4.2% (traditional system: 15.7%); average early warning time 4.6 hours (traditional system: no early warning capability); positioning accuracy ±0.65 meters (traditional system: ±3.2 meters);

[0322] Target detection and tracking: Personnel detection accuracy was 94.2% (traditional system: 85.1%); Equipment detection accuracy was 92.7% (traditional system: 79.8%); Tracking stability (MOTA) was 83.5% (traditional system: 61.2%); Cross-camera tracking accuracy was 78.9% (traditional system: 45.3%).

[0323] Abnormal event response: The average event detection time is 1.8 seconds (traditional system: 15.6 seconds); the accuracy rate of identifying the responsible party is 89.3% (traditional system: no such function); the overall system reliability is 99.2% (runtime ratio).

[0324] Through the practical application of this system, the construction site accident rate has decreased by 47.8%, equipment maintenance costs have decreased by 32.5%, and workers' safety awareness and sense of responsibility have significantly improved. In particular, the system's function of identifying responsible parties makes on-site responsibility tracking more accurate and fair, greatly reducing liability disputes and improving management efficiency.

[0325] By integrating cable monitoring and video analytics technologies and co-processing multi-source data, timely detection and precise location of hidden faults such as broken cores are achieved. A cable anomaly detection method based on empirical mode decomposition and Hilbert spectral analysis can capture minute changes in cable parameters, significantly improving fault early warning capabilities. An anomaly pattern recognition method combining multidimensional phase space and density clustering can effectively distinguish different types of cable anomalies, providing targeted guidance for maintenance. A dynamic region recognition algorithm combining spatiotemporal Markov random fields and improved optical flow estimation significantly improves target recognition and tracking performance in complex construction site environments. The introduction of multi-feature fusion saliency analysis and deep feature scene segmentation technology enables high-quality keyframe extraction, supporting rapid location and analysis of anomaly events. A responsibility identification mechanism based on spatiotemporal correlation and behavioral analysis enables automatic tracking of responsibility for anomaly events, providing objective evidence for construction site safety management.

[0326] It should be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present invention will not describe the various possible combinations separately.

Claims

1. A method for intelligent site management based on photoelectric and video fence, characterized in that, The method comprises the following steps: Step S1, obtain real-time data of the construction site every predetermined period, and pre-process to form a pre-processed data set, including enhanced video data, cable operation data, environmental state parameters, and equipment switch state; Step S2, read the enhanced video data, identify and obtain the current construction site state, obtain dynamic areas and mark them to form a dynamic area data set; Step S3, read the cable operation data and equipment switch state, extract cable abnormal parameters therefrom, and evaluate the severity of the abnormality, output an abnormality level index and compare it with a threshold value, if the threshold value is exceeded, record the time stamp of the abnormality occurrence time and the corresponding detection area coordinates; Step S4, according to the time stamp of the abnormality occurrence time, retrieve a video segment of a predetermined length from the dynamic area data set, determine the key frame position range and form a key frame candidate set; Step S5, identify and track the target based on the key frame candidate set and cross-frame match the target to generate a target trajectory, and obtain a list of responsible subjects according to the target trajectory; Step S4 specifically comprises: Step S41, read the dynamic area data set and abnormality detection result, and extract the corresponding video segment from the dynamic area data set; Step S42, re-optimize the video segment to output an optimized video segment set; Step S43, for each optimized video segment set, obtain and calculate the motion saliency and visual saliency according to the corresponding motion analysis result, and output the saliency sequence of each video segment; Step S44, based on the optimized video segment set, calculate the time sequence saliency, obtain the potential key frame position range, and determine the potential key frame, and output the key frame candidate set; Step S44 specifically comprises: Step S44a, read each enhanced video segment and the corresponding saliency sequence, as well as the construction area safety level map and the operation risk database; based on the construction area safety level map, regionally weight the saliency sequence to form a safety risk weighted saliency sequence S_r(t); Step S44b, identify the standard procedure and non-standard operation in the video: extract the deep features f_t of each frame; use a pre-trained construction site operation specification identification model M_op to classify the features to obtain the operation type O_t and the specification compliance score C_t; label the frames with a specification compliance score below a threshold value T_c as potential risk frames; calculate the operation risk weighting factor R_op(t) = β·(1-C_t)·I(O_t), where I(O_t) is the inherent risk index of the operation type, and β is an adjustable parameter; Step S44c, generate a safety-oriented cumulative saliency curve CSC_safety(t) = ∑_{i=1}^t [α·S_r(i)+ (1-α)·R_op(i)]; where α is a balance parameter between saliency and operation risk, which is dynamically adjusted according to the engineering stage; Step S44d, use a multi-level threshold detection algorithm to identify key frames: apply an adaptive three-level threshold {T_low, T_med, T_high} to the CSC_safety curve; for high-risk operation areas (engineering key parts), use the T_low threshold; For medium-risk areas, T_med threshold is adopted; For low-risk areas, T_high threshold is adopted; Identify the inflection point meeting the corresponding threshold as a key frame candidate; Step S44e, adopt the redundant frame filtering method based on the construction site knowledge graph, calculate the construction semantic similarity between the candidate frames, retain the key frames representative in construction semantics, and output the final key frame candidate set.

2. The method of claim 1, wherein, Step S1 is specifically: Step S11, every predetermined period, obtain the original video data of the construction site and perform enhancement processing to obtain enhanced video data; Step S12, obtain cable monitoring data and modal decomposition to form original cable operation data, including cable normal mode, abnormal mode and cable parameter value; Step S13, obtain the initial environmental state parameters of the environmental state parameter sensor and perform fusion processing to obtain the environmental state parameters; Step S14, collect the working parameters of the construction equipment intelligent switch, construct the original equipment data, and perform preprocessing to form the equipment switch state; Step S15, based on the output data of steps S11 to S14, construct a multi-modal data set, and perform time alignment through a multi-dimensional dynamic time warping algorithm to form a preprocessed data set.

3. The method of claim 2, wherein, Step S2 is specifically: Step S21, read the enhanced video data, extract the video frames, and for each video frame, perform region division, extract feature vectors and clustering to obtain an initial scene segmentation result; Step S22, for the initial scene segmentation result, adopt a spatiotemporal Markov random field module for optimization and minimize the energy function to obtain a spatiotemporal consistent scene segmentation result; Step S23, based on the spatiotemporal consistent scene segmentation result, calculate the motion vector field, identify the dynamic region, and segment out the motion salient region; track each motion in the motion salient region to obtain the dynamic region tracking result, i.e. the dynamic region data set; Step S24, label and hierarchically represent the dynamic region tracking result, construct a comprehensive dynamic region description data set and output.

4. The method of claim 3, wherein, Step S3 is specifically: Step S31, obtain the comprehensive dynamic region description data set, perform multivariate anomaly detection and dimension reduction to obtain a processed multi-source synchronous data set; Step S32, read the cable operation data, and preliminarily detect whether there is an abnormal mode; based on the processed multi-source synchronous data set, calculate and output the cable abnormal parameter set; Step S33, fuse the cable abnormal parameter set with the equipment switch state and environmental state parameters in the multi-source synchronous data set to output the fused abnormal feature data set, i.e. the anomaly detection result; Step S34, call the preconfigured fuzzy inference system to evaluate each data point of the anomaly detection result to obtain an anomaly severity evaluation result; Step S35, record the timestamp and coordinate information of the anomaly severity evaluation result, and output the spatiotemporal positioning result.

5. The method of claim 1, wherein, Step S5 is specifically: Step S51, read the key frame data set, the dynamic region data set, the anomaly detection result and the preconfigured registration database; adopt YOLOvx target detection algorithm for identification and tracking; obtain a target list; Step S52, cross-frame matching is performed on the target results one by one; target trajectory data is generated to form a target trajectory; Step S53, in combination with the target trajectory, a time-space correlation analysis algorithm is used to determine the target closest to the detection area at the time point when the anomaly occurs, and a list of potential responsible subjects is output.

6. The method of claim 2, wherein, In step S11, the process of enhancing the original video data of the construction site is as follows: In step S111, the original video data of the construction site is obtained, and video frame images are extracted. For each video frame image, discrete wavelet transform is used for decomposition to obtain four frequency subbands. In step S112, for each frequency subband, the local mean and standard deviation of the frequency subband are calculated; and based on the local mean, the standard deviation and the pre-stored adjustable parameters, an enhancement function is constructed. The enhancement function is applied to each pixel in the frequency subband. In step S113, the enhanced frequency subband is reconstructed by using a weighted summation method to obtain an enhanced image, and the enhanced images are spliced to form enhanced video data.

7. The method of claim 3, wherein, Step S21 is specifically as follows: In step S211, each frame image of the enhanced video data is read and divided into a plurality of initial regions, and a region center is set. For each pixel, the distance between the pixel and the superpixel center is iteratively calculated, and the pixel attribution and the superpixel center are updated until convergence or the maximum number of iterations is reached to obtain a superpixel segmentation result. In step S212, for each superpixel region in the superpixel segmentation result, a color histogram is calculated, a texture feature is extracted using an improved SURF algorithm, a shape descriptor is calculated, and a comprehensive feature vector is combined. In step S213, all comprehensive feature vectors are clustered, a distance matrix between the feature vectors is calculated, a hierarchical tree is constructed, and the feature hierarchical tree is cut according to a preset threshold to obtain an initial scene segmentation result.

8. The method of claim 3, wherein, Step S22 is specifically as follows: In step S221, the initial scene segmentation results of a plurality of consecutive frames are read, and a space-time Markov random field model is constructed, including an energy function of data items and smoothing items. In step S222, a Gaussian mixture model is used to estimate the probability that a pixel belongs to each label, and the data items in the energy function are calculated. In step S223, a weight is calculated based on the color similarity and spatial distance of adjacent pixels to obtain the smoothing items in the energy function. In step S224, a graph cut algorithm is applied to minimize the energy function to obtain an optimized scene segmentation label, and a space-time consistent scene segmentation result is generated and output according to the scene segmentation label.

9. A smart construction site management system based on photoelectric and video fence, characterized in that, The control module is connected with an information collection device, and the information collection device includes a camera device, a cable monitoring device, an environmental state parameter sensor, and a construction equipment intelligent switch. The cable monitoring device is arranged at a predetermined length along the direction in which the cable extends. The control module includes at least one processor and a memory in communication with the at least one processor. The memory stores instructions executable by the processor, and the instructions are executed by the processor to implement the intelligent construction site management method based on the photoelectric and video fence. ​

Citation Information

Patent Citations

  • A systematic maintenance management system and method for electrical equipment suitable for construction sites

    CN116243072B

  • Behavior identification method and system based on buried leakage cable and multi-mode sensing

    CN118747321A

  • Connection system for high-voltage cable fault early warning and positioning

    CN217846513U