Electric power construction site video tracking and early warning method

By combining mask autoencoders and monocular depth estimation algorithms, the three-dimensional spatial structure reconstruction and multimodal feature fusion of power construction sites are realized, solving the problems of insufficient three-dimensional spatial perception and inaccurate early warning in existing technologies, and improving the safety management efficiency and intelligence level of power safety supervision sites.

CN121095847BActive Publication Date: 2026-04-21DEYANG POWER SUPPLY COMPANY STATE GRID SICHUAN ELECTRIC POWER
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DEYANG POWER SUPPLY COMPANY STATE GRID SICHUAN ELECTRIC POWER
Filing Date
2025-09-17
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing video monitoring systems at power construction sites lack comprehensive perception of three-dimensional spatial structures and fusion of multimodal features, resulting in insufficient accuracy, stability, and real-time performance in target detection and tracking. They are unable to cope with multi-source information interaction and hidden risk identification in dynamic operation scenarios, and suffer from false alarms, missed alarms, and slow early warning response.

Method used

A masked autoencoder is used for self-supervised learning and monocular depth estimation algorithm to reconstruct the three-dimensional spatial structure. Combined with multimodal feature fusion, it can achieve efficient identification, real-time tracking and intelligent early warning of personnel, equipment and vehicles at the work site and spatial risk behavior.

Benefits of technology

It significantly improves the efficiency and intelligence level of safety management at power construction sites, enhances the accuracy of identifying high-risk areas and abnormal behaviors, and enables continuous and accurate tracking and real-time intelligent early warning in complex working conditions and dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095847B_ABST
    Figure CN121095847B_ABST
Patent Text Reader

Abstract

This invention discloses a video tracking and early warning method for power construction sites, comprising: collecting raw video data and preprocessing it to obtain standardized video frame data; extracting high-dimensional apparent feature data using a masked autoencoder model; predicting depth information using a monocular depth estimation algorithm to generate depth feature data; fusing multimodal features, performing spatial alignment, multi-target detection and tracking to generate target tracking data; analyzing target spatial behavior, identifying abnormal behavior, and automatically generating early warning data; and storing the tracking and early warning data in a database for continuous training and optimization. This invention achieves intelligent video tracking and risk early warning for personnel and equipment at power construction sites, effectively improving the automation and precision of on-site safety management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of safety monitoring technology, and in particular to a method for video tracking and early warning at power construction sites. Background Technology

[0002] With the continuous expansion of the power industry and the ongoing improvement of automation levels, safety risk prevention and control and operation management at power construction sites are facing increasingly complex challenges. Traditional power construction site supervision mainly relies on manual inspections or video monitoring using ordinary cameras. These methods have prominent problems such as limited on-site perception capabilities, low target recognition accuracy, and insufficient adaptability to complex operating environments. In recent years, multi-target detection, target tracking, and abnormal behavior analysis methods based on computer vision and artificial intelligence have been gradually applied to power safety supervision scenarios. Target detection algorithms identify on-site personnel, equipment, and vehicles, and target tracking technology is used for continuous monitoring and risk analysis of operational behavior. However, existing technologies mostly focus on single-modal visual information processing, lacking comprehensive perception of the three-dimensional spatial structure of the work site and deep fusion of multi-modal features.

[0003] Most commonly used video analytics systems rely on two-dimensional image feature extraction and analysis, which cannot effectively reconstruct the three-dimensional spatial relationships of the scene. This results in limitations in the accuracy, stability, and real-time performance of target detection and tracking under conditions such as multi-target occlusion, dense distribution of personnel and equipment, and complex changes in working conditions. Existing systems mainly rely on simple rule-based judgments and single feature thresholds for identifying abnormal behavior and safety risks. They are ill-equipped to handle the interaction of multi-source information in dynamic work scenarios, the timely identification of hidden risks, and the need for intelligent early warning of high-risk areas. They lack a focus on key areas and dangerous actions, as well as regional intelligent discrimination mechanisms, which can easily lead to false alarms, missed alarms, and slow early warning responses.

[0004] Therefore, how to provide video tracking and early warning methods for power construction operations is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a video tracking and early warning method for power construction sites. This invention fully utilizes the self-supervised learning capability of a mask autoencoder for the apparent features of work site images and the accurate reconstruction of three-dimensional spatial structures using a monocular depth estimation algorithm. Through multimodal feature fusion, spatial semantic enhancement, multi-target detection, and three-dimensional consistency tracking technologies, it achieves efficient identification, real-time tracking, and intelligent early warning of spatial risk behaviors for personnel, equipment, and vehicles at power construction sites. This invention possesses advantages such as strong on-site perception capabilities, accurate spatial risk identification, timely response to abnormal behaviors, and a high degree of intelligent early warning, thereby improving the efficiency and intelligence level of safety management at power construction sites.

[0006] The power construction site video tracking and early warning method according to an embodiment of the present invention includes:

[0007] Raw video data from power construction sites is collected, and the raw video data is preprocessed to obtain standardized video frame data.

[0008] For standardized video frame data, a mask autoencoder model is used for random masking processing. Part of the image blocks are input into the encoder for feature encoding, and the decoder reconstructs the masked area to extract high-dimensional appearance features of the work site and generate appearance feature data.

[0009] For standardized video frame data, a monocular depth estimation algorithm is used to predict depth information, obtain the depth value of each pixel in each frame, reconstruct the three-dimensional spatial structure of the work site, and generate depth feature data.

[0010] Multi-level feature fusion is performed on apparent feature data and deep feature data to obtain fused multimodal feature representation, which is then used for multi-target detection and tracking to obtain target tracking data;

[0011] Based on the spatial location and movement trajectory of the operator and equipment, the system monitors the relative distance, boundary crossing behavior and abnormal approach events in real time, performs early warning judgment, and triggers early warning signals and generates early warning data for detected violations or dangerous behaviors according to the set early warning threshold.

[0012] Target tracking data and early warning data are stored in the database, and the mask autoencoder model and monocular depth estimation algorithm are continuously trained and optimized based on newly acquired video data.

[0013] Optionally, the original video data specifically includes a continuous sequence of dynamic image frames covering the personnel, equipment, vehicles, and panoramic view of the work area at the power construction site.

[0014] Optionally, the preprocessing of the original video data specifically includes noise reduction, color normalization, and size standardization of the original video data.

[0015] Optionally, the generation of appearance feature data specifically includes:

[0016] A mask autoencoder model is constructed, which includes a dynamic risk-aware mask generation module, an encoder, a region feature enhancement dual-branch encoding module, a feature fusion layer, a decoder, and a risk-weighted reconstruction decoding module. The dynamic risk-aware mask generation module generates an adaptive mask distribution scheme based on the regional risk level of the video frame at the work site. The encoder performs feature encoding on the unmasked image blocks. The region feature enhancement dual-branch encoding module includes a high-risk branch subnetwork and a normal region branch subnetwork. The feature fusion layer is used to fuse the features of the two branches. The decoder reconstructs the complete video frame based on the fused features and the mask distribution. The risk-weighted reconstruction decoding module includes a risk weight allocation unit and a reconstruction loss calculation unit.

[0017] Standardized video frame data is input into the dynamic risk perception mask generation module. The dynamic risk perception mask generation module includes a risk area identification unit and a mask allocation unit. The risk area identification unit divides the video frame into regions based on the personnel distribution, equipment status and historical high-risk area distribution results in the standardized video frame data. The mask allocation unit dynamically sets the mask ratio of each region according to the region risk level, reduces the mask ratio of high-risk regions and increases the mask ratio of ordinary regions, and outputs masked video frame data.

[0018] Unmasked high-risk area image blocks and unmasked ordinary area image blocks are respectively input into the encoder and the dual-branch encoding module for region feature enhancement. The encoder performs preliminary feature encoding on all unmasked image blocks, the high-risk area feature extraction branch performs deep feature encoding on the unmasked high-risk area image blocks, and the ordinary area feature extraction branch performs feature encoding on the unmasked ordinary area image blocks. The output feature vectors are concatenated or weighted fused by the feature fusion layer to obtain the multidimensional appearance feature encoding result of the video frame.

[0019] The multidimensional appearance feature encoding result and mask position information are input to the decoder and the risk-weighted reconstruction decoding module. The decoder reconstructs the complete video frame based on the multidimensional appearance feature encoding result and mask position information. The risk weight allocation unit assigns different reconstruction loss weights to high-risk areas and ordinary areas. The reconstruction loss calculation unit performs reconstruction operations on all occluded image blocks and outputs the reconstructed video frame.

[0020] Based on the reconstructed video frames, the reconstruction error of all occluded image blocks is statistically analyzed. Occluded image blocks in high-risk areas are weighted according to the high-risk error weight threshold, and occluded image blocks in ordinary areas are weighted according to the ordinary area error weight threshold. The reconstruction accuracy of different areas is evaluated by weighted accumulation, and the evaluation results are used as the target of training and optimization.

[0021] The mask autoencoder model is trained, and the reconstruction error is used as the optimization target to continuously optimize the parameters of the dynamic risk perception mask generation module, encoder, region feature enhancement dual-branch encoding module, feature fusion layer, decoder and risk weighted reconstruction decoding module.

[0022] Using the trained masked autoencoder model, appearance feature encoding is performed on the newly acquired standardized video frame data, and the appearance feature data of the operation site video frames is output.

[0023] Optionally, the generation of deep feature data specifically includes:

[0024] The monocular depth estimation algorithm includes a risk region-guided multi-scale feature enhancement encoder, a spatial relation context attention module, a feature decoder, and a region weight adaptive depth regression loss branch.

[0025] Standardized video frame data is input into a risk area-guided multi-scale feature enhancement encoder. Based on the location information of high-risk and ordinary areas, high-resolution and multi-scale feature extraction is performed on the high-risk area image blocks, and standard-scale feature extraction is performed on the ordinary area image blocks, to obtain high-risk area features and ordinary area features respectively.

[0026] High-risk area features and ordinary area features are input into the spatial relationship context attention module. The module receives high-risk area features and ordinary area features, calculates the correlation between each area feature and other area features, obtains the correlation score between area features, normalizes the correlation score and uses it as attention weight, performs weighted fusion of area features, and outputs fused spatial features.

[0027] The fused spatial features are input into the feature decoder, which upsamples and reconstructs the fused spatial features to generate a pixel-level depth map with the same size as the input video frame. Each pixel value corresponds to the depth estimate at the same position in the input video frame.

[0028] Spatial consistency correction is performed on each frame of pixel-level depth map. The spatial error caused by camera motion or viewpoint change is corrected by using inter-frame correspondence and geometric constraints to obtain the corrected depth map sequence.

[0029] The corrected depth map sequence is converted into a three-dimensional spatial point cloud structure of the work site through a point cloud mapping method. Each point in the point cloud contains three-dimensional coordinates.

[0030] A region-weighted adaptive depth regression loss branch is adopted, which assigns high-risk loss weight thresholds to pixels in high-risk regions and ordinary region loss weight thresholds to pixels in ordinary regions. Based on the weighted depth regression error of all pixels, the parameters of the multi-scale feature enhancement encoder, spatial relationship context attention module and feature decoder are optimized, and the optimized and corrected depth feature data is output.

[0031] Optionally, obtaining the target tracking data specifically involves:

[0032] The high-dimensional appearance feature data and the deep feature data are aligned according to spatial coordinates, and the spatial position corresponding to each frame of the image is uniformly encoded to form a multimodal feature alignment dataset.

[0033] For each spatial location in the multimodal feature alignment dataset, the apparent feature vector is concatenated with the deep feature vector to generate a fused feature vector;

[0034] For each candidate detection target, the fused feature vector is connected with the corresponding spatial location coordinates and regional risk label values ​​to form an extended feature vector. The extended feature vector is then input into the target detection network, and the appearance features, depth features, spatial location information and regional risk information contained in the extended feature vector are used to jointly classify and locate the target.

[0035] By using a multi-scale spatial neighborhood discrimination method, during the candidate target detection process, for each detected target, the distribution of fusion features of the surrounding neighborhood is analyzed. The detection confidence is improved for regions with significant changes in spatial features, and the false detection probability is reduced for regions with continuous spatial features. Finally, the category label, location coordinates and confidence score of each target in each frame are output.

[0036] In the process of multi-target tracking, for the detected targets in consecutive frames, the three-dimensional motion consistency association method is adopted by combining the fused feature vector and the three-dimensional spatial coordinates of the target. The targets with the smallest spatial distance, consistent motion trend and highest fused feature similarity in adjacent frames are identified as the same target, assigned the same identity label, and the spatial trajectory is updated to generate target tracking data.

[0037] Spatial behavior analysis is performed on target tracking data. Using three-dimensional spatial trajectories, the real-time position, speed and historical movement path of workers and equipment are calculated. High-risk operation, unauthorized movement and boundary crossing behavior characteristics are extracted to form spatial behavior data for on-site safety monitoring.

[0038] The multi-target detection results after fusing feature vectors and spatial semantic injection are output in a unified manner with the target tracking data after being correlated with 3D motion consistency.

[0039] Optionally, the generation of early warning data specifically includes:

[0040] Real-time analysis of target tracking data and spatial behavior data is performed to statistically analyze the real-time position, velocity, acceleration, and historical trajectory of each target in three-dimensional space.

[0041] The spatial distance between operators and equipment targets is monitored in real time. When the distance between any two targets is less than the set spatial safety threshold, it is judged as abnormal approach behavior.

[0042] The system compares the regional attributes of the target at the work site in real time and judges whether the target has crossed the boundary, stayed for a long time, or entered a high-risk area based on the regional risk label. When the target stays in a high-risk area for more than a preset threshold, it is judged as a behavior of staying beyond the boundary.

[0043] By combining the changes in the target's speed and acceleration, and through behavioral data analysis, high-risk operations, unauthorized movements, and sudden stops are identified as abnormal behavioral events, and corresponding abnormal behavior indicators are output.

[0044] Multi-level early warning conditions are set for abnormal behavior events of different risk types. The events are divided into three levels: low, medium and high according to their risk level. Early warning signals are automatically generated for targets or behaviors that meet the early warning conditions. The occurrence time, involved objects, spatial location and event description of all abnormal behavior events and early warning signals are structured and organized to finally generate early warning data for on-site safety management and decision analysis.

[0045] The beneficial effects of this invention are:

[0046] The video tracking and early warning method for power construction sites, based on the deep fusion of a mask autoencoder model and a monocular depth estimation algorithm, provided by this invention, effectively compensates for the shortcomings of existing technologies in 3D spatial perception, fine target tracking, and multimodal risk early warning. This invention introduces a mask autoencoder to achieve self-supervised learning of complex visual features at the work site, and combines this with a monocular depth estimation algorithm for accurate reconstruction of the 3D spatial structure of the site, significantly improving the accuracy of identifying high-risk areas, abnormal behaviors, and violations. The deep fusion of multimodal features and innovative spatial behavior analysis process enable the system to continuously and accurately track personnel, equipment, and vehicles even under complex working conditions, multi-target occlusion, and dynamic environmental changes, and to provide real-time intelligent early warnings for risk factors such as abnormal spatial distances, boundary crossings, and high-risk behaviors. Through this invention, the safety risk perception capability, automated response efficiency, and management intelligence level of power safety monitoring sites are significantly improved, providing strong technical support for safe production and intelligent operation and maintenance in the power industry. Attached Figure Description

[0047] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0048] Figure 1 This is a flowchart of the video tracking and early warning method for power construction operations proposed in this invention;

[0049] Figure 2 This is a schematic diagram of the masked autoencoder model structure and its appearance feature extraction process for the power construction site video tracking and early warning method proposed in this invention. Detailed Implementation

[0050] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0051] refer to Figure 1 and Figure 2 Methods for video tracking and early warning at power construction sites include:

[0052] Raw video data from power construction sites is collected, and the raw video data is preprocessed to obtain standardized video frame data.

[0053] For standardized video frame data, a mask autoencoder model is used for random masking processing. Part of the image blocks are input into the encoder for feature encoding, and the decoder reconstructs the masked area to extract high-dimensional appearance features of the work site and generate appearance feature data.

[0054] For standardized video frame data, a monocular depth estimation algorithm is used to predict depth information, obtain the depth value of each pixel in each frame, reconstruct the three-dimensional spatial structure of the work site, and generate depth feature data.

[0055] Multi-level feature fusion is performed on apparent feature data and deep feature data to obtain fused multimodal feature representation, which is then used for multi-target detection and tracking to obtain target tracking data;

[0056] Based on the spatial location and movement trajectory of the operator and equipment, the system monitors the relative distance, boundary crossing behavior and abnormal approach events in real time, performs early warning judgment, and triggers early warning signals and generates early warning data for detected violations or dangerous behaviors according to the set early warning threshold.

[0057] Target tracking data and early warning data are stored in the database, and the mask autoencoder model and monocular depth estimation algorithm are continuously trained and optimized based on newly acquired video data.

[0058] In this embodiment, the original video data specifically includes a continuous dynamic image frame sequence covering the personnel, equipment, vehicles, and panoramic view of the work area at the power construction site.

[0059] In this embodiment, the preprocessing of the original video data specifically includes noise reduction, color normalization, and size standardization of the original video data.

[0060] In this embodiment, the generation of appearance feature data specifically refers to:

[0061] A mask autoencoder model is constructed, which includes a dynamic risk-aware mask generation module, an encoder, a region feature enhancement dual-branch encoding module, a feature fusion layer, a decoder, and a risk-weighted reconstruction decoding module. The dynamic risk-aware mask generation module generates an adaptive mask distribution scheme based on the regional risk level of the video frame at the work site. The encoder performs feature encoding on the unmasked image blocks. The region feature enhancement dual-branch encoding module includes a high-risk branch subnetwork and a normal region branch subnetwork. The feature fusion layer is used to fuse the features of the two branches. The decoder reconstructs the complete video frame based on the fused features and the mask distribution. The risk-weighted reconstruction decoding module includes a risk weight allocation unit and a reconstruction loss calculation unit.

[0062] Standardized video frame data is input into the dynamic risk perception mask generation module. The dynamic risk perception mask generation module includes a risk area identification unit and a mask allocation unit. The risk area identification unit divides the video frame into regions based on the personnel distribution, equipment status and historical high-risk area distribution results in the standardized video frame data. The mask allocation unit dynamically sets the mask ratio of each region according to the region risk level, reduces the mask ratio of high-risk regions and increases the mask ratio of ordinary regions, and outputs masked video frame data.

[0063] Unmasked high-risk area image blocks and unmasked ordinary area image blocks are respectively input into the encoder and the dual-branch encoding module for region feature enhancement. The encoder performs preliminary feature encoding on all unmasked image blocks, the high-risk area feature extraction branch performs deep feature encoding on the unmasked high-risk area image blocks, and the ordinary area feature extraction branch performs feature encoding on the unmasked ordinary area image blocks. The output feature vectors are concatenated or weighted fused by the feature fusion layer to obtain the multidimensional appearance feature encoding result of the video frame.

[0064] The multidimensional appearance feature encoding result and mask position information are input to the decoder and the risk-weighted reconstruction decoding module. The decoder reconstructs the complete video frame based on the multidimensional appearance feature encoding result and mask position information. The risk weight allocation unit assigns different reconstruction loss weights to high-risk areas and ordinary areas. The reconstruction loss calculation unit performs reconstruction operations on all occluded image blocks and outputs the reconstructed video frame.

[0065] Based on the reconstructed video frames, the reconstruction error of all occluded image blocks is statistically analyzed. Occluded image blocks in high-risk areas are weighted according to the high-risk error weight threshold, and occluded image blocks in ordinary areas are weighted according to the ordinary area error weight threshold. The reconstruction accuracy of different areas is evaluated by weighted accumulation, and the evaluation results are used as the target of training and optimization.

[0066] The mask autoencoder model is trained, and the reconstruction error is used as the optimization target to continuously optimize the parameters of the dynamic risk perception mask generation module, encoder, region feature enhancement dual-branch encoding module, feature fusion layer, decoder and risk weighted reconstruction decoding module.

[0067] Using the trained masked autoencoder model, appearance feature encoding is performed on the newly acquired standardized video frame data, and the appearance feature data of the operation site video frames is output.

[0068] In this embodiment, the generation of depth feature data specifically refers to:

[0069] The monocular depth estimation algorithm includes a risk region-guided multi-scale feature enhancement encoder, a spatial relation context attention module, a feature decoder, and a region weight adaptive depth regression loss branch.

[0070] Standardized video frame data is input into a risk area-guided multi-scale feature enhancement encoder. Based on the location information of high-risk and ordinary areas, high-resolution and multi-scale feature extraction is performed on the high-risk area image blocks, and standard-scale feature extraction is performed on the ordinary area image blocks, to obtain high-risk area features and ordinary area features respectively.

[0071] High-risk area features and ordinary area features are input into the spatial relationship context attention module. The module receives high-risk area features and ordinary area features, calculates the correlation between each area feature and other area features, obtains the correlation score between area features, normalizes the correlation score and uses it as attention weight, performs weighted fusion of area features, and outputs fused spatial features.

[0072] The fused spatial features are input into the feature decoder, which upsamples and reconstructs the fused spatial features to generate a pixel-level depth map with the same size as the input video frame. Each pixel value corresponds to the depth estimate at the same position in the input video frame.

[0073] Spatial consistency correction is performed on each frame of pixel-level depth map. The spatial error caused by camera motion or viewpoint change is corrected by using inter-frame correspondence and geometric constraints to obtain the corrected depth map sequence.

[0074] The corrected depth map sequence is converted into a three-dimensional spatial point cloud structure of the work site through a point cloud mapping method. Each point in the point cloud contains three-dimensional coordinates.

[0075] A region-weighted adaptive depth regression loss branch is adopted, which assigns high-risk loss weight thresholds to pixels in high-risk regions and ordinary region loss weight thresholds to pixels in ordinary regions. Based on the weighted depth regression error of all pixels, the parameters of the multi-scale feature enhancement encoder, spatial relationship context attention module and feature decoder are optimized, and the optimized and corrected depth feature data is output.

[0076] In this embodiment, obtaining the target tracking data specifically refers to:

[0077] The high-dimensional appearance feature data and the deep feature data are aligned according to spatial coordinates, and the spatial position corresponding to each frame of the image is uniformly encoded to form a multimodal feature alignment dataset.

[0078] For each spatial location in the multimodal feature alignment dataset, the apparent feature vector is concatenated with the deep feature vector to generate a fused feature vector;

[0079] For each candidate detection target, the fused feature vector is connected with the corresponding spatial location coordinates and regional risk label values ​​to form an extended feature vector. The extended feature vector is then input into the target detection network, and the appearance features, depth features, spatial location information and regional risk information contained in the extended feature vector are used to jointly classify and locate the target.

[0080] By using a multi-scale spatial neighborhood discrimination method, during the candidate target detection process, for each detected target, the distribution of fusion features of the surrounding neighborhood is analyzed. The detection confidence is improved for regions with significant changes in spatial features, and the false detection probability is reduced for regions with continuous spatial features. Finally, the category label, location coordinates and confidence score of each target in each frame are output.

[0081] In the process of multi-target tracking, for the detected targets in consecutive frames, the three-dimensional motion consistency association method is adopted by combining the fused feature vector and the three-dimensional spatial coordinates of the target. The targets with the smallest spatial distance, consistent motion trend and highest fused feature similarity in adjacent frames are identified as the same target, assigned the same identity label, and the spatial trajectory is updated to generate target tracking data.

[0082] Spatial behavior analysis is performed on target tracking data. Using three-dimensional spatial trajectories, the real-time position, speed and historical movement path of workers and equipment are calculated. High-risk operation, unauthorized movement and boundary crossing behavior characteristics are extracted to form spatial behavior data for on-site safety monitoring.

[0083] The multi-target detection results after fusing feature vectors and spatial semantic injection are output in a unified manner with the target tracking data after being correlated with 3D motion consistency.

[0084] In this embodiment, generating early warning data specifically refers to:

[0085] Real-time analysis of target tracking data and spatial behavior data is performed to statistically analyze the real-time position, velocity, acceleration, and historical trajectory of each target in three-dimensional space.

[0086] The spatial distance between operators and equipment targets is monitored in real time. When the distance between any two targets is less than the set spatial safety threshold, it is judged as abnormal approach behavior.

[0087] The system compares the regional attributes of the target at the work site in real time and judges whether the target has crossed the boundary, stayed for a long time, or entered a high-risk area based on the regional risk label. When the target stays in a high-risk area for more than a preset threshold, it is judged as a behavior of staying beyond the boundary.

[0088] By combining the changes in the target's speed and acceleration, and through behavioral data analysis, high-risk operations, unauthorized movements, and sudden stops are identified as abnormal behavioral events, and corresponding abnormal behavior indicators are output.

[0089] Multi-level early warning conditions are set for abnormal behavior events of different risk types. The events are divided into three levels: low, medium and high according to their risk level. Early warning signals are automatically generated for targets or behaviors that meet the early warning conditions. The occurrence time, involved objects, spatial location and event description of all abnormal behavior events and early warning signals are structured and organized to finally generate early warning data for on-site safety management and decision analysis.

[0090] In this embodiment, storing target tracking data and early warning data in a database, and continuously training and optimizing the mask autoencoder model and monocular depth estimation algorithm based on newly acquired video data, specifically involves:

[0091] The target tracking data and early warning data are organized and labeled with the identity number, category, spatial location, behavioral status and corresponding early warning level of each target;

[0092] Each target tracking data and early warning data is assigned a unique timestamp. The target tracking data and early warning data are indexed according to spatial location, time order and object category, and stored in a structured database. The database includes a target tracking information table, an early warning event table and a spatial behavior log table.

[0093] It supports multi-condition joint retrieval based on identity number, spatial location, warning level and behavior type, enabling efficient query and statistical analysis of historical behavior, risk events and spatial anomalies;

[0094] Newly acquired video data is periodically or as needed input into the mask autoencoder and monocular depth estimation algorithm to complete incremental training and adaptive optimization of the mask autoencoder model parameters, thereby improving the ability of feature extraction and spatial structure modeling.

[0095] Based on the latest tracking and early warning information stored in the database, the mask autoencoder model parameters, spatial risk thresholds, and early warning triggering rules are dynamically adjusted to achieve adaptive safety management for diverse on-site working conditions. Example

[0096] To verify the feasibility of this invention in practice, it was applied to a maintenance work site at a 500 kV substation. The existing video surveillance system used traditional two-dimensional image feature extraction and analysis methods, mainly relying on the contours and textures of personnel and equipment for target detection. Due to the lack of ability to reconstruct three-dimensional spatial structures, the system frequently experienced missed alarms and false alarms in environments with multiple target obstructions, complex lighting, and densely populated personnel. Especially in high-risk areas, two-dimensional analysis could not accurately determine the actual spatial distance between workers and live equipment, often leading to delayed or completely ineffective warnings.

[0097] The method of this invention is deployed in a scenario, acquiring multi-angle video streams from the scene and performing denoising, normalization, and size standardization to obtain standardized video frames. A mask autoencoder is used to perform random occlusion and reconstruction learning on the standardized video frames, extracting high-dimensional appearance feature data. Simultaneously, a monocular depth estimation algorithm is used to predict the depth value of each pixel in each frame, generating depth feature data and reconstructing the 3D spatial structure. Appearance and depth features are fused at multiple levels and input into the target detection and tracking steps to achieve multi-target recognition and trajectory tracking of personnel, equipment, and vehicles. The system further performs spatial behavior analysis on the target tracking data, monitoring personnel crossing boundaries, abnormal approach, and illegal actions in real time, and automatically triggering multi-level warnings. All tracking and warning data are ultimately stored in a database for training optimization and safety decision-making.

[0098] During a 21-day maintenance period in April 2025, the method of this invention processed 3012 hours of video data, identifying 115 workers and 27 pieces of equipment. Results showed that in complex working environments, the system could accurately distinguish multiple workers who were obscured and calculate their three-dimensional distances from the equipment in real time.

[0099] Table 1. Comparison of the effectiveness of the method of the present invention and the traditional two-dimensional video analysis system during substation maintenance.

[0100] As shown in Table 1, the method of this invention significantly outperforms the traditional two-dimensional video analysis system in several key performance indicators. The two methods maintain consistency in total video processing time, number of workers identified, and number of devices identified, indicating that the comparison conditions are comparable. However, in identifying high-risk events such as lingering beyond designated boundaries, abnormal approach, and unauthorized actions, the accuracy rates of the method of this invention reach 100%, 98.6%, and 95.2%, respectively, while the traditional methods only achieve 72%, 78%, and 69%, a significant difference. This demonstrates that the present invention can more comprehensively and accurately capture on-site safety hazards, reducing missed detections and false alarms.

[0101] In terms of multi-target continuous tracking capability, the accuracy of the method of this invention reaches 98.4%, far exceeding the 82.7% of the traditional system, demonstrating its strong recognition and tracking capabilities even in densely populated areas and when targets are occluded. Furthermore, this invention achieves a three-dimensional spatial position error of only 0.18 meters through monocular depth estimation, while traditional two-dimensional methods cannot provide spatial coordinates, highlighting the unique advantages of this invention in three-dimensional environment reconstruction and risk distance determination.

[0102] In terms of real-time performance, the average alarm delay of this invention is only 0.8 seconds, while the traditional method is 82 seconds, representing a nearly 100-fold improvement in response speed. This effectively shortens the safety intervention time and prevents the accident from escalating. Furthermore, this invention did not miss any events during the entire maintenance cycle, while the traditional system missed as many as 11 events. Regarding the false alarm rate for safety violations, this invention has a false alarm rate of only 2.3%, while the traditional system has a false alarm rate as high as 12.1%, greatly improving the reliability and practicality of the early warning system.

[0103] Overall, this invention achieves accurate reconstruction of three-dimensional spatial structures, high accuracy in continuous multi-target tracking, and rapid early warning of abnormal events through the deep integration of mask autoencoder and monocular depth estimation algorithm. It solves the problems of insufficient three-dimensional spatial perception, high false alarm and false negative rates, and slow response in traditional two-dimensional video analysis systems, providing more efficient and reliable technical support for power safety monitoring operations.

[0104] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for video tracking and early warning at power construction sites, characterized in that, include: Raw video data from power construction sites is collected, and the raw video data is preprocessed to obtain standardized video frame data. For standardized video frame data, a mask autoencoder model is used for random masking processing. Part of the image blocks are input into the encoder for feature encoding, and the decoder reconstructs the masked area to extract high-dimensional appearance features of the work site and generate appearance feature data. For standardized video frame data, a monocular depth estimation algorithm is used to predict depth information, obtain the depth value of each pixel in each frame, reconstruct the three-dimensional spatial structure of the work site, and generate depth feature data. Multi-level feature fusion is performed on apparent feature data and deep feature data to obtain fused multimodal feature representation, which is then used for multi-target detection and tracking to obtain target tracking data; Based on the spatial location and movement trajectory of the operator and equipment, the system monitors the relative distance, boundary crossing behavior and abnormal approach events in real time, performs early warning judgment, and triggers early warning signals and generates early warning data based on the set early warning threshold for detected violations or dangerous behaviors. Target tracking data and early warning data are stored in the database, and the mask autoencoder model and monocular depth estimation algorithm are continuously trained and optimized based on newly acquired video data; The generation of appearance feature data specifically includes: A mask autoencoder model is constructed, which includes a dynamic risk-aware mask generation module, an encoder, a region feature enhancement dual-branch encoding module, a feature fusion layer, a decoder, and a risk-weighted reconstruction decoding module. The dynamic risk-aware mask generation module generates an adaptive mask distribution scheme based on the regional risk level of the video frame at the work site. The encoder performs feature encoding on the unmasked image blocks. The region feature enhancement dual-branch encoding module includes a high-risk branch subnetwork and a normal region branch subnetwork. The feature fusion layer is used to fuse the features of the two branches. The decoder reconstructs the complete video frame based on the fused features and the mask distribution. The risk-weighted reconstruction decoding module includes a risk weight allocation unit and a reconstruction loss calculation unit. Standardized video frame data is input into the dynamic risk perception mask generation module. The dynamic risk perception mask generation module includes a risk area identification unit and a mask allocation unit. The risk area identification unit divides the video frame into regions based on the personnel distribution, equipment status and historical high-risk area distribution results in the standardized video frame data. The mask allocation unit dynamically sets the mask ratio of each region according to the region risk level, reduces the mask ratio of high-risk regions and increases the mask ratio of ordinary regions, and outputs masked video frame data. Unmasked high-risk area image blocks and unmasked ordinary area image blocks are respectively input into the encoder and the dual-branch encoding module for region feature enhancement. The encoder performs preliminary feature encoding on all unmasked image blocks, the high-risk area feature extraction branch performs deep feature encoding on the unmasked high-risk area image blocks, and the ordinary area feature extraction branch performs feature encoding on the unmasked ordinary area image blocks. The output feature vectors are concatenated or weighted fused by the feature fusion layer to obtain the multidimensional appearance feature encoding result of the video frame. The multidimensional appearance feature encoding result and mask position information are input to the decoder and the risk-weighted reconstruction decoding module. The decoder reconstructs the complete video frame based on the multidimensional appearance feature encoding result and mask position information. The risk weight allocation unit assigns different reconstruction loss weights to high-risk areas and ordinary areas. The reconstruction loss calculation unit performs reconstruction operations on all occluded image blocks and outputs the reconstructed video frame. Based on the reconstructed video frames, the reconstruction error of all occluded image blocks is statistically analyzed. Occluded image blocks in high-risk areas are weighted according to the high-risk error weight threshold, while occluded image blocks in ordinary areas are weighted according to the ordinary area error weight threshold. The reconstruction accuracy of different areas is evaluated by weighted accumulation, and the evaluation results are used as the target for training and optimization. The mask autoencoder model is trained, and the reconstruction error is used as the optimization target to continuously optimize the parameters of the dynamic risk perception mask generation module, encoder, region feature enhancement dual-branch encoding module, feature fusion layer, decoder and risk weighted reconstruction decoding module. Using the trained masked autoencoder model, appearance feature encoding is performed on the newly acquired standardized video frame data, and the appearance feature data of the operation site video frames is output.

2. The method for video tracking and early warning at power construction sites according to claim 1, characterized in that, The original video data specifically includes a continuous sequence of dynamic image frames covering the personnel, equipment, vehicles, and panoramic view of the work area at the power construction site.

3. The method for video tracking and early warning at power construction sites according to claim 1, characterized in that, The preprocessing of the raw video data specifically includes noise reduction, color normalization, and size standardization.

4. The method for video tracking and early warning at power construction sites according to claim 1, characterized in that, The generation of deep feature data specifically includes: The monocular depth estimation algorithm includes a risk region-guided multi-scale feature enhancement encoder, a spatial relation context attention module, a feature decoder, and a region weight adaptive depth regression loss branch. Standardized video frame data is input into a risk area-guided multi-scale feature enhancement encoder. Based on the location information of high-risk and ordinary areas, high-resolution and multi-scale feature extraction is performed on the high-risk area image blocks, and standard-scale feature extraction is performed on the ordinary area image blocks, to obtain high-risk area features and ordinary area features respectively. High-risk area features and ordinary area features are input into the spatial relationship context attention module. The module receives high-risk area features and ordinary area features, calculates the correlation between each area feature and other area features, obtains the correlation score between area features, normalizes the correlation score and uses it as attention weight, performs weighted fusion of area features, and outputs fused spatial features. The fused spatial features are input into the feature decoder, which upsamples and reconstructs the fused spatial features to generate a pixel-level depth map with the same size as the input video frame. Each pixel value corresponds to the depth estimate at the same position in the input video frame. Spatial consistency correction is performed on each frame of pixel-level depth map. Inter-frame correspondence and geometric constraints are used to correct spatial errors caused by camera motion or viewpoint changes, and a corrected depth map sequence is obtained. The corrected depth map sequence is converted into a three-dimensional spatial point cloud structure of the work site through a point cloud mapping method. Each point in the point cloud contains three-dimensional coordinates. The method employs a region-weighted adaptive depth regression loss branch, assigning high-risk loss weight thresholds to pixels in high-risk regions and ordinary region loss weight thresholds to pixels in ordinary regions. Based on the weighted depth regression error of all pixels, the parameters of the multi-scale feature enhancement encoder, spatial relation context attention module, and feature decoder are optimized, and the optimized and corrected depth feature data is output.

5. The method for video tracking and early warning at power construction sites according to claim 1, characterized in that, The obtained target tracking data specifically includes: The high-dimensional appearance feature data and the deep feature data are aligned according to spatial coordinates, and the spatial position corresponding to each frame of the image is uniformly encoded to form a multimodal feature alignment dataset. For each spatial location in the multimodal feature alignment dataset, the apparent feature vector is concatenated with the deep feature vector to generate a fused feature vector; For each candidate detection target, the fused feature vector is connected with the corresponding spatial location coordinates and regional risk label values ​​to form an extended feature vector. The extended feature vector is then input into the target detection network, and the appearance features, depth features, spatial location information and regional risk information contained in the extended feature vector are used to jointly classify and locate the target. By using a multi-scale spatial neighborhood discrimination method, during the candidate target detection process, for each detected target, the fusion feature distribution of the surrounding neighborhood is analyzed. The detection confidence is improved for regions with significant changes in spatial features, and the false detection probability is reduced for regions with continuous spatial features. Finally, the category label, location coordinates and confidence score of each target in each frame are output. In the process of multi-target tracking, for the detected targets in consecutive frames, the three-dimensional motion consistency association method is adopted by combining the fused feature vector and the three-dimensional spatial coordinates of the target. The targets with the smallest spatial distance, consistent motion trend and highest fused feature similarity in adjacent frames are identified as the same target, assigned the same identity label, and the spatial trajectory is updated to generate target tracking data. Spatial behavior analysis is performed on target tracking data. Using three-dimensional spatial trajectories, the real-time position, speed and historical movement path of workers and equipment are calculated. High-risk operation, unauthorized movement and boundary crossing behavior characteristics are extracted to form spatial behavior data for on-site safety monitoring. The multi-target detection results after fusing feature vectors and spatial semantic injection are output in a unified manner with the target tracking data after being correlated with 3D motion consistency.

6. The method for video tracking and early warning at power construction sites according to claim 1, characterized in that, The generation of early warning data specifically includes: Real-time analysis of target tracking data and spatial behavior data is performed to statistically analyze the real-time position, velocity, acceleration, and historical trajectory of each target in three-dimensional space. The spatial distance between operators and equipment targets is monitored in real time. When the distance between any two targets is less than the set spatial safety threshold, it is judged as abnormal approach behavior. The system compares the regional attributes of the target at the work site in real time and judges whether the target has crossed the boundary, stayed for a long time, or entered a high-risk area based on the regional risk label. When the target stays in a high-risk area for more than a preset threshold, it is judged as a behavior of staying beyond the boundary. By combining the changes in the target's speed and acceleration, and through behavioral data analysis, high-risk operations, unauthorized movements, and sudden stops are identified as abnormal behavioral events, and corresponding abnormal behavior indicators are output. Multi-level early warning conditions are set for abnormal behavior events of different risk types. The events are divided into three levels: low, medium and high according to their risk level. Early warning signals are automatically generated for targets or behaviors that meet the early warning conditions. The occurrence time, involved objects, spatial location and event description of all abnormal behavior events and early warning signals are structured and organized to finally generate early warning data for on-site safety management and decision analysis.

Citation Information

Patent Citations

  • Monocular three-dimensional human body reconstruction method and device and electronic equipment

    CN116486009A

  • Monocular depth estimation system and estimation method based on self-supervised learning

    CN117541636A