Intelligent early warning method for approaching of high-risk equipment of thermal power plant
Through multi-source sensor data fusion and reverse optical modeling, the problems of low positioning accuracy and poor warning reliability caused by strong electromagnetic and optical interference in thermal power plants have been solved, and continuous reliable tracking and safety warnings in extreme environments have been achieved.
Patent Information
- Application Number
- CN202510950776.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-26
Smart Images

Figure CN120708359A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to power electronics technology, in particular to an intelligent early warning method for approaching high-risk equipment in a thermal power plant. Background Art
[0002] Currently, there are multiple technical paths for personnel positioning and safety warning in industrial environments. Positioning technologies based on wireless radio frequency, such as ultra-wideband (UWB) or radio frequency identification (RFID), achieve location tracking by having personnel wear tags and deploying positioning base stations or card readers in the factory. UWB technology can provide centimeter-level positioning accuracy, while RFID technology is often used for regional access management and identity recognition. In the field of visual positioning, some solutions use conventional RGB cameras for two-dimensional (2D) image processing and use target detection algorithms to identify people in the picture. A more advanced solution uses three-dimensional (3D) depth cameras such as binocular stereo vision or time of flight (ToF) to directly calculate the three-dimensional spatial coordinates of personnel by obtaining the depth information of the scene, thereby achieving more accurate positioning.
[0003] However, in power plant environments, such as large generators and transformers, short circuits can cause two difficult-to-handle problems:
[0004] The strong transient magnetic field generated by the device can cause nonlinear distortion in the depth camera's CMOS sensor, resulting in "depth drift" in the depth map at the moment the device starts and stops. This can lead to errors in the calculated distance to a person of up to 30-500cm. Existing technologies rely on static calibration or simple filtering, which cannot cope with dynamic distortion caused by rapid changes in magnetic field intensity (at the μs level).
[0005] The intense arc light (>10,000 lux) generated by electrical equipment failures can saturate camera sensors, creating a 0.5-2 second "blind spot." During this period, frightened individuals may make unpredictable escape movements, necessitating continuous tracking. Summary of the Invention
[0006] The purpose of the invention is to provide an intelligent early warning method for the approach of high-risk equipment in thermal power plants in order to solve the problems of low positioning accuracy and poor early warning reliability in the existing technology under strong electromagnetic and strong optical interference environments.
[0007] The technical solution is an intelligent early warning method for approaching high-risk equipment in thermal power plants, including:
[0008] Collect multi-source sensor data covering the operating area, identify and obtain the type of environmental interference, and extract RGB image streams;
[0009] When the environmental interference type is identified as optical interference, virtual depth reconstruction is performed based on the RGB image stream and historical trajectory data through reverse optical modeling and motion continuity constraints to generate a virtual depth map;
[0010] Fuse the virtual depth map and RGB image stream to determine the 3D spatial position of the person;
[0011] Combined with the pre-stored high-risk equipment location database, the dangerous distance between personnel and high-risk equipment is calculated, the risk level is assessed and an early warning signal is generated.
[0012] The beneficial effect is that it can effectively deal with complex interference in thermal power plants and significantly improve the robustness and safety of personnel positioning. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a flow chart of the present invention.
[0014] Figure 2 This is a flow chart of the virtual depth reconstruction of the present invention.
[0015] Figure 3 It is a flow chart of the present invention for inferring human body contour and corresponding contour confidence distribution.
[0016] Figure 4 It is a flow chart of the present invention for solving the inverse problem optimization model.
[0017] Figure 5 It is a flow chart of the present invention for predicting future motion to obtain a predicted trajectory sequence and corresponding motion uncertainty. DETAILED DESCRIPTION
[0018] In order to better understand the present invention, the present invention will be described in more detail below in conjunction with specific embodiments. It should be noted that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention. All other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0019] First, the technical concept of the present invention will be described.
[0020] In a preferred embodiment of the present invention, to address extreme scenarios where events such as electrical faults within a thermal power plant may simultaneously trigger combined interference such as strong magnetic fields and strong arc flashes, the system first identifies the current state of the dual interference environment through a comprehensive analysis of multi-source sensor data. Once successfully identified, the system activates a multimodal fusion mode, launching two independent technical processing paths in parallel: first, a physical model-based dual-time-scale dynamic magnetic field distortion compensation module is activated for the raw depth image stream and magnetic field intensity data, aiming to generate a compensated depth map that eliminates the effects of magnetic field distortion; second, completely independent of the depth sensor, a virtual depth reconstruction module based on inverse optical modeling is activated for the RGB image stream, aiming to infer and generate a virtual depth map from overexposed areas caused by strong light.
[0021] After generating two independent depth maps (i.e., the compensated depth map and the virtual depth map) in parallel, the system enters the intelligent selection and fusion processing link. For areas in the image where the difference in depth values between the two depth maps is small and there is no significant conflict, the system uses a Bayesian fusion framework to perform weighted fusion based on their respective confidence maps to obtain the statistically optimal depth estimation result. For conflicting areas where the difference in depth values between the two depth maps exceeds a preset threshold due to reasons such as arc light saturation, the system suspends fusion and instead executes a decision-making selection mechanism based on the type of interference. For example, in an area that is completely saturated by arc light, the system will determine that the virtual depth map is more credible in the area based on the identification of optical interference, and directly adopt its depth value.
[0022] The system seamlessly integrates the depth results of Bayesian fusion (applied to non-conflict areas) and the depth results of decision selection (applied to conflict areas) to generate a complete and logically consistent final fusion depth map. The complete processing flow from identification, parallel decoupling to conflict arbitration and fusion enables the present invention to maintain continuous and reliable tracking of personnel in harsh working conditions where both magnetic field and optical physical interferences exist simultaneously and are coupled with each other. It solves the technical pain points of the existing technology that causes perception blind spots or depth drifts due to the failure of a single compensation method under complex interference. Through the collaborative work of multiple paths and multiple modes, the safety redundancy and robustness of the early warning system at the most dangerous moments are improved.
[0023] It should be noted that if only one situation occurs, it can also be solved independently, that is, executing some steps, as described in the following embodiments.
[0024] Example 1: Describe the overall process of the intelligent early warning method for approaching high-risk equipment in a thermal power plant.
[0025] In an optional embodiment, the early warning method of the present invention can be implemented within a system consisting of a distributed sensor network and a central processing server. The sensor network covers key operational areas within a thermal power plant and may include an RGB camera, a depth camera (e.g., a Time-of-Flight (ToF) camera), a three-axis magnetic field intensity sensor, and a multi-point light intensity sensor. The central processing server is used to execute the early warning method of the present invention.
[0026] The intelligent early warning method of the present invention mainly includes the following steps:
[0027] Step 1: Collect multi-source sensor data covering the work area, identify the interference type of the current environment, and obtain the environmental interference type identification and RGB image stream.
[0028] Specifically, the distributed sensors throughout the plant are first synchronized and calibrated at the nanosecond level using the IEEE 1588 precision time protocol. This assigns a unified synchronized timestamp to all collected data, outputting time-aligned multi-source data streams. In parallel, the system collects raw RGB images (e.g., 1920×1080 resolution, 30fps), raw depth images (e.g., 640×480 resolution, 30fps), three-axis magnetic field intensity values (e.g., 100Hz sampling), and multi-point light intensity data (e.g., 1000Hz sampling) from each sensor. After acquisition, the data is preprocessed, including mapping the data from different sensors to a unified world coordinate system through multi-view geometric transformations, and performing denoising and filtering on each type of data.
[0029] The IEEE 1588 clock synchronization process is as follows: Upon system startup, a single clock source is selected as the Grandmaster Clock (GrandmasterClock) within the plant's sensor network using the Precision Time Protocol (PTP) "Best Master Clock Algorithm." Subsequently, each slave sensor periodically exchanges PTP messages with the master clock, including "Sync," "Follow_Up," "Delay_Req," and "Delay_Resp." By parsing the timestamp information in these messages, each slave clock independently calculates the time offset from the master clock and the network transmission path delay. Based on this information, each slave clock continuously adjusts its local clock frequency and phase, ultimately achieving nanosecond-level synchronization with the master clock and providing a unified time reference for all subsequently collected data.
[0030] The multi-view geometric transformation algorithm is as follows: During the system deployment phase, camera calibration techniques such as the checkerboard calibration method are used to pre-obtain the intrinsic parameter matrix K (containing focal length, principal point coordinates, etc.) and extrinsic parameter matrix [R|T] (containing the rotation matrix R and translation vector T relative to the world coordinate system) for each camera (including RGB cameras and depth cameras). At runtime, any three-dimensional point P_cam in the camera coordinate system can be converted to a point P_world in the unified world coordinate system using the coordinate transformation formula P_world = R*P_cam + T, thereby achieving spatial alignment of sensor data from different positions and postures.
[0031] Furthermore, by comprehensively analyzing the filtered magnetic field data (for example, calculating the spatial gradient gradB and time rate of change dB / dt of the magnetic field intensity) and stable light intensity data (for example, detecting sudden changes in light intensity), the interference type of the current environment is identified in real time, and a composite interference type identifier (for example, normal / magnetic interference / light interference / dual interference) is generated. In this embodiment, the accurate identification of environmental interference can provide a decision-making basis for the subsequent adaptive selection of the optimal processing path, thereby improving the robustness of the entire early warning system.
[0032] Furthermore, it also includes calculating the eigenvalues of the magnetic field gradient tensor field. Specifically, the system first calculates all nine partial derivatives (such as dBx / dx, dBx / dy, etc.) of the three components of the magnetic field Bx, By, and Bz in the three coordinate axis directions based on the distributed magnetic field sensor data through the difference method. Then, a 3x3 magnetic field gradient tensor grad B is constructed. In order to calculate the eigenvalues of this tensor, the system needs to solve its characteristic equation det(grad B-λI)=0, where I is the unit matrix and λ is the eigenvalue to be calculated. This is a cubic polynomial equation about λ. Solving this equation can obtain three eigenvalues {λ1, λ2, λ3}, which respectively characterize the spatial rate of change of the magnetic field in three mutually orthogonal main directions.
[0033] Step 2: When the environmental interference type is identified as optical interference, virtual depth reconstruction is performed based on the RGB image stream and historical trajectory data through inverse optical modeling and motion continuity constraints to generate a virtual depth map and uncertainty distribution.
[0034] When the system detects optical interference such as strong light or arc light in the current environment, traditional depth cameras will fail due to sensor saturation. At this point, the system will no longer rely on the failed depth image and will instead activate the virtual depth reconstruction module. This module uses the unaffected RGB image stream to infer the outline of the person by analyzing the morphological features of the overexposed area, and combines it with the person's historical movement trajectory for constraints, ultimately generating a virtual depth map that does not rely on direct depth measurement. The detailed implementation process of this step will be expanded in Example 2.
[0035] Step 3: Fuse the virtual depth map, uncertainty distribution, and RGB image stream to determine the 3D spatial position of the person.
[0036] In this step, the system fuses the depth information generated in the previous step (whether from virtual depth reconstruction or from magnetic field distortion compensation described in subsequent embodiments) with the RGB image stream. For example, the YOLOv8 model can be used to detect people in the RGB image stream to obtain a 2D bounding box of the person. This 2D bounding box is then projected onto the depth map, and the RANSAC algorithm can be used to fit the point cloud within the projected area to eliminate background interference, ultimately accurately calculating the 3D spatial coordinates (X, Y, Z) of the person's center of mass.
[0037] Step 4: Combined with the pre-stored high-risk equipment location database, calculate the dangerous distance between personnel and high-risk equipment, and evaluate the risk level.
[0038] The system pre-stores the 3D location of high-risk equipment and its danger radius, R_danger (e.g., 5 meters for high-voltage equipment and 3 meters for rotating equipment). This step reads the 3D position of the individual and calculates the Euclidean distance, D_risk, between them and the nearest high-risk equipment. By comparing D_risk with R_danger, for example by calculating the relative risk S = R_danger / D_risk, the current situation can be quantified into a specific risk level, such as 1-5.
[0039] Step 5: Generate and output corresponding warning signals based on the risk level.
[0040] This step aims to translate the assessed risk level into specific early warning actions. For example, the system can trigger a hierarchical early warning response based on the risk level: Level 1 (Safety) only logs, Level 2 (Caution) triggers a yellow alert, Level 3 (Warning) triggers an orange alert, Level 4 (Danger) triggers a red alert accompanied by an audible and visual alarm, and Level 5 (Emergency) directly outputs an emergency stop command. In this way, a complete early warning process from hazard perception to closed-loop control is implemented, ensuring the safety of operators.
[0041] Existing vision technologies are unable to effectively handle the complete information loss caused by extreme optical overload conditions. The following describes the situation: During welding operations, switchgear operation, or faults in thermal power plants, high-voltage arcs are generated, generating a momentary burst of intense light that is too bright for the human eye to perceive. The brightness far exceeds the capabilities of conventional HDR (High Dynamic Range) cameras. This extreme optical overload can cause large areas of the camera sensor to saturate, resulting in a pale, overexposed area. All real-world scene information within this area, including the outlines and positions of people, is completely obscured and irrecoverable. Traditional approaches attempt to recover this information through image post-processing or simply mark the overexposed area as invalid, but both approaches fail. The former is impossible due to the physical loss of information, while the latter results in the early warning system being blinded by this information loss at the most dangerous moment of the arc, creating a critical safety blind spot and preventing effective detection and warning of people near the intense light.
[0042] Example 2: Based on the above example, the specific steps of performing virtual depth reconstruction under optical interference conditions are described.
[0043] In one embodiment of the present invention, the step of performing virtual depth reconstruction includes:
[0044] First, the RGB image stream is analyzed to identify the overexposed areas in the image to generate an overexposed area mask, and the precise boundaries of the overexposed areas are extracted to obtain the overexposed boundary features.
[0045] Specifically, the system identifies overexposed areas in RGB images by setting a higher brightness threshold (for example, for 8-bit grayscale images, the brightness value is greater than 250) and combining it with connected domain analysis. The identified areas are marked to generate overexposed area masks. Then, the Canny edge detection algorithm can be used to accurately extract the boundaries of these overexposed areas and calculate the geometric and optical features of the boundaries, such as curvature and gradient distribution. These features are collectively referred to as overexposed boundary features. In this embodiment, traditional methods usually regard overexposed areas as invalid information and discard them, while the present invention innovatively regards them as key clues to infer the contours of obscured objects.
[0046] Secondly, using the overexposed area mask and overexposed boundary features, an inverse optical diffusion model is constructed and solved to reversely infer the human contour of the obscured person from the boundary shape of the overexposed area, and output the inferred human contour and contour confidence distribution.
[0047] To infer the true human contour from the overexposure result, the present invention establishes an inverse problem optimization model that represents the observed overexposed saturated image I_saturated as the convolution result of the true human contour I_real and an optical diffusion kernel function PSF.
[0048] Preferably, the objective function J(I real ) can be constructed as:
[0049] J(I real )=||I saturated -PSF OI real |||2+λ1×TV(I real )+λ2×||grad I real ∣∣1;
[0050] Among them: I saturated Represents the observed overexposed and saturated image.
[0051] I real represents the true human silhouette image to be solved without overexposure. PSF stands for the optical diffusion kernel function, which describes the diffusion effect of a point light source in the imaging system. In an optional embodiment, the PSF can be adaptively constructed based on the ambient light intensity L to simulate the overexposure diffusion behavior under different lighting conditions. O represents the convolution operation. || I saturated -PSF OI real || 2 Represents the data fidelity term, whose goal is to minimize the boundary error between the saturated image predicted by the model and the saturated image actually observed. TV(I real ) represents the total variation regularization term, which is used to encourage the restored contour edges to be smooth and continuous. ||grad I real ||1 represents the L1 norm regularization term of the gradient, which helps produce sparse gradients and thus obtain clearer contour edges. λ1 and λ2 are regularization parameters used to balance the weight of the data fidelity term and the regularization term during the optimization process. For example, they can be set to λ1 = 0.01 and λ2 = 0.005.
[0052] Where, PSF(u,v)=(1 / (2πσ x σ γ sqrt(1-ρ 2 )))×exp(-Q(u,v) / (2(1-ρ 2 )));
[0053] Q(u,v)=((u-u0) / σ x ) 2 -2ρ((u-u0) / σ x )((v-v0) / σ γ )+((v-v0) / σ γ ) 2 ;
[0054]
[0055] u, v are the pixel coordinates in the image coordinate system; u0, v0 are the PSF center coordinates (the position of the overexposure source in the image); σ x ,σ γ is the standard deviation (diffusion radius) in the x and y directions; ρ (rho) is the correlation coefficient, ranging from -1 to 1, describing the degree of rotation of the ellipse; L is the ambient light intensity; L0 is the reference light intensity; θ (theta) is the angle of the light source relative to the camera; is the basic diffusion radius, typical value Pixels, Pixel; ρ0 is the basic correlation coefficient, with a typical value of 0.1; κ x ,κ γ is the light intensity sensitivity coefficient, typical value κ x =0.25,κ γ =0.20; is the light intensity sensitivity of the correlation coefficient, with a typical value of 0.15; δ x ,δ γ is the angle sensitivity coefficient, typical value δ x =0.30,δ γ =0.25; is the angular sensitivity of the correlation coefficient, with a typical value of 0.10.
[0056] When high light intensity saturation correction is performed, PSF_saturated(u,v)=PSF(u,v)×[1+η×tanh((L-L_sat) / L_trans)];
[0057] η is the saturation enhancement factor, with a typical value of 0.5; L_sat is the saturation light intensity threshold, with a typical value of 15000 lux; L_trans is the transition interval width, with a typical value of 3000 lux.
[0058] Actual arcs are not ideal point light sources; their shapes are often elongated or have specific directionality. This results in the overexposed areas on the camera sensor not being perfectly circular, but rather exhibiting a certain degree of anisotropy. The correction factor A(θ) = 1 + 0.2 × cos(2θ) mathematically models this directional effect. It gives the point spread function (PSF) different diffusion widths in different directions, thus more realistically simulating the elliptical or spindle-shaped overexposed spots caused by strong directional light sources such as arcs.
[0059] To solve the aforementioned inverse optimization model, in a preferred embodiment of the present invention, the alternating direction method of multipliers (ADMM) is used to iteratively solve the model to generate an intermediate human body contour. The ADMM algorithm efficiently decomposes the original complex optimization problem into multiple, more easily solvable subproblems.
[0060] Furthermore, during the ADMM's iterative solution, portions of the intermediate human silhouette that do not conform to predefined body shape characteristics (for example, shoulder width should be within the range of 30-80 cm) are periodically corrected (for example, every 10 iterations) based on a predefined prior model of human shape. This correction incorporates prior knowledge of the physical world into the mathematical solution, effectively avoiding unreasonable solutions and improving the accuracy of the silhouette inference.
[0061] When the iteration meets the convergence condition (for example, the relative change rate of two adjacent iteration results is less than the preset threshold value 0.001), the corrected contour is output as the final inferred human body contour I final .
[0062] After obtaining the inferred human body contour, the system will evaluate its credibility. Specifically, based on the boundary consistency between the inferred human body contour and the optical features of the overexposed area, the shape rationality with the human body shape prior model, and the temporal consistency of the contours of adjacent frames, a comprehensive evaluation is performed to generate the contour confidence distribution C. total The confidence distribution will serve as the weight basis for subsequent data fusion.
[0063] The boundary consistency score S_boundary (weight 0.4) is given the highest weight, directly reflecting the degree of agreement between the reconstructed contour and the observational evidence in the original image, and is the most basic guarantee of reliability. The shape rationality score S_shape (weight 0.3) is second, leveraging prior knowledge to ensure the physical plausibility of the result. The temporal consistency score S_temporal (weight 0.3) is also given equal weight, as it leverages motion continuity to smooth the result and suppress occasional errors in a single frame.
[0064] Subsequent steps, such as trajectory prediction based on historical trajectory data and ultimately generating a virtual depth map, will be further described in other embodiments.
[0065] When solving the inverse problem under the ADMM framework, the original problem is decomposed. Among them, the subproblem containing the I_real quadratic term (I-subproblem) is a large quadratic programming problem, which can be solved efficiently and iteratively using the conjugate gradient method (CG). The CG method searches for the optimal solution by constructing a series of conjugate directions, and its convergence speed is faster than ordinary gradient descent. The subproblem containing the TV or L1 regularization term (Z-subproblem) has an analytical solution and can be solved in one step using a soft-thresholding operator. The operator is in the form of soft(x,T)=sign(x)*max(|x|-T,0), and its function is to shrink the input signal, setting the components whose absolute value is less than the threshold T to zero, thereby achieving sparsity.
[0066] Example 3: Based on the above example, the specific steps of predicting future motion based on historical trajectory data to obtain a predicted trajectory sequence and corresponding motion uncertainty are described.
[0067] In one embodiment of the present invention, the step of predicting future motion based on historical trajectory data specifically includes:
[0068] First, the system analyzes the speed and acceleration changes in the historical trajectory data to identify the person's current motion state label, which includes stationary, straight or turning. Specifically, the system reads the historical trajectory data of the person's three-dimensional position located in the previous frame (for example, the first 5 frames, covering about 500ms)
[0069] {P1,P2,...,P5}. Based on these positions and the corresponding timestamps, the system calculates the instantaneous velocity sequence
[0070] V i and acceleration sequence a i By analyzing these dynamic parameters, for example, when the velocity amplitude
[0071] When |V| is less than 0.2 m / s, the current motion state can be determined to be stationary; when the velocity amplitude is between 0.2 and 1.5 m / s and the direction change rate dθ / dt is less than 30° / s, it can be determined to be moving straight; when the direction change rate dθ / dt is greater than 30° / s, it is determined to be turning. In a preferred embodiment, a hidden Markov model (HMM) can be used to smooth the initially identified state sequence to output a more stable and logical current motion state label.
[0072] Secondly, based on the current motion state label, a state-adaptive human kinematic model is constructed. This model sets different motion parameters for different motion states. In this embodiment, in order to make the trajectory prediction closer to the real physical world, the model parameters will be adaptively adjusted according to the motion state identified in the previous step. For example, the model will set different upper limits for motion parameters for different states:
[0073] When the state is stationary, set the maximum acceleration a_max to 1m / s 2 When the state is straight, set the maximum acceleration a_max to 3m / s 2 When the state is turning, set the maximum acceleration a_max to 5m / s 2 , allowing for faster speed changes. Furthermore, the model can account for the impact of equipment like helmets and overalls on motion flexibility, for example by reducing motion parameters by 10%. This state-adaptive model, constructed in this way, can more accurately constrain subsequent trajectory predictions than traditional models using fixed parameters.
[0074] Thirdly, the extended Kalman filter (EKF) is used in combination with the state-adaptive human kinematics model to process the historical trajectory data and generate a predicted trajectory sequence and a predicted covariance matrix sequence. The system uses the extended Kalman filter to fuse historical trajectory information and predict the future. The state vector x of the EKF can include the three-dimensional position, velocity and acceleration components of the person, that is, In its state transfer equation x(k+1)=F×x(k)+G×u(k)+w(k), the state transfer matrix F will contain the constraints of the state adaptive kinematic model constructed in the previous step. EKF performs the prediction step: x*(k+1|k)=F×x*(k|k) and P(k+1|k)=F×P(k|k)×F T +Q, iteratively predict the motion trajectory for a period of time in the future (for example, 2 seconds), thereby generating a predicted trajectory sequence {P pred (t+0.1s),...,P pred (t+2s)} and the corresponding prediction covariance matrix sequence.
[0075] Finally, based on the prediction covariance matrix sequence, the motion uncertainty is modeled. The prediction covariance matrix P represents the uncertainty of the prediction results. To convert it into a more intuitive representation, the system extracts the position-related 3x3 submatrix Σ_pos from the prediction covariance matrix P at each moment. By performing eigenvalue decomposition on this position covariance submatrix: Σ_pos = V × Λ × V T , its eigenvalues (diagonal elements of Λ) and eigenvectors (columns of V) can be obtained. Based on these decomposition results, the geometric parameters of a three-dimensional confidence ellipsoid (or uncertainty ellipse in a two-dimensional plane) can be calculated. For example, the major axis a and minor axis b of a 95% confidence ellipse can be determined by the maximum and minimum eigenvalues λ_max and λ_min, respectively, and its orientation is determined by the corresponding eigenvectors. This ellipse intuitively represents the possible spatial range of the predicted person's position, which is the motion uncertainty generated by the model.
[0076] According to one aspect of the present application, the ellipse parameter calculation is specifically as follows: performing eigenvalue decomposition on the position covariance submatrix Σ_pos=VΛV TAfter that, the diagonal elements {λ_max,λ_min,...} of the resulting diagonal matrix Λ are the squares of the variances of the ellipse (or ellipsoid) in the directions of the various principal axes, and the column vectors of the eigenvector matrix V define the directions of these principal axes. The lengths of the semi-major and semi-minor axes of the confidence ellipse are proportional to the square roots of the eigenvalues, respectively. For example, in two dimensions, the semi-major axis a and semi-minor axis b of the 95% confidence ellipse are calculated as a = c*sqrtλ_max and b = c*aqrtλ_min, where c is a constant related to the confidence level. For a two-dimensional Gaussian distribution, the c value corresponding to a 95% confidence level is approximately 2.447. The angle of rotation of the ellipse is determined by the direction of the eigenvector v_max corresponding to the largest eigenvalue.
[0077] Example 4: Based on the above examples, the whole process of combining the inferred human body contour and the predicted trajectory sequence to finally generate a virtual depth map with safety assurance is described.
[0078] In one embodiment of the present invention, the steps of generating a virtual depth map by combining the inferred human body contour and the predicted trajectory sequence, and fusing the contour confidence distribution with the motion uncertainty to generate an uncertainty distribution include:
[0079] First, combined with the 3D scene model, an initial virtual depth map is generated by raycasting the inferred human body contour and constraining it using the predicted trajectory sequence. Specifically, for each boundary point (u i ,v i ), the system will combine the known camera intrinsic parameter matrix K and the camera center position C to construct a 3D ray starting from the camera center and passing through the point: r(t)=C+t×K-1×[u i ,v i ,1] T At the same time, the system reads the pre-stored three-dimensional scene model of the power plant (including triangular meshes of equipment, walls, and ground), and combines scene constraints such as human height (e.g., 1.5m-2.0m) and ground position (e.g., z=0) to limit the effective search range of each ray.
[0080] In order to determine the most likely depth value for each point on the contour, the system uses the predicted trajectory sequence {P_pred} generated in Example 3 as a strong constraint. For example, for each ray r(t), the system calculates the minimum distance d_min between it and the currently predicted position of the person P_current. The maximum a posteriori estimation (MAP) is used to solve an optimal depth value, which should make its corresponding point on the ray closest to the predicted position while satisfying the scene constraints. After performing this operation for all points on the contour, an initial virtual depth map D_virtual can be generated.
[0081] Secondly, the contour confidence distribution and motion uncertainty are integrated to calculate a comprehensive uncertainty, which is the uncertainty distribution. The system will comprehensively evaluate the uncertainty in the entire virtual depth reconstruction process. This comprehensive uncertainty σ_total can be composed of multiple parts,
[0082] For example: σ_total=sqrt(σ 2 _motion+σ 2 _recon+σ 2 _system);
[0083] Wherein: σ_motion is the motion uncertainty generated by the third embodiment (eg, the size of the uncertainty ellipse).
[0084] σ_recon is the contour inference uncertainty generated by the second embodiment (related to the contour confidence C_total).
[0085] σ_system is a constant representing the inherent error of the system, for example 0.05 meters.
[0086] Finally, based on the initial virtual depth map and the comprehensive uncertainty, a preset safety factor is applied to calculate the safe envelope depth, and this safe envelope depth is used as the final output virtual depth map. This step is the key to ensuring safety redundancy in the present invention. The system does not directly use the initial virtual depth map D_virtual, but instead introduces the concept of a safe envelope. The specific calculation formula is: D_safe = D_virtual - k × σ_total;
[0087] Where: D_safe is the final output safety envelope depth. D_virtual is the initial virtual depth map. σ_total is the combined uncertainty calculated in the previous step. k is a preset safety factor. This factor can be dynamically adjusted based on the risk level. For example, k = 1.5 in general areas and k = 2.0 or higher in high-risk areas to provide a larger safety margin.
[0088] In this embodiment, by subtracting a safety distance proportional to the uncertainty from the geometrically inferred depth value, the resulting safe envelope depth D_safe is a conservative but more reliable estimate. This design philosophy of erring on the side of conservatism over risk ensures that even in the presence of various uncertainties, the depth information output by the system can maximize personnel safety. This is of vital practical significance in high-risk operating environments such as thermal power plants.
[0089] According to one aspect of the present application, the problem is that existing visual positioning technologies that rely on semiconductor sensors lack the inherent physical mechanism to compensate for strong dynamic electromagnetic field interference, as described below. Thermal power plant equipment, such as generators, main transformers, and high-voltage switchgear, generates magnetic fields with intensities reaching hundreds of gauss and rapidly varying with load. This magnetic field is not simply background noise; it can directly act on photogenerated charge carriers within the depth camera's CMOS / CCD sensor through the Lorentz force, deflecting their motion trajectory and causing spatial distortion and crosstalk in the pixel signal. For ToF cameras, the magnetic field also affects the response characteristics of their internal circuits, introducing additional phase measurement delays. Existing technologies typically attempt to smooth erroneous depth data using general algorithms such as Kalman filtering. However, this does not address the root cause of the problem and cannot eliminate distortion at the physical level. Furthermore, when the magnetic field changes rapidly, the compensation effect is severely delayed, resulting in meter-level deviations in positioning results, seriously compromising the reliability of the early warning system. To this end, the following embodiments are provided.
[0090] Embodiment 5: After the above embodiment describes the scenario of coping with optical interference, another process is described in detail: when the environmental interference type is identified as magnetic field interference, a specific processing method.
[0091] In one embodiment of the present invention, the method further includes: when the environmental interference type is identified as magnetic field interference, performing dual-time-scale dynamic magnetic field distortion compensation on the collected original depth image based on magnetic field intensity data to generate a compensated depth map and a depth confidence map.
[0092] In a thermal power plant environment, generators, transformers, and other equipment generate strong, rapidly changing magnetic fields (e.g., with intensities of 50-500 gauss and rates exceeding 10 gauss / second). These magnetic fields can directly affect the charge transfer paths of the CMOS chips within depth sensors (such as ToF cameras) or introduce circuit delays in phase measurements during time-of-flight ranging, resulting in severe distortion of depth measurements.
[0093] Specifically, the steps of performing dual-time-scale dynamic magnetic field distortion compensation include:
[0094] First, in the fast channel, a simplified linear model is used to perform immediate coarse compensation on the original depth image based on the magnetic field intensity data to generate a fast compensated depth map and a fast channel confidence map.
[0095] This channel aims to achieve millisecond-level instantaneous response to sudden magnetic field changes. Specifically, it reads the real-time magnetic field strength value B(t) and rate of change dB / dt and quickly calculates the compensation using a simplified linear model, for example: ΔD_fast = B(t) × K_instant × sign(dB / dt).
[0096] Where ΔD_fast is the depth distortion compensation calculated by the fast channel, K_instant is an empirical constant (e.g., 0.15 cm / Gauss), and sign(dB / dt) indicates the direction of magnetic field change. This calculation process can be completed in a very short time (e.g., within 5 milliseconds), providing a rapid, rough compensation for the entire depth map, generating a fast-compensated depth map D_fast. Simultaneously, this channel assesses its uncertainty based on the rate of magnetic field change, for example, σ_fast = |dB / dt| × 0.02, and generates a corresponding fast channel confidence map.
[0097] Secondly, in the precise channel, based on the magnetic field intensity data, the original depth image is accurately compensated by iteratively solving the magnetic-optical coupling physical model to generate a precise compensated depth map and a precise channel confidence map.
[0098] The goal of this channel is to pursue the highest accuracy of compensation. It does not use a simplified model, but is based on a magneto-optical coupling physical model established from the first principles of physics to describe the unified nonlinear relationship between the depth distortion and the magnetic field strength and direction (the specific construction will be described in detail in Example 7). The system uses numerical optimization methods such as the improved Levenberg-Marquardt (LM) algorithm to iteratively solve this complex physical model to obtain a highly accurate distortion compensation value, thereby generating a precise compensated depth map D_precise. Although this process has high accuracy, it requires a certain amount of calculation time (for example, convergence requires 50 milliseconds).
[0099] Finally, a time-varying weight function is used to combine the fast channel confidence map and the precise channel confidence map to perform weighted fusion on the fast compensated depth map and the precise compensated depth map to generate the compensated depth map and depth confidence map.
[0100] To balance the real-time performance of the fast channel with the accuracy of the precise channel, the present invention designs an adaptive fusion mechanism. Specifically, a weight function w(t) = exp(-t / τ) that decays exponentially with time is used to calculate the fusion weights of the two channels, where t is the time elapsed since the magnetic field change began and τ is a time constant (e.g., 20 milliseconds). The final fusion depth value D_final is obtained through weighted averaging: D_final = w(t) × D_fast + (1-w(t)) × D_precise.
[0101] In this embodiment, when the magnetic field begins to change (t is very small), w(t) approaches 1, and the system relies primarily on the results of the fast channel, ensuring a fast response. Over time, the calculations of the precise channel gradually complete, and its results become more reliable. At this point, w(t) decreases, and the weight of the precise channel (1-w(t)) increases, allowing the system to smoothly transition to using more accurate compensation results. Through this dual-channel collaborative operation and adaptive fusion mechanism, the present invention achieves an effective balance between real-time performance and accuracy when responding to dynamic magnetic field interference.
[0102] Accordingly, in subsequent steps, determining the person's 3D spatial position involves intelligently selecting and fusing the compensated depth map and the virtual depth map, along with their corresponding depth confidence maps and uncertainty distributions, to determine the person's 3D spatial position. Details of this fusion step will be described in subsequent embodiments.
[0103] Example 6: Based on the above-mentioned Example 5, a preferred implementation of the precise channel is described in detail, especially how to overcome the calculation delay through the pre-compensation concept.
[0104] In one embodiment of the present invention, in the precise channel, the step of iteratively solving a magneto-optical coupling physical model to generate a precise compensation depth map and a precise channel confidence map includes:
[0105] Firstly, based on the temporal changes of magnetic field intensity data, a magnetic field predictor is constructed and applied to generate a sequence of magnetic field predictions at future moments.
[0106] In order to overcome the inherent delay caused by the iterative solution of the precise channel (for example, 50 milliseconds), the present invention innovatively introduces a magnetic field prediction mechanism. Specifically, the system reads the filtered magnetic field data sequence {B(ti×Δt)} within a recent period of time (for example, 100 milliseconds) and uses an autoregressive model AR(3) to fit the temporal variation of the magnetic field. Based on the fitted model, the system can calculate the first-order derivative dB / dt and the second-order derivative d 2 B / dt 2 , and use a second-order Taylor expansion predictor to predict the magnetic field strength at the next τ time: B_pred(t+τ)=B(t)+dB / dt×τ+0.5×d 2 B / dt 2 ×τ 2 .
[0107] By using this formula, the system can predict the magnetic field prediction sequence {B_pred} for the next 10 to 50 milliseconds. In this embodiment, the prediction capability is the key to achieving pre-compensation and eliminating delay.
[0108] Secondly, an improved Levenberg-Marquardt algorithm is adopted, and the magnetic field prediction sequence is used as input to iteratively solve the magneto-optical coupling physical model, so as to calculate the distortion compensation value in advance and obtain an accurate compensated depth map.
[0109] In conventional methods, the LM algorithm uses the magnetic field data at the current time t for calculation. However, by the time the calculation is complete, the actual magnetic field has changed, resulting in compensation lag. In a preferred embodiment of the present invention, the improved LM algorithm uses the magnetic field prediction sequence {B_pred} generated in the previous step for future time points as input.
[0110] Specifically, the LM algorithm solves the linear system (J_reg+μ k × I) × δ = -F(ΔD k ) to obtain the update step size δ, where J_reg is the regularized Jacobian matrix and μk is the damping factor. The algorithm adaptively adjusts μk based on the gain ratio ρ to accelerate convergence while ensuring stability. Because the input is magnetic field data at a future time, the entire iterative solution process is actually a pre-calculation of the distortion compensation value required at that future time. When the iteration is completed at the current time, the calculated compensation value corresponds exactly to the magnetic field state at a certain time in the future, effectively offsetting the algorithm's inherent computational time and achieving zero or low-latency compensation.
[0111] The gain ratio ρ is used to evaluate the cost-effectiveness of the model improvement in one iteration of the LM algorithm. Its calculation formula is: ρ=(F(ΔD k )-F(ΔD k +δ)) / (L(0)-L(δ)), where the numerator F(ΔD k )-F(ΔD k +δ) represents the actual error reduction, and F(ΔD) is the residual sum of squares. The denominator L(0)-L(δ) represents the error reduction predicted by the linear model, which can be expanded to δ T ×(μ k ×δ-J T ×F(ΔD k )). Here δ is the update step size obtained by solving the linear system, μ k is the current damping factor, and J is the Jacobian matrix. The calculation result of the gain ratio ρ directly determines the adjustment strategy of the damping factor μ in the next step.
[0112] Finally, according to the quality of the iterative solution, an accurate channel confidence map is generated.
[0113] The confidence of the accurate compensation result depends not only on the depth value after compensation, but also on the reliability of the calculation process itself. Therefore, the system will evaluate the confidence based on the convergence quality of the iterative solution. For example, the compensation confidence C_precise of each pixel can be calculated by combining the iteration quality index Q_iter (reflecting the stability and convergence speed of the iterative process) and the error standard deviation σ_pred of the magnetic field prediction. A compensation result that is smooth, converges quickly and is based on high-precision prediction will have a higher confidence. This confidence map provides key weight information for the dual-channel fusion process described in the subsequent Example 5.
[0114] Example 7: Based on the above examples, a detailed and fundamental explanation is given of the magneto-optical coupling physical model underlying the precise compensation channel described in Examples 5 and 6. Constructing a model based on the first principles of physics is the basis for achieving high-precision magnetic field distortion compensation.
[0115] In one embodiment of the present invention, the magneto-optical coupling physical model used in the precise channel includes:
[0116] The spatial offset submodule is used to model the spatial offset effect of the magnetic field on the charge transfer path inside the depth sensor through the Lorentz force.
[0117] Specifically, this submodule starts from the microscopic physical process inside the sensor. In a CMOS or CCD image sensor, incident photons generate electron-hole pairs in the photosensitive layer of the silicon wafer. Under the action of the electric field E inside the sensor, these photogenerated electrons drift toward the collecting electrode with a drift velocity of v. When there is an external magnetic field B, the moving electrons are affected by the Lorentz force F = e × (E + v × B), causing their motion trajectory to be deflected. By numerically integrating the electron's equation of motion, the spatial offset Δr accumulated by the charge during the entire drift process can be calculated. This spatial offset within the pixel will cause part of the charge that should have been collected by a pixel to leak to adjacent pixels, or collect charges from adjacent areas, thereby causing distortion of the image signal. This is one of the fundamental reasons why the magnetic field affects depth images.
[0118] The circuit delay submodule is used to model the circuit delay effect introduced by the magnetic field on the phase measurement during the time-of-flight (ToF) ranging process.
[0119] Specifically, this submodule targets the ranging principle of ToF type depth cameras. ToF cameras calculate distance by measuring the phase difference between emitted light and reflected light. The magnetic field not only introduces distortion through the above-mentioned spatial offset effect, but also introduces a delay to the phase measurement itself. This delay Δt can be divided into two parts: one is the charge transfer delay Δt_charge caused by the change in the charge transfer path, and the other is the circuit response delay Δt_circuit caused by the influence of the magnetic field on the response characteristics of the sensor readout circuit. The total phase measurement offset Can be modeled as Where f_mod is the modulation frequency of the camera. This unintended phase shift will be incorrectly interpreted by the camera as a depth measurement error ΔD, the magnitude of which can be determined by Calculated, where c is the speed of light.
[0120] The compensation algorithm for ToF phase measurement is as follows: The compensation algorithm is based on the phase offset model Δφ=2π×f_mod×(Δt_charge+Δt_circuit) and the depth error model The compensation process first calculates the expected depth measurement error ΔD_err based on the measured magnetic field strength and the pre-calibrated delay coefficients k1 and k2. This error is then subtracted from the raw depth value D_raw output by the ToF camera, resulting in the compensated depth value D_compensated = D_raw - ΔD_err. This process is performed pixel by pixel, resulting in a depth map that has been compensated for the phase measurement error.
[0121] Numerical integration method of charge transfer path of CMOS sensor, specifically: charge motion equation m_e*d 2 r / dt² = e*(E + dr / dt × B) is a second-order ordinary differential equation. To solve this equation, the system uses numerical integration methods, such as the fourth-order Runge-Kutta (RK4) method, which is preferred for its stability and accuracy. This method divides the entire drift time of the charge, t_drift, into multiple small time steps h. Within each time step, the position r and velocity v of the charge are updated by calculating the weighted average of the slopes (k1, k2, k3, k4) at four different points. This iteratively and accurately tracks the complete motion trajectory of the charge under the combined action of the electric and magnetic fields, ultimately determining its total offset Δr on the collecting electrode.
[0122] The unified parameterization submodule is used to integrate the spatial offset effect with the circuit delay effect to establish a unified nonlinear relationship between the depth distortion and the magnetic field strength and direction.
[0123] This submodule aims to integrate the above two physical effects and establish an end-to-end mathematical model from the macroscopic magnetic field input (intensity and direction) to the final depth distortion output. Since the physical processes of the two effects are highly nonlinear, direct solution is very complicated. Therefore, in a preferred embodiment, the unified relationship is parameterized as a multivariate polynomial function whose variables are the three-dimensional components of the magnetic field B_x, B_y, B_z and the direction angle of the magnetic field For example, the model can be expressed as: Among them, the coefficient {a ijk The distortion compensation parameter matrix A is constructed using the parameterized model. This parameter matrix A can be obtained through a one-time calibration of the camera in a laboratory environment. After calibration, the matrix is embedded in the system. In subsequent actual operation, the precision compensation channel uses this parameterized model and measured magnetic field data to calculate accurate depth distortion compensation values through an iterative solution.
[0124]
[0125] Among them, θ (theta) is the angle between the magnetic field direction and the horizontal plane (pitch angle), ranging from [0,π], unit: radian; φ (phi) is the angle between the projection of the magnetic field direction on the horizontal plane and the x-axis (azimuth angle), ranging from [0,2π], unit: radian; α1, α2, α3 are pitch angle influence coefficients, reflecting the influence of the vertical component of the magnetic field on the depth distortion; β1, β2, β3 are azimuth angle influence coefficients, reflecting the influence of the horizontal component of the magnetic field on the depth distortion; γ1, γ2 are cross-coupling coefficients, reflecting the composite effect of the magnetic field direction; A(θ) is the pitch angle response function; B(φ) is the azimuth angle response function; C(θ,φ) is the cross-coupling function;
[0126] Specifically, the polynomial parameter fitting is as follows: Place the depth camera in an environment where the magnetic field strength and direction can be precisely controlled, and collect a large number of (depth true value, depth measurement value) data pairs under different magnetic field conditions. Then, the polynomial distortion model (where i+j+k≤5) is used as the fitting function, and the least squares method is used to solve the parameter {a ijk To improve robustness, Ridge Regression with L2 regularization can be used for fitting to prevent overfitting. The final 15 polynomial coefficients are the distortion compensation parameter matrix.
[0127] The least-squares update method works as follows: When each new magnetic field data sample point B(t) arrives, the RLS algorithm uses a recursive formula to update the covariance matrix and model parameters without recalculating the entire historical data set. It first calculates the prediction error for the new data. Then, based on this error and the Kalman gain vector, it adjusts the parameter estimates from the previous moment and updates the covariance matrix. This algorithm incorporates a forgetting factor, which assigns greater weight to recent data, enabling the model to quickly adapt to the dynamic evolution of magnetic field patterns.
[0128] Example 8. After the previous examples describe how to obtain a virtual depth map and a compensated depth map under optical interference and magnetic field interference respectively, this example will elaborate in detail on how the system intelligently selects and fuses two sources of depth information (for example, in a complex interference scenario) to generate an optimal fused depth map that is ultimately used for positioning.
[0129] In one embodiment of the present invention, the steps of intelligent selection and fusion include:
[0130] Firstly, under the Bayesian fusion framework, the optimal depth estimation is performed on the compensated depth map and the virtual depth map based on the depth confidence map and the uncertainty distribution as weights to obtain the Bayesian fusion depth.
[0131] Specifically, for non-conflicting areas in the image, the system uses a Bayesian fusion framework to integrate two sources of depth information. The two information sources are: the compensated depth map D_comp and its confidence C_comp from the magnetic field compensation path, and the virtual depth map D_virtual and its confidence (uncertainty distribution) C_virtual from the optical reconstruction path. The system regards these two depth maps as two independent noisy observations of the true depth D. By constructing a Bayesian model and setting a reasonable likelihood function and prior (for example, the depth of the previous frame can be introduced as a time prior D_prior to enhance temporal smoothness), the mean D_bayes of the posterior probability distribution can be analytically solved, which is the optimal depth estimate at the current moment. In an optional embodiment, the posterior mean can be calculated by the following formula:
[0132] D_bayes=(D_comp×Q_comp / σ 2 _comp+D_virtual×Q_virtual / σ 2 _virtual+D_prior / σ 2 _prior) / (Q_comp / σ 2 _comp+Q_virtual / σ 2 _virtual+1 / σ 2 _prior),
[0133] Where Q is the comprehensive quality score, σ 2 is the variance.
[0134] Secondly, the depth value difference between the compensated depth map and the virtual depth map is detected. When the difference exceeds a preset threshold, it is marked as a conflict area, and one of the depth maps is selected as the conflict-processed depth map of the area based on the interference type and confidence level.
[0135] Bayesian fusion is suitable for situations where two information sources have similar estimates of the true value. However, when the difference between the two is huge, the fusion result may be unreliable. Therefore, the system calculates the difference between the two depth maps.
[0136] ΔD = |D_comp - D_virtual|. When this difference exceeds a dynamic threshold (e.g., 3 × max(σ_comp, σ_virtual)), the pixel is marked as a conflicting region. For conflicting regions, the system no longer performs fusion but instead makes a decision based on the current interference type and the confidence level of each information source.
[0137] For example, if the system identifies that the current interference is mainly optical interference and the confidence C_virtual of the virtual depth map is very high (for example, greater than 0.7), the system will choose to use the virtual depth map D_virtual as the result of the area.
[0138] If the system identifies that the current interference is mainly magnetic field interference and the confidence C_comp of the compensated depth map is very high (for example, greater than 0.8), the system will choose to use the compensated depth map D_comp as the result for this area.
[0139] The conflict resolution mechanism based on expert rules enables the system to adopt the one that is more likely to be correct when there are serious disagreements among information sources, thus avoiding incorrect fusion.
[0140] Finally, the Bayesian fused depth is integrated with the conflict-processed depth map to generate a fused depth map for subsequent localization.
[0141] This step aims to generate a complete and seamless final depth map. The system uses the conflict area marker generated in the previous step as the mask M_conflict. The final fused depth map D_final is synthesized in the following way: D_final = M_conflict × D_conflict + (1-M_conflict) × D_bayes. This means that in the non-conflict area of the image, the depth value comes from the statistically optimal Bayesian fusion result; while in the conflict area, the depth value comes from the result of the decision selection in the previous step. In this way, the present invention generates a smooth and reliable fused depth map, providing high-quality input data for the subsequent three-dimensional spatial positioning of personnel.
[0142] Example 9: Based on all the above examples, part of the process of the early warning method is described: after accurately obtaining the three-dimensional spatial position of the personnel, how to perform dynamic risk assessment and generate hierarchical early warning signals that match the risk level.
[0143] In one embodiment of the present invention, the steps of evaluating the risk level and generating a corresponding warning signal include:
[0144] First, based on the three-dimensional spatial position of the personnel and the high-risk equipment position database, the danger distance between the personnel and the high-risk equipment is calculated. Specifically, the system will read the fused three-dimensional position list of the personnel output by the aforementioned embodiment (for example, embodiment eight), which contains the precise three-dimensional coordinates of one or more personnel in the working area. At the same time, the system will access a pre-stored high-risk equipment position database, which stores the precise position, size and geometric shape of various types of high-risk equipment (such as high-voltage switchgear, rotating steam turbines, etc.) in a unified world coordinate system in the form of a three-dimensional model. This step obtains a quantified danger distance D_risk by calculating the shortest Euclidean distance from the center of mass of the personnel position to the surface of the nearest high-risk equipment.
[0145] Secondly, compare the danger distance with the pre-stored device danger radius to quantify the current situation into one of multiple risk levels. To give the danger distance a clear safety meaning, one or more danger radii R_danger are also pre-stored in the high-risk equipment location database for each device. For example, the arc danger radius of high-voltage equipment may be set to 5 meters, while the physical collision danger radius of large rotating equipment may be set to 3 meters. In this step, compare the danger distance D_risk calculated in the previous step with the danger radius R_danger of the corresponding device, thereby converting the abstract distance into a specific risk level. In a preferred embodiment, the relative risk degree S = R_danger / D_risk can be calculated, and the current situation can be divided into 1 to 5 risk levels according to the value range of S. Further, to achieve a more predictive risk assessment, the system can also combine the personnel movement speed and direction obtained in Embodiment 3 to predict the danger distance after a few seconds (e.g., after T seconds) in the future, and evaluate the risk level based on this predicted distance.
[0146] Finally, trigger a hierarchical warning response based on the divided risk levels, where a lower risk level corresponds to a prompt warning, and the highest risk level corresponds to an emergency shutdown instruction to generate a warning signal. This step is the execution end of the safety closed-loop control of the present invention. The system will automatically trigger a corresponding hierarchical warning strategy according to the risk level evaluated in the previous step. A specific example of the hierarchical warning response is as follows:
[0147] Risk level 1 (D_risk > 2R_danger): Safe state. The system only records the personnel trajectory and safety state in the background and does not generate any external warnings.
[0148] Risk level 2 (R_danger < D_risk < 2R_danger): Attention state. The system triggers a yellow warning on the visual interface of the central monitoring room to prompt the monitoring personnel to pay attention to this area.
[0149] Risk level 3 (0.5R_danger < D_risk < R_danger): Warning state. The system triggers an orange warning. In addition to prompting on the monitoring interface, it may also send a prompt warning message to the personal terminal (such as a smart bracelet) of the person concerned.
[0150] Risk level 4 (0.2R_danger < D_risk < 0.5R_danger): Dangerous state. The system triggers the highest-level red warning, immediately drives the on-site sound and light alarm to sound, and clearly indicates the specific location of the hazard source and the recommended evacuation direction in the warning message.
[0151] Risk Level 5 (D_risk < 0.2R_danger): Emergency. In addition to triggering a red alert and on-site audible and visual alarms, the system automatically generates and sends an emergency shutdown command to the control systems of related equipment (such as circuit breakers for high-risk equipment) to minimize accidents.
[0152] In this way, the present invention can not only see the danger, but also make appropriate and timely responses according to the severity of the danger, thereby forming a complete and intelligent safety protection system.
[0153] Example 10. This example aims to provide a specific and reproducible numerical calculation case to illustrate in detail how the core dual-time-scale dynamic magnetic field distortion compensation method (as described in Examples 5 and 6) of the present invention works when dealing with dynamic magnetic field interference.
[0154] Scenario setting and initial conditions: Assume that at time t = 0, a depth camera located in an area of a thermal power plant is measuring the distance of an operator, but the location is interfered by the strong magnetic field from a nearby large motor.
[0155] The base depth value of a pixel in the area (which can be interpolated from surrounding undisturbed pixels or obtained based on scene priors) is D_base=3.0 meters.
[0156] The magnetic field strength B(t) collected by the system at the current moment is 60 gauss.
[0157] The magnetic field is increasing rapidly, with a rate of change of dB / dt = +20 gauss / second.
[0158] The system calculates the second-order rate of change d of the magnetic field based on historical data 2 B / dt 2 =-50 Gauss / second 2 .
[0159] The relevant parameters of the system are set as follows: fast channel empirical constant K_instant = 0.15 cm / gauss, fusion weight time constant τ = 20 milliseconds, and magnetic field change starts at t0 = -10 milliseconds.
[0160] The calculation process consists of two steps:
[0161] The fast channel's immediate coarse compensation aims to respond quickly within 5 milliseconds. Specifically, the system uses a simplified linear model to calculate the depth distortion compensation:
[0162] ΔD_fast=B(t)×K_instant×sign(dB / dt)
[0163] Substitute the values: ΔD_fast = 60 gauss × 0.15 cm / gauss × (+1) = 9.0 cm. The depth value after fast compensation is: D_fast = D_base - ΔD_fast = 300 cm - 9.0 cm = 291.0 cm. At the same time, calculate the uncertainty of this channel: σ_fast = |dB / dt| × 0.02 = 20 × 0.02 = 0.4 cm
[0164] Pre-compensation calculation of the precision channel, the goal of this channel is to complete the accurate solution within 50 milliseconds and use prediction to offset the delay. First, apply the magnetic field predictor to predict the magnetic field strength after 50 milliseconds (τ_pred = 0.05s):
[0165] B_pred(t+τ_pred)=B(t)+dB / dt×τ_pred+0.5×d 2 B / dt 2 ×τ_pred 2 .
[0166] Substitute the values: B_pred(t+0.05s)=60+20×0.05+0.5×(-50)×(0.05) 2 =60+1-0.0625=60.9375 gauss.
[0167] The system uses the predicted future magnetic field value, B_pred, as input to an improved LM algorithm based on a magneto-optical coupling physics model for iterative solution. Assuming the algorithm converges after three iterations, a more accurate depth distortion compensation value, ΔD_precise, of 8.2 cm is obtained.
[0168] The depth value after precise compensation is: D_precise = D_base - ΔD_precise = 300 cm - 8.2 cm = 291.8 cm. Assume that based on the convergence quality of this iteration, the uncertainty of this channel is evaluated to be σ_precise = 0.15 cm.
[0169] For dual-channel adaptive fusion, at time t = 0, Δt = t - t0 = 10 milliseconds have passed since the magnetic field change began, time t0. The time-varying fusion weight is calculated as: w(t) = exp(-Δt / τ) = exp(-10 / 20) = exp(-0.5) ≈ 0.6065.
[0170] The results of the two channels are weightedly fused to obtain the final depth value: D_final = w(t) × D_fast + (1-w(t)) × D_precise.
[0171] D_final = 0.6065 × 291.0 + (1 - 0.6065) × 291.8 ≈ 176.49 + 114.65 = 291.14 cm.
[0172] Using the dual-channel compensation and fusion method of the present invention, the system ultimately outputs a compensated depth of 291.14 cm for a distance measurement point subject to dynamic strong magnetic field interference. This result effectively combines the low latency of the fast channel with the high precision of the precise channel. During the initial phase of magnetic field changes, it provides a highly reliable depth measurement that quickly and gradually approaches a more accurate value, fully demonstrating the robustness and advancement of the present invention in complex electromagnetic environments.
[0173] According to one aspect of the present application, the prediction may specifically be:
[0174] Instant automatic protection: 50ms to trigger automatic power off and equipment emergency stop; safety interlock system that requires no manual intervention; the first line of defense to prevent accidents from escalating.
[0175] Progressive warning: Risk level 1-2, 2 seconds advance warning (trajectory prediction); risk level 3-4, immediate sound and light alarm; risk level 5, automatic shutdown (50ms is sufficient).
[0176] Magnetic field prediction can be divided into multiple time scales: short-term prediction (50ms-200ms) for automatic protection, medium-term prediction (0.5s-2s) for manual warning, and long-term prediction (2s-10s) for operation planning.
[0177] According to one aspect of the present application, the following are examples of key usage scenarios:
[0178] The moment of electrical equipment failure: From the occurrence of the fault to the formation of the arc: 10-100ms; 50ms prediction can trigger protection before the arc occurs, avoiding more serious explosion accidents.
[0179] Accidental entry into high-voltage areas: Based on trajectory prediction, workers are given a 5-10 second advance warning, allowing them to change their direction.
[0180] In another embodiment of the present application, a state transition process is also included, which includes four states: Normal, Magnetic, Optical, and Dual. The state transition is triggered by a clear threshold condition:
[0181] From the "normal" state, when the magnetic field strength is greater than 50 gauss and the rate of change dB / dt is greater than 10 gauss / second, it switches to the "magnetic interference" state.
[0182] From the "Normal" state, when an overexposed area with a light intensity greater than 10,000 lux is detected, it switches to the "Light Interference" state.
[0183] From the "magnetic interference" or "optical interference" state, when the trigger condition of another interference source is also met, it switches to the "double interference" state.
[0184] When the intensity or rate of change of the corresponding interference source falls back below the threshold and lasts for a period of time, the state machine will return to a state with lower interference level or the "normal" state.
[0185] By constructing and solving an inverse optical diffusion model with strong constraints on human shape and motion trajectory, the system is able to effectively locate people even when extreme optical overload causes large-area sensor saturation. This solves the problem of traditional vision technology attempting to recover lost information and instead uses the overexposed RGB image stream (input parameter) itself as a usable feature signal. Specifically, the system combines the observed overexposed area boundaries (technical features) with the PSF kernel function (algorithmic features) that describes the physical diffusion process of light, and solves an ill-posed inverse problem using the ADMM algorithm to reversely infer the outline of people obscured by strong light. To ensure the uniqueness and rationality of the solution, the solution process is subject to two strong constraints: one is a priori model of human shape based on ergonomic data (static constraint), and the other is a motion continuity constraint (dynamic constraint) generated by predicting historical trajectories based on the extended Kalman filter. In thermal power plant scenarios, when dangerous events such as switchgear arcs or welding operations generate instantaneous strong light, both traditional depth cameras and standard cameras will be blinded. However, the present invention utilizes this blinded, overexposed area to reconstruct a virtual depth map of the personnel. Combined with the calculation of the safety envelope (Claim 6), this provides a conservative but reliable position estimate. This overcomes the fatal flaw of existing technologies, which create monitoring blind spots at the most dangerous moments, and significantly improves the safety of users (operators) in extremely dangerous situations.
[0186] The present invention functionally achieves optimal integration of depth data from two different technical paths, magnetic field compensation and optical reconstruction, by designing an intelligent fusion algorithm that combines a Bayesian framework with conflict decision rules. This embodiment can intelligently cope with complex composite interference environments and generate a seamless, coherent, and highly confident final depth map. Specifically, the algorithm uses the compensated depth map and the virtual depth map and their respective confidence maps (input parameters) as evidence. In non-conflicting areas, it uses a Bayesian fusion framework (algorithm feature) for weighted averaging to obtain a statistically optimal depth estimate, improving the internal performance of the system and making the depth map transition smoothly between different areas. More importantly, when a conflicting area with a significant difference between the two depth data sources is detected, the system activates the decision rule (business rules and method features) and abandons fusion based on the main interference type identifier (input parameter) of the current environment and the confidence level of each data source. Instead, it makes a two-choice decision. For example, in an overexposed area caused by arc light, the system will determine that the virtual depth map has a higher weight. This fusion strategy, combining statistical optimization with expert rules, avoids erroneous averaging of conflicting data and ensures the logical consistency and reliability of the output. In a thermal power plant scenario, faced with a complex working surface characterized by both background magnetic fields and localized strong light, this method can process depth information in a zoned and strategic manner, ultimately outputting a high-quality fused depth map, providing a solid data foundation for subsequent hazard assessments.
[0187] The preferred embodiments of the present invention are described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the scope of protection of the present invention.
Claims
1. An intelligent early warning method for approaching high-risk equipment in a thermal power plant, characterized by: include: Collect multi-source sensor data covering the operating area, identify and obtain the type of environmental interference, and extract RGB image streams; When the environmental interference type is identified as optical interference, virtual depth reconstruction is performed based on the RGB image stream and historical trajectory data through reverse optical modeling and motion continuity constraints to generate a virtual depth map; Fuse the virtual depth map and RGB image stream to determine the 3D spatial position of the person; Combined with the pre-stored high-risk equipment location database, the dangerous distance between personnel and high-risk equipment is calculated, the risk level is assessed and an early warning signal is generated.
2. The method according to claim 1, characterized in that The steps for virtual depth reconstruction include: Analyze the overexposed areas in the RGB image stream and reversely infer the person's human body contour and the corresponding contour confidence distribution based on its optical characteristics; Based on historical trajectory data, future motion is predicted to obtain a predicted trajectory sequence and the corresponding motion uncertainty; The inferred human contour and predicted trajectory sequence are combined to generate a virtual depth map; and the contour confidence distribution and motion uncertainty are fused to generate an uncertainty distribution.
3. The method according to claim 2, characterized in that Based on the optical features, the inferred human body contour and the corresponding contour confidence distribution are inferred, including: The inverse problem optimization model is called to express the observed overexposed and saturated image as the convolution result of the true human body contour to be determined and the optical diffusion kernel function; Under the constraints of the preset human body shape prior model, the inverse problem optimization model is solved by minimizing the boundary error between the saturated image predicted by the model and the actual observed saturated image, and the inferred human body contour is output; A contour confidence distribution is generated based on the boundary consistency between the inferred human contour and the optical features of the overexposed area, as well as the shape rationality with the prior model of the human shape.
4. The method according to claim 3, characterized in that Solve inverse optimization models, including: The alternating direction multiplier method is used to iteratively solve the inverse problem optimization model to generate the intermediate human body contour; During the iterative solution process, the parts of the intermediate human body contour that do not conform to the preset body shape characteristics are periodically corrected based on the human body shape prior model; After the iteration converges, the corrected contour is output as the inferred human body contour.
5. The method according to claim 2, characterized in that Based on historical trajectory data, future motion is predicted to obtain a predicted trajectory sequence and the corresponding motion uncertainty, including: Analyze the speed and acceleration changes in historical trajectory data to identify the person's current motion state label, including stationary, straight, or turning; According to the current motion state label, a state-adaptive human kinematic model is constructed, and different motion parameters are set for different motion states; Kalman filtering is used in combination with the human kinematics model to process historical trajectory data and generate predicted trajectory sequences and predicted covariance matrix sequences; Based on the prediction covariance matrix sequence, the motion uncertainty is modeled.
6. The method according to claim 2, characterized in that The inferred human silhouette and predicted trajectory sequence are combined to generate a virtual depth map, and the silhouette confidence distribution is fused with the motion uncertainty to generate an uncertainty distribution, including: Combined with the 3D scene model, the initial virtual depth map is generated by raycasting the inferred human body contour and constraining it using the predicted trajectory sequence; The profile confidence distribution and motion uncertainty are integrated to calculate the comprehensive uncertainty, which is the uncertainty distribution. Based on the initial virtual depth map and the comprehensive uncertainty, the preset safety factor is applied to calculate the safety envelope depth, which is used as the final virtual depth map.
7. The method according to claim 1, characterized in that The method further includes: When the environmental interference type is identified as magnetic field interference, the collected original depth image is compensated for dual-time-scale dynamic magnetic field distortion based on the magnetic field intensity data to generate a compensated depth map and a depth confidence map; Intelligently select and fuse the compensated depth map and virtual depth map, as well as their corresponding depth confidence map and uncertainty distribution, to determine the 3D spatial position of the person.
8. The method according to claim 7, characterized in that Perform dual-time-scale dynamic magnetic field distortion compensation, including: In the fast channel, based on the magnetic field intensity data, a linear model is used to perform immediate coarse compensation on the original depth image to generate a fast compensated depth map and a fast channel confidence map; In the precise channel, based on the magnetic field intensity data, the original depth image is accurately compensated by iteratively solving the magnetic-optical coupling physical model to generate a precise compensated depth map and a precise channel confidence map; A time-varying weight function is used to combine the fast and precise channel confidence maps to perform weighted fusion on the fast and precise compensated depth maps to generate the compensated depth map and depth confidence map.
9. The method according to claim 8, characterized in that In the precise channel, the magneto-optical coupling physical model is iteratively solved to generate an accurate compensated depth map and an accurate channel confidence map, including: Based on the temporal changes of magnetic field intensity data, a magnetic field predictor is constructed and applied to generate a sequence of magnetic field predictions at future moments; The improved Levenberg-Marquardt algorithm is used, and the magnetic field prediction sequence is used as input to iteratively solve the magnetic-optical coupling physical model, calculate the distortion compensation value in advance, and obtain an accurate compensation depth map; Based on the convergence quality of the iterative solution, an accurate channel confidence map is generated.
10. The method according to claim 8, characterized in that The magneto-optical coupling physical model used in the precise channel includes: The spatial offset submodule is used to model the spatial offset effect of the magnetic field on the charge transfer path inside the depth sensor through the Lorentz force; The circuit delay submodule is used to model the circuit delay effect introduced by the magnetic field on the phase measurement during time-of-flight ranging; The unified parameterization submodule is used to integrate the spatial offset effect with the circuit delay effect to establish a unified nonlinear relationship between the depth distortion and the magnetic field strength and direction.
11. The method according to claim 7, characterized in that The steps of intelligent selection and fusion include: In the Bayesian fusion framework, the optimal depth estimation is performed on the compensated depth map and the virtual depth map based on the depth confidence map and the uncertainty distribution as weights to obtain the Bayesian fusion depth. Detect the depth value difference between the compensated depth map and the virtual depth map. When the difference exceeds a preset threshold, mark it as a conflict area and select one of the depth maps as the conflict-resolved depth map of the area based on the interference type and confidence level. The Bayesian fusion depth is integrated with the depth map after conflict resolution to generate a fused depth map for subsequent positioning.
12. The method according to claim 1, characterized in that The steps to assess the risk level and generate corresponding early warning signals include: Calculate the danger distance between personnel and high-risk equipment based on the personnel's three-dimensional spatial position and the high-risk equipment location database; Compare the danger distance with the pre-stored equipment danger radius and quantify the current situation into one of multiple risk levels; Based on the risk levels, a hierarchical early warning response is triggered.