A computer vision-based security engineering construction safety monitoring method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-08-11
AI Technical Summary
现有图像增强方法通常将整帧图像视为均匀退化场进行处理,未考虑由局部扬尘散射、逆光过曝、暗区噪声造成的空间非均匀视觉不适,导致复原结果在提升亮度的同时丢失关键纹理细节,影响后续检测精度
[0047] 1. This invention employs a reaction-diffusion equation guided by visual imperfection, using KL divergence to measure local visual imperfection, thus achieving a dynamic balance between edge-preserving diffusion and imperfection-driven reaction. Compared to existing global enhancement methods, this equation automatically strengthens the reaction term in high-imperfection regions such as glare and dark areas to regress the observed values, denoises only by diffusion in flat areas, and automatically stops diffusion at edges to protect structural texture.
Smart Images

Figure CN122551280A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and construction safety monitoring technology, and in particular relates to a computer vision-based method for monitoring the construction safety of security engineering projects. Background Technology
[0002] Security engineering construction sites are characterized by numerous high-altitude operations, overlapping heavy machinery, dense temporary structures, and complex ambient lighting, making them high-risk areas for safety accidents such as falls from heights, falling objects, and mechanical injuries. Traditional safety management mainly relies on manual inspections and fixed-rule video surveillance, which has fundamental limitations such as delayed response, strong subjective dependence, and inability to provide full coverage.
[0003] Construction sites are often affected by multiple degradation factors, such as glare, dust, dense shadows, and low illumination at night. Existing image enhancement methods typically treat the entire image as a uniform degradation field, without considering the spatial non-uniformity and visual discomfort caused by local dust scattering, backlight overexposure, and noise in dark areas. This results in the restoration result losing key texture details while improving brightness, affecting the accuracy of subsequent detection. Summary of the Invention
[0004] This invention addresses the shortcomings of existing security engineering construction safety monitoring methods in six aspects: image restoration, small target detection, behavior recognition, hazard assessment, 3D positioning, and multi-camera scheduling. It provides a computer vision-based security engineering construction safety monitoring method.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0006] S1. Acquire video streams of the construction scene captured by multiple cameras, and perform adaptive restoration of each frame image using the reaction-diffusion equation guided by visual inappropriateness to obtain an enhanced image;
[0007] S2. Input the enhanced image into a convolutional neural network with a field-aligned dynamic receptive field to detect small targets such as people and safety equipment in the image;
[0008] S3. Extract the skeletal key point sequence from the detected personnel, calculate the Riemann torsional energy spectrum of the local transformation between adjacent frames on the joint rotation group manifold, and combine the peak value and standard deviation of the Riemann torsional energy spectrum to form a violation score. Compare the violation score with a preset first threshold, a second threshold, and a third threshold, where the first threshold < the second threshold < the third threshold. When the violation score is less than the first threshold, it is judged as normal; when it is between the first and second thresholds, it is identified as a minor violation and triggered. A level-two response is triggered when the level falls between the second and third thresholds, identifying it as a moderate violation. A level-3 response is triggered when the level is greater than or equal to the third threshold, which is identified as a serious violation, i.e., the highest level of violation. Level 1 response;
[0009] S4. Using mobile mechanical equipment and static hazardous boundaries as source terms, the risk potential diffusion equation is solved in real time through a physical information neural network to obtain the dynamic risk potential field at the current and future moments.
[0010] S5. Using the neural radiation field of structural constraints, the foot points of the same person detected from multiple perspectives are reconstructed in three dimensions to obtain their precise spatial coordinates in the world coordinate system.
[0011] S6. Based on the dynamic risk potential field, calculate the material derivative of the dynamic risk potential field along the direction of movement for the positions of all detected personnel. If the material derivative exceeds the safety limit, trigger a dangerous intrusion warning.
[0012] S7. Input the dynamic risk potential field distribution and the precise spatial coordinates of the personnel into the distributed mirror descent coordinator, drive multiple PTZ cameras to conduct collaborative patrols with the goal of minimizing the uncovered risk potential and field of view overlap. When the highest level of violation exists, forcibly assign the PTZ camera closest to the personnel to perform evidence-gathering zoom, and reconstruct the coverage of the remaining cameras.
[0013] Preferably, the formula for adaptive restoration in S1 using the response-diffusion equation guided by visual inappropriateness is as follows:
[0014] ,
[0015] in, The instantaneous rate of change in image brightness values during their evolution. The standard deviation is Gaussian kernel, In the pseudo-moment Pixel position The brightness value of the evolved image at that location, Pixel position The initial observed image brightness value, These are the global weighting coefficients. For natural scene statistical models, The brightness of a local image patch centered on the current pixel. This is a sharpness parameter that controls the sensitivity of visual asymmetry to divergence. The KL divergence between the local distribution and the natural statistical model. In pixels The gradient statistical distribution of a local image patch centered on the image. For natural scene statistical models, For divergence operators, In the pseudo-moment Pixel position Spatial gradient of the evolution image at that location.
[0016] Preferably, in the convolutional neural network with field-aligned dynamic receptive fields in S2, the sampling position of the convolutional kernel is anisotropically deformed through the local structure tensor field, and its offset formula is:
[0017] ,
[0018] , ,
[0019] in, For the deformed first The coordinate offset of each sampling point relative to the kernel center It is an orthogonal matrix composed of two eigenvectors. The anisotropic strength coefficient, The local structural coherence coefficient, For the standard convolution kernel, the first The original coordinate vector of each sampling point relative to the kernel center For structure tensor The first major feature, For structure tensor The second major feature, To prevent division by zero constant, for The corresponding unit eigenvector, for The corresponding unit eigenvector, The coordinates of the current convolution kernel center in the image are the two-dimensional pixel coordinates.
[0020] Preferably, the formula for calculating the Riemann torsional energy spectrum of the local transformation between adjacent frames in S3 is as follows:
[0021] , ,
[0022] , ,
[0023] in, The Riemann torsional energy spectrum is the result of local transformation between adjacent frames. It is a six-dimensional torsion vector. For measuring tensors Induced Riemann inner product, A six-dimensional torsion vector transpose, For the Riemannian metric tensor, For the first The absolute pose matrix of the joint at frame time. From The 3×3 rotation matrix extracted from it. Rotational sub-blocks of Riemannian metric tensors For the translation sub-block of the Riemannian metric tensor, The transformation matrix is... The matrix representation in Lie algebra encodes the instantaneous motion information of the joint. For the first Frame to the The local rigid body transformation matrix of a specific joint between frames. For standard linear isomorphism, For instantaneous angular velocity components, For instantaneous linear velocity components, Using the time index of the video frame, the formula for calculating the violation score is:
[0024] ,
[0025] ,
[0026] in, To score points for violations, The Riemann torsional energy spectrum is the result of local transformation between adjacent frames. The weights are positive real number combinations. The total number of frames in the action sequence. Let be the arithmetic mean energy of the sequence.
[0027] Preferably, the risk level classification rule in S3 is as follows:
[0028] ,
[0029] in, To score points for violations, The first threshold, The second threshold, The third threshold, , , Based on risk level, the tiered response measures are as follows: The system only sends discreet vibrations or alerts to the smart safety devices worn by the violators, without disturbing other workers. The system plays a pre-set voice warning to the area via directional loudspeakers at the construction site, and simultaneously pushes screenshots of violations and their locations to the handheld terminal of the on-duty safety officer. The system immediately sends an emergency deceleration or shutdown command to the construction equipment controller associated with the area; injects the location coordinates of the violator as an additional potential energy term into the source term of the risk potential diffusion equation described in step S4; forcibly schedules the PTZ camera closest to the violator to turn to that location and perform maximum optical zoom recording; the remaining PTZ cameras are automatically redistributed to fill scheduling gaps through the distributed mirror descent coordinator described in step S7.
[0030] Preferably, the formula for calculating the risk potential diffusion equation in S4 is as follows:
[0031] ,
[0032] in, For dynamic risk potential field, Let be the local partial derivative of the dynamic risk potential field with respect to time. The spatial diffusion coefficient is... Let be the spatial gradient vector of the dynamic risk potential field. For indexing mobile hazard sources, For the first A mobile machine at any time instantaneous dangerous rate, It is a spatial position vector. For the first A mobile machine at any time Spatial position vector, Let be the Dirac distribution function. This is the attenuation coefficient.
[0033] Preferably, the formula for calculating the precise spatial coordinates in the world coordinate system in S5 is as follows:
[0034] ,
[0035] ,
[0036] ,
[0037] in, Provides the precise spatial coordinates of personnel in the world coordinate system at the current moment. For the initial fusion coordinates, Let the unit normal vector be the plane in which the person is standing. For the constant term of the plane equation, Let Euclidean norm be the normal vector. For camera number indexing, For the first Confidence weights for each perspective For the first The 3D estimated coordinates of the foot points independently recovered by each camera For the first The three-dimensional coordinates of the optical center of a camera in the world coordinate system. The number of spatial points uniformly sampled on the light source. For sampling point index, The distance between two sampling points. To constrain the neural radiation field at the sampling point The volume density value, It is a spatial position vector. Let be the unit vector pointing the light ray to the scene. For the first The Euclidean distance from each sampling point to the optical center of the camera.
[0038] Preferably, the formula for calculating the material derivative of the dynamic risk potential field in S6 is:
[0039] ,
[0040] ,
[0041] in, The material derivative of the dynamic risk potential field. For personnel at all times The three-dimensional position coordinates in the world coordinate system. Let be the local partial derivative of the dynamic risk potential field with respect to time. Let be the spatial gradient vector of the dynamic risk potential field. Let be the instantaneous velocity vector of the person. The time interval between two adjacent frames. For the same person at any time The three-dimensional position coordinates.
[0042] Preferably, in the distributed mirror descent coordinator of S7, the attitude update rule for each PTZ camera is as follows:
[0043] ,
[0044] ,
[0045] in, For the first The PTZ camera in the first Update the pose vector after the next iteration. For the PTZ camera's index number, For distributed iteration steps, For camera Feasible attitude space For the loss function in gradient at, For the first The PTZ camera in the first Update the pose vector after the next iteration. For the mirror descent step size, Let Bregman divergence be the property of the divergence. Let be the Bregman potential function. For candidate pose vectors, Candidate pose vector The potential function value at that point, For attitude vector The potential function value at that point, For attitude vector The gradient vector at that point.
[0046] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0047] 1. This invention employs a reaction-diffusion equation guided by visual imperfection, using KL divergence to measure local visual imperfection, thus achieving a dynamic balance between edge-preserving diffusion and imperfection-driven reaction. Compared to existing global enhancement methods, this equation automatically strengthens the reaction term in high-imperfection regions such as glare and dark areas to regress the observed values, denoises only by diffusion in flat areas, and automatically stops diffusion at edges to protect structural texture.
[0048] 2. The field-aligned dynamic receptive field module proposed in this invention uses the eigenvalue decomposition of the local structural tensor to obtain the dominant direction and coherence, enabling the convolution kernel to adaptively stretch along the strong structural direction of steel pipes, rods, etc., and contract in the orthogonal direction.
[0049] 3. This invention calculates the Riemann torsional energy spectrum on a joint rotation group manifold, maps the rigid body transformation between adjacent frames to Lie algebras, calculates the torsional energy density under the metric tensor, and constructs a violation score through a weighted combination of peak value and standard deviation. This avoids angle period jumps and distance distortion, effectively identifying violations such as climbing and vaulting. Simultaneously, by using three threshold levels to classify risk levels L1 to L3, a gradient-based approach is implemented, from private alerts and area warnings to forced equipment speed reduction and shutdown, potential field injection, and mandatory evidence collection. This avoids excessive alarms while ensuring timely intervention for serious threats.
[0050] 4. This invention establishes a partial differential equation for risk potential diffusion, using the instantaneous dangerous power of mobile machinery as a point source, and employing a space-dependent diffusion coefficient and attenuation terms to describe the risk propagation and dissipation process. A physical information neural network is used for real-time solution. Combined with the material derivative warning criterion along the direction of personnel movement, even if the current distance is still far, if personnel rush towards a high-potential area at high speed or rapidly approach a high-risk source, the system can still trigger an early warning. Attached Figure Description
[0051] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 This is a schematic diagram of a computer vision-based method for monitoring the safety of security engineering construction. Detailed Implementation
[0053] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described below in conjunction with the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0054] Numerous specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways than those described herein, and therefore the invention is not limited to the specific embodiments disclosed in the following specification.
[0055] In this embodiment, video streams of a construction scene captured by multiple cameras are acquired. For each frame, an adaptive restoration is performed using a reaction-diffusion equation guided by visual asymmetry to obtain an enhanced image. The enhanced image is then input into a convolutional neural network with a field-aligned dynamic receptive field to detect small targets such as personnel and safety equipment. The reaction-diffusion equation guided by visual asymmetry uses an edge-preserving diffusion term to protect key textures such as steel pipe edges, and a reaction term driven by KL divergence to automatically strengthen and correct high-asymmetry areas such as glare and dark areas to prevent over-smoothing. This solves the texture loss problem caused by non-uniform composite degradation such as glare, dust, and dark area noise in construction scenes, improving the signal-to-noise ratio of the restored image. The improvement in image quality and texture retention is enhanced. Furthermore, by utilizing a local structure tensor field to drive the convolution kernel to adaptively stretch along the dominant directions of steel pipes and rods and contract vertically, a streamlined receptive field aligned with the linear structure of the scene is formed. This solves the problem of small target edges being diluted by the background and high false negative rates caused by the mismatch between the fixed square receptive field and the linear structure of the construction scene, thus improving the average accuracy of detecting safety helmets and safety belts smaller than 30×30 pixels. The synergy of these two methods achieves joint optimization from pixel-level image quality to target-level feature extraction, addressing the challenges of construction scenes and providing a stable and reliable front-end perception foundation for subsequent behavior recognition and risk warning. The formula for adaptive restoration using the reaction-diffusion equation guided by visual inappropriateness in S1 is as follows:
[0056] ,
[0057] in, The instantaneous rate of change in image brightness values during their evolution. The standard deviation is Gaussian kernel, In the pseudo-moment Pixel position The brightness value of the evolved image at that location, Pixel position The initial observed image brightness value, These are the global weighting coefficients. For natural scene statistical models, The brightness of a local image patch centered on the current pixel. This is a sharpness parameter that controls the sensitivity of visual asymmetry to divergence. The KL divergence between the local distribution and the natural statistical model. In pixels The gradient statistical distribution of a local image patch centered on the image. For natural scene statistical models, For divergence operators, In the pseudo-moment Pixel position The spatial gradient of the evolving image is used as the equation. The equations employ an evolutionary partial differential equation framework because degradations in construction images, such as glare, dark area noise, and dust scattering, are spatially non-uniform and cannot be simultaneously corrected by a single global transformation. Therefore, each pixel must evolve independently along pseudo-time to adapt to the restoration requirements of spatial variations. The framework consists of two main collaborative mechanisms: In the edge-preserving diffusion term, the image is first smoothed with a Gaussian kernel before calculating the gradient to suppress the misclassification of strong noise in dark areas as edges, leading to incorrect diffusion shutdown; the squared gradient magnitude is used as a measure of edge strength, which is better than the first-order magnitude at distinguishing strong structural edges from weak texture fluctuations; the edge stopping function is used as the diffusion coefficient, automatically approaching zero at strong edges such as steel pipes and safety helmet outlines to protect key textures, and approaching one in flat areas for sufficient noise reduction; a divergence operator is applied to the current evolving image gradient modulated by the edge stopping function, forming an anisotropic diffusion flow that smooths along the edge direction and stops across the edge direction, denoising without blurring the structure. In the reaction term, the residual term anchors the restored result to the constraints of the original observation information, preventing the image from being over-smoothed in strongly degraded areas and becoming a false texture; the local image patch statistical distribution quantifies whether the current pixel neighborhood is a natural texture or an abnormal state such as glare or noise; the natural scene statistical model, as a visually normal, non-reference benchmark, is obtained by offline training with a large number of high-quality images, making the judgment scene-adaptive and independent of manual thresholds; KL divergence is used to measure the degree of deviation of the local distribution from the natural model, because KL divergence is highly sensitive to outliers in long-tailed distributions and can keenly capture glare and white cast. Extreme degradation signals such as dark area clipping are detected. The divergence is mapped to a visual discomfort coefficient driving the intensity of the response term using an exponential saturation function. This S-shaped function responds rapidly to moderate to severe degradation and tends to saturate at extreme values to avoid oscillations. At high discomfort levels, the strong response term pulls the image back to observation to preserve real details, while at low discomfort levels, the response term automatically shuts off, allowing the diffusion term to dominate denoising. A sharpness parameter controls the sensitivity of discomfort to divergence, facilitating flexible adjustment under the degradation statistical characteristics of different construction site scenes. Global weights balance the overall contribution ratio of diffusion and response, ensuring the stability of iterative convergence. All elements work together to form a physically driven restoration that preserves edges in structural regions, denoises flat regions, and adaptively corrects visual discomfort areas, providing clear, high-signal-to-noise ratio image input for subsequent field alignment detection. In the convolutional neural network with field alignment dynamic receptive field, the sampling position of the convolutional kernel undergoes anisotropic deformation through a local structural tensor field, with the offset formula being:
[0058] ,
[0059] , ,
[0060] in, For the deformed first The coordinate offset of each sampling point relative to the kernel center It is an orthogonal matrix composed of two eigenvectors. The anisotropic strength coefficient, The local structural coherence coefficient, For the standard convolution kernel, the first The original coordinate vector of each sampling point relative to the kernel center For structure tensor The first major feature, For structure tensor The second major feature, To prevent division by zero constant, for The corresponding unit eigenvector, for The corresponding unit eigenvector, The coordinates of the current convolution kernel center in the image are given as two-dimensional pixels. In the field-aligned dynamic receptive field offset formula, the local structure tensor is first obtained by calculating the outer product of the image gradients and then smoothing it with Gaussian weights. This robustly statistically analyzes the dominant distribution of gradient directions within the neighborhood of each pixel in noisy construction images, overcoming the weakness of single-pixel gradient directions being easily affected by noise. The sampling position of the convolution kernel is adaptively deformed. First, a small neighborhood is taken around each pixel in the image, and the direction and intensity of the steepest brightness change at each point in this neighborhood are calculated. The direction vector of each point is then outer-productted with itself to obtain a matrix that reflects the gradient direction distribution at that point. Then, a Gaussian weighted window is used to weight and average these matrices for all points in the neighborhood to obtain the local structure tensor. This tensor encodes the dominant gradient direction and the strength of direction consistency in the region surrounding the pixel. Gaussian smoothing suppresses the interference of image noise on direction estimation, making the structure analysis more stable and reliable. Eigenvalue decomposition of the local structure tensor yields two non-negative eigenvalues and two mutually perpendicular unit eigenvectors. Larger eigenvalues indicate stronger gradient energy in that direction, and their corresponding eigenvectors point to the direction of the most dramatic change in local brightness, i.e., the dominant direction. Smaller eigenvalues indicate gradient energy in the vertical direction, and their corresponding eigenvectors are perpendicular to the dominant direction. Local structural coherence is defined by dividing the difference between these two eigenvalues by their sum. It normalizes the strength of the directionality in the region; coherence approaches one at sharp, straight edges and approaches zero in flat, unstructured regions. The minimal constant added to the denominator is solely to prevent calculation errors when both eigenvalues are zero in flat regions. An anisotropy intensity coefficient is introduced to allow the system to flexibly adjust the degree of deformation under construction scenarios with different structural densities; a larger value indicates more severe kernel deformation. A rotation matrix is formed by placing the two eigenvectors side-by-side, and its transpose rotates the sampling points of the standard convolution kernel from the original image coordinate system to a local natural coordinate system with the dominant and orthogonal directions as axes. In this coordinate system, the sampling stride along the dominant direction is magnified to "one times the inherent deformation tolerance multiplied by the coherence" of the original stride, stretching the convolution kernel along the edge extension direction of the steel pipe and rod, expanding the receptive field to capture the surrounding context of small safety equipment spanning the structure. The sampling stride along the orthogonal direction is proportionally reduced to the reciprocal of the stretching factor, narrowing the kernel in the vertical edge direction, preventing background texture from adjacent parallel structures from being mixed into the convolution, thereby enhancing the focusing ability on the edges of small targets. The stretching and shrinking are reciprocals of each other, ensuring that the total sampling area of the convolution kernel remains approximately unchanged, and the amount of information is basically conserved. After deformation, a rotation matrix is used to transform the sampling points from the local natural coordinate system back to the original image coordinate system to obtain the final offset position.
[0061] For detected individuals, skeletal keypoint sequences are extracted. The Riemann torsional energy spectrum of the local transformation between adjacent frames is calculated on the joint rotation group manifold. The peak value and standard deviation of the Riemann torsional energy spectrum are combined to form a violation score. This violation score is compared with preset first, second, and third thresholds, where the first threshold < the second threshold < the third threshold. When the violation score is less than the first threshold, it is considered normal; when it falls between the first and second thresholds, it is identified as a minor violation and triggered. A level-two response is triggered when the level falls between the second and third thresholds, identifying it as a moderate violation. A level-3 response is triggered when the level is greater than or equal to the third threshold, which is identified as a serious violation, i.e., the highest level of violation. The design adopts a progressive architecture of manifold torsional energy spectrum, violation score, and three-layer threshold graded response because the core feature of violations in security construction is not simply joint position displacement, but abnormal torsion of limbs in three-dimensional space. Climbing, vaulting and other actions will generate high-frequency, high-intensity torsional energy pulses at joints such as knees and elbows, while the torsional energy distribution of compliant actions such as normal walking and bending over is relatively stable. Traditional methods use joint coordinates to perform graph convolution in Euclidean space, which will cause distance distortion due to the periodicity of angles, making it difficult to distinguish actions with similar amplitudes but completely different properties. This invention maps the rigid body transformation of joints between adjacent frames to a manifold of a special Euclidean group SE(3), extracts a six-dimensional torsional vector in Lie algebra, and assigns anisotropic weights to each component using a Riemannian metric tensor that depends on the current posture, and calculates the torsional energy density to form a time energy spectrum. This manifold metric method naturally encodes the nonlinear mechanical resistance of large-angle torsion, avoiding topological mismatch in Euclidean space. Based on this, a violation score is constructed by weighting the peak value and standard deviation of the energy spectrum. The peak value reflects the instantaneous extreme torsional intensity during the action, while the standard deviation reflects the fluctuation and instability of torsional energy over time. The combination of the two can simultaneously capture severe instantaneous violations and continuous unstable anomalies, eliminating the blind spots of single peak or mean indicators. The violation score is compared with a three-tiered incremental threshold to classify the risk level because different violations have different degrees of urgency: minor violations such as briefly letting one hand leave the grip only require a private reminder to avoid disrupting the construction rhythm; moderate violations such as leaning out of the edge require area-specific warnings to attract the attention of the person and those around them; while serious violations such as completely climbing over the guardrail or climbing on the outside of a height pose an immediate fall threat and must trigger active interventions such as equipment slowing down and stopping, injecting high-potential energy terms into the risk potential field to warn the entire area, and mandatory PTZ recording. This classification mechanism avoids over-alarming leading to personnel ignoring the issue and ensures that the highest level of threat receives a millisecond-level multi-dimensional response, integrating behavior recognition, risk assessment, and physical intervention into a closed loop. The formula for calculating the Riemann torsional energy spectrum of the local transformation between adjacent frames is as follows:
[0062] , ,
[0063] , ,
[0064] in, The Riemann torsional energy spectrum is the result of local transformation between adjacent frames. It is a six-dimensional torsion vector. For measuring tensors Induced Riemann inner product, A six-dimensional torsion vector transpose, For the Riemannian metric tensor, For the first The absolute pose matrix of the joint at frame time. From The 3×3 rotation matrix extracted from it. Rotational sub-blocks of Riemannian metric tensors For the translation sub-block of the Riemannian metric tensor, The transformation matrix is... The matrix representation in Lie algebra encodes the instantaneous motion information of the joint. For the first Frame to the The local rigid body transformation matrix of a specific joint between frames. For standard linear isomorphism, For instantaneous angular velocity components, For instantaneous linear velocity components, Using the time index of the video frame, the formula for calculating the violation score is:
[0065] ,
[0066] ,
[0067] in, To score points for violations, The Riemann torsional energy spectrum is the result of local transformation between adjacent frames. The weights are positive real number combinations. The total number of frames in the action sequence. The arithmetic mean energy value of the sequence. The risk level classification rule is as follows:
[0068] ,
[0069] in, To score points for violations, The first threshold, The second threshold, The third threshold, , , Based on risk level, the tiered response measures are as follows: The system only sends discreet vibrations or alerts to the smart safety devices worn by the violators, without disturbing other workers. The system plays a pre-set voice warning to the area via directional loudspeakers at the construction site, and simultaneously pushes screenshots of violations and their locations to the handheld terminal of the on-duty safety officer. The system immediately sends an emergency deceleration or shutdown command to the construction equipment controller associated with the area; injects the coordinates of the violator's location as an additional potential energy term into the source term of the risk potential diffusion equation described in step S4; forcibly schedules the PTZ camera closest to the violator to turn to that location and perform maximum optical zoom recording; the remaining PTZ cameras automatically redistribute their coverage areas through the distributed mirror descent coordinator described in step S7 to fill scheduling gaps. The selection of each computational element in this design scheme is based on the core requirement of "how to accurately measure the severity and instability of joint torsion on a curved manifold and map it to differentiated treatment matching the urgency of the risk." First, the joint motion between adjacent frames is represented as a local rigid body transformation matrix on a special Euclidean group. The motion of human joints in three-dimensional space is essentially a rigid body motion coupled with rotation and translation. Relying solely on Euclidean coordinate differences will lose the structural information between the rotational and translational components, failing to fully describe a torsional motion. By introducing logarithmic mappings and linear isomorphisms on the group, elements in the curved manifold are projected onto Lie algebras and transformed into a six-dimensional torsional vector. This preserves the local structure of group multiplication in a flat vector space while obtaining an algebraic object capable of linear combination and quadratic form operations. This avoids distance jumps caused by angular periodicity and fully encodes instantaneous angular and linear velocities. A Riemannian metric tensor dependent on the current joint posture is used as the weighting matrix because the torsional range and mechanical resistance of the joint change nonlinearly under different postures. Constructing it as a block diagonal structure allows for anisotropic weighting of rotational and linear displacement components, resulting in higher energy values at large-angle limit torsions. This physically more closely approximates the actual effort exerted by the joint, enhancing the ability to identify violations. The six-dimensional torsional information is compressed into a scalar energy spectrum value through quadratic forms, measuring the intensity of instantaneous torsion and facilitating the formation of time series for statistical feature extraction. The peak value of the energy spectrum sequence is selected to capture the instantaneous extreme torsion during the movement, while the standard deviation is used to measure the intensity of the fluctuation in torsional energy throughout the entire movement segment. Normal periodic movements exhibit stable energy spectrum fluctuations, while violations, due to their non-periodic nature and multi-joint switching, show abrupt high and low fluctuations, resulting in a significantly increased standard deviation. The two are combined using adjustable weights to construct a violation score, complementing the blind spots of relying solely on peak values (which may miss persistent abnormalities) and relying solely on standard deviations (which may miss single extreme movements). Furthermore, the score can be flexibly adjusted by the subject operating characteristic curve according to safety regulations, balancing recall and false alarm rates. Finally, the violation score is compared with a three-tiered, incremental threshold to classify risk levels and trigger tiered responses, as the required level of social intervention and automatic handling varies drastically depending on the severity of the violation: the first threshold excludes inherent fluctuations in normal construction; the second threshold separates minor violations requiring only private reminders from moderate violations requiring regional warnings; and the third threshold ensures zero false alarms for serious violations posing an immediate threat to life. Send private vibrations only to individuals to avoid excessive disruption to the construction schedule. Targeted voice alerts and push notifications to safety personnel should be sent to draw necessary attention. This directly triggers equipment to slow down and shut down, injects high-potential energy into the risk field to warn the entire system, and forcibly assigns the nearest camera for evidence collection and collaborative reconstruction of all cameras. Overall, these elements together constitute a complete mathematical mapping link from the manifold geometric representation of joint motion to physical safety intervention at the construction site, ensuring that the measurement of violations has physical consistency, the criteria have statistical completeness, and the response has risk proportionality.
[0070] Using mobile machinery and static hazardous boundaries as source terms, this invention employs a physical information neural network to solve the risk potential diffusion equation in real time, obtaining the dynamic risk potential field at the current and future moments. The reason for using mobile machinery and static hazardous boundaries as the sources of the risk potential field is that their risk generation mechanisms are fundamentally different: the hazard of mobile machinery changes constantly with its position, mass, and speed, requiring injection as a time-varying point source in the equation; the static boundary, on the other hand, reflects its isolation effect by controlling the distribution of the diffusion coefficient at the boundary. By introducing a partial differential equation for risk potential diffusion, the risk is modeled as a continuous physical field that can propagate, attenuate, and superimpose, rather than an isolated geometric distance detection. This allows the risk to diffuse from a high-potential region to the surrounding area, be blocked by obstacles, and naturally attenuate after the source disappears, conforming to the dynamic propagation laws of hazards in the real world.
[0071] The physical information neural network is used to solve this equation in real time because traditional numerical methods require re-dividing and iterating the entire three-dimensional space at each time step, resulting in huge computational overhead and making it difficult to meet the millisecond-level real-time monitoring requirements. In contrast, the physical information neural network takes spatiotemporal coordinates as input and directly outputs the potential field value. Once the network completes online fine-tuning, a single inference can instantly obtain the potential field value and its partial derivative at any spatial point. Simultaneously, its automatic differentiation mechanism can analytically provide the precise time partial derivative and spatial gradient required for the material derivative, providing a continuously differentiable computational basis for subsequent early warning, thus enabling the early warning to possess true forward predictive capabilities. The calculation formula for the risk potential diffusion equation is as follows:
[0072] ,
[0073] in, For dynamic risk potential field, Let be the local partial derivative of the dynamic risk potential field with respect to time. The spatial diffusion coefficient is... Let be the spatial gradient vector of the dynamic risk potential field. For indexing mobile hazard sources, For the first A mobile machine at any time instantaneous dangerous rate, It is a spatial position vector. For the first A mobile machine at any time Spatial position vector, Let be the Dirac distribution function. This is the attenuation coefficient.
[0074] Using a structurally constrained neural radiation field, the foot points of the same person detected from multiple perspectives are reconstructed in three dimensions to obtain their precise spatial coordinates in the world coordinate system; the formula for calculating the precise spatial coordinates in the world coordinate system in S5 is as follows:
[0075] ,
[0076] ,
[0077] ,
[0078] in, Provides the precise spatial coordinates of personnel in the world coordinate system at the current moment. For the initial fusion coordinates, Let the unit normal vector be the plane in which the person is standing. For the constant term of the plane equation, Let Euclidean norm be the normal vector. For camera number indexing, For the first Confidence weights for each perspective For the first The 3D estimated coordinates of the foot points independently recovered by each camera For the first The three-dimensional coordinates of the optical center of a camera in the world coordinate system. The number of spatial points uniformly sampled on the light source. For sampling point index, The distance between two sampling points. To constrain the neural radiation field at the sampling point The volume density value, It is a spatial position vector. Let be the unit vector pointing the light ray to the scene. For the first The Euclidean distance from each sampling point to the camera's optical center. The spatial diffusion coefficient uses different diffusion intensities in different areas; a large diffusion coefficient in open areas indicates that the risk easily spreads to the surroundings; a small diffusion coefficient at fences and obstacles indicates that the risk propagation is hindered. This spatial dependence allows the potential field to reflect the constraint effect of the actual physical layout of the construction site on hazard propagation. The spatial gradient of the potential field points in the direction of the fastest increase in risk potential value; its magnitude represents the degree of spatial non-uniformity of the risk and is the potential difference driving diffusion. The combination of divergence and gradient allows the risk potential energy to flow naturally from high-potential areas to low-potential areas, simulating the process by which risk gradually weakens and is diluted by the surrounding environment as distance increases in space, rather than simply setting a hard boundary distance threshold. The index of the moving hazard source and the instantaneous hazard power are considered. The hazard at the construction site is not constant; the threat level posed by the same crane when stationary and rotating at full speed is drastically different. The hazard power is calculated in real time by the equipment mass and current speed, converting the physical kinetic energy of the machinery into the injection intensity of the risk potential field, dynamically linking the potential field source strength with the actual movement state of the equipment. The spatial location vector, with the real-time position of each mobile device provided by the target tracking module, enables the source term to move synchronously with the device. The Dirac distribution function, which focuses risk injection strictly on the current spatial point of the machine, is an idealized point source model. The summation sign accumulates the contributions of multiple moving hazard sources in the scene, allowing the potential field to comprehensively reflect the complex working condition of multiple sources coexisting. The product of the attenuation coefficient and the potential field causes the risk potential to decay exponentially over time when there is no continuous source injection. This reflects the time-sensitivity of the hazard: after a piece of equipment stops moving, the threat it poses is not permanent but gradually decreases over time and with environmental absorption. The attenuation coefficient controls the dissipation rate, preventing the infinite accumulation of historical risk potentials from interfering with current judgments. A physical information neural network is chosen as the solver because traditional numerical methods require re-dividing and iterating the entire three-dimensional space at each time step, resulting in excessive computational overhead and failing to meet the millisecond-level real-time monitoring requirements. The physical information neural network trains by embedding equation residuals, initial conditions, and boundary conditions into a loss function. Once the network converges, the potential field value and partial derivatives at any spatiotemporal point can be obtained instantaneously through a single forward propagation. More importantly, its automatic differentiation mechanism can analytically provide the time partial derivative and spatial gradient required for material derivative early warning, without the need for numerical difference approximation, resulting in high computational accuracy and no truncation error. The diffusion term allows the risk to propagate reasonably from the hazard source to the surrounding area, the source term allows the risk to be dynamically injected with mobile devices, and the attenuation term allows the risk to dissipate naturally after the source disappears. Together, these three factors make the risk potential field a kind of physical continuous field. The physical information neural network ensures that this complex partial differential equation can be solved in real time at the millisecond level.
[0079] Based on the dynamic risk potential field, the material derivative of the dynamic risk potential field is calculated along the direction of movement for all detected personnel. If the material derivative exceeds the safety limit, a dangerous intrusion warning is triggered. The formula for calculating the material derivative of the dynamic risk potential field in S6 is as follows:
[0080] ,
[0081] ,
[0082] in, The material derivative of the dynamic risk potential field. For personnel at all times The three-dimensional position coordinates in the world coordinate system. Let be the local partial derivative of the dynamic risk potential field with respect to time. Let be the spatial gradient vector of the dynamic risk potential field. Let be the instantaneous velocity vector of the person. The time interval between two adjacent frames. For the same person at any time The three-dimensional position coordinates are given. This invention uses the material derivative of the potential field as the early warning criterion. The material derivative consists of the sum of two terms. The first term is the local time partial derivative of the potential field, which represents the rate of change of the risk potential value with time at a fixed spatial point. It reflects the dynamic change of the hazard source itself, and the risk potential field around it is continuously increasing. Even if the person is stationary, the perceived risk is increasing. The second term is the inner product of the person's velocity vector and the spatial gradient of the potential field. It represents the additional rate of change of potential value felt by the person when moving in the direction of increasing risk potential value due to traversing the spatially uneven potential field. This term is positive and large when the person rushes towards the high potential area, and negative when moving away from the high potential area. The sum of the two terms fully describes the superposition of the effects of the potential field itself changing and the person moving in the potential field. Physically, it is equivalent to the instantaneous rate of change of risk potential energy actually felt by the person along their own movement trajectory. When the material derivative exceeds the preset safety limit, an early warning is immediately triggered. This means that even if the risk potential of the personnel's current location is low, the system can issue an early warning as long as they are rushing towards a high-risk area at a dangerous high speed, or if a high-risk area is approaching the personnel at a dangerous speed. The local temporal partial derivative and spatial gradient are both analytically obtained by the physical information neural network at the personnel's current position and time through an automatic differentiation mechanism. This eliminates the need for numerical difference approximation, ensuring high accuracy and no truncation error, thus guaranteeing the continuity and real-time performance of derivative calculations. The personnel's instantaneous velocity vector is obtained by multi-frame positional difference of skeletal key points and smoothed by Kalman filtering. This reflects the true direction of movement while suppressing velocity estimation jitter caused by detection noise. The safety limit threshold is a preset positive real number, serving as a sensitivity adjustment parameter for triggering the early warning. It can be flexibly set according to construction safety specifications, achieving a leap from passive ranging to active extrapolation.
[0083] The dynamic risk potential field distribution and the precise spatial coordinates of the personnel are input into the distributed mirror descent coordinator, driving multiple PTZ cameras to conduct collaborative patrols with the goal of minimizing uncovered risk potential and field-of-view overlap. When the highest level of violation occurs, the PTZ camera closest to the personnel is forcibly assigned to perform evidence-gathering zoom, while the remaining cameras reconstruct coverage. The attitude update rule for each PTZ camera is as follows:
[0084] ,
[0085] ,
[0086] in, For the first The PTZ camera in the first Update the pose vector after the next iteration. For the PTZ camera's index number, For distributed iteration steps, For camera Feasible attitude space For the loss function in gradient at, For the first The PTZ camera in the first Update the pose vector after the next iteration. For the mirror descent step size, Let Bregman divergence be the property of the divergence. Let be the Bregman potential function. For candidate pose vectors, Candidate pose vector The potential function value at that point, For attitude vector The potential function value at that point, For attitude vector The gradient vector at the location. The distributed mirror descent coordinator and attitude update rules in this step adopt these factors to achieve optimal coverage and energy consumption balance of multiple PTZ cameras over the dynamic risk potential field without the need for a central control node. This fundamentally overcomes the three major defects commonly found in traditional independent scheduling or single-agent methods: overlapping and redundant field of view, coexisting coverage blind spots, and lack of relay evidence collection capabilities for serious violations. Existing PTZ scheduling typically relies on fixed patrol routes, manual joysticks, or single-agent reinforcement learning. Fixed routes are disconnected from the actual risk distribution; high-risk areas may be left unattended while low-risk areas are repeatedly patrolled. Manual joysticks rely on human intervention and cannot be globally optimized. Single-agent reinforcement learning treats the joint action space of all cameras as an exponentially exponentially expanding high-dimensional space, making convergence difficult and unable to constrain overlapping field of view. When serious violations occur, these methods are even less able to automatically coordinate the collaborative division of labor, with one camera collecting close-up evidence and the other cameras filling in the gaps.
[0087] This invention employs a distributed mirror descent coordinator, with the core design being: a distributed architecture where each camera independently maintains its own local loss function, exchanging information only with neighbors within its communication range, without needing to transmit the global state back to the central server. This avoids single-point failures and communication bottlenecks at the central node, making it suitable for dynamically changing network conditions at construction sites. The local loss function consists of two parts: the risk potential quality of the area not covered within the camera's field of view, and the penalty for overlap with neighboring cameras' fields of view. Bregman divergence and the Bregman potential function are used because the PTZ camera gimbal operates on a spherical / toroidal manifold. If ordinary Euclidean distance is used as the update constraint, the difference between 359° and 1° will be incorrectly treated as a large distance. Bregman divergence, induced by a convex potential function suitable for spherical geometry, accurately reflects the distance between two angles measured by the true rotation amount, ensuring that optimization is always performed under the correct manifold constraints. Mirror descent iterative updates calculate a linear approximation loss at each step based on the current gradient, using Bregman divergence as a penalty term for approaching the current pose, ultimately solving a closed-form optimization problem to obtain the new pose. This conservative yet stable update strategy avoids the disruption of video analysis stability caused by drastic attitude changes. The neighbor consensus term ensures that each camera asymptotically approaches a globally consistent optimum under distributed communication constraints through exchanging Lagrange multipliers with its neighbors, eliminating the need for a central node. Mathematically, this guarantees convergence to a stationary point in the global risk coverage function. Forced evidence collection and coverage reconstruction occur when the system identifies the highest level of serious violation. The camera forced to collect evidence temporarily exits the distributed iteration, and its neighbors immediately remove that area from their own loss functions and reassign coverage responsibility through the consensus term, achieving seamless replacement. This collaborative relay is something that a single agent or independent scheduling cannot accomplish.
[0088] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments for application in other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A computer vision-based method for monitoring the safety of security engineering construction, characterized in that, Includes the following steps: S1. Acquire video streams of the construction scene captured by multiple cameras, and perform adaptive restoration of each frame image using the reaction-diffusion equation guided by visual inappropriateness to obtain an enhanced image; S2. Input the enhanced image into a convolutional neural network with a field-aligned dynamic receptive field to detect small targets such as people and safety equipment in the image; S3. Extract the skeletal key point sequence from the detected personnel, calculate the Riemann torsional energy spectrum of the local transformation between adjacent frames on the joint rotation group manifold, and combine the peak value and standard deviation of the Riemann torsional energy spectrum to form a violation score. Compare the violation score with preset first threshold, second threshold, and third threshold, where the first threshold < second threshold < third threshold. When the violation score is less than the first threshold, it is judged as normal; when it is between the first and second thresholds, it is identified as a minor violation and triggered. A level-two response is triggered when the level falls between the second and third thresholds, identifying it as a moderate violation. A level-3 response is triggered when the level is greater than or equal to the third threshold, which is identified as a serious violation, i.e., the highest level of violation. Level response; S4. Using mobile mechanical equipment and static hazardous boundaries as source terms, the risk potential diffusion equation is solved in real time through a physical information neural network to obtain the dynamic risk potential field at the current and future moments. S5. Using the neural radiation field of structural constraints, the foot points of the same person detected from multiple perspectives are reconstructed in three dimensions to obtain their precise spatial coordinates in the world coordinate system. S6. Based on the dynamic risk potential field, calculate the material derivative of the dynamic risk potential field along the direction of movement for the positions of all detected personnel. If the material derivative exceeds the safety limit, trigger a dangerous intrusion warning. S7. Input the dynamic risk potential field distribution and the precise spatial coordinates of the personnel into the distributed mirror descent coordinator, drive multiple PTZ cameras to conduct collaborative patrols with the goal of minimizing the uncovered risk potential and field of view overlap. When the highest level of violation exists, forcibly assign the PTZ camera closest to the personnel to perform evidence-gathering zoom, and reconstruct the coverage of the remaining cameras.
2. The method for monitoring the safety of security engineering construction based on computer vision according to claim 1, characterized in that, The formula for adaptive restoration in S1 using the response-diffusion equation guided by visual inadequacy is as follows: , in, The instantaneous rate of change in image brightness values during their evolution. The standard deviation is Gaussian kernel, In the pseudo-moment Pixel position The evolution of the image brightness value at that location, Pixel position The initial observed image brightness value, These are the global weighting coefficients. For natural scene statistical models, The brightness of a local image patch centered on the current pixel. This is a sharpness parameter that controls the sensitivity of visual asymmetry to divergence. The KL divergence between the local distribution and the natural statistical model. In pixels The gradient statistical distribution of a local image patch centered on the image. For natural scene statistical models, For divergence operators, In the pseudo-moment Pixel position Spatial gradient of the evolution image at that location.
3. The method for monitoring the safety of security engineering construction based on computer vision according to claim 1, characterized in that, In the convolutional neural network with field-aligned dynamic receptive fields in S2, the sampling position of the convolutional kernel undergoes anisotropic deformation through the local structure tensor field, and its offset formula is: , , , in, For the deformed first The coordinate offset of each sampling point relative to the kernel center It is an orthogonal matrix composed of two eigenvectors. The anisotropic strength coefficient, The local structural coherence coefficient, For the standard convolution kernel, the first The original coordinate vector of each sampling point relative to the kernel center For structure tensor The first major feature, For structure tensor The second major feature, To prevent division by zero constant, for The corresponding unit eigenvector, for The corresponding unit eigenvector, The coordinates of the current convolution kernel center in the image are the two-dimensional pixel coordinates.
4. The computer vision-based security engineering construction safety monitoring method according to claim 1, characterized in that, The formula for calculating the Riemann torsional energy spectrum of the local transformation between adjacent frames in S3 is as follows: , , , , in, The Riemann torsional energy spectrum is the result of local transformation between adjacent frames. It is a six-dimensional torsion vector. For measuring tensors Induced Riemann inner product, A six-dimensional torsion vector transpose, For the Riemannian metric tensor, For the first The absolute pose matrix of the joint at frame time. From The 3×3 rotation matrix extracted from it. Rotational sub-blocks of Riemannian metric tensors For the translation sub-block of the Riemannian metric tensor, Let be the transformation matrix. The matrix representation in Lie algebra encodes the instantaneous motion information of the joint. For the first Frame to the The local rigid body transformation matrix of a specific joint between frames. For standard linear isomorphism, For instantaneous angular velocity components, For instantaneous linear velocity components, Using the time index of the video frame, the formula for calculating the violation score is: , , in, To score points for violations, The Riemann torsional energy spectrum is the result of local transformation between adjacent frames. The weights are positive real number combinations. The total number of frames in the action sequence. Let be the arithmetic mean energy of the sequence.
5. The method for monitoring the safety of security engineering construction based on computer vision according to claim 1, characterized in that, The risk level classification rules in S3 are as follows: , in, To score points for violations, The first threshold, The second threshold, The third threshold, , , Based on risk level, the tiered response measures are as follows: The system only sends discreet vibrations or alerts to the smart safety devices worn by the violators, without disturbing other workers. The system plays a pre-set voice warning to the area via directional loudspeakers at the construction site, and simultaneously pushes screenshots of violations and their locations to the handheld terminal of the on-duty safety officer. The system immediately sends an emergency deceleration or shutdown command to the construction equipment controller associated with the area; injects the location coordinates of the violator as an additional potential energy term into the source term of the risk potential diffusion equation described in step S4; forcibly schedules the PTZ camera closest to the violator to turn to that location and perform maximum optical zoom recording; the remaining PTZ cameras are automatically redistributed to fill scheduling gaps through the distributed mirror descent coordinator described in step S7.
6. The method for monitoring the safety of security engineering construction based on computer vision according to claim 1, characterized in that, The formula for calculating the risk potential diffusion equation in S4 is as follows: , in, For dynamic risk potential field, Let be the local partial derivative of the dynamic risk potential field with respect to time. The spatial diffusion coefficient is... Let be the spatial gradient vector of the dynamic risk potential field. For indexing mobile hazard sources, For the first A mobile machine at any time instantaneous dangerous rate, It is a spatial position vector. For the first A mobile machine at any time Spatial position vector, Let be the Dirac distribution function. This is the attenuation coefficient.
7. The method for monitoring the safety of security engineering construction based on computer vision according to claim 1, characterized in that, The formula for calculating precise spatial coordinates in the world coordinate system in S5 is as follows: , , , in, Provides the precise spatial coordinates of personnel in the world coordinate system at the current moment. As the initial fusion coordinates, Let the unit normal vector be the plane in which the person is standing. For the constant term of the plane equation, Let Euclidean norm be the normal vector. For camera number indexing, For the first Confidence weights for each perspective For the first The 3D estimated coordinates of the foot points independently recovered by each camera For the first The three-dimensional coordinates of the optical center of a camera in the world coordinate system. The number of spatial points uniformly sampled on the light source. For sampling point index, The distance between two sampling points. To constrain the neural radiation field at the sampling point The volume density value, It is a spatial position vector. Let be the unit vector pointing the light ray to the scene. For the first The Euclidean distance from each sampling point to the optical center of the camera.
8. The method for monitoring the safety of security engineering construction based on computer vision according to claim 1, characterized in that, The formula for calculating the material derivative of the dynamic risk potential field in S6 is as follows: , , in, The material derivative of the dynamic risk potential field. For personnel at all times The three-dimensional position coordinates in the world coordinate system. Let be the local partial derivative of the dynamic risk potential field with respect to time. Let be the spatial gradient vector of the dynamic risk potential field. Let be the instantaneous velocity vector of the person. The time interval between two adjacent frames. For the same person at any time The three-dimensional position coordinates.
9. A computer vision-based method for monitoring the safety of security engineering construction, as described in claim 1, is characterized in that, The attitude update rule for each PTZ camera in the distributed mirror descent coordinator of S7 is as follows: , , in, For the first The PTZ camera in the first Update the pose vector after the next iteration. For the PTZ camera's index number, For distributed iteration steps, For camera Feasible attitude space For the loss function in gradient at, For the first The PTZ camera in the first Update the pose vector after the next iteration. For the mirror descent step size, Let Bregman divergence be the property of the divergence. Let be the Bregman potential function. For candidate pose vectors, Candidate pose vector The potential function value at that point, For attitude vector The potential function value at that point, For attitude vector The gradient vector at that point.