Substation video fusion monitoring method and system based on 3DGS
By constructing a 3DGS-based substation video fusion monitoring system, the system establishes a mapping relationship between equipment spatial coordinates and cameras using LiDAR point clouds and multi-view camera images. This solves the problems of inaccurate multi-view stitching and positioning in substation monitoring, and realizes real-time monitoring and interactive analysis at the equipment level, thereby improving monitoring efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QINHUANGDAO POWER SUPPLY COMPANY OF STATE GRID JIBEI ELECTRIC POWER COMPANY
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-05
AI Technical Summary
In existing substation monitoring systems, multi-view video surveillance suffers from fragmented perspectives, inconsistent spatial positioning, and weak correlation between camera images, making it difficult to achieve device-level "what you see is what you get" monitoring and interactive analysis. Furthermore, the process of splicing and fusion of multi-view videos is prone to boundary misalignment and spatial mapping deviation.
By constructing a 3DGS-based substation video fusion monitoring system, a three-dimensional Gaussian splash model is built using LiDAR point clouds and multi-view camera images. A mapping relationship library between equipment spatial coordinates and monitoring cameras is established. The system uses confidence adjustment of image stitching boundaries and deviation correction pointers of spatial coordinate mapping for fusion training, dynamically selects the optimal camera view, and realizes the linkage monitoring of the three-dimensional model and real-time video.
It improves the accuracy of image stitching and equipment positioning in substation equipment monitoring, enhances monitoring efficiency and status recognition accuracy, and ensures stable fusion and accurate analysis of multi-view videos.
Smart Images

Figure CN121982218A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of 3D modeling technology, specifically to a method and system for video fusion monitoring of substations based on 3DGS. Background Technology
[0002] With the continuous expansion of power systems, substations, as critical infrastructure in the power grid, have a significant impact on grid stability due to their operational safety and intelligent maintenance level. Existing substations typically deploy multiple fixed cameras and PTZ cameras to provide 24 / 7 video monitoring of primary and secondary equipment, combined with manual inspections or simple video analysis to achieve equipment status awareness. However, due to the complex spatial structure of substations, the diverse types of equipment, and severe obstruction, traditional two-dimensional video monitoring methods struggle to accurately represent the three-dimensional spatial relationships between devices, resulting in problems such as fragmented perspectives, inconsistent spatial positioning, and weak correlation between cross-camera images.
[0003] In recent years, 3D modeling technology that combines LiDAR point clouds with multi-view vision has been gradually applied to power scenarios. However, existing 3D reconstruction methods mostly focus on geometric modeling or static display, making it difficult to deeply integrate with real-time video monitoring and achieve device-level "what you see is what you get" monitoring and interactive analysis. Furthermore, in multi-camera collaborative monitoring scenarios, camera viewpoint selection and switching typically rely on manual configuration or simple rules, making it difficult to dynamically select the optimal viewpoint based on device spatial location, occlusion, and real-time image quality, thus affecting monitoring efficiency and status recognition accuracy. In addition, due to calibration errors, pose drift, and scene differences between different camera viewpoints, boundary misalignment and spatial mapping deviations are prone to occur during multi-view video stitching and fusion, further restricting accurate monitoring and intelligent analysis based on 3D models. Summary of the Invention
[0004] This application provides a 3DGS-based method and system for video fusion monitoring of substations, which solves the technical problems of poor image stitching and equipment positioning accuracy caused by the limitations of multiple perspectives and field of view in substation equipment monitoring.
[0005] The first aspect of this application provides a substation video fusion monitoring method based on 3DGS, the method comprising: Obtain the lidar point cloud and multi-view camera images of the substation, and construct a three-dimensional Gaussian splash model containing device geometry and appearance information; based on P substation 3DGS samples at the source view and Q substation 3DGS samples at the target view, establish a mapping relationship library between the device spatial coordinates and the monitoring cameras; based on the mapping relationship library, use the confidence adjustment pointer for the picture stitching boundary and the deviation correction pointer for the spatial coordinate mapping to perform fusion training; when the user clicks on the target device in the three-dimensional Gaussian splash model, call the optimal camera dynamically filtered based on the improved line-of-sight cone analysis algorithm in the on-site deployment environment, receive the video stream of the optimal camera, and determine the device status recognition result with spatial coordinates.
[0006] In the second aspect of this application, a substation video fusion monitoring system based on 3DGS is provided, and the system includes: Model construction module: Obtain the lidar point cloud and multi-view camera images of the substation, and construct a three-dimensional Gaussian splash model containing device geometry and appearance information; Mapping relationship establishment module: Based on P substation 3DGS samples at the source view and Q substation 3DGS samples at the target view, establish a mapping relationship library between the device spatial coordinates and the monitoring cameras; Fusion training module: Based on the mapping relationship library, use the confidence adjustment pointer for the picture stitching boundary and the deviation correction pointer for the spatial coordinate mapping to perform fusion training; Result screening module: When the user clicks on the target device in the three-dimensional Gaussian splash model, call the optimal camera dynamically filtered based on the improved line-of-sight cone analysis algorithm in the on-site deployment environment, receive the video stream of the optimal camera, and determine the device status recognition result with spatial coordinates.
[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages: First, fuse the lidar point cloud and multi-view camera images of the substation, and construct a three-dimensional Gaussian splash model that simultaneously represents the three-dimensional structure and appearance characteristics of the device. Subsequently, select multiple groups of 3DGS samples from different perspectives, establish the corresponding mapping relationship between the device spatial coordinates and each monitoring camera, and on this basis, introduce a picture stitching confidence adjustment and spatial coordinate deviation correction mechanism to perform joint fusion training on the multi-view video. In practical applications, when the user selects a certain device in the three-dimensional model, the most suitable monitoring perspective is automatically selected from the on-site cameras through the improved line-of-sight cone analysis algorithm, the corresponding video stream is obtained, and the device status recognition result consistent with the device spatial position is output, realizing the linkage monitoring of the three-dimensional model and the real-time video. Description of the Drawings
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a schematic diagram of the substation video fusion monitoring method based on 3DGS provided in the embodiments of this application.
[0010] Figure 2 This is a schematic diagram of the structure of a 3DGS-based substation video fusion monitoring system provided in an embodiment of this application.
[0011] Figure labeling: Model building module 11, mapping relationship establishment module 12, fusion training module 13, result filtering module 14. Detailed Implementation
[0012] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0013] Example 1, as Figure 1 As shown, this application provides a 3DGS-based substation video fusion monitoring method, wherein the method includes: Acquire point cloud data from LiDAR and multi-view camera images of the substation, and construct a 3D Gaussian splash model containing equipment geometry and appearance information.
[0014] In this embodiment, a lidar system and multi-view cameras are first deployed at the substation site. These multi-view cameras include fixed cameras and PTZ cameras. During the same inspection cycle, the system synchronously collects data on the substation equipment and environment. The lidar outputs 3D point cloud data containing spatial coordinates and reflection intensity information, while the multi-view cameras output video frames covering the main equipment area, busbar area, and bay area, collected from different directions. Subsequently, the collected data is preprocessed, including denoising and downsampling of the point cloud to remove floating points and dynamic interference points; distortion correction and exposure / white balance normalization are performed on the images, and time-series alignment of the point cloud and images is achieved using timestamps. Afterward, sensor extrinsic parameter calibration and coordinate unification are performed, i.e., the coordinate systems of each camera and the radar are transformed to the substation's unified world coordinate system, obtaining the intrinsic parameter matrix and extrinsic parameter pose of each camera, as well as the spatial distribution of the point cloud in the world coordinate system. Based on the unified coordinate system, the system constructs a 3D Gaussian splash model using the geometric constraints provided by the point cloud and the texture / color information provided by the multi-view images. Specifically, the scene is represented as a set of 3D Gaussian primitives, each corresponding to a spatial center point. Its parameters include at least the center coordinates, anisotropic covariance, opacity, color, and spherical harmonic coefficients (used to represent view-dependent appearance). The anisotropic covariance describes the local geometry and orientation; opacity represents occlusion and visibility; and color and spherical harmonic coefficients represent view-dependent appearance. Initially, point cloud points can be used as the initial value for the Gaussian center, and the initial covariance value is estimated based on the principal direction of the local neighborhood of the point cloud. The appearance parameters are obtained by projecting the Gaussian onto the imaging plane of each camera and sampling the initial color value at the corresponding pixel. Then, a 3D Gaussian training process based on differentiable rendering is used to project and render the Gaussian primitives at various viewpoints to generate synthetic images. These images are then aligned with the real-world images for photometric consistency. The spatial position, covariance, appearance, and opacity parameters of the Gaussian are iteratively optimized to ensure the rendering results are consistent with the real images across multiple viewpoints. During optimization, point cloud constraints, such as distance constraints from the Gaussian center to the point cloud or sparse regularization, can be incorporated to stabilize the geometric structure. After training convergence, a 3D Gaussian splash model (3DGS) of the substation is obtained, which contains information on the geometric shape and appearance texture of the equipment. This 3D Gaussian splash model of the substation can be rendered in real time from any viewing angle and provides a unified 3D reference for subsequent equipment spatial coordinate positioning and video fusion.
[0015] Based on P 3DGS samples of substations from the source perspective and Q 3DGS samples of substations from the target perspective, a mapping relationship library between equipment spatial coordinates and monitoring cameras is established.
[0016] In one embodiment, after constructing the 3D Gaussian splash model of the substation, to achieve a stable correlation between equipment spatial coordinates and monitoring camera viewpoints, the monitoring cameras within the substation are first divided into a source viewpoint set and a target viewpoint set according to their installation location and field of view coverage. The source viewpoint set provides a high-confidence benchmark for geometry and appearance, while the target viewpoint set covers camera viewpoints that may be invoked during actual operation monitoring. Subsequently, P 3DGS samples are selected from the source viewpoint set, each corresponding to a camera's rendering observation of the 3D Gaussian splash model in a specific frame and its associated real image. Then, Q 3DGS samples are selected from the target viewpoint set, each corresponding to a rendering observation and real image from the target camera's viewpoint. For each sample, its camera intrinsic and extrinsic parameters (pose, field of view, focal length, pan / tilt angle, and timestamp) are uniformly recorded. Afterward, using the spatial coordinates of each device in the 3D Gaussian splash model as a device spatial coordinate entry, cross-view projection calculations are performed on each entry to obtain the predicted pixel position and scale range of the device in the corresponding camera's view. To enhance mapping accuracy, P samples from the source view are used as geometric consistency constraints to align the projected area of the device in the source view with the device detection box / segmentation mask in the real image. This reverses the correction of minor deviations in camera extrinsic parameters or device coordinate offsets, and the corrected parameters are transferred to Q samples in the target view. Cross-view alignment is achieved by optimizing device status recognition accuracy and multi-view stitching error. Finally, a mapping record of "device spatial coordinates - pixel projection parameters" is formed for each camera, and the confidence index of the mapping, the preset position information of the camera, and the callable range are stored simultaneously. The confidence index can be quantified based on projection residuals and occlusion ratios. The above mapping record is written into the mapping relationship database with the device ID or device spatial coordinates as the primary key and the camera ID, view type, or preset position as the index. This allows the system to quickly retrieve the matching set of candidate cameras and the corresponding pixel positioning results based on the 3D spatial coordinates of any device during subsequent runtime, providing basic data support for multi-view fusion training and online optimal camera selection.
[0017] Furthermore, the method includes: A multimodal sensing sample library for substations is constructed; the equipment spatial coordinates in the three-dimensional Gaussian splash model are used as the main query feature tensor, and the real-time PTZ camera video stream is used as the feedback correction signal; based on the feedback correction signal, a pose compensation controller is used to automatically correct the preset position offset of the PTZ camera.
[0018] Preferably, to achieve precise linkage control between the 3D Gaussian splash model and the PTZ camera, multimodal sensing sample entries are established for each type of equipment in the substation and added to the substation's multimodal sensing sample library. Each sample entry in this library includes the equipment's spatial coordinates in the 3D Gaussian splash model, the corresponding multi-view video frames, the point set corresponding to the equipment in the LiDAR point cloud, the equipment type label, and the operating status label. During operation, when monitoring or viewpoint correction of a target device is required, the system reads the spatial information of the target device from the 3D Gaussian splash model and encodes it into a main query feature tensor. This main query feature tensor includes at least the device's 3D position vector, spatial scale vector, direction feature vector obtained by Gaussian covariance matrix decomposition, and equipment type embedding vector. This main query feature tensor serves as a unified query entry point, used to quickly retrieve samples matching the device's spatial position and geometric features, along with the corresponding camera preset position reference information, from the multimodal sensing sample library. Subsequently, the PTZ camera associated with the target device is invoked, and its current video stream is received. In consecutive video frames, the corresponding real-time observation results are retrieved to obtain the pixel center coordinates, pixel scale, bounding box stability, and recognition confidence of the target device in the current frame. Then, based on the camera intrinsic and extrinsic parameters recorded in the mapping database, the spatial coordinates of the target device in the 3D Gaussian splash model are projected onto the imaging plane of the current PTZ camera to calculate the theoretical pixel position. By comparing the real-time observed pixel position with the theoretical pixel position, pixel deviation vectors Δu and Δv, as well as scale deviation Δs, are calculated, and these deviations, along with the recognition confidence, form a real-time feedback correction signal. After receiving this feedback correction signal, the pose compensation controller maps the deviation in the pixel plane to the compensation in the gimbal angle space based on the camera's imaging model and the gimbal motion model. The camera imaging model is established based on the camera's intrinsic and extrinsic parameters. The intrinsic parameters include at least the focal length parameter, principal point position, and distortion parameters, while the extrinsic parameters include the camera's position and attitude in the substation's unified coordinate system. These parameters are determined by the camera's own nominal parameters and calibration results. The gimbal motion model describes the correspondence between the PTZ camera's gimbal control parameters and the camera's optical axis direction. It is established based on the gimbal's mechanical structure parameters and control interface. It can map the offset in the pixel plane to the compensation for the gimbal's horizontal angle, pitch angle, and zoom magnification. That is, it maps the horizontal pixel deviation Δu to the gimbal's horizontal angle compensation Δθ, the vertical pixel deviation Δv to the pitch angle compensation Δφ, and the scale deviation Δs to the zoom compensation Δζ. Then, the pose compensation controller introduces a proportional-integral-derivative (PID) adjustment control strategy to smooth the compensation command, and sets the maximum angular velocity and maximum zoom rate thresholds according to the camera's mechanical constraints to generate the final gimbal control command.Finally, the system sends the generated PTZ control commands to the PTZ camera, driving it to perform minor adjustments and continuously receiving new video frames during the adjustment process, forming a closed-loop feedback. When the pixel deviation of the target device is less than a preset threshold in several consecutive frames, and the recognition confidence is consistently higher than the threshold, the current PTZ pose is determined to be in an effective correction state. At this time, the system records the current PTZ horizontal angle, pitch angle, and zoom ratio as new preset position parameters, and writes these parameters, along with the corresponding device spatial coordinates and observation conditions, back to the multimodal sensing sample library for subsequent rapid retrieval and recalibration. Through the above steps, an automatic preset position correction process for the PTZ camera is realized, based on a 3D Gaussian splash model, supported by a multimodal sensing sample library, and driven by real-time video feedback, thereby improving the stability, accuracy, and intelligence level of the monitoring perspective of substation equipment.
[0019] Furthermore, the method includes: The CNN layer extracts local texture features and edge structure features of video frames, and the Transformer layer captures the geometric correlation between multiple perspectives. At the same time, the equipment status recognition accuracy and multi-view stitching error corresponding to the substation multimodal perception sample library are used as optimization targets.
[0020] Optionally, for multi-view video frames from different surveillance cameras, a feature modeling network consisting of convolutional neural network (CNN) layers and Transformer layers is constructed to jointly model the local appearance features and cross-view geometric correlation features of the device. In actual operation, video streams from each camera are sampled at fixed time intervals to obtain a set of multi-view video frames, and each frame image undergoes resolution unification, pixel normalization, and brightness correction preprocessing. The preprocessed single-frame image is then input into the CNN layer for extraction of local texture and edge structure features. In the first convolution stage of the CNN, multiple small-sized convolution kernels, such as 3×3, are used to perform convolution operations on the input image, calculating local weighted sums pixel by pixel to extract the basic texture information and brightness variation features of the image. The convolution result is processed by a non-linear activation function to form a low-level feature map, which is used to characterize the color distribution, subtle textures, and local grayscale differences on the device surface. In the intermediate convolutional layers, low-level feature maps are further abstracted and combined by stacking convolutional layers and introducing deeper channel dimensions. This stage, through multiple convolutions and activation operations, enables the network to produce significant responses to device surface edges, corners, connections, and local structural morphology, thereby extracting stable edge structure features. Simultaneously, pooling operations or strided convolutions can be combined to reduce the feature map resolution, enhancing feature robustness and suppressing background noise. In the deeper layers of the CNN, convolutional kernels with expanded receptive fields or dilated convolutions are introduced to spatially aggregate the mid-layer features, forming a high-dimensional semantic feature tensor containing the overall device outline, key component layout, and local geometric trends. This feature tensor, indexed by spatial location, preserves local texture, edge structure, and preliminary geometric information, providing a foundation for subsequent cross-view modeling. Finally, the CNN layer outputs a multi-channel feature map for each video frame, where each feature vector corresponds to a visual description of a local region within the video frame.
[0021] After obtaining CNN feature maps from different camera viewpoints, they are input into the Transformer layer. The Transformer layer expands the CNN feature maps from each viewpoint into a sequence of feature vectors according to their spatial location, and unifies the features from different viewpoints into an embedding space of the same dimension through linear mapping. Subsequently, positional encoding and viewpoint encoding are superimposed on each feature vector. The positional encoding is used to represent the relative spatial position of the feature in the image, while the viewpoint encoding is composed of the camera number, installation orientation, and device spatial coordinates derived from the 3D Gaussian splash model, used to explicitly introduce 3D geometric priors. In the Transformer encoder, each feature embedding is linearly mapped to generate a query vector, a key vector, and a value vector. For any query vector, the dot product similarity between it and all key vectors is calculated, and attention weights are obtained through normalization operations. These attention weights are used to measure the correlation between features from different viewpoints and spatial locations. This attention mechanism enables the network to automatically focus on the corresponding regions of the same device under different viewpoints, strengthens the feature aggregation of geometrically consistent regions across viewpoints, and reduces the weight of features in irrelevant backgrounds or occluded regions. Subsequently, the calculated attention weights are used to perform a weighted summation of the value vectors, generating a feature representation that integrates multi-view information. Through a multi-head attention parallel mechanism, the network can simultaneously model the geometric correspondence at the overall device contour level and the fine-grained geometric relationships at the key component level. Next, the fusion result is input into a feedforward neural network for nonlinear transformation, and combined with residual connections and layer normalization operations, outputting stable cross-view joint features.
[0022] Then, the cross-view joint features output by the Transformer are fed into the device state recognition branch and the multi-view stitching evaluation branch, respectively. On the one hand, the device state recognition loss function is calculated by comparing it with the device state labels marked in the multimodal perception sample library to improve the state recognition accuracy of different device types. On the other hand, the correspondence of the same device in the multi-view images is predicted based on the joint features, and compared with the actual stitching boundary or projection result to calculate the multi-view stitching error loss, which is used to constrain cross-view geometric consistency. By jointly optimizing the device state recognition accuracy and multi-view stitching error, the parameters of the CNN layer and Transformer layer are updated in reverse, enabling the network to fully perceive the local texture and structural features of the device while stably capturing the geometric correlation between multiple views. This provides a stable and accurate feature foundation for subsequent video fusion and spatially consistent device state recognition.
[0023] Furthermore, the method includes: The type dimension of the equipment status recognition accuracy is set as a hierarchical calculation logic according to the coverage ratio of each equipment type in the substation multimodal perception sample library, and the single-type recognition accuracy and overall recognition accuracy of each equipment type are marked.
[0024] Optionally, firstly, based on the substation multimodal sensing sample library, the system performs type statistics on the equipment included in the sample library to determine the coverage ratio of each equipment type. These equipment types include, but are not limited to, transformers, circuit breakers, disconnect switches, current transformers, voltage transformers, and surge arresters. The system counts the number of samples of each equipment type in the sample library and calculates the proportion of that type of sample to the total number of samples in the sample library, obtaining the coverage ratio parameter corresponding to each equipment type, which is used as a weight reference for subsequent hierarchical calculations. Subsequently, a hierarchical calculation logic for equipment status recognition accuracy is constructed. For each equipment type, the system selects the predicted status corresponding to that type of equipment from the model's output recognition results and compares it one by one with the actual status annotations of that equipment type in the sample library. The system counts the number of correctly identified samples under that type and the total number of samples, and calculates the single-type recognition accuracy of that equipment type by the ratio of the two. This single-type recognition accuracy reflects the independent performance of the model's ability to recognize the status of a specific equipment type. After calculating the single-type recognition accuracy for each device type, the system performs weighted fusion of these accuracy rates based on the coverage ratio of each device type in the multimodal perception sample library to calculate the overall device status recognition accuracy. Specifically, the single-type recognition accuracy for each device type is multiplied by its corresponding coverage ratio weight, and the weighted results for all device types are summed to obtain the overall recognition accuracy reflecting the model's comprehensive performance across all devices. This method ensures that device types with a larger sample size have an influence commensurate with their actual coverage in the overall evaluation, while avoiding the excessive amplification or reduction of the overall indicator by a small number of sample types. Finally, during training and evaluation, the system records and labels the single-type recognition accuracy for each device type and the calculated overall recognition accuracy, and stores these indicators in association with the corresponding training rounds, model parameter versions, and scene information. In subsequent fusion training and system operation, these labeling results can serve as input for confidence adjustment and bias correction pointers, dynamically adjusting the weights of different device types in multi-view fusion and spatial mapping optimization, thereby achieving refined control and continuous optimization of device status recognition performance.
[0025] Furthermore, the method includes: The positioning benchmark for the multi-view stitching error is set as a graded threshold standard based on the pixel distance between the actual stitching boundary and the ideal stitching boundary of the image, and the stitching error level and the maximum deviation area of each stitching area are marked.
[0026] Optionally, to avoid bias in model evaluation and optimization caused by uneven sample sizes across different device types, the theoretical overlap region between different camera viewpoints is first calculated in a unified world coordinate system based on a 3D Gaussian splash model and a mapping library of device spatial coordinates and surveillance cameras. This overlap region is then projected onto the imaging planes of each camera to obtain the theoretical stitching boundary under ideal conditions of no pose error, no occlusion, and no distortion. This ideal stitching boundary can be represented as a boundary curve composed of a set of continuous pixels, serving as the geometric reference for multi-view image fusion. During multi-view video fusion or stitching, the system performs boundary detection and change analysis on the fused image. For example, through brightness gradient abrupt changes, texture discontinuities, or feature matching residual detection, the actual stitching boundary of different viewpoints in the fusion result is extracted. This actual stitching boundary is also represented as a set of pixel-level coordinates and lies in the same pixel coordinate system as the ideal stitching boundary. Subsequently, the correspondence between the ideal stitching boundary and the actual stitching boundary is matched point by point, and the pixel distance deviation between the two in the direction perpendicular to the boundary is calculated to obtain the local offset distribution of the stitching boundary. Based on this, multiple pixel distance threshold ranges are pre-set, such as low deviation threshold, medium deviation threshold, and high deviation threshold, to distinguish different levels of stitching error. If the pixel distance deviation of a certain boundary region falls within the corresponding threshold range, the region is determined to belong to the corresponding stitching error level. Then, the system divides the entire stitching boundary into several stitching regions according to spatial continuity, and calculates the maximum pixel deviation for each region. Based on the statistical results, the system labels the stitching region with the corresponding stitching error level. Simultaneously, the region with the largest pixel distance deviation among all stitching regions is selected and labeled as the maximum deviation region, and its corresponding camera combination, spatial location, and associated device information are recorded. Finally, the error level labels, pixel deviation statistics, and location results of the maximum deviation region for each stitching region are stored as important inputs for the confidence adjustment pointer and spatial coordinate deviation correction pointer during subsequent fusion training. During model training and online runtime, the system can assign higher optimization weights to high error levels or maximum deviation regions, guiding the model to prioritize correcting unstable stitching regions, thereby gradually reducing multi-view stitching errors and improving the overall consistency of the fused image.
[0027] Furthermore, the method includes: Based on the substation multimodal sensing sample library, P 3DGS samples in the source view and Q 3DGS samples in the target view are defined, where P ≥ 5Q; based on the P 3DGS samples and the Q 3DGS samples, the mapping relationship library is established by referring to the multi-view geometric consistency constraint under the cross-scene adaptation mechanism.
[0028] Optionally, the camera viewpoints are first functionally divided based on a multimodal sensing sample library to obtain a source viewpoint set and a target viewpoint set. The source viewpoints are used to provide high-confidence, all-around geometric and appearance benchmarks, typically covering multiple typical observation directions of the device. Their number is significantly greater than that of the target viewpoints, and they include at least the following five basic viewpoints: First, the viewpoint directly above the device, where the camera's optical axis is basically vertically downward and aligned with the spatial center of the device, used to acquire the top structure and overall outline of the device; second, the viewpoint directly in front of the device, where the optical axis is horizontally pointing towards the center of the device, used to acquire the front structural features of the device; third, the viewpoint directly behind the device, where the optical axis is horizontally pointing towards the center of the device, used to supplement the structural information of the back of the device; fourth, the viewpoint directly to the left of the device, where the optical axis is horizontally pointing towards the center of the device, used to describe the left outline and lateral components of the device; and fifth, the viewpoint directly to the right of the device, where the optical axis is horizontally pointing towards the center of the device, used to describe the right outline and symmetrical structure of the device. In addition to the above basic viewpoints, the source viewpoint set may also include auxiliary viewpoints such as obliquely upward and obliquely forward to further enhance geometric coverage. The target viewpoint is used to cover monitoring views that need to be dynamically invoked or switched during actual operation. It is selected from the actual deployed PTZ or fixed cameras, and its number is relatively small to meet the needs of online monitoring and scheduling. Subsequently, P 3DGS samples are selected from the source viewpoint set and Q 3DGS samples are selected from the target viewpoint set, satisfying the quantity constraint of P≥5Q. Each 3DGS sample includes the camera intrinsic and extrinsic parameters, 3D Gaussian splash model parameters, rendered image, and observation results aligned with the real image at the corresponding viewpoint. Through the above quantity constraints, it is ensured that the source viewpoint samples provide sufficient redundancy support for the target viewpoint in terms of geometric direction, scale variation, and visibility, thereby reducing the impact of instability or occlusion of a single observation at the target viewpoint on the establishment of the mapping relationship.
[0029] Subsequently, a multi-view geometric consistency constraint under a cross-scene adaptation mechanism is introduced to jointly align the source view samples and the target view samples. Specifically, using P 3DGS samples from the source view as geometric references, the Gaussian ellipsoid center position, covariance matrix, and appearance parameters corresponding to the device are extracted to construct a stable 3D geometric description of the device under the source view. Then, the Gaussian parameters of the same device in Q 3DGS samples from the target view are compared with the source view references to calculate the geometric consistency error under different viewpoints and scene conditions. For example, the position deviation of the Gaussian center in a unified world coordinate system, the difference in the covariance matrix, and the overlap of the outline after rendering and projection. By setting a multi-view geometric consistency constraint, the 3D representations of the corresponding devices in the source and target views are kept consistent in spatial position, shape scale, and directional distribution. Then, the consistency error is optimized and corrected under the cross-scene adaptation mechanism. Specifically, based on the error distribution between the source and target views, the camera pose parameters, device spatial coordinate mapping parameters, and 3DGS local geometric parameters corresponding to the target view are jointly adjusted. This allows the 3D Gaussian representation under the target view to gradually converge to the geometric reference of the source view, thereby eliminating geometric deviations introduced by camera installation errors, PTZ drift, or scene changes. Finally, based on the results optimized by consistency constraints and cross-scene adaptation, a mapping relationship library between device spatial coordinates and surveillance cameras is established. This library uses device spatial coordinates or device ID as the primary key and records pixel projection relationships, visibility scores, and geometric consistency confidence levels under different cameras, different viewpoints, and different preset position conditions. The source view mapping serves as a high-confidence reference item, and the target view mapping serves as a callable mapping item. Through this method, a stable and reliable mapping relationship library is formed under multi-view and multi-scene conditions, providing fundamental support for subsequent multi-view fusion training and online optimal camera selection.
[0030] Based on the aforementioned mapping relationship library, fusion training is performed using the confidence adjustment pointer of the image stitching boundary and the deviation correction pointer of the spatial coordinate mapping.
[0031] In one embodiment, after constructing the mapping relationship library between device spatial coordinates and surveillance cameras, the system retrieves the corresponding mapping records from the mapping relationship library by device ID, camera ID, and viewpoint type. This retrieves the device spatial coordinates, camera intrinsic and extrinsic parameters, 3DGS parameters, and the predicted pixel projection position of the device in each camera's view. Simultaneously, it loads multi-view video frames at the corresponding time points, forming a fusion training sample set containing device 3D representation, multi-view observations, and mapping relationships. Subsequently, a confidence adjustment pointer for the image stitching boundary is introduced. This pointer is jointly determined by the error level of the stitching boundary, boundary stability indicators, and historical fusion consistency statistics. During fusion training, when performing feature alignment or pixel-level fusion on images from different viewpoints, the system dynamically adjusts the fusion weight of each viewpoint within the stitching area based on the confidence adjustment pointer. Viewpoints with smaller stitching errors and higher confidence are given higher weights, while views with larger errors or instability have their participation reduced, thereby suppressing the interference of unreliable viewpoints on the fusion result. Finally, a deviation correction pointer for spatial coordinate mapping is introduced for geometric correction. This pointer indicates the correction direction and magnitude of the current mapping relationship in terms of spatial location, scale, or orientation. During the fusion training process, when there is a spatial deviation between the rendered result of the 3D Gaussian splash model and the actual video observation, the system performs reverse fine-tuning of the Gaussian center position, covariance parameters, or camera extrinsic parameters based on the deviation correction pointer. This gradually aligns the 3D projection result with the actual observation, thereby reducing spatial mapping errors. Subsequently, the system performs multiple rounds of training with the joint optimization of viewpoint quality loss, occlusion suppression loss, etc. In each training iteration, it considers both image stitching quality and spatial mapping accuracy, and updates model parameters, 3DGS parameters, and mapping relationship parameters through backpropagation. Finally, as the fusion training progresses, the system continuously monitors the changing trends of stitching error and spatial deviation, and updates the confidence index, deviation statistics, and corresponding adjustment pointer parameters in the mapping relationship library based on the latest training results. When a mapping relationship exhibits stable low stitching error and low spatial deviation across multiple training rounds, its confidence level is increased and it is retained as a priority mapping term. Through the aforementioned dual-pointer-based fusion training mechanism, multi-view video and the 3D Gaussian splash model are collaboratively optimized at the image and spatial layers, providing a highly consistent and stable fusion foundation for subsequent online monitoring and equipment status recognition.
[0032] Furthermore, the method includes: Based on P 3DGS samples from the source viewpoint, a source viewpoint set is determined. A Gaussian ellipsoid covariance reference tensor is extracted from the source viewpoint set. Based on Q 3DGS samples from the target viewpoint, a target viewpoint set is determined. Alignment optimization is performed based on the camera video stream corresponding to the PTZ camera, combined with the target viewpoint set, to determine the geometric deformation vector relative to the covariance reference tensor. A cross-scene adaptation mechanism is set up: the Frobenius norm difference between the covariance matrices of the Gaussian ellipsoids of the corresponding devices under the source and target viewpoints is calculated. The Frobenius norm value is used to quantify the degree of geometric distortion in 3DGS reconstruction. A viewpoint-geometry joint alignment network is used for cross-domain optimization to correct the spatial position, covariance parameters, and spherical harmonic coefficients of the Gaussian ellipsoid, thereby offsetting the reconstruction deviation caused by the viewpoint difference. In the shared representation space, a viewpoint discriminator and a gradient inversion unit are introduced for unsupervised adaptation. Simultaneously, a multi-viewpoint geometric consistency constraint term is constructed based on the feedback correction signal.
[0033] Optionally, based on the substation multimodal sensing sample library, P 3DGS samples at the source viewpoints are selected from the samples with completed 3D Gaussian modeling. These source viewpoints are then used to form a source viewpoint set, covering multiple typical observation directions of the equipment. Each source viewpoint contains a set of Gaussian ellipsoid parameters corresponding to the equipment. These Gaussian ellipsoid parameters include at least the center coordinates of the Gaussian ellipsoid, the 3D covariance matrix, and the spherical harmonic coefficients. For each equipment instance in the source viewpoint set, the system registers its corresponding Gaussian ellipsoids in different source viewpoints under a unified world coordinate system, extracts its covariance matrix one by one, performs eigenvalue decomposition on the extracted covariance matrix, extracts the scale distribution along each principal axis, and assigns weights to different source viewpoint samples based on the viewpoint coverage and observation quality. Finally, a Gaussian ellipsoid covariance reference tensor characterizing the standard geometry of the equipment is generated through a weighted summation method, serving as a reference benchmark for subsequent target viewpoint geometric correction and consistency constraints. Subsequently, a target viewpoint set is determined based on Q 3DGS samples at the target viewpoint. This target viewpoint set corresponds to the viewpoints of PTZ cameras or fixed cameras that can be called upon during actual operation. While its number is relatively limited, it is real-time. The system reads the intrinsic and extrinsic parameters of each camera in the target viewpoint set, as well as the current pan-tilt angle information, and binds them to the target domain 3DGS model to form an initial three-dimensional Gaussian representation of the target domain. To ensure the stability of subsequent alignment optimization, the system constrains the Gaussian ellipsoid of the target domain during the initialization phase, ensuring that its center position remains consistent with the device space coordinates under the source viewpoint reference, allowing only adjustments to its covariance and appearance parameters within a limited range.
[0034] During alignment optimization, the system utilizes the extrinsic parameters of each camera in the target view set to rasterize and render the 3DGS model of the target domain, generating a composite view consistent with the current imaging conditions of the PTZ camera. This composite view maintains consistency with the real-time video stream from the PTZ camera in terms of resolution, field of view, and projection model. Next, the generated composite view is aligned pixel-by-pixel with the real-time video stream acquired by the PTZ camera, and the differences between the two in color, brightness, or structural feature space are calculated to form a pixel-level error map. This pixel-level error map serves as the input to the loss function, which, through a backpropagation mechanism, applies to the 3DGS model parameters of the target domain, updating only the Gaussian ellipsoidal covariance matrix related to the target device, gradually approximating its projected contour under the target view to the actual observation. After several rounds of iterative updates, the system performs element-by-element subtraction between the updated target domain Gaussian ellipsoidal covariance matrix and the source view covariance reference tensor to obtain a geometric deformation vector describing the direction and magnitude of geometric changes in the target view. This geometric deformation vector characterizes the degree of deviation of the 3DGS model from the standard geometry under the target view.
[0035] After obtaining the geometric deformation vector, the system calculates the Frobenius norm difference between the Gaussian ellipsoidal covariance matrices of the corresponding devices under the source and target viewpoints. That is, the Frobenius norm is calculated from the difference matrix of the two covariance matrices to obtain a single scalar value, used to quantify the degree of geometric distortion of the 3DGS reconstruction under the target viewpoint relative to the source viewpoint baseline. When the Frobenius norm value exceeds a preset geometric distortion threshold, the system activates a cross-scene adaptation mechanism, feeding the geometric deformation vector, covariance difference, and camera viewpoint parameters as inputs into a viewpoint-geometric joint alignment network. This network models the coupling relationship between viewpoint changes and geometric deformation within a unified feature space and performs joint correction on the 3D Gaussian splash model parameters. The network input consists of the geometric deformation vector, the difference representation of the Gaussian ellipsoidal covariance matrix, and the camera viewpoint parameters. These inputs are normalized and scale-aligned before entering the network and mapped to a unified-dimensional feature representation. Structurally, the viewpoint-geometric joint alignment network includes at least a viewpoint feature encoding branch, a geometric feature encoding branch, and a feature fusion and regression branch. The viewpoint feature encoding branch encodes camera viewpoint parameters through several layers of fully connected networks, converting continuous pose and gimbal parameters into high-dimensional viewpoint embedding vectors. This allows the network to learn the impact patterns of different viewpoint changes on 3D reconstruction errors. The geometric feature encoding branch models geometric deformation vectors and covariance differences, extracting high-level feature representations reflecting the degree and direction of geometric distortion layer by layer through linear mapping and nonlinear activation. After completing their respective encodings, features from the viewpoint and geometric branches are concatenated or weighted in the fusion layer. An attention mechanism adaptively adjusts the relative importance of viewpoint and geometric features, enabling the network to dynamically prioritize viewpoint or geometric information under different scene conditions.
[0036] After feature fusion, the joint features are fed into a regression decoding branch to predict the correction amounts for the parameters of the 3D Gaussian splash model. This regression decoding branch outputs the displacement correction amount of the Gaussian ellipsoid center position, the incremental correction term of the covariance matrix, and the adjustment parameters of the appearance spherical harmonic coefficients through multiple layers of nonlinear mapping. The correction for the center position compensates for the overall spatial drift caused by camera pose shift or projection errors; the correction for the covariance parameters corrects local geometric stretching, compression, or orientation deviations; and the correction for the spherical harmonic coefficients compensates for appearance inconsistencies caused by changes in viewpoint. During joint optimization, the network aims to minimize the geometric consistency loss and appearance consistency loss between the 3D Gaussian rendering results from the source and target viewpoints. Simultaneously, a covariance difference regularization term defined by the Frobenius norm is introduced to constrain the predicted parameter update magnitude. Subsequently, through the backpropagation mechanism, the weight parameters of the view-geometric joint alignment network, as well as the center, covariance, and spherical harmonic coefficients of the target domain Gaussian ellipsoid, are jointly updated. This allows the 3D Gaussian representation under the target viewpoint to gradually converge to the source viewpoint reference in terms of geometric structure and appearance, thereby effectively offsetting the 3D reconstruction deviation caused by viewpoint changes, gimbal offset, or scene condition differences.
[0037] During cross-scene adaptation, the system introduces a view discriminator into the shared representation space to distinguish whether the current Gaussian features originate from the source or target view domain. Simultaneously, a gradient inversion unit is introduced between the backbone network and the view discriminator, enabling the backbone network to gradually eliminate differences in view-related features during training, achieving consistency in the feature distributions of the source and target views. Meanwhile, based on feedback correction signals from the PTZ camera, the system constructs a multi-view geometric consistency constraint term in the shared representation space. This term applies consistency constraints to the center position and covariance parameters of the corresponding Gaussian ellipsoids of the same device under different viewpoints, ensuring stable spatial relationships across different perspectives. Through these joint constraints and optimizations, the 3D Gaussian splash model maintains geometric and appearance consistency across viewpoints and scenes, providing a high-precision geometric foundation for subsequent mapping relation library construction and multi-view fusion training.
[0038] Furthermore, the method includes: In the shared representation space, the confidence adjustment pointer of the image stitching boundary is set by combining the single-type recognition accuracy and the overall recognition accuracy; in the shared representation space, the deviation correction pointer of the spatial coordinate mapping is set by combining the positioning error level and the maximum visual blind zone area of each device type.
[0039] Optionally, after completing multi-view feature modeling and cross-scene adaptation, the system introduces a pointer generation mechanism for fusion control in the shared representation space to construct confidence adjustment pointers for image stitching boundaries and deviation correction pointers for spatial coordinate mapping. First, in the shared representation space, the single-type recognition accuracy and the weighted overall recognition accuracy for each device type are used as performance prior parameters and introduced into the shared representation space for correlation modeling with the joint features after multi-view fusion. Specifically, the single-type recognition accuracy is mapped to device-level confidence weights to reflect the reliability of the current model's recognition results for that device type, and the overall recognition accuracy is used as a global stability factor to characterize the overall credibility of the current fusion model across all devices in the system. The system fuses the aforementioned performance priors with the feature vectors in the shared representation space through linear mapping or gating mechanisms to generate intermediate confidence feature representations describing the fusion credibility under different devices and perspectives. Based on this, the system sets a confidence adjustment pointer for the stitching boundary of multi-view images. Specifically, it locates the feature subspace corresponding to the stitching boundary region in the shared representation space and performs joint analysis on the features in this subspace with the aforementioned confidence feature representation. When the single-type recognition accuracy of a certain device type is high and the overall recognition accuracy is within a stable range, the fusion confidence level of that device in its corresponding stitching boundary region is increased, thereby generating a higher confidence adjustment pointer. Conversely, when the single-type or overall recognition accuracy decreases, the confidence adjustment pointer value for that stitching boundary region is correspondingly decreased. This confidence adjustment pointer is used to dynamically adjust the fusion weights of different viewpoint images at the stitching boundary during subsequent fusion training or online fusion, ensuring that high-confidence viewpoints dominate in the boundary region, thereby suppressing stitching artifacts caused by recognition instability. Subsequently, in the same shared representation space, a deviation correction pointer for spatial coordinate mapping is set, combining the spatial positioning error level and maximum visual blind zone information of each device type. Specifically, the positioning error levels statistically obtained from different device types during multi-view mapping are encoded as discrete or continuous error intensity features. The location, scale, and corresponding camera viewpoint information of the maximum visual blind spot are encoded as spatial uncertainty features. These features are fused with the device geometric representation in the shared representation space to assess the reliability of the current device spatial coordinate mapping. When a device type has a high positioning error level or is located in the maximum visual blind spot area at a specific viewpoint or region, the system generates a deviation correction pointer pointing to that device and its corresponding mapping relationship. The pointer's value and direction indicate the spatial coordinate mapping items that require focused correction. Finally, the system stores the confidence adjustment pointer of the image stitching boundary and the deviation correction pointer of the spatial coordinate mapping in a structured form in the shared representation space, and associates them with the corresponding device ID, camera ID, and viewpoint information.In subsequent fusion training and online operation, the two types of pointers mentioned above serve as control signals to participate in the allocation of image fusion weights and the updating of spatial mapping parameters, respectively. This guides the system to enhance fusion strength in reliable areas and prioritize correction in geometrically unstable areas, thereby achieving a synergistic improvement in the quality of multi-view video fusion and the spatial positioning accuracy of the device.
[0040] When a user clicks on a target device in a 3D Gaussian splash model, the optimal camera, dynamically selected based on an improved line-of-sight cone analysis algorithm, is invoked in the on-site deployment environment. The video stream from the optimal camera is received, and the device status recognition result with spatial coordinates is determined.
[0041] In one embodiment, when a user clicks on a target device through an interactive interface in a 3D Gaussian splash model, the system obtains the device's spatial coordinates corresponding to the user's click location. This includes the device's 3D center position in a unified world coordinate system, its spatial enclosing range, and the local geometric orientation information provided by the 3D Gaussian splash model. These spatial coordinates serve as the starting point for line-of-sight analysis. Simultaneously, the system reads the parameter information of all currently online surveillance cameras from a mapping database. Subsequently, the system performs visibility determination on candidate cameras based on an improved line-of-sight cone analysis algorithm. Specifically, using the target device's spatial coordinates as the vertex and the optical center position of each candidate camera as the center of the cone's base, a reverse line-of-sight cone is constructed pointing from the device to the camera. The opening angle of this line-of-sight cone is determined by the camera's current field of view and zoom ratio. Inside the cone, the system combines the Gaussian ellipsoid opacity, depth sorting results, and scene occlusion information from the 3D Gaussian splash model to sample and analyze the spatial path between the device and the camera, calculating the target device's visibility ratio, potential occlusion level, and effective projected area from the camera's perspective. Unlike traditional line-of-sight analysis based solely on geometric connections, this improved line-of-sight cone analysis algorithm further incorporates a continuum representation of a 3D Gaussian splash model. This eliminates the occlusion judgment's reliance on discrete triangular meshes, instead assessing visibility by accumulating Gaussian opacity along the line-of-sight direction, thus enhancing analysis accuracy in complex, densely populated scenes. After completing the visibility analysis, the system guides the 3D Gaussian rendering engine to achieve collaborative optimization of high-confidence device identification and low spatial drift in the target scene to determine the optimal camera. Once the optimal camera is determined, the system sends a call command to the corresponding camera. If it is a PTZ camera, the system performs necessary gimbal rotation and zoom adjustments based on the device spatial coordinates-panel parameter mapping results recorded in the mapping database, ensuring the target device is located within the area of interest in the image. Then, the system receives the real-time video stream returned by the optimal camera and performs device detection and status recognition processing within the video frames, binding the identified device status results with the device's spatial coordinates in the 3D Gaussian splash model. Finally, the system synchronously displays the device status recognition results with spatial coordinates in the 3D Gaussian splash model and the monitoring interface. Specifically, the corresponding device is highlighted in the 3D model and its current status information is overlaid. Simultaneously, the device identification area and status label are marked in the video feed. These results can also be recorded as monitoring events with timestamps and spatial coordinates for subsequent inspection analysis and maintenance decisions. Through this process, users can automatically access the optimal monitoring perspective by clicking on the device in the 3D model and obtain spatially consistent and highly reliable device status recognition results.
[0042] Furthermore, the method includes: Based on the mapping relationship library, the confidence adjustment pointer and the deviation correction pointer are used to perform fusion training from the source view to the target view: the model parameter update is driven by the joint optimization objective, which includes view quality loss, occlusion suppression loss, multi-view geometric consistency loss and pose correction regularization term; the pose correction regularization term is dynamically weighted according to the statistical characteristics of the largest visual blind zone region, guiding the 3D Gaussian rendering engine to achieve collaborative optimization of high confidence device recognition and low spatial drift in the target scene, and automatically generating a 3D inspection report containing alarm markers and status data.
[0043] Optionally, based on the established mapping database, the system uses source-view 3DGS samples as high-confidence geometric and appearance references and target-view 3DGS samples as optimization objects. It introduces confidence adjustment pointers and bias correction pointers to perform fusion training on the 3D Gaussian splash model and camera mapping relationship from the source view to the target view, and updates model parameters by jointly optimizing the target. At the beginning of training, the system first constructs source-target view sample pairs based on the device ID and view index recorded in the mapping database. For the same device instance, the source view sample provides its Gaussian ellipsoid parameters, spherical harmonic coefficients, and corresponding rendering results that are stably converged under multi-directional, high-coverage conditions, while the target view sample provides the video stream acquired by PTZ or a fixed camera under actual monitoring conditions and its initial 3DGS representation. The system aligns these samples in a unified world coordinate system and uses the device space coordinates as anchor points for cross-view association. During fusion training, the confidence adjustment pointer is first applied to the feature and pixel fusion layer. The system continuously adjusts the contribution weights of the source and target viewpoints in different spatial regions, especially within the image stitching boundary region, based on the confidence level adjustment pointer. When the device type corresponding to a certain stitching boundary has a high single-type recognition accuracy and the overall recognition accuracy is within a stable range, the features of the corresponding viewpoint in that region are given higher weight during the fusion process; conversely, their weight is reduced, weakening the network's reliance on unreliable stitching regions during backpropagation. In this way, the confidence level adjustment pointer directly affects the proportion of multi-view features in the fusion result, thereby suppressing stitching artifacts and boundary instability during the training phase. Simultaneously, the bias correction pointer acts on the geometric parameter update path. Based on the device type positioning error level and the maximum visual blind zone information indicated by the bias correction pointer, the system applies differentiated update constraints to the Gaussian ellipsoid center position, covariance matrix, and camera extrinsic parameters of the corresponding device during backpropagation. For mapping relationships in areas with high error levels or the largest visual blind spots, the deviation correction pointer guides the model to prioritize correcting spatial projection deviations during updates and restricts the direction of parameter updates to converge towards the source viewpoint geometric reference. For areas with smaller positioning errors, the model is allowed to make more flexible parameter adjustments to improve its expressive ability under the target viewpoint.
[0044] In each training iteration, the system simultaneously calculates and weights a multi-variable loss function. This multi-variable loss function is defined by a joint optimization objective and includes view quality loss, occlusion suppression loss, multi-view geometric consistency loss, and pose correction regularization term. Among them, view quality loss measures the difference between the 3D Gaussian rendering result and the ideal observation in terms of sharpness, effective coverage area, and device saliency under the target viewpoint. Its gradient mainly acts on the appearance parameters and spherical harmonic coefficient updates of the Gaussian ellipsoid. Occlusion suppression loss is based on the cumulative opacity distribution along the line of sight in the 3D Gaussian splash model and penalizes the rendering results of occluded or partially occluded areas. Its gradient mainly acts on the covariance parameter and opacity-related parameters to guide the model to reduce occlusion interference. Multi-view geometric consistency loss ensures that the center position, principal axis direction of covariance, and scale of the Gaussian ellipsoid of the same device under the source viewpoint and the target viewpoint are consistent. Its gradient directly acts on the geometric parameter update path to ensure cross-view geometric stability. Furthermore, the pose correction regularization term plays a constraining and stabilizing role in the joint optimization. Based on the characteristics of the maximum visual blind zone region statistically obtained from the mapping relation library, the system sets dynamic weights for the pose correction regularization term. When the spatial coordinates of the device involved in the training sample are located in the maximum visual blind zone region, or when its historical positioning error level is high, the system automatically increases the weight of the pose correction regularization term, imposing stronger constraints on the update amplitude of the camera extrinsic parameters and the center position of the Gaussian ellipsoid, preventing spatial drift under weak observation conditions. When the sample is in an area with good visibility and sufficient geometric constraints, the weight of this regularization term is correspondingly reduced, enabling the model to have stronger fitting ability while ensuring stability. Through this dynamic weighting mechanism based on the statistical characteristics of the visual blind zone, the 3D Gaussian rendering engine is guided to achieve collaborative optimization between high-confidence device recognition and low spatial drift in the target scene.
[0045] After the fusion training is completed and the preset convergence conditions are met, the system solidifies the optimized 3D Gaussian splash model parameters, mapping relationship parameters, and pointer states into a running version. During actual deployment, the system uses this running version to identify the device status of the real-time video stream. When the identification result triggers an anomaly threshold or alarm rule, the system automatically summarizes the spatial coordinates, status category, occurrence time, and associated viewpoint information of the abnormal device, and generates a corresponding alarm marker in the 3D Gaussian splash model. At the same time, the system integrates the above-mentioned spatialized status data, alarm records, and key viewpoint screenshots to automatically generate a 3D inspection report containing 3D spatial positioning, device status description, and historical change information, thereby realizing intelligent inspection and maintenance decision-making directly supported by the fusion training results.
[0046] In summary, the embodiments of this application have at least the following technical effects: First, point clouds from the substation's LiDAR and images from multi-view cameras are acquired to construct a 3D Gaussian splash model containing equipment geometry and appearance information. Next, based on P 3DGS samples from the source view and Q 3DGS samples from the target view, a mapping library between equipment spatial coordinates and monitoring cameras is established. Then, based on this mapping library, fusion training is performed using a confidence adjustment pointer for image stitching boundaries and a deviation correction pointer for spatial coordinate mapping. Finally, when a user clicks on a target device in the 3D Gaussian splash model, the optimal camera dynamically selected based on an improved line-of-sight cone analysis algorithm is invoked in the field deployment environment. The optimal camera's video stream is received, and the device status recognition result with spatial coordinates is determined. This solves the technical problem of poor image stitching and equipment positioning accuracy in substation equipment monitoring due to limitations in multi-view and field of view. It achieves the technical effect of improving the monitoring accuracy and efficiency of substation equipment by dynamically selecting the optimal camera and spatial coordinates for device status recognition.
[0047] Example 2 is based on the same inventive concept as the 3DGS-based substation video fusion monitoring method in the previous examples, such as... Figure 2 As shown, this application provides a 3DGS-based substation video fusion monitoring system, wherein the system includes: Model building module 11: Acquires substation lidar point cloud and multi-view camera images to construct a 3D Gaussian splash model containing equipment geometry and appearance information; Mapping relationship establishment module 12: Based on P substation 3DGS samples from the source view and Q substation 3DGS samples from the target view, establishes a mapping relationship library between equipment spatial coordinates and monitoring cameras; Fusion training module 13: Based on the mapping relationship library, uses the confidence adjustment pointer of the image stitching boundary and the deviation correction pointer of the spatial coordinate mapping to perform fusion training; Result filtering module 14: When the user clicks on the target device in the 3D Gaussian splash model, the optimal camera obtained by dynamically filtering based on the improved line-of-sight cone analysis algorithm is called in the on-site deployment environment, the video stream of the optimal camera is received, and the equipment status recognition result with spatial coordinates is determined.
[0048] Furthermore, the mapping relationship establishment module 12 is used to perform the following method: A multimodal sensing sample library for substations is constructed; the equipment spatial coordinates in the three-dimensional Gaussian splash model are used as the main query feature tensor, and the real-time PTZ camera video stream is used as the feedback correction signal; based on the feedback correction signal, a pose compensation controller is used to automatically correct the preset position offset of the PTZ camera.
[0049] Furthermore, the mapping relationship establishment module 12 is used to perform the following method: The CNN layer extracts local texture features and edge structure features of video frames, and the Transformer layer captures the geometric correlation between multiple perspectives. At the same time, the equipment status recognition accuracy and multi-view stitching error corresponding to the substation multimodal perception sample library are used as optimization targets.
[0050] Furthermore, the mapping relationship establishment module 12 is used to perform the following method: The type dimension of the equipment status recognition accuracy is set as a hierarchical calculation logic according to the coverage ratio of each equipment type in the substation multimodal perception sample library, and the single-type recognition accuracy and overall recognition accuracy of each equipment type are marked.
[0051] Furthermore, the mapping relationship establishment module 12 is used to perform the following method: The positioning benchmark for the multi-view stitching error is set as a graded threshold standard based on the pixel distance between the actual stitching boundary and the ideal stitching boundary of the image, and the stitching error level and the maximum deviation area of each stitching area are marked.
[0052] Furthermore, the mapping relationship establishment module 12 is used to perform the following method: Based on the substation multimodal sensing sample library, P 3DGS samples in the source view and Q 3DGS samples in the target view are defined, where P ≥ 5Q; based on the P 3DGS samples and the Q 3DGS samples, the mapping relationship library is established by referring to the multi-view geometric consistency constraint under the cross-scene adaptation mechanism.
[0053] Furthermore, the fusion training module 13 is used to perform the following methods: Based on P 3DGS samples from the source viewpoint, a source viewpoint set is determined. A Gaussian ellipsoid covariance reference tensor is extracted from the source viewpoint set. Based on Q 3DGS samples from the target viewpoint, a target viewpoint set is determined. Alignment optimization is performed based on the camera video stream corresponding to the PTZ camera, combined with the target viewpoint set, to determine the geometric deformation vector relative to the covariance reference tensor. A cross-scene adaptation mechanism is set up: the Frobenius norm difference between the covariance matrices of the Gaussian ellipsoids of the corresponding devices under the source and target viewpoints is calculated. The Frobenius norm value is used to quantify the degree of geometric distortion in 3DGS reconstruction. A viewpoint-geometry joint alignment network is used for cross-domain optimization to correct the spatial position, covariance parameters, and spherical harmonic coefficients of the Gaussian ellipsoid, thereby offsetting the reconstruction deviation caused by the viewpoint difference. In the shared representation space, a viewpoint discriminator and a gradient inversion unit are introduced for unsupervised adaptation. Simultaneously, a multi-viewpoint geometric consistency constraint term is constructed based on the feedback correction signal.
[0054] Furthermore, the fusion training module 13 is used to perform the following methods: In the shared representation space, the confidence adjustment pointer of the image stitching boundary is set by combining the single-type recognition accuracy and the overall recognition accuracy; in the shared representation space, the deviation correction pointer of the spatial coordinate mapping is set by combining the positioning error level and the maximum visual blind zone area of each device type.
[0055] Furthermore, the result filtering module 14 is used to perform the following method: Based on the mapping relationship library, the confidence adjustment pointer and the deviation correction pointer are used to perform fusion training from the source view to the target view: the model parameter update is driven by the joint optimization objective, which includes view quality loss, occlusion suppression loss, multi-view geometric consistency loss and pose correction regularization term; the pose correction regularization term is dynamically weighted according to the statistical characteristics of the largest visual blind zone region, guiding the 3D Gaussian rendering engine to achieve collaborative optimization of high confidence device recognition and low spatial drift in the target scene, and automatically generating a 3D inspection report containing alarm markers and status data.
[0056] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A substation video fusion monitoring method based on 3DGS, characterized in that, The method includes: Acquire point cloud data from LiDAR and multi-view camera images of the substation, and construct a 3D Gaussian splash model containing equipment geometry and appearance information; Based on P 3DGS samples of substations from the source perspective and Q 3DGS samples of substations from the target perspective, a mapping relationship library between equipment spatial coordinates and monitoring cameras is established. Based on the aforementioned mapping relationship library, fusion training is performed using the confidence adjustment pointer of the image stitching boundary and the deviation correction pointer of the spatial coordinate mapping. When a user clicks on a target device in a 3D Gaussian splash model, the optimal camera, dynamically selected based on an improved line-of-sight cone analysis algorithm, is invoked in the on-site deployment environment. The video stream from the optimal camera is received, and the device status recognition result with spatial coordinates is determined.
2. The substation video fusion monitoring method based on 3DGS as described in claim 1, characterized in that, The method includes: Construct a multimodal sensing sample library for substations; The device space coordinates in the three-dimensional Gaussian splash model are used as the main query feature tensor, and the real-time PTZ camera video stream is used as the feedback correction signal. Based on the feedback correction signal, the pose compensation controller automatically corrects the preset position offset of the PTZ camera.
3. The substation video fusion monitoring method based on 3DGS as described in claim 2, characterized in that, The method includes: CNN layers are used to extract local texture features and edge structure features of video frames, while Transformer layers are used to capture the geometric relationships between multiple viewpoints. Meanwhile, the equipment status recognition accuracy and multi-view stitching error corresponding to the substation multimodal sensing sample library are used as optimization targets.
4. The substation video fusion monitoring method based on 3DGS as described in claim 3, characterized in that, The method includes: The type dimension of the equipment status recognition accuracy is set as a hierarchical calculation logic according to the coverage ratio of each equipment type in the substation multimodal perception sample library, and the single-type recognition accuracy and overall recognition accuracy of each equipment type are marked.
5. The substation video fusion monitoring method based on 3DGS as described in claim 4, characterized in that, The method includes: The positioning benchmark for the multi-view stitching error is set as a graded threshold standard based on the pixel distance between the actual stitching boundary and the ideal stitching boundary of the image, and the stitching error level and the maximum deviation area of each stitching area are marked.
6. The substation video fusion monitoring method based on 3DGS as described in claim 4, characterized in that, The method includes: Based on the substation multimodal sensing sample library, P 3DGS samples in the source view and Q 3DGS samples in the target view are defined, where P ≥ 5Q; Based on the P 3DGS samples and the Q 3DGS samples, and in accordance with the multi-view geometric consistency constraints under the cross-scene adaptation mechanism, the mapping relationship library is established.
7. The substation video fusion monitoring method based on 3DGS as described in claim 6, characterized in that, The method includes: Determine the source view set based on P 3DGS samples at the source view; The Gaussian ellipsoid covariance reference tensor is extracted from the aforementioned source view set. Determine the target viewpoint set based on Q 3DGS samples from the target viewpoint; Based on the camera video stream corresponding to the PTZ camera, alignment optimization is performed in conjunction with the target viewpoint set to determine the geometric deformation vector relative to the covariance reference tensor. Set up a cross-scene adaptation mechanism: calculate the difference in Frobenius norm of the covariance matrix of the Gaussian ellipsoid of the corresponding device under the source view and the target view, quantify the degree of geometric distortion of 3DGS reconstruction with the Frobenius norm value, and use the view-geometry joint alignment network for cross-domain optimization to correct the spatial position, covariance parameter and spherical harmonic coefficient of the Gaussian ellipsoid to offset the reconstruction deviation caused by the view difference. In the shared representation space, a view discriminator and a gradient inversion unit are introduced for unsupervised adaptation. At the same time, a multi-view geometric consistency constraint term is constructed based on the feedback correction signal.
8. The substation video fusion monitoring method based on 3DGS as described in claim 7, characterized in that, The method includes: In the shared representation space, the confidence adjustment pointer of the image stitching boundary is set by combining the single-type recognition accuracy and the overall recognition accuracy. In the shared representation space, a deviation correction pointer for the spatial coordinate mapping is set based on the positioning error level and the maximum visual blind zone area of each device type.
9. The substation video fusion monitoring method based on 3DGS as described in claim 8, characterized in that, The method includes: Based on the mapping relationship library, the confidence adjustment pointer and the bias correction pointer are used to perform fusion training from the source view to the target view: the model parameter update is driven by the joint optimization objective, which includes view quality loss, occlusion suppression loss, multi-view geometric consistency loss and pose correction regularization term; The pose correction regularization term is dynamically weighted based on the statistical characteristics of the largest visual blind spot area, guiding the 3D Gaussian rendering engine to achieve collaborative optimization of high-confidence device recognition and low spatial drift in the target scene, and automatically generating a 3D inspection report containing alarm markers and status data.
10. A 3DGS-based substation video fusion monitoring system, characterized in that, The system is used to implement the 3DGS-based substation video fusion monitoring method according to any one of claims 1-9, the system comprising: Model building module: acquires point cloud data from LiDAR and multi-view camera images of the substation, and builds a 3D Gaussian splash model containing equipment geometry and appearance information; Mapping relationship establishment module: Based on P 3DGS samples of substations from the source view and Q 3DGS samples of substations from the target view, establish a mapping relationship library between equipment spatial coordinates and monitoring cameras; Fusion Training Module: Based on the aforementioned mapping relationship library, fusion training is performed using the confidence adjustment pointer of the image stitching boundary and the deviation correction pointer of the spatial coordinate mapping. Result filtering module: When a user clicks on a target device in the 3D Gaussian splash model, the optimal camera obtained by dynamically filtering based on the improved line-of-sight cone analysis algorithm is called in the on-site deployment environment. The video stream of the optimal camera is received, and the device status recognition result with spatial coordinates is determined.
Citation Information
Cited By
Power grid dispatching mobile intelligent page adaptive generation method based on multi-form data
CN122248135A