Substation 3D Gaussian point cloud-driven camera automatic point distribution method

By using a high-precision 3D Gaussian point cloud model and an AI vision model, combined with a knowledge base of deployment rules and real-time rendering technology, the optimal deployment scheme for substation cameras is generated. This solves the problems of low deployment efficiency and high cost of substation cameras, and achieves accurate, efficient deployment results and verifiability.

CN121962518APending Publication Date: 2026-05-01JIANGSU HAOHAN INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU HAOHAN INFORMATION TECH
Filing Date
2026-01-04
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, the deployment of substation cameras relies on manual experience, which is inefficient and costly. Furthermore, it is difficult to accurately capture the spatial characteristics of key parts of the equipment in complex 3D environments, and there is a lack of efficient virtual verification methods, resulting in long deployment cycles and high costs.

Method used

Based on a high-precision 3D Gaussian point cloud model, combined with an AI visual model and a point placement rule knowledge base, the optimal camera placement scheme is generated through spatial geometric calculation and intelligent optimization decision-making. High-fidelity virtual view images are then generated using real-time rendering technology to verify the placement effect.

Benefits of technology

It achieves accurate, efficient, and verifiable automatic deployment of cameras in substations, significantly improving the deployment efficiency and reliability of security monitoring, and reducing the cost and error risk of manual deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962518A_ABST
    Figure CN121962518A_ABST
Patent Text Reader

Abstract

The invention provides a transformer substation 3D Gaussian point cloud driven camera automatic point distribution method, which comprises the following steps: based on a high-precision 3D Gaussian point cloud model of a transformer substation, automatically identifying key parts and three-dimensional space characteristics of equipment through an AI visual model; according to a preset point distribution rule knowledge base, according to the key parts of the equipment and the three-dimensional space characteristics of the key parts, space geometric calculation and intelligent optimization decision are executed, and an optimal camera point distribution scheme is obtained; and performing real-time rendering according to parameters in the optimal camera stationing scheme by using a high-precision 3D Gaussian point cloud model, and generating a high-fidelity virtual view angle picture to verify a stationing effect. According to the transformer substation 3D Gaussian point cloud driven camera automatic point distribution method, based on the high-precision 3D Gaussian point cloud model, the AI visual model, the point distribution rule knowledge base and the real-time rendering technology are combined, and the accuracy, high efficiency and verifiability of transformer substation camera automatic point distribution are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Gaussian point cloud technology, and in particular to an automatic camera placement method driven by 3D Gaussian point clouds in substations. Background Technology

[0002] Currently, with the rapid development and increasing intelligence of power systems, the importance of safety monitoring in substations, as core nodes of power transmission and distribution, is becoming increasingly prominent. Traditional substation camera deployment relies heavily on manual experience, determining camera installation locations, orientations, and parameters through on-site surveys and manual design. However, this method suffers from low efficiency, high cost, and susceptibility to human interference, especially in complex 3D environments where it struggles to accurately capture the spatial characteristics of critical equipment components and optimize deployment schemes. In recent years, 3D point cloud technology, using LiDAR or structured light scanning equipment to acquire high-precision 3D data, has made substation scene modeling possible. However, how to utilize point cloud data to achieve automated and intelligent deployment remains a technical challenge. Existing technologies include attempts to identify equipment through point cloud segmentation, but these lack precise extraction of the 3D spatial characteristics (spatial coordinates, dimensions, orientation) of critical equipment components and fail to integrate with power security standards to form a systematic knowledge base of deployment rules. Furthermore, traditional deployment scheme verification typically relies on post-installation effect testing, lacking efficient virtual verification methods, resulting in long deployment cycles and high costs.

[0003] To address the aforementioned issues, there is an urgent need for a method based on a high-precision 3D Gaussian point cloud model, combined with AI visual models, a point placement rule knowledge base, and real-time rendering technology, to automatically identify key parts of equipment, optimize camera placement schemes, and verify the placement effect through high-fidelity virtual images, thereby improving the deployment efficiency and reliability of substation security monitoring. Summary of the Invention

[0004] The purpose of this invention is to provide an automatic camera placement method driven by 3D Gaussian point cloud in substations, so as to solve the problems pointed out in the background art.

[0005] The automatic camera placement method driven by 3D Gaussian point cloud in substations provided in this embodiment of the invention includes: Based on a high-precision 3D Gaussian point cloud model of a substation, AI visual model is used to automatically identify key parts of the equipment and their three-dimensional spatial characteristics. Based on a pre-set knowledge base of placement rules, and according to the key parts of the equipment and their three-dimensional spatial characteristics, spatial geometric calculations and intelligent optimization decisions are performed to obtain the optimal camera placement scheme. Using a high-precision 3D Gaussian point cloud model, high-fidelity virtual view images are generated in real time based on the parameters in the optimal camera placement scheme to verify the placement effect.

[0006] Optionally, the three-dimensional spatial characteristics include: spatial coordinates, dimensions, and orientation.

[0007] Optionally, the deployment rule knowledge base includes: coverage rules, obstruction avoidance rules, key part focus rules, and installation feasibility rules formed by digitizing a large number of power security standards.

[0008] Optionally, the optimal camera placement scheme includes: camera model, installation location, orientation angle, and focal length parameters.

[0009] Optionally, the step of using a high-precision 3D Gaussian point cloud model to perform real-time rendering based on parameters in the optimal camera placement scheme to generate a high-fidelity virtual perspective image includes: By calling the 3D Gaussian point cloud data and simulating the resolution, focal length, and viewing angle parameters of a real camera according to the parameters in the optimal camera placement scheme, a high-quality virtual image with pixel-level precision and consistent with the imaging effect of a real camera is generated, thus obtaining a high-fidelity virtual perspective image.

[0010] Optionally, the automatic camera placement method driven by 3D Gaussian point clouds in substations also includes: When using a high-precision 3D Gaussian point cloud model to render in real time based on the parameters in the optimal camera placement scheme, if the user inputs a personalized camera placement intention scheme, the reason path for the user's input of the personalized camera placement intention scheme can be inferred based on the information displayed to the user during the rendering process within the time window before the user inputs the personalized camera placement intention scheme. Based on the underlying reasons, determine the strategy to confirm users' strong interest in personalized camera deployment solutions, and implement the corresponding confirmation. If the verification is successful, the multimodal cost of accepting and responding to the personalized camera placement intention plan is determined based on the first remaining process after the verification moment in the rendering process. The optimal display strategy for showcasing multimodal costs during the remaining process of embedding the plan; among which, the optimal display strategy can guide users to make a choice about personalized camera placement options in the fastest way; Based on the optimal presentation strategy, the multimodal cost is displayed to the user during the first remaining process execution; When the user completes the selection of their preferred personalized camera placement scheme, the selected scheme is added to the rendering process, completing the second remaining process after the rendering time.

[0011] Optionally, the reasoning path for inferring the user's input of a personalized camera placement intention scheme includes: Through the interaction log analysis module of the real-time rendering engine, the triggering time of the user's input of the personalized camera placement intention plan is accurately captured, and the time window threshold is adaptively determined based on the dynamic change rate of the user's interaction history data. The time window threshold is dynamically adjusted by calculating the weighted entropy value of the user's gaze density, view zoom operation frequency and dwell time variance of the key equipment area during the pre-rendering process. The weight coefficient of the weighted entropy value is dynamically configured according to the preset mapping table of substation equipment type complexity. Based on the time window threshold, extract all multimodal information displayed to the user during the rendering process within the time window before the trigger time, including the device area heat map of the high-fidelity virtual view image, the dynamic curve of the view coverage index, and the spatial geometric conflict warning mark. Semantic parsing of multimodal information is performed using a temporal convolutional neural network model to generate a reason path for the user's input of personalized camera placement intention scheme. The reason path is represented as a user intent tracing map, which includes the focusing trajectory of key parts of the device, the spatial occlusion avoidance motivation sequence, and the viewpoint optimization preference weight distribution. The viewpoint optimization preference weight distribution is dynamically quantified by analyzing the correlation coefficient between the user's interaction intensity with different device areas and the rendering frame rate fluctuation.

[0012] Optionally, the strategy for confirming a strong user preference for a personalized camera deployment plan includes: Real-time monitoring of users' subsequent operational behavior after inputting personalized camera placement intentions, and determination of intention strength through a dual verification mechanism; The dual verification mechanism includes: The first verification involves calculating the cosine similarity between the direction vector of subsequent parameter adjustment operations and the camera pose parameters in the personalized camera placement intention scheme. When the cosine similarity is consistently higher than the preset similarity threshold and the operation frequency exceeds the preset frequency threshold, preliminary confirmation is triggered. The second layer of verification involves analyzing the frequency of intent keywords in the user's voice or text input using a natural language processing model, and then cross-validating this with the click hotspot matching degree of the highlighted area in the rendered interface. When the joint confidence of keyword frequency and hotspot matching degree exceeds the threshold confidence threshold, the user's intent is determined to be strong. If the verification fails, it will automatically revert to the optimal camera placement plan and generate optimization suggestions.

[0013] Optionally, the multimodal cost of determining the acceptance and response to the personalized camera deployment intention plan includes: During the first residual process after the rendering time is confirmed, multimodal costs are quantified in real time, including real-time rendering latency increment, resource usage volatility, and critical device area coverage loss rate. The real-time rendering latency increment is predicted based on a nonlinear regression model of GPU frame generation time and point cloud data volume. The resource usage volatility is calculated by monitoring the instantaneous changes in CUDA core utilization and memory bandwidth. The critical device area coverage loss rate is measured based on the deviation between the visibility probability matrix of critical parts of the device in the 3D Gaussian point cloud model and a preset safety threshold.

[0014] Optionally, the optimal display strategy for multimodal costs within the embedded remaining process of the planning includes: A Markov decision process model is constructed using a deep reinforcement learning algorithm. With the goal of minimizing user decision time, the model dynamically plans the optimal display sequence and interaction form for multimodal costs. The Markov decision process model discretizes the time axis of the first residual process into decision step sizes, takes the priority weight vector of multimodal costs as input, and outputs the optimal display strategy, including priority sorting of highlighted warning icons, overlaying layers of cost impact visualization heatmaps, and timing of voice guidance.

[0015] The present invention has achieved the following beneficial effects: Based on a high-precision 3D Gaussian point cloud model, combined with AI visual models, a deployment rule knowledge base, and real-time rendering technology, the system achieves accurate, efficient, and verifiable automatic deployment of cameras in substations. Step A accurately identifies the 3D spatial characteristics of key equipment components using an AI visual model, providing a data foundation for deployment. Step B utilizes spatial geometric calculations and intelligent optimization decisions to generate a deployment plan that meets security standards. Step C1 generates high-fidelity virtual images through ray tracing rendering to verify the deployment effect. Overall, this approach ensures the coverage, focus, and installation feasibility of the deployment plan, significantly improving the deployment efficiency and reliability of substation security monitoring and reducing the cost and error risk of manual deployment.

[0016] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.

[0017] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of an automatic camera placement method driven by a 3D Gaussian point cloud in a substation, as described in an embodiment of the present invention. Figure 2 This is another schematic diagram of the automatic camera placement method driven by 3D Gaussian point cloud in substations in an embodiment of the present invention. Detailed Implementation

[0019] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0020] The research and development approach of this application is based on a high-precision 3D Gaussian point cloud model, combined with AI visual models, a deployment rule knowledge base, and real-time rendering technology, to construct an automatic deployment method for substation cameras, achieving accurate, efficient, and verifiable monitoring deployment. First, the AI ​​visual model performs semantic segmentation and feature extraction on the 3D Gaussian point cloud data, automatically identifying the three-dimensional spatial characteristics (spatial coordinates, dimensions, and orientation) of key parts of substation equipment, providing a data foundation for deployment. Second, a deployment rule knowledge base is constructed based on digital power security standards. Through spatial geometric calculations and intelligent optimization decisions, an optimal deployment scheme that meets the requirements of coverage, occlusion avoidance, focus, and installation feasibility is generated. Finally, ray tracing rendering technology is used to simulate real camera imaging based on point cloud data, generating high-fidelity virtual view images to verify the deployment effect. The entire process closely revolves around the security needs of substations, ensuring the scientific and practical nature of the deployment scheme, significantly improving deployment efficiency and reducing labor costs.

[0021] Figure 1 A flowchart of an automatic camera placement method driven by a 3D Gaussian point cloud in a substation is provided for embodiments of this application, such as... Figure 1 As shown, the method includes: A. Based on a high-precision 3D Gaussian point cloud model of the substation, an AI visual model is used to automatically identify key parts of the equipment and their three-dimensional spatial characteristics. These three-dimensional spatial characteristics include: spatial coordinates, dimensions, and orientation.

[0022] The working principle of this step is to capture the spatial distribution information of equipment within the substation using a high-precision 3D Gaussian point cloud model, and then use an AI vision model to perform semantic segmentation and feature extraction on the point cloud data to automatically identify key parts of the equipment and their three-dimensional spatial characteristics. Three-dimensional spatial characteristics include spatial coordinates (defined as the three-dimensional position of the key parts of the equipment in the point cloud coordinate system, in meters, usually with the origin of the point cloud model as a reference), dimensions (defined as the length, width, and height of the geometric bounding box of the key parts, in meters, calculated through point cloud clustering), and orientation (defined as the rotation angle of the key parts relative to the principal axes of the point cloud coordinate system, in degrees, determined through principal component analysis). The high-precision 3D Gaussian point cloud model refers to the three-dimensional point cloud data of the substation acquired through LiDAR or structured light scanning equipment. Each point in the point cloud contains its position (x, y, z) and Gaussian distribution parameters (mean and covariance matrix), used to describe the uncertainty of the point's position. The AI ​​vision model employs a deep learning architecture (such as PointNet++ or DGCNN) to perform semantic segmentation on point cloud data using a pre-trained neural network. This divides the point cloud into substation equipment areas (such as transformers and circuit breakers) and non-equipment areas (ground and walls), and further extracts the geometric features of key components (such as transformer heat sinks and circuit breaker control panels). Specifically, the spatial coordinates are obtained by using the origin of the point cloud model as a reference, and calculating the centroid coordinates (x, y) of the key component point set identified through semantic segmentation. c ,y c ,z c The dimensions are obtained by fitting bounding boxes to the key component point set and calculating the length, width, and height of the bounding boxes. The orientation is obtained by performing principal component analysis (PCA) on the key component point set and extracting the angle between the principal direction vector and the principal axes of the coordinate system. The construction process of the AI ​​vision model includes: collecting a cloud dataset of substation sites, labeling the categories of key equipment components, training based on the PointNet++ architecture, and optimizing the loss function (such as cross-entropy loss) to improve semantic segmentation accuracy. Finally, the model outputs the three-dimensional spatial characteristic data of the key components, providing a foundation for subsequent point placement.

[0023] B. Based on a pre-set deployment rule knowledge base, and considering the key components of the equipment and their three-dimensional spatial characteristics, spatial geometric calculations and intelligent optimization decisions are performed to obtain the optimal camera deployment scheme. The deployment rule knowledge base includes: coverage rules, obstruction avoidance rules, key component focus rules, and installation feasibility rules, all derived from the digitization of numerous power security standards.

[0024] The working principle of this step is to utilize a pre-set deployment rule knowledge base, combined with the three-dimensional spatial characteristics of key equipment components, to generate a camera deployment scheme that meets the requirements of power security through spatial geometric calculations and intelligent optimization algorithms. The deployment rule knowledge base refers to a digitally stored set of power security specifications, including coverage rules (defined as the proportion of key equipment components that the camera's viewpoint must cover, at least 90%, calculated using the field of view and equipment size), occlusion avoidance rules (defined as no obstructions between the camera's line of sight and the key components, verified through ray tracing algorithms), key component focus rules (defined as the pixel percentage of the key component in the camera's image, at least 50%, calculated using focal length and distance), and installation feasibility rules (defined as the camera installation location must meet physical constraints, such as a height of 2-5 meters above the ground, filtered through spatial coordinates). Three-dimensional spatial characteristics (spatial coordinates, dimensions, orientation) are obtained from step A and used for spatial geometric calculations. Spatial geometric calculations include: calculating the camera's field of view coverage area using the centroid of key components as the target (based on camera focal length and viewing angle parameters, the coverage area is calculated using trigonometric formulas); using a ray tracing algorithm, rays are emitted from candidate camera positions to key components to determine if occlusion exists; and the pixel percentage is calculated based on the size and distance of the key components to verify focus. Intelligent optimization decision-making employs a genetic algorithm to construct an objective function (combining coverage, occlusion rate, focus, and installation feasibility, with weights of 0.4, 0.3, 0.2, and 0.1 respectively), and iteratively optimizes to select the optimal camera position, orientation, and focal length parameters. The genetic algorithm initializes a population of 100 placement schemes (each scheme includes camera position, orientation, and focal length), iterates for 100 generations, and selects the placement scheme that satisfies all rules. The final output is the optimal camera placement scheme, including camera model (defined as a device model that meets the resolution and focal length range, such as 1080p, 10-50mm focal length), installation position (3D coordinates), orientation angle (angle relative to the principal axis of the coordinate system), and focal length parameter (in millimeters).

[0025] C. Using a high-precision 3D Gaussian point cloud model, render in real time according to the parameters in the optimal camera placement scheme to generate a high-fidelity virtual view image to verify the placement effect. The optimal camera placement scheme includes: camera model, installation position, orientation angle, and focal length parameters. Step C specifically includes the following sub-steps: C1. Call the 3D Gaussian point cloud data, and based on the parameters in the optimal camera placement scheme, simulate the resolution, focal length, and viewing angle parameters of the real camera to generate a high-quality virtual image with pixel-level precision that is consistent with the imaging effect of the real camera, thereby obtaining a high-fidelity virtual perspective image.

[0026] This step works by using a high-precision 3D Gaussian point cloud model and real-time rendering technology to simulate the imaging process of a real camera, generating a high-fidelity virtual view image consistent with the actual camera effect to verify the effectiveness of the point placement scheme. The 3D Gaussian point cloud data contains the positions (x, y, z) of all points within the substation and Gaussian distribution parameters (mean and covariance matrix), used to describe the scene's geometry and surface characteristics. The optimal camera placement scheme provides camera parameters, including model (defined as camera hardware specifications, such as 1080p resolution), installation position (3D coordinates, in meters), orientation angle (angle relative to the principal axis of the coordinate system, in degrees), and focal length parameters (in millimeters, determining the field of view). The simulation process is based on a ray tracing rendering algorithm, constructing a triangular mesh on the point cloud data (generated through Delaunay triangulation) and assigning texture information (through point cloud color or preset materials) to each point. The rendering algorithm simulates light emanating from the camera position based on camera resolution (defined as the number of pixels, such as 1920×1080), focal length (calculated using the perspective projection formula, field of view = 2×arctan(sensor size / (2×focal length))), and viewing angle parameters (determined by orientation angle and position). It calculates the intersection points with the point cloud triangular mesh to generate pixel-level accurate virtual images. Resolution parameters are obtained from the camera model specifications, while focal length and viewing angle parameters are directly read from the point layout scheme. During rendering, the uncertainty of the Gaussian distribution of the point cloud is considered, and Gaussian blur is used to smooth the edges to ensure image realism. The final output high-fidelity virtual viewpoint image (defined as an RGB image consistent with the real camera image, with the same resolution as the camera) is used to verify the coverage and focus of the point layout scheme. The rendering algorithm is accelerated on a GPU, with single-frame rendering time controlled within 1 second.

[0027] Through the above steps, based on a high-precision 3D Gaussian point cloud model, combined with AI visual models, a deployment rule knowledge base, and real-time rendering technology, the automatic deployment of substation cameras is achieved with precision, efficiency, and verifiability. Step A accurately identifies the three-dimensional spatial characteristics of key equipment components through the AI ​​visual model, providing a data foundation for deployment; Step B utilizes spatial geometric calculations and intelligent optimization decisions to generate a deployment scheme that meets security standards; Step C1 generates high-fidelity virtual images through ray tracing rendering to verify the deployment effect. Overall, these steps collectively ensure the coverage, focus, and installation feasibility of the deployment scheme, significantly improving the deployment efficiency and reliability of substation security monitoring, and reducing the cost and error risk of manual deployment.

[0028] like Figure 2 As shown, in some embodiments, the automatic camera placement method driven by 3D Gaussian point clouds in substations further includes: D. During the real-time rendering process using a high-precision 3D Gaussian point cloud model and based on the parameters of the optimal camera placement scheme, if a user inputs a personalized camera placement intention scheme, the reason path for the user's input can be inferred based on the information displayed to the user during the rendering process within the time window prior to the user's input of the personalized camera placement intention scheme. Step D specifically includes the following sub-steps: D1. Through the interaction log analysis module of the real-time rendering engine, the triggering time of the user's input of the personalized camera placement intention plan is accurately captured, and the time window threshold is adaptively determined based on the dynamic change rate of the user's interaction history data. The time window threshold is dynamically adjusted by calculating the weighted entropy value of the user's gaze density, view zoom operation frequency and dwell time variance of the key equipment area during the pre-rendering process. The weight coefficient of the weighted entropy value is dynamically configured according to the preset mapping table of substation equipment type complexity.

[0029] This step works by using the interaction log analysis module of the real-time rendering engine to accurately capture the trigger moment of the user's input of a personalized camera placement intention plan. Based on the dynamic change rate of the user's historical interaction data, it adaptively determines a time window threshold to define the retrospective time range for analyzing user intent. The trigger moment is defined as the point in time when the user submits the personalized camera placement intention plan through interface operations (such as mouse clicks, touch input, or voice commands), accurate to milliseconds, and captured by the rendering engine's logging module. The time window threshold is defined as a dynamic time range (in seconds, typically between 5-30 seconds) preceding the trigger moment, used to extract user interaction behavior to infer intent; its value is dynamically adjusted using a weighted entropy value. Weighted entropy is a measure of the complexity of user interaction behavior, calculated based on three key parameters: gaze density (defined as the number of interaction points of the user in key device areas of the rendered interface per unit time, in points / second, recorded by interface clicks or touch coordinates), view zoom operation frequency (defined as the number of times the user adjusts the viewpoint per unit time, in times / second, counted by the operation log of the rendering engine), and dwell time variance (defined as the statistical variance of the user's interaction dwell time in different device areas, in seconds², calculated by timestamp differences). The weighted entropy is calculated as follows: after normalizing the gaze density, view zoom operation frequency, and dwell time variance, multiply them by weighting coefficients, sum them, and then calculate using the entropy function (based on the information entropy formula). The weighting coefficients are dynamically configured according to a preset mapping table of substation equipment type complexity. Equipment type complexity is defined as a comprehensive score of equipment geometry and monitoring priority (range 0-1, based on a preset table lookup for equipment types such as transformers and circuit breakers). The mapping table stores the correspondence between device types and weights. For example, transformers have high complexity (score 0.8), with weights allocated as follows: gaze point density 0.5, viewpoint zoom operation frequency 0.3, and dwell time variance 0.2; circuit breakers have medium complexity (score 0.5), with weights allocated as 0.4, 0.4, and 0.2. The interaction log analysis module records the timestamps and coordinates of user operations in real time, calculates the dynamic change rate (defined as the change in interaction frequency per unit time, in times / second², obtained by analyzing log data through a sliding window), and adjusts the sensitivity of entropy calculation based on the change rate to ensure that the time window threshold accurately covers the key behavioral sequences that form user intent.

[0030] In a practical implementation example, in a real substation monitoring scenario, suppose a user submits a personalized camera placement plan (e.g., adjusting a camera position to the top of the transformer, coordinates (10,5,3) meters) via mouse clicks on the real-time rendering interface. The interaction log analysis module of the real-time rendering engine captures the trigger time as the 120th second of rendering, with the timestamp accurate to milliseconds. The system extracts the interaction data from the first 30 seconds of the interaction log and calculates the gaze density: the user clicked 20 times in the transformer area, with a time window of 30 seconds, resulting in a gaze density of 20 / 30 = 0.67 points / second; the viewpoint zoom operation frequency is 5 times the user adjusts the viewpoint, resulting in a frequency of 5 / 30 = 0.167 times / second; the dwell time variance is calculated by recording the user's dwell time in the transformer area (e.g., 2 seconds, 3 seconds, 1.5 seconds), with a variance of 0.25 seconds². Assuming the transformer equipment type complexity score is 0.8, a pre-defined mapping table is consulted. The weighting coefficients are: gaze density 0.5, viewpoint zoom operation frequency 0.3, and dwell time variance 0.2. The normalized weighted entropy value is calculated to be 0.72 (processed using the entropy function). The dynamic change rate is analyzed using a sliding window (window size 5 seconds) to examine log data, revealing an interaction frequency change of 0.1 times / second², indicating stable user interaction. The system maps the entropy value of 0.72 to a time window threshold of 15 seconds (based on a pre-defined linear mapping rule: entropy value 0-1 corresponds to a window of 5-30 seconds). Ultimately, the system determines the 15 seconds before the trigger moment as the time window threshold for subsequent analysis of user intent. This process runs in real-time within the rendering engine's log module, with log data stored in a 1MB memory buffer (overlaying 30 seconds of interaction data) to ensure low-latency capture and processing.

[0031] D2. Based on the time window threshold, extract all multimodal information displayed to the user during the rendering process within the time window before the trigger time, including the device area heat map of the high-fidelity virtual view image, the dynamic curve of the view coverage index, and the spatial geometric conflict warning mark.

[0032] This step works by extracting all multimodal information displayed to the user within the time window before the trigger moment from the output data of the real-time rendering engine, based on the time window threshold determined in step D1. This information is then used to infer user intent. The multimodal information includes a device area heatmap (defined as a two-dimensional distribution map reflecting the intensity of user interaction on a high-fidelity virtual viewpoint image, measured in interaction counts per pixel, generated by statistical analysis of interface click coordinates), a dynamic curve of viewpoint coverage (defined as the percentage of key device areas covered by the camera's viewpoint on the timeline, measured in %, calculated using ray tracing), and spatial geometric conflict warning markers (defined as locations of occlusion conflicts between the camera's line of sight and the device, highlighted in red, detected using ray tracing algorithms). The time window threshold (obtained from D1, measured in seconds) defines the scope of the extracted data. The extraction process is implemented through the rendering engine's frame buffer and log database: Device area heatmaps are generated using Gaussian kernel density estimation by recording the distribution of user click coordinates (x, y) on a virtual viewpoint image (resolution such as 1920×1080), with heatmap intensity values ​​ranging from 0 to 1; the dynamic curve of viewpoint coverage is calculated using a ray tracing algorithm to determine the area of ​​key device regions covered by the camera's field of view in each frame (based on a triangular mesh projection of point cloud data), and the ratio to the total device area generates the coverage (%), recorded over time to form a curve; spatial geometric conflict warning markers are created by ray tracing, emitting rays from the camera position to the key device regions, detecting intersections of obstructions (such as walls, other devices), and marking conflict areas (represented by pixel coordinates). Multimodal information is stored in memory as structured data (approximately 1MB per frame, covering all frames within the time window) for subsequent semantic parsing. The extraction process ensures data integrity, maintaining a frame rate of 60 FPS to match real-time rendering requirements.

[0033] Continuing with the above implementation example, in a substation monitoring scenario, assuming step D1 determines the time window threshold to be 15 seconds, and the user submits a personalized deployment plan at the 120th second, the system extracts multimodal information from the rendering engine's frame buffer and log database for the period from the 105th to the 120th second. The process of generating the equipment area heatmap is as follows: the user's click coordinates on the high-fidelity virtual view image (resolution 1920×1080) are recorded. For example, if the user clicks 30 times in the transformer area (pixel range 500×300), a heatmap is generated using Gaussian kernel density estimation (bandwidth 10 pixels). The highest intensity value is 0.85, concentrated in the transformer heat sink area. The dynamic curve generation process for the view coverage index is as follows: The ray tracing algorithm calculates the ratio of the area of ​​the transformer heatsink covered by the camera's field of view (e.g., 1000 square centimeters) in each frame (900 frames in total, 15 seconds × 60 FPS) to the total area of ​​the heatsink (1200 square centimeters), resulting in a coverage of 83.3%, which is recorded as a curve point on the time axis. The overall curve shows that the coverage fluctuates between 80% and 85%. The spatial geometric conflict warning marker generation process is as follows: Ray tracing detects obstructions between the camera's line of sight and the transformer (such as a cable tray located at coordinates (8,4,2) meters), and marks them as red highlighted areas (pixel range 200×150) on the virtual view image. This information is stored in a memory buffer (total size approximately 900MB, covering 900 frames) and extracted through the rendering engine's API interface, taking approximately 0.5 seconds to ensure real-time performance. This data provides a rich information foundation for subsequent intent inference. D3. Semantic parsing of multimodal information is performed using a temporal convolutional neural network model to generate the reason path of the user's input personalized camera placement intention scheme. The reason path is represented as a user intent tracing map, which includes the focusing trajectory of key parts of the device, the spatial occlusion avoidance motivation sequence, and the viewpoint optimization preference weight distribution. The viewpoint optimization preference weight distribution is dynamically quantified by analyzing the correlation coefficient between the user's interaction intensity with different device areas and the rendering frame rate fluctuation.

[0034] The working principle of this step is to use a temporal convolutional neural network (TCN) model to semantically parse the multimodal information extracted in step D2, generating a causal path for the user's personalized camera placement intention scheme, representing the reasons for the formation of the user's intention. The causal path is defined as a user intention tracing map, which includes the focus trajectory of key equipment parts (defined as the sequence of user interaction behavior on key equipment areas on the time axis, represented by coordinate sequences, obtained through heatmap analysis), the spatial occlusion avoidance motivation sequence (defined as the user's action sequence of adjusting the viewpoint to avoid occlusion, represented by timestamps and adjustment parameters, and analyzed by correlation through conflict warning markers), and the viewpoint optimization preference weight distribution (defined as the user's attention priority to different equipment areas, ranging from 0 to 1, calculated by the correlation coefficient between interaction intensity and rendering frame rate fluctuations). The TCN model adopts a multi-layer one-dimensional convolutional structure, and the input is a time series of multimodal information (including heatmap intensity sequence, coverage curve values, and conflict warning marker sequence). The data dimensions of each frame are heatmap (compressed to 256×256), coverage (1-dimensional scalar), and conflict markers (a set of conflict point coordinates). The model extracts temporal features using convolutional kernels (size 3, stride 1), and combines residual connections and dilated convolutions to handle long sequence dependencies, outputting the cause path. The focus trajectory of key equipment parts is extracted from the heatmap intensity peak sequence and mapped to equipment region coordinates (annotated via point cloud semantic segmentation). The spatial occlusion avoidance motivation sequence is generated by analyzing the correlation between the timestamps of conflict warning markers and user viewpoint adjustment operations (obtained from logs). The viewpoint optimization preference weight distribution is normalized to weights (range 0-1) by calculating the Pearson correlation coefficient between interaction intensity (defined as the mean heatmap intensity, in times / pixel) and rendering frame rate fluctuation (defined as the standard deviation of frame generation time, in milliseconds, obtained through GPU performance monitoring). The TCN model is pre-trained on a substation station cloud interaction dataset, with the optimized loss function being mean squared error. Inference time after training is approximately 0.1 seconds / frame. The cause path is stored in structured data (approximately 100KB) for subsequent steps.

[0035] Continuing with the above implementation example, in a substation scenario, assume that D2 extracts 15 seconds (900 frames) of multimodal information, and the TCN model (convolution kernel size 3, number of layers 5, dilation factor 2) parses the data. The input data includes: heatmap (compressed to 256×256, intensity value 0-1), coverage curve (value per frame 80-85%), and conflict warning marker (transformer area 3 times occlusion, coordinates (8,4,2) meters). After model processing, the following cause paths are generated: the focus trajectory of key equipment parts shows that the user is focused on the transformer heatsink (coordinates (10,5,3) meters), with a peak heatmap intensity of 0.85 lasting for 10 seconds; the spatial occlusion avoidance motivation sequence shows that the user adjusts their viewpoint (tilt angle change of 5 degrees) at 110 and 115 seconds, corresponding to the timestamps of the conflict warning markers; the viewpoint optimization preference weight distribution is calculated by taking the correlation coefficient of 0.9 between interaction intensity (heatmap mean 0.7 times / pixel) and frame rate fluctuation (standard deviation 5 milliseconds), resulting in a weight of 0.6 for the transformer heatsink, 0.3 for the circuit breaker, and 0.1 for other areas. The TCN model is inferred on the GPU, with a single frame time of 0.08 seconds and a total time of 72 seconds. The cause paths are stored as structured data (80KB in size), including focus trajectories (10 coordinate sequences), motivation sequences (5 adjustment actions), and weight distributions (weights for 3 areas), providing accurate evidence for subsequent verification of user intent. During implementation, the model was pre-trained on 1,000 sets of cloud-based interactive data from substations, achieving an accuracy rate of 95%.

[0036] E. Based on the underlying reasons, determine the strategy for confirming users' strong interest in personalized camera deployment plans, and execute the corresponding confirmation. Step E specifically includes the following sub-steps: E1. Real-time monitoring of the user's subsequent operational behavior after inputting a personalized camera placement plan, and determination of the intensity of intent through a dual verification mechanism.

[0037] This step works by continuously monitoring the user's subsequent actions after inputting a personalized camera placement plan through the real-time rendering engine's interaction log module. A dual verification mechanism is used to determine the strength of the user's intention regarding this plan. The subsequent action flow is defined as the user's sequence of interactions within a certain period (default 10 seconds) after the trigger moment (obtained from D1), including click coordinates, viewpoint adjustment parameters (pitch angle, yaw angle, in degrees), and input text or voice commands, recorded in the rendering engine's log database (approximately 1KB per second). Intention strength is defined as the user's persistence with the personalized placement plan, categorized as strong (confidence > 90%), moderate (50-90%), and weak (< 50%) to initiate verification. The verification result is output as a confidence score (0-100%) for subsequent steps to confirm the user's intention. The monitoring and verification process runs in the rendering engine's real-time thread, ensuring a latency of less than 0.1 seconds.

[0038] E2. The dual authentication mechanism includes: E21. First verification: Calculate the cosine similarity between the direction vector of subsequent parameter adjustment operations and the camera pose parameters in the personalized camera placement intention scheme. When the cosine similarity is continuously higher than the preset similarity threshold and the operation frequency exceeds the preset frequency threshold, preliminary confirmation is triggered.

[0039] The working principle of this step is to evaluate the consistency between the operation and the plan by calculating the cosine similarity between the direction vector of the user's subsequent parameter adjustment operation and the camera pose parameters in the personalized camera placement intention plan, which serves as the first layer of verification and initial confirmation. The direction vector is defined as the spatial change vector of the user's subsequent adjustment operation, including position changes (Δx, Δy, Δz, in meters) and angle changes (Δpitch angle, Δyaw angle, in degrees), extracted from the interaction log (through operation timestamps and parameter records). The camera pose parameters are defined as the camera position (x, y, z, in meters) and orientation (pitch angle, yaw angle, in degrees) in the personalized placement plan, directly obtained from the user's input plan. The cosine similarity is defined as the cosine value of the angle between the direction vector and the pose parameter vector, ranging from 0 to 1, calculated by dividing the vector dot product by the modulus product. The preset similarity threshold is defined as 0.85 (based on experience to ensure high consistency), and the preset frequency threshold is defined as 2 times per second (based on log statistics of operation counts, in times / second). The verification process extracts operation data from the most recent 10 seconds using a log analysis module, calculates the direction vector for each frame (60 FPS) (using coordinate and angle differences), compares it with the proposed pose vector, and generates a similarity sequence (60 values ​​per second). When the mean of the similarity sequence remains above 0.85 for at least 5 seconds and the operation frequency exceeds 2 times / second, a preliminary confirmation signal (a Boolean value, stored in memory, approximately 1KB) is triggered. The calculation process runs on the CPU, with each frame taking approximately 0.01 seconds.

[0040] Continuing with the above implementation example, in a substation scenario, the user-submitted personalized camera placement plan specifies a camera position of (10, 5, 3) meters, a pitch angle of 30 degrees, and a yaw angle of 45 degrees. The operation log for the following 10 seconds (120-130 seconds) shows: at second 121, the position was adjusted to (10.1, 5.1, 3.1) meters, with a pitch angle of 32 degrees; at second 122, the pitch angle was adjusted to 33 degrees. The system extracts the direction vector: position change (0.1, 0.1, 0.1) meters, angle change (2, 0) degrees. The pose vector of the plan is (10, 5, 3, 30, 45), and the cosine similarity calculated through normalization is 0.93 (position vector similarity 0.95, angle vector similarity 0.91, average 0.93). The operation frequency within 10 seconds is 5 times (0.5 times / second), and the average similarity sequence (600 frames) is 0.92, consistently higher than 0.85. The frequency was less than 2 times / second, initially confirming that it was not triggered. The system recorded the result (confidence level 0.92, unconfirmed status) and stored it in memory (1KB). This example shows that the user's operation is highly consistent with the plan, but the frequency is insufficient, requiring a second verification.

[0041] E22. The second layer of verification involves analyzing the frequency of intent keywords in the user's voice or text input using a natural language processing model, and then performing cross-validation by combining the click hotspot matching degree of the highlighted area in the rendered interface. When the joint confidence of keyword frequency and hotspot matching degree exceeds the threshold confidence threshold, it is determined that the user's intent is strong.

[0042] This step works by analyzing the frequency of intent keywords in the user's speech or text input using a Natural Language Processing (NLP) model. This, combined with the click heatmap matching degree of the highlighted area in the rendered interface, is used to assess the strength of the user's intent through joint confidence, serving as a second layer of verification. Intent keyword frequency is defined as the number of times relevant keywords (such as "transformer" or "heat sink") appear in the user's speech or text input, measured in times, extracted by the NLP model (based on the existing BERT architecture). Click heatmap matching degree is defined as the overlap ratio between the user's click coordinates and the highlighted area of ​​the rendered interface (generated by a D2 heatmap), ranging from 0 to 1, calculated using coordinate geometry. Joint confidence is defined as a weighted average of keyword frequency and matching degree (weights of 0.6 and 0.4, based on experience), ranging from 0 to 100%. The threshold confidence is defined as 90% (to ensure high confidence). The NLP model performs word segmentation and semantic analysis on the speech (via a speech-to-text API) or text input, extracting keywords and counting their frequencies (using word frequency vectors). Click heatmap matching score is calculated by comparing the overlap area between the click coordinates (obtained from logs) and the boundary of the highlighted area (extracted from the heatmap, range such as 500×300 pixels), normalized to 0-1. After joint confidence calculation, if it exceeds 90%, the intention is considered strong (Boolean value, stored in memory, approximately 1KB).

[0043] Continuing the above implementation example, in a substation scenario, the user inputs the voice command "Adjust to the top of the heat sink" at 123 seconds. The NLP model (BERT, pre-trained on a power industry corpus) parses the speech, extracting the keywords "heat sink" (frequency 1) and "top" (frequency 1), with the frequency normalized to 0.8 (based on the maximum expected frequency of 2). The highlighted area on the rendered interface is the transformer heat sink (pixel range 500×300). The user's click coordinates (510, 310) fall within this area, resulting in a matching score of 0.95. The joint confidence score is calculated as 0.8×0.6+0.95×0.4=0.86 (86%), which is below 90%, indicating no strong intent. The system records the result (confidence score 86%, unconfirmed status) and stores it in memory (1KB). If the subsequent voice input "confirm heatsink again" increases in frequency to 2, and is normalized to 1.0, the joint confidence level is 1.0×0.6+0.95×0.4=0.98 (98%), which exceeds 90%, indicating a strong intention.

[0044] E23. If the verification fails, it will automatically revert to the optimal camera placement plan and generate optimization suggestions.

[0045] The working principle of this step is that if the dual verification (E21, E22) fails to confirm the user's intention, the system automatically reverts to the optimal camera placement plan generated in step B and generates optimization suggestions to guide the user to adjust the plan to improve the placement effect. Failure to confirm is defined as the verification result of E21 or E22 not reaching the threshold (similarity < 0.85 or confidence < 90%). The optimal camera placement plan includes camera position, orientation, and focal length parameters (obtained from step B). The reverting process loads the parameters of the optimal plan through the rendering engine, updates the rendering pipeline, and restores the original virtual view image (resolution such as 1920×1080). The optimization suggestions are defined as text suggestions generated based on user interaction data and a placement rule knowledge base, including a problem description (e.g., "too much occlusion") and improvement measures (e.g., "adjust to a higher position"). Suggestions are generated by comparing the differences between the user's plan and the optimal plan (position deviation, coverage loss), combined with placement rules (coverage range, occlusion avoidance), and stored in memory (approximately 10KB). The prompt is displayed as a pop-up window (text length < 100 characters) through the rendered interface, taking approximately 0.1 seconds. Backspace and prompt generation ensure system stability and avoid invalid interactions.

[0046] Continuing with the above implementation example, in the substation scenario, the user's solution (location (10,5,3) meters, elevation angle 30 degrees) has a similarity of 0.93 but a frequency of 0.5 times / second (<2 times / second) in E21 verification, and an E22 confidence level of 86% (<90%), thus failing verification. The system reverts to the optimal solution (location (9,4,2.5) meters, elevation angle 25 degrees), loads parameters through the rendering engine, and restores the virtual view image (takes 0.5 seconds). Comparing the user's solution and the optimal solution, the location deviation is 1.22 meters, and the coverage loss is 5% (user's solution 80%, optimal solution 85%). Combining the occlusion avoidance rules, a suggestion is generated: "The current solution occludes the transformer heat sink. It is recommended to adjust the location to (9,4,2.5) meters to improve coverage." The suggestion is displayed as a pop-up window on the interface (resolution 1920×1080, centered), and the text is stored in memory (8KB).

[0047] F. If the verification is successful, based on the first remaining process after the verification moment in the rendering process, determine the multimodal cost of accepting and responding to the personalized camera placement intention plan.

[0048] Step F specifically includes the following sub-steps: F1. In the first remaining process after the rendering time is confirmed, the multimodal cost is quantified in real time, including the real-time rendering latency increment, resource usage volatility, and critical device area coverage loss rate; wherein, the real-time rendering latency increment is predicted based on a nonlinear regression model of GPU frame generation time and point cloud data volume; the resource usage volatility is calculated by monitoring the instantaneous changes in CUDA core utilization and memory bandwidth; the critical device area coverage loss rate is measured based on the deviation between the visibility probability matrix of critical parts of the device in the 3D Gaussian point cloud model and a preset safety threshold.

[0049] The working principle of this step is to quantify the multimodal cost of the personalized point placement scheme in real time during the first remaining process after the confirmation moment (the time point when E22 determines strong intent) (defined as the time period from confirmation to the end of rendering, usually 10-30 seconds), providing a basis for subsequent optimization. Multimodal costs include real-time rendering latency increment (defined as the increase in frame generation time of the personalized scheme compared to the optimal scheme, in milliseconds, predicted by a nonlinear regression model), resource usage volatility (defined as the instantaneous change in GPU resource utilization, in %, calculated using CUDA core utilization and memory bandwidth), and critical device area coverage loss rate (defined as the percentage decrease in device visibility of the personalized scheme compared to the optimal scheme, in %, calculated using the point cloud visibility matrix). The rendering latency increment is predicted by a pre-trained nonlinear regression model (based on multinomial regression, inputting point cloud data volume (number of points, in millions) and scheme parameters, outputting frame generation time). The point cloud data volume is obtained from a 3D Gaussian point cloud model (by point counting), and the scheme parameters are extracted from user input. Resource usage volatility is measured in real-time using GPU monitoring tools (such as NVIDIA Nsight), which collect CUDA core utilization (range 0-100%) and memory bandwidth (unit GB / s). The standard deviation within a 10-second window is calculated and normalized to 0-100%. Coverage loss rate is calculated using ray tracing to obtain the visibility probability matrix of key device areas under the personalized solution (based on point cloud triangular mesh, matrix elements are 0-1). This matrix is ​​compared with the matrix of the optimal solution, and the deviation is divided by a safety threshold (90%) to obtain the loss rate. Cost data is stored in memory (approximately 50KB), and the calculation time is approximately 0.2 seconds per frame.

[0050] For example, in a substation scenario, the confirmation time is at 130 seconds, and the first remaining process is from 130 to 150 seconds. The personalized solution (location (10,5,3) meters) has 5 million point cloud data points. The nonlinear regression model (trained on 1000 sets of point cloud data, R²=0.95) predicts a frame generation time of 20 milliseconds, an increase of 5 milliseconds compared to the optimal solution (15 milliseconds). Resource usage volatility is monitored via Nsight: CUDA core utilization fluctuates between 10-90% (standard deviation 25%), memory bandwidth fluctuates between 2-8 GB / s (standard deviation 2 GB / s), and the normalized volatility is 30%. Coverage loss rate is calculated using ray tracing to determine the visibility probability matrix of the transformer heatsink (average 0.80 over 900 frames). The optimal solution has a value of 0.85, resulting in a loss rate of (0.85-0.80) / 0.90=5.56%. Cost data (latency increment 5 milliseconds, volatility 30%, loss rate 5.56%) is stored in memory (40KB).

[0051] G. Develop the optimal display strategy for showcasing multimodal costs during the remaining process of embedding the system; the optimal display strategy should guide users to quickly choose their preferred personalized camera placement options. Step G specifically includes the following sub-steps: G1. A Markov decision process model is constructed using a deep reinforcement learning algorithm. The optimization objective is to minimize the user's decision time. The model dynamically plans the optimal display sequence and interaction form for multimodal costs. The Markov decision process model discretizes the time axis of the first residual process into decision step size. It takes the priority weight vector of multimodal costs as input and outputs the optimal display strategy, including the priority sorting of highlighted warning icons, the overlay level of the cost impact visualization heatmap, and the timing of voice guidance.

[0052] This step works by constructing a Markov Decision Process (MDP) model using Deep Reinforcement Learning (DRL) algorithm. The goal is to minimize the user's decision time (defined as the time it takes for the user to complete a trade-off, in seconds), dynamically planning the optimal display sequence and interaction format for multimodal costs. The MDP model discretizes the time axis of the first residual process (obtained from F1, 10-30 seconds) into decision step sizes (defined as 0.5 seconds, based on human-computer interaction response time). The state space represents multimodal costs (latency increment, volatility, loss rate, ranging from 0-100%), and the action space represents combinations of display formats (highlighted warning icons, heatmap overlay levels, voice guidance timing). The priority weight vector is defined as the importance allocation of multimodal costs (range 0-1, based on D3 perspective optimization of preference weights and device criticality, with a weight sum of 1), generated through a weighted average. The DRL algorithm uses DeepQ-Network (DQN), taking the state (cost value + weight vector) and actions as inputs, and outputting the Q-value (predicted decision time). The optimization objective is to minimize the cumulative decision time (defined by a reward function, where the reward is the negative decision time). The priority order of highlighted warning icons is defined as the display order of cost items (e.g., latency > volatility > loss rate), the heatmap overlay level is defined as the cost visualization transparency (0-1), and the voice guidance timing is defined as the time point when the voice prompt is triggered (in seconds). The model is pre-trained on a simulated interaction dataset (1000 user decision scenarios), the inference time is 0.1 seconds / step, and the output display strategy is stored in memory (approximately 20KB).

[0053] For example, in a substation scenario, the first remaining process lasts 130-150 seconds, with a decision step size of 0.5 seconds (40 steps). The multimodal costs are a latency increment of 5 milliseconds, volatility of 30%, and a loss rate of 5.56%. The priority weight vectors (based on D3) are 0.4, 0.3, and 0.2. The DQN model (3 hidden layers, 256 neurons) takes the costs and weights as input and outputs the policy display: highlighted warning icons are sorted in the order of latency > loss rate > volatility (latency icon is red and displayed at the top); the heatmap overlay levels are latency 0.8, loss rate 0.6, and volatility 0.4 (transparency decreasing sequentially); voice guidance occurs at 131 seconds ("Attention rendering latency increased by 5 milliseconds") and 132 seconds ("Coverage decreased by 5.56%"). In implementation, the model was trained on 5000 sets of simulated data, and the reward function optimization reduced the decision time from 15 seconds to 8 seconds. This example demonstrates the efficiency and user-friendliness of policy planning.

[0054] H. Based on the optimal presentation strategy, display the multimodal cost to the user during the first remaining process execution.

[0055] This step works by using the optimal display strategy generated by G1. Within the first remaining period (10-30 seconds), the rendering engine presents multimodal costs (latency increment, volatility, loss rate) to the user, guiding decision-making through intuitive visualization and interaction. The optimal display strategy includes prioritizing highlighted warning icons (defined as the display order of cost items, stored as an ordered list), overlaying layers of cost impact visualization heatmaps (defined as the transparency of cost values ​​on the virtual view image, ranging from 0-1), and voice guidance timing (defined as the time point when voice is triggered, in seconds). The display process is implemented through the UI module of the rendering engine: highlighted warning icons are overlaid on the interface as colored icons (red, yellow, and green represent high, medium, and low costs) (resolution 1920×1080, fixed position in the upper right corner); the heatmap is overlaid on the virtual view image with a semi-transparent color (red represents high cost), mapping transparency based on cost values; voice guidance plays prompts via audio APIs (such as TTS engines) (text length < 50 characters, volume 80 decibels). Display data is acquired from F1 and updated in real time (60 FPS per frame), stored in the frame buffer (approximately 1 MB / frame). The display intensity is dynamically adjusted based on the frequency of user interaction (acquired from logs, in times / second) (the transparency is reduced by 0.1 when the frequency is greater than 1 time / second) to avoid information overload. Processing time is approximately 0.05 seconds / frame.

[0056] For example, in a substation scenario, the first remaining process lasts 130-150 seconds. The optimal display strategy is as follows: the delay icon (red) has the highest priority; the heatmap transparency is 0.8 for delay, 0.6 for loss rate, and 0.4 for volatility; voice guidance is triggered at 131 and 132 seconds. The rendering engine displays a red delay icon (position 1700×100 pixels) on the virtual view image (1920×1080). The heatmap covers the transformer area (500×300 pixels) in red (delay, transparency 0.8), and covers other areas in yellow (loss rate, 0.6) and green (volatility, 0.4). The voice prompt "Rendering delay increased by 5 milliseconds" is played via TTS at 131 seconds. The user interaction frequency is 0.8 times / second, and the transparency remains unchanged. Display data is stored in the frame buffer (900MB, 900 frames), taking 0.04 seconds / frame, for a total time of 3.6 seconds.

[0057] I. When the user completes the selection of the personalized camera placement intention plan, the selected personalized camera placement intention plan is added to the rendering process to complete the second residual process after the rendering time.

[0058] The working principle of this step is as follows: After the user completes the selection of their personalized camera placement plan (defined as the user confirms or modifies the plan through the interface, generating final parameters: position, orientation, and focal length), the selected plan is integrated into the rendering pipeline. In the second remaining process (defined as the subsequent stage after rendering, typically 10-20 seconds), high-fidelity virtual perspective images are generated. The selection is triggered by an interface confirmation event (clicking the "Confirm" button), recording a timestamp and final parameters (position (x, y, z) meters, pitch angle, yaw angle, focal length in millimeters). The rendering pipeline, based on a 3D Gaussian point cloud model (millions of points, including position and texture), generates virtual perspective images (resolution such as 1920×1080) using a ray tracing algorithm. The second remaining process initiates an adaptive point cloud simplification algorithm (based on octree segmentation, simplification rate 0-50%), dynamically adjusting the point cloud detail level based on the plan parameters and GPU load (monitored via Nsight, memory usage <80%) (preserving millimeter-level precision for critical parts of the device, with an error <1 mm). The generation process outputs a deployment effect verification report, which includes the view coverage improvement rate (defined as the ratio of the coverage after trade-offs to the optimal solution, %) and the multimodal cost optimization ratio (defined as the cost reduction ratio, %), stored in memory (approximately 50KB). Processing time is approximately 1 second per frame, and report generation time is 0.5 seconds.

[0059] For example, in a substation scenario, the user confirms the choice at 150 seconds: location (10, 5, 2.8) meters, pitch angle 28 degrees, focal length 40 mm. The rendering pipeline loads a virtual view image (1920×1080) based on 5 million point cloud data. An adaptive point cloud simplification algorithm simplifies the non-critical area point cloud by 20% (preserving millimeter-level accuracy for the transformer heatsink), using 75% of the video memory. Ray tracing generates 900 frames at 0.8 seconds per frame, for a total of 720 seconds. The validation report shows a coverage improvement of 90% / 85% = 105.88%, and a cost optimization (latency reduced from 5 milliseconds to 4 milliseconds) of 20%. The report is stored in memory (40KB) and takes 0.4 seconds.

[0060] The aforementioned automatic camera deployment method driven by 3D Gaussian point clouds in substations significantly improves the intelligence and user-friendliness of camera deployment schemes through multi-step collaborative work. Its core advantage lies in combining high-precision 3D Gaussian point cloud models with real-time rendering technology. Through advanced algorithms such as interaction log analysis, temporal convolutional neural networks, and deep reinforcement learning, it accurately infers users' personalized deployment intentions, dynamically optimizes the deployment scheme, and quantifies multimodal costs (such as rendering latency, resource fluctuations, and coverage loss) in real time. This method not only adaptively adjusts time window thresholds to capture user behavior characteristics but also ensures accurate determination of user intentions through a dual verification mechanism, ultimately generating an efficient display strategy to guide users in making rapid decisions. In practical applications, this solution achieves improved coverage and reduced rendering latency in substation monitoring scenarios while maintaining millimeter-level accuracy and low resource consumption, significantly improving deployment efficiency, system stability, and user interaction experience. It is suitable for the intelligent monitoring needs of complex industrial scenarios.

[0061] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for automatic camera placement driven by 3D Gaussian point clouds in substations, characterized in that, include: Based on a high-precision 3D Gaussian point cloud model of a substation, AI visual model is used to automatically identify key parts of the equipment and their three-dimensional spatial characteristics. Based on a pre-set knowledge base of placement rules, and according to the key parts of the equipment and their three-dimensional spatial characteristics, spatial geometric calculations and intelligent optimization decisions are performed to obtain the optimal camera placement scheme. Using a high-precision 3D Gaussian point cloud model, high-fidelity virtual view images are generated in real time based on the parameters in the optimal camera placement scheme to verify the placement effect.

2. The automatic camera placement method driven by 3D Gaussian point cloud in substations as described in claim 1, characterized in that, The three-dimensional spatial characteristics include: spatial coordinates, dimensions, and orientation.

3. The automatic camera placement method driven by 3D Gaussian point cloud in substations as described in claim 1, characterized in that, The deployment rule knowledge base includes: coverage rules, obstruction avoidance rules, key area focus rules, and installation feasibility rules, which are formed by digitizing a large number of power security standards.

4. The automatic camera placement method driven by 3D Gaussian point cloud in substations as described in claim 1, characterized in that, The optimal camera placement scheme includes: camera model, installation location, orientation angle, and focal length parameters.

5. The automatic camera placement method driven by 3D Gaussian point cloud in substations as described in claim 1, characterized in that, The process involves using a high-precision 3D Gaussian point cloud model and rendering it in real time based on parameters from the optimal camera placement scheme to generate high-fidelity virtual viewpoint images, including: By calling the 3D Gaussian point cloud data and simulating the resolution, focal length, and viewing angle parameters of a real camera according to the parameters in the optimal camera placement scheme, a high-quality virtual image with pixel-level precision and consistent with the imaging effect of a real camera is generated, thus obtaining a high-fidelity virtual perspective image.

6. The automatic camera placement method driven by 3D Gaussian point cloud in substations as described in claim 1, characterized in that, Also includes: When using a high-precision 3D Gaussian point cloud model to render in real time based on the parameters in the optimal camera placement scheme, if the user inputs a personalized camera placement intention scheme, the reason path for the user's input of the personalized camera placement intention scheme can be inferred based on the information displayed to the user during the rendering process within the time window before the user inputs the personalized camera placement intention scheme. Based on the underlying reasons, determine the strategy to confirm users' strong interest in personalized camera deployment solutions, and implement the corresponding confirmation. If the verification is successful, the multimodal cost of accepting and responding to the personalized camera placement intention plan is determined based on the first remaining process after the verification moment in the rendering process. The optimal display strategy for showcasing multimodal costs during the remaining process of embedding the plan; among which, the optimal display strategy can guide users to make a choice about personalized camera placement options in the fastest way; Based on the optimal presentation strategy, the multimodal cost is displayed to the user during the first remaining process execution; When the user completes the selection of their preferred personalized camera placement scheme, the selected scheme is added to the rendering process, completing the second remaining process after the rendering time.

7. The automatic camera placement method driven by 3D Gaussian point cloud in substations as described in claim 6, characterized in that, The reasoning path for inferring the user's input of a personalized camera placement intention includes: Through the interaction log analysis module of the real-time rendering engine, the triggering time of the user's input of the personalized camera placement intention plan is accurately captured, and the time window threshold is adaptively determined based on the dynamic change rate of the user's interaction history data. The time window threshold is dynamically adjusted by calculating the weighted entropy value of the user's gaze density, view zoom operation frequency and dwell time variance of the key equipment area during the pre-rendering process. The weight coefficient of the weighted entropy value is dynamically configured according to the preset mapping table of substation equipment type complexity. Based on the time window threshold, extract all multimodal information displayed to the user during the rendering process within the time window before the trigger time, including the device area heat map of the high-fidelity virtual view image, the dynamic curve of the view coverage index, and the spatial geometric conflict warning mark. Semantic parsing of multimodal information is performed using a temporal convolutional neural network model to generate a reason path for the user's input of personalized camera placement intention scheme. The reason path is represented as a user intent tracing map, which includes the focusing trajectory of key parts of the device, the spatial occlusion avoidance motivation sequence, and the viewpoint optimization preference weight distribution. The viewpoint optimization preference weight distribution is dynamically quantified by analyzing the correlation coefficient between the user's interaction intensity with different device areas and the rendering frame rate fluctuation.

8. The automatic camera placement method driven by 3D Gaussian point cloud in substations as described in claim 6, characterized in that, The strategy for confirming strong user interest in personalized camera deployment schemes includes: Real-time monitoring of users' subsequent operational behavior after inputting personalized camera placement intentions, and determination of intention strength through a dual verification mechanism; The dual verification mechanism includes: The first verification involves calculating the cosine similarity between the direction vector of subsequent parameter adjustment operations and the camera pose parameters in the personalized camera placement intention scheme. When the cosine similarity is consistently higher than the preset similarity threshold and the operation frequency exceeds the preset frequency threshold, preliminary confirmation is triggered. The second layer of verification involves analyzing the frequency of intent keywords in the user's voice or text input using a natural language processing model, and then cross-validating this with the click hotspot matching degree of the highlighted area in the rendered interface. When the joint confidence of keyword frequency and hotspot matching degree exceeds the threshold confidence threshold, the user's intent is determined to be strong. If the verification fails, it will automatically revert to the optimal camera placement plan and generate optimization suggestions.

9. The automatic camera placement method driven by 3D Gaussian point cloud in substations as described in claim 6, characterized in that, The multimodal cost of determining and responding to personalized camera deployment intention schemes includes: During the first residual process after the rendering time is confirmed, multimodal costs are quantified in real time, including real-time rendering latency increment, resource usage volatility, and critical device area coverage loss rate. The real-time rendering latency increment is predicted based on a nonlinear regression model of GPU frame generation time and point cloud data volume. The resource usage volatility is calculated by monitoring the instantaneous changes in CUDA core utilization and memory bandwidth. The critical device area coverage loss rate is measured based on the deviation between the visibility probability matrix of critical parts of the device in the 3D Gaussian point cloud model and a preset safety threshold.

10. The automatic camera placement method driven by 3D Gaussian point cloud in substations as described in claim 6, characterized in that, The optimal presentation strategy for displaying multimodal costs within the embedded remaining process of the planning includes: A Markov decision process model is constructed using a deep reinforcement learning algorithm. With the goal of minimizing user decision time, the model dynamically plans the optimal display sequence and interaction form for multimodal costs. The Markov decision process model discretizes the time axis of the first residual process into decision step sizes, takes the priority weight vector of multimodal costs as input, and outputs the optimal display strategy, including priority sorting of highlighted warning icons, overlaying layers of cost impact visualization heatmaps, and timing of voice guidance.