View generation method and device of transformer substation and computer program product
By generating semantic label maps and 3D Gaussian point clouds of substations, and combining them with augmented reality equipment processing, the problem of unclear views caused by incomplete substation point cloud data was solved, achieving efficient 3D modeling and a clear visualization experience.
Patent Information
- Application Number
- CN202511037171.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies for 3D information acquisition in substations suffer from incomplete point cloud data quality, resulting in unclear views. In particular, dense point cloud computing consumes high cloud computing resources while sparse point cloud lacks details and has limited semantic segmentation capabilities.
By acquiring color images and text descriptions of the substation, a semantic tag map is generated. This map is then combined with 3D Gaussian point clouds for semantic rendering and adjustment. Finally, augmented reality devices are used to process the target Gaussian point cloud set, resulting in an augmented reality view of the substation.
It improves the modeling accuracy and semantic annotation efficiency of substation 3D models, reduces the workload in the initial modeling stage, achieves seamless integration of virtual models and real views, and enhances the visualization experience and interactivity of substations.
Smart Images

Figure CN120953490A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the metaverse domain, and more specifically, to a method, apparatus, and computer program product for generating views of a substation. Background Technology
[0002] Against the backdrop of power system transformation and upgrading, the accurate acquisition and modeling of 3D information for substations has become particularly important. High-quality 3D models can not only comprehensively demonstrate the layout of substation equipment and lines, but also provide strong data support for remote operation and maintenance, fault detection, and intelligent inspection. However, substation 3D information acquisition technology faces a series of challenges, especially in terms of point cloud data quality and processing efficiency.
[0003] With the power industry's push for digital transformation, substation 3D information acquisition technology is constantly evolving. Among related technologies, 3D point cloud reconstruction based on multi-view images and LiDAR is used to generate sparse or dense 3D point clouds. Then, deep learning technology is used for semantic segmentation of the point clouds, labeling the equipment category to which each point belongs, thereby constructing a 3D scene model with semantic information. However, in the modeling methods of related technologies, dense point clouds can provide higher modeling accuracy during reconstruction, but due to the massive amount of data and computational requirements, real-time processing is difficult, limiting its application in dynamic scenes. On the other hand, while sparse point clouds reduce the computational burden, the lack of detail prevents fine semantic segmentation, thus limiting the model's semantic expressive power.
[0004] The related technologies employ a step-by-step approach to reconstruction and semantic segmentation, performing geometric reconstruction first and then semantic annotation. This two-stage method lacks an effective joint optimization mechanism between reconstruction and segmentation. As a result, semantic boundaries are unclear, misclassification frequently occurs during category recognition, and the accuracy and practicality of the 3D model are reduced.
[0005] There is currently no effective solution to the problem of unclear substation views caused by incomplete point cloud data collection in related technologies. Summary of the Invention
[0006] The main objective of this application is to provide a method, apparatus, and computer program product for generating substation views, in order to solve the problem in related technologies where the substation views are unclear due to incomplete point cloud data collection.
[0007] To achieve the above objectives, according to one aspect of this application, a method for generating a view of a substation is provided. The method includes: acquiring a color image of the substation; determining textual description information of the substation, wherein the textual description information includes description information of N devices in the substation, where N is a positive integer; determining a semantic label map of the substation based on the color image and the textual description information, wherein the semantic label map includes a category label of the device to which each pixel of the color image belongs; determining a three-dimensional Gaussian point cloud of the substation based on the color image; performing semantic rendering on the three-dimensional Gaussian point cloud based on the semantic label map to obtain a semantic rendering map; adjusting the three-dimensional Gaussian point cloud based on the semantic rendering map to obtain a target Gaussian point cloud set; and processing the target Gaussian point cloud set using an augmented reality device to obtain an augmented reality view of the substation.
[0008] Optionally, determining the semantic label map of the substation based on color images and text description information includes: inputting the color images and text description information into an object detection model to obtain bounding boxes of N objects and category labels to which the bounding boxes belong, wherein each object represents a device; inputting the color images and bounding boxes of the N objects into an image segmentation model to obtain a mask for each object, and adding the category label of the corresponding device to the mask of each object, wherein the mask is used to represent the outline of the object; and merging and overlaying the masks of the N objects to obtain the semantic label map.
[0009] Optionally, determining the 3D Gaussian point cloud of the substation based on the color image includes: inputting the color image into 3D reconstruction software to obtain a 3D sparse point cloud set and camera pose; initializing each point cloud in the 3D sparse point cloud set as a Gaussian ellipsoid to obtain a 3D Gaussian point cloud, wherein the Gaussian ellipsoid contains at least one of the following 3D spatial information: position, covariance matrix, opacity, color, and semantic feature vector, wherein the semantic feature vector is the feature vector corresponding to the category label.
[0010] Optionally, before semantic rendering of the 3D Gaussian point cloud based on the semantic label map, the method further includes: determining multiple observation rays based on the camera pose; for each observation ray, determining the depth of the Gaussian ellipsoid intersecting the observation ray, where the depth is the distance between the center of the Gaussian ellipsoid and the camera; sorting the Gaussian ellipsoids intersecting the observation rays in descending order of depth to obtain a list of Gaussian ellipsoids for the observation rays; and performing layer-by-layer blending of the Gaussian ellipsoids in the list of Gaussian ellipsoids for each observation ray using a preset transparency processing method to obtain a color-rendered image of the 3D Gaussian point cloud after color rendering.
[0011] Optionally, semantic rendering of the 3D Gaussian point cloud based on the semantic label map to obtain the semantic rendering map includes: determining the semantic probability of the Gaussian ellipsoid traversed by each observation ray, wherein the semantic probability is the probability that the Gaussian ellipsoid belongs to each category label; for each pixel position in the 3D Gaussian point cloud, determining the target observation ray passing through the pixel position, determining the category probability vector of the pixel position based on the semantic probability of the Gaussian ellipsoid intersecting the target observation ray, wherein the category probability vector contains the probability that the pixel position belongs to each category label; and determining the semantic rendering map based on the category probability vectors of all pixel positions.
[0012] Optionally, adjusting the 3D Gaussian point cloud based on the semantic rendering map to obtain the target Gaussian point cloud set includes: comparing the color image with the color rendering image to obtain a color loss term; comparing the semantic rendering map with the semantic label map to obtain a semantic loss term; determining a target loss function based on the color loss term and the semantic loss term; adjusting the features of the Gaussian ellipsoid in the 3D Gaussian point cloud with the goal of minimizing the loss value of the target loss function to obtain the target Gaussian point cloud set, wherein the features include at least one of the following: geometric, appearance, and semantic feature vectors, and the appearance includes color and opacity.
[0013] Optionally, processing the target Gaussian point cloud set using an augmented reality device to obtain an augmented reality view of the substation includes: converting the coordinate system of the target Gaussian point cloud set to the target coordinate system of the augmented reality device; performing perspective mapping on each point cloud in the target Gaussian point cloud set under the target coordinate system to obtain the two-dimensional image coordinates of each point cloud; determining the point cloud set corresponding to each category label in the target Gaussian point cloud set; setting the rendering parameters for the point cloud set corresponding to each category label; rendering the point cloud set corresponding to each category label using the rendering parameters to obtain a rendered image, wherein the rendering parameters include at least one of the following: color and opacity; blending the rendered image with the real-time acquired image of the augmented reality device according to the opacity to obtain an augmented reality view; and displaying the augmented reality view on the augmented reality device according to the two-dimensional image coordinates of each point cloud.
[0014] To achieve the above objectives, according to another aspect of this application, a substation view generation apparatus is provided. The apparatus includes: an acquisition unit, configured to acquire a color image of the substation and determine textual description information of the substation, wherein the textual description information includes description information of N devices in the substation, where N is a positive integer; a first determination unit, configured to determine a semantic label map of the substation based on the color image and the textual description information, wherein the semantic label map includes a category label of the device to which each pixel of the color image belongs; a rendering unit, configured to determine a three-dimensional Gaussian point cloud of the substation based on the color image, and perform semantic rendering on the three-dimensional Gaussian point cloud based on the semantic label map to obtain a semantic rendering map; an adjustment unit, configured to adjust the three-dimensional Gaussian point cloud based on the semantic rendering map to obtain a target Gaussian point cloud set; and a processing unit, configured to process the target Gaussian point cloud set using an augmented reality device to obtain an augmented reality view of the substation.
[0015] To achieve the above objectives, according to another aspect of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the substation view generation method described in various embodiments of this application.
[0016] This application employs the following steps: acquiring a color image of a substation; determining the text description information of the substation, where the text description information includes descriptions of N devices in the substation, where N is a positive integer; determining a semantic label map of the substation based on the color image and text description information, where the semantic label map includes the category label of the device to which each pixel in the color image belongs; determining a 3D Gaussian point cloud of the substation based on the color image; performing semantic rendering on the 3D Gaussian point cloud based on the semantic label map to obtain a semantic rendering map; adjusting the 3D Gaussian point cloud based on the semantic rendering map to obtain a target Gaussian point cloud set; and processing the target Gaussian point cloud set using an augmented reality device to obtain an augmented reality view of the substation. This solves the problem in related technologies where incomplete point cloud data acquisition leads to unclear substation views. By inputting the text description information and color image of the devices, the semantic label map is automatically detected and generated, eliminating the need for predefined categories or manual annotation, significantly reducing the workload in the initial modeling stage and improving the efficiency and accuracy of semantic annotation. By combining the semantic information of the color image and text description, the modeling of a 3D substation model based on a 3D Gaussian point cloud is realized. This approach improves modeling accuracy and overcomes the high computational resource consumption problem of dense point clouds. By processing the target Gaussian point cloud set using augmented reality devices, it achieves a seamless integration of the virtual model and the real view. This enables maintenance personnel to intuitively and clearly identify substation equipment, enhancing the visualization experience and interactivity of the substation, thereby improving the clarity of the substation view. Attached Figure Description
[0017] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0018] Figure 1 This is a flowchart of a substation view generation method provided according to an embodiment of this application;
[0019] Figure 2 This is a schematic diagram of an optional substation view generation method provided according to an embodiment of this application;
[0020] Figure 3 This is a schematic diagram of a substation view generation device provided according to an embodiment of this application;
[0021] Figure 4 This is a schematic diagram of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0022] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0025] The present invention will now be described in conjunction with preferred implementation steps. Figure 1 This is a flowchart of a substation view generation method provided according to an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:
[0026] Step S101: Obtain a color image of the substation and determine the text description information of the substation. The text description information contains description information of N devices in the substation, where N is a positive integer.
[0027] In step S101, an RGB (Red Green Blue camera, used to capture information from the red, green, and blue color channels) camera is used to take multi-angle photos around the substation to obtain multi-view RGB images, i.e., color images. A set of multi-view RGB images of the substation can be used to... i The color images contain rich details of the substation environment. Textual descriptions of equipment within the substation are collected and determined, including but not limited to equipment names, types, locations, and functional descriptions. This information guides automated semantic annotation. Equipment may include transformers, circuit breakers, disconnect switches, relay protection devices, reactive power compensation devices, insulators, and conductors, etc.
[0028] Step S102: Determine the semantic label map of the substation based on the color image and text description information, wherein the semantic label map contains the category label of the device to which each pixel of the color image belongs.
[0029] In step S102, the color image and text description information are input into the object detection model, which can locate and identify objects in the image based on natural language descriptions. The object detection model understands the text description information to detect the bounding boxes of each device in the color image. These bounding boxes, along with the color image, are then used as prompts to input into the image segmentation model to generate pixel-level masks for each device. These masks are then superimposed onto the original color image to form a semantic label map that corresponds one-to-one with the substation equipment type. Each pixel is assigned a category label for the equipment type it represents.
[0030] Step S103: Determine the three-dimensional Gaussian point cloud of the substation based on the color image, and perform semantic rendering on the three-dimensional Gaussian point cloud based on the semantic label map to obtain a semantic rendering map.
[0031] In step S103, sparse point clouds and camera poses of the substation scene are reconstructed from multi-view color images using 3D reconstruction software. Each point in the sparse point cloud is initialized as a Gaussian ellipsoid, and its geometric properties such as position and covariance matrix, appearance properties such as color and opacity, and semantic feature vectors are defined. The dimension of the semantic feature vectors is consistent with the number of device categories. The Gaussian ellipsoids are sorted by depth according to the camera poses, and color rendering maps and semantic rendering maps are calculated using alpha blending (a transparency processing technique).
[0032] Step S104: Adjust the 3D Gaussian point cloud based on the semantic rendering graph to obtain the target Gaussian point cloud set.
[0033] In step S104, the Gaussian ellipsoid parameters are adjusted using the stochastic gradient descent algorithm, combining the color reprojection error (L2 loss) and the semantic segmentation error (cross-entropy loss), and the Gaussian point cloud .ply file with semantic labels is output, which is the target Gaussian point cloud set.
[0034] Step S105: Process the target Gaussian point cloud set using augmented reality equipment to obtain an augmented reality view of the substation.
[0035] In step S105, the intrinsic and extrinsic parameters of the AR (Augmented Reality) device are calibrated using the SLAM (Simultaneous Localization and Mapping) module, mapping the Gaussian point cloud from the world coordinate system to the AR camera coordinate system. The point cloud is then filtered by device category, with differentiated color and opacity settings, and blended with the real-time camera image from the AR device to obtain an augmented reality view, achieving high-confidence target highlighting. Through multimodal data fusion and optimization, semantic 3D reconstruction and AR interactive visualization of substation equipment are realized, applicable to power inspection and equipment management scenarios.
[0036] The substation view generation method provided in this application involves acquiring a color image of the substation, determining its text description information (where N is a positive integer), and then determining a semantic label map based on the color image and text description information. This semantic label map includes the category label of each pixel in the color image. A 3D Gaussian point cloud of the substation is then determined based on the color image. Semantic rendering of the 3D Gaussian point cloud is performed based on the semantic label map to obtain a semantic rendering map. The 3D Gaussian point cloud is adjusted based on the semantic rendering map to obtain a target Gaussian point cloud set. Finally, the target Gaussian point cloud set is processed using an augmented reality device to obtain an augmented reality view of the substation. This method solves the problem of unclear substation views due to incomplete point cloud data acquisition in related technologies. By inputting the text description information and color image of the devices, the method automatically detects and generates semantic label maps without predefined categories or manual annotation, significantly reducing the workload in the initial modeling stage and improving the efficiency and accuracy of semantic annotation. By combining the semantic information of the color image and text description, the method achieves the modeling of a 3D substation model based on a 3D Gaussian point cloud. This approach improves modeling accuracy and overcomes the high computational resource consumption problem of dense point clouds. By processing the target Gaussian point cloud set using augmented reality devices, it achieves a seamless integration of the virtual model and the real view. This enables maintenance personnel to intuitively and clearly identify substation equipment, enhancing the visualization experience and interactivity of the substation, thereby improving the clarity of the substation view.
[0037] To improve the efficiency and accuracy of semantic annotation, object detection models and image segmentation models are used to generate semantic label maps. Optionally, in the substation view generation method provided in this application embodiment, determining the semantic label map of the substation based on color images and text description information includes: inputting color images and text description information into an object detection model to obtain bounding boxes of N objects and category labels to which the bounding boxes belong, wherein each object represents a device; inputting color images and bounding boxes of N objects into an image segmentation model to obtain a mask for each object, and adding the category label of the corresponding device to the mask of each object, wherein the mask is used to represent the outline of the object; merging and superimposing the masks of N objects to obtain a semantic label map.
[0038] In some embodiments, in order to obtain a semantic label map, the acquired RGB image {I} is input into the object detection model. iGiven a set of text descriptions, such as transformers, circuit breakers, disconnectors, relay protection devices, reactive power compensation devices, insulators, and conductors, the object detection model outputs bounding boxes and corresponding category labels for each object in the image. The bounding boxes output by the object detection model are used as prompts and input to the image segmentation model. The image segmentation model extracts a fine pixel-level mask for each object based on these prompts, thus segmenting the precise contour of each object. Each mask region obtained by the image segmentation model is assigned a corresponding category label, such as transformer 1, circuit breaker 2, disconnector 3, relay protection device 4, reactive power compensation device 5, insulator 6, conductor 7, etc., with k categories. Unlabeled areas are defaulted to background 0. All masks are merged and superimposed to construct a pixel-level semantic label map {S} with the same size as the original image. i}
[0039] This embodiment employs automated semantic annotation using object detection and image segmentation models. Users only need to input a text description of the device name to automatically output the corresponding object's bounding box. No predefined category labels or pre-trained detection models are required, resulting in greater versatility and easier application to various substation scenarios. This solves the problem of semantic annotation relying on manual annotation and prior information, significantly reducing the workload in the initial modeling stage and improving the efficiency and accuracy of semantic annotation.
[0040] After obtaining the color image, the three-dimensional Gaussian point cloud of the substation is determined by three-dimensional reconstruction software. Optionally, in the substation view generation method provided in this application embodiment, determining the three-dimensional Gaussian point cloud of the substation based on the color image includes: inputting the color image into the three-dimensional reconstruction software to obtain a three-dimensional sparse point cloud set and camera pose; initializing each point cloud in the three-dimensional sparse point cloud set as a Gaussian ellipsoid to obtain a three-dimensional Gaussian point cloud, wherein the Gaussian ellipsoid contains at least one of the following three-dimensional spatial information: position, covariance matrix, opacity, color, and semantic feature vector, and the semantic feature vector is the feature vector corresponding to the category label.
[0041] In some embodiments, by acquiring a set of multi-view RGB images of a substation {I i Input into the 3D reconstruction software, output a 3D sparse point cloud set P = {P} of the substation scene. j The camera pose T and the 3D sparse point cloud P provide an initial estimate of the substation scene geometry. Each point cloud P in the 3D sparse point cloud P is then... j Initialize it as a Gaussian ellipsoid, the geometric properties of which are defined by a Gaussian function G. j (x) represents each Gaussian function G j (x) is defined by the following parameters:
[0042]
[0043] The position of the Gaussian ellipsoid is represented by μ. j It means that P j The coordinates (x, y, z); the covariance matrix is represented by Σ. j This indicates that the shape of the ellipsoid is controlled. In addition to geometric properties, opacity α is also used. j and color c j To represent the appearance attributes of the Gaussian ellipsoid. In addition, semantic attributes are added, assigning a semantic feature vector f to each Gaussian ellipsoid. j Based on the number k of equipment category labels in the substation, determine the semantic feature vector f. j The dimension is k, and it is randomly initialized. The position is μ. j Covariance matrix Σ j Opacity α j Color c j Semantic feature vector f j The parameters represent the three-dimensional Gaussian ellipsoid.
[0044] This embodiment employs 3D Gaussian splashing technology to reduce computational resource consumption while supporting real-time AR interaction. Semantic attributes, geometric attributes, and appearance attributes are jointly embedded in the 3D Gaussian point cloud, achieving unified optimization of geometric reconstruction and semantic understanding, and realizing accurate matching of semantic tags with the substation's geometric structure.
[0045] After obtaining the 3D Gaussian point cloud, color rendering is performed on the 3D Gaussian point cloud. Optionally, in the substation view generation method provided in this application embodiment, before performing semantic rendering on the 3D Gaussian point cloud based on the semantic tag map, the method further includes: determining multiple observation rays based on the camera pose; for each observation ray, determining the depth of the Gaussian ellipsoid intersecting with the observation ray, wherein the depth is the distance between the center of the Gaussian ellipsoid and the camera; sorting the Gaussian ellipsoids intersecting with the observation rays in descending order of depth to obtain a list of Gaussian ellipsoids for the observation rays; and performing layer-by-layer blending of the Gaussian ellipsoids in the list of Gaussian ellipsoids for each observation ray using a preset transparency processing method to obtain a color-rendered image of the 3D Gaussian point cloud after color rendering.
[0046] In some embodiments, multiple observation rays r are determined based on the camera pose T obtained by the 3D reconstruction software. For each observation ray, the intersecting Gaussian ellipsoids along the ray r are sorted by depth, and alpha blending (i.e., a preset transparency processing method) is used to blend all Gaussian ellipsoids passed through by each camera ray r layer by layer to obtain a color rendering image.
[0047] The rendering formula for color-rendered images is as follows:
[0048] C(r)=∑ j∈N(r) (cj ·α j ·T j );
[0049] Where N(r) is a list of Gaussian ellipsoids ordered by depth on the observed ray r; c j For the Gaussian ellipsoid G j Color in the direction of the ray, α j For the Gaussian ellipsoid G j Opacity; T j =∏ i<j (1-α i The cumulative transmittance represents the probability of reaching the i-th Gaussian ellipsoid without being occluded. The rendered color image is obtained by calculating the colors on all rays.
[0050] This embodiment utilizes observation ray generation based on camera pose to dynamically render 3D Gaussian point clouds from different perspectives, providing users with an immersive AR experience. Depth sorting of the Gaussian ellipsoid ensures the correct rendering order of foreground and background devices, avoiding information confusion caused by improper object occlusion and improving the realism of the visualization. The alpha blending method allows for the natural presentation of semi-transparent or transparent devices without affecting or minimizing occlusion of underlying devices. The combination of layer-by-layer blending and depth sorting significantly reduces rendering computation, enabling the rapid generation of high-quality rendered images to meet the needs of real-time applications. The final color-rendered image not only preserves the device's geometry but also enhances its recognizability through color rendering. Even devices with similar structures can be clearly distinguished by color differences, improving the efficiency of inspection and management.
[0051] After color rendering of the 3D Gaussian point cloud, semantic rendering of the 3D Gaussian point cloud is also required. Optionally, in the substation view generation method provided in this application embodiment, semantic rendering of the 3D Gaussian point cloud based on the semantic label map to obtain the semantic rendering map includes: determining the semantic probability of the Gaussian ellipsoid traversed by each observation ray, wherein the semantic probability is the probability that the Gaussian ellipsoid belongs to each category label; for each pixel position in the 3D Gaussian point cloud, determining the target observation ray passing through the pixel position, determining the category probability vector of the pixel position based on the semantic probability of the Gaussian ellipsoid intersecting the target observation ray, wherein the category probability vector contains the probability that the pixel position belongs to each category label; and determining the semantic rendering map based on the category probability vectors of all pixel positions.
[0052] In some embodiments, each Gaussian ellipsoid carries a semantic feature vector with dimensions consistent with the number of device categories. Each element represents the membership or confidence level of the Gaussian ellipsoid to the corresponding device category. For each observation ray, its intersection with the Gaussian ellipsoid is detected. The semantic feature vectors of the intersecting Gaussian ellipsoids are converted into probabilities using a softmax function, forming the semantic probability of that Gaussian ellipsoid. For example, the semantic rendering formula is shown below:
[0053]
[0054] Where P(s=k|r) is the probability that the final rendering result of the target observation ray r is of category k; For the Gaussian ellipsoid G j The probability of belonging to category k is the semantic feature vector f. j The result calculated using the softmax function can be expressed by the following formula:
[0055]
[0056] For each pixel location (u,v) in the image, emit a target observation ray r, and calculate the class probability vector for each pixel location. This yields the probability map of the k channels of the entire image. That is, semantic rendering graph.
[0057] This embodiment employs probabilistic semantic rendering and integrates semantic confidence weights in alpha blending to achieve end-to-end alignment between the rendering results and semantic labels. By combining semantic feature vectors with color and depth information, it achieves accurate semantic rendering in complex scenes, improving the robustness and accuracy of device recognition.
[0058] To improve the accuracy of the generated view, after color rendering and semantic rendering, it is also necessary to adjust the features of the Gaussian ellipsoid in the 3D Gaussian point cloud. Optionally, in the substation view generation method provided in this application embodiment, adjusting the 3D Gaussian point cloud based on the semantic rendering map to obtain the target Gaussian point cloud set includes: comparing the color image with the color rendering image to obtain a color loss term; comparing the semantic rendering map with the semantic label map to obtain a semantic loss term; determining a target loss function based on the color loss term and the semantic loss term; adjusting the features of the Gaussian ellipsoid in the 3D Gaussian point cloud with the goal of minimizing the loss value of the target loss function to obtain the target Gaussian point cloud set, wherein the features include at least one of the following: geometric, appearance, and semantic feature vectors, and the appearance includes color and opacity.
[0059] In some embodiments, the color-rendered image is compared with the captured color image, and the color loss is calculated, with the color loss term L. rgbThe calculation formula is as follows:
[0060]
[0061] By comparing the semantic rendering graph and the semantic label graph, a semantic cross-entropy loss is constructed, with the semantic loss term L. sem The calculation formula is as follows:
[0062]
[0063] Combining color error and semantic error, the target loss function is as follows:
[0064] L = L rgb +(1-λ)·L sem ;
[0065] Where λ represents the weighting of color reprojection error and semantic segmentation error, and can be set to 0.5. Stochastic gradient descent is used to minimize the joint loss function, while simultaneously adjusting the geometric, appearance, and semantic feature vectors of the Gaussian ellipsoid. After optimization, each Gaussian ellipsoid G... j Encoded to include geometric attributes (position μ) j Covariance matrix Σ j Appearance attributes (color c) j Opacity α j Structured data including semantic labels (category numbers). All Gaussian ellipsoids constitute the final target Gaussian point cloud set G. sem All data is saved as .ply files in a unified data format that supports subsequent retrieval and rendering.
[0066] The training strategy proposed in this embodiment, which jointly optimizes semantic and geometric features, avoids the decoupling problem between semantic and geometric features and improves the semantic understanding capability of substations. The 3D Gaussian point cloud model not only accurately matches the equipment in terms of geometric shape, but also achieves a high degree of consistency in color and semantic information, effectively improving the practicality, accuracy, and user experience of the substation AR visualization system.
[0067] After obtaining the target Gaussian point cloud set, the target Gaussian point cloud set is converted into an augmented reality view. Optionally, in the substation view generation method provided in this application embodiment, the substation augmented reality view is obtained by processing the target Gaussian point cloud set using an augmented reality device, including: converting the coordinate system of the target Gaussian point cloud set into the target coordinate system of the augmented reality device; performing perspective mapping on each point cloud in the target Gaussian point cloud set under the target coordinate system to obtain the two-dimensional image coordinates of each point cloud; determining the point cloud set corresponding to each category label in the target Gaussian point cloud set; setting the rendering parameters of the point cloud set corresponding to each category label; rendering the point cloud set corresponding to each category label using the rendering parameters to obtain a rendered image, wherein the rendering parameters include at least one of the following: color and opacity; mixing the rendered image with the real-time acquired image of the augmented reality device according to the opacity to obtain an augmented reality view; and displaying the augmented reality view on the augmented reality device according to the two-dimensional image coordinates of each point cloud.
[0068] In some embodiments, after completing the target Gaussian point cloud set G sem After optimization and storage, to overlay and interactively display the geometric and semantic information of substation equipment in an AR environment in real time, it is necessary to implement steps such as coordinate system transformation and camera calibration, perspective projection and pixel mapping, and semantic information-based rendering and visualization. The target coordinate system can be the camera coordinate system of the AR glasses. First, the extrinsic parameters (R) of the AR device are obtained through the AR device's built-in SLAM calibration module. AR ,t AR The intrinsic parameter matrix contains (f) x ,f y ,c x ,c y Secondly, the target Gaussian point cloud set G sem The transformation from the world coordinate system to the camera coordinate system of the AR glasses is shown below:
[0069] X cam =R AR X world +t AR ;
[0070] Among them, X cam For the target Gaussian point cloud set G sem The center position in the coordinate system of the AR device; X world For the target Gaussian point cloud set G sem At the center of the world coordinate system μ j ;R AR The rotation matrix is the extrinsic parameter of the AR device; t AR Let G be the translation vector of the extrinsic parameters of the AR device. For each Gaussian point cloud G in the AR device's camera coordinate system... semPerform perspective projection and map it to 2D image pixel coordinates (u,v), the formula is as follows:
[0071]
[0072] Among them, (f x ,f y (c) represents the camera focal length of the AR device, obtained from the AR device parameters. x ,c y The drift of the main point is obtained through camera calibration on an AR device, and (x, y, z) is the calculated X. cam The position in, i.e., X cam = (x, y, z). To fully utilize the semantic Gaussian point cloud G sem The system carries information on substation equipment categories, enabling targeted visual interaction for specific objects. First, semantic filtering and classification are performed, based on the Gaussian ellipsoid G... j The probability of belonging to category k Determine its main semantic tags. Let k1 be a target category (e.g., "transformer"), then the set of Gaussian points corresponding to this category is as follows:
[0073]
[0074] Secondly, the rendering parameters are set for the filtered Gaussian point set. Differentiate its rendering properties such as color c j Opacity α j For color, a highlight color c is preset for each semantic category k. k Then, based on the confidence level p of that point for that category... j (k) Mix the primary color and the highlight color, as shown in the formula below:
[0075]
[0076] Among them, c j ′ The color is the result of mixing the primary color and the highlight color. c is the primary color for this category of devices. k This is the highlighted color set. For opacity, it is linearly mapped to the interval [α] based on confidence level. min ,α max ]:
[0077] α j =α j +(α max -α min )p j (k);
[0078] Where, αj α is the current opacity value. max The maximum value of the mapped interval is preset by the user, α. min The minimum value of the mapping interval is preset by the user. Points with higher confidence levels are closer to the semantic highlight color and more opaque, while points with lower confidence levels are more inclined towards the primary color and semi-transparent, thus intuitively distinguishing different semantic elements. Finally, virtual object generation and fusion are performed, using rendering parameters and the AR device's built-in 3D rendering engine to... Projected into pixel space and fused with real-time camera images from the AR device. For any ray r, its pixel color is calculated as:
[0079] C(r)=∑ j∈N(r) (c j ′ ·α j ·T j );
[0080] T j =∏ j<i (1-α j );
[0081] Where N(r) is the sorted set of visible elements on the ray; C(r) is the color rendered by ray r; c j ′ The color resulting from mixing a primary color and a highlight color; α j The current opacity value calculated using the above formula; T j This is the cumulative transmittance. Finally, the rendered image C will be compared with the AR device's camera image C. cam Blend by opacity ω to obtain the image rendered in the AR view.
[0082] C AR (u,v)=ω(r)C(r)+(1-ω(r))C cam (u,v);
[0083] ω(r)=1-∏ j∈N(r) (1-α j );
[0084] Among them, C AR (u,v) represents the color at (u,v), ω(r) represents the final composite opacity of ray r; C(r) is obtained through the previous formula; C cam (u,v) is obtained based on the AR device. This allows for quick rendering of substation equipment in the AR view for specified semantic categories (such as transformers) and visualization of substation equipment by semantic highlighting.
[0085] This embodiment proposes an adaptive rendering parameter configuration based on semantic confidence to ensure stable semantic highlighting effects on AR devices. In the AR environment, Gaussian point sets can be dynamically filtered for any device category, and highlight color blending and linear enhancement can be applied to the devices based on confidence, realizing confidence-driven semantic highlighting visualization. This makes it possible to support real-time, interactive substation operation and maintenance assistance and inspection.
[0086] According to another embodiment of this application, an optional method for generating views of a substation is also provided. Figure 2 This is a schematic diagram of an optional substation view generation method provided according to an embodiment of this application. For example... Figure 2 As shown, the method includes four steps: data acquisition and preprocessing, 3D Gaussian point cloud initialization, 3D Gaussian rendering and optimization, and AR visualization.
[0087] Specifically, multi-view image acquisition of the substation is performed. Then, semantic labels are generated and sparse 3D point clouds are initialized from the acquired color images. The 3D Gaussian point cloud is initialized and then rendered in color and semantically. The 3D Gaussian point cloud is adjusted through loss calculation and joint optimization to obtain a target Gaussian point cloud set. The target Gaussian point cloud set is then subjected to coordinate system transformation and camera marking using an AR device. The point clouds in the target Gaussian point cloud set are then subjected to perspective projection and pixel mapping. Finally, the image mapped from the target Gaussian point cloud set is rendered and visualized based on semantic information to obtain a view of the substation.
[0088] This embodiment innovatively embeds semantic attributes along with traditional geometric and appearance attributes into a 3D Gaussian point cloud through an optional substation view generation method, achieving unified optimization of geometric reconstruction and semantic understanding. Simultaneously, based on automated semantic annotation, a semantic label map is generated through text prompts, solving the problem of traditional methods relying on manual annotation and prior information. Probabilistic semantic rendering is employed, and semantic confidence weights are integrated into alpha blending to achieve end-to-end alignment between rendering results and semantic labels. The proposed training strategy for joint optimization of semantic and geometric features avoids the decoupling problem between semantic and geometric features in traditional methods, improving the semantic understanding capability of substations. Through an AR visualization interaction enhancement method, Gaussian point sets can be dynamically selected for any equipment category in an AR environment. Highlight color blending and linear enhancement are applied to the equipment based on confidence, achieving confidence-driven semantic highlight visualization, making real-time, interactive substation operation and maintenance assistance and inspection possible.
[0089] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0090] This application also provides a substation view generation apparatus. It should be noted that the substation view generation apparatus of this application can be used to execute the substation view generation method provided in this application. The substation view generation apparatus provided in this application will be described below.
[0091] Figure 3 This is a schematic diagram of a substation view generation device provided according to an embodiment of this application. For example... Figure 3 As shown, the device includes:
[0092] The acquisition unit 301 is used to acquire a color image of the substation and determine the text description information of the substation, wherein the text description information contains description information of N devices in the substation, where N is a positive integer;
[0093] The first determining unit 302 is used to determine the semantic label map of the substation based on the color image and text description information, wherein the semantic label map contains the category label of the device to which each pixel of the color image belongs;
[0094] The rendering unit 303 is used to determine the three-dimensional Gaussian point cloud of the substation based on the color image, and to perform semantic rendering on the three-dimensional Gaussian point cloud based on the semantic label map to obtain the semantic rendering map.
[0095] The adjustment unit 304 is used to adjust the three-dimensional Gaussian point cloud based on the semantic rendering map to obtain the target Gaussian point cloud set;
[0096] The processing unit 305 is used to process the target Gaussian point cloud set through an augmented reality device to obtain an augmented reality view of the substation.
[0097] The substation view generation device provided in this application embodiment acquires a color image of the substation through an acquisition unit 301, determines text description information of the substation, wherein the text description information includes description information of N devices in the substation, where N is a positive integer; a first determination unit 302 determines a semantic tag map of the substation based on the color image and the text description information, wherein the semantic tag map includes the category tag of the device to which each pixel of the color image belongs; and a rendering unit 303 determines a three-dimensional Gaussian point cloud of the substation based on the color image, and performs semantic rendering on the three-dimensional Gaussian point cloud based on the semantic tag map. The process involves: a semantic rendering map; an adjustment unit 304 adjusting the 3D Gaussian point cloud based on the semantic rendering map to obtain a target Gaussian point cloud set; and a processing unit 305 processing the target Gaussian point cloud set using an augmented reality device to obtain an augmented reality view of the substation. This solves the problem of unclear substation views due to incomplete point cloud data collection in related technologies. By inputting text descriptions and color images of the devices, semantic label maps are automatically detected and generated, eliminating the need for predefined categories or manual annotation, significantly reducing the workload in the initial modeling stage and improving the efficiency and accuracy of semantic annotation. Combining the semantic information of color images and text descriptions, 3D substation modeling based on 3D Gaussian point clouds is achieved. This improves modeling accuracy, overcomes the high computational resource consumption problem of dense point clouds, and achieves seamless integration of the virtual model and the real view by processing the target Gaussian point cloud set using an augmented reality device. This allows maintenance personnel to intuitively and clearly identify substation equipment, enhancing the visualization experience and interactivity of the substation, thereby improving the clarity of the substation view.
[0098] Optionally, in the substation view generation device provided in this application embodiment, the first determining unit 302 includes: a first input module, used to input a color image and text description information into an object detection model to obtain bounding boxes of N objects and category labels to which the bounding boxes belong, wherein each object represents a device; a second input module, used to input a color image and bounding boxes of N objects into an image segmentation model to obtain a mask for each object, and add a category label of the corresponding device to the mask of each object, wherein the mask is used to represent the outline of the object; and an overlay module, used to merge and overlay the masks of N objects to obtain a semantic label map.
[0099] Optionally, in the substation view generation device provided in this application embodiment, the rendering unit 303 includes: a third input module, used to input a color image into three-dimensional reconstruction software to obtain a three-dimensional sparse point cloud set and a camera pose; and an initialization module, used to initialize each point cloud in the three-dimensional sparse point cloud set as a Gaussian ellipsoid to obtain a three-dimensional Gaussian point cloud, wherein the Gaussian ellipsoid contains at least one of the following three-dimensional spatial information: position, covariance matrix, opacity, color, and semantic feature vector, wherein the semantic feature vector is the feature vector corresponding to the category label.
[0100] Optionally, in the substation view generation device provided in this application embodiment, the device further includes: a second determining unit, configured to determine multiple observation rays based on the camera pose, and for each observation ray, determine the depth of the Gaussian ellipsoid intersecting with the observation ray, wherein the depth is the distance between the center of the Gaussian ellipsoid and the camera; a sorting unit, configured to sort the Gaussian ellipsoids intersecting with the observation rays in descending order of depth to obtain a list of Gaussian ellipsoids for the observation rays; and a mixing unit, configured to perform layer-by-layer mixing of the Gaussian ellipsoids in the list of Gaussian ellipsoids for each observation ray using a preset transparency processing method to obtain a color-rendered image of the three-dimensional Gaussian point cloud after color rendering.
[0101] Optionally, in the substation view generation apparatus provided in this application embodiment, the rendering unit 303 includes: a first determining module, configured to determine the semantic probability of the Gaussian ellipsoid traversed by each observation ray, wherein the semantic probability is the probability that the Gaussian ellipsoid belongs to each category label; a second determining module, configured to, for each pixel position in the three-dimensional Gaussian point cloud, determine the target observation ray passing through the pixel position, and determine the category probability vector of the pixel position based on the semantic probability of the Gaussian ellipsoid intersecting the target observation ray, wherein the category probability vector contains the probability that the pixel position belongs to each category label; and a third determining module, configured to determine a semantic rendering map based on the category probability vectors of all pixel positions.
[0102] Optionally, in the substation view generation device provided in this application embodiment, the adjustment unit 304 includes: a first comparison module, used to compare the color image with the color rendering image to obtain a color loss term; a second comparison module, used to compare the semantic rendering image with the semantic label image to obtain a semantic loss term; and a fourth determination module, used to determine a target loss function based on the color loss term and the semantic loss term, and to adjust the features of the Gaussian ellipsoid in the three-dimensional Gaussian point cloud with the goal of minimizing the loss value of the target loss function to obtain a target Gaussian point cloud set, wherein the features include at least one of the following: geometric, appearance, and semantic feature vectors, and the appearance includes color and opacity.
[0103] Optionally, in the substation view generation device provided in this application embodiment, the processing unit 305 includes: a conversion module, used to convert the coordinate system of the target Gaussian point cloud set into the target coordinate system of the augmented reality device; a mapping module, used to perform perspective mapping on each point cloud in the target Gaussian point cloud set under the target coordinate system to obtain the two-dimensional image coordinates of each point cloud; a fifth determination module, used to determine the point cloud set corresponding to each category label in the target Gaussian point cloud set; a rendering module, used to set the rendering parameters of the point cloud set corresponding to each category label, and render the point cloud set corresponding to each category label through the rendering parameters to obtain a rendered image, wherein the rendering parameters include at least one of the following: color and opacity; and a display module, used to mix the rendered image with the real-time acquired image of the augmented reality device according to the opacity to obtain an augmented reality view, and display the augmented reality view in the augmented reality device according to the two-dimensional image coordinates of each point cloud.
[0104] The substation view generation device includes a processor and a memory. The aforementioned acquisition unit 301, first determination unit 302, rendering unit 303, adjustment unit 304, and processing unit 305 are all stored in the memory as program units. The processor executes the aforementioned program units stored in the memory to realize the corresponding functions.
[0105] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured, and adjusting kernel parameters can improve the clarity of the substation view.
[0106] The memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0107] This invention provides a computer-readable storage medium storing a program that, when executed by a processor, implements a method for generating views of a substation.
[0108] This invention provides a processor for running a program, wherein the program executes a substation view generation method during runtime.
[0109] Figure 4 This is a schematic diagram of an electronic device provided according to an embodiment of this application. For example... Figure 4As shown, electronic device 401 includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it performs the following steps: acquiring a color image of the substation; determining textual description information of the substation, wherein the textual description information contains description information of N devices in the substation, where N is a positive integer; determining a semantic label map of the substation based on the color image and the textual description information, wherein the semantic label map contains the category label of the device to which each pixel of the color image belongs; determining a three-dimensional Gaussian point cloud of the substation based on the color image; performing semantic rendering on the three-dimensional Gaussian point cloud based on the semantic label map to obtain a semantic rendering map; adjusting the three-dimensional Gaussian point cloud based on the semantic rendering map to obtain a target Gaussian point cloud set; and processing the target Gaussian point cloud set using an augmented reality device to obtain an augmented reality view of the substation. The device in this paper can be a server, PC, PAD, mobile phone, etc.
[0110] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing an initialization program with the following method steps: acquiring a color image of a substation; determining textual description information of the substation, wherein the textual description information contains description information of N devices in the substation, where N is a positive integer; determining a semantic label map of the substation based on the color image and the textual description information, wherein the semantic label map contains a category label of the device to which each pixel of the color image belongs; determining a three-dimensional Gaussian point cloud of the substation based on the color image; performing semantic rendering on the three-dimensional Gaussian point cloud based on the semantic label map to obtain a semantic rendering map; adjusting the three-dimensional Gaussian point cloud based on the semantic rendering map to obtain a target Gaussian point cloud set; and processing the target Gaussian point cloud set through an augmented reality device to obtain an augmented reality view of the substation.
[0111] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0112] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0113] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0114] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0115] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0116] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0117] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0118] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0119] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0120] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for generating views of a substation, characterized in that, include: Acquire a color image of the substation and determine the text description information of the substation, wherein the text description information contains description information of N devices in the substation, where N is a positive integer; A semantic label map of the substation is determined based on the color image and the text description information, wherein the semantic label map contains the category label of the equipment to which each pixel of the color image belongs; The three-dimensional Gaussian point cloud of the substation is determined based on the color image, and the three-dimensional Gaussian point cloud is semantically rendered based on the semantic label map to obtain a semantic rendering map. The three-dimensional Gaussian point cloud is adjusted based on the semantic rendering graph to obtain the target Gaussian point cloud set; The target Gaussian point cloud set is processed by an augmented reality device to obtain an augmented reality view of the substation.
2. The method according to claim 1, characterized in that, Determining the semantic label map of the substation based on the color image and the text description information includes: The color image and the text description information are input into the object detection model to obtain bounding boxes of N objects and category labels to which the bounding boxes belong, wherein each object represents a device; The color image and the bounding boxes of the N objects are input into the image segmentation model to obtain a mask for each object, and a category label of the corresponding device is added to the mask of each object, wherein the mask is used to characterize the outline of the object; The masks of the N objects are merged and superimposed to obtain the semantic label map.
3. The method according to claim 1, characterized in that, Determining the three-dimensional Gaussian point cloud of the substation based on the color image includes: The color image is input into the 3D reconstruction software to obtain a 3D sparse point cloud set and camera pose. Each point cloud in the three-dimensional sparse point cloud set is initialized as a Gaussian ellipsoid to obtain the three-dimensional Gaussian point cloud, wherein the Gaussian ellipsoid contains at least one of the following three-dimensional spatial information: position, covariance matrix, opacity, color, and semantic feature vector, wherein the semantic feature vector is the feature vector corresponding to the category label.
4. The method according to claim 3, characterized in that, Before performing semantic rendering of the 3D Gaussian point cloud based on the semantic label map, the method further includes: Based on the camera pose, multiple observation rays are determined. For each observation ray, the depth of the Gaussian ellipsoid intersecting with the observation ray is determined, wherein the depth is the distance between the center of the Gaussian ellipsoid and the camera. The Gaussian ellipsoids intersecting the observation ray are sorted in descending order of depth to obtain a list of Gaussian ellipsoids for the observation ray. By using a preset transparency processing method, the Gaussian ellipsoids in the Gaussian ellipsoid list for each observation ray are blended layer by layer to obtain a color-rendered image of the three-dimensional Gaussian point cloud after color rendering.
5. The method according to claim 4, characterized in that, Based on the semantic label map, semantic rendering is performed on the 3D Gaussian point cloud to obtain a semantic rendering map, including: Determine the semantic probability of the Gaussian ellipsoid traversed by each observed ray, wherein the semantic probability is the probability that the Gaussian ellipsoid belongs to each of the category labels; For each pixel location in the 3D Gaussian point cloud, a target observation ray passing through the pixel location is determined, and a category probability vector for the pixel location is determined based on the semantic probability of the Gaussian ellipsoid intersecting the target observation ray. The category probability vector contains the probability that the pixel location belongs to each category label. The semantic rendering graph is determined based on the category probability vectors of all pixel locations.
6. The method according to claim 1, characterized in that, The 3D Gaussian point cloud is adjusted based on the semantic rendering map to obtain the target Gaussian point cloud set, which includes: The color image is compared with the color-rendered image to obtain the color loss term; The semantic rendering map is compared with the semantic label map to obtain the semantic loss term; A target loss function is determined based on the color loss term and the semantic loss term. The features of the Gaussian ellipsoid in the 3D Gaussian point cloud are adjusted with the goal of minimizing the loss value of the target loss function to obtain the target Gaussian point cloud set. The features include at least one of the following: geometric, appearance and semantic feature vectors, and the appearance includes color and opacity.
7. The method according to claim 1, characterized in that, The augmented reality view of the substation is obtained by processing the target Gaussian point cloud set using augmented reality devices, including: The coordinate system of the target Gaussian point cloud set is converted into the target coordinate system of the augmented reality device; Perform perspective mapping on each point cloud in the target Gaussian point cloud set under the target coordinate system to obtain the two-dimensional image coordinates of each point cloud; Determine the point cloud set corresponding to each category label in the target Gaussian point cloud set; The rendering parameters for the point cloud set corresponding to each category label are set, and the point cloud set corresponding to each category label is rendered using the rendering parameters to obtain a rendered image. The rendering parameters include at least one of the following: color and opacity. The rendered image is blended with the real-time image acquired by the augmented reality device according to the opacity to obtain the augmented reality view, and the augmented reality view is displayed on the augmented reality device according to the two-dimensional image coordinates of each point cloud.
8. An image generation device for a substation, characterized in that, include: The acquisition unit is used to acquire a color image of the substation and determine the text description information of the substation, wherein the text description information contains description information of N devices of the substation, where N is a positive integer; The first determining unit is configured to determine a semantic label map of the substation based on the color image and the text description information, wherein the semantic label map contains a category label of the equipment to which each pixel of the color image belongs; The rendering unit is used to determine the three-dimensional Gaussian point cloud of the substation based on the color image, and to perform semantic rendering on the three-dimensional Gaussian point cloud based on the semantic label map to obtain a semantic rendering map. An adjustment unit is used to adjust the three-dimensional Gaussian point cloud based on the semantic rendering map to obtain a target Gaussian point cloud set; The processing unit is used to process the target Gaussian point cloud set through an augmented reality device to obtain an augmented reality view of the substation.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the substation view generation method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, It includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the substation view generation method according to any one of claims 1 to 7.