Scene modeling method and device for electric power facility, equipment and medium
By generating Gaussian volumes of power facilities using SFM and 3DGS algorithms and adjusting the Gaussian density in conjunction with texture structure, the problem of difficulty in restoring texture details in traditional modeling methods is solved, thus achieving high-precision power facility scene modeling.
Patent Information
- Application Number
- CN202511771073.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-02-27
AI Technical Summary
Traditional power facility modeling methods struggle to accurately reproduce texture details, resulting in poor accuracy of 3D models during drone inspections.
SFM technology is used to generate point cloud data, combined with the 3DGS algorithm to generate Gaussian volume, and the Gaussian density is adjusted by calculating the gradient through backpropagation, and precise adaptation is performed according to the texture structure type.
The model improves the detail reproduction and texture realism of the power facility scene model, making the model more consistent with the physical characteristics and visual appearance of actual power facilities.
Smart Images

Figure CN121582473A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of modeling, and in particular to a method, apparatus, equipment and medium for scene modeling of power facilities. Background Technology
[0002] Power grid inspection is a core component of ensuring the safe operation of the power grid. Especially in high-voltage transmission lines, high-frequency, high-precision defect detection is required for critical components such as towers, conductors, and insulators. With the expansion and increasing complexity of the power grid, traditional manual inspection methods are no longer sufficient due to their low efficiency, high cost, and high risk. Unmanned aerial vehicle (UAV) inspection technology, with its flexibility, safety, and low cost, has become the mainstream solution for power grid inspection.
[0003] However, the massive amounts of image data collected by drones need to be converted into high-precision 3D models to assist in defect identification, which places stringent demands on modeling technology. Traditional modeling methods (such as photogrammetry and LiDAR) have the problem of failing to reproduce the fine texture structure of power facilities, resulting in poor accuracy of inspection methods that rely on 3D models.
[0004] Therefore, there is an urgent need for a new method for scene modeling of power facilities to improve the modeling accuracy. Summary of the Invention
[0005] This application provides a method, apparatus, equipment, and medium for scene modeling of power facilities, in order to improve the modeling accuracy of power facilities.
[0006] In a first aspect, embodiments of this application provide a scene modeling method for power facilities, including:
[0007] Acquire multiple frames of real images of power facilities taken from multiple perspectives;
[0008] SFM technology is used to process the multiple frames of real images to obtain point cloud data of the power facilities;
[0009] Using the 3DGS algorithm, multiple Gaussian volumes are generated based on the point cloud data, and the multiple Gaussian volumes are projected and rasterized from different viewpoints to obtain multi-frame rendered images corresponding to the viewpoints of the multiple frames of real images.
[0010] For each frame of rendered image, based on the difference between the rendered image and the real image under the corresponding viewpoint, the gradient of each Gaussian body in the rendered image is obtained by backpropagation; the gradient includes gradient value and gradient direction.
[0011] When the gradient value of any Gaussian body is greater than a preset gradient threshold, the Gaussian density of the spatial region corresponding to the Gaussian body is adjusted according to the covariance matrix of the Gaussian body and the texture structure type of the power facility at the position corresponding to the Gaussian body in the real image corresponding to the rendered image.
[0012] A scene model of the power facility is obtained based on the adjusted Gaussian density.
[0013] In one possible implementation, when the gradient value of any Gaussian body is greater than a preset gradient threshold, adjusting the Gaussian density of the spatial region corresponding to the Gaussian body based on the covariance matrix of the Gaussian body and the texture structure type at the position corresponding to the Gaussian body in the real image corresponding to the rendered image includes:
[0014] Geometric feature recognition is performed on the real image corresponding to the rendered image to obtain the texture structure type of each region in the real image;
[0015] The texture structure type at the location in the real image corresponding to the Gaussian volume is determined as the texture structure type of the Gaussian volume;
[0016] If the texture structure type of the Gaussian body is a complex texture region, and the covariance matrix in the Gaussian parameters of the Gaussian body is less than a first preset threshold, then the Gaussian body is cloned along the gradient direction in the corresponding spatial region of the Gaussian body.
[0017] If the texture structure type of the Gaussian body is a flat texture region, and the covariance matrix in the Gaussian parameters of the Gaussian body is greater than a second preset threshold, then the Gaussian body is split.
[0018] In one possible implementation, the step of employing a 3DGS algorithm to generate multiple Gaussian volumes based on the point cloud data, and projecting and rasterizing the multiple Gaussian volumes from different viewpoints to obtain multi-frame rendered images corresponding to the viewpoints of the multiple real images, includes:
[0019] Based on the point cloud data, multiple Gaussian volumes are constructed and the Gaussian parameters of each Gaussian volume are obtained. The Gaussian parameters include the spatial position, covariance matrix, spherical harmonic coefficients, and opacity of the Gaussian volume.
[0020] For each frame of real image, using the camera parameters corresponding to the real image, the multiple Gaussian bodies are projected onto a two-dimensional plane with a viewpoint corresponding to any real image, resulting in multiple projected Gaussian bodies.
[0021] Based on the colors and opacities of the multiple Gaussian volumes, the color of each pixel in the two-dimensional plane is determined, resulting in a rendered image corresponding to the real image viewpoint.
[0022] In one possible implementation, the step of constructing multiple Gaussian volumes based on the point cloud data and obtaining the Gaussian parameters of each Gaussian volume includes:
[0023] The spatial location of each point cloud in the point cloud data is taken as the spatial location of a Gaussian body, and a Gaussian body is constructed at the spatial location of the Gaussian body.
[0024] Based on the preset prior geometric constraints of the power facilities, the covariance matrix of each Gaussian body is determined;
[0025] Based on the color of each pixel in the multi-frame real images, the spherical harmonic coefficients and opacity of the Gaussian body are obtained.
[0026] In one possible implementation, the step of calculating the gradient of each Gaussian body in the rendered image through backpropagation based on the difference between the rendered image and the real image at the corresponding viewpoint includes:
[0027] The loss function of the rendering function is calculated based on the difference between the rendered image and the real image at the corresponding viewpoint.
[0028] For each Gaussian volume in the rendered image, the partial derivative of the loss function with respect to each Gaussian parameter corresponding to the Gaussian volume is calculated to obtain the gradient of the Gaussian volume.
[0029] In one possible implementation, calculating the loss function of the rendering function based on the difference between the rendered image and the real image at the corresponding viewpoint includes:
[0030] Based on the color and brightness differences of corresponding pixels between the rendered image and the real image at the corresponding viewpoint, the multi-view luminance loss of the rendered image is calculated.
[0031] Based on the geometric features of each region in the rendered image and the prior geometric rules of the power facilities, the power prior geometric loss of the rendered image is determined;
[0032] The loss function is obtained by weighting and summing the multi-view photometric loss and the power prior geometric loss according to the preset loss weights.
[0033] In one possible implementation, obtaining the scene model of the power facility based on the adjusted Gaussian density includes:
[0034] Based on the adjusted Gaussian density, multiple adjusted Gaussian bodies are obtained, and the Gaussian parameters of the multiple adjusted Gaussian bodies are updated using a gradient descent optimization strategy.
[0035] Based on the Gaussian parameters of the adjusted Gaussian volumes, the adjusted Gaussian volumes are re-projected and rasterized to obtain multiple new rendered images corresponding to the viewpoints of the multiple real images.
[0036] For each new rendered image, based on the difference between the new rendered image and the real image under the corresponding viewpoint, the gradient of each Gaussian body in the new rendered image is recalculated through backpropagation. When the gradient value of any Gaussian body is greater than a preset gradient threshold, the Gaussian density of the spatial region corresponding to the Gaussian body is adjusted according to the covariance matrix of the Gaussian body and the texture structure type at the position corresponding to the Gaussian body in the real image corresponding to the new rendered image.
[0037] Repeat this step until the gradient value of any Gaussian body is less than or equal to the preset gradient threshold, and then obtain the scene model of the power facility based on the multiple new rendered images.
[0038] Secondly, embodiments of this application provide a scene modeling device for power facilities, comprising:
[0039] The first acquisition module is used to acquire multiple frames of real images of power facilities taken from multiple perspectives.
[0040] The first processing module uses SFM technology to process the multiple frames of real images to obtain point cloud data of the power facilities.
[0041] The second processing module is used to generate multiple Gaussian bodies based on the point cloud data using the 3DGS algorithm, and to project and rasterize the multiple Gaussian bodies from different viewpoints to obtain multiple frame rendered images corresponding to the viewpoints of the multiple frames of real images.
[0042] The calculation module is used to calculate the gradient of each Gaussian body in the rendered image for each frame of rendered image based on the difference between the rendered image and the real image under the corresponding viewpoint through backpropagation; the gradient includes gradient value and gradient direction.
[0043] The adjustment module is used to adjust the Gaussian density of the spatial region corresponding to the Gaussian body according to the covariance matrix of the Gaussian body and the texture structure type at the position corresponding to the Gaussian body in the real image corresponding to the rendered image when the gradient value of any Gaussian body is greater than a preset gradient threshold.
[0044] The second acquisition module is used to acquire a scene model of the power facility based on the adjusted Gaussian density.
[0045] In one possible implementation, the adjustment module is specifically used for:
[0046] Geometric feature recognition is performed on the real image corresponding to the rendered image to obtain the texture structure type of each region in the real image;
[0047] The texture structure type at the location in the real image corresponding to the Gaussian volume is determined as the texture structure type of the Gaussian volume;
[0048] If the texture structure type of the Gaussian body is a complex texture region, and the covariance matrix in the Gaussian parameters of the Gaussian body is less than a first preset threshold, then the Gaussian body is cloned along the gradient direction in the corresponding spatial region of the Gaussian body.
[0049] If the texture structure type of the Gaussian body is a flat texture region, and the covariance matrix in the Gaussian parameters of the Gaussian body is greater than a second preset threshold, then the Gaussian body is split.
[0050] In one possible implementation, the second processing module includes:
[0051] The first processing unit is used to construct multiple Gaussian bodies based on the point cloud data and obtain the Gaussian parameters of each Gaussian body. The Gaussian parameters include the spatial position, covariance matrix, spherical harmonic coefficients, and opacity of the Gaussian body.
[0052] The second processing unit is used to project the multiple Gaussian bodies onto a two-dimensional plane with a viewpoint corresponding to any real image for each frame of real image, using the camera parameters corresponding to the real image, to obtain the multiple Gaussian bodies after projection.
[0053] The third processing unit is used to determine the color of each pixel in the two-dimensional plane based on the color and opacity of the plurality of Gaussian volumes, so as to obtain a rendered image corresponding to the real image viewpoint.
[0054] In one possible implementation, the first processing unit is specifically used for:
[0055] The spatial location of each point cloud in the point cloud data is taken as the spatial location of a Gaussian body, and a Gaussian body is constructed at the spatial location of the Gaussian body.
[0056] Based on the preset prior geometric constraints of the power facilities, the covariance matrix of each Gaussian body is determined;
[0057] Based on the color of each pixel in the multi-frame real images, the spherical harmonic coefficients and opacity of the Gaussian body are obtained.
[0058] In one possible implementation, the computing module includes:
[0059] The first calculation unit is used to calculate the loss function of the rendering function based on the difference between the rendered image and the real image under the corresponding viewpoint.
[0060] The second calculation unit is used to calculate the partial derivative of the loss function with respect to each Gaussian parameter corresponding to the Gaussian body for each Gaussian body in the rendered image, so as to obtain the gradient of the Gaussian body.
[0061] In one possible implementation, the first computing unit is specifically used for:
[0062] Based on the color and brightness differences of corresponding pixels between the rendered image and the real image at the corresponding viewpoint, the multi-view luminance loss of the rendered image is calculated.
[0063] Based on the geometric features of each region in the rendered image and the prior geometric rules of the power facilities, the power prior geometric loss of the rendered image is determined;
[0064] The loss function is obtained by weighting and summing the multi-view photometric loss and the power prior geometric loss according to the preset loss weights.
[0065] In one possible implementation, the second acquisition unit is specifically used for:
[0066] Based on the adjusted Gaussian density, multiple adjusted Gaussian bodies are obtained, and the Gaussian parameters of the multiple adjusted Gaussian bodies are updated using a gradient descent optimization strategy.
[0067] Based on the Gaussian parameters of the adjusted Gaussian volumes, the adjusted Gaussian volumes are re-projected and rasterized to obtain multiple new rendered images corresponding to the viewpoints of the multiple real images.
[0068] For each new rendered image, based on the difference between the new rendered image and the real image under the corresponding viewpoint, the gradient of each Gaussian body in the new rendered image is recalculated through backpropagation. When the gradient value of any Gaussian body is greater than a preset gradient threshold, the Gaussian density of the spatial region corresponding to the Gaussian body is adjusted according to the covariance matrix of the Gaussian body and the texture structure type at the position corresponding to the Gaussian body in the real image corresponding to the new rendered image.
[0069] Repeat this step until the gradient value of any Gaussian body is less than or equal to the preset gradient threshold, and then obtain the scene model of the power facility based on the multiple new rendered images.
[0070] Thirdly, embodiments of this application provide a computer device, including: a memory and a processor;
[0071] The memory stores computer-executed instructions;
[0072] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0073] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0074] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0075] The power facility scene modeling method, apparatus, equipment, and medium provided in this application embodiment first acquire multi-view real photos of the power facility, then generate point clouds using SFM technology and obtain rendered images based on the 3DGS algorithm, then compare the rendered images with the real images to calculate the gradients of each Gaussian volume, and finally adjust the Gaussian density according to the gradient and texture structure. This method makes the adjustment of Gaussian density no longer a blind modification based solely on the gradient, but rather a precise adaptation combined with the texture structure types of different locations of the power facility. This significantly improves the detail reproduction and texture realism of the power facility scene model, making the model more closely match the physical characteristics and visual appearance of the actual power facility. Attached Figure Description
[0076] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0077] Figure 1 This is a flowchart illustrating a scene modeling method for power facilities provided in Embodiment 1 of this application;
[0078] Figure 2 A schematic diagram illustrating the principle of the scene modeling method for power facilities provided in this application embodiment;
[0079] Figure 3 One of the schematic diagrams illustrating the scene model of a power facility provided in the application;
[0080] Figure 4 The second illustration shows the effect of a scene model of a power facility provided for the application;
[0081] Figure 5This is a schematic diagram of the structure of a scene modeling device for power facilities provided in Embodiment 3 of this application;
[0082] Figure 6 A schematic diagram of the structure of the computer device provided in this application.
[0083] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0084] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0085] To facilitate understanding of the technical content of this solution, the background technology is described in detail below:
[0086] Ensuring the reliability and safety of power transmission networks is the cornerstone of modern society. Power line inspection, as a key link in ensuring power grid stability, has the core task of timely detection of defects in transmission lines, such as damaged insulators. Traditional manual inspection is labor-intensive, inefficient, and exposes workers to the risks of electric shock and falls from heights. While helicopter inspection has been introduced, it is costly, noisy, and has limited mobility in mountainous areas, making it difficult to carry out on a large scale and routinely.
[0087] In recent years, with the development of technologies such as autonomous driving, drones have become a revolutionary technology for power line inspection. They can approach power facilities at low cost, high flexibility, and high safety to collect high-resolution data. However, how to automatically and accurately construct 3D models from massive amounts of image data and identify defects has become the core bottleneck of current technological development.
[0088] To achieve comprehensive digitization of power facilities, the industry has explored the application of various 3D reconstruction technologies, but all have limitations in power inspection scenarios. Photogrammetry is cost-effective, requiring only a standard red-green-blue (RGB) camera, and can generate realistic 3D models. However, due to the lack of texture in key components of power facilities and their sensitivity to lighting, the reconstructed models are prone to breakage and voids, and cannot accurately analyze defects in components such as insulators. LiDAR (Light Detection and Ranging) offers high geometric accuracy, is unaffected by lighting, can operate in all weather conditions, and penetrates vegetation, but it is costly and heavy, and its point clouds lack color and texture details, making it unable to identify defects that rely on appearance.
[0089] Therefore, there is an urgent need for a modeling method that can highly reproduce the texture details of power facilities.
[0090] To address the problems mentioned above, the inventors discovered during their research that, to overcome the difficulty of accurately reproducing texture details using traditional modeling methods, they utilized 3D Gaussian Splatting (3DGS) technology to generate an explicit model containing millions of three-dimensional Gaussian primitives. Furthermore, by dynamically adjusting the Gaussian volume density based on the texture type of each region in the power inspection video, the reconstruction accuracy of key components of power facilities was improved to the centimeter level. This precisely balances the geometric accuracy of power facilities with photorealistic qualities, solving the problem of distortion in weak texture structures in traditional point cloud or photogrammetry.
[0091] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0092] Figure 1 This is a flowchart illustrating a scene modeling method for power facilities provided in Embodiment 1 of this application, as shown below. Figure 1 As shown, it includes:
[0093] S101. Acquire multiple frames of real images of power facilities taken from multiple perspectives.
[0094] In this step, in order to achieve 3D modeling of the power facilities, multiple frames of real images of the power facilities taken from multiple different perspectives will be acquired to provide multi-view input for subsequent modeling.
[0095] In practical applications, drones can be used to capture real-time surround video of power facilities and break down the inspection video into a continuous image sequence frame by frame, with each frame corresponding to the instantaneous state of the power facility from a certain perspective.
[0096] For example, multiple frames of real-world images contain power facilities such as poles, conductors, and insulator strings.
[0097] S102. Using SFM technology, multiple frames of real images are processed to obtain point cloud data of power facilities.
[0098] In this step, SFM technology will first be used to process multiple frames of real images into point clouds to obtain point cloud data of power facilities.
[0099] Specifically, the acquired multi-frame real images are processed using the Structure from Motion (SfM) algorithm: First, ORB (Oriented Fast and Rotated BRIEF) feature points are extracted from the images, and preliminary matching of feature points between images from different viewpoints is performed; then, the Random Sample Consensus (RANSAC) algorithm is used to filter out mismatched feature points to ensure the accuracy of the matching results; finally, based on the purified feature point matching relationship, the three-dimensional spatial coordinates of the feature points are obtained through triangulation calculation, thereby constructing a sparse point cloud of power facilities.
[0100] S103. Using the 3DGS algorithm, multiple Gaussian volumes are generated based on point cloud data. The Gaussian volumes are then projected and rasterized from different viewpoints to obtain multi-frame rendered images corresponding to the viewpoints of multiple real images.
[0101] It should be understood that in the power sector, the point cloud data constructed using SfM technology only initially outlines the macroscopic structure of power facilities such as the location of towers and the direction of conductors. However, due to the scarcity of feature points, the point cloud density is low and the accuracy is poor in some details, which still require enhancement using 3DGS technology.
[0102] In this step, the 3DGS algorithm is also required to construct multiple 3D Gaussian volumes based on the generated point cloud data. Then, these multiple Gaussian volumes are projected and rasterized from different viewpoints to obtain multi-frame rendered images corresponding to the viewpoints of multiple real images.
[0103] In this process, the perspective from which the Gaussian volume is projected each time is consistent with the perspective from which the real image of the corresponding frame is captured. Consequently, each of the acquired multi-frame rendered images has a real image corresponding to its perspective.
[0104] S104. For each frame of rendered image, based on the difference between the rendered image and the real image under the corresponding viewpoint, the gradient of each Gaussian body in the rendered image is obtained by backpropagation.
[0105] The gradient includes the gradient value and the gradient direction.
[0106] In this step, for each frame, starting from the difference between the rendered image and the real image at the corresponding viewpoint, the contribution of each Gaussian parameter of each Gaussian body to the difference is calculated, and finally the gradient of each Gaussian body is obtained, which is used to perform reverse optimization on the rendered image.
[0107] Specifically, the magnitude of the gradient is used to indicate the size of the contribution of the Gaussian body to the difference at each Gaussian parameter; the gradient direction is used to indicate the direction in which each Gaussian parameter of the Gaussian body needs to be adjusted.
[0108] S105. When the gradient value of any Gaussian body is greater than the preset gradient threshold, adjust the Gaussian density of the spatial region corresponding to the Gaussian body according to the covariance matrix after the Gaussian body is projected and the texture structure type at the position corresponding to the Gaussian body in the real image corresponding to the rendered image.
[0109] In this step, if the gradient value of any Gaussian volume exceeds a preset gradient threshold, it indicates that the geometric representation of the spatial region containing the Gaussian volume is insufficient. At this point, a real image corresponding to the viewpoint of the rendered image is acquired. Based on the pixel position of the Gaussian volume in the rendered image, the texture structure type at the location corresponding to the pixel position of the Gaussian volume in the real image is obtained. Then, the Gaussian density of the spatial region corresponding to the Gaussian volume is adjusted based on the covariance matrix in the Gaussian parameters of the Gaussian volume and the texture structure type.
[0110] It should be understood that the covariance matrix of a Gaussian body represents the distribution, scale, and orientation of the three-dimensional Gaussian body in space, reflecting the geometric representation of the Gaussian body to its spatial region. Meanwhile, in real images, the texture structure type at the location corresponding to the Gaussian body is determined based on the real structure of the power facilities, and is used to clarify the visual detail reproduction requirements of the corresponding spatial region (for example, the detail reproduction requirements of complex texture structures are higher).
[0111] Therefore, based on the actual geometric representation of the Gaussian body in its region and the actual detail reproduction requirements of that region, a basis can be provided for adjusting the density of the Gaussian body.
[0112] S106. Obtain a scene model of the power facilities based on the adjusted Gaussian density.
[0113] In this step, the projection and rasterization operations on the newly obtained Gaussian volumes will be re-executed based on the adjusted Gaussian density to obtain multiple new rendered images and a scene model of the power facilities. It should be understood that these multiple new rendered images are closer to the real representation of the power facilities, and correspondingly, the scene model obtained is more accurate.
[0114] The scene modeling method for power facilities provided in this application first acquires multi-view real photos of the power facilities, then generates point clouds using SFM technology and obtains rendered images based on the 3DGS algorithm, then compares the rendered images with the real images to calculate the gradients of each Gaussian volume, and finally adjusts the Gaussian density according to the gradient and texture structure. This method makes the adjustment of Gaussian density no longer a blind modification based solely on the gradient, but rather a precise adaptation combined with the texture structure types of different locations of the power facilities. This significantly improves the detail reproduction and texture realism of the power facility scene model, making the model more closely match the physical characteristics and visual appearance of the actual power facilities.
[0115] Furthermore, Embodiment 2 of this application provides a scene modeling method for power facilities. Based on the above embodiments, this embodiment provides a detailed description of the specific implementation methods of the above steps, including:
[0116] Step 1: Acquire multiple frames of real images of the power facilities from multiple perspectives.
[0117] Step 2: Using SFM technology, process multiple frames of real images to obtain point cloud data of power facilities.
[0118] Step 3: Using the 3DGS algorithm, multiple Gaussian volumes are generated based on the point cloud data. The Gaussian volumes are then projected and rasterized from different viewpoints to obtain multi-frame rendered images corresponding to the viewpoints of multiple real images.
[0119] Specifically, step 3 includes the following steps 3.1 to 3.3:
[0120] Step 3.1: Based on the point cloud data, construct multiple Gaussian volumes and obtain the Gaussian parameters of each Gaussian volume. The Gaussian parameters include the spatial location of the Gaussian volume, the covariance matrix, the spherical harmonic coefficients, and the opacity.
[0121] In this step, a 3D Gaussian volume will be initialized for each point cloud or its neighborhood based on the point cloud data, and the Gaussian parameters will be obtained.
[0122] Among them, spatial position represents the three-dimensional coordinates of the Gaussian volume in the three-dimensional world coordinate system, used to determine the specific location of the Gaussian volume in the power facility scenario; the covariance matrix of the Gaussian volume is used to characterize the geometric characteristics of the three-dimensional Gaussian volume in space, such as its distribution shape, scale, and orientation; the spherical harmonic coefficient is used to characterize the color and light reflection characteristics of the Gaussian volume, used to simulate the color change of the Gaussian volume under different viewing angles and different lighting conditions; and opacity is used to characterize the Gaussian volume's ability to block light. The spatial position, covariance matrix, spherical harmonic coefficient, and opacity of the Gaussian parameters define the properties of the Gaussian volume from four dimensions.
[0123] In detail, from a mathematical perspective, the 3DGS algorithm reasonably simplifies the traditional 3D Gaussian formula. The 3DGS technology simplifies the 3D Gaussian distribution formula as follows: since the Gaussian distribution is centered on the target point (point cloud), its mean has already been centered, so the mean vector μ is set to 0; at the same time, the pre-normalization coefficient used to ensure the integral value of the probability density function is 1 is discarded, allowing the scale of the Gaussian distribution to be freely adjusted. Its specific mathematical definition relies on the mean vector μ and the covariance matrix Σ, and the Gaussian function... The definition is as follows:
[0124]
[0125] in, Let be the coordinates in the 3D Gaussian projection space. Covariance matrix. ,in, Let be the rotation matrix and be the scaling matrix. Physically speaking, the covariance matrix Σ of a 3D Gaussian is similar to the shape of an ellipsoid.
[0126] In one specific implementation, the Gaussian volume can be constructed and the Gaussian parameters obtained using the methods described in steps 3.1.1 to 3.1.3:
[0127] Step 3.1.1: Treat the spatial location of each point cloud in the point cloud data as the spatial location of a Gaussian body, and construct a Gaussian body at the spatial location of the Gaussian body.
[0128] In practical applications, a three-dimensional Gaussian volume can be initialized for each point cloud or its neighborhood.
[0129] Step 3.1.2: Determine the covariance matrix of each Gaussian body based on the preset prior geometric constraints of the power facilities.
[0130] In this step, to further improve the modeling accuracy, the covariance matrix of each Gaussian body can be determined based on the preset prior geometric constraints of the power facilities.
[0131] Specifically, the covariance matrix used to express the spatial distribution morphology of the Gaussian body will be determined based on the pre-defined a priori geometric constraints of the power facilities corresponding to the spatial region where the Gaussian body is located (e.g., the geometric characteristics of the conductor are linear, the geometric characteristics of the insulator are cylindrical, the tower body has vertical characteristics, and the crossarm has symmetrical characteristics).
[0132] Step 3.1.3: Based on the color of each pixel in multiple frames of real images, obtain the spherical harmonic coefficients and opacity of the Gaussian volume.
[0133] The methods used to obtain the spherical harmonic coefficients and opacity of the Gaussian body in this step are the same as those used in traditional 3DGS modeling methods, and will not be repeated here.
[0134] The method provided in this implementation uses the spatial location of the point cloud as the basis for Gaussian volume localization, combines the prior geometric constraints of the power facility to determine the covariance matrix, and then extracts the spherical harmonic coefficients and opacity from the pixel colors of the real image to construct the Gaussian volume. This makes the spatial distribution of the Gaussian volume no longer an irregular initial assignment, but a precise definition that strictly follows the geometric structural characteristics of the power facility. This ensures that the initial Gaussian volume model fits the physical structural attributes of the power facility from the construction stage, greatly reducing the correction cost in the subsequent optimization stage.
[0135] In step 3.2: For each frame of real image, using the camera parameters corresponding to the real image, project multiple Gaussian bodies onto a two-dimensional plane at the viewpoint corresponding to any real image to obtain multiple projected Gaussian bodies.
[0136] The camera parameters include both intrinsic and extrinsic parameters.
[0137] Specifically, the coverage area of the Gaussian volume on the image is calculated by deriving the covariance matrix and the projection matrix, thus obtaining multiple Gaussian volumes after projection.
[0138] In detail, this step utilizes the anisotropic volume "sputtering" effect of the 3D Gaussian to achieve mapping rendering from three-dimensional space to a two-dimensional image plane. When rendering from a specific viewpoint, the 3D Gaussian is projected onto the image plane to form a 2D Gaussian. The mean of the 2D Gaussian is determined by the projection matrix, and its corresponding covariance matrix is approximately expressed as follows when projected onto the two-dimensional plane:
[0139]
[0140] in, This represents the view transformation matrix from the world coordinate system to the camera coordinate system. The Jacobian matrix representing the affine approximation of perspective projection transformation. Let represent the covariance matrix after the Gaussian projection.
[0141] Step 3.3: Based on the colors and opacities of multiple Gaussian volumes, determine the color of each pixel in the two-dimensional plane to obtain a rendered image corresponding to the real image viewpoint.
[0142] In detail, when rendering an image from a given viewpoint, for each pixel, N Gaussian bodies overlapping at that pixel are obtained, and these N Gaussian bodies are sorted according to the opacity value of each Gaussian body. Based on the sorted N Gaussian bodies, the color of the pixel is calculated.
[0143] For pixels in an image Its color The calculation formula is:
[0144]
[0145] in, is the color represented by the spherical harmonic coefficients in Gaussian body i; 1 to N represent the mixing order of each Gaussian function; The opacity of the Gaussian body i is calculated by considering the point. The Gaussian function is obtained.
[0146] It should be understood that steps 3.1 to 3.3 above construct a Gaussian volume containing spatial location, covariance matrix, spherical harmonic coefficients, and opacity based on point cloud data, and then project it onto a two-dimensional plane in combination with camera parameters and generate a rendered image by determining the pixel color based on color and opacity. This method enables the spatial features and visual attributes of the three-dimensional Gaussian volume to be accurately mapped to the two-dimensional image perspective, achieving a high degree of fidelity in the rendered image in terms of geometric structure and color texture, and providing an accurate visual comparison benchmark for subsequent model optimization.
[0147] It should be understood that step 3 uses the snowballing algorithm to differentially accumulate the color and opacity of the Gaussian volume to the image pixels according to the three-dimensional probability density function, making the rasterization rendering process differentiable. This differentiability of the rendering process provides the foundation for optimizing the Gaussian parameters of the Gaussian volume through backpropagation via gradient calculation.
[0148] Step 4: For each frame of the rendered image, based on the difference between the rendered image and the real image under the corresponding viewpoint, calculate the gradient of each Gaussian body in the rendered image through backpropagation.
[0149] The gradient includes the gradient value and the gradient direction.
[0150] Specifically, step 4 includes the following steps 4.1 to 4.2:
[0151] Step 4.1: Calculate the loss function of the rendering function based on the difference between the rendered image and the real image at the corresponding viewpoint.
[0152] The loss function can be, for example, the difference between the rendered image and the real image in terms of pixel color, lighting consistency, etc.
[0153] In one possible implementation, the loss function can be obtained using steps 4.1.1 to 4.1.3 as follows:
[0154] Step 4.1.1: Calculate the multi-view luminance loss of the rendered image based on the color and brightness differences of corresponding pixels between the rendered image and the real image at the corresponding viewpoint.
[0155] It should be understood that multi-view photometric loss is used to measure the differences between the rendered image and the real image in terms of pixel-level color, brightness, and texture, ensuring that the rendered image generated after inverse optimization based on the loss function visually approximates the real appearance of power facilities.
[0156] For example,
[0157] Where H and W are the height and width of the image, , These are the color brightness values at pixel (i,j) of the rendered image and the real image, respectively.
[0158] Step 4.1.2: Determine the power prior geometric loss of the rendered image based on the geometric features of each region in the rendered image and the prior geometric rules of the power facilities.
[0159] In this step, key structures of power facilities (such as towers, crossarms, and conductors) in the rendered image can be identified through algorithms such as semantic segmentation and edge detection, and their geometric features (such as the main direction of the tower and the plane of symmetry of the crossarm) can be extracted from the rendered image. Then, based on the preset prior geometric rules of power facilities, the difference loss between the features of the identified key structures and the prior geometry is calculated.
[0160] For example, if the identified power facility is a tower, the angle loss between the main direction of the tower in the rendered image and the Z-axis of the world coordinate system is calculated from the verticality of the power facility in the rendered image to obtain the power prior geometric loss.
[0161] For example, if the identified power facility is a crossarm, the chamfer distance loss of the left and right side structures of the crossarm in the rendered image is calculated to obtain the power prior geometric loss.
[0162] It should be understood that this scheme, by setting a prior geometric loss for power, enables the gradient optimization algorithm based on the loss function to correct the problem that the geometric features of power facilities in the rendered image deviate from the real form.
[0163] Step 4.1.3: Based on the preset loss weights, perform a weighted summation of the multi-view photometric loss and the power prior geometric loss to obtain the loss function.
[0164] Specifically, the loss function The calculation formula is as follows:
[0165]
[0166] in, This represents the preset loss weights of the loss function; Indicates multi-view photometric loss; This represents the prior geometric loss of electricity.
[0167] The method provided in this implementation calculates the multi-view photometric loss of the rendered image and the real image separately, calculates the geometric loss by combining the prior geometric rules of the power facility, and then obtains the loss function by weighted summation of the two types of losses according to preset weights. This allows the model optimization objective to simultaneously take into account the visual reproduction accuracy and physical structural specifications of the power facility. The rendered image not only matches the appearance of the real power facility in terms of pixel-level color and brightness, but also follows the power engineering design rules such as vertical tower and symmetrical crossarms in terms of geometric structure. Ultimately, this greatly improves the accuracy and rationality of the 3D reconstruction of the power facility.
[0168] Step 4.2: For each Gaussian volume in the rendered image, calculate the partial derivative of the loss function with respect to each Gaussian parameter corresponding to the Gaussian volume to obtain the gradient of the Gaussian volume.
[0169] For example, the gradient vector of a Gaussian body is
[0170] in, The partial derivative of the loss function with respect to the spatial location μ is given by the loss function. The partial derivative of the loss function with respect to the covariance matrix Σ; The partial derivative of the loss function with respect to the spherical harmonic coefficients S; Let be the partial derivative of the loss function with respect to the opacity α.
[0171] Step 5: When the gradient value of any Gaussian body is greater than the preset gradient threshold, adjust the Gaussian density of the spatial region corresponding to the Gaussian body according to the covariance matrix of the Gaussian body and the texture structure type at the position corresponding to the Gaussian body in the real image corresponding to the rendered image.
[0172] Specifically, step 5 may include steps 5.1 to 5.4 as follows:
[0173] Step 5.1: Perform geometric feature recognition on the real image corresponding to the rendered image to obtain the texture structure type of each region in the real image.
[0174] In this step, geometric feature recognition and texture analysis algorithms (such as semantic segmentation, edge density statistics, and gray-level co-occurrence matrix) are used to extract the texture attributes of each region in the real image from the same viewpoint as the rendered image.
[0175] Step 5.2: Determine the texture structure type of the location in the real image corresponding to the Gaussian volume as the texture structure type of the Gaussian volume.
[0176] In this step, the projection relationship of Gaussian volume is used to find the two-dimensional pixel region corresponding to each Gaussian volume in the real image, and the texture structure type of the pixel region is directly determined as the texture structure type of the corresponding three-dimensional Gaussian volume.
[0177] Step 5.3: If the texture structure type of the Gaussian body is a complex texture region, and the covariance matrix in the Gaussian parameters of the Gaussian body is less than the first preset threshold, then clone the Gaussian body along the gradient direction in the corresponding spatial region of the Gaussian body.
[0178] It should be understood that texture structure types include complex texture areas, that is, areas with dense edges, dramatic grayscale changes, and rich details, such as the patterns on power insulators, bolt groups at tower connections, tower edges, and crossarm connection points.
[0179] In this step, if the Gaussian body's texture structure is a complex texture region, it indicates a higher demand for detail reproduction. If the covariance matrix of the Gaussian body is less than a first preset threshold, it means the Gaussian body's spatial range is small and its coverage is limited, and the current quantity is insufficient to reproduce the details. Therefore, it is necessary to clone the Gaussian body along the gradient direction in the corresponding spatial region.
[0180] Specifically, the cloned Gaussian body is spatially similar to the original Gaussian body, and inherits parameters such as the covariance matrix, spherical harmonic coefficients, and opacity of the original Gaussian body, with only its spatial position finely adjusted along the gradient.
[0181] For example, if the covariance of a certain Gaussian body in the insulator pattern area is small, it means that it can only cover a small part of the pattern. After cloning, increasing the number of Gaussian bodies can completely restore the complex details of the pattern and avoid blurring of the pattern in the rendered image.
[0182] Step 5.4: If the texture structure type of the Gaussian body is a flat texture region, and the covariance matrix in the Gaussian parameters of the Gaussian body is greater than the second preset threshold, then split the Gaussian body.
[0183] It should be understood that texture structure types also include flat texture areas, which are areas with sparse edges, uniform gray levels, and no obvious details, such as the smooth sides of metal towers and the main body of conductors.
[0184] In this step, if the texture structure type of the Gaussian body is a flat texture region, and if the covariance matrix of the Gaussian body is greater than the second preset threshold, it indicates that the Gaussian body has a large spatial range, loose shape, and excessive coverage. This region is over-reconstructed, which can easily lead to blurring of the flat region. Therefore, it is necessary to split the Gaussian body.
[0185] Specifically, the split Gaussian bodies are evenly distributed within the spatial range of the original Gaussian body; the covariance matrix of each Gaussian body is reduced, the shape is more compact, and it inherits the spherical harmonic coefficients and opacity of the original Gaussian body (to ensure the color consistency of the smooth texture).
[0186] For example, if the smooth surface of a metal tower is represented by a large Gaussian volume, the rendered surface will be blurry and lack texture. Splitting it into multiple small Gaussian volumes can ensure that the entire tower surface is covered, avoid blurring, and reduce the computational redundancy caused by a single large Gaussian volume.
[0187] It should be understood that the methods provided in steps 5.1 to 5.4 above first obtain the texture structure type of each region by performing geometric feature recognition on the real image and matching it to the corresponding Gaussian body. Then, based on the dual conditions of texture structure type and covariance matrix threshold, the method clones Gaussian bodies with complex textures and small covariance matrices along the gradient direction and splits Gaussian bodies with flat textures and large covariance matrices. This constructs an adaptive density control mechanism based on gradient backpropagation and covariance matrix analysis, so that the number and distribution of Gaussian bodies can accurately adapt to the texture features and geometric expression requirements of different regions of power facilities. This achieves accurate restoration of details in complex texture regions (such as insulator patterns and bolt groups) by cloning and supplementing Gaussian bodies, and avoids redundant calculations and blur artifacts in flat texture regions (such as the surface of metal towers and conductor bodies) by splitting and simplifying Gaussian bodies. This achieves a precise balance of Gaussian densities that are not visible in different regions, and ultimately improves the detail accuracy and computational efficiency of the scene model for the 3D reconstruction of power facilities.
[0188] Step 6: Obtain a scene model of the power facilities based on the adjusted Gaussian density.
[0189] Step 6 includes the following steps 6.1 to 6.4:
[0190] Step 6.1: Based on the adjusted Gaussian density, obtain multiple adjusted Gaussian bodies, and use the gradient descent optimization strategy to update the Gaussian parameters of the multiple adjusted Gaussian bodies.
[0191] The adjusted Gaussian bodies include the original Gaussian bodies that were not adjusted, as well as newly added Gaussian bodies after adjustment.
[0192] In this step, a gradient descent optimization strategy will be used to optimize the Gaussian parameters of each Gaussian body based on the gradient optimization direction.
[0193] Step 6.2: Based on the Gaussian parameters of the adjusted Gaussian volumes, re-project and rasterize the adjusted Gaussian volumes to obtain multiple new rendered images corresponding to the viewpoints of multiple real images.
[0194] In this step, the rendered image will be regenerated based on the updated Gaussian volume parameters, providing a new reference object for comparing the rendered image with the real image in the next round.
[0195] Step 6.3: For each new rendered image, based on the difference between the new rendered image and the real image under the corresponding viewpoint, recalculate the gradient of each Gaussian body in the new rendered image through backpropagation. When the gradient value of any Gaussian body is greater than the preset gradient threshold, adjust the Gaussian density of the spatial region corresponding to the Gaussian body according to the covariance matrix of the Gaussian body and the texture structure type at the position corresponding to the Gaussian body in the real image corresponding to the new rendered image.
[0196] In this step, the gradient determined based on the difference between the real image and the rendered image will be iteratively used to check the deviation of the rendered image after a new round of parameter updates. If there is still a significant deviation (i.e. the gradient exceeds the threshold), the Gaussian density will be adjusted again to further optimize the model.
[0197] Step 6.4: Repeat steps 6.1 to 6.3 above until the gradient value of any Gaussian body is less than or equal to the preset gradient threshold. Based on multiple new rendered images, obtain the scene model of the power facility.
[0198] This step refers to outputting a 3D scene model of the power facility after ensuring that the final scene model can meet the preset accuracy requirements, with the iteration termination condition being "the gradient value of any Gaussian body is less than or equal to the preset gradient threshold".
[0199] It should be understood that steps 6.1 to 6.2 above involve first obtaining a Gaussian volume based on the adjusted Gaussian density and updating its Gaussian parameters using a gradient descent strategy, then reprojecting and rasterizing to generate a new rendered image. Subsequently, the gradient is recalculated based on the difference between the new rendered image and the real image, and the Gaussian density is adjusted as needed. Finally, the process iterates until all Gaussian volume gradients converge to generate a power facility scene model. This method allows the obtained scene model to continuously correct the parameters and density distribution of the Gaussian volume through closed-loop iteration, enabling the optimization process to accurately adapt to the texture features and geometric constraints of the power facility. This gradually reduces the rendering error of the power facility scene model to an acceptable range. The final generated model accurately restores complex details such as insulator patterns and bolt groups, while strictly adhering to the geometric rules of power engineering, such as tower verticality and crossarm symmetry, achieving the effect of high-precision digital reconstruction of the three-dimensional scene of the power facility.
[0200] The scene modeling method for power facilities provided in this application embodiment effectively improves the modeling accuracy and scene adaptability of power scenes. It overcomes the defects of model breakage and void caused by the complex texture and light sensitivity of power facilities in photogrammetry. It avoids the shortcomings of LiDAR, such as high cost and lack of color texture, which makes it unable to identify appearance defects. Through 3DGS technology and power prior constraints, it accurately restores the geometric shape of key components such as conductors and insulators, while retaining color information to support the identification of appearance defects.
[0201] Figure 2 A schematic diagram illustrating the principle of the scene modeling method for power facilities provided in this application embodiment, as shown below. Figure 2 As shown, this solution uses real-time inspection videos collected by inspection drones as input. After slicing the inspection videos to obtain multiple frames of realistic images, a closed-loop process of "SfM point cloud construction → 3D Gaussian volume initialization → projection and differentiable rasterization → adaptive density control → iterative optimization" is employed to achieve high-precision 3D reconstruction of power facilities. The core logic is to explicitly represent the scene using a 3D Gaussian volume, and accurately restore components such as conductors and insulators through differentiable rendering and adaptive optimization. The flowchart corresponds to a collaborative mechanism between the workflow (solid line) and gradient flow (dashed line) to support iterative optimization of the model.
[0202] Furthermore, it should be noted that existing technologies also include modeling techniques with high modeling accuracy, such as Neural Radiance Fields (NeRF). While these techniques can generate realistic rendered images and have a strong ability to model complex optical phenomena, their computational costs are staggering, their processing speed is slow, and the generated geometric structures lack metric-level accuracy, making it difficult to meet engineering needs.
[0203] To address this issue, the modeling method provided in this solution can also meet the real-time modeling requirements of power line inspection.
[0204] Specifically, during the processing of video slices from drone inspections, after a certain number of image frames have been calculated, the current iteration's power facility scene model is output in real time, providing users with a real-time preview of the 3D form. Based on this, by continuously receiving new video slice inputs and cyclically executing the core processing flow of "projection-rasterization-density control-optimization," dynamic incremental updates of the scene model can be achieved, thereby accurately adapting to the core real-time requirements of drone inspection scenarios.
[0205] In each frame-fixing process, a targeted parameter update strategy is adopted: only the scene region corresponding to the newly added image frame is subjected to iterative update of Gaussian volume parameters, while the parameters of the historical scene region that has been optimized are frozen. This greatly reduces unnecessary repetitive calculation load and achieves a fast response from video stream input to 3D model output.
[0206] It should be understood that the method provided by this solution, in terms of interactive efficiency, overcomes the drawbacks of NeRF's offline rendering and inability to interact in real time. It relies on 3DGS to achieve real-time rendering, supporting immersive virtual roaming for inspection personnel, allowing them to observe equipment details from any angle, significantly improving defect detection efficiency. In the data utilization and processing workflow, a hybrid route strategy is designed to ensure data acquisition quality, and multi-source data is integrated through standardized processes to achieve stable and high-quality modeling. Regarding functional integration and industry adaptability, it integrates interactive inspection, geometric measurement, and defect identification functions, specifically addressing pain points in power line inspection, providing comprehensive decision support, creating an efficient and practical solution for power line integrity diagnosis, and promoting the upgrade of power line inspection from a traditional, inefficient model to a digital and intelligent one.
[0207] Figure 3 and Figure 4 These are all schematic diagrams illustrating the scene model of a power facility provided in the application. Figure 3 and Figure 4 The model shown is built using the scene modeling method for power facilities provided in this solution. During the modeling process, incremental modeling driven by video slicing, combined with lightweight SfM and adaptive density control, achieves a single-frame processing time of ≤80ms (RTX 4090 graphics card). With a 30fps video stream input, the model latency is ≤1s, supporting real-time viewing of the 3D model on-site during inspections and assisting maintenance personnel in quickly locating equipment defects (such as tower tilting, insulator damage, etc.). Unlike traditional point cloud / photogrammetry, which is prone to distortion in weakly textured areas such as conductors and insulators, this solution, based on 3DGS technology, uses dynamic splitting to encrypt high volumes and prior constraints to correct morphology, ensuring conductor reconstruction errors are ≤5cm and insulator geometric deviations are ≤3cm, meeting the accuracy requirements of digital operation and maintenance of power facilities.
[0208] Figure 5 This is a schematic diagram of the structure of a scene modeling device for power facilities provided in Embodiment 3 of this application, as shown below. Figure 5 As shown, the scene modeling device 20 for power facilities provided in this embodiment includes:
[0209] The first acquisition module 201 is used to acquire multiple frames of real images of power facilities taken from multiple perspectives.
[0210] The first processing module 202 uses SFM technology to process multiple frames of real images to obtain point cloud data of power facilities.
[0211] The second processing module 203 is used to generate multiple Gaussian bodies based on point cloud data using the 3DGS algorithm, and to project and rasterize the multiple Gaussian bodies from different perspectives to obtain multi-frame rendered images corresponding to the perspectives of multiple real images.
[0212] The calculation module 204 is used to calculate the gradient of each Gaussian body in the rendered image for each frame of the rendered image based on the difference between the rendered image and the real image under the corresponding viewpoint through backpropagation; the gradient includes the gradient value and the gradient direction.
[0213] The adjustment module 205 is used to adjust the Gaussian density of the spatial region corresponding to the Gaussian body according to the covariance matrix of the Gaussian body and the texture structure type at the position corresponding to the Gaussian body in the real image corresponding to the rendered image when the gradient value of any Gaussian body is greater than the preset gradient threshold.
[0214] The second acquisition module 206 is used to acquire a scene model of the power facility based on the adjusted Gaussian density.
[0215] In one possible implementation, the adjustment module 205 is specifically used for:
[0216] Geometric feature recognition is performed on the real image corresponding to the rendered image to obtain the texture structure type of each region in the real image;
[0217] The texture structure type at the location in the real image corresponding to the Gaussian volume is determined as the texture structure type of the Gaussian volume.
[0218] If the texture structure type of the Gaussian body is a complex texture region, and the covariance matrix in the Gaussian parameters of the Gaussian body is less than the first preset threshold, then the Gaussian body is cloned along the gradient direction in the corresponding spatial region of the Gaussian body.
[0219] If the texture structure type of the Gaussian body is a flat texture region, and the covariance matrix in the Gaussian parameters of the Gaussian body is greater than the second preset threshold, then the Gaussian body is split.
[0220] In one possible implementation, the second processing module 202 includes:
[0221] The first processing unit is used to construct multiple Gaussian volumes based on point cloud data and obtain the Gaussian parameters of each Gaussian volume. The Gaussian parameters include the spatial location, covariance matrix, spherical harmonic coefficients, and opacity of the Gaussian volume.
[0222] The second processing unit is used to project multiple Gaussian bodies onto a two-dimensional plane with a viewpoint corresponding to any real image for each frame of real image, using the camera parameters corresponding to the real image, to obtain multiple Gaussian bodies after projection.
[0223] The third processing unit is used to determine the color of each pixel in the two-dimensional plane based on the color and opacity of multiple Gaussian volumes, so as to obtain a rendered image corresponding to the real image viewpoint.
[0224] In one possible implementation, the first processing unit is specifically used for:
[0225] The spatial location of each point in the point cloud data is taken as the spatial location of a Gaussian body, and a Gaussian body is constructed at the spatial location of the Gaussian body.
[0226] Based on the pre-defined prior geometric constraints of the power facilities, the covariance matrix of each Gaussian body is determined;
[0227] Based on the color of each pixel in multiple frames of real images, the spherical harmonic coefficients and opacity of the Gaussian volume are obtained.
[0228] In one possible implementation, the computing module 204 includes:
[0229] The first calculation unit is used to calculate the loss function of the rendering function based on the difference between the rendered image and the real image under the corresponding viewpoint.
[0230] The second computational unit is used to calculate the partial derivative of the loss function with respect to each Gaussian parameter of the Gaussian body for each Gaussian body in the rendered image, and to obtain the gradient of the Gaussian body.
[0231] In one possible implementation, the first computing unit is specifically used for:
[0232] Based on the color and brightness differences of corresponding pixels between the rendered image and the real image at the corresponding viewpoint, calculate the multi-view luminance loss of the rendered image;
[0233] Based on the geometric features of each region in the rendered image and the prior geometric rules of the power facilities, the power prior geometric loss of the rendered image is determined.
[0234] Based on the preset loss weights, the multi-view photometric loss and the power prior geometric loss are weighted and summed to obtain the loss function.
[0235] In one possible implementation, the second acquisition unit is specifically used for:
[0236] Based on the adjusted Gaussian density, multiple adjusted Gaussian bodies are obtained, and the Gaussian parameters of the multiple adjusted Gaussian bodies are updated using a gradient descent optimization strategy.
[0237] Based on the Gaussian parameters of the adjusted Gaussian volumes, the adjusted Gaussian volumes are re-projected and rasterized to obtain multiple new rendered images corresponding to the viewpoints of multiple real images.
[0238] For each new rendered image, based on the difference between the new rendered image and the real image under the corresponding viewpoint, the gradient of each Gaussian body in the new rendered image is recalculated through backpropagation. When the gradient value of any Gaussian body is greater than the preset gradient threshold, the Gaussian density of the spatial region corresponding to the Gaussian body is adjusted according to the covariance matrix of the Gaussian body and the texture structure type at the position corresponding to the Gaussian body in the real image corresponding to the new rendered image.
[0239] Repeat this step until the gradient value of any Gaussian body is less than or equal to the preset gradient threshold, and obtain the scene model of the power facility based on multiple new rendered images.
[0240] The scene modeling device 20 for power facilities provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0241] Figure 6 A schematic diagram of the structure of the computer device provided in this application. Figure 5 As shown, the computer device 30 provided in this embodiment includes at least one processor 301 and a memory 302. Optionally, the device 30 further includes a communication component 303. The processor 301, memory 302, and communication component 303 are connected via a bus 304.
[0242] In a specific implementation, at least one processor 301 executes computer execution instructions stored in memory 302, causing at least one processor 301 to perform the above-described method.
[0243] The specific implementation process of processor 301 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0244] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0245] The memory may include read-only memory and random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Sync Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0246] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0247] This application also provides a computer program product, including a computer program that, when executed, implements the above-described method.
[0248] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed, implement the above-described method.
[0249] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as SRAM, EEPROM, EPROM, PROM, ROM, magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0250] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside within an ASIC. Alternatively, the processor and the readable storage medium can exist as discrete components in a device.
[0251] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0252] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0253] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0254] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0255] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0256] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A method for scene modeling of power facilities, characterized in that, include: Acquire multiple frames of real images of power facilities taken from multiple perspectives; SFM technology is used to process the multiple frames of real images to obtain point cloud data of the power facilities; Using the 3DGS algorithm, multiple Gaussian volumes are generated based on the point cloud data, and the multiple Gaussian volumes are projected and rasterized from different viewpoints to obtain multi-frame rendered images corresponding to the viewpoints of the multiple frames of real images. For each frame of rendered image, based on the difference between the rendered image and the real image under the corresponding viewpoint, the gradient of each Gaussian body in the rendered image is obtained by backpropagation; the gradient includes gradient value and gradient direction. When the gradient value of any Gaussian body is greater than a preset gradient threshold, the Gaussian density of the spatial region corresponding to the Gaussian body is adjusted according to the covariance matrix of the Gaussian body and the texture structure type at the position corresponding to the Gaussian body in the real image corresponding to the rendered image. A scene model of the power facility is obtained based on the adjusted Gaussian density.
2. The method according to claim 1, characterized in that, When the gradient value of any Gaussian body is greater than a preset gradient threshold, adjusting the Gaussian density of the spatial region corresponding to the Gaussian body based on the covariance matrix of the Gaussian body and the texture structure type at the position corresponding to the Gaussian body in the real image corresponding to the rendered image includes: Geometric feature recognition is performed on the real image corresponding to the rendered image to obtain the texture structure type of each region in the real image; The texture structure type at the location in the real image corresponding to the Gaussian volume is determined as the texture structure type of the Gaussian volume; If the texture structure type of the Gaussian body is a complex texture region, and the covariance matrix in the Gaussian parameters of the Gaussian body is less than a first preset threshold, then the Gaussian body is cloned along the gradient direction in the corresponding spatial region of the Gaussian body. If the texture structure type of the Gaussian body is a flat texture region, and the covariance matrix in the Gaussian parameters of the Gaussian body is greater than a second preset threshold, then the Gaussian body is split.
3. The method according to claim 1 or 2, characterized in that, The method employs a 3DGS algorithm to generate multiple Gaussian volumes based on the point cloud data. These Gaussian volumes are then projected and rasterized from different viewpoints to obtain multi-frame rendered images corresponding to the viewpoints of the multiple real-world images, including: Based on the point cloud data, multiple Gaussian volumes are constructed and the Gaussian parameters of each Gaussian volume are obtained. The Gaussian parameters include the spatial position, covariance matrix, spherical harmonic coefficients, and opacity of the Gaussian volume. For each frame of real image, using the camera parameters corresponding to the real image, the multiple Gaussian bodies are projected onto a two-dimensional plane with a viewpoint corresponding to any real image, resulting in multiple projected Gaussian bodies. Based on the colors and opacities of the multiple Gaussian volumes, the color of each pixel in the two-dimensional plane is determined, resulting in a rendered image corresponding to the real image viewpoint.
4. The method according to claim 3, characterized in that, The process of constructing multiple Gaussian volumes based on the point cloud data and obtaining the Gaussian parameters of each Gaussian volume includes: The spatial location of each point cloud in the point cloud data is taken as the spatial location of a Gaussian body, and a Gaussian body is constructed at the spatial location of the Gaussian body. Based on the preset prior geometric constraints of the power facilities, the covariance matrix of each Gaussian body is determined; Based on the color of each pixel in the multi-frame real images, the spherical harmonic coefficients and opacity of the Gaussian body are obtained.
5. The method according to claim 1 or 2, characterized in that, The step of obtaining the gradient of each Gaussian volume in the rendered image through backpropagation based on the difference between the rendered image and the real image at the corresponding viewpoint includes: The loss function of the rendering function is calculated based on the difference between the rendered image and the real image at the corresponding viewpoint. For each Gaussian volume in the rendered image, the partial derivative of the loss function with respect to each Gaussian parameter corresponding to the Gaussian volume is calculated to obtain the gradient of the Gaussian volume.
6. The method according to claim 5, characterized in that, The step of calculating the loss function of the rendering function based on the difference between the rendered image and the real image at the corresponding viewpoint includes: Based on the color and brightness differences of corresponding pixels between the rendered image and the real image at the corresponding viewpoint, the multi-view luminance loss of the rendered image is calculated. Based on the geometric features of each region in the rendered image and the prior geometric rules of the power facilities, the power prior geometric loss of the rendered image is determined; The loss function is obtained by weighting and summing the multi-view photometric loss and the power prior geometric loss according to the preset loss weights.
7. The method according to claim 1 or 2, characterized in that, The process of obtaining the scene model of the power facility based on the adjusted Gaussian density includes: Based on the adjusted Gaussian density, multiple adjusted Gaussian bodies are obtained, and the Gaussian parameters of the multiple adjusted Gaussian bodies are updated using a gradient descent optimization strategy. Based on the Gaussian parameters of the adjusted Gaussian volumes, the adjusted Gaussian volumes are re-projected and rasterized to obtain multiple new rendered images corresponding to the viewpoints of the multiple real images. For each new rendered image, based on the difference between the new rendered image and the real image under the corresponding viewpoint, the gradient of each Gaussian body in the new rendered image is recalculated through backpropagation. When the gradient value of any Gaussian body is greater than a preset gradient threshold, the Gaussian density of the spatial region corresponding to the Gaussian body is adjusted according to the covariance matrix of the Gaussian body and the texture structure type at the position corresponding to the Gaussian body in the real image corresponding to the new rendered image. Repeat this step until the gradient value of any Gaussian body is less than or equal to the preset gradient threshold, and then obtain the scene model of the power facility based on the multiple new rendered images.
8. A scene modeling device for power facilities, characterized in that, include: The first acquisition module is used to acquire multiple frames of real images of power facilities taken from multiple perspectives. The first processing module is used to process the multiple frames of real images using SFM technology to obtain point cloud data of the power facilities. The second processing module is used to generate multiple Gaussian bodies based on the point cloud data using the 3DGS algorithm, and to project and rasterize the multiple Gaussian bodies from different viewpoints to obtain multiple frame rendered images corresponding to the viewpoints of the multiple frames of real images. The calculation module is used to calculate the gradient of each Gaussian body in the rendered image for each frame of rendered image based on the difference between the rendered image and the real image under the corresponding viewpoint through backpropagation; the gradient includes gradient value and gradient direction. The adjustment module is used to adjust the Gaussian density of the spatial region corresponding to the Gaussian body according to the covariance matrix of the Gaussian body and the texture structure type at the position corresponding to the Gaussian body in the real image corresponding to the rendered image when the gradient value of any Gaussian body is greater than a preset gradient threshold. The second acquisition module acquires a scene model of the power facility based on the adjusted Gaussian density.
9. A computer device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-7.
Citation Information
Cited By
Low-computing-power rapid three-dimensional modeling method and system based on 3DGS
CN122023680A