Method and device for determining scene model of cable tunnel and electronic equipment

By acquiring the three-dimensional point cloud data of the cable tunnel and generating a virtual image, a high-precision cable tunnel scene model was constructed, which solved the problem of low model accuracy in the existing technology and achieved higher accuracy and detail expression.

CN120672993APending Publication Date: 2025-09-19STATE GRID BEIJING ELECTRIC POWER CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510693754.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately construct scene models of cable tunnels, resulting in low model accuracy.

Method used

By acquiring the three-dimensional point cloud data of the cable tunnel, the three-dimensional space model is determined, and virtual images are generated through multiple virtual shooting poses, ultimately constructing a high-precision scene model.

Benefits of technology

The accuracy and detail of the cable tunnel scene model are improved, solving the problem of low model accuracy in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672993A_ABST
    Figure CN120672993A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for determining a scene model of a cable tunnel and electronic equipment. The method comprises the following steps: acquiring three-dimensional point cloud data corresponding to a target cable tunnel; determining a three-dimensional space model corresponding to the target cable tunnel according to the three-dimensional point cloud data; determining a plurality of virtual shooting poses corresponding to the three-dimensional space model; virtual images corresponding to the multiple virtual shooting poses are determined according to the three-dimensional point cloud data, the corresponding virtual images are determined according to projection of the three-dimensional point cloud data on the corresponding two-dimensional image coordinate systems, and the multiple virtual shooting poses correspond to the multiple two-dimensional image coordinate systems one to one; and determining a scene model corresponding to the target cable tunnel according to the three-dimensional point cloud data and the plurality of virtual images. According to the invention, the technical problem of low accuracy of the constructed scene model due to complex internal environment and single geometrical shape of the cable tunnel in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a method, device and electronic equipment for determining a scene model of a cable tunnel. Background Art

[0002] Establishing an accurate scene model for cable tunnels is crucial for their operation and maintenance, safety, and future planning. Currently, 3D Gaussian rendering is the primary method used to build these scene models. However, this technique requires extracting image features to train the scene model. However, the complex interior environment and simple geometry of cable tunnels make feature extraction difficult, resulting in a scene model with low accuracy.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] The embodiments of the present invention provide a method, device and electronic device for determining a scene model of a cable tunnel, so as to at least solve the technical problem in related technologies of low accuracy of the constructed scene model due to the complex internal environment and single geometric shape of the cable tunnel.

[0005] According to one aspect of an embodiment of the present invention, a method for determining a scene model of a cable tunnel is provided, comprising: acquiring three-dimensional point cloud data corresponding to a target cable tunnel; determining a three-dimensional space model corresponding to the target cable tunnel based on the three-dimensional point cloud data; determining multiple virtual shooting poses corresponding to the three-dimensional space model; determining virtual images corresponding to the multiple virtual shooting poses based on the three-dimensional point cloud data, wherein the corresponding virtual images are determined based on the projection of the three-dimensional point cloud data on a corresponding two-dimensional image coordinate system, and the multiple virtual shooting poses correspond one-to-one to multiple two-dimensional image coordinate systems; determining the scene model corresponding to the target cable tunnel based on the three-dimensional point cloud data and the multiple virtual images.

[0006] Optionally, determining the virtual images corresponding to the multiple virtual shooting poses respectively based on the three-dimensional point cloud data includes: determining adjustment parameters corresponding to the multiple virtual shooting poses respectively; determining multiple adjustment coordinates corresponding to the three-dimensional point cloud data based on the multiple adjustment parameters, wherein the multiple adjustment coordinates are the coordinates corresponding to the three-dimensional point cloud data in multiple three-dimensional shooting coordinate systems, and the multiple three-dimensional shooting coordinate systems correspond one-to-one to the multiple virtual shooting poses; determining device parameters corresponding to the virtual shooting device; determining pixel coordinates corresponding to the multiple virtual shooting poses respectively based on the device parameters and the multiple adjustment coordinates, wherein the corresponding pixel coordinates are the projection coordinates of the corresponding adjustment coordinates on the corresponding two-dimensional image coordinate system; and determining the virtual images corresponding to the multiple virtual shooting poses respectively based on the multiple pixel coordinates.

[0007] Optionally, determining the scene model corresponding to the target cable tunnel based on the three-dimensional point cloud data and multiple virtual images includes: determining multiple simulated object models based on the three-dimensional point cloud data, wherein the multiple simulated object models are used to represent three-dimensional models corresponding to multiple target objects in the target cable tunnel; updating model coefficients corresponding to the multiple simulated object models based on the multiple virtual images to obtain multiple updated models; and determining the scene model corresponding to the target cable tunnel based on the multiple updated models.

[0008] Optionally, updating the model coefficients corresponding to the multiple simulated object models respectively based on the multiple virtual images to obtain multiple updated models includes: determining the environmental adjustment parameters corresponding to the multiple virtual shooting postures respectively; updating the model coefficients corresponding to the multiple simulated object models respectively based on the multiple virtual images and the environmental adjustment parameters corresponding to the multiple virtual images respectively to obtain the multiple updated models.

[0009] Optionally, determining the scene model corresponding to the target cable tunnel based on the multiple updated models includes: determining test images corresponding to the multiple virtual shooting poses respectively based on the multiple updated models, wherein the corresponding test images are images determined based on the projections of the multiple updated models on the corresponding two-dimensional image coordinate systems; determining an error value based on the multiple test images and the virtual images corresponding to the multiple test images; and determining the scene model corresponding to the target cable tunnel based on the multiple updated models when the error value is lower than an error threshold.

[0010] Optionally, determining the test images corresponding to the multiple virtual shooting poses respectively based on the multiple updated models includes: determining the model projections of the multiple updated models on the corresponding two-dimensional image coordinate system; determining the pixel information corresponding to the multiple model projections respectively; and determining the test images corresponding to the multiple virtual shooting poses respectively based on the pixel information corresponding to the multiple model projections respectively.

[0011] Optionally, before obtaining the three-dimensional point cloud data corresponding to the target cable tunnel, the method further includes: obtaining scene data corresponding to the target cable tunnel, wherein the scene data includes initial point cloud data and a regional image; determining a plurality of target pixel points corresponding to the initial point cloud data from a plurality of initial pixel points included in the regional image; determining image information corresponding to the plurality of target pixel points respectively, wherein the corresponding image information is used to represent pixel feature information of the regional image at the corresponding target pixel points; determining the three-dimensional point cloud data based on a plurality of image information and the initial point cloud data, wherein a plurality of data points in the three-dimensional point cloud data respectively carry corresponding image information.

[0012] According to one aspect of an embodiment of the present invention, a device for determining a scene model of a cable tunnel is provided, comprising: an acquisition module for acquiring three-dimensional point cloud data corresponding to a target cable tunnel; a first determination module for determining a three-dimensional space model corresponding to the target cable tunnel based on the three-dimensional point cloud data; a second determination module for determining a plurality of virtual shooting poses corresponding to the three-dimensional space model; a third determination module for determining virtual images corresponding to the plurality of virtual shooting poses based on the three-dimensional point cloud data, wherein the corresponding virtual images are determined based on the projection of the three-dimensional point cloud data on a corresponding two-dimensional image coordinate system, and the plurality of virtual shooting poses correspond one-to-one to the plurality of two-dimensional image coordinate systems; and a fourth determination module for determining the scene model corresponding to the target cable tunnel based on the three-dimensional point cloud data and the plurality of virtual images.

[0013] According to one aspect of an embodiment of the present invention, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement any of the above methods for determining a scene model of a cable tunnel.

[0014] According to one aspect of an embodiment of the present invention, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute any of the above methods for determining a scene model of a cable tunnel.

[0015] In an embodiment of the present invention, a method is adopted in which three-dimensional point cloud data corresponding to a target cable tunnel is obtained; a three-dimensional space model corresponding to the target cable tunnel is determined based on the three-dimensional point cloud data; a plurality of virtual shooting poses corresponding to the three-dimensional space model are determined; a virtual image corresponding to each of the plurality of virtual shooting poses is determined based on the three-dimensional point cloud data, wherein the corresponding virtual image is determined based on the projection of the three-dimensional point cloud data on the corresponding two-dimensional image coordinate system, and the plurality of virtual shooting poses correspond one-to-one to the plurality of two-dimensional image coordinate systems; a scene model corresponding to the target cable tunnel is determined based on the three-dimensional point cloud data and the plurality of virtual images. By determining the virtual images corresponding to each of the plurality of virtual shooting poses based on the three-dimensional point cloud data, the purpose of determining the scene model corresponding to the target cable tunnel based on the three-dimensional point cloud data and the plurality of virtual images is achieved. Since the plurality of virtual images can reflect the characteristics of objects inside the cable tunnel, the accuracy of the determined scene model is improved, thereby solving the technical problem in the related art that the accuracy of the constructed scene model is low due to the complex internal environment and the single geometric shape of the cable tunnel. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0017] Figure 1 is a flow chart of a method for determining a scene model of a cable tunnel according to an embodiment of the present invention;

[0018] Figure 2 is a flowchart of scene model construction provided by an optional embodiment of the present invention;

[0019] Figure 3 is a flow chart of a training update model provided by an optional embodiment of the present invention;

[0020] Figure 4 It is a structural block diagram of a device for determining a scene model of a cable tunnel according to an embodiment of the present invention. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0022] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0023] Example 1

[0024] According to an embodiment of the present invention, an embodiment of a method for determining a scene model of a cable tunnel is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0025] Figure 1 : is a flow chart of a method for determining a scene model of a cable tunnel according to an embodiment of the present invention, such as Figure 1 As shown, the method includes the following steps:

[0026] Step S102: Acquire three-dimensional point cloud data corresponding to the target cable tunnel.

[0027] In step S102 provided in the present application, three-dimensional point cloud data corresponding to the target cable tunnel is obtained.

[0028] Among them, a target cable tunnel is involved, and the target cable tunnel refers to a cable tunnel that needs to be three-dimensionally reconstructed.

[0029] Among them, three-dimensional point cloud data is involved. Three-dimensional point cloud data refers to three-dimensional spatial data composed of a series of discrete three-dimensional coordinate points. Three-dimensional coordinate points are usually surface points of objects or points in scenes collected in space, and carry additional information such as color, normal, etc., for more detailed scene description.

[0030] In this step, 3D point cloud data of the cable tunnel interior is collected using technologies such as laser scanning, structured light scanning, or other 3D data acquisition methods. This 3D point cloud data contains high-precision geometric information about the cable tunnel interior, as well as information about the cable tunnel surface, such as color and material. This data is used for 3D scene reconstruction and rendering, forming the foundation for subsequent modeling and analysis.

[0031] Step S104: determining a three-dimensional space model corresponding to the target cable tunnel based on the three-dimensional point cloud data.

[0032] In step S104 provided in the present application, a three-dimensional space model corresponding to the target cable tunnel is determined.

[0033] Among them, a three-dimensional space model is involved. The three-dimensional space model refers to a mathematical representation used to describe and reconstruct three-dimensional scenes in the real world, which can reflect the volume of the target cable tunnel.

[0034] In this step, a 3D spatial model reflecting the cable tunnel volume is constructed using the 3D point cloud data obtained from the target cable tunnel. This process typically involves technical steps such as data preprocessing, feature extraction, coordinate transformation, point cloud registration, and model optimization. This step determines the spatial data of the cable tunnel, such as its length, width, height, and volume. This abstract data is converted into an intuitive image, providing a direct understanding of the cable tunnel's spatial layout.

[0035] Step S106: determining a plurality of virtual shooting poses corresponding to the three-dimensional space model.

[0036] In step S106 provided in the present application, a plurality of virtual shooting poses corresponding to the three-dimensional space model are determined.

[0037] Among them, virtual shooting pose is involved, which refers to the camera position and direction under different perspectives simulated in a three-dimensional space model.

[0038] This step creates a series of virtual camera positions and orientations within the constructed 3D cable tunnel model. These virtual camera poses cover multiple angles within and around the model. Actual photography may encounter physical limitations, such as difficult-to-reach locations or difficult lighting conditions. Virtual camera poses overcome these limitations and generate images from perspectives that are difficult to obtain in physical photography. This allows for the efficient generation of large amounts of training data, saving time and costs while ensuring both quality and quantity.

[0039] It's important to note that generating high-quality training images requires a reasonable assessment of the virtual camera poses used to capture the image. This avoids rendering angles that don't exist or aren't visible in the model, which could render the training data invalid or misleading. To achieve smooth transitions from one viewpoint to another, it's crucial to ensure that the virtual camera poses are spatially continuous, meaning that the changes between adjacent poses are minimal. This is crucial for generating continuous image sequences and achieving smooth video rendering.

[0040] In step S108, virtual images corresponding to the plurality of virtual shooting poses are determined based on the three-dimensional point cloud data, wherein the corresponding virtual images are determined based on the projection of the three-dimensional point cloud data onto the corresponding two-dimensional image coordinate system, and the plurality of virtual shooting poses correspond one-to-one to the plurality of two-dimensional image coordinate systems.

[0041] In step S108 provided in the present application, virtual images corresponding to the plurality of virtual shooting postures are determined.

[0042] This involves virtual images, which are two-dimensional images generated by projecting three-dimensional point cloud data onto a virtual two-dimensional image coordinate system. These virtual images simulate the view captured by a camera when observing a cable tunnel from a corresponding angle and can serve as input data for model training.

[0043] Among them, the two-dimensional image coordinate system is involved. The two-dimensional image coordinate system refers to a coordinate system used to locate and describe the pixel positions in the image. The two-dimensional image coordinate system is used to project points in three-dimensional space onto a two-dimensional image. This process usually involves the camera's intrinsic parameters (such as focal length, image size, and image principal point position) and extrinsic parameters (such as the camera's position and orientation in the world coordinate system). When three-dimensional point cloud data is projected onto a two-dimensional image coordinate system through a specified virtual shooting pose, the generated image is the scene seen from the perspective of the virtual pose. The projection process involves geometric transformations, such as perspective projection or orthogonal projection, which converts points in three-dimensional space into pixel positions and color values ​​in a two-dimensional image coordinate system through a camera model.

[0044] In this step, a virtual image is generated for each virtual shooting pose based on the 3D point cloud data. These virtual images are obtained by projecting the point cloud data onto the corresponding 2D image coordinate system when the virtual camera observes the 3D point cloud model from different positions and orientations. This process generates multi-angle views by simulating observations from different perspectives, enriching the perspectives and data for model training.

[0045] This step generates virtual images for each virtual shooting pose, significantly increasing the amount of data used for model training and improving the model's generalization capabilities. Even when actual data is limited, data augmentation techniques can simulate a rich range of perspectives, enabling the model to better learn the scene's 3D structure and surface details. In cable tunnel scenarios, this helps the model overcome practical limitations of shooting, such as lighting conditions, confined spaces, and object occlusion, improving the robustness and accuracy of model reconstruction and rendering.

[0046] It’s important to note that virtual images can also be combined with real-world images to verify and optimize the model. By comparing the similarity between virtual images and actual captured images, the quality of the model’s reconstruction and rendering can be assessed, allowing for further adjustments to model parameters.

[0047] Step S110 : determining a scene model corresponding to a target cable tunnel based on the three-dimensional point cloud data and a plurality of virtual images.

[0048] In step S110 provided in the present application, a scene model corresponding to the target cable tunnel is determined.

[0049] This involves scene models, which are three-dimensional models that accurately represent the geometry, texture, color, and lighting effects of a specific scene. The target cable tunnel scene model not only includes the geometry of the cable tunnel, but also accurately represents environmental characteristics such as the distribution of cables inside, the texture of the tunnel walls, and lighting conditions.

[0050] In this step, 3D point cloud data directly obtained from the interior of the cable tunnel using technologies such as 3D laser scanning, as well as virtual images generated by simulating camera shots from different perspectives, are used to construct and optimize a 3D scene model that fully and realistically represents the target cable tunnel. The 3D point cloud data provides direct geometric information about the interior of the cable tunnel, ensuring high precision in the scene model's spatial positioning and geometric features. The use of virtual images allows the model to learn texture, color, and lighting information from different perspectives, further improving the visual realism and detailed expression of the scene model. This step integrates multiple data sources, enriching the data used for model training and enhancing the model's learning capabilities, enabling it to adapt to the complex and changing internal environments of cable tunnels.

[0051] Through the above steps S102-S106, the three-dimensional point cloud data corresponding to the target cable tunnel can be obtained; based on the three-dimensional point cloud data, the three-dimensional space model corresponding to the target cable tunnel can be determined; multiple virtual shooting poses corresponding to the three-dimensional space model can be determined; based on the three-dimensional point cloud data, virtual images corresponding to the multiple virtual shooting poses can be determined, wherein the corresponding virtual images are determined based on the projection of the three-dimensional point cloud data on the corresponding two-dimensional image coordinate system, and the multiple virtual shooting poses correspond one-to-one to the multiple two-dimensional image coordinate systems; based on the three-dimensional point cloud data and the multiple virtual images, the scene model corresponding to the target cable tunnel can be determined. By determining the virtual images corresponding to the multiple virtual shooting poses based on the three-dimensional point cloud data, the purpose of determining the scene model corresponding to the target cable tunnel based on the three-dimensional point cloud data and the multiple virtual images is achieved. Since the multiple virtual images can reflect the characteristics of objects inside the cable tunnel, the accuracy of the determined scene model is improved, thereby solving the technical problem in the related art that the accuracy of the constructed scene model is low due to the complex internal environment and single geometric shape of the cable tunnel.

[0052] As an optional embodiment, determining virtual images corresponding to multiple virtual shooting poses respectively based on three-dimensional point cloud data includes: determining adjustment parameters corresponding to the multiple virtual shooting poses respectively; determining multiple adjustment coordinates corresponding to the three-dimensional point cloud data based on the multiple adjustment parameters, wherein the multiple adjustment coordinates are coordinates corresponding to the three-dimensional point cloud data in multiple three-dimensional shooting coordinate systems, and the multiple three-dimensional shooting coordinate systems correspond one-to-one to the multiple virtual shooting poses; determining device parameters corresponding to the virtual shooting device; determining pixel coordinates corresponding to the multiple virtual shooting poses respectively based on the device parameters and the multiple adjustment coordinates, wherein the corresponding pixel coordinates are the projection coordinates of the corresponding adjustment coordinates on the corresponding two-dimensional image coordinate system; and determining virtual images corresponding to the multiple virtual shooting poses respectively based on the multiple pixel coordinates.

[0053] In this embodiment, specific steps of determining virtual images corresponding to a plurality of virtual shooting postures respectively based on three-dimensional point cloud data are described.

[0054] Among them, adjustment parameters are involved. Adjustment parameters refer to the parameters required when converting three-dimensional point cloud data from the original three-dimensional world coordinate system to the corresponding three-dimensional shooting coordinate system.

[0055] Among them, it involves adjusting coordinates, which refers to the coordinates of the three-dimensional point cloud data in the corresponding three-dimensional shooting coordinate system.

[0056] Among them, the three-dimensional shooting coordinate system is involved. The three-dimensional shooting coordinate system refers to the coordinate system used to represent and locate the camera position and direction in three-dimensional space. Each virtual shooting pose defines an independent three-dimensional shooting coordinate system.

[0057] Among them, device parameters are involved. Device parameters refer to the configuration information of the virtual shooting device, such as focal length, image size, image principal point position, etc. These parameters determine the imaging characteristics and quality of the virtual image.

[0058] Among them, pixel coordinates are involved. Pixel coordinates refer to the corresponding projection coordinates of the three-dimensional point cloud data projected onto the two-dimensional image coordinate system.

[0059] In this step, a series of adjustment parameters are first determined for each virtual shooting pose. These adjustment parameters are then used to determine the adjusted coordinates of the 3D point cloud data at each virtual shooting pose—that is, the position of the point cloud in the virtual camera coordinate system—to ensure accurate projection of each point from different perspectives. Next, the device parameters of the virtual shooting device, such as focal length and principal point position, are defined. These parameters affect the imaging effect of the virtual image and ensure image clarity and accuracy. Based on the device parameters and adjustment coordinates, the 3D point cloud data is projected onto the corresponding 2D image coordinate system, obtaining the pixel coordinates of each point in the virtual image and a visual representation of the point cloud distribution on the image. Finally, based on the multiple pixel coordinate information, virtual images corresponding to the multiple virtual shooting poses are generated. These images simulate the effect of observing the cable tunnel from different angles, providing rich perspective information for training and optimizing the 3D model.

[0060] Through this step, the projection coordinates of the three-dimensional point cloud data on the two-dimensional image coordinate system corresponding to multiple virtual shooting postures are determined. Based on the projection coordinates, the virtual images corresponding to the multiple virtual shooting postures are determined to ensure that the virtual images are closer to the real scene in terms of both visual effects and physical authenticity.

[0061] As an optional embodiment, the scene model corresponding to the target cable tunnel is determined based on three-dimensional point cloud data and multiple virtual images, including: determining multiple simulated object models based on the three-dimensional point cloud data, wherein the multiple simulated object models are used to represent the three-dimensional models corresponding to multiple target objects in the target cable tunnel; updating the model coefficients corresponding to the multiple simulated object models based on the multiple virtual images to obtain multiple updated models; and determining the scene model corresponding to the target cable tunnel based on the multiple updated models.

[0062] In this embodiment, specific steps of determining a scene model corresponding to a target cable tunnel based on three-dimensional point cloud data and a plurality of virtual images are described.

[0063] This involves simulation object models, which represent different objects in cable tunnels in three-dimensional space. These models are abstracted and generalized from point sets in 3D point cloud data and are used to more systematically represent and analyze various physical objects in cable tunnels.

[0064] Here, target objects are involved, and target objects refer to objects in the target cable tunnel, such as cables, tunnel walls, supporting structures, etc.

[0065] This involves model coefficients, which refer to the set of parameters used to describe the mathematical properties and physical characteristics of simulated object models in 3D reconstruction and rendering algorithms. These include, but are not limited to, coefficients for geometry, texture, color, and lighting response. These coefficients are updated during model training and optimization to improve model accuracy and realism.

[0066] Among them, the updated model is involved. The updated model refers to a model that updates the model coefficients through virtual image training on the basis of the simulated object model, which can reflect the fine structure, texture, color and other characteristics of the corresponding target object.

[0067] In this step, various objects within the cable tunnel, such as cables, tunnel walls, and support structures, are first identified based on 3D point cloud data. Then, an initial simulated object model is constructed for each identified target object, closely reproducing the target object's true geometry and surface features. Next, the model coefficients of the simulated object model are trained and adjusted using virtual images captured from multiple virtual camera positions. The virtual images provide information about the target object's appearance from various viewpoints, including color, texture, and lighting. Using a backpropagation algorithm, the model coefficients are continuously optimized to more accurately match the visual features of the target object in the virtual images. After multiple updates and optimizations of the model coefficients, the resulting updated models are integrated to construct a complete, highly accurate cable tunnel scene model. This model not only incorporates the geometry of the cable tunnel but also meticulously reproduces the texture, color, and lighting of objects within it. This model can generate realistic views from any viewpoint, supporting applications such as cable tunnel maintenance, planning, and safety assessment.

[0068] Through this step, using a simulated object model to represent each object in the cable tunnel can more accurately capture and express the details of the object, such as the arrangement of cables and the texture of the tunnel wall, thereby improving the realism and detail of the scene model.

[0069] It should be noted that by using three-dimensional point cloud data, the point cloud can be converted into a Gaussian ellipsoid to build the basic framework of the model and obtain a simulated object model. Then, through training and technical optimization, such as the use of deep learning algorithms, the resolution and realism of the model can be refined and improved.

[0070] As an optional embodiment, based on multiple virtual images, model coefficients corresponding to multiple simulated object models are updated to obtain multiple updated models, including: determining environmental adjustment parameters corresponding to multiple virtual shooting postures; based on multiple virtual images and the environmental adjustment parameters corresponding to the multiple virtual images, model coefficients corresponding to multiple simulated object models are updated to obtain multiple updated models.

[0071] In this embodiment, specific steps of updating model coefficients corresponding to a plurality of simulation object models respectively according to a plurality of virtual images to obtain a plurality of updated models are described.

[0072] This involves environmental adjustment parameters, which are a set of parameters used to adjust the environmental conditions used when generating virtual images, such as lighting, atmospheric effects, and color balance. These parameters can be considered additional conditions beyond the virtual shooting pose, used to simulate specific environments and lighting effects, enhancing the realism and quality of virtual images.

[0073] In this step, a set of environmental adjustment parameters is determined for each virtual shooting pose. These parameters simulate environmental factors such as lighting conditions, color balance, and atmospheric effects in the cable tunnel at that viewing angle. This step helps ensure that the generated virtual images reflect the real-world visual effects under different environmental conditions. The multiple generated virtual images and their corresponding environmental adjustment parameters are then used to train and optimize the model coefficients of the simulated object model. By comparing the virtual images with the real-world effects, the model coefficients are adjusted to more accurately match the geometry, texture, color, and lighting information in the images, minimizing the difference between the model-generated and virtual images. After multiple rounds of training and model coefficient updates, multiple updated simulated object models are generated. These models are closer to the real-world conditions of the cable tunnel in terms of geometry, texture, and lighting response, improving the accuracy and visual realism of the scene reconstruction.

[0074] Through this step, environmental adjustment parameters are introduced, allowing the simulated object model to simulate various lighting and environmental conditions during training, optimizing the detailed description of the simulated object model so that it can accurately reflect the appearance of the object under different lighting and environmental conditions, thereby improving the realism and detail performance of the scene model.

[0075] As an optional embodiment, the scene model corresponding to the target cable tunnel is determined based on multiple updated models, including: determining test images corresponding to multiple virtual shooting poses based on the multiple updated models, wherein the corresponding test images are images determined based on the projections of the multiple updated models on the corresponding two-dimensional image coordinate systems; determining an error value based on the multiple test images and the virtual images corresponding to the multiple test images; when the error value is lower than the error threshold, determining the scene model corresponding to the target cable tunnel based on the multiple updated models.

[0076] In this embodiment, specific steps of determining a scene model corresponding to a target cable tunnel based on a plurality of updated models are described.

[0077] This involves test images, which are images generated by projecting the updated model onto a 2D image coordinate system during the 3D reconstruction and model verification phase. These test images are generated to evaluate the rendering quality and accuracy of the updated model in different virtual shooting poses.

[0078] This involves error values, which are quantitative metrics used to measure the difference between a test image and a virtual image during model validation. Error values ​​can be pixel-level color differences or structural distortion, and are used to assess the reconstruction quality and visual realism of scene models.

[0079] This involves an error threshold, which is a standard set during model validation to determine whether the error value is acceptably low. When the error value is below the error threshold, the scene model is considered to meet the reconstruction quality requirements and can be used in practical applications.

[0080] In this step, first, for each updated model, the updated model is projected onto the corresponding two-dimensional image coordinate system according to the set virtual shooting pose to generate a series of test images. Then, the generated test image is compared with the virtual image previously used for training, and the difference between the two, namely the error value, is calculated. The calculation of the error value may involve multiple indicators, such as color difference, structural distortion, texture matching, etc., which is a key indicator for evaluating the quality of model reconstruction. When the error value is lower than the pre-set error threshold, it means that the performance of the updated model under multiple perspectives matches the virtual image well, and the scene reconstructed by the model is accurate and realistic enough. At this point, all updated models can be integrated to generate a complete target cable tunnel scene model, which can provide high-quality views from any perspective to meet the needs of cable tunnel management and operation and maintenance.

[0081] By comparing the differences between the test image and the virtual image in this step, the reconstruction accuracy and detail expression ability of the updated model can be effectively verified, ensuring that the scene model can accurately reflect the actual conditions of the cable tunnel, avoiding inaccuracies or unreality caused by overfitting or underfitting, and ensuring the high quality and reliability of the scene model.

[0082] As an optional embodiment, test images corresponding to multiple virtual shooting poses are determined based on multiple updated models, including: determining the model projections of the multiple updated models on the corresponding two-dimensional image coordinate system; determining the pixel information corresponding to the multiple model projections; and determining the test images corresponding to the multiple virtual shooting poses based on the pixel information corresponding to the multiple model projections.

[0083] In this embodiment, specific steps of determining test images corresponding to a plurality of virtual shooting poses respectively based on a plurality of updated models are described.

[0084] Among them, model projection is involved. Model projection refers to updating the projection of the model on the corresponding two-dimensional image coordinate system. Model projection takes into account the camera's intrinsic parameters (such as focal length, image size) and extrinsic parameters (i.e., virtual shooting posture), and converts the points and surfaces in the three-dimensional model into pixel information in the two-dimensional image coordinate system.

[0085] Among them, pixel information is involved. Pixel information refers to data representing the color, depth, texture and other information of each pixel in a two-dimensional image coordinate system.

[0086] In this step, first, for each updated model, the model is projected onto a corresponding two-dimensional image coordinate system based on its corresponding virtual shooting pose. The projection process involves a conversion from three-dimensional to two-dimensional, and it is necessary to consider the internal and external parameters of the camera to ensure the accurate presentation of the model from different perspectives. After the model projection is completed, the color, depth, and other attribute information of each pixel of the model projection are determined. This information constitutes the pixel data of the test image and is used for subsequent image generation. Finally, a test image is generated based on the pixel information of each model projection. These test images cover the effects of observing the target cable tunnel from multiple virtual perspectives, providing intuitive feedback information for model verification and optimization.

[0087] Through this step, test images of the updated model in the corresponding virtual shooting pose are generated. The generation of test images provides a circular feedback mechanism for model training. By comparing the differences between test images and virtual images or real images, the model parameters can be continuously adjusted, more refined iterative optimization can be performed, and the reconstruction quality of the model can be improved.

[0088] It should be noted that during the model projection phase, lighting and environmental simulation can be incorporated to more realistically reproduce the lighting conditions and environmental characteristics inside the cable tunnel, enhancing the visual realism of the test images. To improve the quality and detail of the test images, each updated model can be projected multiple times. Image fusion techniques (such as weighted averaging and multi-view stereo disparity correction) are then used to generate the final test image, ensuring image coherence and consistency across different viewpoints.

[0089] As an optional embodiment, before obtaining the three-dimensional point cloud data corresponding to the target cable tunnel, it also includes: obtaining scene data corresponding to the target cable tunnel, wherein the scene data includes initial point cloud data and a regional image; determining a plurality of target pixel points corresponding to the initial point cloud data from a plurality of initial pixel points included in the regional image; determining image information corresponding to the plurality of target pixel points, wherein the corresponding image information is used to represent pixel feature information of the regional image at the corresponding target pixel point; determining the three-dimensional point cloud data based on the plurality of image information and the initial point cloud data, wherein the plurality of data points in the three-dimensional point cloud data respectively carry corresponding image information.

[0090] In this embodiment, specific steps of determining three-dimensional point cloud data are described.

[0091] Among them, scene data is involved. Scene data refers to the original observation information of the target scene, such as point cloud data, image data, etc. It is the basis for subsequent 3D reconstruction, 3D point cloud generation and scene understanding.

[0092] This involves initial point cloud data, which refers to a collection of unprocessed point data acquired in the early stages of 3D reconstruction using technologies such as laser scanning. Each point contains coordinate information that describes the geometric features of the target cable tunnel.

[0093] This involves regional images, which are images captured in a specific area or from a specific perspective. They typically contain surface details and environmental information about the target cable tunnel. Regional images can come from panoramic cameras, pinhole cameras, or other imaging devices, and are used to supplement the texture and color information of point cloud data.

[0094] Among them, the initial pixel point is involved. The initial pixel point refers to the basic unit that constitutes the regional image. Each initial pixel point carries information such as color and texture on the image.

[0095] This involves target pixels, which are pixels selected from the initial pixel set and directly or indirectly associated with points in the initial point cloud data. Target pixels are used to assign color and texture to the point cloud data, ensuring the consistency of the geometric and textural features of the point cloud data.

[0096] Among them, image information is involved. Image information refers to the data used to represent pixel feature information on the target pixel point in three-dimensional reconstruction, such as color value, texture, lighting response, etc.

[0097] In this step, initial point cloud data of the target cable tunnel is first acquired using technologies such as laser scanning. Simultaneously, images of the area are captured or collected. These images may include surface details and environmental information of the cable tunnel, captured from various angles. From the collected regional images, image processing and analysis are used to identify target pixels that are directly or indirectly associated with points in the initial point cloud data. The selection of target pixels is often based on the geometric position of the point cloud and the pixel coordinates of the image, ensuring the correspondence between the point cloud data and the image information. Next, corresponding image information is extracted from each target pixel, including color values ​​and texture details, to enrich and refine the visual features of the initial point cloud data. Finally, the extracted image information is associated with the points in the initial point cloud data to generate three-dimensional point cloud data with additional attributes such as color and texture. This dataset not only contains the geometric information of the object but also incorporates surface texture and color information, providing a more comprehensive description of the detailed features of the target cable tunnel.

[0098] This step associates the image information with the initial point cloud data, generating a 3D point cloud that contains rich color and texture information. This helps improve the accuracy of the point cloud data's detailed representation, enabling subsequent 3D reconstruction and new perspective synthesis to present more realistic visuals. This is particularly true in cable tunnels, where lighting conditions are complex and texture features are limited. Image information can complement the point cloud data and enhance the integrity and authenticity of the 3D reconstruction. Furthermore, directly integrating image information into the point cloud data simplifies the data processing for 3D reconstruction and new perspective synthesis, avoiding complex image matching and feature extraction, and reducing the technical complexity and computational cost of 3D reconstruction.

[0099] Based on the above embodiment and optional embodiment, an optional implementation manner is provided, which is described in detail below.

[0100] There are many problems in the 3D reconstruction of underground spaces, especially cable tunnel scenes. First, the surface materials inside cable tunnels are often relatively simple, such as concrete or metal. The textures of these materials are relatively simple and lack sufficient diversity. This characteristic makes it difficult to obtain feature information with sufficient discrimination when extracting and matching texture features. Secondly, the overall structure inside cable tunnels is relatively simple, usually presenting a regular circular or rectangular structure. This lack of diversity in geometric features further increases the difficulty of feature extraction and matching. During the reconstruction process, due to the scarcity of significant feature points, various algorithms often find it difficult to accurately locate and match corresponding points, thus affecting the accuracy and reliability of 3D reconstruction.

[0101] Among traditional 3D reconstruction techniques, structure from motion (SFM) and multi-view stereo (MVS) undoubtedly play a pivotal role. SFM extracts and matches feature points from image sequences to reconstruct a sparse point cloud and camera parameters, while MVS leverages image information from multiple perspectives to precisely identify matching relationships between images, further generating dense point cloud data. Both methods rely on feature matching and geometric principles, resulting in excellent performance in scenes with rich textures, dense views, and high overlap. However, when faced with scenes such as cable tunnels with monotonous textures and geometric features, complex lighting, and confined spaces, feature point extraction becomes extremely difficult, and the acquisition of geometric structure information is incomplete. This significantly reduces computational accuracy, and the reconstruction results are often unsatisfactory, manifesting in numerous defects such as incomplete structure, severe distortion, frequent holes, and blurred colors.

[0102] With the continuous advancement of 3D laser scanning technology, efficiently and accurately acquiring 3D models of objects or complex scenes has become a reality. Due to its superior performance, 3D laser scanning technology is capable of capturing information on virtually all objects and scenes, with the exception of certain challenging objects such as glass, water, and specular surfaces. This ensures the high geometric accuracy and integrity of the resulting 3D models. Notably, laser point cloud technology not only captures the geometric structure of objects in detail but also effectively records and expresses color information, enabling the precise reproduction of the scene's appearance and details. Laser point cloud technology demonstrates its unique advantages in modeling cable tunnel interiors. It can comprehensively and meticulously depict the geometry and texture characteristics of cable tunnel interiors, providing a high-precision, high-fidelity data foundation for subsequent engineering analysis, maintenance management, and virtual reality applications. However, while laser point cloud data offers significant advantages in representing complex scenes, its inherent unstructured, discrete point format poses challenges for realistic rendering. This data format lacks intuitiveness and is not conducive to direct observation and understanding by the human eye, thus limiting the widespread application of three-dimensional point cloud data (3D point cloud data) in visualization and interactive applications.

[0103] Compared with the above methods, Neural Radiance Field (NeRF) and subsequent related studies have achieved remarkable results in new view synthesis. These methods have the characteristics of continuity, differentiability and smoothness. However, due to its high computational complexity, both training and inference time are long. Studies have optimized the training and inference speed of NeRF and improved memory efficiency by dividing the space into grids for calculation. Some studies have also added deep priors to solve the geometric ambiguity problem under sparse view input, improved reconstruction quality, optimized the NeRF algorithm, expanded the scale of the research scene, and attempted to reconstruct large-scale scenes in reality. Overall, NeRF is a computationally intensive algorithm with long training time, high computational complexity, strict requirements on shooting conditions, and difficulty in balancing rendering quality, scene scale, training and inference speed. It is difficult to achieve real-time high-quality rendering, which is not suitable for 3D reconstruction of cable tunnels.

[0104] 3D Gaussian Rendering (3DGS) utilizes anisotropic 3D Gaussian ellipsoids (3D Gaussian ellipsoids) to explicitly represent scenes. Compared to NeRF, it demonstrates superior rendering performance and faster training and inference speeds. Further research has improved its ability to address over-reconstruction issues caused by Gaussian densification, quality degradation caused by ambiguous input training views, and aliasing issues when changing the sampling rate. Building on this, efforts have also been made to reconstruct dynamic scenes and scale them. Due to its advantages in real-time and high-quality rendering, 3DGS has also achieved remarkable results in several downstream applications, such as 3D surface mesh generation (3D surface mesh generation) and 3D scene editing (3D scene editing). While these methods have certain advantages, most of them often encounter several problems in practical applications. First, they heavily rely on dense views for training. Although effective rendering can be achieved from training views with sparse or unevenly distributed view inputs, rendering quality from new views becomes unstable, resulting in geometric ambiguity and image blur. This instability stems from the limited training views and the inability of the model to fully learn the structure and distribution of the global scene, as well as the complexity of the nonlinear optimization process, which may cause it to fit feature details, noise, and eventually fall into local optimality or overfitting. Secondly, most methods rely on SFM to obtain camera poses and sparse point clouds. However, for many real-world scenes, especially the interiors of cable tunnels, there are usually weak textures, monotonous geometric features, mutually occluded structures, and narrow or strip-shaped spatial distributions. This complexity poses a considerable obstacle to SFM calculations, resulting in a significant reduction in the accuracy of the initial data. Therefore, these challenges make it difficult for the reconstruction results to meet the strict standards required for high-quality rendering of new views of the scene with six degrees of freedom (6-DOF).

[0105] In view of this, an optional embodiment of the present invention provides a cable tunnel 3DGS real scene generation technology based on a mobile scanning system. Figure 2 This is a flow chart of scene model construction provided by an optional embodiment of the present invention, such as Figure 2 As shown, a virtual image can be determined based on the point cloud data, and a detailed model can be determined based on the virtual image. This can solve the problems of the 3DGS new view synthesis technology's dependence on dense input views, and the initial input data being subject to various limitations of SFM calculation accuracy and scene conditions. It can also improve the robustness of the algorithm in rendering new perspectives in cable tunnel scenes.

[0106] The following describes in detail the method steps provided by the optional embodiment of the present invention.

[0107] S1. Obtain the three-dimensional point cloud data corresponding to the target cable tunnel.

[0108] A1. Obtain scene data corresponding to the target cable tunnel.

[0109] 1) Multimodal Data Acquisition. A mobile laser SLAM scanning system was used to acquire a 3D laser point cloud (similar to the initial point cloud data described above) and a panoramic image (similar to the regional image described above) of the cable tunnel interior scene. After preprocessing, the required initial color point cloud, pinhole camera image, and corresponding camera parameters were obtained. This process specifically involves the following steps.

[0110] A mobile laser scanning system was used for data collection. Prior to data collection, the route was planned based on the actual conditions of the cable tunnel. After sensor calibration, data collection began, acquiring laser point clouds and panoramic images of the cable tunnel interior.

[0111] The data acquisition equipment is a lightweight mobile laser SLAM scanning system equipped with a multi-line laser radar and a 360-degree panoramic camera. The acquisition method is manual backpack walking scanning. This scanning method is suitable for high-efficiency data acquisition in complex scenes such as the interior of cable tunnels, and can obtain high-quality laser point clouds. Before scanning, ensure that the sensors of the equipment are calibrated. Before the operation, it is necessary to plan the scanning route reasonably to ensure full coverage and overlap of the scanning area. The planning should take into account factors such as the single scanning route and the characteristics of the interior of the cable tunnel to avoid omissions. Scan according to the planned route to ensure that intersections with the previous route are generated during the scanning process to improve mapping accuracy. Set a reasonable image acquisition point interval according to the conditions of the scene. During the scanning process, the scanning path is displayed through the touch screen controller, and the scanning progress and completeness are displayed in real time.

[0112] 2) Perform pre-processing such as denoising, filtering, and registration on the point cloud to obtain a point cloud with consistent global coordinates.

[0113] During the point cloud data acquisition process, noise is generated due to factors such as equipment accuracy, environmental factors, and human disturbance. This noise can affect the accuracy of the point cloud and the effectiveness of subsequent processing, necessitating denoising. Common point cloud denoising methods include bilateral filtering, Gaussian filtering, and statistical filtering. The appropriate denoising method should be selected based on the characteristics of the scene and data.

[0114] If the same scene is captured multiple times, point cloud registration is required. Multiple scans of the target must have a certain degree of regional overlap. This is achieved by finding points with the same name in the overlapping region and stitching them together. The identification of feature points in the overlapping region is directly related to the quality of the registration results. Using the correct correspondence, the rigid body transformation is estimated to produce a globally consistent point cloud.

[0115] 3) Use spherical interpolation to obtain the pose of the panoramic image, and project the panoramic image into images of 6 lenses according to the pinhole camera model as the original image for training.

[0116] The camera and laser were calibrated before operation. The transformation relationship between panoramic images at different perspectives or positions can be calculated through spherical linear interpolation of quaternions, thereby obtaining the position and posture of each panoramic image, including the rotation matrix and translation vector.

[0117] Training requires undistorted pinhole camera images. Panoramic images must be projected according to the pinhole camera model. This method customizes the focal length, image size, and principal point position of the pinhole camera model. After unfolding the panoramic image plane, projection is performed in six directions: front, back, left, right, up, and down, based on the scanning direction.

[0118] 4) Assign three primary colors (RGB color values) to the laser point cloud according to the panoramic image and posture to obtain a color point cloud (same as the above three-dimensional point cloud data).

[0119] The laser point cloud is transformed based on the image's pose information and projected onto the image plane frame by frame. The corresponding pixel of the laser point cloud projection point is found on the image plane. The RGB color value is extracted from the found pixel and assigned to the corresponding point in the laser point cloud. Color smoothing is performed on the colored point cloud to reduce color noise and improve color consistency.

[0120] S2. Determine the three-dimensional spatial model corresponding to the target cable tunnel based on the three-dimensional point cloud data.

[0121] According to all point cloud data points (p i =[x i ,y i ,z i ,r i ,g i ,b i ] T ,pi ∈P), determine the three-dimensional space range of the scene (V according to the maximum and minimum values ​​of the point cloud data point coordinates 3 , the same as the above three-dimensional space model) is:

[0122] (V 3 =X∈R 3 :p x_min ≤x≤p x_max ,p y_min ≤y≤p y_max ,p z_min ≤z≤p z_max )

[0123] Among them, p i is the data point in the three-dimensional point cloud data, P is the total number of data points, [x i ,y i ,z i ] is the three-dimensional coordinate value of the corresponding data point, [r i ,g i ,b i ] is the color parameter of the corresponding data point, R is a real number set, p x_min 、p y_min and p z_min They are the minimum coordinate values ​​of all data points in the three-dimensional point cloud data on the corresponding three-dimensional coordinate axis, p x_max 、p y_max and p z_max They are the maximum coordinate values ​​of all data points in the three-dimensional point cloud data on the corresponding three-dimensional coordinate axes, and [x, y, z] are the coordinates of the three-dimensional space model corresponding to the target cable tunnel.

[0124] S3. Determine multiple virtual shooting poses corresponding to the three-dimensional space model.

[0125] For the official 3DGS, two factors significantly impact reconstruction quality. The first is the accuracy of the initial data. Research and analysis suggest that the input colored point cloud is a rough approximation of the scene's true distribution, representing global low-frequency information. This information is fundamental to the entire optimization process, playing a crucial role in guiding the learning of high-frequency information and preventing training from prematurely converging to local minima. The second factor is the necessity of dense input views. Given the complexity of nonlinear optimization processes, sparse view input provides limited information, which can easily lead to overfitting or convergence to local minima during training.

[0126] An optional embodiment of the present invention addresses these two factors by using a laser point cloud instead of the SFM sparse point cloud as the initial Gaussian input. This avoids the poor SFM calculation accuracy caused by monotonous geometry and texture in cable tunnel scenes, thereby improving the accuracy of the input data. The use of virtual views in pre-training meets the 3DGS's requirement for dense training views, enhancing the model's learning of the scene's global information.

[0127] Due to the narrow and long characteristics of the tunnel, in order to facilitate processing, this space range V 3 Divide into multiple sub-areas, and randomly generate a large number of virtual poses (V R_poses ) to ensure that the generated poses are continuous in space. For each pose, a bounding box algorithm is used to perform point cloud collision detection, and the rationality of the pose orientation and height is judged and filtered to retain the poses that meet the requirements. The formula for screening virtual poses is:

[0128] Filter(V R_poses )→V R_poses′

[0129] Among them, Filter() is the filtering function, V R_poses′ is the final determined virtual shooting pose.

[0130] S4. Determine virtual images corresponding to the plurality of virtual shooting postures according to the three-dimensional point cloud data.

[0131] For each point p in the point cloud set P i =[x i ,y i ,z i ,r i ,g i ,b i ] T , according to the pose and camera intrinsic parameter matrix, the point cloud is projected in the corresponding camera coordinate system (the same as the above three-dimensional shooting coordinate system) and the adjusted coordinates (p c )for:

[0132] p c =R v ·p i +T v

[0133] Among them, R v is the rotation matrix of the virtual pose, T v is the translation vector of the virtual pose, R v 、T v constitutes the above adjustment parameters.

[0134] In the virtual pose V R_poses′ On the corresponding image, the pixel coordinates u corresponding to the point cloudi for:

[0135] u i =K·p c

[0136] Where K is the device parameter.

[0137] Assign pixels by pixel coordinates and point cloud colors. The color assignment formula is:

[0138] u i [r,g,b]=[r i ,g i ,b i ]

[0139] Among them, u i [r,g,b] is the color parameter corresponding to the pixel.

[0140] When multiple point clouds are projected onto the same pixel, the closest point is selected based on the point cloud depth. Blank pixels in the image are processed using interpolation. The image is then smoothed, denoised, and filtered.

[0141] S5. Determine a scene model corresponding to the target cable tunnel based on the three-dimensional point cloud data and the multiple virtual images.

[0142] Figure 3 This is a flow chart of a training update model provided by an optional embodiment of the present invention. Figure 3 As shown, based on the virtual image, the initial Gaussian model is continuously trained. This step is described in detail below.

[0143] B1. Determine multiple simulation object models based on 3D point cloud data.

[0144] The color laser point cloud is initialized as a three-dimensional Gaussian ellipsoid (the same as the simulated object model mentioned above) to obtain high-precision initial data. The model is pre-trained using the virtual image projected by the color point cloud as the true value to increase the diversity of the training views, fully learn the global information of the scene, and obtain a rough Gaussian model.

[0145] The initial value of the mean of each three-dimensional Gaussian ellipsoid is the coordinate of the three-dimensional laser point cloud, and the initial value of the mean of the three-dimensional Gaussian ellipsoid is μ i for:

[0146] μi=p i (x i ,y i ,z i )

[0147] Gaussian function G i (x) is:

[0148]

[0149] Among Gaussian's third-order spherical harmonic coefficients, the initial value of the first-order spherical harmonic coefficient is the color of the laser point cloud, and the initial values ​​of other coefficients are 0.

[0150] The covariance is expressed by the scaling coefficient and rotation quaternion of the three axes of the ellipsoid. The initial value of the scaling coefficient is obtained by calculating the average distance between the point cloud and its neighboring point clouds, and the initial value of the rotation quaternion is 1, 0, 0, 0.

[0151] Opacity is obtained by inputting the inverse of the activation function (sigmoid) with a value of 0.1.

[0152] B2. Based on the multiple virtual images, update the model coefficients corresponding to the multiple simulation object models to obtain multiple updated models.

[0153] Using the above virtual image V R_poses′ As the true value, the three-dimensional Gaussian ellipsoid is trained. During the training process, the mean value μ of the Gaussian i It never participates in the update, only updates the Gaussian covariance, spherical harmonic coefficients, opacity and other properties.

[0154] The process of projecting the three-dimensional Gaussian onto the two-dimensional plane during the training process is as follows: Assume that in the world coordinate system, the mean of the three-dimensional Gaussian is μ W , the covariance matrix is ​​∑ W , the transformation from the world coordinate system to the camera coordinate system is linear, then in the camera coordinate system, the mean μ of the three-dimensional Gaussian is C for:

[0155] μ C =R′·μ W +T

[0156] Covariance matrix ∑ C for:

[0157] ∑ C =R′·∑ W ·R T

[0158] Among them, R′ and T are the rotation matrix and translation vector from the world coordinate system to the camera coordinate system, respectively.

[0159] The transformation from the camera coordinate system to the pixel coordinate system is a nonlinear transformation, and the pixel coordinate variable U is:

[0160] U=K·Y

[0161] Among them, Y is the camera coordinate variable and K is the intrinsic parameter transformation matrix.

[0162] Using Taylor's formula to expand at the mean, the pixel coordinate variable U is:

[0163] U≈F(μ C )+J·(Y-μ C )=K·μ C +J·(Y-μ C )

[0164] Among them, F() is the projection transformation function from the camera coordinate system to the pixel coordinate system, and J is the Jacobian matrix of the nonlinear mapping.

[0165] Then, calculate the mean of the pixel plane:

[0166] μ U =K·μ C =K·(R′·μ W +T)

[0167] The covariance matrix is:

[0168] ∑ U =J·∑C·J T =J·R′·∑ W ·R′ T ·J T

[0169] Since the plane formed by the X and Y axes of the camera coordinate system is parallel to the pixel plane, the first two rows and columns of the matrix are the covariance matrix of the two-dimensional Gaussian on the pixel plane.

[0170] When updating the spherical harmonic coefficients, the initial values ​​of the first-order spherical harmonic coefficients are directly related to the RGB values ​​of the color point cloud. Therefore, during the pre-training phase, the first-order coefficients are prioritized for training and updating. As the number of iterations increases, the spherical harmonic coefficients gradually increase to 16 coefficients of the third order, expressing color values ​​in different directions on the Gaussian sphere. The use of large-scale virtual images as training images enables the Gaussian to learn richer feature information about the scene. The pre-trained Gaussian model with a rough output has a certain degree of generalization in scene representation and good geometric form, capable of rendering good results from all reasonable perspectives, but the rendering quality is relatively rough.

[0171] The Gaussian rendering process is as follows:

[0172] 3DGS performs raster rendering by projecting onto a two-dimensional image. The opacity of the pixels at the projection is calculated using a Gaussian function. The opacity σ' of the projection corresponding to the i-th Gaussian model at position U is:

[0173]

[0174] Among them, σ is the opacity of the Gaussian model at position U, μ U is the mean of the Gaussian projection onto the two-dimensional plane, ∑ U is the covariance matrix of the two-dimensional Gaussian.

[0175] The final color C rendered at pixel position U U for:

[0176]

[0177] Among them, the image pixel U has N Gaussians for raster rendering, c i is the color of the i-th Gaussian.

[0178] B3. Determine test images corresponding to the multiple virtual shooting poses based on the multiple updated models, and determine error values ​​based on the multiple test images and the virtual images corresponding to the multiple test images.

[0179] The error value is determined by determining the loss function, and the loss function still uses L1 loss and image structure similarity difference L D-SSIM The total loss function L is:

[0180] L=(1-λ)L1+λL D-SSIM

[0181] Where λ = 0.2, and L1 is the average of the absolute differences in pixel values ​​between the model output image and the target real image.

[0182] B4. When the error value is lower than the error threshold, determine the scene model corresponding to the target cable tunnel based on the multiple updated models.

[0183] The rough Gaussian model is read and trained and fine-tuned using images of the pinhole camera model projected from panoramic images to learn the detailed information of the scene.

[0184] Due to the limited training data, the learning rates of all parameters were kept low during training to avoid destroying the global scene information learned during pre-training. To maintain the stability of the Gaussian ellipsoid's color representation during fine-tuning, a small initial learning rate of 0.001 was set, and a step-wise decay strategy was used to adjust the learning rate to ensure stable convergence of color attributes and avoid overfitting.

[0185] The projection, rendering, and training loss functions of the 3D Gaussian during fine-tuning are the same as those described in step B2.

[0186] The lighting conditions inside the cable tunnel are relatively complex due to the closed space, artificial lighting, etc., which often cause geometric ambiguity in the synthesis of new 3DGS views, specifically manifested as floating objects and artifacts in the air. In the present invention, a lightweight attention-based sequence model (Transformer network) is added to capture the continuity and differential properties of the light distribution inside the cable tunnel, and to establish a connection between the camera pose and the scene light changes to eliminate the negative impact of light changes. The essence of appearance information is a mapping relationship. The image pose is input into the network to obtain the corresponding appearance adjustment vector. During Gaussian rasterization rendering, the appearance adjustment vector is a learned parameter used to adjust the spherical harmonic coefficients stored in the Gaussian to obtain new spherical harmonic coefficients related to the camera extrinsic parameters. The new spherical harmonic coefficients are involved in calculating the color of the Gaussian.

[0187] After training and learning in the above steps, the obtained model can render high-quality 6-DOF perspective images inside the cable tunnel.

[0188] Through the above optional implementation, at least the following beneficial effects can be achieved:

[0189] (1) Initialize the laser point cloud as a 3D Gaussian ellipsoid, fix the coordinates of the point cloud during training, and use a partial parameter optimization method. This eliminates the reliance on SFM to obtain initial data, and does not affect the accuracy of the initial data due to weak textures and geometric monotony in the cable tunnel scene. At the same time, the accurate geometric structure is also ensured during the optimization process.

[0190] (2) Based on the step-by-step training of 3DGS, the global information and details of the scene are learned successively. The use of virtual views satisfies the strict requirements of 3DGS for dense training views, and enhances the robustness of the reconstruction process without increasing the time cost and difficulty of the operation acquisition;

[0191] (3) Since the continuous and subtle changes of light in the interior space of the cable tunnel affect the view synthesis, prompting the Gaussian to fit the changes of light and produce floating artifacts in the air, a lightweight appearance model is added to the framework to express the smooth and continuous changes of brightness, thereby reducing the floating artifacts caused by these effects.

[0192] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0193] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.

[0194] Example 2

[0195] According to an embodiment of the present invention, a device for implementing the above-mentioned method for determining the scene model of the cable tunnel is also provided. Figure 4 is a structural block diagram of a device for determining a scene model of a cable tunnel according to an embodiment of the present invention, such as Figure 4 As shown, the device includes: an acquisition module 402, a first determination module 404, a second determination module 406, a third determination module 408 and a fourth determination module 410. The device will be described in detail below.

[0196] An acquisition module 402 is used to acquire three-dimensional point cloud data corresponding to the target cable tunnel; a first determination module 404 is connected to the above-mentioned acquisition module 402, and is used to determine the three-dimensional space model corresponding to the target cable tunnel based on the three-dimensional point cloud data; a second determination module 406 is connected to the above-mentioned first determination module 404, and is used to determine multiple virtual shooting poses corresponding to the three-dimensional space model; a third determination module 408 is connected to the above-mentioned second determination module 406, and is used to determine virtual images corresponding to the multiple virtual shooting poses based on the three-dimensional point cloud data, wherein the corresponding virtual images are determined based on the projection of the three-dimensional point cloud data on the corresponding two-dimensional image coordinate system, and the multiple virtual shooting poses correspond one-to-one to the multiple two-dimensional image coordinate systems; a fourth determination module 410 is connected to the above-mentioned third determination module 408, and is used to determine the scene model corresponding to the target cable tunnel based on the three-dimensional point cloud data and the multiple virtual images.

[0197] It should be noted here that the above-mentioned acquisition module 402, first determination module 404, second determination module 406, third determination module 408 and fourth determination module 410 correspond to steps S102 to S110 in the method for determining the scenario model for implementing a cable tunnel. The instances and application scenarios implemented by the multiple modules and corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned embodiment 1.

[0198] Example 3

[0199] According to another aspect of an embodiment of the present invention, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions, wherein the processor is configured to execute the instructions to implement any of the above methods for determining a scene model of a cable tunnel.

[0200] Example 4

[0201] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute any of the above methods for determining a scene model of a cable tunnel.

[0202] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0203] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0204] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0205] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0206] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0207] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.

[0208] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A method for determining a scene model of a cable tunnel, characterized in that: include: Obtain the three-dimensional point cloud data corresponding to the target cable tunnel; Determining a three-dimensional spatial model corresponding to the target cable tunnel based on the three-dimensional point cloud data; Determining a plurality of virtual shooting poses corresponding to the three-dimensional space model; Determining, based on the three-dimensional point cloud data, virtual images corresponding to the plurality of virtual shooting poses, respectively, wherein the corresponding virtual images are determined based on projections of the three-dimensional point cloud data onto corresponding two-dimensional image coordinate systems, and the plurality of virtual shooting poses correspond one-to-one to the plurality of two-dimensional image coordinate systems; A scene model corresponding to the target cable tunnel is determined based on the three-dimensional point cloud data and the multiple virtual images.

2. The method according to claim 1, characterized in that Determining virtual images corresponding to the plurality of virtual shooting postures respectively based on the three-dimensional point cloud data includes: Determining adjustment parameters corresponding to the plurality of virtual shooting postures respectively; Determining, based on a plurality of adjustment parameters, a plurality of adjustment coordinates corresponding to the three-dimensional point cloud data, wherein the plurality of adjustment coordinates are coordinates corresponding to the three-dimensional point cloud data in a plurality of three-dimensional shooting coordinate systems, and the plurality of three-dimensional shooting coordinate systems correspond one-to-one to the plurality of virtual shooting poses; Determine device parameters corresponding to the virtual shooting device; Determining pixel coordinates corresponding to the plurality of virtual shooting poses respectively based on the device parameters and the plurality of adjustment coordinates, wherein the corresponding pixel coordinates are projection coordinates of the corresponding adjustment coordinates on the corresponding two-dimensional image coordinate system; According to the plurality of pixel coordinates, virtual images corresponding to the plurality of virtual shooting postures are determined.

3. The method according to claim 1, characterized in that Determining a scene model corresponding to the target cable tunnel based on the three-dimensional point cloud data and the multiple virtual images includes: Determining a plurality of simulated object models based on the three-dimensional point cloud data, wherein the plurality of simulated object models are used to represent three-dimensional models corresponding to a plurality of target objects in the target cable tunnel; updating model coefficients corresponding to the plurality of simulation object models respectively according to the plurality of virtual images to obtain a plurality of updated models; A scene model corresponding to the target cable tunnel is determined according to the multiple updated models.

4. The method according to claim 3, characterized in that The updating of model coefficients corresponding to the plurality of simulation object models respectively based on the plurality of virtual images to obtain a plurality of updated models includes: Determining environment adjustment parameters corresponding to the plurality of virtual shooting postures respectively; According to the multiple virtual images and the environment adjustment parameters respectively corresponding to the multiple virtual images, the model coefficients respectively corresponding to the multiple simulation object models are updated to obtain the multiple updated models.

5. The method according to claim 3, characterized in that Determining the scene model corresponding to the target cable tunnel based on the multiple updated models includes: Determining, based on the multiple updated models, test images corresponding to the multiple virtual shooting poses, respectively, wherein the corresponding test images are images determined based on projections of the multiple updated models onto the corresponding two-dimensional image coordinate systems; determining an error value based on a plurality of test images and virtual images corresponding to the plurality of test images; When the error value is lower than an error threshold, a scene model corresponding to the target cable tunnel is determined according to the multiple updated models.

6. The method according to claim 5, characterized in that The determining, based on the multiple updated models, test images corresponding to the multiple virtual shooting poses, respectively, includes: Determining model projections of the multiple updated models on the corresponding two-dimensional image coordinate system; Determine pixel information corresponding to each of the multiple model projections; Determine test images corresponding to the multiple virtual shooting postures respectively according to pixel information respectively corresponding to the multiple model projections.

7. The method according to any one of claims 1 to 6, characterized in that Before obtaining the three-dimensional point cloud data corresponding to the target cable tunnel, the method further includes: Acquiring scene data corresponding to the target cable tunnel, wherein the scene data includes initial point cloud data and a regional image; Determining a plurality of target pixel points corresponding to the initial point cloud data from a plurality of initial pixel points included in the regional image; Determining image information corresponding to each of the plurality of target pixel points, wherein the corresponding image information is used to represent pixel feature information of the regional image at the corresponding target pixel point; The three-dimensional point cloud data is determined based on a plurality of image information and the initial point cloud data, wherein a plurality of data points in the three-dimensional point cloud data respectively carry corresponding image information.

8. A device for determining a scene model of a cable tunnel, characterized in that: include: An acquisition module is used to obtain three-dimensional point cloud data corresponding to the target cable tunnel; A first determining module is configured to determine a three-dimensional space model corresponding to the target cable tunnel based on the three-dimensional point cloud data; A second determining module is used to determine a plurality of virtual shooting poses corresponding to the three-dimensional space model; a third determining module, configured to determine, based on the three-dimensional point cloud data, virtual images corresponding to the plurality of virtual shooting poses, respectively, wherein the corresponding virtual images are determined based on projections of the three-dimensional point cloud data onto corresponding two-dimensional image coordinate systems, and the plurality of virtual shooting poses correspond one-to-one to the plurality of two-dimensional image coordinate systems; The fourth determining module is used to determine the scene model corresponding to the target cable tunnel based on the three-dimensional point cloud data and multiple virtual images.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method for determining the scene model of the cable tunnel according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method for determining a scene model of a cable tunnel according to any one of claims 1 to 7.