Three-dimensional reconstruction method and system for pollution insulator of monocular depth-sensor-free
Through the three-dimensional reconstruction method of a single-dimensional insulator without depth sensor, the three-dimensional reconstruction is performed using a monocular image and image synthesis model, which solves the problems of high operation difficulty, high cost and insufficient precision in the prior art, and realizes high-precision three-dimensional modeling and fault judgment.
Patent Information
- Application Number
- CN202510356241.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-01
AI Technical Summary
The existing three-dimensional reconstruction method requires additional lidar, depth sensors and multiple photos, which makes operational and maintenance personnel difficult, costly and insufficiently accurate.
The three-dimensional reconstruction method of a monocular depth-free sensor is adopted to achieve high-precision three-dimensional modeling of insulators by acquiring monocular images, subject extraction, multi-view image prediction and image synthesis models.
No depth sensors and multi-view image shooting are required, which reduces the operational difficulty and cost of operation and maintenance personnel, and realizes high-precision three-dimensional modeling of filthy insulators, which can clearly observe the filthy location and degree of accumulation, and assists operation and maintenance personnel to determine the fault status.
Smart Images

Figure CN120236013A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional reconstruction, and particularly to a three-dimensional reconstruction method and system for contaminated insulators without a depth sensor using a single camera. Background Art
[0002] To ensure the stability of the power system, regular operation and maintenance of insulators is of great significance. Since the operation and maintenance of contaminated insulators have long relied on visual methods such as manual inspection and drone inspection, and faults are monitored through image recognition methods, there are problems such as unclear fault information and undefined fault severity.
[0003] The detection method of contaminated insulators based on image recognition mainly relies on machine learning. For example, patent application CN116071368A discloses a method and device for detecting multi-angle images of insulator contamination and fineness analysis. The method includes: obtaining multi-angle real-time images and insulator parameter information, unifying the illumination intensity of the multi-angle real-time images and matching them with the insulator parameter information to obtain a preliminary model of multi-angle insulator contamination, complementing the preliminary model according to prior information to obtain an imaging model, obtaining the reflectivity of the insulator disk surface according to the color information in the imaging model, and dividing the contamination levels of various parts of the insulator surface according to the reflectivity and prior information. This method is not accurate enough for some cases and is also difficult to predict the subsequent development of faults. Through the three-dimensional reconstruction method, a three-dimensional visualization model of the contaminated insulator can be obtained, and the contamination position and contamination degree of the contaminated insulator can be observed more clearly, assisting maintenance personnel in making a more accurate judgment on the fault state of the insulator.
[0004] Currently, related three-dimensional reconstruction methods often have problems such as the need for additional lidar, depth sensors, and multiple photos, which have high requirements for maintenance personnel, drone inspection line algorithms, etc. Summary of the Invention
[0005] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a three-dimensional reconstruction method and system for contaminated insulators without a depth sensor using a single camera, which does not require a depth sensor and does not require a large number of multi-view images of insulators to be taken, reducing the operation difficulty of maintenance personnel, reducing the number of sensors, reducing costs, and achieving high-precision three-dimensional modeling of insulators.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] A three-dimensional reconstruction method for contaminated insulators without a depth sensor using a single camera includes the following steps:
[0008] Obtain a single-view image of a contaminated insulator in an arbitrary pose;
[0009] Perform body extraction on the monocular image of the contaminated insulator to obtain a background-free contaminated insulator image;
[0010] Perform multi-view image prediction on the background-free contaminated insulator image according to the implicit transformation through a multi-view image prediction model to obtain a multi-view contaminated insulator image;
[0011] Perform 3D reconstruction on the multi-view contaminated insulator image through an image synthesis model to obtain a 3D model of the contaminated insulator.
[0012] Further, the specific steps for performing body extraction on the monocular image of the contaminated insulator include:
[0013] Extract features from the monocular image of the contaminated insulator to obtain key information in the image;
[0014] Obtain the distance from each pixel point in the monocular image of the contaminated insulator to the camera through depth estimation;
[0015] Identify the main object and its position in the monocular image of the contaminated insulator through object detection;
[0016] Generate a body mask corresponding to the main body and its position in the monocular image of the contaminated insulator;
[0017] Use morphological operations to remove noise and holes and improve the quality of the body mask;
[0018] Segment the main body from the background according to the body mask through image segmentation technology to obtain a background-free contaminated insulator image.
[0019] Further, the morphological operations include multiple ones among erosion, dilation, opening operation, and closing operation.
[0020] Further, the implicit transformation is:
[0021]
[0022] In the formula, is the predicted image, f is the implicit transformation, x is the input RGB image, R is the rotation angle of the camera, and T is the position of the camera.
[0023] Further, the implicit transformation is implemented based on a diffusion model, and the loss function of the diffusion model is:
[0024]
[0025] In the formula, E Z~ε(x),t,∈~N(0,1)To optimize the expectation, that is, to minimize the average of the squared L2 norm between the predicted noise and the actual noise. ∈ is the noise variable, ∈~N(0,1) indicates that the noise variable ∈ follows the standard normal distribution. zt is the latent variable, t is the frame, x is the input RGB image, R is the rotation angle of the camera, T is the position of the camera, θ is the parameter to be optimized by the denoising network, and c(x,R,T) is the embedding obtained by fusing the input image and the relative camera extrinsic parameters (R,T), which is used as the conditional information of the diffusion model to guide the diffusion model to generate images corresponding to the perspective.
[0026] Furthermore, the image synthesis model is a 3D model, and its training process includes the NeRf training stage and the mesh training stage.
[0027] Furthermore, in the NeRf training stage of the image synthesis model, image loss and mask loss are adopted, and the loss function is:
[0028]
[0029] In the formula, is the rendered image of the i-th image, is the real image of the i-th image, λ lpips is the LPIPS loss weight, C lpips is the perceptual loss between the rendered image and the real image of the i-th image, is the rendered mask of the i-th image, is the real mask of the i-th image, λ mask is the weight of the mask loss term.
[0030] Furthermore, in the mesh training stage of the image synthesis model, depth estimation loss, normal vector estimation loss, and regularization loss are adopted to construct the loss function as:
[0031]
[0032] In the formula, is the rendered depth of the i-th image, is the real depth of the i-th image, is the rendered normal of the i-th image, is the real normal of the i-th image, λ depth is the weight of the depth loss term, λ normal is the weight of the normal loss term, M gt is the real mask, λ reg is the weight of the regularization loss term, Loss reg is the regularization loss.
[0033] Further, post - process the three - dimensional model of the contaminated insulator, where the post - processing is to smooth the rough surface of the three - dimensional model of the contaminated insulator.
[0034] According to another aspect of the present invention, a three - dimensional reconstruction system for contaminated insulators without a depth sensor of a single - eye type is provided, including:
[0035] A single - eye image acquisition module, configured to acquire single - eye images of contaminated insulators in any pose;
[0036] A main body extraction module, configured to extract the main body from the single - eye image of the contaminated insulator that has undergone image pre - processing to obtain a background - free contaminated insulator image;
[0037] A multi - perspective contaminated insulator image acquisition module, configured to perform multi - perspective image prediction on the background - free contaminated insulator image according to an implicit transformation through a multi - perspective image prediction model to obtain multi - perspective contaminated insulator images;
[0038] A three - dimensional reconstruction module, configured to perform three - dimensional reconstruction on the multi - perspective contaminated insulator images through an image synthesis model to obtain a three - dimensional model of the contaminated insulator.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] 1. The present invention extracts the main body from the single - eye image of the contaminated insulator to obtain a background - free contaminated insulator image, performs multi - perspective image prediction on the background - free contaminated insulator image through a multi - perspective image prediction model to obtain multi - perspective contaminated insulator images, and performs three - dimensional reconstruction on the multi - perspective contaminated insulator images through an image synthesis model to obtain a three - dimensional model of the contaminated insulator. Without a depth sensor and taking multi - perspective images of the insulator, it realizes high - precision three - dimensional modeling of the contaminated insulator, reduces the operation difficulty of maintenance personnel, and reduces costs.
[0041] 2. The present invention post - processes the three - dimensional model of the contaminated insulator, smooths the rough surface of the three - dimensional model of the contaminated insulator, improves the accuracy of the three - dimensional model of the contaminated insulator. On the premise of ensuring the model accuracy and visualization ability, the contaminated position and contamination degree of the contaminated insulator can be observed more clearly, assisting maintenance personnel to make a more accurate judgment on the fault state of the insulator. Description of the Drawings
[0042] Figure 1 It is a schematic flowchart of a three - dimensional reconstruction method for contaminated insulators without a depth sensor of a single - eye type proposed by the present invention;
[0043] Figure 2 It is a schematic diagram of a single - eye image of a contaminated insulator;
[0044] Figure 3 Schematic diagram of a multi - perspective contaminated insulator image
[0045] Figure 4 Schematic diagram of a 3D model of a contaminated insulator Detailed implementation manners
[0046] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to the following embodiments.
[0047] English abbreviations involved:
[0048] Learned Perceptual Image Patch Similarity: LPIPS
[0049] Neural Radiance Fields: NeRF
[0050] Embodiment 1
[0051] This embodiment provides a three - dimensional reconstruction method for a contaminated insulator without a depth sensor in a single - view, as Figure 1 shown, including the following steps:
[0052] S1. Obtain a single - view image of a contaminated insulator in an arbitrary pose.
[0053] In practical applications, a single - view image of a contaminated insulator in an arbitrary pose is obtained by photographing the contaminated insulator with a camera as the input for three - dimensional reconstruction. The single - view image of the contaminated insulator is as Figure 2 shown. In this embodiment, the reconstruction of the contaminated insulator mainly focuses on the contamination condition of the insulator, and the spatio - temporal state of the insulator is appropriately ignored.
[0054] S2. Extract the main body from the single - view image of the contaminated insulator to obtain an image of the contaminated insulator without a background.
[0055] The specific steps for extracting the main body from the single - view image of the contaminated insulator include:
[0056] Extract features from the single - view image of the contaminated insulator. Traditional image - processing methods (such as edge detection, texture feature extraction, etc.) or deep - learning methods (such as feature extraction by convolutional neural networks) can be used. The purpose of this step is to extract key information in the image and provide a basis for subsequent depth estimation and object detection.
[0057] The distance from each pixel point in the monocular image of the contaminated insulator to the camera is obtained through monocular depth estimation. The methods of monocular depth estimation include learning-based methods (such as training models using convolutional neural networks), geometric and physical constraints (such as using rules like the convergence of parallel lines and the change in object size), and self-supervised and weakly-supervised learning (guiding the model to learn depth information through image transformations such as optical flow and disparity).
[0058] Use an object detection algorithm (such as YOLOv8) to identify the main object and its location in the monocular image of the contaminated insulator. A main object mask is generated corresponding to the main object and its location in the monocular image of the contaminated insulator. The main object mask is a binary image, where the pixel value of the main object area is 1 (white) and the pixel value of the background area is 0 (black).
[0059] Use morphological operations to remove noise and holes and improve the quality of the main object mask. The morphological operations include multiple ones such as erosion, dilation, opening, and closing. Morphological operations can effectively remove small noise points and holes in the image, making the main object area more complete and clear.
[0060] Segment the main object from the background according to the generated main object mask through image segmentation technology to obtain a background-free contaminated insulator image.
[0061] S3. Perform multi-view image prediction on the background-free contaminated insulator image according to the implicit transformation through a multi-view image prediction model to obtain multi-view contaminated insulator images.
[0062] Through a pre-trained multi-view image prediction model, perform multi-view image prediction on the background-free contaminated insulator image. A set of multi-view contaminated insulator images based on the pre-trained model, such as Figure 3 shown.
[0063] The training steps of the multi-view image prediction model specifically include:
[0064] When generating a 3D model, use a relatively large number of contaminated insulator pictures from different perspectives. The angular interval of the multi-view pictures strictly conforms to a pre-determined numerical relationship. The quality and quantity of the pictures are greater than those of general 3D reconstruction methods, and a more refined 3D model can be obtained. In this embodiment, monocular images of contaminated insulators collected from 6 groups of camera perspectives separated by 60 degrees are input, and 6 groups are repeated to achieve image prediction with different parameters.
[0065] The multi-view image prediction is based on the following implicit transformation:
[0066]
[0067] In the formula, For the predicted image, f is the implicit transformation, x is the input RGB image, R is the rotation angle of the camera, and T is the position of the camera.
[0068] The specific steps of the implicit transformation are as follows: Input an RGB image. At the same time, input a T, which represents the distance between the camera and the object. By default, the image has no rotation at this time, and the rotation angle of the camera is 0. Input the rotation angle R of the camera, and through the implicit transformation, output the image after rotating the RGB image with a rotation angle of 0 by R degrees.
[0069] The implicit transformation is implemented based on an improved diffusion model, and its improvement method is to replace the loss function with:
[0070]
[0071] In the formula, E Z~ε(x),t,∈~N(0,1) is the optimization expectation, that is, to minimize the average of the squared L2 norm between the predicted noise and the actual noise. ∈ is the noise variable, ∈~N(0,1) means that the noise variable ∈ follows the standard normal distribution. zt is the latent variable, t is the frame, x is the input RGB image, R is the rotation angle of the camera, T is the position of the camera, θ is the parameter to be optimized of the denoising network, and c(x,R,T) is the embedding obtained by fusing the input image and the relative camera extrinsic parameters (R,T), which is used as the conditional information of the diffusion model to guide the diffusion model to generate images corresponding to the perspective.
[0072] When the value of the loss function is less than the preset value, it is considered that the training of this model has converged. Applying this loss function not only improves the training speed but also ensures the accuracy of model training.
[0073] S4. Perform 3D reconstruction on the multi-view contaminated insulator images through an image synthesis model to obtain a 3D model of the contaminated insulator.
[0074] For the generated multi-view contaminated insulator images, through the pre-trained synthesis model for synthesis, a 3D reconstruction model of the contaminated insulator as shown in Figure 4 can be obtained. This model is beneficial for technicians to judge the state of insulators in the field of insulator operation and maintenance.
[0075] The image synthesis model is a 3D model, and its training process includes the NeRf training stage and the mesh training stage.
[0076] In the NeRf training stage of the image synthesis model, image loss and mask loss are adopted, and the loss function is:
[0077]
[0078] In the formula, is the rendered image of the i-th image, is the ground truth image of the i-th image, λ lpips is the LPIPS loss weight, C lpips is the perceptual loss between the rendered image and the ground truth image of the i-th image, is the rendered mask of the i-th image, is the ground truth mask of the i-th image, λ mask is the weight of the mask loss term.
[0079] In the grid training stage of the image synthesis model, a loss function is constructed using depth estimation loss, normal vector estimation loss, and regularization loss as follows:
[0080]
[0081] In the formula, is the rendered depth of the i-th image, is the ground truth depth of the i-th image, is the rendered normal of the i-th image, is the ground truth normal of the i-th image, λ depth is the weight of the depth loss term, λ normal is the weight of the normal loss term, M gt is the ground truth mask, λ reg is the weight of the regularization loss term, Loss reg is the regularization loss.
[0082] When Loss1 and Loss2 are less than the preset values, it is considered that the model training is completed.
[0083] In another preferred embodiment, it further includes:
[0084] S5. Post-process the three-dimensional model of the contaminated insulator to obtain a more refined three-dimensional model of the insulator.
[0085] By smoothing the rough surface of the three-dimensional model of the contaminated insulator, the post-processing of the three-dimensional reconstruction model is realized. This operation is beneficial for finite element simulation analysis or other predictions of the obtained three-dimensional model. The reliability of the finite element simulation results is higher than that of traditional algorithms, and it has a certain predictive ability for the subsequent development of faults.
[0086] Embodiment 2
[0087] This embodiment provides a three-dimensional reconstruction system for contaminated insulators without a depth sensor, including:
[0088] A contaminated insulator monocular image acquisition module for acquiring monocular images of contaminated insulators in any pose;
[0089] A background-free polluted insulator image acquisition module, which is used to extract the main body of the monocular image of the polluted insulator to obtain a background-free polluted insulator image;
[0090] A multi-view polluted insulator image acquisition module, which is used to perform multi-view image prediction on the background-free polluted insulator image through a multi-view image prediction model to obtain a multi-view polluted insulator image;
[0091] A 3D reconstruction module, which is used to perform 3D reconstruction on the multi-view polluted insulator image through a synthesis model to obtain a 3D model of the polluted insulator.
[0092] The rest is the same as in Embodiment 1.
[0093] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.
Claims
1. A monocular 3D reconstruction method for contaminated insulators without depth sensor, characterized in that: The following steps are involved: Obtain monocular images of contaminated insulators in any pose; Extracting the main body of the contaminated insulator monocular image to obtain a contaminated insulator image without background; Performing multi-view image prediction on the background-free contaminated insulator image according to implicit transformation using a multi-view image prediction model to obtain a multi-view contaminated insulator image; The multi-view contaminated insulator image is three-dimensionally reconstructed through an image synthesis model to obtain a contaminated insulator three-dimensional model.
2. The monocular depth sensor-free contaminated insulator 3D reconstruction method according to claim 1 is characterized in that: The specific steps of extracting the subject of the contaminated insulator monocular image include: Extract features from the monocular image of the contaminated insulator to obtain key information in the image; Obtaining the distance from each pixel in the monocular image of the contaminated insulator to the camera through depth estimation; Identify the main object and its position in the monocular image of the contaminated insulator by target detection; Generate a subject mask according to the subject and its position in the monocular image of the contaminated insulator; Using morphological operations to remove noise and holes to improve the quality of the subject mask; The subject is segmented from the background by image segmentation technology according to the subject mask to obtain an insulator image without background contamination.
3. The monocular depth sensor-free contaminated insulator 3D reconstruction method according to claim 2 is characterized in that: The morphological operations include multiple ones of erosion, dilation, opening and closing operations.
4. The monocular depth sensor-free contaminated insulator 3D reconstruction method according to claim 1, characterized in that: The implicit transformation is: In the formula, is the predicted image, f is the implicit transformation, x is the input RGB image, R is the rotation angle of the camera, and T is the position of the camera.
5. The monocular depth sensor-free contaminated insulator 3D reconstruction method according to claim 4, characterized in that: The multi-view image prediction model is a diffusion model, the implicit transformation is implemented based on the diffusion model, and the loss function of the diffusion model is: In the formula, E Z~ε(x),t,∈~N(0,1) To optimize the expectation, that is, to minimize the average value of the L2 norm square between the predicted noise and the actual noise, ∈ is the noise variable, ∈~N(0,1) indicates that the noise variable ∈ obeys the standard normal distribution, zt is the latent variable, t is the frame, x is the input RGB image, R is the rotation angle of the camera, T is the position of the camera, θ is the parameter to be optimized of the denoising network, c(x,R,T) is the embedding obtained by fusing the input image and the relative camera extrinsic parameters (R,T), which serves as the conditional information of the diffusion model to guide the diffusion model to generate images of the corresponding perspective.
6. The monocular depth sensor-free contaminated insulator 3D reconstruction method according to claim 1, characterized in that: The image synthesis model is a 3D model, and its training process includes a NeRf training stage and a grid training stage.
7. The monocular depth sensor-free contaminated insulator 3D reconstruction method according to claim 6, characterized in that: In the NeRf training stage of the image synthesis model, image loss and mask loss are used, and the loss function is: In the formula, is the rendered image of the i-th image, is the true image of the i-th image, λ lpips is the LPIPS loss weight, C lpips is the perceptual loss between the rendered image and the real image of the i-th image, is the rendering mask of the i-th image, is the true mask of the i-th image, λ mask is the weight of the mask loss term.
8. The monocular depth sensor-free contaminated insulator 3D reconstruction method according to claim 7, characterized in that: In the grid training stage of the image synthesis model, the loss function is constructed using depth estimation loss, normal vector estimation loss and regularization loss: In the formula, is the rendering depth of the i-th image, is the true depth of the i-th image, is the rendering normal of the ith image, is the true normal of the ith image, λ depth is the weight of the depth loss term, λ normal is the weight of the normal loss term, M gt is the true mask, λ reg is the weight of the regularization loss term, Loss reg is the regularization loss.
9. The monocular depth sensor-free contaminated insulator 3D reconstruction method according to claim 1, characterized in that: The three-dimensional model of the contaminated insulator is post-processed, and the post-processing is to smooth the rough surface of the three-dimensional model of the contaminated insulator.
10. A monocular 3D reconstruction system for contaminated insulators without depth sensor, characterized in that: include: Monocular image acquisition module, used to acquire monocular images of contaminated insulators in any posture; A subject extraction module, used for performing subject extraction on the contaminated insulator monocular image after image preprocessing to obtain a contaminated insulator image without background; A multi-view contaminated insulator image acquisition module is used to perform multi-view image prediction on the background-free contaminated insulator image according to implicit transformation through a multi-view image prediction model to obtain a multi-view contaminated insulator image; The three-dimensional reconstruction module is used to perform three-dimensional reconstruction on the multi-view contaminated insulator image through an image synthesis model to obtain a contaminated insulator three-dimensional model.
Citation Information
Patent Citations
Insulator pollution multi-angle image detection and fineness analysis method and device
CN116071368A