Power transmission line detection method based on three-dimensional reconstruction and related equipment
By acquiring multiple images of the transmission line and its azimuth perspective, using a multi-assumption joint view selection algorithm and motion structure recovery algorithm to calculate confidence and sparse depth, combined with a three-dimensional reconstruction model for three-dimensional reconstruction and training, the problem of low accuracy of three-dimensional reconstruction detection in the existing technology is solved, and a higher accuracy transmission line detection is achieved.
Patent Information
- Application Number
- CN202510211407.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-06
AI Technical Summary
The transmission line detection method based on three-dimensional reconstruction in the prior art has low accuracy and is difficult to accurately identify defects and corrosion problems of transmission lines.
By acquiring multiple images of the transmission line and its azimuth perspective, the confidence of the image is calculated using a multi-assumption joint view selection algorithm, the motion structure recovery algorithm calculates the sparse depth, and the transmission line is three-dimensionally reconstructed in combination with the three-dimensional reconstruction model. The three-dimensional reconstruction model is then trained, and the trained three-dimensional reconstruction model is generated for detection.
The accuracy of three-dimensional reconstruction and the quality of three-dimensional model are improved, the accuracy of detection results is enhanced, and the defects and corrosion problems of transmission lines can be more precisely identified.
Smart Images

Figure CN120107485A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of three-dimensional reconstruction, and in particular to a power transmission line detection method based on three-dimensional reconstruction and related equipment. Background Art
[0002] Transmission lines are an important part of the power system, and their safety and reliability directly affect the stability of power supply. At present, traditional transmission line inspection methods mainly rely on manual inspection, drone inspection, and automated inspection methods based on image processing. However, these methods have some significant disadvantages:
[0003] Manual inspection: Manual inspection is not only inefficient, but also poses great safety hazards, especially for high-voltage transmission lines, where manual contact poses a great risk.
[0004] Drone inspection: Although drone inspection can improve efficiency, its image information is relatively simple and is usually limited to two-dimensional (2D) plane images, making it difficult to accurately capture the three-dimensional structural information of lines and equipment. In addition, the flight environment of drones is limited, especially in complex weather conditions, which greatly affects flight stability and data quality.
[0005] Automatic detection by image processing: When performing structural fault detection, the detection method based on image processing can usually only process 2D plane images, and it is difficult to accurately reconstruct the three-dimensional structure. In addition, when the image quality is poor, the processing effect is not ideal.
[0006] Traditional 3D reconstruction technology: Traditional 3D reconstruction technologies such as LiDAR and traditional multi-view stereo geometry (MVS) have been applied in power transmission line detection and other fields, but they also have some limitations. For example, LiDAR cannot obtain the color texture of objects, and traditional MVS may have holes and texture aliasing during the reconstruction process. In the face of these challenges, a new and more superior 3D reconstruction method is urgently needed for power operation and maintenance.
[0007] Neural Radiance Field (NeRF) is a three-dimensional (3D) reconstruction method based on deep learning, which generates realistic new perspective images by inferring the density and color distribution of 3D scenes from images of multiple perspectives. However, traditional NeRF model training relies on learning the details of the scene from sparse image data, which usually requires a lot of computing resources and high-quality input images. Especially when the depth information is incomplete or inaccurate, the performance of NeRF may be limited, resulting in poor 3D modeling and inability to accurately identify defects, corrosion and other problems of transmission lines. It can be seen that there is a problem of low accuracy of transmission line detection based on 3D reconstruction. Summary of the invention
[0008] The present application provides a power transmission line detection method based on three-dimensional reconstruction and related equipment, which can solve the problem of low accuracy of power transmission line detection based on three-dimensional reconstruction.
[0009] In a first aspect, an embodiment of the present application provides a power transmission line detection method based on three-dimensional reconstruction, the power transmission line detection method comprising:
[0010] Acquire multiple images of the target transmission line and the azimuth viewing angle of each image, and acquire the three-dimensional positions of multiple spatial points of the target transmission line;
[0011] The confidence of all images is calculated using the multi-hypothesis joint view selection algorithm, and the sparse depth of spatial points is calculated using the motion structure recovery algorithm;
[0012] Using the 3D reconstruction model, the target transmission line is 3D reconstructed based on the 3D positions of all spatial points, all azimuth viewing angles, sparse depths, and confidence levels to obtain a 3D model of the target transmission line.
[0013] The 3D reconstruction model is trained according to the 3D model and all the images to obtain a trained 3D reconstruction model;
[0014] The trained 3D reconstruction model is used to perform 3D reconstruction on the transmission line to be inspected to obtain a final 3D model of the transmission line to be inspected. The transmission line to be inspected is inspected based on the final 3D model to obtain a detection result; the detection result is used to describe the defect condition of the transmission line to be inspected.
[0015] Optionally, the sparse depth is calculated using a structure from motion algorithm, including:
[0016] Use the motion structure recovery algorithm to determine multiple key points from all spatial points, and build optimization targets based on all key points;
[0017] The optimization objective is minimized to obtain the sparse depth.
[0018] Optionally, the optimization goal is:
[0019]
[0020] in, represents the set of all key points, a represents the scale, b represents the displacement scalar, Z SfM represents the sparse depth to be solved, Z MVS Represents a dense depth map.
[0021] Optionally, a three-dimensional reconstruction model is used to perform three-dimensional reconstruction on the target transmission line based on the three-dimensional positions of all spatial points, all azimuth viewing angles, sparse depths, and confidence levels to obtain a three-dimensional model of the target transmission line, including:
[0022] Based on all images, multiple spatial rays in the three-dimensional space where all spatial points are located are obtained;
[0023] For each spatial ray, the spatial ray is sampled according to the confidence and sparse depth to obtain multiple sampling points, and the color is calculated based on all sampling points, the three-dimensional positions of all spatial points and all azimuth viewing angles to obtain the color function of the spatial ray;
[0024] All spatial rays are rendered based on all color functions to obtain a three-dimensional model of the target transmission line.
[0025] Optionally, the spatial rays are sampled according to the confidence and the sparse depth to obtain multiple sampling points, including:
[0026] By formula:
[0027] t i ~N(Z,σ 2 )
[0028] σ=Z · clamp(e,α l ,α h )
[0029] Get the i-th sampling point t on the spatial ray i ;
[0030] Among them, Z represents the sparse depth, σ represents the standard deviation, e represents the confidence, and α l Represents the lower bound of confidence, α h Represents the upper limit of the confidence level, i=1,2,...,N, where N represents the number of sampling points on the spatial ray.
[0031] Optionally, color calculation is performed based on all sampling points, the three-dimensional positions of all spatial points, and all azimuth viewing angles to obtain a color function of the spatial ray, including:
[0032] By formula:
[0033]
[0034] Get the color function C(r) of the spatial ray;
[0035] Among them, T i represents the transmittance at the i-th sampling point, σ i Represents the high-frequency encoding of the 3D position and the corresponding azimuth viewing angle of the i-th sampling point, δ irepresents the distance between the i-th sampling point and the i+1-th sampling point, δ j represents the distance between the jth sampling point and the j+1th sampling point, c i Indicates the color of the i-th sampling point.
[0036] Optionally, a 3D reconstruction model is trained according to the 3D model and all the images to obtain a trained 3D reconstruction model, including:
[0037] Construct a loss function based on the 3D model and all images;
[0038] Determine whether the value of the loss function is less than the preset value of the loss function;
[0039] If so, the three-dimensional reconstruction model is used as the trained three-dimensional reconstruction model;
[0040] Otherwise, adjust the parameters of the 3D reconstruction model, and return to the step of using the 3D reconstruction model to perform 3D reconstruction on the target transmission line based on the 3D positions of all spatial points, all azimuth viewing angles, sparse depths, and confidence levels to obtain a 3D model of the target transmission line.
[0041] Optionally, the loss function is:
[0042]
[0043] in, represents the value of the loss function, P represents the set of spatial rays, Indicates color loss, represents the depth loss, represents the true color value in the image, C(r) represents the color function of the three-dimensional model, represents the sparse depth of the 3D model and Z(r) represents the depth estimate.
[0044] In a second aspect, an embodiment of the present application provides a transmission line detection device based on three-dimensional reconstruction, comprising:
[0045] An acquisition module, used to acquire multiple images of the target transmission line and the azimuth viewing angle of each image, and to acquire the three-dimensional positions of multiple spatial points of the target transmission line;
[0046] A calculation module, used to calculate the confidence of all images using a multi-hypothesis joint view selection algorithm, and to calculate the sparse depth of spatial points using a motion structure recovery algorithm;
[0047] A three-dimensional reconstruction module is used to use the three-dimensional reconstruction model to perform three-dimensional reconstruction on the target transmission line based on the three-dimensional positions of all spatial points, all azimuth viewing angles, sparse depths, and confidence levels to obtain a three-dimensional model of the target transmission line;
[0048] A training module, used for training the 3D reconstruction model according to the 3D model and all the images to obtain a trained 3D reconstruction model;
[0049] The detection module is used to use the trained three-dimensional reconstruction model to perform three-dimensional reconstruction on the transmission line to be detected, obtain the final three-dimensional model of the transmission line to be detected, and detect the transmission line to be detected based on the final three-dimensional model to obtain the detection result; the detection result is used to describe the defect status of the transmission line to be detected.
[0050] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned three-dimensional reconstruction-based power transmission line detection method when executing the above-mentioned computer program.
[0051] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned power transmission line detection method based on three-dimensional reconstruction.
[0052] The above solution of the present application has the following beneficial effects:
[0053] In an embodiment of the present application, multiple images of the target transmission line and the azimuth viewing angle of each image are obtained, and the three-dimensional positions of multiple spatial points of the target transmission line are obtained, and then the confidence of all images is calculated using a multi-hypothesis joint view selection algorithm, and the sparse depth of the spatial point is calculated using a motion structure recovery algorithm, and then a three-dimensional reconstruction model is used to perform three-dimensional reconstruction on the target transmission line based on the three-dimensional positions of all spatial points, all azimuth viewing angles, sparse depths, and confidences to obtain a three-dimensional model of the target transmission line, and then the three-dimensional reconstruction model is trained according to the three-dimensional model and all images to obtain a trained three-dimensional reconstruction model, and finally the trained three-dimensional reconstruction model is used to perform three-dimensional reconstruction on the transmission line to be detected to obtain a final three-dimensional model of the transmission line to be detected, and the transmission line to be detected is detected according to the final three-dimensional model to obtain a detection result. Among them, using confidence and sparse depth for three-dimensional reconstruction can improve the accuracy of three-dimensional reconstruction and the quality of the three-dimensional model. Training the three-dimensional reconstruction model can make the performance of the three-dimensional reconstruction model meet expectations. Using the trained three-dimensional reconstruction model for three-dimensional reconstruction can further improve the accuracy of three-dimensional reconstruction. Detecting power lines based on high-precision three-dimensional models effectively improves the accuracy of detection results.
[0054] Other beneficial effects of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0056] Figure 1 A flowchart of a power transmission line detection method based on three-dimensional reconstruction provided in one embodiment of the present application;
[0057] Figure 2 A schematic diagram of Gaussian distribution provided in an embodiment of the present application;
[0058] Figure 3 A sampling point distribution diagram provided in an embodiment of the present application;
[0059] Figure 4 A schematic diagram of the structure of a power transmission line detection device based on three-dimensional reconstruction provided in one embodiment of the present application;
[0060] Figure 5 A schematic diagram of the structure of a terminal device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0061] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0062] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0063] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0064] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0065] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0066] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0067] In response to the problem of low accuracy of existing transmission line detection based on three-dimensional reconstruction, an embodiment of the present application provides a transmission line detection method based on three-dimensional reconstruction. The transmission line detection method obtains multiple images of the target transmission line and the azimuth viewing angle of each image, and obtains the three-dimensional positions of multiple spatial points of the target transmission line, and then uses a multi-hypothesis joint view selection algorithm to calculate the confidence of all images, and uses a motion structure recovery algorithm to calculate the sparse depth of the spatial point, and then uses a three-dimensional reconstruction model to perform three-dimensional reconstruction on the target transmission line based on the three-dimensional positions of all spatial points, all azimuth viewing angles, sparse depths, and confidences to obtain a three-dimensional model of the target transmission line, and then trains the three-dimensional reconstruction model based on the three-dimensional model and all images to obtain a trained three-dimensional reconstruction model, and finally uses the trained three-dimensional reconstruction model to perform three-dimensional reconstruction on the transmission line to be detected to obtain a final three-dimensional model of the transmission line to be detected, and detects the transmission line to be detected based on the final three-dimensional model to obtain a detection result. Among them, using confidence and sparse depth for three-dimensional reconstruction can improve the accuracy of three-dimensional reconstruction and the quality of the three-dimensional model. Training the three-dimensional reconstruction model can make the performance of the three-dimensional reconstruction model meet expectations. Using the trained three-dimensional reconstruction model for three-dimensional reconstruction can further improve the accuracy of three-dimensional reconstruction. Detecting power lines based on high-precision three-dimensional models effectively improves the accuracy of detection results.
[0068] Next, the power transmission line detection method based on three-dimensional reconstruction provided by the present application is exemplified.
[0069] like Figure 1 As shown, the power transmission line detection method based on three-dimensional reconstruction provided by the present application includes the following steps:
[0070] Step 11, obtaining multiple images of the target transmission line and the azimuth viewing angle of each image, and obtaining the three-dimensional positions of multiple spatial points of the target transmission line.
[0071] In some embodiments of the present application, the image of the target transmission line is obtained by taking pictures by a camera, a drone or other equipment, and the azimuth viewing angle is the shooting direction when the image is taken.
[0072] It should be noted that the above-mentioned spatial points are points in the three-dimensional modeling space when three-dimensional modeling is performed on the target transmission line, and the three-dimensional position is the position in the three-dimensional coordinate system of the three-dimensional modeling. If the target transmission line model is rebuilt using three-dimensional modeling software, the three-dimensional coordinate system is the coordinate system used by the three-dimensional modeling software.
[0073] Step 12: Calculate the confidence of all images using the multi-hypothesis joint view selection algorithm, and calculate the sparse depth of the spatial points using the motion structure recovery algorithm.
[0074] In an embodiment of the present application, the steps of calculating the confidence of all images using the multi-hypothesis joint view selection algorithm and calculating the sparse depth of the spatial points using the motion structure recovery algorithm include:
[0075] In the first step, the confidence scores of all images are calculated using a multi-hypothesis joint view selection algorithm.
[0076] For example, in traditional MVS reconstruction, the reconstructed model is usually obtained through feature extraction and matching, stereo matching, parallax smoothing and filling, surface reconstruction and three-dimensional reconstruction, texture mapping and rendering. However, due to parallax discontinuity, occlusion problems and computational complexity, the reconstructed results interfere with defect detection. In traditional MVS reconstruction, there is an adaptive chessboard sampling and multi-hypothesis joint view selection algorithm (ACMH, Adaptive Multi-Hypothesis Joint View Selection) to obtain a robust multi-view matching cost for each pixel. For a pixel p in the reference image (an image selected from all images by traditional MVS reconstruction technology), the corresponding matching cost is calculated using the propagated hypothesis and embedded into the cost matrix:
[0077]
[0078] Where m j,i is the matching cost of the ith hypothesis scored by the jth image, where the hypothesis is a randomly generated normal vector for each pixel of the reference image, 8 means that for each pixel, 8 local optimal hypotheses are selected from 8 regions in the entire image, and N is the number of input images. The matching cost is calculated using bilateral weighted adaptation of normalized cross-correlation, and the calculation formula is as follows, which describes the photometric consistency between the reference and source blocks:
[0079]
[0080] Where I is the reference image, T is other images other than the reference image, and a window W(p) centered on pixel p is selected, whose size is generally k×k. and are the means of the pixels in the windows of I and T respectively.
[0081] To infer the aggregate view from the cost matrix above, record decisions are applied in each column to determine whether the view is suitable. For a bad view, its corresponding eight matching costs are always high. In contrast, a good view always has some smaller matching costs. In addition, a good matching cost will reduce the iterations with the algorithm. Therefore, a good matching cost boundary is defined as:
[0082]
[0083] Where t represents the tth iteration, τ 0 is the initial matching cost threshold, and α is a constant. In addition, a fixed bad matching cost threshold τ1 is defined (τ1>τ(t)). In each iteration, among all matching costs corresponding to an image, if the condition is met: m j,i <τ(t) The number of matching costs is greater than the first preset number n1 and satisfies the condition: m j,i >τ1 is less than the second preset number n2. Then the image is merged into the current view selection set St in the tth iteration. In the next iteration, the image that is not in the set St is calculated, and the corresponding matching cost is calculated by using the propagation hypothesis and embedded into the cost matrix. After the preset number of iterations, the matching costs of all views in the current view selection set St that meet the condition (m j,i <τ(t) and m j,i >τ1) and substitute the average of all matching costs into the following formula to calculate the confidence level C(m ij ):
[0084]
[0085] In the above formula, m ij represents the average value in the above text, and β represents a constant.
[0086] In the second step, the sparse depth of the spatial points is calculated using the motion structure recovery algorithm.
[0087] Specifically, a structure from motion (SFM) algorithm is used to determine multiple key points from all spatial points, and an optimization target is constructed based on all the key points. The optimization target is then minimized to obtain a sparse depth.
[0088] The above optimization objectives are:
[0089]
[0090] in, represents the set of all key points, a represents the scale, b represents the displacement scalar, Z SfM represents the sparse depth to be solved, Z MVS Represents a dense depth map.
[0091] The depth error between the sparse depth and dense depth maps is expressed as:
[0092]
[0093] in, Represents the sparse depth obtained by the solution.
[0094] Exemplarily, a solution algorithm such as a gradient descent method and a cosine annealing algorithm may be used to minimize the optimization target and obtain the sparse depth.
[0095] Step 13, using the three-dimensional reconstruction model, based on the three-dimensional positions of all spatial points, all azimuth viewing angles, sparse depths, and confidence levels, the target transmission line is three-dimensionally reconstructed to obtain a three-dimensional model of the target transmission line.
[0096] In some embodiments of the present application, the above-mentioned step of using the three-dimensional reconstruction model to perform three-dimensional reconstruction on the target transmission line based on the three-dimensional positions of all spatial points, all azimuth viewing angles, sparse depths, and confidence levels to obtain the three-dimensional model of the target transmission line includes:
[0097] The first step is to obtain multiple spatial rays in the three-dimensional space where all spatial points are located based on all images.
[0098] Specifically, for each image, the camera is calibrated to obtain the intrinsic and extrinsic parameters of the camera, so as to map the image coordinates to the coordinates in the three-dimensional space and obtain the surface depth information of the object in each image. Then, the images and depth data under different viewing angles are registered and aligned. In the three-dimensional space, each pixel in the image is used as the starting point of a spatial ray, and the direction of the azimuth viewing angle of the image is used as the direction of these spatial rays to obtain the spatial rays.
[0099] In the second step, for each spatial ray, the spatial ray is sampled according to the confidence and sparse depth to obtain multiple sampling points, and the color is calculated based on all sampling points, the three-dimensional positions of all spatial points and all azimuth viewing angles to obtain the color function of the spatial ray.
[0100] Specifically, through the formula:
[0101] t i ~N(Z,σ 2 )
[0102] σ=Z · clamp(e,α l ,α h )
[0103] Get the i-th sampling point t on the spatial ray i ;
[0104] Among them, Z represents the sparse depth, σ represents the standard deviation, e represents the confidence, and α l Represents the lower bound of confidence, α h Represents the upper limit of the confidence level, i=1,2,...,N, where N represents the number of sampling points on the spatial ray.
[0105] By formula:
[0106]
[0107] Get the color function C(r) of the spatial ray;
[0108] Among them, T i represents the transmittance at the i-th sampling point, σ i Represents the high-frequency encoding of the 3D position and the corresponding azimuth viewing angle of the i-th sampling point, δ i represents the distance between the i-th sampling point and the i+1-th sampling point, δ j represents the distance between the jth sampling point and the j+1th sampling point, c i Indicates the color of the i-th sampling point.
[0109] The calculation formula for the above high frequency coding is:
[0110] γ(p)=(sin(2 0 πp),cos(2 0 πp),...,sin(2 L-1 πp),cos(2 L-1 πp))
[0111] Among them, γ(p) represents the high-frequency code of the spatial ray, and p represents the three-dimensional position of all spatial points on the spatial ray or the azimuth view of the image corresponding to the spatial ray. The high-frequency code of the sampling point is obtained by discretizing the high-frequency code of the spatial ray. When inputting the three-dimensional position, L can be 8, and for the azimuth view, L can be 5.
[0112] It should be noted that the color function C(r) of the above-mentioned spatial ray describes the color value (the value of the three color channels RGB) of each sampling point on the spatial ray.
[0113] The third step is to render all spatial rays based on all color functions to obtain the three-dimensional model of the target transmission line.
[0114] Specifically, according to the color function, each sampling point is rendered into a corresponding color to obtain a three-dimensional model of the target transmission line.
[0115] It should be noted that the above-mentioned 3D reconstruction model can be a NeRF model that uses the above-mentioned formula for obtaining the sampling points on the spatial ray for sampling. In the traditional NeRF model, the position (x, y, z) of the 3D point and the corresponding azimuth view (θ, φ) are used as input, and this information is processed by a multi-layer perceptron (MLP) to predict the color (color, c) and voxel density (volume density, σ) of the point, thereby constructing an "implicit representation" of the 3D scene. Specifically, the input 3D coordinate X is first processed through an 8-layer fully connected network (256 channels per layer, using the ReLU activation function), and the voxel density σ and a 256-dimensional feature vector are output. Next, this feature vector is combined with the direction of the camera's view (light direction) and passed to another fully connected layer (using ReLU activation and 128 channels), and finally the RGB color related to the view is output.
[0116] NeRF uses the classic volume rendering principle to render the color of each ray passing through the scene. Specifically, the volume density σ(x) can be regarded as the probability density of encountering an infinitesimal particle when the ray terminates at position x. For a spatial ray r(t) = o + td from the near boundary tn to the far boundary tf, its expected color C(r) is calculated as follows:
[0117]
[0118] The function T(t) represents the cumulative transmittance along the ray from tn to t, that is, the probability that the ray does not encounter any particles on the path from tn to t.
[0119] In practical applications, NeRF cannot directly estimate continuous 3D point information, so a numerical approximation method is needed. In the area where integration is required, approximate calculations are performed by uniformly sampling N discrete points. However, this method may cause MLP to only learn the information of these discrete points, thereby limiting the resolution of NeRF, and the generated results may not be clear enough. To solve this problem, the integration area of the ray is first divided into N small parts, and uniform random sampling is performed in each small area. This method can maintain the continuity of the sampling position while using only discrete points. Each sampling point can be expressed as:
[0120]
[0121] This sampling method has two disadvantages: 1. Points within the sampling range have the same sampling probability and cannot be concentrated on surface points. 2. Since a sampling range is set, the model has no chance to learn points outside the range, which will reduce the diversity of training samples. These two disadvantages will reduce depth accuracy and view synthesis quality.
[0122] Therefore, in the 3D reconstruction model of the present application, the method of sampling the spatial rays with Gaussian distribution according to the confidence and sparse depth can effectively improve the depth accuracy and synthesis quality. Figure 2 As shown in the figure, the curve represents the Gaussian distribution of the space point on the space ray, the horizontal axis represents the position of the space point, and the vertical axis represents the distribution probability. The sampling points obtained by sampling according to the Gaussian distribution are as follows Figure 3 ,The dotted lines in the figure represent spatial rays, and the dots represent sampling points obtained by Gaussian sampling.
[0123] The traditional NeRF model uses the idea of hierarchical representation rendering to improve rendering efficiency by more effectively sampling free space and occluded regions in the scene, which should not contribute to the rendering results. This is done by simultaneously optimizing two neural networks: "coarse" and "fine".
[0124] First use stratified sampling to get N C Points are calculated by coarse rendering equation:
[0125]
[0126] to w i Normalize to get the piecewise constant probability density function, and then get N by inverse transform sampling f points, add to the original N C One of the points is used for fine rendering.
[0127] Although layered sampling improves sampling efficiency by combining a coarse network with a fine network, it still requires multiple sampling and calculations in each network. This results in a significant increase in computational cost compared to traditional single-layer sampling. The rendering process of each layer requires a lot of calculations, especially when dealing with more complex or more detailed scenes, which may lead to longer rendering times.
[0128] It is worth mentioning that using confidence and sparse depth for 3D reconstruction can improve the accuracy of 3D reconstruction and the quality of the 3D model.
[0129] Step 14: training the 3D reconstruction model according to the 3D model and all the images to obtain a trained 3D reconstruction model.
[0130] In some embodiments of the present application, the step of training the 3D reconstruction model according to the 3D model and all the images to obtain the trained 3D reconstruction model includes:
[0131] In the first step, a loss function is constructed based on the 3D model and all images.
[0132] Specifically, the loss function is:
[0133]
[0134] in, represents the value of the loss function, R represents the set of spatial rays, Indicates color loss, represents the depth loss, represents the true color value in the image, C(r) represents the color function of the three-dimensional model, represents the sparse depth of the 3D model and Z(r) represents the depth estimate.
[0135] The second step is to determine whether the value of the loss function is less than the preset value of the loss function.
[0136] If so, the 3D reconstruction model is used as the trained 3D reconstruction model.
[0137] Otherwise, adjust the parameters of the 3D reconstruction model, and return to the step of using the 3D reconstruction model to perform 3D reconstruction on the target transmission line based on the 3D positions of all spatial points, all azimuth viewing angles, sparse depths, and confidence levels to obtain a 3D model of the target transmission line.
[0138] Exemplarily, the value of the loss function is 2.3, and the preset value of the loss function is 1.5. The parameters of the 3D reconstruction model are adjusted, and the step of using the 3D reconstruction model to perform 3D reconstruction on the target transmission line based on the 3D positions of all spatial points, all azimuth viewing angles, sparse depth, and confidence levels to obtain a 3D model of the target transmission line is returned. At this time, the value of the loss function is 1.2, which is less than the preset value of the loss function. The training is completed, and the 3D reconstruction model at this time is used as the trained 3D reconstruction model.
[0139] Step 15, using the trained 3D reconstruction model to perform 3D reconstruction on the transmission line to be detected, to obtain a final 3D model of the transmission line to be detected, and detecting the transmission line to be detected based on the final 3D model to obtain a detection result.
[0140] The above detection results are used to describe the defect status of the transmission line to be detected. The transmission line to be detected is the transmission line that needs to be detected (such as a substation, a transmission cable, etc.).
[0141] Specifically, the confidence of all images of the transmission line to be detected is calculated using a multi-hypothesis joint view selection algorithm, and the sparse depth of the spatial points of the three-dimensional model of the transmission line to be detected is calculated using a motion structure recovery algorithm; the trained three-dimensional reconstruction model is used to perform three-dimensional reconstruction of the target transmission line based on the three-dimensional positions of all spatial points, all azimuth viewing angles, sparse depth, and confidence, to obtain a three-dimensional model of the target transmission line. The three-dimensional model is input into a model capable of defect detection, such as a three-dimensional convolutional neural network (3DCNN), for processing to obtain a detection result.
[0142] It should be noted that the 3D model provides a more comprehensive display of detailed information about the equipment, helping maintenance personnel gain a deeper understanding of the equipment's geometry, surface features, and potential defects. With the help of this 3D reconstruction, surface defect detection and analysis of equipment can be more refined, and problems such as cracks, corrosion, and leakage can be identified and monitored, thereby achieving early warning and taking necessary maintenance measures.
[0143] It is worth mentioning that using confidence and sparse depth for 3D reconstruction can improve the accuracy of 3D reconstruction and the quality of the 3D model. Training the 3D reconstruction model can make the performance of the 3D reconstruction model meet expectations. Using the trained 3D reconstruction model for 3D reconstruction can further improve the accuracy of 3D reconstruction. Detecting power lines based on high-precision 3D models effectively improves the accuracy of detection results.
[0144] In addition, the method of the present application can provide more accurate depth information by introducing depth prior (i.e., sparse depth), helping NeRF to better infer the three-dimensional geometric structure of the scene. This additional depth information can reduce blur, improve the clarity and detail of reconstruction, and is particularly suitable for objects with complex structures. At the same time, it provides initial depth information for NeRF and reduces the process of learning depth from scratch for the model. This can shorten the training time and enable the model to converge to high-quality three-dimensional reconstruction results faster. Using Gaussian sampling, the probability density of the sampling points can be concentrated in more informative areas, while reducing sampling in blank or meaningless areas. This sampling method focusing on high-probability areas can significantly improve rendering efficiency and enable the model to converge to high-quality reconstruction results faster. Since Gaussian sampling samples densely in key areas and reduces attention to irrelevant areas, it can reduce the noise caused by sampling in invalid areas. This is particularly useful in occlusion processing and accurate rendering of object boundaries, which helps to obtain smoother and more realistic rendering effects. Depth loss can provide direct depth information supervision to help NeRF learn the geometric structure in the scene more accurately. By adding depth loss, NeRF can more accurately estimate the depth of each point in the three-dimensional scene, making the reconstructed objects more consistent in space. Depth loss reduces depth uncertainty by constraining the model's depth prediction, reducing blur in the rendering process and depth errors between perspectives. This is particularly important for objects with complex or multi-layer structures (such as scenes with occlusion relationships) because depth loss helps the model distinguish foreground and background more clearly in these areas.
[0145] The following is an exemplary description of the power transmission line detection device based on three-dimensional reconstruction provided by the present application.
[0146] like Figure 3 As shown, the embodiment of the present application provides a power transmission line detection device based on three-dimensional reconstruction, and the power transmission line detection device 400 includes:
[0147] An acquisition module 401 is used to acquire multiple images of a target transmission line and the azimuth viewing angle of each image, and acquire the three-dimensional positions of multiple spatial points of the target transmission line;
[0148] A calculation module 402 is used to calculate the confidence of all images using a multi-hypothesis joint view selection algorithm, and calculate the sparse depth of the spatial point using a motion structure recovery algorithm;
[0149] The three-dimensional reconstruction module 403 is used to use the three-dimensional reconstruction model to perform three-dimensional reconstruction on the target transmission line based on the three-dimensional positions of all spatial points, all azimuth viewing angles, sparse depths, and confidence levels to obtain a three-dimensional model of the target transmission line;
[0150] A training module 404 is used to train the 3D reconstruction model according to the 3D model and all the images to obtain a trained 3D reconstruction model;
[0151] The detection module 405 is used to use the trained three-dimensional reconstruction model to perform three-dimensional reconstruction on the transmission line to be detected, obtain the final three-dimensional model of the transmission line to be detected, and detect the transmission line to be detected based on the final three-dimensional model to obtain the detection result; the detection result is used to describe the defect status of the transmission line to be detected.
[0152] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0153] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0154] like Figure 5 As shown, an embodiment of the present application provides a terminal device, and the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 5 Only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 implements the steps of any of the above-mentioned method embodiments when executing the computer program D102.
[0155] Specifically, when the processor D100 executes the computer program D102, it obtains multiple images of the target transmission line and the azimuth viewing angle of each image, and obtains the three-dimensional positions of multiple spatial points of the target transmission line, and then uses a multi-hypothesis joint view selection algorithm to calculate the confidence of all images, and uses a motion structure recovery algorithm to calculate the sparse depth of the spatial point, and then uses a three-dimensional reconstruction model to perform three-dimensional reconstruction on the target transmission line based on the three-dimensional positions of all spatial points, all azimuth viewing angles, sparse depths, and confidences to obtain a three-dimensional model of the target transmission line, and then trains the three-dimensional reconstruction model according to the three-dimensional model and all images to obtain a trained three-dimensional reconstruction model, and finally uses the trained three-dimensional reconstruction model to perform three-dimensional reconstruction on the transmission line to be detected to obtain a final three-dimensional model of the transmission line to be detected, and detects the transmission line to be detected according to the final three-dimensional model to obtain a detection result. Among them, using confidence and sparse depth for three-dimensional reconstruction can improve the accuracy of three-dimensional reconstruction and the quality of the three-dimensional model. Training the three-dimensional reconstruction model can make the performance of the three-dimensional reconstruction model meet expectations. Using the trained three-dimensional reconstruction model for three-dimensional reconstruction can further improve the accuracy of three-dimensional reconstruction. Detecting power lines based on high-precision three-dimensional models effectively improves the accuracy of detection results.
[0156] The processor D100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0157] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk, a smart memory card (SMC, SmartMedia Card), a secure digital (SD, Secure Digital) card, a flash card (Flash Card), etc. equipped on the terminal device D10. Further, the memory D101 may also include both an internal storage unit of the terminal device D10 and an external storage device. The memory D101 is used to store an operating system, an application program, a boot loader (BootLoader), data and other programs, such as the program code of the computer program, etc. The memory D101 may also be used to temporarily store data that has been output or is to be output.
[0158] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0159] An embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0160] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium that can carry the computer program code to the power transmission line detection method device / terminal device based on three-dimensional reconstruction. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0161] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0162] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0163] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles described in the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A transmission line detection method based on three-dimensional reconstruction, characterized in that: include: Acquire multiple images of a target transmission line and the azimuth viewing angle of each image, and acquire three-dimensional positions of multiple spatial points of the target transmission line; The confidence of all images is calculated using the multi-hypothesis joint view selection algorithm, and the sparse depth of spatial points is calculated using the motion structure recovery algorithm; Using the three-dimensional reconstruction model, based on the three-dimensional positions of all spatial points, all azimuth viewing angles, the sparse depth, and the confidence level, a target transmission line is three-dimensionally reconstructed to obtain a three-dimensional model of the target transmission line; Training the three-dimensional reconstruction model according to the three-dimensional model and all images to obtain a trained three-dimensional reconstruction model; The trained three-dimensional reconstruction model is used to perform three-dimensional reconstruction on the transmission line to be inspected to obtain a final three-dimensional model of the transmission line to be inspected, and the transmission line to be inspected is inspected according to the final three-dimensional model to obtain a detection result; the detection result is used to describe the defect condition of the transmission line to be inspected.
2. The power transmission line detection method according to claim 1, characterized in that: The sparse depth is calculated using the motion structure recovery algorithm, including: Use the motion structure recovery algorithm to determine multiple key points from all spatial points, and build optimization targets based on all key points; The optimization objective is minimized to obtain the sparse depth.
3. The power transmission line detection method according to claim 2, characterized in that: The optimization goal is: in, represents the set of all key points, a represents the scale, b represents the displacement scalar, Z sfM represents the sparse depth to be solved, Z MVS Represents a dense depth map.
4. The power transmission line detection method according to claim 1, characterized in that: The three-dimensional reconstruction model is used to perform three-dimensional reconstruction on the target transmission line based on the three-dimensional positions of all spatial points, all azimuth viewing angles, the sparse depth, and the confidence level to obtain the three-dimensional model of the target transmission line, including: Based on all images, multiple spatial rays in the three-dimensional space where all spatial points are located are obtained; For each spatial ray, sampling the spatial ray according to the confidence and the sparse depth to obtain a plurality of sampling points, performing color calculation based on all sampling points, the three-dimensional positions of all spatial points and all azimuth viewing angles to obtain a color function of the spatial ray; All spatial rays are rendered based on all color functions to obtain a three-dimensional model of the target transmission line.
5. The power transmission line detection method according to claim 4, characterized in that: The sampling of the spatial ray according to the confidence and the sparse depth to obtain a plurality of sampling points includes: By formula: t i ~N(Z,σ 2 ) σ=Z·clamp(e,α l ,a h ) Get the i-th sampling point t on the spatial ray i ; Among them, Z represents the sparse depth, σ represents the standard deviation, e represents the confidence, and α l Represents the lower bound of confidence, α h Represents the upper limit of the confidence level, i=1,2,...,N, where N represents the number of sampling points on the spatial ray.
6. The power transmission line detection method according to claim 5, characterized in that: The color calculation is performed based on all sampling points, the three-dimensional positions of all spatial points and all azimuth viewing angles to obtain the color function of the spatial ray, including: By formula: Get the color function C(r) of the spatial ray; Among them, T i represents the transmittance at the i-th sampling point, σ i Represents the high-frequency encoding of the 3D position and the corresponding azimuth viewing angle of the i-th sampling point, δ i represents the distance between the i-th sampling point and the i+1-th sampling point, δ j represents the distance between the jth sampling point and the j+1th sampling point, c i Represents the color of the i-th sampling point.
7. The power transmission line detection method according to claim 1, characterized in that: The step of training the three-dimensional reconstruction model according to the three-dimensional model and all the images to obtain the trained three-dimensional reconstruction model comprises: Constructing a loss function based on the three-dimensional model and all images; Determine whether the value of the loss function is less than a preset value of the loss function; If yes, the three-dimensional reconstruction model is used as the trained three-dimensional reconstruction model; Otherwise, adjust the parameters of the three-dimensional reconstruction model, and return to the step of using the three-dimensional reconstruction model to perform three-dimensional reconstruction on the target transmission line based on the three-dimensional positions of all spatial points, all azimuth viewing angles, the sparse depth, and the confidence level to obtain the three-dimensional model of the target transmission line.
8. The power transmission line detection method according to claim 7, characterized in that: The loss function is: in, represents the value of the loss function, R represents the set of spatial rays, Indicates color loss, represents the depth loss, represents the true color value in the image, C(r) represents the color function of the three-dimensional model, represents the sparse depth of the 3D model, and Z(r) represents the depth estimate.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the three-dimensional reconstruction-based transmission line detection method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the three-dimensional reconstruction-based power transmission line detection method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Defect detection method, device, equipment and medium
CN120765656A