Electric power engineering three-dimensional model rapid generation method and system based on multi-angle photos
By using multi-angle photo and deep learning technology in power engineering, combined with drone acquisition and convolutional neural network processing, the problems of low efficiency, high cost and limited accuracy in the existing technology are solved, and fast, accurate and low-cost three-dimensional model generation is achieved.
Patent Information
- Application Number
- CN202510215444.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has problems such as inefficient, high cost and limited accuracy in complex environments in the construction of three-dimensional models of power engineering. In particular, the monocular camera method is greatly affected by light and has weak feature extraction effect.
A drone equipped with visible light and infrared cameras collects multi-angle photos from different angles, uses a dual-branch convolutional neural network and attention mechanism to fuse visible light and infrared image features, and uses twin neural networks and cross attention mechanisms to match and image stitching, and combines deep learning surface reconstruction algorithm to optimize the three-dimensional point cloud model.
It realizes the rapid, efficient and low-cost generation of three-dimensional power engineering models, improves the accuracy and completeness of the model, is suitable for complex environments, and reduces equipment costs.
Smart Images

Figure CN120070767A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of 3D modeling technology, and more specifically, to a method and system for quickly generating a 3D model of a power project based on multi-angle photos. Background Art
[0002] In the field of power engineering, 3D models are crucial for accurate planning, efficient construction, and equipment maintenance management throughout the entire life cycle. There are mainly two traditional ways to construct 3D models: one is that designers manually model and combine each device and component one by one. This method not only requires a large amount of manpower and time, but also has a cumbersome modeling process and extremely low efficiency; the other is to use laser point cloud scanning. Although it can quickly obtain the 3D information of an object, the generated model file is extremely large, and the subsequent processing and storage costs are high. And in some complex environments, such as narrow spaces and high electromagnetic interference areas, the accuracy and applicability of laser point cloud scanning will be severely limited. Therefore, there is an urgent need for a new method to quickly, efficiently, and low-costly generate 3D models of power projects.
[0003] The prior art such as the Chinese patent application with the publication number "CN113838191A" discloses a 3D reconstruction method based on an attention mechanism and monocular multi-view. The method includes: S1: Taking pictures of the scene to be measured by a camera and collecting the image data of the scene to be measured; S2: Sequencing the image data, and sequentially extracting feature points and matching feature points to obtain feature point matching pairs; S3: Calibrating the camera to obtain the camera pose information, camera internal parameters, and the structure information of the scene to be measured, and performing sparse point cloud reconstruction; S4: Inputting the camera pose information, camera internal parameters, the structure information of the scene to be measured, and the images into a monocular multi-view reconstruction network with a preset attention mechanism to obtain several depth estimation maps of the scene to be measured; S5: Fusing the several depth estimation maps to obtain a dense point cloud model.
[0004] The problems existing in the above prior art are that this method only relies on a monocular camera to collect image data, which is greatly affected by lighting factors and may affect the accuracy and integrity of the reconstructed model; the feature extraction and fusion effects may be weak, and it is difficult to fully mine the image information. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a method and system for quickly generating a 3D model of a power project based on multi-angle photos.
[0006] The technical solution of the present invention is as follows:
[0007] The present invention proposes a method for quickly generating a 3D model of a power project based on multi-angle photos, including the following steps:
[0008] Step S1: Use a drone equipped with visible light and infrared cameras to take pictures of the equipment and scenes in the power project from different angles, and obtain visible light and infrared images from multiple angles;
[0009] Step S2: Use a dual-branch convolutional neural network CNN to extract the features of visible light and infrared images respectively, and use the attention mechanism to dynamically fuse the features of visible light and infrared images to obtain comprehensive image features;
[0010] Step S3: Take the comprehensive image feature vectors of the images taken from different angles as the input of the Siamese neural network, confirm adjacent images according to the similarity matching results output by the Siamese neural network, and obtain the matching feature points of adjacent images through the cross-attention mechanism;
[0011] Step S4: Stitch adjacent images based on the matching feature points of adjacent images, and perform preliminary 3D reconstruction through the principle of triangulation to obtain a sparse 3D point cloud model;
[0012] Step S5: Optimize and refine the sparse 3D point cloud model based on the surface reconstruction algorithm of deep learning to generate a complete and smooth 3D model.
[0013] As a preferred embodiment, the drone equipped with visible light and infrared cameras is used to take pictures of the equipment and scenes in the power project from different angles to obtain visible light and infrared images from multiple angles. Among them, the visible light and infrared images correspond one by one, and the overlap rate between adjacent images from different angles reaches more than 60%.
[0014] As a preferred embodiment, the process of using the dual-branch convolutional neural network CNN to extract the features of visible light and infrared images respectively specifically includes:
[0015] Construct a scale space with a Gaussian convolution kernel, and the specific calculation formula is:
[0016] L(x, y, σ) = G(x, y, σ) * I(x, y);
[0017] In the formula: L(x, y, σ) is the image in different scale spaces; σ is the scale factor; G(x, y, σ) is the Gaussian kernel function; I(x, y) is the original image;
[0018] Detect the extreme points of the difference of Gaussian, and the specific calculation formula is:
[0019] D(x, y, σ) = L(x, y, kσ) - L(x, y, σ);
[0020] In the formula: D(x, y, σ) is the difference-of-Gaussian image; k is the scale interval factor.
[0021] As a preferred embodiment, the visible light and infrared image features are dynamically fused using the attention mechanism to obtain comprehensive image features, and the feature spaces are aligned through a loss function. The specific calculation formula of the feature space alignment loss function is as follows:
[0022]
[0023] Dynamically fuse multi-modal features using the attention mechanism; the formula for adaptive feature weight allocation is:
[0024] f fusion = αf v + (1 - α)·f ir ;
[0025] α = σ(W·[f v ·f ir + b);
[0026] Where: L align is the feature space alignment loss function; N is the number of samples participating in the calculation of the loss function; is the feature vector of the i-th visible light image; is the feature vector of the i-th infrared image; f fusion is the fused comprehensive image feature vector; α is the adaptive weight coefficient; f v is the feature vector of the visible light image; f ir is the feature vector of the infrared image; σ is the activation function; W is the weight matrix; b is the bias term.
[0027] As a preferred embodiment, in the deep learning-based surface reconstruction algorithm, when optimizing and refining a sparse three-dimensional point cloud model to generate a complete and smooth three-dimensional model, deep learning is combined with multi-view stereo matching to improve the point cloud density and accuracy; the specific formula for generating a dense point cloud is:
[0028] Depth Map = softmax(Conv3D(I 1 , I 2 , …, I n ));
[0029] Where: Depth Map is the depth map generated by three-dimensional convolution of images taken from multiple different angles; Conv3D is the three-dimensional convolution operation; (I 1 , I 2 , …, I n ) are the images taken from different angles.
[0030] As a preferred embodiment, during the process of generating a dense point cloud, statistical filtering is used to remove outliers. The specific calculation formula is:
[0031]
[0032] Retain the points that satisfy |d j -μ| < 3δ;
[0033] Where: μ is the mean value of the data in a certain dimension of the point cloud data, which is used to measure the central tendency of the data in this dimension and serves as a reference benchmark when removing outliers; J is the number of points participating in the statistics; d j is the value of the j-th point in a certain dimension; δ is the standard deviation of the data in a certain dimension of the point cloud data, which is used to measure the fluctuation of the data and serves as a threshold reference when judging outliers.
[0034] As a preferred implementation, in the process of optimizing and refining the sparse three-dimensional point cloud model to generate a complete and smooth three-dimensional model, a pre-trained 3D-GAN of ShapeNet is used as the basic model, and the basic model is fine-tuned through fine-tuning with power equipment data. The loss function formula of the specific fine-tuning process is:
[0035] L total = L GAN + λL recin ;
[0036] Where:
[0037]
[0038] Where: L total is the total loss function of the fine-tuning process; L GAN is the loss of the generative adversarial network; λ is the weight coefficient; L recon is the reconstruction loss; P is the generated point cloud; Q is the real point cloud; p is the point in the generated point cloud; q is the point in the real point cloud.
[0039] On the other hand, the present invention also provides a fast generation system for a three-dimensional model of a power project based on multi-angle photos, including:
[0040] A data acquisition module that uses a drone equipped with visible light and infrared cameras to photograph the equipment and scenes in the power project from different angles to obtain multi-angle visible light and infrared images;
[0041] A feature extraction and fusion module that uses a dual-branch convolutional neural network CNN to extract the features of visible light and infrared images respectively, and uses an attention mechanism to dynamically fuse the features of visible light and infrared images to obtain comprehensive image features;
[0042] The feature point matching module uses the comprehensive image feature vectors of the images taken at different angles as the input of the Siamese neural network to calculate the similarity of the images taken at different angles, and obtains the matching feature points of the images taken at different angles through the cross-attention mechanism;
[0043] The image stitching and preliminary modeling module stitches the captured images based on the matching feature points of the images taken at different angles, and performs preliminary 3D reconstruction through the principle of triangulation to obtain a sparse 3D point cloud model;
[0044] The model generation module optimizes and refines the sparse 3D point cloud model based on the surface reconstruction algorithm of deep learning to generate a complete and smooth 3D model.
[0045] On the other hand, the present invention also provides an electronic device with a computer program stored thereon, and when the computer program is executed by a processor, it implements the method for quickly generating a 3D model of a power project based on multi-angle photos as described in any embodiment of the present invention.
[0046] On the other hand, the present invention also provides a computer-readable medium for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method for quickly generating a 3D model of a power project based on multi-angle photos as described in any embodiment of the present invention.
[0047] The present invention has the following beneficial effects:
[0048] 1. High efficiency and speed: By using a drone equipped with dual cameras to collect images and combining with a fine-tuned 3D-GAN model pre-trained on ShapeNet, the whole process from image collection to 3D model generation can be quickly completed, improving the modeling efficiency.
[0049] 2. Fusion of multi-modal information: Using a dual-branch convolutional neural network and an attention mechanism to fuse visible light and infrared image features, integrating multi-modal information, and improving the accuracy and integrity of the model.
[0050] 3. Improvement of point cloud quality: Generating a dense point cloud by combining deep learning and multi-view stereo matching, and using statistical filtering to remove outliers, improving the point cloud density and accuracy, and optimizing the model quality.
[0051] 4. Strong model optimization ability: Using a pre-trained 3D-GAN as the basic model and fine-tuning it can better fit the power equipment data and generate a more realistic complete and smooth 3D model.
[0052] 5. Low cost and strong adaptability: Compared with laser point cloud scanning, the equipment cost is low, and it is not restricted by complex environments such as narrow spaces and high electromagnetic interference areas, and has a wide range of applicability. Description of the Drawings
[0053] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0054] Figure 1 It is a schematic diagram of the method flow of the present invention. Specific embodiments
[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0056] It should be understood that the step numbers used in the text are only for convenient description and do not limit the execution order of the steps.
[0057] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0058] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0059] The term "and / or" refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0060] Embodiment 1:
[0061] To make the purpose, technical solutions and advantages of the present invention clearer, the following will combine specific embodiments of the present application and refer to the attached Figure 1 , and clearly and completely describe the technical solutions of the present invention.
[0062] To solve the problems of the prior art, the present invention provides a method for quickly generating a three-dimensional model of a power project based on multi-angle photos, including the following steps:
[0063] Step S1: Use a drone equipped with visible light and infrared cameras to capture images of the equipment and scenes in the power project from different angles, obtaining visible light and infrared images from multiple angles;
[0064] When using a drone equipped with visible light and infrared cameras to capture images of the equipment and scenes in the power project from different angles to obtain visible light and infrared images from multiple angles, the visible light and infrared images correspond one by one, and the overlap rate between adjacent images from different angles reaches more than 60%.
[0065] The collected images need to contain sufficient feature information and ensure that there is a certain overlapping area between adjacent photos for subsequent feature matching. To ensure the comprehensiveness and accuracy of the collected images, before shooting, it is necessary to accurately plan the shooting route and shooting angle of the drone according to the distribution, scale of the power project facilities and the on-site environmental conditions through a path planning algorithm (such as the drone shooting path planning optimized based on the Dijkstra algorithm or A* algorithm, etc.). For example, for a large substation, a hierarchical and zonal shooting strategy can be adopted to ensure that each device and component can be clearly photographed.
[0066] Step S2: Use a dual-branch convolutional neural network CNN to extract the features of visible light and infrared images respectively, and use the attention mechanism to dynamically fuse the features of visible light and infrared images to obtain comprehensive image features;
[0067] When using the attention mechanism to dynamically fuse the features of visible light and infrared images to obtain comprehensive image features, the feature space is aligned through a loss function. The specific calculation formula of the feature space alignment loss function is:
[0068]
[0069] Dynamically fuse multi-modal features using the attention mechanism; the adaptive feature weight allocation formula is:
[0070] f fusion =αf v +(1 - α)·f ir ;
[0071] α = σ(W·[f v ·f ir +b);
[0072] In the formula: L align is the feature space alignment loss function; N is the number of samples participating in the calculation of the loss function; is the feature vector of the i-th visible light image; is the feature vector of the i-th infrared image; f fusion is the fused comprehensive image feature vector; α is the adaptive weight coefficient; f vis the feature vector of the visible light image; f ir is the feature vector of the infrared image; σ is the activation function; W is the weight matrix; b is the bias term.
[0073] Step S3: Use the combined image feature vectors of the images taken at different angles as the input of the Siamese neural network. Confirm adjacent images according to the similarity matching results output by the Siamese neural network, and obtain the matching feature points of the adjacent images through the cross-attention mechanism;
[0074] Using the Siamese network structure, take the feature vectors of images taken at different angles as the input, and output the matching results of adjacent images through the Siamese network. Let the feature vectors of two images be f 1 and f 2 , the output of the Siamese network is the similarity score s. The output of the Siamese network is the similarity score s. Train the network by minimizing the loss function. The loss function is defined as:
[0075]
[0076] In the formula, t is the label, t = 1 indicates that the feature vectors match, t = 0 indicates non-matching, m is the margin parameter, and d(f 1 , f 2 ) is the distance between the two feature vectors.
[0077] Based on the matching results of adjacent images, obtain the matching feature points of adjacent images through the cross-attention mechanism; specifically, calculate the attention weights between image features and focus on the key areas; the calculation formula of the attention weights is:
[0078]
[0079] In the formula: a g,h is the attention weight; are the features at the g-th position in image 1 and the features at the h-th position in image 2 respectively; ε is the dimension of the vector features; f match is the feature vector of the matching feature points obtained through the cross-attention mechanism; A g,h is the element in the attention weight matrix, corresponding to a g,h , and is used to represent the weight relationship between features at different positions.
[0080] Step S4: Stitch adjacent images based on the matching feature points of adjacent images, and perform preliminary 3D reconstruction through the principle of triangulation to obtain a sparse 3D point cloud model;
[0081] Based on the matched feature point pairs, 3D reconstruction is performed using the principle of triangulation. Let the projection matrix of Image 1 be P 1 , and the projection matrix of Image 2 be P 2 . For a pair of matched feature points x 1 and x 2 , the corresponding 3D point X satisfies the following relationship:
[0082] x 1 = P 1 X;
[0083] x 2 = P 2 X;
[0084] By solving the above equations, the coordinates of the 3D points can be obtained. 3D reconstruction is performed on all the matched feature point pairs to obtain a sparse 3D point cloud model. When solving the equations, considering the possible noise and errors in actual measurements, optimization algorithms such as the least squares method are used to estimate the 3D point coordinates to improve the accuracy of 3D reconstruction. At the same time, in order to improve the computational efficiency, parallel computing technologies such as GPU acceleration are used to perform parallel processing on a large number of feature point pairs.
[0085] Step S5, based on the deep learning-based surface reconstruction algorithm, optimize and refine the sparse 3D point cloud model to generate a complete and smooth 3D model.
[0086] In the process of optimizing and refining the sparse 3D point cloud model based on the deep learning-based surface reconstruction algorithm to generate a complete and smooth 3D model, deep learning is combined with multi-view stereo matching to improve the point cloud density and accuracy; the specific formula for generating a dense point cloud is:
[0087] Depth Map = softmax(Conv3D(I 1 , I 2 , …, I n ));
[0088] Where: Depth Map is the depth map generated after three-dimensional convolution of images taken from multiple different angles; Conv3D is the three-dimensional convolution operation; (I 1 , I 2 , …, I n ) are the images taken from different angles.
[0089] In the process of generating the dense point cloud, statistical filtering is used to remove outliers, and the specific calculation formula is:
[0090]
[0091] Retain those that satisfy |d jPoints where |-μ| < 3δ;
[0092] Where: μ is the mean of the data in a certain dimension of the point cloud data, used to measure the central tendency of the data in this dimension, and is used as a reference benchmark when removing outliers; J is the number of points participating in the statistics; d j is the value of the j-th point in a certain dimension; δ is the standard deviation of the data in a certain dimension of the point cloud data, used to measure the fluctuation of the data, and is used as a threshold reference when judging outliers.
[0093] In the process of optimizing and refining the sparse three-dimensional point cloud model to generate a complete and smooth three-dimensional model, a pre-trained 3D-GAN of ShapeNet is used as the basic model, and the basic model is fine-tuned through fine-tuning of power equipment data. The loss function formula of the specific fine-tuning process is:
[0094] L total = L GAN + λL recin ;
[0095] Where:
[0096]
[0097] Where: L total is the total loss function of the fine-tuning process; L GAN is the loss of the generative adversarial network; λ is the weight coefficient; L recon is the reconstruction loss; P is the generated point cloud; Q is the real point cloud; p is the point in the generated point cloud; q is the point in the real point cloud.
[0098] Example 2:
[0099] This embodiment provides a fast generation system for a three-dimensional model of a power project based on multi-angle photos, including:
[0100] A data acquisition module, which uses a drone equipped with visible light and infrared cameras to photograph the equipment and scenes in the power project from different angles to obtain multi-angle visible light and infrared images;
[0101] A feature extraction and fusion module, which uses a dual-branch convolutional neural network CNN to extract the features of visible light and infrared images respectively, and uses an attention mechanism to dynamically fuse the features of visible light and infrared images to obtain comprehensive image features;
[0102] A feature point matching module, which uses the comprehensive image feature vectors of the images taken from different angles as the input of the Siamese neural network to calculate the similarity of the images taken from different angles, and obtains the matching feature points of the images taken from different angles through a cross-attention mechanism;
[0103] The image stitching and preliminary modeling module stitches the captured images based on the matching feature points of the images captured from different angles, and performs preliminary 3D reconstruction through the principle of triangulation to obtain a sparse 3D point cloud model;
[0104] The model generation module optimizes and refines the sparse 3D point cloud model based on the surface reconstruction algorithm of deep learning to generate a complete and smooth 3D model.
[0105] Embodiment 3:
[0106] This embodiment provides an electronic device with a computer program stored thereon. When the computer program is executed by a processor, it implements the method for quickly generating a 3D model of a power project based on multi-angle photos as described in any embodiment of the present invention.
[0107] Embodiment 4:
[0108] This embodiment provides a computer-readable medium for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method for quickly generating a 3D model of a power project based on multi-angle photos as described in any embodiment of the present invention.
[0109] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent the case where A exists alone, A and B exist simultaneously, or B exists alone. Where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, and c may represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c may be single or multiple.
[0110] Those of ordinary skill in the art can realize that the units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0111] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.
[0112] In several embodiments provided by the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (hereinafter referred to as ROM), random access memory (hereinafter referred to as RAM), magnetic disks, or optical discs that can store program codes.
[0113] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for rapidly generating a three-dimensional model of an electric power engineering project based on multi-angle photos, characterized in that: The following steps are involved: Step S1, using a drone equipped with visible light and infrared cameras to photograph equipment and scenes in the power project from different angles to obtain multi-angle visible light and infrared images; Step S2, using a dual-branch convolutional neural network (CNN) to extract visible light and infrared image features respectively, and using an attention mechanism to dynamically fuse the visible light and infrared image features to obtain a comprehensive image feature; Step S3, using the obtained comprehensive image feature vectors of the images taken at different angles as the input of the Siamese neural network, confirming the adjacent images according to the similarity matching results output by the Siamese neural network, and obtaining the matching feature points of the adjacent images through the cross attention mechanism; Step S4, stitching adjacent images based on matching feature points of adjacent images, performing preliminary 3D reconstruction through triangulation principle, and obtaining a sparse 3D point cloud model; Step S5: Optimize and refine the sparse 3D point cloud model based on a deep learning surface reconstruction algorithm to generate a complete and smooth 3D model.
2. The method for rapidly generating a three-dimensional model of an electric power engineering project based on multi-angle photos according to claim 1 is characterized in that: The method uses a drone equipped with visible light and infrared cameras to shoot equipment and scenes in power engineering from different angles to obtain multi-angle visible light and infrared images, wherein the visible light and infrared images correspond one to one, and the overlap rate between adjacent images at different angles reaches more than 60%.
3. The method for rapidly generating a three-dimensional model of an electric power engineering project based on multi-angle photos according to claim 1 is characterized in that: The dual-branch convolutional neural network CNN is used to extract visible light and infrared image features respectively. The specific process includes: Gaussian convolution kernel constructs scale space, and the specific calculation formula is: L(x,y,σ)=G(x,y,σ)*I(x,y); Where: L(x,y,σ) is the image in different scale spaces; σ is the scale factor; G(x,y,σ) is the Gaussian kernel function; I(x,y) is the original image; Gaussian difference extreme point detection, the specific calculation formula is: D(x,y,σ)=L(x,y,kσ)-L(x,y,σ); Where: D(x, y, σ) is the Gaussian difference image; k is the scale interval factor.
4. The method for rapidly generating a three-dimensional model of an electric power engineering project based on multi-angle photos according to claim 1, characterized in that: The attention mechanism is used to dynamically fuse the visible light and infrared image features to obtain comprehensive image features, and the feature space is aligned through the loss function. The specific calculation formula of the alignment loss function is: Use attention mechanism to dynamically fuse multimodal features; Adaptive The feature weight allocation formula is: f fusion =αf v +(1-a)·f ir ; α=σ(W·[f v ·f ir ]+b); Where: L align is the feature space alignment loss function; N is the number of samples involved in the loss function calculation; is the feature vector of the i-th visible light image; is the feature vector of the i-th infrared image; f fusion is the integrated image feature vector after fusion; α is the adaptive weight coefficient; f v is the feature vector of the visible light image; f ir is the feature vector of the infrared image; σ is the activation function; W is the weight matrix; b is the bias term.
5. The method for rapidly generating a three-dimensional model of an electric power engineering project based on multi-angle photos according to claim 1, characterized in that: In the surface reconstruction algorithm based on deep learning, the sparse 3D point cloud model is optimized and refined to generate a complete and smooth 3D model. Deep learning and multi-view stereo matching are combined to improve the point cloud density and accuracy. The specific formula for generating dense point cloud is: Depth Map=softmax(Conv3D(I1,I2,…,I n )); Where: Depth Map is the depth map generated by three-dimensional convolution of images taken at multiple angles; Conv3D is the three-dimensional convolution operation; (I1, I2, …, I n ) are images taken at different angles.
6. The method for rapidly generating a three-dimensional model of an electric power engineering project based on multi-angle photos according to claim 5 is characterized in that: In the process of generating dense point cloud, statistical filtering is used to remove outliers. The specific calculation formula is: Keep satisfied |d j - points where μ|<3δ; Where: μ is the mean value of a certain dimension of the point cloud data, which is used to measure the central tendency of the data in this dimension and serves as a reference for removing outliers; J is the number of points involved in the statistics; d j is the value of the jth point in a certain dimension; δ is the standard deviation of the data in a certain dimension in the point cloud data, which is used to measure the fluctuation of the data and serves as a threshold reference when judging outliers.
7. The method for rapidly generating a three-dimensional model of an electric power engineering project based on multi-angle photos according to claim 1, characterized in that: In the process of optimizing and refining the sparse three-dimensional point cloud model to generate a complete and smooth three-dimensional model, the 3D-GAN pre-trained by ShapeNet is used as the basic model, and the basic model is fine-tuned by fine-tuning the power equipment data. The loss function formula of the specific fine-tuning process is: THE total =L GAN +λL recon ; in: Where: L total is the total loss function of the fine-tuning process; L GAN is the loss of the generated adversarial network; λ is the weight coefficient; L recon is the reconstruction loss; P is the generated point cloud; Q is the real point cloud; p is the point in the generated point cloud; q is the point in the real point cloud.
8. A rapid generation system of three-dimensional models of electric power engineering projects based on multi-angle photos, characterized in that: include: The data acquisition module uses drones equipped with visible light and infrared cameras to shoot equipment and scenes in power engineering from different angles to obtain multi-angle visible light and infrared images; The feature extraction and fusion module uses a dual-branch convolutional neural network (CNN) to extract visible light and infrared image features respectively, and uses the attention mechanism to dynamically fuse the visible light and infrared image features to obtain comprehensive image features; The feature point matching module uses the comprehensive image feature vectors of images taken at different angles as the input of the Siamese neural network to calculate the similarity of images taken at different angles, and obtains the matching feature points of images taken at different angles through the cross-attention mechanism; Image stitching and preliminary modeling module stitches the images based on the matching feature points of images taken at different angles, performs preliminary 3D reconstruction through the principle of triangulation, and obtains a sparse 3D point cloud model; The model generation module optimizes and refines the sparse 3D point cloud model based on the surface reconstruction algorithm of deep learning to generate a complete and smooth 3D model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for quickly generating a three-dimensional model of an electric power engineering project based on multi-angle photos as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for quickly generating a three-dimensional model of an electric power engineering project based on multi-angle photos as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Three-dimensional reconstruction method based on attention mechanism and monocular multi-view angle
CN113838191A