Three-dimensional reconstruction method, semantic recognition method and system of multi-view 2D image
By collecting and processing multi-view two-dimensional images of power scenes, generating dynamic masks, extracting features and deriving depth information, the problems of three-dimensional reconstruction accuracy and integrity in complex power facility scenarios are solved, and efficient and accurate three-dimensional reconstruction is achieved.
Patent Information
- Application Number
- CN202510199936.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-03
AI Technical Summary
Existing multi-view stereo reconstruction, neural radiation field (NeRF) and self-supervised learning methods are difficult to ensure the accuracy and integrity of three-dimensional reconstruction in complex scenarios of power facilities, especially when dynamic objects exist.
By collecting two-dimensional images from different perspectives of the power scene, a dynamic mask is generated to eliminate dynamic objects, perform two-dimensional feature extraction and depth value spatial mapping, generate a depth feature map and deduce three-dimensional spatial coordinates, and finally generate a three-dimensional point cloud.
The accuracy and efficiency of three-dimensional reconstruction of power scenes are improved, the accuracy of three-dimensional reconstruction is ensured, and the impact of dynamic objects on reconstruction is reduced.
Smart Images

Figure CN120088407A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power inspection, and particularly to a three-dimensional reconstruction method, semantic recognition method, and system for multi-view Figure 2 D images. Background Art
[0002] With the continuous expansion of the scale of power facilities and the increasing extensiveness of their distribution areas, the traditional manual inspection method is unable to meet the requirements of efficient, safe, and comprehensive coverage of power facility inspections. Against this background, the drone inspection technology has gradually become an indispensable important inspection tool in the power industry due to its high flexibility, fast response speed, and wide coverage ability. However, it is difficult for the drone inspection technology to directly obtain three-dimensional point cloud data rich in geometric features. Therefore, the research on deriving three-dimensional geometric information from two-dimensional images from multiple perspectives is crucial.
[0003] Currently, mainly through multi-view stereo reconstruction methods, three-dimensional reconstruction through Neural Radiance Fields (NeRF), and self-supervised learning three-dimensional reconstruction methods, but these methods all have their defects. Among them, the multi-view stereo reconstruction method is difficult to complete a complete three-dimensional reconstruction while maintaining high accuracy in complex power facility scenarios, especially when dealing with dynamic objects; while the method of three-dimensional reconstruction through Neural Radiance Fields (NeRF) is difficult to eliminate the influence of dynamic objects on three-dimensional reconstruction, and the accuracy of its three-dimensional reconstruction depends on a large amount of labeled data, so it is difficult to ensure the accuracy of three-dimensional reconstruction; while the self-supervised learning three-dimensional reconstruction method trains a depth estimation model through the photometric consistency loss in multi-view images, and it is easy to produce multi-view inconsistency problems in dynamic scenes, affecting the accuracy of three-dimensional reconstruction and further affecting the subsequent semantic recognition results.
[0004] Application Content
[0005] This application provides a three-dimensional reconstruction method, semantic recognition method, and system for multi-view Figure 2 D images. The three-dimensional reconstruction method of this application can improve the accuracy and efficiency of three-dimensional reconstruction of power scenes, and the semantic recognition method based on the three-dimensional reconstruction method can improve the accuracy of three-dimensional semantic recognition of power scenes.
[0006] In a first aspect, this application provides a three-dimensional reconstruction method for multi-view Figure 2 D images, including:
[0007] Collect two-dimensional images of different perspectives of a power scene to obtain a first image set, where the first image set contains several two-dimensional images of different perspectives;
[0008] Generate a corresponding dynamic mask for each dynamic object in the first image set, and filter the pixel values in the first image set based on the dynamic mask to obtain a second image set;
[0009] Extract features from the second image set to obtain two-dimensional features corresponding to each two-dimensional image, map the two-dimensional features to the depth value space to obtain a preliminary feature map, and determine a corresponding depth feature map based on the preliminary feature map;
[0010] Determine the three-dimensional space coordinates corresponding to each two-dimensional coordinate based on the depth feature map, and generate a three-dimensional point cloud corresponding to the power scene based on each three-dimensional space coordinate.
[0011] In the embodiment of the present application, by collecting two-dimensional images of different perspectives of the power scene, it is convenient to process the two-dimensional images subsequently and provide a data basis for subsequent three-dimensional reconstruction; generating a corresponding dynamic mask for each dynamic object in the first image set can effectively monitor the dynamic objects in the first image set, facilitate subsequent removal of the dynamic objects in the first image set, and obtain a second image set that only contains static objects, which can prevent the influence of dynamic objects on subsequent three-dimensional reconstruction, and thus ensure the accuracy of three-dimensional reconstruction; by extracting features from the second image set, the two-dimensional features corresponding to each two-dimensional image in the second image set can be extracted, which is convenient for subsequent three-dimensional reconstruction; mapping the two-dimensional features to the depth value space can obtain more effective data and effectively solve the problem of insufficient three-dimensional annotation data. Determining a corresponding depth feature map based on the preliminary feature map can accurately infer a depth feature map containing image depth information and improve the accuracy of subsequent three-dimensional reconstruction; determining the three-dimensional space coordinates corresponding to each two-dimensional coordinate based on the depth feature map can accurately deduce the corresponding three-dimensional space coordinates, which is convenient for subsequent rapid three-dimensional reconstruction; generating a three-dimensional point cloud corresponding to the power scene based on each three-dimensional space coordinate can automatically, accurately and quickly perform three-dimensional reconstruction based on the three-dimensional space coordinates. Compared with the prior art, the three-dimensional reconstruction method of the present application can improve the accuracy and efficiency of three-dimensional reconstruction of the power scene.
[0012] Further, the generating a corresponding dynamic mask for each dynamic object in the first image set is specifically:
[0013] Calculate the motion vector corresponding to each pixel in the first image set;
[0014] Input each motion vector and the first image set into a preset neural network model to detect the dynamic objects in the first image set and obtain a dynamic object detection result;
[0015] Generate a corresponding dynamic mask based on the dynamic object detection result.
[0016] In this way, by generating corresponding dynamic masks for each dynamic object in the first image set, the dynamic objects in the first image set can be effectively monitored, facilitating the subsequent removal of the dynamic objects in the first image set to obtain a second image set containing only static objects, preventing the influence of dynamic objects on subsequent 3D reconstruction, and thus ensuring the accuracy of 3D reconstruction.
[0017] Further, the mapping of the two-dimensional features to the depth value space to obtain a preliminary feature map is specifically as follows:
[0018] Fuse the two-dimensional features corresponding to each two-dimensional image to obtain corresponding fused features;
[0019] Perform feature embedding processing on the fused features to obtain corresponding embedded features;
[0020] Perform position encoding and depth encoding on the embedded features to obtain position encoding features and depth encoding features respectively;
[0021] Based on a preset depth regression network model, map the embedded features, the position encoding features, and the depth encoding features to the depth value space to obtain corresponding preliminary feature maps.
[0022] In this way, by mapping the two-dimensional features to the depth value space, more effective data can be obtained, effectively solving the problem of insufficient 3D annotation data, and thus improving the accuracy of subsequent 3D reconstruction.
[0023] Further, the determination of the corresponding depth feature map based on the preliminary feature map is specifically as follows:
[0024] Input the preliminary feature map into a preset depth estimation model, and infer the depth information corresponding to each pixel in the preliminary feature map based on photometric consistency constraints;
[0025] Calculate a photometric consistency loss function based on the depth information, determine the minimum loss value, and determine the optimal depth information corresponding to each pixel based on the minimum loss value;
[0026] Determine the corresponding depth feature map based on the optimal depth information.
[0027] In this way, determining the corresponding depth feature map based on the preliminary feature map can accurately infer a depth feature map containing image depth information and improve the accuracy of subsequent 3D reconstruction.
[0028] Further, the generation of a 3D point cloud corresponding to the power scene based on each of the 3D spatial coordinates is specifically as follows:
[0029] Extract the feature points in each depth feature map and match the feature points to obtain feature point pairs;
[0030] Determine the optimal three-dimensional spatial coordinates corresponding to each two-dimensional coordinate based on the feature point pairs;
[0031] Perform multi-view fusion on the three-dimensional spatial coordinates and the optimal three-dimensional spatial coordinates to obtain a three-dimensional point cloud corresponding to the power scene.
[0032] In this way, generating a three-dimensional point cloud corresponding to the power scene based on each of the three-dimensional spatial coordinates can automatically, accurately, and quickly perform three-dimensional reconstruction based on the three-dimensional spatial coordinates.
[0033] Further, after generating the three-dimensional point cloud corresponding to the power scene based on each of the three-dimensional spatial coordinates, it further includes:
[0034] Perform point cloud alignment on the three-dimensional point cloud based on a preset rigid transformation matrix to obtain an initial point cloud optimization result, where the rigid transformation matrix includes a first rotation matrix and a first translation matrix;
[0035] Calculate a second rotation matrix and a second translation matrix for each point cloud in the initial point cloud optimization result based on the least squares method, and perform alignment on the point clouds in the initial point cloud optimization result based on the second rotation matrix and the second translation matrix to obtain a target three-dimensional point cloud.
[0036] In this way, by performing multiple alignments on the three-dimensional point cloud, it can be ensured that the generated target three-dimensional point cloud has a small error from the actual power scene and is more in line with the actual power scene.
[0037] In a second aspect, the present application also provides a method for semantic recognition of multi-view Figure 2 D images, including:
[0038] According to the three-dimensional reconstruction method of the multi-view Figure 2 D images described in the present application, generate a three-dimensional point cloud corresponding to the power scene;
[0039] Perform semantic segmentation on the three-dimensional point cloud to obtain each power facility in the power scene.
[0040] According to the three-dimensional reconstruction method of the multi-view Figure 2 D images described in the embodiments of the present application, three-dimensional reconstruction can be accurately and quickly completed to obtain a target three-dimensional point cloud that is more in line with the actual power scene; by performing semantic segmentation on the three-dimensional point cloud, each power facility in the power scene can be accurately identified, realizing precise annotation of the three-dimensional point cloud, and providing an intelligent three-dimensional semantic recognition ability for inspection drones. Compared with the prior art, the present application can improve the accuracy of three-dimensional semantic recognition of the power scene.
[0041] Further, the performing semantic segmentation on the three-dimensional point cloud to obtain each power facility specifically is:
[0042] Voxelize the three-dimensional point cloud to obtain a number of voxel blocks;
[0043] Extract features for each of the voxel blocks to obtain a number of corresponding local spatial features and global spatial features;
[0044] Based on the local spatial features and the global spatial features, determine the label probabilities of each of the voxel blocks, and obtain each power facility in the power scenario based on the label probabilities.
[0045] In this way, by extracting features for each voxel block and generating the label probability corresponding to each voxel block, the three-dimensional point cloud can be accurately semantically segmented, and then each power facility in the power scenario can be accurately identified, realizing the precise annotation of the three-dimensional point cloud and providing the inspection UAV with intelligent three-dimensional semantic recognition capabilities.
[0046] In a third aspect, the present application provides a semantic recognition system for multi-view Figure 2 3D images, including: a three-dimensional reconstruction module and a segmentation module;
[0047] The three-dimensional reconstruction module is used to generate a three-dimensional point cloud corresponding to the power scenario according to the three-dimensional reconstruction method of the multi-view Figure 2 3D images of the present application;
[0048] The segmentation module is used to perform semantic segmentation on the three-dimensional point cloud to obtain each power facility in the power scenario.
[0049] According to the three-dimensional reconstruction method of the multi-view Figure 2 3D images of the present application, the embodiments of the present application can accurately and quickly complete three-dimensional reconstruction to obtain a target three-dimensional point cloud that better conforms to the actual power scenario; by performing semantic segmentation on the three-dimensional point cloud, each power facility in the power scenario can be accurately identified, realizing the precise annotation of the three-dimensional point cloud and providing the inspection UAV with intelligent three-dimensional semantic recognition capabilities. Compared with the prior art, the present application can improve the accuracy of three-dimensional semantic recognition of the power scenario.
[0050] Further, the segmentation module includes: a first processing unit, an extraction unit, and a second processing unit;
[0051] The first processing unit is used to voxelize the three-dimensional point cloud to obtain a number of voxel blocks;
[0052] The extraction unit is used to extract features for each of the voxel blocks to obtain a number of corresponding local spatial features and global spatial features;
[0053] The second processing unit is configured to determine the label probabilities of the voxel blocks based on the local spatial features and the global spatial features, and obtain each power facility in the power scene based on the label probabilities. Description of the Drawings
[0054] Figure 1 is a flowchart of an embodiment of a method for three-dimensional reconstruction of multi-view Figure 2 D images provided by the present application;
[0055] Figure 2 is Figure 1 a flowchart of step S102 in
[0056] Figure 3 is Figure 1 a flowchart of step S103 in
[0057] Figure 4 is Figure 3 a flowchart of step S302 in
[0058] Figure 5 is Figure 3 a flowchart of step S303 in
[0059] Figure 6 is a flowchart of an embodiment of a method for semantic recognition of multi-view Figure 2 D images provided by the present application;
[0060] Figure 7 is Figure 6 a flowchart of step S602 in
[0061] Figure 8 is a flowchart of an embodiment of a method for semantic recognition of multi-view Figure 2 D images provided by the present application; Detailed Embodiments
[0062] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0063] It should be understood that the step numbers used in the text are only for convenience of description and do not limit the execution order of the steps.
[0064] It should be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an" and "the" are intended to include plural forms unless the context clearly indicates otherwise.
[0065] The terms “include” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0066] The term "and / or" means and includes any and all possible combinations of one or more of the associated listed items.
[0067] As power facilities expand in scale and become more widely distributed, manual inspections are unable to meet the needs of efficient and safe inspections. Drone inspection technology has become an indispensable and important inspection tool for the power industry due to its flexibility, rapid response and wide coverage. However, it is difficult for drones to directly obtain three-dimensional point cloud data, so research on deriving three-dimensional geometric information from two-dimensional images is crucial. Current three-dimensional reconstruction methods include multi-view stereo reconstruction, neural radiance field (NeRF) and self-supervised learning, but all three have defects. Among them, multi-view stereo reconstruction is difficult to ensure accuracy and integrity in complex scenes; NeRF is difficult to eliminate the impact of dynamic objects and relies on a large amount of labeled data; and self-supervised learning is prone to multi-perspective inconsistencies in dynamic scenes, which will affect three-dimensional reconstruction and subsequent semantic recognition.
[0068] Next, the nouns involved in this application are analyzed:
[0069] 3D reconstruction is a process of creating a 3D model from 2D images or scanned data using computer technology. This technology involves a series of complex computational steps, including data acquisition, image processing, geometric analysis, and model optimization. The core idea is to use multi-view geometry to infer the 3D structure of a scene or object using multiple 2D images from different perspectives.
[0070] Optical Flow Computation (FlowNet), a convolutional neural network for optical flow estimation, can estimate motion vectors by learning the pixel displacement between image pairs.
[0071] Position encoding can introduce the relative position of each feature pixel in space so that the network can perceive the differences between different positions and better understand the perspective information.
[0072] The photometric consistency constraint is based on the assumption that the photometric values (brightness or color) of the same scene should remain consistent under different viewpoints. By comparing the photometric values of images from different viewpoints, the model can infer the depth information of the objects in the scene.
[0073] Voxelization refers to dividing the three-dimensional space into grids, where each voxel (i.e., grid cell) represents a region in the three-dimensional space.
[0074] Based on this, the embodiments of the present application provide a three-dimensional reconstruction method, a semantic recognition method, and a system for multi-view Figure 2 D images. The three-dimensional reconstruction method of the present application can improve the accuracy and efficiency of the three-dimensional reconstruction of the power scene, and the semantic recognition method based on the three-dimensional reconstruction method can improve the accuracy of the three-dimensional semantic recognition of the power scene.
[0075] The three-dimensional reconstruction method, semantic recognition method, and system for multi-view Figure 2 D images provided by the embodiments of the present application are specifically described through the following embodiments. First, the three-dimensional reconstruction method for multi-view Figure 2 D images in the embodiments of the present application is described.
[0076] The three-dimensional reconstruction method for multi-view Figure 2 D images provided by the embodiments of the present application relates to the field of power inspection. The three-dimensional reconstruction method for multi-view Figure 2 D images provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can also be software running on the terminal or the server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application for implementing the three-dimensional reconstruction method for multi-view Figure 2 D images, etc., but is not limited to the above forms.
[0077] This application can also be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0078] Embodiment 1
[0079] Please refer to Figure 1 , Figure 1 which Figure 2 is a schematic flowchart of an embodiment of the 3D reconstruction method for multi-view
[0080] Step S101: Collect two-dimensional images of different perspectives of the power scene to obtain a first image set, where the first image set contains several two-dimensional images of different perspectives;
[0081] In some embodiments, multiple shooting points or the positions of image acquisition devices can be determined according to the characteristics and requirements of power facilities in the power scene, and two-dimensional images of the power scene can be collected from different perspectives through the image acquisition devices to obtain a first image set containing several images of different perspectives. Among them, the shooting points should cover the entire scene as much as possible while avoiding overlapping perspectives or blind spots.
[0082] In some embodiments, by deploying multiple cameras on an unmanned aerial vehicle and controlling the unmanned aerial vehicle to reach the designated shooting points to capture two-dimensional images of the power scene from different perspectives, a first image set containing several images of different perspectives can be obtained, laying a data foundation for subsequent 3D reconstruction.
[0083] In some embodiments, the image acquisition device can be, but is not limited to, a camera, a high-resolution camera, etc.
[0084] In some embodiments, the power scene should include key facilities such as power towers and transmission lines.
[0085] Step S102: Generate a corresponding dynamic mask for each dynamic object in the first image set, and filter the pixel values in the first image set based on the dynamic mask to obtain a second image set;
[0086] Please refer to Figure 2 , step S102 includes but is not limited to steps S201 to S203;
[0087] Step S201: Calculate the motion vector corresponding to each pixel in the first image set;
[0088] In some embodiments, after obtaining the first image set containing several images from different perspectives, it is necessary to calculate the pixel displacement between adjacent frame images to obtain the motion vector corresponding to each pixel.
[0089] In some embodiments, the optical flow can be determined by comparing the pixel intensity changes of adjacent frame images, and then the optical flow field F(x, y) ∈ R between adjacent image pairs can be determined H×W×2 , where H and W are the height and width of the corresponding image, to capture the motion information in the time sequence, that is, the motion vector of the present application. Among them, the calculation formula of the optical flow field is: F(x, y) = (u(x, y), v(x, y)), where u(x, y) and v(x, y) respectively represent the displacement amounts in the horizontal and vertical directions.
[0090] Step S202: Input each of the motion vectors and the first image set into a preset neural network model to detect the dynamic objects in the first image set and obtain a dynamic object detection result;
[0091] In some embodiments, after obtaining the motion vector corresponding to each pixel, it is necessary to input each of the motion vectors and the first image set into a preset neural network model to detect dynamic objects. At this time, after receiving the motion vector, the neural network model will calculate the modulus of the motion vector, that is, the motion distance of each pixel, and judge each of the motion distances with a predefined preset threshold. When the motion distance exceeds the preset threshold, it is considered that the area corresponding to the pixel belongs to a dynamic object, otherwise it is a static object. Among them, the calculation formula of the neural network model is: Dynamic(x, y) is the dynamic object detection result, ||F(x, y)|| is the modulus of the motion scale, that is, the motion distance, T is the preset threshold, 1 represents a dynamic object, and 0 is a static object.
[0092] In some embodiments, the neural network model is based on the combination of a convolutional neural network (CNN) and a long short-term memory network (LSTM). Among them, the CNN is used to extract image features, and the LSTM is used to capture temporal information. The neural network model formed by the combination of the two is used to learn the spatio-temporal features in the image and detect dynamic objects.
[0093] In some embodiments, the dynamic objects can be, but are not limited to, moving vehicles or workers, etc.
[0094] Step S203: Generate a corresponding dynamic mask based on the dynamic object detection result.
[0095] In some embodiments, after obtaining the dynamic object detection result Dynamic(x, y), a dynamic mask M(x, y) can be generated, that is, M(x, y) = Dynamic(x, y). The dynamic mask M(x, y) represents the dynamic object area. When M(x, y) = 1, it is a dynamic object, and when M(x, y) = 0, it is a static object, which is convenient for ignoring these areas during subsequent 3D reconstruction.
[0096] In some embodiments, after obtaining the dynamic mask, the pixel values in the first image set are screened based on the dynamic mask, that is, the pixel values at the positions where the dynamic objects are located are set to 0 and do not participate in subsequent feature extraction, while the pixel values of the static objects are retained for subsequent feature extraction. After removing all the dynamic masks in the first image set, the second image set can be obtained. The relevant formula is: I'(x, y) = I(x, y) · (1 - M(x, y)), where I′ is the second image set, I is the first image set, and M(x, y) is the dynamic mask.
[0097] In this way, by generating a corresponding dynamic mask for each dynamic object in the first image set, the dynamic objects in the first image set can be effectively monitored, which is convenient for removing the dynamic objects in the first image set subsequently, obtaining the second image set that only contains static objects, preventing the influence of dynamic objects on subsequent 3D reconstruction, and thus ensuring the accuracy of 3D reconstruction.
[0098] Step S103: Extract features from the second image set to obtain the two-dimensional features corresponding to each two-dimensional image, map the two-dimensional features to the depth value space to obtain a preliminary feature map, and determine the corresponding depth feature map based on the preliminary feature map;
[0099] Please refer to Figure 3 , step S103 includes, but is not limited to, steps S301 to S303;
[0100] Step S301: Extract features from the second image set to obtain the two-dimensional features corresponding to each two-dimensional image;
[0101] Step S302: Map the two-dimensional features to the depth value space to obtain a preliminary feature map;
[0102] Step S303: Determine the corresponding depth feature map based on the preliminary feature map;
[0103] In step S301 of some embodiments, the second image set is input into a convolutional neural network for feature extraction to obtain feature representations at different levels. The convolutional neural network is composed of multiple convolutional layers and multiple pooling layers stacked alternately. Each convolutional layer performs a convolution operation on the second image through multiple convolutional kernels to extract local low-level features. Through convolution processing, initial two-dimensional features can be obtained, and the initial two-dimensional features are input into the pooling layer to extract higher-level features. Through the alternating processing of multiple convolutional layers and multiple pooling layers, two-dimensional features at different levels can be obtained.
[0104] In some embodiments, the convolution calculation formula is: In the formula, is the element at position (i, j) and depth d in the initial two-dimensional features, W u,v,c,d is the convolutional kernel parameter, representing the weight of the d-th convolutional kernel at (u, v) and input channel c, I′ s·i+u,s·j+v,c is the element of the second image set, representing the element at position (s·i + u, s·j + v) and channel c in the second image set I′, b d is the bias term related to the d-th convolutional kernel. The size of the convolutional kernel is k×k, and the stride is s.
[0105] In some embodiments, the pooling layer can be but is not limited to max pooling or average pooling, etc. The calculation formula for the max pooling operation is: In the formula, are two-dimensional features.
[0106] It should be noted that the convolutional layer includes two types: shallow convolutional layer and deep convolutional layer. Among them, the shallow convolutional layer mainly extracts low-level features such as edges and textures, while the deep convolutional layer can capture high-level features such as the shape and category of objects.
[0107] In this way, by performing feature extraction on the second image set, the two-dimensional features corresponding to each two-dimensional image in the second image set can be extracted, which is convenient for subsequent 3D reconstruction.
[0108] Please refer to Figure 4 , step S302 includes but is not limited to steps S401 to S404;
[0109] Step S401: Perform feature fusion on the two-dimensional features corresponding to each two-dimensional image to obtain the corresponding fused features;
[0110] In some embodiments, since the second image set includes two-dimensional images from different perspectives, and when extracting features from the two-dimensional images, the same or similar features may be extracted multiple times. Therefore, it is necessary to fuse the two-dimensional features from different perspectives, which can be understood as fusing the two-dimensional features extracted from different images to obtain fused features. Among them, the fusion method is weighted fusion, and the fusion calculation formula is: F fusion = α·F 1 +(1 - α)·F 2 ; In the formula, F fusion is the fused feature, α is the weight parameter for controlling feature fusion, and F 1 and F 2 are any two-dimensional images in the second image set respectively.
[0111] It should be noted that the weight parameter can be obtained through learning or manually set according to the specific application scenario.
[0112] It should be noted that when performing feature fusion, an attention mechanism needs to be added to adjust the attention weight, enhance key information while suppressing repeated or irrelevant features. The formula for adjusting the attention weight is: F′ = F fusion ·A; In the formula, F′ is the adjusted fused feature, F fusion is the fused feature before adjustment, and A is the attention weight matrix, which can be obtained through learning and is used to perform weighted adjustment on the feature F.
[0113] Step S402: Perform feature embedding processing on the fused feature to obtain the corresponding embedded feature;
[0114] In some embodiments, the fused feature F′ can be, but is not limited to, subjected to feature embedding processing through linear mapping or non-linear transformation to extract the features of the two-dimensional image into a higher-dimensional feature space, obtaining the corresponding embedded feature F″, which provides a basis for subsequent position encoding and depth encoding.
[0115] In some embodiments, the calculation formula for feature embedding through linear projection is: F″ = F′×W + b; In the formula, F″ is the embedded feature, F′∈R H×W×D , H and W are the height and width of the feature map, D is the number of feature channels; F′ is the fused feature, F′∈R H×W×C , H and W are the height and width of the corresponding map of the fused feature, C is the number of feature channels, where D > C; W is the weight matrix, W∈R C ×D ; b is the bias vector, b∈R D .
[0116] It should be noted that the embedded features obtained after the dimensionality increase of the feature embedding are not real geometric three dimensions, but the dimensionality of the image features is increased so that the representation in the feature space can contain additional information such as depth and perspective.
[0117] Step S403: Perform position encoding and depth encoding on the embedded features to obtain position-encoded features and depth-encoded features respectively.
[0118] In some embodiments, after obtaining the embedded feature F″, in order to include information such as depth and perspective in the feature space, it is necessary to perform position encoding and depth encoding on it to obtain position-encoded features and depth-encoded features respectively. Among them, position encoding can add position information to the feature representation, facilitating the subsequent model to recognize the position differences of different features, and then representing perspective-related information, while depth encoding can directly add the depth information as an additional channel to the feature representation.
[0119] In some embodiments, the embedded feature F″ can be position-encoded through the sine and cosine position encoding formula to obtain position-encoded features, and its calculation formula is: In the formula, pos is the spatial position of the feature, such as the (x, y) coordinates of the image; i is the dimension index of the position encoding; D is the total embedding dimension.
[0120] In some embodiments, the depth-encoded features can be estimated through a depth estimation model, and its calculation formula is: F″′ = concat(F″, D), where F″′ is the depth-encoded feature, F″ ∈ R H×W(D+1) , H and W are the height and width of the encoded feature corresponding map, D is the number of feature channels; F″ is the embedded feature; concat is the concatenation operation in the feature dimension.
[0121] In some embodiments, position encoding and depth encoding can be performed through a sliding transformer. Specifically: First, divide the map of the embedded feature into multiple non-overlapping windows of size M×M, and each window contains M 2 pixels. Assuming the size of the image feature map is H×W, the divided image representation windows. Then, apply the self-attention mechanism to each window separately instead of calculating for the entire image, and the calculation formula is: where Q = XW Q , K = XW K , V = XW V are the query, key, and value matrices, X is the feature within the window, W Q , W K , W Vis a linear transformation matrix. Again, the sliding window transformer adopts a hierarchical architecture with different window sizes at each layer to extract multi-scale features of the image at different scales. This sliding window block processing method not only improves the coding efficiency but also retains local information, making it suitable for processing high-resolution images.
[0122] Step S404: Map the embedded feature, the position encoding feature, and the depth encoding feature to the depth value space based on a preset deep regression network model to obtain a corresponding preliminary feature map.
[0123] It can be understood that after obtaining the embedded feature, the position encoding feature, and the depth encoding feature, they need to be input into a preset neural regression network model together to map the embedded feature, the position encoding feature, and the depth encoding feature to the depth value space to obtain a preliminary feature map. The calculation formula is D out = f(F1; θ); where D out is the preliminary feature map; f(·) is the mapping function of the deep regression network, usually a neural network. F1 is the input feature map extracted from two-dimensional images of multiple perspectives, including the embedded feature, the position encoding feature, and the depth encoding feature; θ is the parameter of the deep regression network.
[0124] It should be noted that the deep regression network model usually consists of several convolutional layers, pooling layers, fully connected layers, and activation functions, which process the input feature map F1 and map it to a preliminary feature map. Among them, the convolutional layer (Convolutional Layers): is used to extract the spatial features of the input feature map, and the calculation formula is: F1′ = Conv(F1, W) + b; where F1′ is the output feature map after the convolutional operation, F is the input feature map, W is the convolutional kernel, and b is the bias term. The pooling layer (Pooling Layers): The pooling operation (such as max pooling or average pooling) is used to reduce the size of the feature map while retaining important information. The fully connected layer (Fully Connected Layers): maps the feature after the convolutional operation and the pooling operation to the depth value space to obtain a preliminary feature map, and the calculation formula is: D out = FC(F1′, W fc , b fc ), where D out is the preliminary feature map, F1′ is the output feature map after the convolutional operation, W fc is the weight matrix, and b fc is the bias.
[0125] It should be noted that the preliminary feature map is usually a linear layer, which is the same size as the original two-dimensional image, and the value of each pixel represents the depth of that pixel.
[0126] It should be noted that in order to ensure the accuracy of the prediction of the depth regression network model, a loss function needs to be set to measure the error between the depth map output by the network and the true depth map. By iterating the loss function, the initial feature map D is finally obtained. out Among them, the loss function can be but is not limited to the mean squared error loss function, the absolute error loss function, and the structural similarity loss function. Among them, the mean squared error (MSE) calculates the depth D of the predicted depth map out (i, j) and the depth D of the true depth map gt (i, j). The calculation formula is: Among them, i and j are the corresponding pixels, and N is the number of depth maps; the absolute error uses the L1 loss to measure the depth D of the predicted depth map out (i, j) and the depth D of the true depth map gt (i, j). The calculation formula is: Among them, i and j are the corresponding pixels, and N is the number of depth maps; while the structural similarity loss is used to measure the similarity of the image structure, which can help retain the details in the depth map; by combining these loss functions, the final loss can be obtained: L = L MSE + L L1 + λ · L SSIM , where L is the final loss, L MSE is the mean squared error loss, L L1 is the absolute error loss, λ is a weight coefficient, and L SSIM is the structural similarity loss. It should be noted that the loss function is not the focus of this application, so it will not be elaborated here.
[0127] In this way, by mapping the two-dimensional features to the depth value space, more effective data can be obtained, effectively solving the problem of insufficient three-dimensional annotation data, and further improving the accuracy of subsequent three-dimensional reconstruction.
[0128] Please refer to Figure 5 , step S303 includes but is not limited to steps S501 to S503;
[0129] Step S501, input the preliminary feature map into a preset depth estimation model, and infer the depth information corresponding to each pixel in the preliminary feature map based on the photometric consistency constraint;
[0130] In some embodiments, first, the preliminary feature map is input into a preset depth estimation model, where the depth estimation model is a deep convolutional neural network that has been trained to infer depth information from the feature map. Second, the inference of the depth value is optimized through the photometric consistency constraint inside the depth estimation model. Specifically, the depth estimation model attempts to find a depth map such that when images from different perspectives are transformed and aligned according to this depth map, the photometric difference between them is minimized. Exemplarily, assume that there are two images I 1 and I 2 captured from two perspectives (perspective 1 and perspective 2), and the corresponding pixel points are p 1 =(x 1 , y 1 ) and p 2 =(x 2 , y 2 ). If p 1 and p 2 are projections of the same object, then under the photometric consistency assumption, there is I 1 (p 1 ) = I 2 (p 2 ), that is, the two pixel points represent the same object or the same part of the scene. According to the photometric consistency assumption, if p 1 and p 2 are projections of the same object, then their photometric values should be similar or equal. To achieve this, it is necessary to align the pixel points in these two perspectives through depth (or disparity). Assume that the projection matrix of the camera is K, the three-dimensional coordinates are (X, Y, Z), and its projection point p 2 in perspective 2 can be expressed as: p 2 = K·R·(K -1 ·p 1 ·Z) + t; where R and t are the rotation and translation matrices from perspective 1 to perspective 2, and Z is the depth value.
[0131] Step S502, calculate the photometric consistency loss function based on the depth information, determine the minimum loss value, and determine the optimal depth information corresponding to each pixel based on the minimum loss value;
[0132] In some embodiments, to infer the depth information of each pixel, it is necessary to calculate the photometric consistency loss function and minimize this function to ensure that the difference between the photometric values of two pixel points is minimized. Specifically, for each pixel point, calculate the difference in photometric values of the pixel point in images from different viewpoints, and accumulate these differences to form a total loss value. Then, continuously adjust the parameters in the depth estimation model through an optimization algorithm to minimize the photometric consistency function. During the optimization process, the depth estimation model will gradually find a depth map that minimizes the photometric consistency loss, and this depth map is the optimal depth information corresponding to each pixel. The relevant formula is as follows: In the formula, p represents the pixel point in Viewpoint 1, and I 2 p 2 represents the photometric value of the pixel point re-projected in Viewpoint 2.
[0133] Step S503: Determine the corresponding depth feature map based on the optimal depth information.
[0134] In some embodiments, to convert the optimal depth information into a depth feature map, it should be noted that the depth feature map is still a two-dimensional image, where the value of each pixel represents the depth (or distance) of the pixel in three-dimensional space. This depth feature map can be used as the input for subsequent three-dimensional reconstruction.
[0135] In this way, determining the corresponding depth feature map based on the preliminary feature map can accurately infer the depth feature map containing the image depth information and improve the accuracy of subsequent three-dimensional reconstruction.
[0136] Step S104: Determine the three-dimensional space coordinates corresponding to each two-dimensional coordinate based on the depth feature map, and generate a three-dimensional point cloud corresponding to the power scene based on each of the three-dimensional space coordinates.
[0137] In some embodiments, after obtaining the depth feature map, the two-dimensional coordinates (x, y) can be converted into three-dimensional space coordinates (X, Y, Z) by using geometric mapping. This conversion depends on the internal parameter matrix of the image acquisition device to map the pixel coordinates to the image acquisition device coordinate system to obtain the three-dimensional space coordinates (X, Y, Z). Exemplarily, assuming the two-dimensional coordinates are (x, y) and the depth is Z, the internal parameter matrix K is: According to the imaging principle, the three-dimensional coordinates can be obtained through the following formula: In the formula, (X, Y, Z) are the three-dimensional space coordinates, and K -1 is the inverse matrix of the internal parameter matrix, which is used to map the pixel coordinates to the image acquisition device coordinate system.
[0138] In some embodiments, generating a three-dimensional point cloud corresponding to the power scenario based on each of the three-dimensional space coordinates includes: extracting feature points from each of the depth feature maps, and matching the feature points to obtain feature point pairs; determining the optimal three-dimensional space coordinates corresponding to each two-dimensional coordinate based on the feature point pairs; performing multi-view fusion on the three-dimensional space coordinates and the optimal three-dimensional space coordinates to obtain a three-dimensional point cloud corresponding to the power scenario. Specifically: First, extract feature points from each depth feature image. Among them, feature points are usually pixel points with significant changes in the image, such as edges, corners, etc., and use the feature matching algorithm SIFT to align the feature points of the same object in different views to form feature point pairs; Second, for the matched feature point pairs, according to the multi-view geometric relationship, calculate the optimal three-dimensional position through triangulation. Exemplarily, if the positions of feature points (x1, y1) and (x2, y2) in three-dimensional space are (X, Y, Z), the optimal three-dimensional position can be obtained by solving the following equations: where f and c are parameters of the image acquisition device.
[0139] In some embodiments, multiple three-dimensional point clouds obtained from different perspectives or different time points are fused to obtain a three-dimensional point cloud corresponding to the power scenario. At the same time, an optimization algorithm (such as Bundle Adjustment) is used to further optimize the three-dimensional point cloud, and duplicate points in the three-dimensional point cloud are removed, and operations such as smoothing are performed.
[0140] In this way, generating a three-dimensional point cloud corresponding to the power scenario based on each of the three-dimensional space coordinates can perform three-dimensional reconstruction automatically, accurately, and quickly based on the three-dimensional space coordinates.
[0141] In some embodiments, after generating the three-dimensional point cloud corresponding to the power scenario based on each of the three-dimensional space coordinates, it further includes: performing point cloud alignment on the three-dimensional point cloud based on a preset rigid transformation matrix to obtain an initial point cloud optimization result, where the rigid transformation matrix includes a first rotation matrix and a first translation matrix; calculating a second rotation matrix and a second translation matrix for each point cloud in the initial point cloud optimization result based on the least squares method, and aligning the point clouds in the initial point cloud optimization result based on the second rotation matrix and the second translation matrix to obtain a target three-dimensional point cloud. Specifically: First, after obtaining the three-dimensional point cloud, it is necessary to pass the rigid transformation matrix T i to transform the point cloud p i under the i-th perspective into a unified world coordinate system to obtain an initial point cloud optimization result. Among them, the expression of the rigid transformation matrix is In the formula, R i is a 3×3 rotation matrix representing the rotation of the perspective, and t i is a 3×1 translation vector representing the translation of the perspective. Transforming the point cloud P iEach point p in i,j is mapped to the world coordinate system by the formula: p' i,j = T i ·p i,j . Secondly, after obtaining the optimized result of the initial point cloud, since there may be errors and redundancies in the point clouds from different perspectives, the least squares method is needed to solve the point cloud with the minimum alignment error to obtain the second rotation matrix and the second translation matrix of each point cloud. The calculation formula is where R and t are the second rotation matrix and the second translation matrix respectively, p i,k and p j,k are matching point pairs. Finally, based on the second rotation matrix and the second translation matrix, the point clouds from different perspectives in the optimized result of the initial point cloud are aligned. After further alignment, the point cloud data under all perspectives can be merged into the complete target three-dimensional point cloud P final , and its calculation formula is: where P final represents the optimized target three-dimensional point cloud unified in the same coordinate system, T i is the rigid transformation matrix, and P i is a single point cloud.
[0142] In this way, by aligning the three-dimensional point cloud multiple times, it can be ensured that the generated target three-dimensional point cloud has a small error from the actual power scene and is more in line with the actual power scene.
[0143] By collecting two-dimensional images from different perspectives of the power scenario in the embodiments of the present application, it is convenient to process the two-dimensional images subsequently and provide a data basis for subsequent three-dimensional reconstruction. For each dynamic object in the first image set, a corresponding dynamic mask is generated, which can effectively monitor the dynamic objects in the first image set, facilitate the subsequent removal of the dynamic objects in the first image set, and obtain a second image set containing only static objects, which can prevent the influence of dynamic objects on subsequent three-dimensional reconstruction, thereby ensuring the accuracy of three-dimensional reconstruction. By extracting features from the second image set, the two-dimensional features corresponding to each two-dimensional image in the second image set can be extracted, which is convenient for subsequent three-dimensional reconstruction. Mapping the two-dimensional features to the depth value space can obtain more effective data, effectively solve the problem of insufficient three-dimensional annotation data, and determine the corresponding depth feature map based on the preliminary feature map, which can accurately infer the depth feature map containing image depth information and improve the accuracy of subsequent three-dimensional reconstruction. Based on the depth feature map, the three-dimensional space coordinates corresponding to each two-dimensional coordinate can be determined, and the corresponding three-dimensional space coordinates can be accurately deduced, which is convenient for subsequent rapid three-dimensional reconstruction. Based on each of the three-dimensional space coordinates, a three-dimensional point cloud corresponding to the power scenario is generated, and three-dimensional reconstruction can be automatically, accurately, and quickly performed based on the three-dimensional space coordinates. Compared with the prior art, the three-dimensional reconstruction method of the present application can improve the accuracy and efficiency of three-dimensional reconstruction of the power scenario.
[0144] Embodiment 2
[0145] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of an embodiment of the semantic recognition method for multi-view Figure 2 D images provided by the present application, including steps S601 to step S602;
[0146] Step S601, according to the three-dimensional reconstruction method for multi-view Figure 2 D images of the present application, generate a three-dimensional point cloud corresponding to the power scenario;
[0147] In some embodiments, generating a three-dimensional point cloud corresponding to the power scenario according to the three-dimensional reconstruction method has been described in detail in Embodiment 1, so it will not be elaborated here.
[0148] Step S602, perform semantic segmentation on the three-dimensional point cloud to obtain each power facility in the power scenario.
[0149] Please refer to Figure 7 , step S602 may but is not limited to including steps S701 to step S703;
[0150] Step S701, perform voxelization processing on the three-dimensional point cloud to obtain a number of voxel blocks;
[0151] In some embodiments, the voxel resolution can be preset. On the premise of ensuring the voxel resolution, the point cloud points in each three-dimensional point cloud are assigned to the corresponding voxel blocks according to their spatial coordinates. Among them, each voxel block represents the point cloud data within a certain spatial range. At this time, for each three-dimensional point cloud p=(X, Y, Z), the voxel index can be calculated as: Where V x , V y , V z are all voxel resolutions.
[0152] Step S702: Extract features for each of the voxel blocks to obtain a number of corresponding local spatial features and global spatial features;
[0153] In some embodiments, a number of voxel blocks are input into a 3D convolutional neural network to extract spatial features by stacking multiple 3D convolutional layers in the voxel blocks, so as to obtain local spatial features and global spatial features. The calculation method of 3D convolution is: In the formula, V' i,j,k is the three-dimensional spatial feature; W is the parameter of the three-dimensional convolution kernel; the convolution kernel size is D'*H'*W', V is the voxel block, V∈R D×H×W×C , D, H, and W respectively represent the depth, height, and width of the voxel, C represents the number of feature channels; b is the bias term.
[0154] Step S703: Based on the local spatial features and the global spatial features, determine the label probabilities of each of the voxel blocks, and obtain each power facility in the power scene based on the label probabilities.
[0155] In some embodiments, after obtaining the local spatial features and the global spatial features, first, input them into a fully connected layer or a classification head to generate the semantic label probability corresponding to each voxel block. The calculation formula of the label probability is: P(y = k|f i ) = softmax(W·f i +b); in the formula, P is the label probability, W is the weight matrix of the classification head, b is the bias term, the softmax function converts the output into the probability of each category, f i is the feature of the i-th voxel block, and k represents the semantic category (such as electric tower, wire, ground, etc.). Finally, by selecting the label with the highest label probability as the semantic label of the corresponding voxel block, its formula is: In the formula, is the semantic label. After classifying each voxel block as described above, the generated semantic labels can be used to label the category information of each voxel block. In the finally output three-dimensional point cloud scene, each voxel block has a corresponding semantic label, thus realizing the accurate recognition of objects such as electric towers, wires, and ground.
[0156] By extracting features from each voxel block and generating the label probability corresponding to each voxel block, the 3D point cloud can be accurately semantically segmented, and then each power facility in the power scene can be accurately identified, realizing the accurate annotation of the 3D point cloud and providing the inspection UAV with intelligent 3D semantic recognition ability.
[0157] According to the 3D reconstruction method of the multi-view Figure 2 D image described in this application, 3D reconstruction can be accurately and quickly completed to obtain a target 3D point cloud that better conforms to the actual power scene; by performing semantic segmentation on the 3D point cloud, each power facility in the power scene can be accurately identified, realizing the accurate annotation of the 3D point cloud and providing the inspection UAV with intelligent 3D semantic recognition ability. Compared with the prior art, this application can improve the accuracy of 3D semantic recognition in the power scene.
[0158] Embodiment III
[0159] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an embodiment of the semantic recognition system for multi-view Figure 2 D images provided by this application, including: a 3D reconstruction module 100 and a segmentation module 200;
[0160] The 3D reconstruction module 100 is used to generate a 3D point cloud corresponding to the power scene according to the 3D reconstruction method of the multi-view Figure 2 D image described in this application;
[0161] The segmentation module 200 is used to perform semantic segmentation on the 3D point cloud to obtain each power facility in the power scene.
[0162] Regarding the information interaction, execution process, etc. between the modules in the above 3D reconstruction system for multi-view Figure 2 D images, since they are based on the same concept as the embodiment of the 3D reconstruction method of the multi-view Figure 2 D image in the second aspect of the present invention, the achieved technical effects are basically the same. For specific content, please refer to the description in Embodiment I of the method of the present invention, which will not be elaborated here.
[0163] The segmentation module 200 includes: a first processing unit, an extraction unit, and a second processing unit; specifically, the first processing unit is used to voxelize the 3D point cloud to obtain a number of voxel blocks; the extraction unit is used to extract features for each of the voxel blocks to obtain a number of corresponding local spatial features and global spatial features; the second processing unit is used to determine the label probability of each voxel block based on the local spatial features and the global spatial features, and obtain each power facility in the power scene based on the label probability.
[0164] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the method of this embodiment.
[0165] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the multi-view Figure 2 3D reconstruction method of D images as described in the first embodiment above.
[0166] Those of ordinary skill in the art can understand that all or part of the processes in the above-described method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like.
[0167] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above description is only for the specific embodiments of this application and is not used to limit the protection scope of this application.
[0168] It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.
Claims
1. A method for 3D reconstruction of multi-view 2D images, characterized in that: include: Collecting two-dimensional images of the power scene at different perspectives to obtain a first image set, wherein the first image set includes a plurality of two-dimensional images at different perspectives; Generating a corresponding dynamic mask for each dynamic object in the first image set, and filtering pixel values in the first image set based on the dynamic mask to obtain a second image set; Performing feature extraction on the second image set to obtain two-dimensional features corresponding to each two-dimensional image, mapping the two-dimensional features to a depth value space to obtain a preliminary feature map, and determining a corresponding depth feature map based on the preliminary feature map; The three-dimensional space coordinates corresponding to each two-dimensional coordinate are determined based on the depth feature map, and a three-dimensional point cloud corresponding to the power scene is generated based on each of the three-dimensional space coordinates.
2. The method for 3D reconstruction of multi-view 2D images according to claim 1, characterized in that: The generating a corresponding dynamic mask for each dynamic object in the first image set is specifically: Calculate the motion vector corresponding to each pixel in the first image set; Inputting each of the motion vectors and the first image set into a preset neural network model to detect dynamic objects in the first image set to obtain a dynamic object detection result; Based on the dynamic object detection result, a corresponding dynamic mask is generated.
3. The method for 3D reconstruction of multi-view 2D images according to claim 1, characterized in that: The two-dimensional features are mapped to the depth value space to obtain a preliminary feature map, specifically: Performing feature fusion on the two-dimensional features corresponding to each two-dimensional image to obtain corresponding fused features; Performing feature embedding processing on the fused features to obtain corresponding embedded features; Performing position encoding and depth encoding on the embedded features to obtain position encoding features and depth encoding features respectively; Based on a preset deep regression network model, the embedded features, the position encoding features and the depth encoding features are mapped to the depth value space to obtain a corresponding preliminary feature map.
4. The method for 3D reconstruction of multi-view 2D images according to claim 1, characterized in that: The determining of the corresponding depth feature map based on the preliminary feature map is specifically: Inputting the preliminary feature map into a preset depth estimation model, and inferring the depth information corresponding to each pixel in the preliminary feature map based on the photometric consistency constraint; Calculating a photometric consistency loss function based on the depth information, determining a minimum loss value, and determining optimal depth information corresponding to each pixel based on the minimum loss value; A corresponding depth feature map is determined based on the optimal depth information.
5. The method for 3D reconstruction of multi-view 2D images according to claim 1, characterized in that: The generating of the three-dimensional point cloud corresponding to the power scene based on each of the three-dimensional space coordinates is specifically: Extracting feature points from each of the depth feature maps, and matching the feature points to obtain feature point pairs; Determine the optimal three-dimensional space coordinates corresponding to each two-dimensional coordinate based on the feature point pairs; The three-dimensional space coordinates and the optimal three-dimensional space coordinates are fused with multiple views to obtain a three-dimensional point cloud corresponding to the power scene.
6. The method for 3D reconstruction of multi-view 2D images according to any one of claims 1 to 5, characterized in that: After generating the three-dimensional point cloud corresponding to the power scene based on each of the three-dimensional space coordinates, the method further includes: Performing point cloud alignment on the three-dimensional point cloud based on a preset rigid change matrix to obtain an initial point cloud optimization result, wherein the rigid change matrix includes a first rotation matrix and a first translation matrix; The second rotation matrix and the second translation matrix of each point cloud in the initial point cloud optimization result are calculated based on the least squares method, and the point clouds in the initial point cloud optimization result are aligned based on the second rotation matrix and the second translation matrix to obtain the target three-dimensional point cloud.
7. A semantic recognition method for multi-view 2D images, characterized in that: include: A three-dimensional point cloud corresponding to an electric power scene is generated according to a three-dimensional reconstruction method of a multi-view 2D image according to any one of claims 1 to 5; The three-dimensional point cloud is semantically segmented to obtain various power facilities in the power scene.
8. The semantic recognition method of multi-view 2D images according to claim 7, characterized in that: The three-dimensional point cloud is semantically segmented to obtain various power facilities, specifically: voxelize the three-dimensional point cloud to obtain a plurality of voxel blocks; Performing feature extraction on each of the voxel blocks to obtain a number of corresponding local spatial features and global spatial features; Based on the local spatial features and the global spatial features, the label probability of each voxel block is determined, and based on the label probability, each power facility in the power scene is obtained.
9. A semantic recognition system for multi-view 2D images, characterized in that: include: 3D reconstruction module and segmentation module; The three-dimensional reconstruction module is used to generate a three-dimensional point cloud corresponding to the power scene according to the three-dimensional reconstruction method of a multi-view 2D image according to any one of claims 1 to 5; The segmentation module is used to perform semantic segmentation on the three-dimensional point cloud to obtain various power facilities in the power scene.
10. The semantic recognition system for multi-view 2D images according to claim 9, characterized in that: The segmentation module includes: a first processing unit, an extraction unit, and a second processing unit; The first processing unit is used to voxelize the three-dimensional point cloud to obtain a plurality of voxel blocks; The extraction unit is used to extract features from each voxel block to obtain a number of corresponding local space features and global space features; The second processing unit is used to determine the label probability of each of the voxel blocks based on the local spatial features and the global spatial features, and obtain each power facility in the power scene based on the label probability.
Citation Information
Cited By
Video compression method, system and equipment based on dynamic perception and medium
CN121691695A
A dynamic perception-based video compression method, system, device and medium
CN121691695B