Feature pyramid network in multi-view three-dimensional reconstruction and multi-view three-dimensional reconstruction method

By using multi-scale feature extraction and global context information fusion through feature pyramid network, the accuracy and efficiency problems of traditional multi-view stereo reconstruction in complex scenes are solved, achieving efficient and high-quality 3D reconstruction suitable for a variety of application scenarios.

CN120876776AInactive Publication Date: 2025-10-31NEW ELEMENTS (ZHANGJIAGANG) DIGITAL TECHNOLOGY CO LTD +3
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510908411.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional multi-view stereo reconstruction methods lack accuracy and completeness in complex scenes, have low computational efficiency, and feature matching and geometric constraints are easily affected by noise, making them difficult to apply in real time.

Method used

A feature pyramid network, including a lightweight convolutional neural network and a lightweight Transformer module, is used to generate high-quality 3D point clouds through multi-scale feature extraction, global context information fusion, feature pyramid construction, matching cost volume regularization, and depth map optimization.

Benefits of technology

It improves the accuracy and robustness of multi-view stereo reconstruction, reduces computational overhead, expands application scenarios, and is suitable for fields such as autonomous driving, augmented reality, and robot navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876776A_ABST
    Figure CN120876776A_ABST
Patent Text Reader

Abstract

The invention provides a feature pyramid network in multi-view three-dimensional reconstruction and a multi-view three-dimensional reconstruction method. Comprising an input image module, a camera parameter module, a shallow feature extraction module, a multi-scale feature extraction module, a global representation sub-network module, a position coding module, a feature fusion module, an FPN encoder module, an FPN decoder module, a matching cost body construction module, a cost body regularization module, a depth sampling range estimation module and a depth map generation module. A depth map optimization module, a point cloud generation module and a loss function module. The invention provides a feature pyramid network in multi-view three-dimensional reconstruction, which comprises the key technologies of multi-scale feature extraction, global context information extraction, feature fusion, feature pyramid construction, matching cost body construction, cost body regularization, depth sampling range estimation, depth map generation and optimization and the like. And high-efficiency and high-quality multi-view three-dimensional reconstruction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and deep learning technology, specifically to a feature pyramid network and a multi-view stereo reconstruction method. Background Technology

[0002] Multi-view Stereo (MVS) reconstruction is an important research direction in computer vision, aiming to reconstruct a 3D model of a scene from images captured from multiple perspectives. Traditional MVS methods mainly rely on manually designed feature matching and geometric constraints, achieving 3D reconstruction through steps such as feature point detection, matching, and triangulation. Although these methods have achieved good results in some simple scenes, the accuracy and completeness of the reconstruction results remain significantly limited in complex scenes (such as those with indistinct textures, severe occlusion, and varying lighting). Traditional MVS methods have the following limitations:

[0003] Feature matching is difficult: Traditional methods rely on manually designed feature point detection and matching algorithms (such as SIFT, SURF, etc.). In areas with indistinct textures or in scenes with severe occlusion, feature point detection and matching are prone to mismatches or missed matches.

[0004] Insufficient geometric constraints: Traditional geometric constraint methods (such as parallax-based stereo matching) are easily affected by the accumulation of noise and errors when processing large scenes or multi-view data, leading to a decrease in the accuracy of reconstruction results.

[0005] Low computational efficiency: Traditional MVS methods typically require a large amount of manual parameter tuning and complex post-processing steps, resulting in low computational efficiency and making them difficult to apply in real time.

[0006] To address this, a feature pyramid network and a multi-view stereo reconstruction method are proposed. Summary of the Invention

[0007] The present invention aims to solve the problems mentioned in the background art by providing a feature pyramid network and a method for multi-view stereo reconstruction.

[0008] The specific technical solution is as follows:

[0009] A feature pyramid network for multi-view stereo reconstruction includes: an input image module for receiving multiple images from different viewpoints; a camera parameter module for storing and managing camera intrinsic and extrinsic parameters; a shallow feature extraction module for extracting shallow features; a multi-scale feature extraction module for extracting feature maps at different scales; a global representation sub-network module for extracting global context information; a position encoding module for generating positional codes; a feature fusion module for fusing shallow features and global context information to form multi-scale feature maps; an FPN encoder module for encoding multi-scale feature maps to generate high-level feature maps; an FPN decoder module for sampling high-level feature maps to generate low-level feature maps to form a feature pyramid; a matching cost body for constructing a matching cost body; a cost body regularization module for regularizing the matching cost body; a depth sampling range estimation module for estimating the depth sampling range; a depth map generation module for generating depth maps; a depth map optimization module for optimizing depth maps; a point cloud generation module for generating 3D point clouds; and a loss function module for defining a loss function for training the network. These modules are connected sequentially to realize feature extraction, matching, depth estimation, and point cloud generation in multi-view stereo reconstruction.

[0010] Preferably, the shallow feature extraction module uses a lightweight convolutional neural network for feature extraction.

[0011] Preferably, the global representation sub-network module uses a lightweight Transformer module to extract global context information, and the position encoding module generates position codes for the Transformer module.

[0012] Preferably, the depth map generation module processes the regularized cost volume using a 3D convolutional neural network to generate a depth map; the depth map optimization module uses bilateral filtering, nonlocal mean filtering, or super-resolution techniques to optimize the depth map.

[0013] Preferably, the loss function module defines a loss function based on the generated 3D point cloud, which is used to train the network and optimize model parameters.

[0014] Preferably, the convolutional neural network includes multiple convolutional layers, activation functions, and pooling layers, wherein:

[0015] Convolutional layers extract local features from the input data through convolution operations. The convolutional kernel slides across the input data, performing weighted summation on each local region to extract features. Activation functions are used to increase the non-linearity of the model, enabling the model to learn complex feature representations. Pooling layers reduce the size of the feature map through downsampling operations.

[0016] Convolution operation:

[0017] Assuming the input data is X, the convolution kernel is W, and the bias is b, then the output Y of the convolutional layer is expressed as:

[0018]

[0019] Where (i,j) is the position on the output feature map, and (m,n) is the position on the convolution kernel;

[0020] Activation function: The result Z of the output of the convolutional layer after passing through the activation function is represented as: Z(i,j)=max(0,Y(i,j));

[0021] Pooling operation: Assuming the pooling window size is k×k and the stride is s, the output P of the pooling layer is expressed as: P(i,j)=max(m,n)∈window Y(is+m,js+n);

[0022] Where (i,j) is the position on the output feature map, and (m,n) is the position within the pooling window.

[0023] Preferably, the Transformer module includes a self-attention mechanism and a feedforward neural network, wherein:

[0024] Self-attention allows the model to simultaneously focus on different positions within the input sequence while processing it, thereby capturing long-range dependencies within the input data. First, it calculates the attention score between each position and other positions in the input sequence. Then, it performs a weighted summation of these scores to obtain the contextual representation of each position. Assume the input sequence is X with dimension d. x The computational process of the self-attention mechanism is then expressed as:

[0025]

[0026] Among them, W Q W K W V These are the weight matrices learned during training, used to map the input sequence X to the query, key, and value spaces, respectively; d k It is the dimension of the key vector, used as a scaling factor to adjust the attention score;

[0027] Feedforward neural networks are used to further process the contextual representation at each location to extract higher-level features, including two linear transformations and a non-linear activation function.

[0028] Assuming the output of the self-attention mechanism is A, the feedforward neural network can be represented as:

[0029] FFN(A) = ReLU(AW1+b1)W2+b2;

[0030] Where W1 and W2 are the weight matrices learned during training, and b1 and b2 are bias terms.

[0031] Preferably, the position encoding generated by the position encoding module includes absolute position encoding and relative position encoding. The absolute position encoding and relative position encoding cooperate with each other. The absolute position encoding is used to assign a fixed and unique coordinate information to each element in the feature map, and the relative position encoding is used to capture the relative distance or relative relationship between elements in the feature map.

[0032] Preferably, the feature fusion module fuses shallow features and global context information through weighted fusion to form a multi-scale feature map.

[0033] Preferably, the FPN encoder module encodes multi-scale feature maps through multiple convolutional layers and normalization layers to generate high-level feature maps.

[0034] Preferably, the FPN decoder module upsamples the high-level feature map through an upsampling layer and skip connections to generate a low-level feature map, forming a feature pyramid.

[0035] Preferably, the matching cost body construction module constructs the matching cost body by using the feature map in the feature pyramid through disparity search and cost aggregation.

[0036] Preferably, the cost body regularization module reduces noise and redundant information by smoothing and normalizing the matching cost body.

[0037] Preferably, the depth sampling range estimation module estimates the depth sampling range based on the matching cost volume through statistical analysis and threshold filtering.

[0038] Preferably, the 3D convolutional neural network includes multiple 3D convolutional layers and activation functions. The 3D convolutional layers are used to perform convolution operations on the 3D matching cost volume to extract features with spatial structure. The convolution operation formula is as follows:

[0039] Assuming the input 3D matching cost volume is C, the convolution kernel is K, and the bias is b, then the output F of the 3D convolutional layer is expressed as:

[0040]

[0041] in:

[0042] (i,j,k) is the position on the output feature map F;

[0043] (m,n,p) is the position on the convolution kernel K;

[0044] C(i+m,j+n,k+p) represents the value of the input matching cost body at position (i+m,j+n,k+p);

[0045] K(m,n,p) represents the weight of the convolution kernel at position (m,n,p);

[0046] b is the bias term;

[0047] The activation function is:

[0048] ReLU(x) = max(0,x);

[0049] After the feature map F passes through the 3D convolutional layer and is activated, the output is:

[0050] F′(i,j,k)=max(0,F(i,j,k)).

[0051] Preferably, the depth map optimization module optimizes the generated depth map through smoothing filtering and edge enhancement to improve the accuracy and completeness of the depth map.

[0052] Preferably, the point cloud generation module generates a 3D point cloud using a back projection method based on a depth map and camera parameters.

[0053] Preferably, the loss function defined by the loss function module includes depth loss, photometric loss, and regularization loss, which are used to train the network and optimize model parameters.

[0054] This invention also proposes a multi-view stereo reconstruction method, comprising the following steps:

[0055] S1. Image Acquisition: Acquire multiple images of the target scene from different perspectives, including at least one reference image and multiple source images;

[0056] S2. Camera Parameter Acquisition: Acquire the camera intrinsic and extrinsic parameters corresponding to each viewpoint image, including focal length, principal point coordinates, rotation matrix, and translation vector, to facilitate subsequent geometric transformations and projection operations.

[0057] S3. Image preprocessing: The acquired images are preprocessed by grayscale conversion and normalization to adapt to the subsequent feature extraction process;

[0058] S4. Shallow Feature Extraction: Use a lightweight convolutional neural network (CNN) to extract shallow features from the preprocessed image;

[0059] S5. Multi-scale feature extraction: Utilize multiple convolutional layers and pooling layers to further extract feature maps of different scales from shallow features;

[0060] S6. Global Context Information Extraction: A lightweight Transformer module is used to extract global context information from multi-scale feature maps;

[0061] S7. Position Encoding Generation: Generates position encodings for the Transformer module to adapt to feature maps of different scales while preserving spatial position information;

[0062] S8. Feature Fusion: Fusion of shallow features and global context information to form a multi-scale feature map containing rich information;

[0063] S9. Feature Pyramid Construction: The feature pyramid is constructed by encoding and upsampling the multi-scale feature maps through the FPN encoder and decoder modules.

[0064] S10. Matching cost body construction: Using the feature maps in the feature pyramid, a matching cost body is constructed through disparity search and cost aggregation.

[0065] S11. Cost body regularization: The constructed matching cost body is smoothed and normalized to reduce the impact of noise and redundant information.

[0066] S12. Depth Sampling Range Estimation: Based on the regularized matching cost volume, the effective range of depth sampling is estimated through statistical analysis and threshold screening.

[0067] S13. Depth Map Generation: A 3D convolutional neural network (3D CNN) is used to process the regularized matching cost volume to generate a preliminary depth map.

[0068] S14. Depth Map Optimization: Perform smoothing filtering and edge enhancement post-processing on the generated depth map to improve its quality;

[0069] S15, 3D point cloud generation: Combine the optimized depth map and camera parameters to generate a 3D point cloud using the inverse projection method;

[0070] S16. Model Training: Define a comprehensive loss function that includes depth loss, photometric loss, and regularization loss to supervise network training and optimize model parameters.

[0071] S17. Application Output: The final generated 3D point cloud can be used for specific application scenarios, such as 3D modeling, virtual reality, augmented reality, or autonomous driving.

[0072] The present invention has the following beneficial effects:

[0073] This invention provides a Feature Pyramid Network (FPN) for multi-view stereo reconstruction. This network achieves efficient and high-quality multi-view stereo reconstruction through key technologies such as multi-scale feature extraction, global context information extraction, feature fusion, feature pyramid construction, matching cost volume construction, cost volume regularization, depth sampling range estimation, depth map generation, and optimization. Specifically, this invention has the following technical effects:

[0074] 1. Improve reconstruction accuracy: By fusing multi-scale feature extraction and global contextual information, richer feature representations are generated, improving the accuracy of depth maps and 3D point clouds.

[0075] 2. Enhance robustness: By reducing the impact of noise and redundant information through cost volume regularization and depth graph optimization, the robustness of the model in complex scenarios is improved.

[0076] 3. Improve computational efficiency: Lightweight CNN and Transformer modules are used, combined with efficient feature pyramid construction and depth map generation methods, to reduce computational overhead.

[0077] 4. Expanding application scenarios: The model demonstrates good generalization ability across multiple resolutions and different scenarios, making it suitable for various application scenarios such as autonomous driving, augmented reality, and robot navigation.

[0078] In summary, this invention significantly improves the performance of multi-view stereo reconstruction through a series of innovative module designs and technical means, providing new ideas and methods for research and application in this field. Attached Figure Description

[0079] Figure 1 This is an architecture diagram of the feature pyramid network in multi-view stereo reconstruction provided in an embodiment of the present invention;

[0080] Figure 2 The flowchart illustrates the application method of the feature pyramid network in multi-view stereo reconstruction provided in this embodiment of the invention. Detailed Implementation

[0081] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0082] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual images. They should not be construed as limiting the scope of this patent. To better illustrate the embodiments of the present invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions of the product. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0083] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "inner," and "outer" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present patent. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0084] In the description of this invention, unless otherwise explicitly specified and limited, the term "connection" or similar designation indicating a connection between components should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral part; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0085] Example

[0086] The feature pyramid network provided in this embodiment for multi-view stereo reconstruction, such as Figure 1 As shown, it includes: an input image module, a camera parameter module, a shallow feature extraction module, a multi-scale feature extraction module, a global representation sub-network module, a location encoding module, a feature fusion module, an FPN encoder module, an FPN decoder module, a matching cost body construction module, a cost body regularization module, a depth sampling range estimation module, a depth map generation module, a depth map optimization module, a point cloud generation module, and a loss function module, wherein:

[0087] The input image module is used to receive multiple images from different perspectives as input, including reference images and source images;

[0088] The camera parameter module is used to store and manage the camera's intrinsic and extrinsic parameters for each image, which are then used for subsequent geometric transformations and projections.

[0089] The shallow feature extraction module is used to extract shallow features from the input image using a lightweight convolutional neural network (CNN).

[0090] The multi-scale feature extraction module is used to extract feature maps of different scales through multiple convolutional layers and pooling layers;

[0091] The global representation subnetwork module is used to extract global context information using the lightweight Transformer module;

[0092] The position encoding module is used to generate position codes for the Transformer module to adapt to images of different scales;

[0093] The feature fusion module is used to fuse shallow features and global context information to form a multi-scale feature map;

[0094] The FPN encoder module is used to encode multi-scale feature maps to generate high-level feature maps;

[0095] The FPN decoder module is used to upsample the high-level feature map to generate the low-level feature map, forming a feature pyramid.

[0096] The matching cost body construction module is used to construct the matching cost body using the feature maps in the feature pyramid;

[0097] The cost body regularization module is used to regularize the matching cost body to reduce noise and redundant information.

[0098] The depth sampling range estimation module is used to estimate the depth sampling range based on the matching cost volume, which is used for subsequent depth map generation.

[0099] The depth map generation module is used to process the regularized cost volume through a 3D convolutional neural network (3D CNN) to generate a depth map.

[0100] The depth map optimization module is used to optimize the generated depth map, improving its accuracy and completeness.

[0101] The point cloud generation module is used to generate 3D point clouds based on the optimized depth map and camera parameters;

[0102] The loss function module is used to define loss functions, train networks, and optimize model parameters;

[0103] The system comprises several modules: an input image module connected to a camera parameter module (providing multiple images from different perspectives, with the camera parameter module managing the intrinsic and extrinsic camera parameters); a shallow feature extraction module that passes images to it for initial feature extraction; a multi-scale feature extraction module that passes the extracted feature maps to it for multi-scale feature extraction; and a global representation sub-network module that passes the multi-scale feature maps to it for global context information extraction. The global representation sub-network module is connected to the position encoding module. The global context information extracted by the global representation sub-network module is passed to the position encoding module to generate position codes. The position encoding module is connected to the feature fusion module. The position codes generated by the position encoding module are passed to the feature fusion module for fusion with multi-scale feature maps. The multi-scale feature extraction module is connected to the feature fusion module. The multi-scale feature maps extracted by the multi-scale feature extraction module are passed to the feature fusion module for fusion with position codes. The feature fusion module is connected to the FPN encoder module. The multi-scale feature maps generated by the feature fusion module are passed to the FPN encoder module for encoding to generate high-level feature maps. The FPN encoder module is connected to the F... The FPN encoder module is connected to the PN decoder module. The high-level feature map generated by the FPN encoder module is passed to the FPN decoder module for upsampling to generate low-level feature maps, forming a feature pyramid. The FPN decoder module is connected to the matching cost body construction module. The feature pyramid generated by the FPN decoder module is passed to the matching cost body construction module to construct the matching cost body. The matching cost body construction module is connected to the cost body regularization module. The matching cost body generated by the matching cost body construction module is passed to the cost body regularization module for regularization processing. The cost body regularization module is connected to the depth sampling range estimation module. The regularized cost body generated by the cost body regularization module is passed to the depth sampling range estimation module for estimation. The system calculates the depth sampling range. The depth sampling range estimation module is connected to the depth map generation module; the depth sampling range generated by the depth sampling range estimation module is passed to the depth map generation module to generate a depth map. The depth map generation module is connected to the depth map optimization module; the depth map generated by the depth map generation module is passed to the depth map optimization module for optimization. The depth map optimization module is connected to the point cloud generation module; the optimized depth map generated by the depth map optimization module is passed to the point cloud generation module, and combined with camera parameters, a 3D point cloud is generated. The point cloud generation module is connected to the loss function module; the 3D point cloud generated by the point cloud generation module is passed to the loss function module to define the loss function, which is used to train the network and optimize model parameters.

[0104] By integrating multiple key components such as the input image module, camera parameter module, and feature extraction module, this embodiment provides an efficient multi-view stereo reconstruction solution. This solution can effectively process image data from different viewpoints, achieve high-quality 3D reconstruction, and is suitable for various application scenarios.

[0105] Specifically, in this embodiment: the convolutional neural network (CNN) includes multiple convolutional layers, activation functions, and pooling layers. The design of multiple convolutional layers and activation functions enhances the model's learning and expressive capabilities while maintaining its lightweight nature, improving computational efficiency and generalization performance. The pooling layers include max pooling and average pooling. By combining max pooling and average pooling, the main features of the image can be preserved while effectively reducing the feature dimension, accelerating the computation process, and improving the model's operating efficiency.

[0106] Convolutional layers extract local features from the input data through convolution operations. The convolutional kernel (or filter) slides across the input data, performing a weighted summation on each local region to extract features; activation functions increase the model's non-linearity, enabling it to learn complex feature representations; pooling layers reduce the size of the feature map through downsampling operations while preserving important features.

[0107] Convolution operation:

[0108] Assuming the input data is X, the convolution kernel is W, and the bias is b, the output Y of the convolutional layer can be expressed as:

[0109]

[0110] Where (i,j) is the position on the output feature map, and (m,n) is the position on the convolution kernel.

[0111] Activation function: Taking ReLU as an example, the output Z of the convolutional layer after passing through the activation function can be expressed as: Z(i,j)=max(0,Y(i,j)).

[0112] Pooling operation (taking max pooling as an example): Assuming the pooling window size is k×k and the stride is s, the output P of the pooling layer can be expressed as: P(i,j)=max(m,n)∈window Y(is+m,js+n)

[0113] Where (i,j) is the position on the output feature map, and (m,n) is the position within the pooling window.

[0114] Specifically, in this embodiment: the Transformer module includes a self-attention mechanism and a feedforward neural network. The introduction of the self-attention mechanism and the feedforward neural network enhances the model's ability to capture long-distance dependencies, helps to extract richer global contextual information, and improves the model's understanding depth and accuracy.

[0115] The self-attention mechanism allows the model to simultaneously focus on different positions within the input sequence, thereby capturing long-range dependencies within the input data. First, it calculates the similarity (or attention score) between each position and other positions in the input sequence. Then, it performs a weighted summation of these scores to obtain the contextual representation of each position. Assuming the input sequence is X with dimension d... x The computational process of the self-attention mechanism can then be expressed as:

[0116] Q,K,V=XW Q XW K XW V

[0117]

[0118] Among them, W Q W K W V These are the weight matrices learned during training, used to map the input sequence X to the query, key, and value spaces, respectively; d k It is the dimension of the key vector, used as a scaling factor to adjust the attention score.

[0119] In this context, the feedforward neural network is used to further process the context representation at each location to extract higher-level features, and typically includes two linear transformations and a non-linear activation function (such as ReLU).

[0120] Assuming the output of the self-attention mechanism is A, the feedforward neural network can be represented as:

[0121] FFN(A)=ReLU(AW1+b1)W2+b2

[0122] Where W1 and W2 are the weight matrices learned during training, and b1 and b2 are bias terms.

[0123] Specifically, in this embodiment: the position encoding module generates position encoding including absolute position encoding and relative position encoding. Absolute and relative position encoding work together. Absolute position encoding assigns a fixed, unique coordinate to each element (such as a pixel or feature vector) in the feature map. This helps the model understand the specific location of each element in the image. Especially when processing multi-scale feature maps, absolute position encoding ensures that the model does not confuse the same location at different scales. Relative position encoding captures the relative distance or relationship between elements in the feature map, not just their absolute positions. This helps the model understand the interactions between elements. Especially when the model needs to capture long-distance dependencies, relative position encoding can provide more contextual information. The combination of absolute and relative position encoding not only provides precise position information for each element in the feature map but also considers the relative positional relationships between elements, enhancing the model's spatial awareness.

[0124] Specifically, in this embodiment: the feature fusion module fuses shallow features and global context information through weighted fusion to form a multi-scale feature map. By adopting weighted fusion, the shallow features and global context information are effectively integrated, so that the generated multi-scale feature map contains both local details and a global perspective, laying a solid foundation for the subsequent generation of depth maps.

[0125] Specifically, in this embodiment: the FPN encoder module encodes multi-scale feature maps through multiple convolutional layers and normalization layers to generate high-level feature maps, thereby achieving effective encoding of multi-scale feature maps. The generated high-level feature maps have stronger abstraction and representativeness, which is beneficial to improving the overall performance of the model.

[0126] Specifically, in this embodiment: the FPN decoder module upsamples the high-level feature map through an upsampling layer and skip connections to generate a low-level feature map, forming a feature pyramid. By combining the upsampling layer and skip connections, the model can maintain high-level information while recovering details, constructing a feature pyramid with distinct layers and rich information, which further improves the reconstruction accuracy of the model.

[0127] Specifically, in this embodiment: the matching cost body construction module constructs the matching cost body using the feature maps in the feature pyramid through disparity search and cost aggregation. By using the disparity search and cost aggregation methods, the matching cost body is effectively constructed, providing an important data foundation for subsequent depth map generation and ensuring the accuracy and reliability of the depth map.

[0128] Specifically, in this embodiment: the cost body regularization module reduces noise and redundant information by smoothing and normalizing the matching cost body. By smoothing and normalizing the matching cost body, the impact of noise and redundant information is reduced, thereby improving the robustness and generalization ability of the model.

[0129] Specifically, in this embodiment: the depth sampling range estimation module estimates the depth sampling range based on the matching cost volume through statistical analysis and threshold filtering. By using statistical analysis and threshold filtering methods, the effective range of depth sampling is accurately estimated, providing accurate guidance for the generation of depth maps and avoiding unnecessary computational overhead.

[0130] Specifically, in this embodiment: the 3D Convolutional Neural Network (3DCNN) includes multiple 3D convolutional layers and activation functions. The 3D convolutional layers are used to perform convolution operations on the 3D matching cost volume to extract features with spatial structure. The convolution operation formula is as follows:

[0131] Assuming the input 3D matching cost volume is C, the convolution kernel is K, and the bias is b, then the output F of the 3D convolutional layer is expressed as:

[0132]

[0133] in:

[0134] (i,j,k) is the position on the output feature map F;

[0135] (m,n,p) is the position on the convolution kernel K;

[0136] C(i+m,j+n,k+p) represents the value of the input matching cost body at position (i+m,j+n,k+p);

[0137] K(m,n,p) represents the weight of the convolution kernel at position (m,n,p);

[0138] b is the bias term;

[0139] The activation function is:

[0140] ReLU(x) = max(0,x);

[0141] After the feature map F passes through the 3D convolutional layer and is activated, the output is:

[0142] F′(i,j,k)=max(0,F(i,j,k));

[0143] The purpose of the activation function is to set negative values ​​to 0 and retain positive values, thereby introducing non-linearity and enabling the model to learn more complex feature representations.

[0144] By utilizing multiple 3D convolutional layers and activation functions, we can deeply mine the spatial structure information in the matching cost volume, generate high-quality depth maps, and significantly improve the accuracy and level of detail of 3D reconstruction.

[0145] Specifically, in this embodiment: the depth map optimization module optimizes the generated depth map through smoothing filtering and edge enhancement to improve the accuracy and completeness of the depth map. Through post-processing methods such as smoothing filtering and edge enhancement, the quality of the depth map is further improved, making it more in line with the needs of the actual scene and enhancing the realism of the 3D reconstruction.

[0146] Specifically, in this embodiment: the point cloud generation module generates a 3D point cloud using the depth map and camera parameters through the inverse projection method. Combined with the optimized depth map and camera parameters, the 3D point cloud is generated through the inverse projection method, providing accurate data support for subsequent applications and expanding the applicability of the system.

[0147] Specifically, in this embodiment: the loss function module defines a loss function including depth loss, photometric loss, and regularization loss, which are used to train the network and optimize model parameters. A comprehensive loss function including depth loss, photometric loss, and regularization loss is defined, which provides effective supervision signals for model training, ensures reasonable optimization of model parameters, and promotes continuous improvement of model performance.

[0148] This embodiment also provides a multi-view stereo reconstruction method, such as Figure 2 As shown, it includes the following steps:

[0149] S1. Image Acquisition: Acquire multiple images of the target scene from different perspectives, including at least one reference image and multiple source images;

[0150] S2. Camera Parameter Acquisition: Acquire the camera intrinsic and extrinsic parameters corresponding to each viewpoint image, including focal length, principal point coordinates, rotation matrix, and translation vector, to facilitate subsequent geometric transformations and projection operations.

[0151] S3. Image preprocessing: The acquired images are preprocessed by grayscale conversion and normalization to adapt to the subsequent feature extraction process;

[0152] S4. Shallow Feature Extraction: Use a lightweight convolutional neural network (CNN) to extract shallow features from the preprocessed image;

[0153] S5. Multi-scale feature extraction: Utilize multiple convolutional layers and pooling layers to further extract feature maps of different scales from shallow features;

[0154] S6. Global Context Information Extraction: A lightweight Transformer module is used to extract global context information from multi-scale feature maps;

[0155] S7. Position Encoding Generation: Generates position encodings for the Transformer module to adapt to feature maps of different scales while preserving spatial position information;

[0156] S8. Feature Fusion: Fusion of shallow features and global context information to form a multi-scale feature map containing rich information;

[0157] S9. Feature Pyramid Construction: The feature pyramid is constructed by encoding and upsampling the multi-scale feature maps through the FPN encoder and decoder modules.

[0158] S10. Matching cost body construction: Using the feature maps in the feature pyramid, a matching cost body is constructed through disparity search and cost aggregation.

[0159] S11. Cost body regularization: The constructed matching cost body is smoothed and normalized to reduce the impact of noise and redundant information.

[0160] S12. Depth Sampling Range Estimation: Based on the regularized matching cost volume, the effective range of depth sampling is estimated through statistical analysis and threshold screening.

[0161] S13. Depth Map Generation: A 3D convolutional neural network (3D CNN) is used to process the regularized matching cost volume to generate a preliminary depth map.

[0162] S14. Depth Map Optimization: Perform smoothing filtering and edge enhancement post-processing on the generated depth map to improve its quality;

[0163] S15, 3D point cloud generation: Combine the optimized depth map and camera parameters to generate a 3D point cloud using the inverse projection method;

[0164] S16. Model Training: Define a comprehensive loss function that includes depth loss, photometric loss, and regularization loss to supervise network training and optimize model parameters.

[0165] S17. Application Output: The final generated 3D point cloud can be used for specific application scenarios, such as 3D modeling, virtual reality, augmented reality, or autonomous driving.

[0166] From image acquisition to final application output, this method provides a complete multi-view stereo reconstruction workflow. Each step has been carefully designed to ensure the efficiency and reliability of the entire process, providing strong support for the practical application of multi-view stereo reconstruction technology.

[0167] In summary, the feature pyramid network and multi-view stereo reconstruction method in this embodiment have the following technical effects:

[0168] 1. Multi-view image processing:

[0169] Input image module: Receives multiple images from different perspectives, including reference images and source images, providing a data foundation for subsequent processing.

[0170] Camera Parameters Module: Stores and manages the camera's intrinsic and extrinsic parameters for each image, ensuring the accuracy of geometric transformations and projection operations.

[0171] 2. Feature extraction and fusion:

[0172] Shallow feature extraction module: Uses a lightweight convolutional neural network (CNN) to extract shallow features from the input image, laying the foundation for subsequent multi-scale feature extraction.

[0173] Multi-scale feature extraction module: Extracts feature maps of different scales through multiple convolutional and pooling layers, enhancing the model's ability to capture features at different scales.

[0174] Global Representation Subnetwork Module: This module utilizes a lightweight Transformer module to extract global contextual information, overcoming the limitation of traditional CNNs which can only capture local information.

[0175] Location encoding module: Generates location codes to help the Transformer module better understand the spatial location information in the feature map.

[0176] Feature fusion module: Combines shallow features and global context information through weighted fusion to form multi-scale feature maps, enhancing the richness and diversity of features.

[0177] 3. Feature Pyramid Construction:

[0178] The FPN encoder module encodes multi-scale feature maps to generate high-level feature maps, improving the abstractness and representativeness of the features.

[0179] The FPN decoder module upsamples high-level feature maps through upsampling layers and skip connections to generate low-level feature maps, constructing a hierarchical feature pyramid that provides rich multi-scale information for subsequent depth map generation.

[0180] 4. Depth map generation and optimization:

[0181] Matching cost body construction module: Utilizing feature maps in the feature pyramid, a matching cost body is constructed through disparity search and cost aggregation methods, providing an important data foundation for depth map generation.

[0182] Cost volume regularization module: Smooths and normalizes the matching cost volume, reducing the impact of noise and redundant information, and improving the robustness and generalization ability of the model.

[0183] Depth sampling range estimation module: It estimates the effective range of depth sampling through statistical analysis and threshold filtering, which improves the efficiency and accuracy of depth map generation.

[0184] Depth map generation module: Uses a 3D convolutional neural network to process the regularized cost volume and generate a preliminary depth map.

[0185] Depth map optimization module: Through post-processing techniques such as smoothing filtering and edge enhancement, the quality of the depth map is further improved, making it more in line with the needs of real-world scenarios.

[0186] 5. 3D point cloud generation and model training:

[0187] Point cloud generation module: Combining the optimized depth map and camera parameters, it generates 3D point clouds through inverse projection, providing accurate data support for subsequent applications.

[0188] The loss function module defines a comprehensive loss function that includes depth loss, photometric loss, and regularization loss. This provides effective supervision signals for model training, ensures reasonable optimization of model parameters, and promotes continuous improvement in model performance.

[0189] The overall technical effects are as follows:

[0190] High-precision reconstruction: Through multi-scale feature extraction and fusion of global contextual information, the model can capture more details and global information, and the generated depth map and 3D point cloud have higher accuracy and completeness.

[0191] High robustness: Cost volume regularization and depth graph optimization modules effectively reduce the impact of noise and redundant information, improving the robustness of the model in complex scenarios.

[0192] Efficiency: The lightweight CNN and Transformer module design, combined with efficient feature pyramid construction and depth map generation methods, enables the model to maintain high performance while having low computational overhead.

[0193] Strong generalization ability: The model exhibits good generalization ability across multiple resolutions and different scenarios, making it suitable for various applications such as autonomous driving, augmented reality, and robot navigation.

[0194] In summary, this invention significantly improves the performance of multi-view stereo reconstruction through a series of innovative module designs and technical means, providing new ideas and methods for research and application in this field.

[0195] The above are merely preferred embodiments of the present invention and are not intended to limit the implementation methods and protection scope of the present invention. Those skilled in the art should recognize that any equivalent substitutions and obvious changes made based on the description and illustrations of the present invention should be included within the protection scope of the present invention.

Claims

1. A feature pyramid network for multi-view stereo reconstruction, characterized in that, include: An input image module for receiving multiple images from different perspectives; A camera parameter module for storing and managing camera intrinsic and extrinsic parameters; a shallow feature extraction module for extracting shallow features; A multi-scale feature extraction module for extracting feature maps at different scales; a global representation sub-network module for extracting global context information; a position encoding module for generating positional encodings; a feature fusion module for fusing shallow features and global context information to form multi-scale feature maps; an FPN encoder module for encoding multi-scale feature maps to generate high-level feature maps; an FPN decoder module for sampling high-level feature maps to generate low-level feature maps to form a feature pyramid; a matching cost body construction module for constructing the matching cost body; a cost body regularization module for regularizing the matching cost body; a depth sampling range estimation module for estimating the depth sampling range; and a depth map generation module for generating depth maps. A depth map optimization module for optimizing depth maps; A point cloud generation module is used to generate 3D point clouds; a loss function module is used to define the loss function for training the network. These modules are connected in sequence to realize feature extraction, matching, depth estimation and point cloud generation in multi-view stereo reconstruction.

2. The feature pyramid network for multi-view stereo reconstruction according to claim 1, characterized in that, The shallow feature extraction module uses a lightweight convolutional neural network for feature extraction.

3. The feature pyramid network for multi-view stereo reconstruction according to claim 1, characterized in that, The global representation subnetwork module uses the lightweight Transformer module to extract global context information, and the position encoding module generates position codes for the Transformer module.

4. The feature pyramid network for multi-view stereo reconstruction according to claim 1, characterized in that, The depth map generation module processes the regularized cost volume using a 3D convolutional neural network to generate a depth map; the depth map optimization module uses bilateral filtering, nonlocal mean filtering, or super-resolution techniques to optimize the depth map.

5. The feature pyramid network for multi-view stereo reconstruction according to claim 1, characterized in that, The loss function module defines a loss function based on the generated 3D point cloud, which is used to train the network and optimize model parameters.

6. The feature pyramid network for multi-view stereo reconstruction according to claim 2, characterized in that, The convolutional neural network includes multiple convolutional layers, activation functions, and pooling layers, wherein: Convolutional layers extract local features from the input data through convolution operations. The convolutional kernel slides across the input data, performing weighted summation on each local region to extract features. Activation functions are used to increase the non-linearity of the model, enabling the model to learn complex feature representations. Pooling layers reduce the size of the feature map through downsampling operations. Convolution operation: Assuming the input data is X, the convolution kernel is W, and the bias is b, then the output Y of the convolutional layer is expressed as: Where (i,j) is the position on the output feature map, and (m,n) is the position on the convolution kernel; Activation function: The result Z of the output of the convolutional layer after passing through the activation function is represented as: Z(i,j)=max(0,Y(i,j)); Pooling operation: Assuming the pooling window size is k×k and the stride is s, the output P of the pooling layer is expressed as: P(i,j)=max(m,n)∈window Y(is+m,js+n); Where (i,j) is the position on the output feature map, and (m,n) is the position within the pooling window.

7. The feature pyramid network for multi-view stereo reconstruction according to claim 3, characterized in that, The Transformer module includes a self-attention mechanism and a feedforward neural network, wherein: Self-attention allows the model to simultaneously focus on different positions within the input sequence while processing it, thereby capturing long-range dependencies within the input data. First, it calculates the attention score between each position and other positions in the input sequence. Then, it performs a weighted summation of these scores to obtain the contextual representation of each position. Assume the input sequence is X with dimension d. x The computational process of the self-attention mechanism is then expressed as: Among them, W Q W K W V These are the weight matrices learned during training, used to map the input sequence X to the query, key, and value spaces, respectively; d k It is the dimension of the key vector, used as a scaling factor to adjust the attention score; Feedforward neural networks are used to further process the contextual representation at each location to extract higher-level features, including two linear transformations and a non-linear activation function. Assuming the output of the self-attention mechanism is A, the feedforward neural network can be represented as: FFN(A) = ReLU(AW1+b1)W2+b2; Where W1 and W2 are the weight matrices learned during training, and b1 and b2 are bias terms.

8. The feature pyramid network for multi-view stereo reconstruction according to claim 1, characterized in that, The position encoding module generates position encoding including absolute position encoding and relative position encoding. The absolute position encoding and relative position encoding work together. The absolute position encoding is used to assign a fixed and unique coordinate information to each element in the feature map, while the relative position encoding is used to capture the relative distance or relative relationship between elements in the feature map.

9. The feature pyramid network for multi-view stereo reconstruction according to claim 1, characterized in that, The feature fusion module fuses shallow features and global context information through weighted fusion to form a multi-scale feature map.

10. The feature pyramid network for multi-view stereo reconstruction according to claim 1, characterized in that, The FPN encoder module encodes multi-scale feature maps through multiple convolutional layers and normalization layers to generate high-level feature maps.

11. The feature pyramid network for multi-view stereo reconstruction according to claim 7, characterized in that, The FPN decoder module upsamples the high-level feature map through an upsampling layer and skip connections to generate a low-level feature map, forming a feature pyramid.

12. The feature pyramid network for multi-view stereo reconstruction according to claim 1, characterized in that, The matching cost body construction module constructs the matching cost body by using the feature map in the feature pyramid through disparity search and cost aggregation.

13. The feature pyramid network for multi-view stereo reconstruction according to claim 1, characterized in that, The cost body regularization module reduces noise and redundant information by smoothing and normalizing the matching cost body.

14. The feature pyramid network for multi-view stereo reconstruction according to claim 1, characterized in that, The depth sampling range estimation module estimates the depth sampling range based on the matching cost volume through statistical analysis and threshold filtering.

15. The feature pyramid network for multi-view stereo reconstruction according to claim 4, characterized in that, The 3D convolutional neural network includes multiple 3D convolutional layers and activation functions. The 3D convolutional layers are used to perform convolution operations on the 3D matching cost volume to extract features with spatial structure. The convolution operation formula is as follows: Assuming the input 3D matching cost volume is C, the convolution kernel is K, and the bias is b, then the output F of the 3D convolutional layer is expressed as: in: (i,j,k) is the position on the output feature map F; (m,n,p) is the position on the convolution kernel K; C(i+m,j+n,k+p) represents the value of the input matching cost body at position (i+m,j+n,k+p); K(m,n,p) represents the weight of the convolution kernel at position (m,n,p); b is the bias term; The activation function is: ReLU(x) = max(0,x); After the feature map F passes through the 3D convolutional layer and is activated, the output is: F′(i,j,k)=max(0,F(i,j,k)).

16. The feature pyramid network for multi-view stereo reconstruction according to claim 1, characterized in that, The depth map optimization module optimizes the generated depth map through smoothing filtering and edge enhancement to improve the accuracy and completeness of the depth map.

17. The feature pyramid network for multi-view stereo reconstruction according to claim 1, characterized in that, The point cloud generation module generates a 3D point cloud using a back projection method based on the depth map and camera parameters.

18. The feature pyramid network for multi-view stereo reconstruction according to claim 1, characterized in that, The loss function module defines loss functions including depth loss, photometric loss, and regularization loss, which are used to train the network and optimize model parameters.

19. A multi-view stereo reconstruction method, implemented using the feature pyramid network for multi-view stereo reconstruction as described in any one of claims 1 to 18, characterized in that, Includes the following steps: S1. Image Acquisition: Acquire multiple images of the target scene from different perspectives, including at least one reference image and multiple source images; S2. Camera Parameter Acquisition: Acquire the camera intrinsic and extrinsic parameters corresponding to each viewpoint image, including focal length, principal point coordinates, rotation matrix, and translation vector; S3. Image preprocessing: Perform grayscale conversion and normalization preprocessing on the acquired images; S4. Shallow Feature Extraction: Use a lightweight convolutional neural network to extract shallow features from the preprocessed image; S5. Multi-scale feature extraction: Utilize multiple convolutional layers and pooling layers to further extract feature maps of different scales from shallow features; S6. Global Context Information Extraction: A lightweight Transformer module is used to extract global context information from multi-scale feature maps; S7. Position Code Generation: Generates position codes for the Transformer module; S8. Feature Fusion: Fusion of shallow features and global context information to form a multi-scale feature map; S9. Feature Pyramid Construction: The feature pyramid is constructed by encoding and upsampling the multi-scale feature maps through the FPN encoder and decoder modules. S10. Matching cost body construction: Using the feature maps in the feature pyramid, a matching cost body is constructed through disparity search and cost aggregation. S11. Cost body regularization: Smoothing and normalizing the constructed matching cost body; S12. Depth Sampling Range Estimation: Based on the regularized matching cost volume, the effective range of depth sampling is estimated through statistical analysis and threshold screening. S13. Depth Map Generation: A 3D convolutional neural network is used to process the regularized matching cost volume to generate a preliminary depth map. S14. Depth Map Optimization: Perform smoothing filtering and edge enhancement post-processing on the generated depth map; S15, 3D point cloud generation: Combine the optimized depth map and camera parameters to generate a 3D point cloud using the inverse projection method; S16. Application Output: The final generated 3D point cloud is used for the required application scenarios.

Citation Information

Cited By

  • Electrical equipment infrared spectrum identification method and system based on feature fusion

    CN121544954A