3D reconstruction methods, apparatus, equipment, readable storage media, and program products
By determining the spatial range and number of depth planes in the 3D reconstruction method, calculating the pixel movement range and unit movement amount, and constructing a non-uniform virtual depth plane, the problem of low depth estimation is solved, and the accuracy and completeness of 3D reconstruction are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHONGQING LANDIAN AUTOMOBILE TECHNOLOGY CO LTD
- Filing Date
- 2025-11-26
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, the overall accuracy of depth estimation in multi-view stereo vision 3D reconstruction methods is not high, which affects the integrity of 3D reconstruction.
By acquiring multi-view images and camera parameters, the spatial range and number of depth planes are determined, the pixel movement range and unit movement amount are calculated, and the pixel movement positions are divided based on this information to construct a non-uniform virtual depth plane for 3D reconstruction.
It significantly improves the reconstruction integrity of distant areas and enhances the overall accuracy and completeness of 3D reconstruction.
Smart Images

Figure CN121213801B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of three-dimensional reconstruction technology, and in particular to a three-dimensional reconstruction method, apparatus, device, readable storage medium, and program product. Background Technology
[0002] With the development of image processing technology, multi-view stereo vision 3D reconstruction technology has emerged. This algorithm constructs a series of depth hypothesis planes, uses multi-view images of the same object and corresponding camera parameter information to generate corresponding depth maps, and then back-projects each pixel into 3D space to generate point clouds, forming the final 3D cost volume.
[0003] In related technologies, the constructed depth assumption plane is uniformly distributed along the principal optical axis of the reference image, which results in low overall accuracy of depth estimation and affects the integrity of 3D reconstruction. Summary of the Invention
[0004] Therefore, it is necessary to provide a three-dimensional reconstruction method, apparatus, device, readable storage medium, and program product that can improve the integrity of three-dimensional reconstruction in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a three-dimensional reconstruction method, including:
[0006] Acquire multi-view images of the target object and camera parameters when acquiring the multi-view images; the camera parameters include intrinsic parameters describing the internal optical and geometric properties of the camera and extrinsic parameters describing the position and orientation of the camera in the world coordinate system;
[0007] Determine the depth plane spatial range and the number of depth planes of the target object;
[0008] Based on the spatial range of the depth plane and the number of depth planes, determine the pixel movement range and unit movement amount of the pixels in the multi-view image;
[0009] Based on the pixel movement range and the unit movement amount, the pixel movement position of the pixel in the multi-view image is determined, the unit movement amount is spaced between adjacent pixel movement positions, and the number of pixel movement positions matches the number of depth planes;
[0010] Each virtual depth plane is determined based on the movement position of each pixel;
[0011] Based on the multi-view images, the camera parameters, and each of the virtual depth planes, a 3D reconstruction is performed on the target object to obtain the 3D reconstruction result of the target object.
[0012] In one embodiment, determining the pixel movement range and unit movement amount of pixels in the multi-view image based on the depth plane spatial range and the number of depth planes includes:
[0013] Based on the first reference point at the minimum depth in the depth plane spatial range and the camera parameters, the first pixel point in the reference image in the multi-view image is determined;
[0014] The second pixel in the reference image is determined based on the second reference point with the maximum depth in the depth plane spatial range and the camera parameters;
[0015] The pixel movement range is determined based on the first pixel and the second pixel.
[0016] The unit movement amount is determined based on the number of depth planes and the pixel movement range.
[0017] In one embodiment, each virtual depth plane is determined based on the movement position of each pixel, including:
[0018] Based on the pixel movement position and the camera parameters, determine the spatial reference point corresponding to each pixel movement position;
[0019] Each virtual depth plane is determined based on the spatial depth of each of the aforementioned spatial reference points.
[0020] In one embodiment, based on the multi-view images, the camera parameters, and each of the virtual depth planes, a 3D reconstruction is performed on the target object to obtain the 3D reconstruction result of the target object, including:
[0021] Feature extraction is performed on the multi-view images to obtain the image features of the multi-view images;
[0022] Based on the camera parameters and the image features, feature projection is performed on the virtual depth plane to obtain the feature volume corresponding to each virtual depth plane;
[0023] By integrating the feature volumes corresponding to the multiple virtual depth planes, a three-dimensional cost volume is obtained;
[0024] Based on the three-dimensional cost volume, a three-dimensional reconstruction is performed on the target object.
[0025] In one embodiment, three-dimensional reconstruction of the target object based on the three-dimensional cost volume includes:
[0026] The three-dimensional cost volume is subjected to three-dimensional depth convolution and three-dimensional point convolution in sequence to obtain the convolution intermediate features;
[0027] Feature aggregation is performed based on the intermediate features of the convolution to obtain a depth map; the depth map is used to describe the probability distribution of the target object in the depth dimension.
[0028] Based on the depth map, a 3D reconstruction is performed on the target object.
[0029] In one embodiment, based on the depth map, a 3D reconstruction of the target object is performed, including:
[0030] Based on the image features of the depth map and the reference image in the multi-view image, a composite feature map is obtained by feature stitching.
[0031] Three-dimensional reconstruction of the target object is performed based on the composite feature map.
[0032] Secondly, this application also provides a three-dimensional reconstruction apparatus, comprising:
[0033] The acquisition module is used to acquire multi-view images of the target object and camera parameters when acquiring the multi-view images; the camera parameters include intrinsic parameters describing the internal optical and geometric characteristics of the camera and extrinsic parameters describing the position and orientation of the camera in the world coordinate system.
[0034] A depth determination module is used to determine the spatial range of the depth plane and the number of depth planes of the target object;
[0035] A pixel determination module is used to determine the pixel movement range and unit movement amount of pixels in the multi-view image based on the spatial range of the depth plane and the number of depth planes.
[0036] The segmentation module is used to determine the pixel movement position of a pixel in the multi-view image based on the pixel movement range and the unit movement amount, wherein the unit movement amount is spaced between adjacent pixel movement positions, and the number of pixel movement positions matches the number of depth planes;
[0037] The plane determination module determines each virtual depth plane based on the movement position of each pixel.
[0038] The reconstruction module is used to perform three-dimensional reconstruction of the target object based on the multi-view images, the camera parameters, and each of the virtual depth planes, to obtain the three-dimensional reconstruction result of the target object.
[0039] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.
[0040] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0041] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0042] The aforementioned 3D reconstruction method, apparatus, device, readable storage medium, and program product determine the pixel movement range and unit movement amount of pixels in a multi-view image by using the spatial range and number of depth planes. Then, based on the pixel movement range and unit movement amount, the pixel movement range is equally divided to obtain multiple pixel movement positions. Each virtual depth plane is determined according to these multiple pixel movement positions. Since the intervals between the multiple pixel movement positions are the same, the search along the image epipolar line is nearly uniform when performing depth projection on virtual depth planes at different depths. This fundamentally ensures that the matching search accuracy is consistent regardless of the distance of the object. Based on multi-view images, camera parameters, and each virtual depth plane, 3D reconstruction is performed on the target object to obtain the 3D reconstruction result. The method provided in this application can significantly improve the reconstruction integrity of distant regions, thereby improving the integrity of 3D reconstruction. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 This is a diagram illustrating the application environment of a 3D reconstruction method in one embodiment.
[0045] Figure 2 This is a flowchart illustrating a three-dimensional reconstruction method in one embodiment;
[0046] Figure 3 This is a flowchart illustrating the process of determining the pixel movement range in one embodiment;
[0047] Figure 4 This is a schematic diagram of the spatial distribution of the virtual depth plane involved in one embodiment;
[0048] Figure 5 This is a schematic diagram of a three-dimensional reconstruction process involved in one embodiment;
[0049] Figure 6This is a structural block diagram of a three-dimensional reconstruction device in one embodiment;
[0050] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0052] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0053] The three-dimensional reconstruction method provided in this application embodiment can be applied to, for example... Figure 1 The application environment shown is illustrated. Terminal 101 communicates with server 102 via a network. A data storage system can store the data that server 102 needs to process. The data storage system can be integrated onto server 102, or it can be located in the cloud or on another network server.
[0054] Terminal 101 can acquire multi-view images of the target object to obtain multi-view images of the target object. Then, the acquired multi-view images and the camera parameters when acquiring the multi-view images are sent to server 102. Server 102 obtains the three-dimensional reconstruction result of the target object by executing the method provided in this application.
[0055] Terminal 101 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include image acquisition devices, smart vehicle devices, video acquisition devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted displays, etc. Head-mounted displays can include virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc.
[0056] Server 102 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.
[0057] In one exemplary embodiment, such as Figure 2 As shown, a three-dimensional reconstruction method is provided, which can be applied to... Figure 1 Taking the server in the example, the explanation includes the following steps 201 to 206. Wherein:
[0058] Step 201: Obtain multi-view images of the target object and camera parameters when acquiring multi-view images; camera parameters include intrinsic parameters describing the internal optical and geometric properties of the camera and extrinsic parameters describing the position and orientation of the camera in the world coordinate system.
[0059] The target object is the object that needs to be reconstructed in three dimensions.
[0060] Multi-view images of a target object refer to images of the target object acquired from multiple perspectives at different shooting angles; further, the camera parameters corresponding to the acquisition of multi-view images refer to the camera parameters corresponding to each image in the multi-view images.
[0061] The camera parameters corresponding to the image include the camera's intrinsic and extrinsic parameters. The intrinsic parameters are a set of parameters describing the camera's internal optical and geometric characteristics, such as focal length, pixel scaling factor, principal point coordinates, etc., and their core function is to establish the mapping relationship between three-dimensional spatial points and two-dimensional pixel coordinates. The extrinsic parameters are parameters of the camera's position and orientation in the world coordinate system, such as rotation matrix, translation vector, etc., used to transform the world coordinate system to the camera coordinate system. Through the camera parameters corresponding to the image, the mapping relationship between each pixel in the image under different viewpoints can be determined.
[0062] Step 202: Determine the spatial range of the depth plane of the target object and the number of depth planes.
[0063] The depth plane spatial range of the target object refers to any reference image contained in the multi-view image. It defines the actual spatial depth range that the target object may exist in under the view of that reference image.
[0064] For example, when a reference image is selected from multiple images (e.g., reference image a), the actual depth range of the target object under that viewpoint (e.g., 1 meter to 5 meters) can be determined based on scene priors or measurement data. This range constitutes the depth plane spatial range of the target object corresponding to the reference image. The depth plane spatial range of the target object is used to constrain the subsequent 3D reconstruction task.
[0065] The number of depth planes refers to the total number of planes divided within a predefined depth plane space. Furthermore, the number of depth planes directly determines the sampling density of depth information in the target scene: a larger number of depth planes means a more refined division of spatial depth, resulting in higher sampling accuracy and a more accurate representation of the 3D structure; however, too many depth planes can lead to an exponential increase in the amount of data to be processed, thus increasing computational complexity and time consumption. Therefore, in practical applications, it is usually necessary to set the number of depth planes reasonably to achieve an optimal balance between data processing efficiency and 3D reconstruction quality.
[0066] In some embodiments, the number of depth planes can be preset based on the accuracy of 3D reconstruction; the higher the accuracy of 3D reconstruction, the more depth planes there are; conversely, the fewer depth planes there are.
[0067] In other embodiments, the number of depth planes can be determined based on the spatial range of the depth planes; for example, the number of depth planes can be positively correlated with the spatial range of the depth planes. For example, if the spatial range of the depth planes is greater than a first threshold, the number of depth planes can be determined as a first number; if the spatial range of the depth planes is not greater than the first threshold, the number of depth planes can be determined as a second number, wherein the second number is less than the first number.
[0068] Step 203: Determine the pixel movement range and unit movement amount of pixels in the multi-view image based on the spatial range of the depth plane and the number of depth planes.
[0069] In some embodiments, one image can be selected from multiple viewpoints as a reference image. The depth plane spatial range is mapped onto the reference image using the camera parameters corresponding to the reference image, thereby obtaining the pixel movement range of the pixels in the reference image. Specifically, the minimum and maximum depth values of the depth plane spatial range are first determined. Then, a point is randomly selected in the spatial plane where the minimum depth value is located, and this point is projected onto the reference image using the camera parameters. Specifically, the three-dimensional spatial coordinates of the point are transformed from the world coordinate system to the camera coordinate system using the camera extrinsic parameters to obtain the camera coordinates of the point. Then, the camera coordinates are projected and transformed using the camera intrinsic parameters to obtain the pixel coordinates. Similarly, a point is randomly selected in the spatial plane where the maximum depth value is located, and this point is projected onto the reference image using the camera parameters. The line connecting the two points in the reference image is the pixel movement range of the pixels in the reference image.
[0070] Next, the pixel movement range is divided equally according to the number of depth planes to obtain the unit movement amount.
[0071] It is understandable that the relationship between the spatial range of the depth plane and the range of pixel movement in the reference image is actually the relationship between the amount of pixel movement in the reference image and the actual depth when the camera acquiring the reference image moves along the optical axis.
[0072] Step 204: Based on the pixel movement range and unit movement amount, determine the pixel movement position of the pixels in the multi-view image, with a unit movement amount between adjacent pixel movement positions, and the number of pixel movement positions matches the number of depth planes.
[0073] Specifically, based on the pixel movement range, the starting point and ending point of the pixel movement are determined. Then, starting from the starting point, the pixel moves one unit towards the ending point to obtain the first pixel movement position; from the first pixel movement position, it moves another one unit towards the ending point to obtain the second pixel movement position, and so on, to obtain multiple pixel movement positions.
[0074] Step 205: Determine each virtual depth plane based on the position of each pixel movement.
[0075] Specifically, based on the camera parameters of the reference image, the movement position of each pixel in the reference image is reverse-mapped into the depth plane space range to obtain their corresponding three-dimensional spatial position points. Then, based on the spatial depth of the three-dimensional spatial position points, a corresponding virtual depth plane is established.
[0076] Step 206: Based on multi-view images, camera parameters, and each virtual depth plane, perform 3D reconstruction of the target object to obtain the 3D reconstruction result of the target object.
[0077] It should be noted that, due to the perspective projection principle of camera imaging, the relationship between the amount of pixel movement in an image and the actual depth is not linear, but non-uniform. When the depth is small, that is, when the object is close to the camera, a small change in depth will cause a drastic change in pixel displacement; when the depth is large, that is, when the object is far away from the camera, a large change in depth may only cause a small change in pixel displacement.
[0078] In related technologies, the depth plane space is equally divided to obtain multiple virtual depth planes that are equally spaced within the virtual depth space. This results in a larger pixel movement for nearby virtual depth planes, leading to a wider search and matching range for multi-view images when making depth assumptions based on nearby virtual depth planes. Conversely, the smaller pixel movement for distant virtual depth planes results in a narrower search and matching range for multi-view images. However, since pixels in an image are more sensitive to changes in the foreground and less sensitive to changes in the background, denser virtual depth planes are needed in the foreground to ensure sufficient detail capture, while sparser virtual depth planes are needed in the background to avoid overly similar search ranges under different depth assumptions, thus preventing misjudgments of depth. Furthermore, the equal division of the depth plane space in related technologies leads to too few depth assumptions in the foreground, resulting in the loss of some detail information; and too many depth assumptions in the background, filling the background with a large number of visually indistinguishable redundant assumptions, ultimately reducing the overall accuracy of depth estimation.
[0079] In this application, the virtual depth plane is determined based on the pixel movement position, so that the pixel movement amount corresponding to the depth between any two adjacent virtual depth planes is a unit movement amount. This makes the constructed virtual depth plane exhibit a non-uniform physical distribution characteristic of being dense in the near and sparse in the far. In this way, uniform pixel-level search can be performed in the reference image, thus ensuring that the matching search accuracy is similar regardless of the distance of the object. This improves the matching accuracy of the near region and avoids the misjudgment of the depth in the far region, thereby improving the overall accuracy of depth estimation.
[0080] For example, suppose equidistant depth division is performed in related technologies, with a virtual depth plane set every 1 meter within a virtual depth range of 1 to 10 meters. Then, a depth interval of 1 to 2 meters in the near area corresponds to a 50-pixel displacement, while an interval of 9 to 10 meters in the far area corresponds to only a 5-pixel displacement. In this case, each depth assumption in the near area needs to cover a 50-pixel range, meaning that the 50-pixel range is sampled only once, resulting in the loss of some detailed information. While the far area is sampled only once within a 5-pixel range, because the pixel differences in the far area are not significant, the depth information obtained from sampling at different depths may be very similar. This not only fails to provide more valuable depth information but also creates a large number of visually indistinguishable redundant assumptions, interfering with the depth estimation algorithm's judgment of the true depth information and further reducing the overall accuracy of depth estimation.
[0081] In this application's scheme, the pixel movement corresponding to each virtual plane interval is equal, assumed to be 2 pixels. Thus, the virtual depth plane interval at nearby locations might be 0.04 meters, meaning the virtual depth planes are denser nearby, while the virtual depth plane interval at distant locations might be 0.4 meters, meaning the virtual depth planes are sparser at distant locations. In this way, by using virtual depth planes of varying densities at each depth level, regardless of proximity, each depth hypothesis corresponds to a precise search range of 2 pixels. This avoids the sparsity of nearby hypotheses and eliminates the redundancy of distant hypotheses, ultimately significantly improving the depth estimation accuracy across the entire scene.
[0082] The aforementioned 3D reconstruction method, apparatus, device, readable storage medium, and program product determine the pixel movement range and unit movement amount of pixels in a multi-view image by using the spatial range and number of depth planes. Then, based on the pixel movement range and unit movement amount, the pixel movement range is equally divided to obtain multiple pixel movement positions. Since the intervals between the multiple pixel movement positions are the same, the search for depth projection on the image epipolar line is nearly uniform, fundamentally ensuring that the matching search accuracy is consistent regardless of the distance of the object. Subsequently, each virtual depth plane is determined based on the multiple pixel movement positions. Based on the multi-view image, camera parameters, and each virtual depth plane, 3D reconstruction is performed on the target object to obtain the 3D reconstruction result of the target object. The method provided in this application can significantly improve the reconstruction integrity of distant regions, thereby improving the integrity of 3D reconstruction.
[0083] In one exemplary embodiment, such as Figure 3 As shown, step 203 includes steps 301 to 304. Wherein:
[0084] Step 301: Based on the first reference point at the minimum depth in the depth plane spatial range and the camera parameters, determine the first pixel point in the reference image in the multi-view image.
[0085] The first reference point can be any spatial point selected from the virtual depth plane with the minimum depth.
[0086] It is understandable that, based on determining the spatial position of the first reference point, the spatial position of the first reference point can be projected onto the reference image using the camera parameters corresponding to the reference image, thus obtaining the first pixel in the reference image.
[0087] Step 302: Based on the second reference point with the maximum depth in the depth plane spatial range and the camera parameters, determine the second pixel point in the reference image.
[0088] In some embodiments, the second reference point may be any spatial point selected from the virtual depth plane of maximum depth.
[0089] In other embodiments, the second reference point can be determined by the first reference point; for example, based on the determination of the first reference point, the line connecting the first reference point and the second reference point is required to be parallel to the optical axis of the reference image, thereby determining the second reference point; wherein, the optical axis of the reference image refers to a straight line that is perpendicular to the imaging plane of the camera corresponding to the reference image and passes through the center of the camera lens.
[0090] It is understandable that, based on determining the spatial position of the second reference point, the spatial position of the second reference point can be projected onto the reference image using the camera parameters corresponding to the reference image, thus obtaining the second pixel in the reference image.
[0091] Step 303: Determine the pixel movement range based on the first pixel and the second pixel.
[0092] The pixel movement range is the number of pixels included between the first pixel and the second pixel.
[0093] In some embodiments, the line connecting the first pixel and the second pixel is defined as the pixel movement range.
[0094] Step 304: Determine the unit movement amount based on the number of depth planes and the pixel movement range.
[0095] In some embodiments, within the pixel movement range, the pixels are divided equally according to the number of depth planes to obtain the unit movement amount.
[0096] For example, if the pixel movement range is 10 pixels and the number of depth planes is 5, then the unit movement amount is 2 pixels.
[0097] In the above embodiments, the pixel movement range is determined by the minimum and maximum depths in the depth plane space, which ensures that the pixel movement range covers the entire depth plane space. Therefore, when determining the virtual depth plane later, the virtual depth plane can cover the entire depth plane space.
[0098] In an exemplary embodiment, step 205 includes steps 401 to 402. Wherein:
[0099] Step 401: Based on the movement position of each pixel and the camera parameters, determine the spatial reference point corresponding to the movement position of each pixel.
[0100] The pixel movement position is the image coordinate in the reference image. Specifically, starting from the first pixel, it moves one unit towards the second pixel to obtain the first pixel movement position; from the first pixel movement position, it moves another one unit towards the second pixel to obtain the second pixel movement position, and so on, to obtain multiple pixel movement positions.
[0101] It is understandable that, since camera parameters can indicate the mapping relationship between image coordinates and world coordinates, the pixel movement position can be mapped to the world coordinate system through the pixel movement position and camera parameters, so as to obtain the spatial reference point corresponding to each pixel movement position.
[0102] Step 402: Determine each virtual depth plane based on the spatial depth of each spatial reference point.
[0103] For example, if the spatial depth of the spatial reference point corresponding to the first pixel is 1 meter, then the first virtual depth plane is also a plane with a spatial depth of 1 meter; if the spatial depth of the spatial reference point corresponding to the second pixel is 1.5 meters, then the first virtual depth plane is also a plane with a spatial depth of 1.5 meters, and so on.
[0104] Through the above embodiments, the virtual depth plane is determined based on the pixel movement position. Since the interval between each pixel movement position is a unit movement amount, according to the relationship between the pixel and the actual depth, multiple virtual depth planes with dense near and sparse far can be obtained, thus making the virtual depth plane more in line with the actual sampling requirements.
[0105] In one embodiment, step 206 may include steps 501 to 504. Wherein:
[0106] Step 501: Extract features from the multi-view images to obtain the image features of the multi-view images.
[0107] In some embodiments, since a multi-view image may include multiple images, feature extraction of a multi-view image actually involves extracting features from each image of the multi-view image. The image features of a multi-view image also include the image features corresponding to each image of the multi-view image.
[0108] In other embodiments, the multi-view images can be preprocessed before feature extraction, such as adjusting multiple images in the multi-view images to the same resolution or normalizing the multi-view images.
[0109] Step 502: Based on camera parameters and image features, perform feature projection on the virtual depth plane to obtain the feature volume corresponding to each virtual depth plane.
[0110] Specifically, for each image in the multi-view image, the image features of the image are projected onto their respective virtual depth plane to obtain the feature map of each view. Then, the transformation relationship between different views is determined by the camera parameters corresponding to each multi-view image, and the feature maps of each view are uniformly projected onto one view to obtain the feature stack of multiple views under the same depth assumption.
[0111] Step 503: Integrate the feature volumes corresponding to multiple virtual depth planes to obtain a three-dimensional cost volume.
[0112] Specifically, the feature volumes corresponding to multiple virtual depth planes are combined in order of decreasing or increasing virtual depth to obtain a three-dimensional cost volume.
[0113] Understandably, the resulting three-dimensional cost volume represents the stacking of image features for each pixel under different depth assumptions.
[0114] Step 504: Perform 3D reconstruction of the target object based on the 3D cost volume.
[0115] Specifically, by stacking image features under different depth assumptions in the three-dimensional cost volume, feature matching under different depth assumptions can be performed. If the image feature matching degree of a pixel under a certain depth assumption is high enough, then the assumed depth is more likely to be the actual spatial depth of the pixel.
[0116] Through the above embodiments, multiple virtual depth planes are used to make depth assumptions, thereby obtaining the feature volume corresponding to each virtual depth plane. The feature volume corresponding to multiple virtual depth planes is then used to determine the three-dimensional cost volume for three-dimensional reconstruction. Since multiple virtual depth planes exhibit the characteristics of being dense at near distances and sparse at far distances in space, it ensures that the constructed cost volume maintains appropriate matching accuracy at each depth level, thereby improving the integrity of the three-dimensional reconstruction.
[0117] In one embodiment, step 504 may include steps 601 to 603. Wherein:
[0118] Step 601: Perform three-dimensional depth convolution and three-dimensional point convolution on the three-dimensional cost volume in sequence to obtain the intermediate features of the convolution.
[0119] Among them, 3D depth convolution refers to performing 3D depth convolution independently on each channel of the 3D cost volume without channel fusion, to obtain the first intermediate feature that is independent of the channel.
[0120] 3D depth convolution refers to combining features from different channels based on the first intermediate feature to obtain the convolutional intermediate feature.
[0121] It should be noted that 3D depthwise convolution and 3D pointwise convolution are performed sequentially, that is, 3D depthwise convolution is performed first, followed by 3D pointwise convolution.
[0122] Step 602: Perform feature aggregation based on the intermediate features of the convolution to obtain a depth map; the depth map is used to describe the probability distribution of the target object in the depth dimension.
[0123] In some embodiments, the feature matching degree of each pixel in the 3D cost volume under different depth assumptions is calculated by convolutional intermediate features, thereby obtaining the probability of the pixel at different depths. Then, based on the depth probability of each pixel, a depth map for the target object is obtained.
[0124] Step 603: Based on the depth map, perform 3D reconstruction of the target object.
[0125] In some embodiments, the depth with the highest depth probability can be determined as the actual depth of the pixel based on the depth probability distribution of each pixel in the depth map. Then, combined with camera parameters, the two-dimensional pixel coordinates are mapped to three-dimensional spatial coordinates through inverse projection transformation, thereby directly generating a dense point cloud of the target object and realizing the three-dimensional reconstruction of the target object.
[0126] In other embodiments, considering that the number of virtual depth planes is limited and the depth probability distribution of each pixel in the depth map is a discrete probability distribution, the depth map can be first subjected to probability regression processing to obtain the depth probability distribution under continuous depth. Then, the actual depth of the pixel is determined by the depth probability distribution under continuous depth. After that, combined with camera parameters, the two-dimensional pixel coordinates are mapped to three-dimensional spatial coordinates through inverse projection transformation, thereby directly generating a dense point cloud of the target object and realizing the three-dimensional reconstruction of the target object.
[0127] Through the above embodiments, the 3D convolution is decomposed into 3D depth convolution and 3D point convolution. The two convolutions are executed sequentially to form a complete convolution, which effectively aggregates the cost volume information. At the same time, it significantly reduces resource consumption and optimizes resource allocation.
[0128] In one embodiment, step 603 may include steps 701 to 702. Wherein:
[0129] Step 701: Based on the image features of the depth map and the reference image in the multi-view image, feature stitching is performed to obtain a composite feature map.
[0130] In some embodiments, the depth map and the feature map of the reference image can be concatenated along the channel dimension to form a composite feature map with four channels. After processing by a depth residual network, the composite feature map can generate a more refined depth map, thereby improving the effect of subsequent 3D reconstruction.
[0131] Step 702: Perform 3D reconstruction of the target object based on the composite feature map.
[0132] It is understandable that the image features of the reference image can provide rich reference information in the process of 3D reconstruction, such as texture and edges. By stitching the image features of the depth map and the reference image together, a composite feature map is obtained. This composite feature map can retain both geometric depth clues and visual image information, ultimately improving the effect of 3D reconstruction.
[0133] Through the above embodiments, composite feature maps are used to perform 3D reconstruction of target objects. Since composite feature maps simultaneously possess geometric depth lines and visual image information, they can effectively improve edge details and enhance the overall accuracy of 3D reconstruction during the 3D reconstruction process.
[0134] For easier understanding, please refer to Figure 4 , Figure 4 A schematic diagram of the distribution of the virtual depth plane involved in the embodiments of this application is shown.
[0135] in, Figure 4 Figure a shows a schematic diagram of obtaining a virtual depth plane by equally dividing the virtual depth space within a related technology.
[0136] For example, Figure 4 Figure a shows the depths X1 to X5 corresponding to multiple virtual depth planes from near to far. It can be seen that the depth interval between adjacent virtual depth planes is the same. However, when projected onto the reference image, the pixel movement distance from X1 to X2 is... arrive The pixel movement distance from X2 to X3 is arrive Similarly, it can be seen that the virtual depth plane in the near area has a large search range in the image, while the virtual depth plane in the far area has a small search range. Combined with what was mentioned earlier, pixels in the image are more sensitive to changes in the near area and less sensitive to changes in the far area. This means that there are fewer effective searches in the near area, resulting in the loss of detailed information, while there will be more similar invalid searches in the far area, leading to depth misjudgment.
[0137] Figure 4 Figure b shows a schematic diagram of the method of this application for obtaining a virtual depth plane.
[0138] Figure 4Figure b shows the depths X1 to X5 corresponding to multiple virtual depth planes from near to far. It can be seen that the depth intervals between adjacent virtual depth planes are different, exhibiting a denser distribution near the edge and a sparser distribution far away. This is because when determining the virtual depth planes, the minimum depth X1 and the maximum depth X5 are determined first, and then the corresponding pixel movement range is calculated. arrive Divide the pixel movement range equally to obtain arrive The five pixels are moved and then back-projected onto the virtual depth plane, resulting in five virtual depth planes X1 to X5. This allows the image to be denser in the foreground and sparser in the background. In this way, combined with what was mentioned earlier, the pixels in the image are more sensitive to changes in the foreground and less sensitive to changes in the background. This means that there are more depth assumptions for the foreground, which can capture more detailed information, while there are fewer depth assumptions for the background, which avoids depth misjudgment caused by different depth assumptions being too similar.
[0139] In some embodiments, please refer to Figure 5 , Figure 5 This paper illustrates a three-dimensional reconstruction process according to an embodiment of the present application, including steps one through five:
[0140] Step 1: Input a set of images and camera parameters from different viewpoints. Specifically, input a set of images from different viewpoints and their corresponding camera calibration parameters. The input images include a reference image and multiple source images. Adjust the images to the same resolution and perform normalization processing.
[0141] Step two involves feature extraction using a convolutional neural network (CNN), generating feature maps of the original size, half the original size, and one-quarter of the original size. Specifically, a neural network consisting of eight 2D convolutional layers is used for feature extraction. Except for the last layer, which consists of only one convolutional layer, all other layers of the CNN include convolutional layers, batch normalization layers, and activation functions. The stride of the third and sixth convolutional layers is set to 2, while the stride of the other layers is 1. After feature extraction, three different sizes of feature maps are generated: the original size, half the original size, and one-quarter of the original size. These three different feature maps retain rich spatial detail information and ensure a larger receptive field, resulting in stronger robustness to noise and weakly textured regions.
[0142] Step 3: Uniform sampling is performed in the inverse depth space, and a cost volume is constructed through differentiable homography transformation. Here, inverse depth is defined as the reciprocal of the actual depth. Since the pixel movement is not linearly related to the actual depth when the camera moves along the optical axis, but approximately linearly related to the inverse depth, uniform sampling in the inverse depth space is equivalent to uniformly dividing pixels along the epipolar lines of the image. This is the same as the pixel movement position determination process mentioned in the previous embodiment, thus achieving near-uniform pixel-level search and ensuring that the matching search accuracy is similar regardless of the object's distance. After feature extraction, the network performs uniform sampling in the inverse depth space based on the minimum inverse depth (corresponding to the reciprocal of the maximum depth), the maximum inverse depth (corresponding to the reciprocal of the minimum depth), and the assumed number of inverse depth samples, generating a series of virtual depth planes parallel to the reference image plane. These planes are uniformly distributed along the principal optical axis of the reference image in the inverse depth dimension. Their corresponding actual depth values are calculated in reverse from the inverse depth. Adjacent depth values correspond to the same pixel movement, thus exhibiting a non-uniform physical distribution characteristic of being denser near the edges and sparser in the distance. These depth planes define sampling points in 3D space that better conform to perspective geometry, used for subsequent feature projection and matching.
[0143] Subsequently, the network constructs the cost body through differentiable homography transformation, the specific process of which is as follows: First, based on the reference camera coordinate system, the homography transformation matrix between each inverse depth sampling plane and each source view is calculated; then, the feature map of each source view is projected onto the reference camera's view frustum through the corresponding homography matrix, forming a feature stack of multiple views under the same inverse depth assumption. In this process, since inverse depth sampling ensures that the pixel displacement projected onto the source view is approximately uniformly distributed, the network uses bilinear interpolation to adjust the projected features, thereby ensuring the spatial consistency of feature projection on different inverse depth planes. The cost body constructed in this way maintains optimal matching accuracy at all depth levels, especially improving the sampling density in the near-range region, and still maintaining sufficient sampling density in the far-range region, laying a solid geometric foundation for subsequent cost body regularization and depth inference.
[0144] Step four involves applying 3D separable convolutions to the cost volume using both 3D depthwise convolutions and 3D pointwise convolutions to aggregate cost volume information. Specifically, depthwise separable convolutions are used instead of convolutional neural networks for cost volume regularization. This method decomposes the 3D convolution into 3D depthwise convolutions and 3D pointwise convolutions. The 3D depthwise convolution is performed independently on the cost volume of each channel, resulting in intermediate feature maps independent of the channels. The 3D pointwise convolution is applied to the feature maps through 3D pointwise convolution, aggregating relevant information from each channel. The two convolutions are executed sequentially to form a complete convolution, effectively aggregating cost volume information and achieving the effects of noise reduction, smoothing depth estimation, and improving matching accuracy. Compared to existing technologies that use multi-scale 3D convolutional neural networks for cost volume regularization, this method reduces computational load and resource consumption.
[0145] Step 5 involves weighting the depth hypotheses along the depth dimension to construct a depth map of the image. Specifically, all possible depth hypotheses are weighted along the depth dimension by multiplying each hypothetical depth value by its corresponding probability weight and summing the results to construct a preliminary depth map. This depth map not only contains the image's 3D geometric information but also effectively fuses multiple depth hypotheses through a weighted averaging strategy, thereby improving the accuracy and stability of the final depth estimation. Since the large-scale perception generated during regularization may lead to over-smoothing in edge regions of the reconstructed 3D model, a depth residual network is used to refine the depth map after generating the preliminary one. The preliminary depth map is concatenated with the feature map of the reference image along the channel dimension to form a composite feature map with four channels. This composite feature map, after processing by the depth residual network, generates a more refined depth map, effectively improving edge details and enhancing the overall accuracy and reliability of the depth estimation.
[0146] Through the above application process, the inverse depth sampling strategy makes the distribution of search points on the epipolar line more uniform under the perspective projection model, ensuring similar pixel-level search accuracy on the image plane regardless of the distance of the object from the camera. This improves the low sampling accuracy problem of traditional uniform depth sampling. Defining depth estimation as a classification problem, the network only needs to select the most likely one from a set of discrete depth hypotheses, rather than precisely regressing a continuous value. This makes the method insensitive to noise and matching ambiguity in the cost volume, avoiding local optima and large prediction biases that are prone to occur in regression models. The combination of multi-scale feature pyramids and 3D regularized networks can capture contextual information from local details to the global scene. When facing repetitive structures, features with large receptive fields can provide crucial scene layout information, thereby effectively distinguishing between true and false matches. At object edges, fine-scale features better preserve edge sharpness and reduce depth confusion between foreground objects and background. Furthermore, using 3D depthwise separable convolutions instead of 3D convolutional neural networks for cost volume regularization significantly reduces computational complexity and resource consumption while maintaining model performance. The cost volume is a four-dimensional tensor, and performing 3D convolution operations on it is extremely computationally intensive. 3D depthwise separable convolutions decouple standard convolution into two efficient steps: pointwise convolution and depthwise convolution, reducing computational cost to 1 / 10 to 1 / 30 of standard 3D convolutions. Simultaneously, due to the simplified computation process, the storage requirements for intermediate activation values are significantly reduced, lowering the demand on GPU memory and other hardware resources. This allows for the processing of high-resolution, deep-sample cost volumes even with limited resources. Similarly, the number of model parameters can be reduced by an order of magnitude. This not only reduces the model's storage requirements, facilitating deployment on storage-constrained devices, but also reduces memory bandwidth requirements during model inference, further improving inference speed.
[0147] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0148] Based on the same inventive concept, this application also provides a three-dimensional reconstruction apparatus for implementing the three-dimensional reconstruction method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the three-dimensional reconstruction apparatus provided below can be found in the limitations of the three-dimensional reconstruction method described above, and will not be repeated here.
[0149] In one exemplary embodiment, such as Figure 6 As shown, a three-dimensional reconstruction device is provided. The three-dimensional reconstruction device 800 includes:
[0150] The acquisition module 801 is used to acquire multi-view images of the target object and camera parameters when acquiring multi-view images; the camera parameters include intrinsic parameters that describe the internal optical and geometric characteristics of the camera and extrinsic parameters that describe the position and attitude of the camera in the world coordinate system.
[0151] The depth determination module 802 is used to determine the spatial range of the depth plane of the target object and the number of depth planes.
[0152] The pixel determination module 803 is used to determine the pixel movement range and unit movement amount of pixels in a multi-view image based on the spatial range of the depth plane and the number of depth planes.
[0153] The segmentation module 804 is used to determine the pixel movement position of pixels in a multi-view image based on the pixel movement range and unit movement amount, with a unit movement amount between adjacent pixel movement positions, and the number of pixel movement positions matches the number of depth planes.
[0154] The plane determination module 805 determines each virtual depth plane based on the movement position of each pixel.
[0155] The reconstruction module 806 is used to perform three-dimensional reconstruction of the target object based on multi-view images, camera parameters and various virtual depth planes, and obtain the three-dimensional reconstruction result of the target object.
[0156] In one embodiment, the pixel determination module 803 is used to determine a first pixel in a reference image in a multi-view image based on a first reference point at the minimum depth in the depth plane spatial range and camera parameters; determine a second pixel in the reference image based on a second reference point at the maximum depth in the depth plane spatial range and camera parameters; determine a pixel movement range based on the first pixel and the second pixel; and determine a unit movement amount based on the number of depth planes and the pixel movement range.
[0157] In one embodiment, the plane determination module 805 is used to determine the spatial reference point corresponding to the movement position of each pixel based on the movement position of each pixel and camera parameters; and to determine each virtual depth plane based on the spatial depth of each spatial reference point.
[0158] In one embodiment, the reconstruction module 806 is used to extract features from multi-view images to obtain image features of multi-view images; to perform feature projection on virtual depth planes based on camera parameters and image features to obtain feature volumes corresponding to each virtual depth plane; to integrate feature volumes corresponding to multiple virtual depth planes to obtain a three-dimensional cost volume; and to perform three-dimensional reconstruction of the target object based on the three-dimensional cost volume.
[0159] In one embodiment, the reconstruction module 806 is specifically used to sequentially perform three-dimensional depth convolution and three-dimensional point convolution on the three-dimensional cost volume to obtain convolution intermediate features; perform feature aggregation based on the convolution intermediate features to obtain a depth map; the depth map is used to describe the probability distribution of the target object in the depth dimension; and perform three-dimensional reconstruction on the target object based on the depth map.
[0160] In one embodiment, the reconstruction module 806 is specifically used to perform feature stitching based on image features of the depth map and reference images in multi-view images to obtain a composite feature map; and to perform three-dimensional reconstruction of the target object based on the composite feature map.
[0161] Each module in the aforementioned 3D reconstruction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0162] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores 3D reconstruction data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network. When the computer program is executed by the processor, it implements a 3D reconstruction method.
[0163] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0164] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described above.
[0165] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0166] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method described above.
[0167] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0168] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0169] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0170] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A three-dimensional reconstruction method, characterized by, The method includes: Acquire multi-view images of the target object and camera parameters when acquiring the multi-view images; the camera parameters include intrinsic parameters describing the internal optical and geometric properties of the camera and extrinsic parameters describing the position and orientation of the camera in the world coordinate system; Determine the depth plane spatial range and the number of depth planes of the target object; Based on the first reference point at the minimum depth in the depth plane spatial range and the camera parameters, the first pixel point in the reference image in the multi-view image is determined; The second pixel in the reference image is determined based on the second reference point with the maximum depth in the depth plane spatial range and the camera parameters; Based on the first pixel and the second pixel, determine the pixel movement range; The unit movement amount is determined based on the number of depth planes and the pixel movement range; Based on the pixel movement range and the unit movement amount, the pixel movement position of the pixel in the multi-view image is determined, the unit movement amount is spaced between adjacent pixel movement positions, and the number of pixel movement positions matches the number of depth planes; Based on the pixel movement position and the camera parameters, determine the spatial reference point corresponding to each pixel movement position; Based on the spatial depth of each of the aforementioned spatial reference points, determine each virtual depth plane; Based on the multi-view images, the camera parameters, and each of the virtual depth planes, a 3D reconstruction is performed on the target object to obtain the 3D reconstruction result of the target object.
2. The method of claim 1, wherein, Based on the multi-view images, the camera parameters, and each of the virtual depth planes, a 3D reconstruction is performed on the target object to obtain the 3D reconstruction result of the target object, including: Feature extraction is performed on the multi-view images to obtain the image features of the multi-view images; Based on the camera parameters and the image features, feature projection is performed on the virtual depth plane to obtain the feature volume corresponding to each virtual depth plane; By integrating the feature volumes corresponding to the multiple virtual depth planes, a three-dimensional cost volume is obtained; Based on the three-dimensional cost volume, a three-dimensional reconstruction is performed on the target object.
3. The method of claim 2, wherein, Based on the three-dimensional cost volume, a three-dimensional reconstruction of the target object is performed, including: The three-dimensional cost volume is subjected to three-dimensional depth convolution and three-dimensional point convolution in sequence to obtain the convolution intermediate features; Feature aggregation is performed based on the intermediate features of the convolution to obtain a depth map; the depth map is used to describe the probability distribution of the target object in the depth dimension. Based on the depth map, a 3D reconstruction is performed on the target object.
4. The method of claim 3, wherein, Based on the depth map, a 3D reconstruction is performed on the target object, including: Based on the image features of the depth map and the reference image in the multi-view image, a composite feature map is obtained by feature stitching. Three-dimensional reconstruction of the target object is performed based on the composite feature map.
5. A three-dimensional reconstruction apparatus, characterized by comprising: The device includes: The acquisition module is used to acquire multi-view images of the target object and camera parameters when acquiring the multi-view images; the camera parameters include intrinsic parameters describing the internal optical and geometric characteristics of the camera and extrinsic parameters describing the position and attitude of the camera in the world coordinate system. The depth determination module is used to determine the spatial range of the depth plane and the number of depth planes of the target object. A pixel determination module is used to determine a first pixel in a reference image of the multi-view image based on a first reference point at the minimum depth in the depth plane spatial range and the camera parameters; determine a second pixel in the reference image based on a second reference point at the maximum depth in the depth plane spatial range and the camera parameters; determine a pixel movement range based on the first pixel and the second pixel; and determine a unit movement amount based on the number of depth planes and the pixel movement range. The segmentation module is used to determine the pixel movement position of a pixel in the multi-view image based on the pixel movement range and the unit movement amount, wherein the unit movement amount is spaced between adjacent pixel movement positions, and the number of pixel movement positions matches the number of depth planes; The plane determination module is used to determine the spatial reference point corresponding to each pixel movement position based on the pixel movement position and the camera parameters; and to determine each virtual depth plane based on the spatial depth of each spatial reference point. The reconstruction module is used to perform three-dimensional reconstruction of the target object based on the multi-view images, the camera parameters, and each of the virtual depth planes, to obtain the three-dimensional reconstruction result of the target object.
6. The apparatus of claim 5, wherein, The reconstruction module is used to extract features from the multi-view images to obtain image features of the multi-view images; to perform feature projection on the virtual depth plane based on the camera parameters and the image features to obtain feature volumes corresponding to each virtual depth plane; to integrate the feature volumes corresponding to multiple virtual depth planes to obtain a three-dimensional cost volume; and to perform three-dimensional reconstruction of the target object based on the three-dimensional cost volume.
7. The apparatus of claim 6, wherein, The reconstruction module is used to sequentially perform 3D depth convolution and 3D point convolution on the 3D cost volume to obtain convolution intermediate features; perform feature aggregation based on the convolution intermediate features to obtain a depth map; the depth map is used to describe the probability distribution of the target object in the depth dimension; and perform 3D reconstruction on the target object based on the depth map. 8.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-7. When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Multi-view three-dimensional network three-dimensional reconstruction method
CN115239871A
Three-dimensional reconstruction method and system based on deep learning
CN118941713A