Three-dimensional point cloud reconstruction method and device and electronic equipment

By extracting multi-stage feature maps of multi-view images and using regularized networks to generate depth maps, the problems of high memory usage and high computing cost in the prior art are solved, and efficient three-dimensional point cloud reconstruction is achieved.

CN119963731AActive Publication Date: 2025-05-09PEKING UNIV

Patent Information

Application Number
CN202510038525.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-09
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The existing multi-image stereo matching method based on deep learning models occupies high video memory and high computing cost in three-dimensional point cloud reconstruction, resulting in inefficient efficiency.

Method used

By acquiring multi-view images and corresponding camera parameters, the multi-stage feature map of each image is extracted, and the matching cost is regularized based on the regularization network to generate a depth map for three-dimensional point cloud reconstruction.

Benefits of technology

It reduces the memory usage and computing costs, and improves the speed and efficiency of three-dimensional point cloud reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963731A_ABST
    Figure CN119963731A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional point cloud reconstruction method and apparatus, and an electronic device. The method comprises the steps of obtaining a first multi-view image associated with a first scene and a corresponding first camera parameter; extracting a first feature map of each first image corresponding to N stages, wherein the first feature map comprises a first depth feature of each pixel point of the corresponding stage; determining the matching cost of a first source image in the first multi-view image and a first reference image in the first multi-view image under each depth value of the Nth stage according to a first feature map of the first image corresponding to the N stages and the first camera parameter; based on the regularization network of the Nth stage, regularizing the matching cost under each depth value of the Nth stage to obtain a first depth map of the first source image; and performing three-dimensional point cloud reconstruction on the first scene according to the first source image and the first depth map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the technical field of three-dimensional point cloud reconstruction, and more specifically, to a three-dimensional point cloud reconstruction method, device and electronic device. Background Art

[0002] 3D point cloud reconstruction is a computer graphics and computer vision technology that constructs a 3D model of an object or scene by processing its point cloud data.

[0003] Multi-image stereo matching (MVS) is a key technology in the field of computer vision for recovering the three-dimensional structure of a scene from multiple images from different perspectives. It reconstructs the three-dimensional model of the scene by analyzing multiple images taken from different perspectives.

[0004] The multi-image stereo matching method based on deep learning model follows the traditional plane scanning algorithm, discretizes the depth range into a limited number of depth value candidates, and constructs a matching cost to measure the difference between the image features of multiple views and each depth value, resulting in high video memory usage and high computational cost. Summary of the invention

[0005] An objective of the embodiments of the present disclosure is to provide a three-dimensional point cloud reconstruction method, device and electronic device.

[0006] According to a first aspect of an embodiment of the present disclosure, a three-dimensional point cloud reconstruction method is provided, comprising:

[0007] Acquire a first multi-view image associated with a first scene and corresponding first camera parameters, wherein the first multi-view image is a first image obtained by photographing the first scene from multiple perspectives using a plurality of first cameras with different postures, and the first camera parameters include posture information of each first camera;

[0008] Extracting a first feature map of N stages corresponding to each first image, wherein the first feature map includes a first depth feature of each pixel point of the corresponding stage; N is a positive integer;

[0009] Determine, according to the first feature map of the first image corresponding to the N stages and the first camera parameter, a matching cost between a first source image in the first multi-view image and a first reference image in the first multi-view image at each depth value in the N stages;

[0010] Based on the regularization network of the Nth stage, regularize the matching cost at each depth value of the Nth stage to obtain a first depth map of the first source image;

[0011] A three-dimensional point cloud is reconstructed for the first scene according to the first source image and the first depth map.

[0012] Optionally, the method further includes:

[0013] traversing a plurality of first images in the first multi-view images, and taking a currently traversed first image as a first source image, and taking at least one first image in the first multi-view images other than the first source image as a first reference image;

[0014] determining a first depth map of the first source image;

[0015] Determine whether there is an untraversed first image in the first multi-view image. If so, continue to perform the step of traversing multiple first images in the first multi-view image; if not, end the traversal and obtain a first depth map corresponding to each first image in the first multi-view image.

[0016] Optionally, determining, according to the first feature map of the first image corresponding to the N stages and the first camera parameter, a matching cost between the first source image in the first multi-view image and the first reference image in the first multi-view image at each depth value in the N stages includes:

[0017] Determine a depth value of the Nth stage according to the first feature map of the first N-1 stages corresponding to the first image and the first camera parameter;

[0018] A matching cost at each depth value of the Nth stage is obtained according to the depth value of the Nth stage, the first feature map of the first image corresponding to the Nth stage, and the first camera parameter.

[0019] Optionally, determining the depth value of the Nth stage according to the first feature map of the first N-1 stages corresponding to the first image and the first camera parameter includes:

[0020] Determine, according to the first feature map of the first N-1 stages corresponding to the first image and the first camera parameter, a matching cost between the first source image and the first reference image at each depth value of the N-1 stage;

[0021] Based on the regularization network of the N-1th stage, regularize the matching cost at each depth value of the N-1th stage to obtain a predicted depth map of the first source image at the N-1th stage;

[0022] Determine a depth value at the Nth stage according to the predicted depth map of the first source image at the N-1th stage.

[0023] Optionally, obtaining the matching cost at each depth value of the Nth stage according to the depth value of the Nth stage, the first feature map of the Nth stage corresponding to the first image, and the first camera parameter includes:

[0024] For each depth value of the Nth stage, project each pixel point corresponding to the Nth stage in the first source image onto the first reference image according to the corresponding depth value according to the first camera parameter to obtain a corresponding projected pixel point;

[0025] The difference in the first depth feature between each pixel point of the first source image at each depth value of the Nth stage and the corresponding projection pixel point is compared to obtain a matching cost at each depth value of the Nth stage.

[0026] Optionally, the regularization network based on the Nth stage performs regularization processing on the matching cost at each depth value in the Nth stage to obtain a first depth map of the first source image, including:

[0027] Based on the regularization network of the Nth stage, regularize the matching cost at each depth value of the Nth stage to obtain the probability distribution of the first source image at each depth value of the Nth stage;

[0028] According to the probability distribution of each depth value of the first source image at the Nth stage, a predicted depth map of the first source image at the Nth stage is obtained as the first depth map.

[0029] Optionally, the method further comprises the step of training a regularized network in N stages, comprising:

[0030] Acquire a second multi-view image associated with a second scene and corresponding second camera parameters, where the second multi-view image is a second image obtained by photographing the second scene from multiple perspectives using a plurality of second cameras in different postures, and the second camera parameters include posture information of each second camera;

[0031] Extracting a second feature map of the second image corresponding to N stages, where the second feature map includes a second depth feature of each pixel point in the corresponding stage;

[0032] Determine, according to the second feature map of the N stages corresponding to the second image and the second camera parameter, a matching cost between a second source image in the second multi-view image and a second reference image in the second multi-view image at each depth value of the N stages;

[0033] Obtaining true depth maps of N stages corresponding to the second source image;

[0034] The N-stage regularized network is trained according to the matching cost at each depth value of the N stages, and the network parameters of the N-stage regularized network are updated by minimizing the loss between the second depth map output by the N-stage regularized network and the corresponding true depth map as the training goal.

[0035] Optionally, the training of the regularized network of the N stages according to the matching cost at each depth value of the N stages includes:

[0036] Taking the network parameters of the regularized network at the N stages as variables, and determining the second depth map of the second source image at the N stages according to the matching cost at each depth value at the N stages;

[0037] For the regularized network of each stage, constructing the first loss function of the regularized network of the corresponding stage according to the second depth map and the true depth map of the corresponding stage;

[0038] According to the weights corresponding to the N stages, the first loss function of each stage is weighted summed to obtain the second loss function;

[0039] Solve the network parameters of the regularized network of N stages when the value of the second loss function is minimum, and complete the training of the regularized network of N stages.

[0040] According to a second aspect of the present disclosure, a three-dimensional point cloud reconstruction device is provided, comprising:

[0041] A first acquisition module is used to acquire a first multi-view image associated with a first scene and corresponding first camera parameters, wherein the first multi-view image is a first image obtained by shooting the first scene from multiple perspectives using a plurality of first cameras with different postures, and the first camera parameters include posture information of each first camera;

[0042] A feature extraction module, used to extract a first feature map corresponding to N stages of each first image, wherein the first feature map includes a first depth feature of each pixel point of the corresponding stage; N is a positive integer;

[0043] a cost determination module, configured to determine, according to the first feature map of the first image corresponding to the N stages and the first camera parameter, a matching cost between a first source image in the first multi-view image and a first reference image in the first multi-view image at each depth value in the N stages;

[0044] A depth estimation module obtains a regularized network based on the Nth stage, performs regularization processing on the matching cost at each depth value of the Nth stage, and obtains a first depth map of the first source image;

[0045] A point cloud reconstruction module is used to reconstruct a three-dimensional point cloud of the first scene according to the first source image and the first depth map.

[0046] According to a third aspect of the present disclosure, an electronic device is provided, comprising a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the method described in the first aspect of the present disclosure under the control of the computer program.

[0047] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described in the first aspect of the present disclosure is implemented.

[0048] Through the embodiments of the present disclosure, first multi-view images and corresponding first camera parameters associated with a first scene are obtained, and each first image corresponds to a first feature map of N stages; according to the first feature map and the first camera parameters corresponding to the first image of N stages, the matching cost of a first source image in the first multi-view image and a first reference image in the first multi-view image at each depth value of the N stage is determined; based on a regularization network of the Nth stage, the matching cost at each depth value of the Nth stage is regularized to obtain a first depth map of the first source image; and three-dimensional point cloud reconstruction of the first scene is performed according to the first source image and the first depth map, which can reduce the occupancy of video memory and the computing cost, and improve the speed of three-dimensional point cloud reconstruction.

[0049] Further features and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the present invention with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.

[0051] Figure 1 is a block diagram showing a hardware configuration of an electronic device that can implement an embodiment of the present disclosure;

[0052] Figure 2 is a flow chart of a three-dimensional point cloud reconstruction method according to an embodiment of the present disclosure;

[0053] Figure 3 is a schematic diagram of an example of a three-dimensional point cloud reconstruction method according to an embodiment of the present disclosure;

[0054] Figure 4 is a block diagram of a three-dimensional point cloud reconstruction device according to an embodiment of the present disclosure;

[0055] Figure 5 is a block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0056] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present invention unless otherwise specifically stated.

[0057] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.

[0058] Technologies, methods and equipment known to persons of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods and equipment should be considered part of the specification.

[0059] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0060] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0061] <Hardware Configuration>

[0062] Figure 1 1 is a block diagram showing a hardware configuration of an electronic device 1000 that can implement an embodiment of the present disclosure.

[0063] The electronic device 1000 may be a portable computer, a desktop computer, a mobile phone, a tablet computer, etc. Figure 1 As shown, the electronic device 1000 may include a processor 1100, a memory 1200, an interface device 1300, a communication device 1400, a display device 1500, an input device 1600, a speaker 1700, a microphone 1800, and the like. Among them, the processor 1100 may be a processor CPU, a microprocessor MCU, and the like. The memory 1200 includes, for example, a ROM (read-only memory), a RAM (random access memory), a non-volatile memory such as a hard disk, and the like. The interface device 1300 includes, for example, a USB interface, a headphone interface, and the like. The communication device 1400 is, for example, capable of wired or wireless communication, and may specifically include Wifi communication, Bluetooth communication, 2G / 3G / 4G / 5G communication, and the like. The display device 1500 is, for example, a liquid crystal display screen, a touch display screen, and the like. The input device 1600 may include, for example, a touch screen, a keyboard, a somatosensory input, and the like. The user may input / output voice information through the speaker 1700 and the microphone 1800.

[0064] Figure 1 The electronic device shown is merely illustrative and does not in any way imply any limitation on the present disclosure, its application or use. In the embodiments of the present disclosure, the memory 1200 of the electronic device 1000 is used to store instructions, and the instructions are used to control the processor 1100 to operate to perform any one of the methods provided in the embodiments of the present disclosure. It should be understood by those skilled in the art that although Figure 1 In the electronic device 1000, multiple devices are shown, but the present disclosure may only involve some of the devices, for example, the electronic device 1000 only involves the processor 1100 and the memory 1200. A technician can design instructions according to the scheme disclosed in the present disclosure. How instructions control the processor to operate is well known in the art, so it will not be described in detail here.

[0065] <Method Example>

[0066] The present disclosure provides a three-dimensional point cloud reconstruction method, which can be implemented by an electronic device. For example, the electronic device can be the electronic device 1000 described above.

[0067] Figure 2 The figure is a flowchart of a three-dimensional point cloud reconstruction method according to an embodiment of the present disclosure.

[0068] like Figure 2 As shown, the three-dimensional point cloud reconstruction method includes steps S2100 to S2500 as shown below:

[0069] Step S2100, obtaining a first multi-view image and corresponding first camera parameters associated with a first scene, wherein the first multi-view image is a first image obtained by photographing the first scene from multiple perspectives using a plurality of first cameras in different postures, and the first camera parameters include posture information of each first camera.

[0070] In this embodiment, multiple cameras may respectively take photos of the first scene to be modeled from multiple viewing angles to obtain images corresponding to each camera.

[0071] In order to ensure the accuracy of modeling, it is required that the overlap between each image is relatively high, that is, the perspective difference between different images cannot be too large, so as to reflect the smoothness of the perspective change.

[0072] In order to ensure the accuracy of modeling, the resolution of each first image may be the same.

[0073] Step S2200, extracting the first feature maps of N stages corresponding to each first image.

[0074] In this embodiment, each first feature map may include the first depth feature of each pixel point in the corresponding stage.

[0075] In one embodiment, a feature extraction network may be used to perform feature extraction processing on each first image to obtain a first feature map corresponding to the first image.

[0076] In this embodiment, the feature extraction network may be any neural network that can extract features from an image, and the specific structure of the feature extraction network is not limited herein.

[0077] In one embodiment, the feature extraction network may be a feature pyramid network to extract multi-scale deep image features.

[0078] In one embodiment, based on the feature pyramid network, first feature maps of N stages corresponding to each first image may be obtained, where N is a positive integer. That is, for each first image, first feature maps corresponding to N different stages may be obtained.

[0079] Furthermore, the size of the first feature map of the j-th stage is smaller than the size of the first feature map of the j+1-th stage, j is a positive integer less than or equal to N, and the size of the first feature map of the N-th stage is the same as the size of the first image.

[0080] Wherein, N may be set in advance according to the number of layers of the cascade structure of the subsequent regularized network, and may be the same as the number of layers of the cascade structure of the regularized network. For example, N may be 3.

[0081] In the example where N is 3, it can be obtained that each first image corresponds to the first feature map of three stages, the size of the first feature map of the first stage is smaller than the size of the first feature map of the second stage, the size of the first feature map of the second stage is smaller than the size of the first feature map of the third stage, and the size of the first feature map of the third stage is the same as the size of the first image. For example, the size of the first feature map of the second stage may be 1 / 2 of the size of the first feature map of the third stage, and the size of the first feature map of the first stage may be 1 / 2 of the size of the first feature map of the second stage. For another example, the size of the first feature map of the second stage may be 1 / 4 of the size of the first feature map of the third stage, and the size of the first feature map of the first stage may be 1 / 4 of the size of the first feature map of the second stage.

[0082] In an example, on the basis of obtaining the first feature maps of N stages corresponding to each first image, the first feature maps of N stages corresponding to each first image can also be fused separately to obtain a fused feature map corresponding to the first image to capture image details and contextual information at different levels.

[0083] Step S2300: Determine the matching cost of a first source image in a first multi-view image and a first reference image in the first multi-view image at each depth value in the Nth stage according to the first feature map and the first camera parameter corresponding to the first image in the Nth stage.

[0084] In one embodiment, the first source image may be any first image in the first multi-view images.

[0085] In one embodiment of the present disclosure, the method further includes: traversing multiple first images in the first multi-view image, and taking the currently traversed first image as the first source image, and taking at least one first image other than the first source image in the first multi-view image as the first reference image; determining a first depth map of the first source image; determining whether there are untraversed first images in the first multi-view image, and if so, continuing to execute the step of traversing multiple first images in the first multi-view image; if not, ending the traversal to obtain a first depth map corresponding to each first image in the first multi-view image.

[0086] In this embodiment, the first depth map of the first source image may be determined through steps S2300 to S2400 of this embodiment.

[0087] By traversing a plurality of first images in the first multi-view images, a first depth map of each first image in the first multi-view images may be obtained.

[0088] In one embodiment of the present disclosure, N is greater than 1, then, according to the first feature map and the first camera parameter corresponding to the first image at the Nth stage, determining the matching cost of the first source image in the first multi-view image and the first reference image in the first multi-view image at each depth value at the Nth stage may include the following steps S2310 to S2320:

[0089] Step S2310, determining a depth value of the Nth stage according to the first feature map and the first camera parameter of the first N-1 stages corresponding to the first image.

[0090] In this embodiment, determining the depth value of the Nth stage according to the first feature map and the first camera parameter of the first N-1 stages corresponding to the first image may include steps S2311 to S2313 as shown below:

[0091] Step S2311, determining the matching cost between the first source image and the first reference image at each depth value at the N-1th stage according to the first feature map and the first camera parameter corresponding to the first image at the first N-1th stage.

[0092] In the case of N=2, the depth value of the N-1th stage may be a set of depth values ​​set according to an application scenario or specific requirements, for example, may be an integer value of 1-M.

[0093] When N is greater than 2, the depth value of the N-1th stage may be determined according to the first feature map and the first camera parameter of the first N-2 stages. For details, please refer to steps S2311 to S2313 in this embodiment, which will not be described in detail here.

[0094] Furthermore, the step of determining the matching cost between the first source image and the first reference image at each depth value in the N-1th stage may refer to steps S2321 to S2322 in this embodiment, which will not be described in detail herein.

[0095] Step S2312: Based on the regularization network of the N-1th stage, regularization processing is performed on the matching cost at each depth value of the N-1th stage to obtain a predicted depth map of the first source image at the N-1th stage.

[0096] In this embodiment, based on the regularization network of the N-1th stage, the matching cost at each depth value of the N-1th stage is regularized to obtain the predicted depth map of the first source image at the N-1th stage, which may include the following steps S23121-S23122:

[0097] Step S23121, based on the regularization network of the N-1th stage, regularize the matching cost at each depth value of the N-1th stage to obtain the probability distribution of the first source image at each depth value of the N-1th stage.

[0098] In this embodiment, N cascaded regularization networks can be pre-set as a hybrid regularization network, which combines the cyclic regularization method of the recurrent neural network (RNN) and the multi-stage cascade structure from coarse to fine. In each regularization stage, the input matching cost is sliced ​​along the depth dimension, and each slice is fed into a 2D TransConv module in turn. The output of each layer of 2D TransConv is connected to a ConvLSTM module, so that the network can effectively process sequence information and maintain spatial continuity. As the depth of the network gradually increases, the number and interval of depth candidates gradually decrease, thereby achieving a gradually refined depth estimation. This is achieved by gradually refining each stage in the cascade structure, allowing the network to optimize and adjust the depth estimation at a higher level. This design not only uses spatial constraints to reduce noise and enhance the fragmented smoothness of the depth map, but also ensures the optimization of calculation efficiency and video memory usage.

[0099] In the embodiment where N is 3, three cascaded regularization networks correspond to three stages from coarse to fine, including a coarse depth estimation stage, an intermediate depth refinement stage, and a detailed depth refinement stage, and each stage focuses on further refining the depth estimation through loop and cascade regularization methods.

[0100] The main task of the rough depth estimation stage corresponding to the first regularization network is to quickly locate possible depth candidate intervals starting from a coarse depth range. This stage does not need to be very accurate, but it should be able to determine the approximate depth distribution area to provide a basis for subsequent stages.

[0101] Specifically, the matching cost can be processed by using 2D TransConv and the connected ConvLSTM module, where 2D TransConv helps capture spatial features, while ConvLSTM processes sequence information in the depth dimension and provides temporal memory function for depth slices. At this stage, the number of depth candidates is relatively large and the depth range is also wide, so the granularity of depth estimation is coarse.

[0102] The intermediate depth refinement stage corresponding to the second regularization network further narrows the depth range and increases the accuracy of the depth estimate based on the coarse depth estimation stage. The goal of this stage is to explore and adjust the potential depth interval in more detail based on the coarse positioning. The regularized depth prediction output by the coarse depth estimation stage is used as input for further refinement. Similarly, each depth slice continues to be processed through modified 2D TransConv and ConvLSTM modules, which now process finer depth slices, with fewer depth candidates and narrower intervals. The output of this stage is a more accurate depth candidate, ready to enter the final refinement stage.

[0103] The task of the detailed depth refinement stage corresponding to the third regularization network is to make the most detailed depth adjustments to ensure that the final depth estimate is as close to the actual value as possible. This stage emphasizes the detail accuracy of the depth and the sensitive capture of depth changes in a small range. The output of the intermediate depth refinement stage continues to be used, but at this time each depth slice is processed more finely, the depth range is further narrowed, and the number of depth candidates is further reduced. The depth slices are processed by further optimized 2DTransConv and ConvLSTM, which are now configured to focus on higher resolution detail information. The final output depth prediction is highly accurate and can be used directly to generate high-quality point cloud data.

[0104] In this embodiment, a recurrent network structure is applied in each regularized network, and a combination of convolutional layers and LSTM units is deployed.

[0105] The input of the regularization network at the N-1th stage can be the matching cost C of the first source image and the first reference image at each depth value at the N-1th stage. N-1 (d1), C N-1 (d2),…,C N-1 (dM).

[0106] The matching cost of the first source image and the first reference image at each depth value at the N-1th stage is processed by a series of convolutional LSTM units, which are designed to process time series data, but are applied here to process depth sequence information. Specifically, the matching cost C of the first source image and the first reference image at each depth value at the N-1th stage is N-1 (d1), C N-1 (d2),…,C N-1 (dM) passes through a series of ConvLSTM units, and after each unit is processed, the output result is passed to the next unit until all depth values ​​of the N-1th stage are processed.

[0107] The internal structure of the ConvLSTM unit includes Hadamard product, Sigmoid and Tanh activation functions, addition, and 2D convolution operations. These components work together to implement LSTM's gated processing of time series data, but here it is for the spatial relationship of deep data. Hadamard product is used for direct multiplication operations between elements, which is usually used to apply gating signals. Sigmoid activation is used to generate gating signals to control the flow of information. Tanh activation produces new candidate values. Addition and 2D convolution are used to merge information and produce new state outputs.

[0108] The processing of the i-th depth value in the N-1th stage depends not only on its own input C N-1 (di) also depends on the output of the processing of the i-1th depth value in the N-1 stage. This design helps the model capture the contextual relationship between depths and improves the continuity and accuracy of depth estimation.

[0109] After the regularized network in the N-1th stage processes all the depth layers, each layer outputs a probability distribution Y N-1 (d1), Y N-1 (d2),…,Y N-1 (dM). These probability distributions represent the probability of each pixel of the first source image at each depth value of the j-1th resolution at the N-1th stage, that is, the probability distribution of the first source image at each depth value of the N-1th stage.

[0110] Step S23122: Obtain a predicted depth map of the first source image at the N-1th stage according to the probability distribution of each depth value of the first source image at the N-1th stage.

[0111] The final predicted depth map is estimated by the regularized probability volume using a Winner-Take-All strategy.

[0112] Specifically, the shape of the probability distribution of the regularized network output at stage N-1 is H N-1 *W N-1 *M, traverse H N-1 *W N-1 pixel positions, corresponding to the output Y for each depth value N-1 (di) Converted to a probability distribution by applying the Softmax function. This step converts the cost of each pixel at the 1st to Mth depth layers into a probability, indicating the likelihood that the pixel is located at that depth.

[0113] The Winner-Take-All strategy selects the depth value with the highest probability from the probability volume as the final depth estimate.

[0114] Step S2313: Determine a depth value at the Nth stage according to the predicted depth map of the first source image at the N-1th stage.

[0115] In the predicted depth map of the first source image at the N-1th stage, the depth value of the i-th pixel is D. Then, the corresponding depth value of the Nth stage may include M values ​​belonging to [D-△D, D+△D], and the difference between two adjacent values ​​among the M values ​​is △D*2 / (M-1).

[0116] Step S2320, obtaining a matching cost at each depth value of the Nth stage according to the depth value of the Nth stage, the first feature map of the first image corresponding to the Nth stage, and the first camera parameter.

[0117] In one embodiment of the present disclosure, obtaining the matching cost at each depth value of the Nth stage according to the depth value of the Nth stage, the first feature map of the first image corresponding to the Nth stage, and the first camera parameter may include the following steps S2321-S2322:

[0118] Step S2321: for each depth value of the Nth stage, according to the first camera parameter, project each pixel point corresponding to the Nth stage in the first source image onto the first reference image according to the corresponding depth value to obtain the corresponding projected pixel point.

[0119] Step S2322, comparing the difference in the first depth feature between each pixel of the first source image at each depth value in the Nth stage and the corresponding projection pixel, to obtain a matching cost at each depth value in the Nth stage.

[0120] In this embodiment, the statistical value of the first depth feature of the projected pixel point corresponding to the k-th pixel point of the first source image at the ith depth value in the Nth stage is first determined, and then the difference between the statistical value and the first depth feature of the k-th pixel point is determined, that is, the first depth feature difference between the k-th pixel point and the corresponding projected pixel point at the ith depth value of the first source image in the Nth stage.

[0121] The first depth feature difference between all pixels of the first source image at each depth value of the Nth stage and the corresponding projected pixels is the matching cost between the first source image and the first reference image at each depth value of the Nth stage.

[0122] The matching cost at each depth value in the Nth stage can be expressed as H N *W N A tensor of N *W N That is the number of pixels in the Nth stage.

[0123] In another embodiment of the present disclosure, when N is 1, determining the matching cost of a first source image in a first multi-view image and a first reference image in the first multi-view image at each depth value of the Nth stage according to a first feature map and a first camera parameter corresponding to the Nth stage of the first image may include: obtaining the depth value of the Nth stage, and obtaining the matching cost at each depth value of the Nth stage according to the depth value of the Nth stage, the first feature map and the first camera parameter corresponding to the Nth stage of the first image.

[0124] Each depth value of the Nth stage may be a set of depth values ​​set according to an application scenario or specific requirements, for example, may be an integer value of 1-M.

[0125] Step S2400: Based on the regularization network of the Nth stage, regularization processing is performed on the matching cost at each depth value of the Nth stage to obtain a first depth map of the first source image.

[0126] In one embodiment of the present disclosure, based on the regularization network of the Nth stage, regularizing the matching cost at each depth value of the Nth stage to obtain the first depth map of the first source image may include the following steps S2410 to S2420:

[0127] Step S2410: Based on the regularization network of the Nth stage, regularization processing is performed on the matching cost at each depth value of the Nth stage to obtain the probability distribution of the first source image at each depth value of the Nth stage.

[0128] In this embodiment, N cascaded regularization networks can be pre-set as a hybrid regularization network, which combines the cyclic regularization method of the recurrent neural network (RNN) and the multi-stage cascade structure from coarse to fine. In each regularization stage, the input matching cost is sliced ​​along the depth dimension, and each slice is fed into a 2D TransConv module in turn. The output of each layer of 2D TransConv is connected to a ConvLSTM module, so that the network can effectively process sequence information and maintain spatial continuity. As the depth of the network gradually increases, the number and interval of depth candidates gradually decrease, thereby achieving a gradually refined depth estimation. This is achieved by gradually refining each stage in the cascade structure, allowing the network to optimize and adjust the depth estimation at a higher level. This design not only uses spatial constraints to reduce noise and enhance the fragmented smoothness of the depth map, but also ensures the optimization of calculation efficiency and video memory usage.

[0129] In the embodiment where N is 3, three cascaded regularization networks correspond to three stages from coarse to fine, including a coarse depth estimation stage, an intermediate depth refinement stage, and a detailed depth refinement stage, and each stage focuses on further refining the depth estimation through loop and cascade regularization methods.

[0130] The main task of the rough depth estimation stage corresponding to the first regularization network is to quickly locate possible depth candidate intervals starting from a coarse depth range. This stage does not need to be very accurate, but it should be able to determine the approximate depth distribution area to provide a basis for subsequent stages.

[0131] Specifically, the matching cost can be processed by using 2D TransConv and the connected ConvLSTM module, where 2D TransConv helps capture spatial features, while ConvLSTM processes sequence information in the depth dimension and provides temporal memory function for depth slices. At this stage, the number of depth candidates is relatively large and the depth range is also wide, so the granularity of depth estimation is coarse.

[0132] The intermediate depth refinement stage corresponding to the second regularization network further narrows the depth range and increases the accuracy of the depth estimate based on the coarse depth estimation stage. The goal of this stage is to explore and adjust the potential depth interval in more detail based on the coarse positioning. The regularized depth prediction output by the coarse depth estimation stage is used as input for further refinement. Similarly, each depth slice continues to be processed through modified 2D TransConv and ConvLSTM modules, which now process finer depth slices, with fewer depth candidates and narrower intervals. The output of this stage is a more accurate depth candidate, ready to enter the final refinement stage.

[0133] The task of the detailed depth refinement stage corresponding to the third regularization network is to make the most detailed depth adjustments to ensure that the final depth estimate is as close to the actual value as possible. This stage emphasizes the detail accuracy of the depth and the sensitive capture of depth changes in a small range. The output of the intermediate depth refinement stage continues to be used, but at this time each depth slice is processed more finely, the depth range is further narrowed, and the number of depth candidates is further reduced. The depth slices are processed by further optimized 2DTransConv and ConvLSTM, which are now configured to focus on higher resolution detail information. The final output depth prediction is highly accurate and can be used directly to generate high-quality point cloud data.

[0134] In this embodiment, a recurrent network structure is applied in each regularized network, and a combination of convolutional layers and LSTM units is deployed.

[0135] The input of the regularization network of the Nth stage can be the matching cost C of the first source image and the first reference image at each depth value of the Nth stage. N (d1), C N (d2),…,C N (dM).

[0136] The matching cost of the first source image and the first reference image at each depth value in the Nth stage is processed by a series of convolutional LSTM units, which are designed to process time series data, but are applied here to process depth sequence information. Specifically, the matching cost C of the first source image and the first reference image at each depth value in the Nth stage is N (d1), C N (d2),…,C N (dM) passes through a series of ConvLSTM units, and after each unit is processed, the output result is passed to the next unit until all depth values ​​of the Nth stage are processed.

[0137] The internal structure of the ConvLSTM unit includes Hadamard product, Sigmoid and Tanh activation functions, addition, and 2D convolution operations. These components work together to implement LSTM's gated processing of time series data, but here it is for the spatial relationship of deep data. Hadamard product is used for direct multiplication operations between elements, which is usually used to apply gating signals. Sigmoid activation is used to generate gating signals to control the flow of information. Tanh activation produces new candidate values. Addition and 2D convolution are used to merge information and produce new state outputs.

[0138] The processing of the i-th depth value in the N-th stage depends not only on its own input C N (di) also depends on the output of the processing of the i-1th depth value in the N stages. This design helps the model capture the contextual relationship between depths and improves the continuity and accuracy of depth estimation.

[0139] After the regularized network in the Nth stage processes all the depth layers, each layer outputs a probability distribution Y N (d1), Y N (d2),…,Y N (dM). These probability distributions represent the probability of each pixel of the first source image corresponding to the j-1th resolution at each depth value at the Nth stage, that is, the probability distribution of the first source image at each depth value at the Nth stage.

[0140] Step S2420: According to the probability distribution of each depth value of the first source image at the Nth stage, a predicted depth map of the first source image at the Nth stage is obtained as a first depth map.

[0141] The final predicted depth map is estimated by the regularized probability volume using a Winner-Take-All strategy.

[0142] Specifically, the shape of the probability distribution of the regularized network output at the Nth stage is H N *W N *M, traverse H N *W N pixel positions, corresponding to the output Y for each depth value N (di) Converted to a probability distribution by applying the Softmax function. This step converts the cost of each pixel at the 1st to Mth depth layers into a probability, indicating the likelihood that the pixel is located at that depth.

[0143] Step S2500: reconstructing a three-dimensional point cloud of a first scene according to a first source image and a first depth map.

[0144] In this embodiment, when a first depth map of at least one first image is obtained, a three-dimensional point cloud of the first scene can be reconstructed based on the first image and the corresponding first depth map; or when first depth maps corresponding to all first images in the first multi-view image are obtained, a three-dimensional point cloud of the first scene can be reconstructed based on all first images and the corresponding first depth maps.

[0145] In this embodiment, for each pixel point of the first image corresponding to the first depth map, the pixel point can be converted from the image coordinate system to the three-dimensional world coordinate system according to its depth value and the first camera parameter to obtain the corresponding three-dimensional world coordinates. All three-dimensional world coordinates are aggregated to obtain a complete three-dimensional point cloud.

[0146] In one embodiment of the present disclosure, before reconstructing the three-dimensional point cloud of the first scene according to the first source image and the first depth map, point cloud construction parameters may be set first, including the accuracy and range of the point cloud to be generated.

[0147] In one embodiment of the present disclosure, after reconstructing the three-dimensional point cloud of the first scene based on the first multi-view image and the first depth map, the method may further include: filtering the three-dimensional point cloud; and / or optimizing the three-dimensional point cloud.

[0148] Filtering can be to apply filtering technology to remove noise in the 3D point cloud and improve the quality of the 3D point cloud. Optimization can be to optimize the point cloud data through methods such as point cloud pruning and downsampling to make it more suitable for subsequent applications.

[0149] In one embodiment of the present disclosure, the finally obtained point cloud data may be saved in a common format, such as PLY or OBJ, to facilitate subsequent use and analysis.

[0150] In one embodiment of the present disclosure, a three-dimensional visualization tool may be used to display the reconstructed point cloud to intuitively show the reconstruction effect.

[0151] Through the embodiments of the present disclosure, first multi-view images and corresponding first camera parameters associated with a first scene are obtained, and each first image corresponds to a first feature map of N stages; according to the first feature map and the first camera parameters corresponding to the first image of N stages, the matching cost of a first source image in the first multi-view image and a first reference image in the first multi-view image at each depth value of the N stage is determined; based on a regularization network of the Nth stage, the matching cost at each depth value of the Nth stage is regularized to obtain a first depth map of the first source image; and three-dimensional point cloud reconstruction of the first scene is performed according to the first source image and the first depth map, which can reduce the occupancy of video memory and the computing cost, and improve the speed of three-dimensional point cloud reconstruction.

[0152] In one embodiment of the present disclosure, the method further includes the step of training a regularized network of N stages, including steps S3100 to S3500 as shown below:

[0153] Step S3100: Acquire a second multi-view image associated with a second scene and corresponding second camera parameters.

[0154] The second multi-view image is a second image obtained by photographing the second scene from multiple perspectives using a plurality of second cameras in different positions, and the second camera parameters include position information of each second camera.

[0155] In this embodiment, the second camera parameters may be the same as the first camera parameters, or may be different from the first camera parameters, which will not be described in detail herein.

[0156] Step S3200: extracting second feature maps corresponding to N stages of the second image, where the second feature maps include second depth features of each pixel point in the corresponding stage.

[0157] For details, please refer to the method of extracting the first feature map in the aforementioned step S2200, which will not be elaborated here.

[0158] Step S3300: Determine the matching cost between the second source image in the second multi-view image and the second reference image in the second multi-view image at each depth value in the N stages according to the second feature map and the second camera parameter corresponding to the second image in the N stages.

[0159] For details, please refer to the method of determining the matching cost in the aforementioned step S2300, which will not be elaborated here.

[0160] Step S3400, obtaining real depth maps of the second source image corresponding to N stages.

[0161] Step S3500, training the regularized network of N stages according to the matching cost at each depth value of N stages, taking minimizing the loss between the second depth map output by the regularized network of N stages and the corresponding true depth map as the training goal, and updating the network parameters of the regularized network of N stages.

[0162] In one embodiment of the present disclosure, the regularized network of N stages is trained according to the matching cost at each depth value of the N stages, including steps S3510 to S3540 as shown below:

[0163] Step S3510, taking the network parameters of the regularized network in N stages as variables, and determining the second depth map of the second source image in N stages according to the matching cost at each depth value in the N stages.

[0164] Step S3520: For the regularized network of each stage, a first loss function of the regularized network of the corresponding stage is constructed according to the second depth map and the true depth map of the corresponding stage.

[0165] In this embodiment, a classification-based cross entropy loss function may be used to construct a first loss function of a regularized network corresponding to N stages, wherein each depth value of a stage is regarded as a predefined class.

[0166] Step S3530, according to the weights corresponding to the N stages, weighted sum is performed on the first loss functions of each stage to obtain a second loss function.

[0167] Step S3540, solving the network parameters of the regularized network of N stages when the value of the second loss function is minimum, and completing the training of the regularized network of N stages.

[0168] In this embodiment, the loss of the regularized network at each stage is weighted according to its contribution to the final depth estimation, so as to achieve the optimal training effect of the regularized network at N stages.

[0169] In this embodiment, the back propagation method may be used to optimize the network parameters of the regularized network in N stages.

[0170] <Example>

[0171] Figure 3 The flowchart is an example of a three-dimensional point cloud reconstruction method according to an embodiment of the present disclosure.

[0172] exist Figure 3In the figure, a process of determining a first depth map of a first source image based on three cascaded regularization networks is shown, which specifically includes: obtaining a preset first stage depth value M1, and obtaining matching costs C1(d1), C1(d2), ..., C1(dM) at each depth value of the first stage according to the depth value of the first stage, the first feature map 11 of the first source image corresponding to the first stage, the first feature map 12 of the first reference image corresponding to the first stage, and the first camera parameters. Based on the regularization network 1 of the first stage, the matching costs C1(d1), C1(d2), ..., C1(dM) at each depth value of the first stage are regularized to obtain the predicted depth map Y1 of the first source image at the first stage, and the depth value M2 of the second stage is determined according to the predicted depth map of the first source image at the first stage. According to the depth value M2 of the second stage, the first feature map 21 of the first source image corresponding to the second stage, the first feature map 22 of the first reference image corresponding to the second stage, and the first camera parameters, the matching costs C2(d1), C2(d2), ..., C2(dM) at each depth value of the second stage are obtained. Based on the regularization network 2 of the second stage, the matching costs C2(d1), C2(d2), ..., C2(dM) at each depth value of the second stage are regularized to obtain the predicted depth map Y2 of the first source image at the second stage, and the depth value M3 of the third stage is determined according to the predicted depth map of the first source image at the second stage. According to the depth value M3 of the third stage, the first feature map 31 of the first source image corresponding to the third stage, the first feature map 32 of the first reference image corresponding to the third stage, and the first camera parameters, the matching costs C3(d1), C3(d2), ..., C3(dM) at each depth value of the third stage are obtained. Based on the regularization network 3 of the third stage, the matching costs C3(d1), C3(d2), ..., C3(dM) at each depth value in the third stage are regularized to obtain the predicted depth map Y3 of the first source image in the third stage as the first depth map of the first source image.

[0173] <Device Example>

[0174] The present disclosure provides a three-dimensional point cloud reconstruction device, such as Figure 4 As shown, the three-dimensional point cloud reconstruction device 4000 includes a first acquisition module 4100 , a feature extraction module 4200 , a cost determination module 4300 , a depth estimation module 4400 and a point cloud reconstruction module 4500 .

[0175] The first acquisition module 4100 is used to acquire a first multi-view image associated with a first scene and corresponding first camera parameters, wherein the first multi-view image is a first image obtained by photographing the first scene from multiple perspectives using a plurality of first cameras in different postures, and the first camera parameters include posture information of each first camera.

[0176] The feature extraction module 4200 is used to extract the first feature map corresponding to N stages of each first image, and the first feature map includes the first depth feature of each pixel point of the corresponding stage; N is a positive integer.

[0177] The cost determination module 4300 is used to determine the matching cost of the first source image in the first multi-view image and the first reference image in the first multi-view image at each depth value in the Nth stage based on the first feature map of the Nth stage corresponding to the first image and the first camera parameters.

[0178] The depth estimation module 4400 obtains a regularized network based on the Nth stage, performs regularization processing on the matching cost at each depth value of the Nth stage, and obtains a first depth map of the first source image.

[0179] The point cloud reconstruction module 4500 is used to reconstruct a three-dimensional point cloud of the first scene according to the first source image and the first depth map.

[0180] In one embodiment of the present disclosure, the 3D point cloud reconstruction device 4000 further includes:

[0181] a traversal module, configured to traverse a plurality of first images in the first multi-view images, and use a currently traversed first image as a first source image, and use at least one first image in the first multi-view images other than the first source image as a first reference image;

[0182] The depth estimation module 4400 is used to determine a first depth map of the first source image;

[0183] A module for determining whether there is an untraversed first image in the first multi-view image; if so, the traversal module continues to execute the step of traversing multiple first images in the first multi-view image; if not, the traversal module ends the traversal and obtains a first depth map corresponding to each first image in the first multi-view image.

[0184] In one embodiment of the present disclosure, the cost determination module 4300 is used to:

[0185] Determine a depth value of the Nth stage according to the first feature map of the first N-1 stages corresponding to the first image and the first camera parameter;

[0186] A matching cost at each depth value of the Nth stage is obtained according to the depth value of the Nth stage, the first feature map of the first image corresponding to the Nth stage, and the first camera parameter.

[0187] In one embodiment of the present disclosure, determining the depth value of the Nth stage according to the first feature map of the first N-1 stages corresponding to the first image and the first camera parameter includes:

[0188] Determine, according to the first feature map of the first N-1 stages corresponding to the first image and the first camera parameter, a matching cost between the first source image and the first reference image at each depth value of the N-1 stage;

[0189] Based on the regularization network of the N-1th stage, regularize the matching cost at each depth value of the N-1th stage to obtain a predicted depth map of the first source image at the N-1th stage;

[0190] Determine a depth value at the Nth stage according to the predicted depth map of the first source image at the N-1th stage.

[0191] In one embodiment of the present disclosure, obtaining the matching cost at each depth value of the Nth stage according to the depth value of the Nth stage, the first feature map of the first image corresponding to the Nth stage, and the first camera parameter includes:

[0192] For each depth value of the Nth stage, project each pixel point corresponding to the Nth stage in the first source image onto the first reference image according to the corresponding depth value according to the first camera parameter to obtain a corresponding projected pixel point;

[0193] The difference in the first depth feature between each pixel point of the first source image at each depth value of the Nth stage and the corresponding projection pixel point is compared to obtain a matching cost at each depth value of the Nth stage.

[0194] In one embodiment of the present disclosure, the depth estimation module 4400 is used to:

[0195] Based on the regularization network of the Nth stage, regularize the matching cost at each depth value of the Nth stage to obtain the probability distribution of the first source image at each depth value of the Nth stage;

[0196] According to the probability distribution of each depth value of the first source image at the Nth stage, a predicted depth map of the first source image at the Nth stage is obtained as the first depth map.

[0197] In one embodiment of the present disclosure, the 3D point cloud reconstruction device 4000 further includes a training module for:

[0198] Acquire a second multi-view image associated with a second scene and corresponding second camera parameters, where the second multi-view image is a second image obtained by photographing the second scene from multiple perspectives using a plurality of second cameras in different postures, and the second camera parameters include posture information of each second camera;

[0199] Extracting a second feature map of the second image corresponding to N stages, where the second feature map includes a second depth feature of each pixel point in the corresponding stage;

[0200] Determine, according to the second feature map of the N stages corresponding to the second image and the second camera parameter, a matching cost between a second source image in the second multi-view image and a second reference image in the second multi-view image at each depth value of the N stages;

[0201] Obtaining true depth maps of N stages corresponding to the second source image;

[0202] The N-stage regularized network is trained according to the matching cost at each depth value of the N stages, and the network parameters of the N-stage regularized network are updated by minimizing the loss between the second depth map output by the N-stage regularized network and the corresponding true depth map as the training goal.

[0203] In one embodiment of the present disclosure, the training of the regularized network of N stages according to the matching cost at each depth value of the N stages includes:

[0204] Taking the network parameters of the regularized network at the N stages as variables, and determining the second depth map of the second source image at the N stages according to the matching cost at each depth value at the N stages;

[0205] For the regularized network of each stage, constructing the first loss function of the regularized network of the corresponding stage according to the second depth map and the true depth map of the corresponding stage;

[0206] According to the weights corresponding to the N stages, the first loss function of each stage is weighted summed to obtain the second loss function;

[0207] Solve the network parameters of the regularized network of N stages when the value of the second loss function is minimum, and complete the training of the regularized network of N stages.

[0208] <Electronic Equipment Embodiment>

[0209] This embodiment provides an electronic device. In one aspect, the electronic device may include the aforementioned three-dimensional point cloud reconstruction device 4000.

[0210] On the other hand, Figure 5 As shown, the electronic device 5000 may include a processor 5100 and a memory 5200, the memory 5200 is used to store a computer program, and the processor 5100 is used to control the electronic device to execute the method of any embodiment of the present disclosure under the control of the computer program.

[0211] <Readable Storage Medium Embodiment>

[0212] This embodiment provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method described in any method embodiment of the present disclosure is executed.

[0213] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0214] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.

[0215] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0216] The computer program instructions for performing the operation of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present invention.

[0217] Various aspects of the present invention are described herein with reference to the flow charts and / or block diagrams of the methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each box of the flow chart and / or block diagram and the combination of each box in the flow chart and / or block diagram can be implemented by computer-readable program instructions.

[0218] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0219] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0220] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a part of a module, a program segment or an instruction, and a part of the module, a program segment or an instruction contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that it is equivalent to implement it by hardware, implement it by software, and implement it by combining software and hardware.

[0221] Embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the marketplace, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.

Claims

1. A three-dimensional point cloud reconstruction method, characterized in that: include: Acquire a first multi-view image associated with a first scene and corresponding first camera parameters, wherein the first multi-view image is a first image obtained by photographing the first scene from multiple perspectives using a plurality of first cameras with different postures, and the first camera parameters include posture information of each first camera; Extracting a first feature map of N stages corresponding to each first image, wherein the first feature map includes a first depth feature of each pixel point of the corresponding stage; N is a positive integer; Determine, according to the first feature map of the first image corresponding to the N stages and the first camera parameter, a matching cost between a first source image in the first multi-view image and a first reference image in the first multi-view image at each depth value in the N stages; Based on the regularization network of the Nth stage, regularize the matching cost at each depth value of the Nth stage to obtain a first depth map of the first source image; A three-dimensional point cloud is reconstructed for the first scene according to the first source image and the first depth map.

2. The method according to claim 1, characterized in that The method further comprises: traversing a plurality of first images in the first multi-view images, and taking a currently traversed first image as a first source image, and taking at least one first image in the first multi-view images other than the first source image as a first reference image; determining a first depth map of the first source image; Determine whether there is an untraversed first image in the first multi-view image. If so, continue to perform the step of traversing multiple first images in the first multi-view image; if not, end the traversal and obtain a first depth map corresponding to each first image in the first multi-view image.

3. The method according to claim 1, characterized in that The determining, according to the first feature map of the first image corresponding to the N stages and the first camera parameter, a matching cost of a first source image in the first multi-view image and a first reference image in the first multi-view image at each depth value in the N stages, comprises: Determine a depth value of the Nth stage according to the first feature map of the first N-1 stages corresponding to the first image and the first camera parameter; A matching cost at each depth value of the Nth stage is obtained according to the depth value of the Nth stage, the first feature map of the first image corresponding to the Nth stage, and the first camera parameter.

4. The method according to claim 3, characterized in that The determining the depth value of the Nth stage according to the first feature map of the first N-1 stages corresponding to the first image and the first camera parameter includes: Determine, according to the first feature map of the first N-1 stages corresponding to the first image and the first camera parameter, a matching cost between the first source image and the first reference image at each depth value of the N-1 stage; Based on the regularization network of the N-1th stage, regularize the matching cost at each depth value of the N-1th stage to obtain a predicted depth map of the first source image at the N-1th stage; Determine a depth value at the Nth stage according to the predicted depth map of the first source image at the N-1th stage.

5. The method according to claim 3, characterized in that: The obtaining, according to the depth value of the Nth stage, the first feature map of the Nth stage corresponding to the first image, and the first camera parameter, a matching cost at each depth value of the Nth stage includes: For each depth value of the Nth stage, project each pixel point corresponding to the Nth stage in the first source image onto the first reference image according to the corresponding depth value according to the first camera parameter to obtain a corresponding projected pixel point; The difference in the first depth feature between each pixel point of the first source image at each depth value of the Nth stage and the corresponding projection pixel point is compared to obtain a matching cost at each depth value of the Nth stage.

6. The method according to claim 1, characterized in that The regularization network based on the Nth stage performs regularization processing on the matching cost at each depth value of the Nth stage to obtain a first depth map of the first source image, including: Based on the regularization network of the Nth stage, regularize the matching cost at each depth value of the Nth stage to obtain the probability distribution of the first source image at each depth value of the Nth stage; According to the probability distribution of each depth value of the first source image at the Nth stage, a predicted depth map of the first source image at the Nth stage is obtained as the first depth map.

7. The method according to claim 1, characterized in that The method further comprises the step of training a regularized network of N stages, comprising: Acquire a second multi-view image associated with a second scene and corresponding second camera parameters, where the second multi-view image is a second image obtained by photographing the second scene from multiple perspectives using a plurality of second cameras in different postures, and the second camera parameters include posture information of each second camera; Extracting a second feature map of the second image corresponding to N stages, where the second feature map includes a second depth feature of each pixel point in the corresponding stage; Determine, according to the second feature map of the N stages corresponding to the second image and the second camera parameter, a matching cost between a second source image in the second multi-view image and a second reference image in the second multi-view image at each depth value of the N stages; Obtaining true depth maps of N stages corresponding to the second source image; The N-stage regularized network is trained according to the matching cost at each depth value of the N stages, and the network parameters of the N-stage regularized network are updated by minimizing the loss between the second depth map output by the N-stage regularized network and the corresponding true depth map as the training goal.

8. The method according to claim 7, characterized in that The training of the regularized network of the N stages according to the matching cost at each depth value of the N stages includes: Taking the network parameters of the regularized network at the N stages as variables, and determining the second depth map of the second source image at the N stages according to the matching cost at each depth value at the N stages; For the regularized network of each stage, constructing the first loss function of the regularized network of the corresponding stage according to the second depth map and the true depth map of the corresponding stage; According to the weights corresponding to the N stages, the first loss function of each stage is weighted summed to obtain the second loss function; Solve the network parameters of the regularized network of N stages when the value of the second loss function is minimum, and complete the training of the regularized network of N stages.

9. A three-dimensional point cloud reconstruction device, characterized in that: include: A first acquisition module is used to acquire a first multi-view image associated with a first scene and corresponding first camera parameters, wherein the first multi-view image is a first image obtained by shooting the first scene from multiple perspectives using a plurality of first cameras with different postures, and the first camera parameters include posture information of each first camera; A feature extraction module, used to extract a first feature map corresponding to N stages of each first image, wherein the first feature map includes a first depth feature of each pixel point of the corresponding stage; N is a positive integer; a cost determination module, configured to determine, according to the first feature map of the first image corresponding to the N stages and the first camera parameter, a matching cost between a first source image in the first multi-view image and a first reference image in the first multi-view image at each depth value in the N stages; A depth estimation module obtains a regularized network based on the Nth stage, performs regularization processing on the matching cost at each depth value of the Nth stage, and obtains a first depth map of the first source image; A point cloud reconstruction module is used to reconstruct a three-dimensional point cloud of the first scene based on the first source image and the first depth map.

10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the method according to any one of claims 1 to 8 under the control of the computer program.

Citation Information

Patent Citations

  • Multi-view three-dimensional network three-dimensional reconstruction method and system

    CN113066168A

  • Multi-view three-dimensional reconstruction method for high-resolution image

    CN116071504A

  • Three-dimensional reconstruction method and system based on improved MVSNet

    CN116912405A

  • Multi-view stereo matching reconstruction method

    CN117132712A

  • Multi-view three-dimensional reconstruction method based on error perception, medium and equipment

    CN118411465A

Cited By

  • Semantic segmentation method, electronic equipment and storage medium

    CN120876845A