Implicit field reconstruction method and system based on view matching and epipolar geometry constraints
By optimizing the camera pose and neural implicit field through the method of perspective matching and epipolar geometry constraints, the problems of blurred rendering results and insufficient geometric details are solved, and high-quality 3D reconstruction and rendering are achieved.
Patent Information
- Application Number
- CN202411613020.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-13
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-11-13
AI Technical Summary
Existing technologies make it difficult to effectively optimize camera pose and neural implicit field expression simultaneously, resulting in blurred rendering results and inability to accurately express the geometric details of objects.
Through a method based on view matching and epipolar geometry constraints, the initial image group is obtained and the initial three-dimensional implicit field is constructed. The sampling point features of the camera ray are extracted, the target key features are generated, the camera pose and implicit field are optimized using the geometric network and texture network, and the rendering value of the matching ray is output.
The accuracy of pose optimization results is improved, the geometric detail expression of the three-dimensional implicit field is enhanced, and high-quality rendering effects are achieved.
Smart Images

Figure CN119152123B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of scene rendering technology, and in particular to an implicit field reconstruction method, system, terminal and computer-readable storage medium based on view matching and epipolar geometry constraints. Background Art
[0002] A camera ray is a ray of light emitted from the camera position and passing through the corresponding image plane pixel, which is used to determine which parts of the scene should be displayed in the image. In the neural implicit field rendering method, the sampling points on the camera ray play a crucial role.
[0003] However, the reconstruction of the neural implicit field depends on an accurate pose input, and the algorithms designed for inaccurate poses find it difficult to effectively optimize both the camera pose and the neural implicit field expression. The resulting new perspective rendering results lack details, and the three-dimensional mesh extracted from the density field is usually overly smooth and cannot well reflect the geometric details of the target object.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The main purpose of the present invention is to provide an implicit field reconstruction method, system, terminal and computer-readable storage medium based on view matching and epipolar geometry constraints, aiming to solve the problem in the prior art that it is difficult to effectively optimize the camera pose and neural implicit field expression at the same time, resulting in blurred rendering results and inaccurate expression of the geometric details of the object.
[0006] To achieve the above object, the present invention provides an implicit field reconstruction method based on view matching and epipolar geometry constraints, the implicit field reconstruction method based on view matching and epipolar geometry constraints comprising the following steps:
[0007] Acquire an initial image group of the target object and construct an initial three-dimensional implicit field, input the photometric error between the initial image group and the multi-view rendering value of the target object into the initial three-dimensional implicit field for training to obtain a three-dimensional implicit field;
[0008] Extracting sampling point features of multiple camera rays, generating target key features based on the multiple sampling point features, inputting the target key features into the three-dimensional implicit field, inputting the initial image group into the three-dimensional implicit field, and outputting a target feature volume;
[0009] Inputting the target feature body into a geometric network, outputting a signed distance field value and an output feature, and obtaining a geometric result of the target object according to the signed distance field value and the output feature;
[0010] Normal information of the target object at each sampling point is obtained according to the geometric result, all the normal information is input into a texture network, and rendering values of matching rays among the multiple camera rays are output.
[0011] Optionally, the implicit field reconstruction method based on view matching and epipolar geometry constraints, wherein the steps of obtaining an initial image group of the target object and constructing an initial 3D implicit field, and inputting the photometric error between the initial image group and the multi-view rendering values of the target object into the initial 3D implicit field for training to obtain the 3D implicit field, specifically include:
[0012] Acquire an initial image group of the target object captured by a target camera, wherein the initial image group is generated by photographing the target object from multiple perspectives by the target camera;
[0013] Construct an initial three-dimensional implicit field and calculate the photometric error between the initial image group and the multi-view rendering value of the target object:
[0014] ;
[0015] ;
[0016] ;
[0017] ;
[0018] in, Indicates sampling point The photometric error, represents the sampling point, represents the image index in the initial image group, Represents the pixels of the image, Indicates the image and target object at pixel points The color error at Indicates the Image at pixel point The RGB value at Indicates that the target object is at the pixel point The first The rendering value of the camera perspective corresponding to the image, Indicates the The camera pose of each image, Indicates the The rotation matrix of the target camera corresponding to the image, Indicates the The displacement vector of the target camera corresponding to the image, represents the three-dimensional rotation group, Represents three-dimensional rational number space;
[0019] The photometric error is input into the initial three-dimensional implicit field for training to obtain a three-dimensional implicit field.
[0020] Optionally, in the implicit field reconstruction method based on view matching and epipolar geometry constraints, the camera rays include: key rays and auxiliary rays;
[0021] The step of extracting sampling point features of multiple camera rays, generating target key features based on the multiple sampling point features, inputting the target key features into the three-dimensional implicit field, inputting the initial image group into the three-dimensional implicit field, and outputting a target feature volume may also include:
[0022] Acquire a plurality of the sampling points through which the plurality of key rays pass, and acquire color information and feature information of each of the sampling points;
[0023] Extract multiple color information and multiple feature information with consistency, select corresponding sampling points, match the key rays corresponding to the selected multiple sampling points, and obtain matching rays:
[0024] ;
[0025] ;
[0026] in, represents the matching ray, represents the camera center of the target camera, represents the number of rays, Represents matching rays The normalized direction of .
[0027] Optionally, the implicit field reconstruction method based on view matching and epipolar geometry constraints, wherein the steps of extracting sampling point features of multiple camera rays, generating target key features based on the multiple sampling point features, inputting the target key features into the three-dimensional implicit field, inputting the initial image group into the three-dimensional implicit field, and outputting the target feature volume, specifically include:
[0028] Extracting key features of key sampling points generated by the matching ray passing through the initial image group, and extracting auxiliary features of multiple auxiliary sampling points generated by multiple auxiliary rays passing through the initial image group:
[0029] ;
[0030] in, Indicates sampling point The characteristics of the place, represents the sampling point, represents the progressive feature mask, represents the initial feature body;
[0031] The key feature is enhanced by utilizing all the auxiliary features to obtain a target key feature, and the target key feature is input into the three-dimensional implicit field to output a target feature body.
[0032] Optionally, the implicit field reconstruction method based on view matching and epipolar geometry constraints, wherein the step of enhancing the key features using the auxiliary features to obtain target key features, and inputting the target key features into the three-dimensional implicit field to output a target feature volume, specifically includes:
[0033] The key feature is combined with all the auxiliary features to obtain the target key feature:
[0034] ;
[0035] in, Indicates key sampling points The key characteristics of the target, Indicates key sampling points The number of auxiliary rays, represents a mathematical function that converts an arbitrary set of real numbers into real numbers representing a probability distribution. represents the key sampling point, represents the auxiliary sampling point, Indicates key sampling points The key features of Indicates auxiliary sampling points Auxiliary features at
[0036] The target key features are input into the three-dimensional implicit field, the camera pose in the three-dimensional implicit field is optimized according to the target key features, and the target key features are output.
[0037] Optionally, the implicit field reconstruction method based on view matching and epipolar geometry constraints, wherein the target feature body is input into a geometric network, a signed distance field value and an output feature are output, and a geometric result of the target object is obtained according to the signed distance field value and the output feature, specifically includes:
[0038] Inputting the target key features of the target feature body into the constructed geometric network, and outputting the signed distance field value and the output features;
[0039] Extracting the geometric result of the target object in three-dimensional space from the signed distance field value according to the output feature:
[0040] ;
[0041] in, Indicates key sampling points The signed distance field value of Indicates key sampling points The output features of Represents a geometric network.
[0042] Optionally, the implicit field reconstruction method based on view matching and epipolar geometry constraints, wherein the normal information of the target object at each sampling point is obtained according to the geometric result, all the normal information is input into a texture network, and rendering values of matching rays among the multiple camera rays are output, specifically includes:
[0043] Calculating the normal of the target object at the sampling point according to the gradient value of the geometric result, and obtaining normal information at the sampling point according to the normal;
[0044] Input the normal information into the constructed texture network and output the RGB value of the sampling point:
[0045] ;
[0046] in, Indicates sampling point RGB values, represents the texture network, Indicates sampling point The output features of Represents matching rays The normalized direction of Indicates sampling point Normal information;
[0047] The color of the matching ray is rendered using the opacity density and RGB values of all sampling points to obtain the rendering value of the matching ray:
[0048] ;
[0049] in, Represents matching rays The rendering value of Indicates sampling point The transmittance at Indicates sampling point opacity.
[0050] In addition, to achieve the above-mentioned object, the present invention further provides an implicit field reconstruction system based on view matching and epipolar geometry constraints, wherein the implicit field reconstruction system based on view matching and epipolar geometry constraints comprises:
[0051] An implicit field optimization module is configured to obtain an initial image set of a target object and construct an initial three-dimensional implicit field, inputting the photometric error between the initial image set and the multi-view rendering values of the target object into the initial three-dimensional implicit field for training to obtain a three-dimensional implicit field;
[0052] a feature volume optimization module, configured to extract sampling point features of multiple camera rays, generate target key features based on the multiple sampling point features, input the target key features into the three-dimensional implicit field, input the initial image group into the three-dimensional implicit field, and output a target feature volume;
[0053] A geometric feature collection module is used to input the target feature body into a geometric network, output a signed distance field value and an output feature, and obtain a geometric result of the target object according to the signed distance field value and the output feature;
[0054] A rendering module is used to obtain normal information of the target object at each sampling point according to the geometric result, input all the normal information into a texture network, and output rendering values of matching rays among the multiple camera rays.
[0055] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and an implicit field reconstruction program based on perspective matching and epipolar geometric constraints stored on the memory and runnable on the processor, wherein the implicit field reconstruction program based on perspective matching and epipolar geometric constraints, when executed by the processor, implements the steps of the implicit field reconstruction method based on perspective matching and epipolar geometric constraints as described above.
[0056] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an implicit field reconstruction program based on perspective matching and epipolar geometric constraints, and when the implicit field reconstruction program based on perspective matching and epipolar geometric constraints is executed by a processor, the steps of the implicit field reconstruction method based on perspective matching and epipolar geometric constraints as described above are implemented.
[0057] In the present invention, an initial image group of the target object is obtained, and an initial three-dimensional implicit field is constructed. The photometric error between the initial image group and the multi-view rendering value of the target object is input into the initial three-dimensional implicit field for training to obtain a three-dimensional implicit field; sampling point features of multiple camera rays are extracted, and target key features are generated based on the multiple sampling point features. The target key features are input into the three-dimensional implicit field, and the initial image group is input into the three-dimensional implicit field to output a target feature body; the target feature body is input into a geometric network, and a signed distance field value and an output feature are output. Based on the signed distance field value and the output feature, a geometric result of the target object is obtained; based on the geometric result, the normal information of the target object at each sampling point is obtained, all the normal information is input into a texture network, and the rendering value of the matching ray among the multiple camera rays is output. The present invention uses a set of images with inaccurate poses as input to reconstruct a three-dimensional implicit field, and jointly optimizes the optimized camera pose and the three-dimensional implicit field. By passing camera rays through key points in the input image, light optimization and matching light consistency are achieved, which can effectively improve the results of pose optimization. The three-dimensional implicit field and camera pose can be further optimized by enhancing point cloud features. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 is a flow chart of a preferred embodiment of the implicit field reconstruction method based on view matching and epipolar geometry constraints of the present invention;
[0059] Figure 2 1 is a flow chart of a preferred embodiment of the implicit field reconstruction method based on view matching and epipolar geometry constraints of the present invention;
[0060] Figure 3 2 is a schematic diagram of feature enhancement of a preferred embodiment of the implicit field reconstruction method based on view matching and epipolar geometry constraints of the present invention;
[0061] Figure 4 1 is a structural diagram of a preferred embodiment of the implicit field reconstruction system based on view matching and epipolar geometry constraints of the present invention;
[0062] Figure 5 Schematic diagram of the operating environment of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION
[0063] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0064] The implicit field reconstruction method based on view matching and epipolar geometry constraints described in the preferred embodiment of the present invention is as follows: Figure 1As shown, the implicit field reconstruction method based on view matching and epipolar geometry constraints includes the following steps:
[0065] Step S10: Obtain an initial image group of the target object and construct an initial 3D implicit field. Input the photometric error between the initial image group and the multi-view rendering value of the target object into the initial 3D implicit field for training to obtain a 3D implicit field.
[0066] The camera rays include key rays and auxiliary rays. The initial image set is a set of images with uncalibrated poses, used as input to capture the 3D target object. The goal of scene reconstruction based on point cloud augmentation is to design a neural network that reconstructs the 3D representation of the target object from a set of images captured from multiple viewpoints, thereby achieving optimized training of the 3D implicit field.
[0067] Specifically, an initial image group of the target object captured by a target camera is obtained, wherein the initial image group is generated by the target camera shooting the target object from multiple perspectives; an initial three-dimensional implicit field is constructed, and a photometric error between the initial image group and the multi-perspective rendering value of the target object is calculated:
[0068] ;
[0069] ;
[0070] ;
[0071] ;
[0072] in, Indicates sampling point The photometric error, represents the sampling point, represents the image index in the initial image group, Represents the pixels of the image, Indicates the image and target object at pixel points The color error at Indicates the Image at pixel point The RGB value at Indicates that the target object is at the pixel point The first The rendering value of the camera perspective corresponding to the image, Indicates the The camera pose of each image, Indicates the The rotation matrix of the target camera corresponding to the image, Indicates the The displacement vector of the target camera corresponding to the image, represents the three-dimensional rotation group, Representing a three-dimensional rational number space; inputting the photometric error into the initial three-dimensional implicit field for training to obtain a three-dimensional implicit field.
[0073] Each image in the initial image set is associated with a camera pose that is assumed to be known or estimated using an algorithm. The initial 3D implicit field is trained by minimizing the photometric error between the input image and the multi-view rendering of the target object, thereby obtaining an optimized 3D implicit field. By optimizing the 3D implicit field, the target object is optimized (i.e., a 3D feature volume that can accurately represent the target object is constructed). By directly optimizing the target, the original opacity or density is changed to a signed distance field, which is itself a way of expressing 3D geometric shapes. This makes the geometric details of the triangular mesh extracted from the reconstructed 3D implicit field finer. By modeling the correspondence between the signed distance field and the density in NeRF (Neural Radiance Fields), the density can be optimized by optimizing the signed distance field, thereby enhancing the rendering results.
[0074] Furthermore, a plurality of sampling points through which the plurality of key rays pass are obtained, and color information and feature information of each sampling point are obtained; a plurality of color information and a plurality of feature information having consistency are extracted, and corresponding sampling points are selected, and the key rays corresponding to the selected plurality of sampling points are matched to obtain matching rays:
[0075] ;
[0076] ;
[0077] in, represents the matching ray, represents the camera center of the target camera, represents the number of rays, Represents matching rays The normalized direction of .
[0078] Among them, Figure 2 As shown, after extracting key points from the input image and matching them using a pre-trained network, the initial 3D implicit field is trained to optimize the 3D feature volume. This feature volume encodes the geometric and photometric information about the target object and can be queried through camera rays to obtain the corresponding color information and feature information for new view synthesis and 3D reconstruction.
[0079] Here, each pixel is actually associated with a specific 3D ray from the object / scene through the pixel center and towards the camera center: ; The color and opacity of the points that can be sampled along the ray The integration is performed to generate the rendering color of the ray. Since the camera pose may be noisy, the reconstructed radiation field or implicit field may not produce a clear rendering with details. In the case of large noise, some methods may not even be able to produce results, so Beyond a neural rendering network that implicitly maps to opacity density (or density) and color, we enhance sample point features by leveraging contextual feature relationships between sample points by matching rays in different images and using multi-visual stereo geometry priors to learn implicit fields.
[0080] Step S20: extract sampling point features of multiple camera rays, generate target key features based on the multiple sampling point features, input the target key features into the three-dimensional implicit field, input the initial image group into the three-dimensional implicit field, and output the target feature body.
[0081] In this embodiment, SuperPoint (Self-Supervised Interest Point Detection and Description, a feature point detection and description method based on deep learning) is used to detect key points on each input image, and SuperGlue (Learning Feature Matching with Graph Neural Networks, a feature matching method based on graph neural networks) is used to perform point-to-point matching between image pairs to extract key points. These steps can obtain a set of sparse ray matching between image pairs, which is used as input for the three-dimensional implicit field, thereby obtaining the target feature volume.
[0082] Among them, the optimization of the feature volume is affected by the photometric loss of rendering along the camera ray passing through the key point (i.e., the key ray). In the key ray feature enhancement module, this optimization can enhance the features of the key sampling point by combining features from auxiliary rays (i.e., rays passing through auxiliary sampling points near the key sampling point in the image).
[0083] Specifically, key features of key sampling points generated by the matching ray passing through the initial image group are extracted, and auxiliary features of multiple auxiliary sampling points generated by multiple auxiliary rays passing through the initial image group are extracted:
[0084] ;
[0085] in, Indicates sampling point The characteristics of the place, represents the sampling point, represents the progressive feature mask, Representing an initial feature body; utilizing all the auxiliary features to enhance the key features to obtain a target key feature, and inputting the target key feature into the three-dimensional implicit field to output a target feature body.
[0086] Among them, the camera rays proposed in this embodiment include key rays and auxiliary rays. Both rays are emitted from the camera through the pixel center and are associated with the camera view. The feature body can obtain the corresponding features from the sampling points on the rays obtained from the camera center passing through the pixels of the input image. The key rays Through the key sampling points detected in the image, these key sampling points usually correspond to surface points with rich texture and geometric features; auxiliary rays These are points around the key sampling point. When learning the features of the sampling points on the key ray in the optimization framework of this embodiment, context or local structure information is provided for point cloud enhancement. The sampling point information on the key ray and the auxiliary ray is used to enhance the sampling point features of the key ray, thereby better optimizing the 3D representation of the target object.
[0087] Furthermore, multiple key rays need to be matched. This involves applying color consistency constraints to the matched key rays through a matching ray consistency enhancement module. Potentially mismatched rays can be identified by comparing the feature similarity of the key ray integrals within the feature volume. When two key rays in the input image are correctly matched and the corresponding camera pose is accurate, the two rays intersect at the object surface captured by the corresponding pixels in 3D space. Therefore, the color values of the two correctly matched key rays integrated in the corresponding 3D implicit field should be consistent. Furthermore, when the 3D implicit field optimization of the target object converges, the opacity of any sampling point on the key ray in 3D space reaches its maximum value when the sampling point is on the object surface, and the color value at that object surface position contributes the most to the corresponding color value integral. Therefore, the color integrals of the two key rays are consistent and correspond to the color value of the object surface. Similar to the color value integral, when the ray matching relationship is correct and the camera pose is correct, the features of the matched key rays should be consistent. Therefore, the feature similarity between the two matched key rays is used to interpolate the color values of the two key rays. In this case, the features between the two matching key rays and the color and opacity of the sampling points can be optimized simultaneously. This method of jointly optimizing the matching rays can adapt the matching accuracy by utilizing the similarity of the ray features.
[0088] in, It is a progressive feature mask used to filter fine-grained features in the early iterations of coarse-to-fine training. In order to take into account the noise in the camera pose, the present invention parameterizes the camera's transformation matrix and jointly optimizes it with the implicit field parameters in the three-dimensional feature volume to obtain the target feature volume.
[0089] Furthermore, the key feature is fused with all the auxiliary features to obtain the target key feature:
[0090] ;
[0091] in, Indicates key sampling points The key characteristics of the target, Indicates key sampling points The number of auxiliary rays, represents a mathematical function that converts an arbitrary set of real numbers into real numbers representing a probability distribution. represents the key sampling point, represents the auxiliary sampling point, Indicates key sampling points The key features of Indicates auxiliary sampling points inputting the target key features into the three-dimensional implicit field, optimizing the camera pose in the three-dimensional implicit field according to the target key features, and outputting the target key features.
[0092] Among them, due to the noise in the camera pose, the key sampling points on the input image, that is, the three-dimensional surface points collected by the corresponding pixel points in the image, may be inconsistent with the two matching key rays. In three-dimensional space, it is not possible to intersect accurately. In this case, auxiliary rays from around the key ray The local neighboring points sampled on The characteristics can greatly facilitate the More stable optimization of related features. Due to the unreliable camera pose, the points on the object surface may not completely intersect with the surface points obtained by the ray corresponding to the key points on the input image. Therefore, in this embodiment, the The context information of the upsampled surface points is used to optimize the camera pose so that the geometric information of the surface points and the corresponding key points can be better learned. This context information can be used to The adjacent rays of the samples are obtained, that is, the key sampling points are fused The target key features and surrounding auxiliary sampling points Auxiliary features at the position generate key sampling points The key features of the target. Among them, the auxiliary ray Features remain unchanged, that is . In the key ray consistency module, the neural field and camera pose are jointly optimized by combining contextual information for point cloud enhancement and using the matching ray enhancement module to enforce geometric and photometric consistency. The sampling points of the key rays are enhanced using the features of the sampling points in the auxiliary rays, resulting in high-quality rendering and reconstruction with fine geometric details. The key ray enhancement module enriches the features of the key rays using the contextual information of the auxiliary rays and performs point cloud enhancement on the sampling points on the key rays to improve the robustness to noisy poses. The matching ray enhancement module learns cross-view information between matching rays to maintain the consistency of ray matching and eliminate ambiguity between input images.
[0093] Among them, Figure 3 As shown in the figure, in the key ray feature enhancement module, the features of the sampling points of the auxiliary rays are used to enhance the point cloud features of the sampling points on the key rays. In the case of inaccurate pose, the spatial context can be used to obtain the surface feature information of the object that correctly corresponds to the key ray. Since the input image is obtained through perspective projection, all rays in the same image should converge to a point in the same camera space.
[0094] Step S30: Input the target feature body into a geometric network, output a signed distance field value and an output feature, and obtain a geometric result of the target object according to the signed distance field value and the output feature.
[0095] Among them, once the initial 3D implicit field and the camera pose are jointly optimized, the expression of the 3D implicit field is realized, that is, the target feature volume can be used for new view synthesis or multi-view 3D reconstruction. The color prediction of the new view synthesis is performed by the texture network. Completed, and the geometric information of the target object (ie, the geometric result) is obtained by the geometric network Come get.
[0096] Specifically, the target key features of the target feature body are input into the constructed geometric network, and the signed distance field value and the output features are output; according to the output features, the geometric result of the target object in the three-dimensional space is extracted from the signed distance field value:
[0097] ;
[0098] in, Indicates key sampling points The signed distance field value of Indicates key sampling points The output features of Represents a geometric network.
[0099] Among them, after the target feature body is input into the geometric network, the target key features at the sampling points along the matching ray passing through the target feature body are obtained, and the geometric network generates a signed distance field value according to the target key features, and then obtains an opacity from the signed distance field value. To render three-dimensional objects and extract three-dimensional reconstruction information, that is, to obtain the geometric results of the target object in three-dimensional space through the geometric network; in this process, the key ray feature enhancement module is used to enhance the features along the key rays through the features of auxiliary ray sampling, and perform point cloud enhancement, thereby improving the robustness of this optimization process.
[0100] Step S40: obtaining normal information of the target object at each sampling point according to the geometric result, inputting all the normal information into a texture network, and outputting rendering values of matching rays among the multiple camera rays.
[0101] Specifically, according to the gradient value of the geometric result, the normal of the target object at the sampling point is calculated, and the normal information at the sampling point is obtained according to the normal; the normal information is input into the constructed texture network, and the RGB value of the sampling point is output:
[0102] ;
[0103] in, Indicates sampling point RGB values, represents the texture network, Indicates sampling point The output features of Represents matching rays The normalized direction of Indicates sampling point Normal information of the matching ray; using the opacity density and RGB values of all sampling points, the color of the matching ray is rendered to obtain the rendering value of the matching ray:
[0104] ;
[0105] in, Represents matching rays The rendering value of Indicates sampling point The transmittance at Indicates sampling point opacity.
[0106] Among them, the function Indicates that along the corresponding ray The integrated transmittance, opaque density It is calculated using the signed distance field value, the normal vector, and the ray direction.
[0107] The gradient of the signed distance field (i.e., the gradient value of the geometric result) is calculated as Normal information at can be used to enhance the geometric network Training, thus achieving iterative training, in each iteration, sampling a pair of matching key rays and , and accompanied by corresponding auxiliary rays, the geometry network, texture network and feature volume optimization are jointly trained end-to-end.
[0108] Among them, the texture network receives the output features, ray directions from the geometric network The normal at point is used as input to predict the point Color In addition, a matching ray consistency enhancement module is designed in this embodiment. It considers the color matching between rays and learns to maintain the consistency between ray matching features. The matching of ray features improves the rendering quality of the optimization results and the final neural implicit field. The matching ray consistency enhancement module can effectively reduce the impact of mismatched rays by eliminating confusion between camera ray matching.
[0109] The present invention uses a set of images with inaccurate poses as input to reconstruct a three-dimensional implicit field, and jointly optimizes the optimized camera pose and the three-dimensional implicit field. By passing camera rays through key points in the input image, light optimization and matching light consistency are achieved, which can effectively improve the results of pose optimization. The three-dimensional implicit field and camera pose can be further optimized by enhancing point cloud features.
[0110] Furthermore, if Figure 4 As shown, based on the above-mentioned implicit field reconstruction method based on view matching and epipolar geometry constraints, the present invention also provides an implicit field reconstruction system based on view matching and epipolar geometry constraints, wherein the implicit field reconstruction system based on view matching and epipolar geometry constraints includes:
[0111] An implicit field optimization module 51 is configured to obtain an initial image set of a target object and construct an initial three-dimensional implicit field, inputting the photometric error between the initial image set and the multi-view rendering values of the target object into the initial three-dimensional implicit field for training to obtain a three-dimensional implicit field;
[0112] a feature volume optimization module 52 for extracting sampling point features of multiple camera rays, generating target key features based on the multiple sampling point features, inputting the target key features into the three-dimensional implicit field, inputting the initial image group into the three-dimensional implicit field, and outputting a target feature volume;
[0113] A geometric feature collection module 53 is configured to input the target feature body into a geometric network, output a signed distance field value and an output feature, and obtain a geometric result of the target object based on the signed distance field value and the output feature;
[0114] The rendering module 54 is configured to obtain normal information of the target object at each sampling point according to the geometric result, input all the normal information into a texture network, and output rendering values of matching rays among the multiple camera rays.
[0115] Furthermore, if Figure 5 As shown, based on the above-mentioned implicit field reconstruction method and system based on view matching and epipolar geometry constraints, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 5 Only some of the components of the terminal are shown, but it should be understood that implementation of all of the shown components is not required, and more or fewer components may be implemented instead.
[0116] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard drive or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, the memory 20 may include both the internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software installed on the terminal and various data, such as program code installed on the terminal. The memory 20 may also be used to temporarily store data that has been output or is about to be output. In one embodiment, the memory 20 stores an implicit field reconstruction program 40 based on view matching and epipolar geometry constraints. This implicit field reconstruction program 40 based on view matching and epipolar geometry constraints can be executed by the processor 10, thereby implementing the implicit field reconstruction method based on view matching and epipolar geometry constraints described in this application.
[0117] In some embodiments, the processor 10 can be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 20, such as executing the implicit field reconstruction method based on perspective matching and epipolar geometry constraints.
[0118] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.
[0119] In one embodiment, when the processor 10 executes the implicit field reconstruction program 40 based on view matching and epipolar geometry constraints in the memory 20, the following steps are implemented:
[0120] Acquire an initial image group of the target object and construct an initial three-dimensional implicit field, input the photometric error between the initial image group and the multi-view rendering value of the target object into the initial three-dimensional implicit field for training to obtain a three-dimensional implicit field;
[0121] Extracting sampling point features of multiple camera rays, generating target key features based on the multiple sampling point features, inputting the target key features into the three-dimensional implicit field, inputting the initial image group into the three-dimensional implicit field, and outputting a target feature volume;
[0122] Inputting the target feature body into a geometric network, outputting a signed distance field value and an output feature, and obtaining a geometric result of the target object according to the signed distance field value and the output feature;
[0123] Normal information of the target object at each sampling point is obtained according to the geometric result, all the normal information is input into a texture network, and rendering values of matching rays among the multiple camera rays are output.
[0124] The method of obtaining an initial image group of the target object and constructing an initial three-dimensional implicit field, inputting the photometric error between the initial image group and the multi-view rendering value of the target object into the initial three-dimensional implicit field for training to obtain the three-dimensional implicit field specifically includes:
[0125] Acquire an initial image group of the target object captured by a target camera, wherein the initial image group is generated by photographing the target object from multiple perspectives by the target camera;
[0126] Construct an initial three-dimensional implicit field and calculate the photometric error between the initial image group and the multi-view rendering value of the target object:
[0127] ;
[0128] ;
[0129] ;
[0130] ;
[0131] in, Indicates sampling point The photometric error, represents the sampling point, represents the image index in the initial image group, Represents the pixels of the image, Indicates the image and target object at pixel points The color error at Indicates the Image at pixel point The RGB value at Indicates that the target object is at the pixel point The first The rendering value of the camera perspective corresponding to the image, Indicates the The camera pose of each image, Indicates the The rotation matrix of the target camera corresponding to the image, Indicates the The displacement vector of the target camera corresponding to the image, represents the three-dimensional rotation group, Represents three-dimensional rational number space;
[0132] The photometric error is input into the initial three-dimensional implicit field for training to obtain a three-dimensional implicit field.
[0133] Wherein, the camera rays include: key rays and auxiliary rays;
[0134] The step of extracting sampling point features of multiple camera rays, generating target key features based on the multiple sampling point features, inputting the target key features into the three-dimensional implicit field, inputting the initial image group into the three-dimensional implicit field, and outputting a target feature volume may also include:
[0135] Acquire a plurality of the sampling points through which the plurality of key rays pass, and acquire color information and feature information of each of the sampling points;
[0136] Extract multiple color information and multiple feature information with consistency, select corresponding sampling points, match the key rays corresponding to the selected multiple sampling points, and obtain matching rays:
[0137] ;
[0138] ;
[0139] in, represents the matching ray, represents the camera center of the target camera, represents the number of rays, Represents matching rays The normalized direction of .
[0140] The step of extracting sampling point features of multiple camera rays, generating target key features based on the multiple sampling point features, inputting the target key features into the three-dimensional implicit field, inputting the initial image group into the three-dimensional implicit field, and outputting a target feature volume specifically includes:
[0141] Extracting key features of key sampling points generated by the matching ray passing through the initial image group, and extracting auxiliary features of multiple auxiliary sampling points generated by multiple auxiliary rays passing through the initial image group:
[0142] ;
[0143] in, Indicates sampling point The characteristics of the place, represents the sampling point, represents the progressive feature mask, represents the initial feature body;
[0144] The key feature is enhanced by utilizing all the auxiliary features to obtain a target key feature, and the target key feature is input into the three-dimensional implicit field to output a target feature body.
[0145] The method of enhancing the key feature by using the auxiliary feature to obtain the target key feature, inputting the target key feature into the three-dimensional implicit field, and outputting the target feature body specifically includes:
[0146] The key feature is combined with all the auxiliary features to obtain the target key feature:
[0147] ;
[0148] in, Indicates key sampling points The key characteristics of the target, Indicates key sampling points The number of auxiliary rays, represents a mathematical function that converts an arbitrary set of real numbers into real numbers representing a probability distribution. represents the key sampling point, represents the auxiliary sampling point, Indicates key sampling points The key features of Indicates auxiliary sampling points Auxiliary features at
[0149] The target key features are input into the three-dimensional implicit field, the camera pose in the three-dimensional implicit field is optimized according to the target key features, and the target key features are output.
[0150] The step of inputting the target feature body into a geometric network, outputting a signed distance field value and an output feature, and obtaining a geometric result of the target object according to the signed distance field value and the output feature specifically includes:
[0151] Inputting the target key features of the target feature body into the constructed geometric network, and outputting the signed distance field value and the output features;
[0152] Extracting the geometric result of the target object in three-dimensional space from the signed distance field value according to the output feature:
[0153] ;
[0154] in, Indicates key sampling points The signed distance field value of Indicates key sampling points The output features of Represents a geometric network.
[0155] The step of obtaining normal information of the target object at each sampling point according to the geometric result, inputting all the normal information into a texture network, and outputting rendering values of matching rays among the multiple camera rays specifically includes:
[0156] Calculating the normal of the target object at the sampling point according to the gradient value of the geometric result, and obtaining normal information at the sampling point according to the normal;
[0157] Input the normal information into the constructed texture network and output the RGB value of the sampling point:
[0158] ;
[0159] in, Indicates sampling point RGB values, represents the texture network, Indicates sampling point The output features of Represents matching rays The normalized direction of Indicates sampling point Normal information;
[0160] The color of the matching ray is rendered using the opacity density and RGB values of all sampling points to obtain the rendering value of the matching ray:
[0161] ;
[0162] in, Represents matching rays The rendering value of Indicates sampling point The transmittance at Indicates sampling point opacity.
[0163] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an implicit field reconstruction program based on perspective matching and epipolar geometric constraints, and when the implicit field reconstruction program based on perspective matching and epipolar geometric constraints is executed by a processor, the steps of the implicit field reconstruction method based on perspective matching and epipolar geometric constraints as described above are implemented.
[0164] In summary, the present invention provides an implicit field reconstruction method based on view matching and epipolar geometry constraints and related equipment, the method comprising: obtaining an initial image group of a target object and constructing an initial three-dimensional implicit field, inputting the photometric error between the initial image group and the multi-view rendering value of the target object into the initial three-dimensional implicit field for training to obtain a three-dimensional implicit field; extracting sampling point features of multiple camera rays, generating target key features based on the multiple sampling point features, inputting the target key features into the three-dimensional implicit field, and inputting the initial image group into the three-dimensional implicit field to output a target feature body; inputting the target feature body into a geometric network, outputting a signed distance field value and an output feature, and obtaining a geometric result of the target object based on the signed distance field value and the output feature; obtaining normal information of the target object at each sampling point based on the geometric result, inputting all the normal information into a texture network, and outputting rendering values of matching rays among the multiple camera rays. The present invention uses a set of images with inaccurate poses as input to reconstruct a three-dimensional implicit field, and jointly optimizes the optimized camera pose and the three-dimensional implicit field. By passing camera rays through key points in the input image, light optimization and matching light consistency are achieved, which can effectively improve the results of pose optimization. The three-dimensional implicit field and camera pose can be further optimized by enhancing point cloud features.
[0165] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal comprising the element.
[0166] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When executed, the program can include the processes in the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.
[0167] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. An implicit field reconstruction method based on view matching and epipolar geometry constraints, characterized in that: The implicit field reconstruction method based on view matching and epipolar geometry constraints includes: Acquire an initial image group of the target object and construct an initial three-dimensional implicit field, input the photometric error between the initial image group and the multi-view rendering value of the target object into the initial three-dimensional implicit field for training to obtain a three-dimensional implicit field; By optimizing the initial three-dimensional implicit field, the target object is optimized to obtain a signed distance field; Extracting sampling point features of multiple camera rays, generating target key features based on the multiple sampling point features, inputting the target key features into the three-dimensional implicit field, inputting the initial image group into the three-dimensional implicit field, and outputting a target feature volume; The step of extracting sampling point features of multiple camera rays, generating target key features based on the multiple sampling point features, inputting the target key features into the three-dimensional implicit field, inputting the initial image group into the three-dimensional implicit field, and outputting a target feature volume specifically includes: Extracting key features of key sampling points generated by matching rays passing through the initial image group, and extracting auxiliary features of multiple auxiliary sampling points generated by multiple auxiliary rays passing through the initial image group: Among them, f(p) represents the feature at the sampling point p, p represents the sampling point, represents the progressive feature mask, represents the initial feature body; utilizing all the auxiliary features to enhance key features to obtain target key features, and inputting the target key features into the three-dimensional implicit field to output a target feature volume; Matching rays in different initial images and using multi-visual stereo geometry priors to learn the three-dimensional implicit field, and enhancing the sampling point features by utilizing the contextual feature relationships between the sampling points; Use the sampling point features of the auxiliary ray to enhance the point cloud of the sampling point features on the key ray; Inputting the target feature body into a geometric network, outputting a signed distance field value and an output feature, and obtaining a geometric result of the target object according to the signed distance field value and the output feature; Inputting the target feature body into a geometric network, outputting a signed distance field value and an output feature, and obtaining a geometric result of the target object according to the signed distance field value and the output feature specifically includes: Inputting the target key features of the target feature body into the constructed geometric network, and outputting the signed distance field value and the output features; Extracting the geometric result of the target object in three-dimensional space from the signed distance field value according to the output feature: [SDF(p k )|f″(p k )]=Φ g (f′(p k )); Among them, SDF(p k ) represents the key sampling point p k The signed distance field value, f″(p k ) represents the key sampling point p k The output features, Φ g represents the geometric network, f′(p k ) represents the key sampling point p k Key characteristics of the target; Obtaining normal information of the target object at each sampling point according to the geometric result, inputting all the normal information into a texture network, and outputting rendering values of matching rays among the multiple camera rays; The texture network is iteratively trained. In each iteration, a pair of matching key rays is sampled and the corresponding auxiliary rays are used to achieve end-to-end joint training of the geometry network, texture network and feature volume optimization.
2. The implicit field reconstruction method based on view matching and epipolar geometry constraints according to claim 1, characterized in that: The method of obtaining an initial image group of the target object and constructing an initial three-dimensional implicit field, inputting the photometric error between the initial image group and the multi-view rendering value of the target object into the initial three-dimensional implicit field for training to obtain the three-dimensional implicit field, specifically includes: Acquire an initial image group of the target object captured by a target camera, wherein the initial image group is generated by photographing the target object from multiple perspectives by the target camera; Construct an initial three-dimensional implicit field and calculate the photometric error between the initial image group and the multi-view rendering value of the target object: R i ∈SO(3); t i ∈R 3 ; Among them, L p Represents the photometric error of the sampling point p, p represents the sampling point, i represents the image index in the initial image group, and x represents the pixel point of the image. Represents the color error between the image and the target object at pixel x, I i (x) represents the RGB value of the i-th image at pixel x, Represents the rendering value of the camera perspective corresponding to the i-th image of the target object at pixel x, represents the camera pose of the i-th image, R i Represents the rotation matrix of the target camera corresponding to the i-th image, t i represents the displacement vector of the target camera corresponding to the i-th image, SO(3) represents the three-dimensional rotation group, R 3 Represents three-dimensional rational number space; The photometric error is input into the initial three-dimensional implicit field for training to obtain a three-dimensional implicit field.
3. The implicit field reconstruction method based on view matching and epipolar geometry constraints according to claim 1, characterized in that: The camera rays include: key rays and auxiliary rays; The step of extracting sampling point features of multiple camera rays, generating target key features based on the multiple sampling point features, inputting the target key features into the three-dimensional implicit field, inputting the initial image group into the three-dimensional implicit field, and outputting a target feature volume may also include: Acquire a plurality of the sampling points through which the plurality of key rays pass, and acquire color information and feature information of each of the sampling points; Extract multiple color information and multiple feature information with consistency, select corresponding sampling points, match the key rays corresponding to the selected multiple sampling points, and obtain matching rays: r(t)=r o +tr d ; t≥0; Among them, r(t) represents the matching ray, r o represents the camera center of the target camera, t represents the displacement vector from any feature point on ray r to the target camera, r d represents the normalized direction of the matching ray r.
4. The implicit field reconstruction method based on view matching and epipolar geometry constraints according to claim 1, characterized in that: The method of enhancing the key feature by using the auxiliary feature to obtain a target key feature, inputting the target key feature into the three-dimensional implicit field, and outputting a target feature body specifically includes: The key feature is combined with all the auxiliary features to obtain the target key feature: Among them, f′(p k ) represents the key sampling point p k The target key feature, N o represents the key sampling point p k The number of auxiliary rays, Softmax represents a mathematical function used to convert a set of arbitrary real numbers into real numbers representing probability distribution, p k represents the key sampling point, q i represents the auxiliary sampling point, f(p k ) represents the key sampling point p k The key feature at i ) represents the auxiliary sampling point q i Auxiliary features at The target key features are input into the three-dimensional implicit field, the camera pose in the three-dimensional implicit field is optimized according to the target key features, and the target feature body is output.
5. The implicit field reconstruction method based on view matching and epipolar geometry constraints according to claim 2, characterized in that: The step of obtaining normal information of the target object at each sampling point according to the geometric result, inputting all the normal information into a texture network, and outputting rendering values of matching rays among the multiple camera rays specifically includes: Calculating the normal of the target object at the sampling point according to the gradient value of the geometric result, and obtaining normal information at the sampling point according to the normal; Input the normal information into the constructed texture network and output the RGB value of the sampling point: c(p)=Φ t (f″(p),r d ,Normal(p)); Among them, c(p) represents the RGB value of sampling point p, Φ t represents the texture network, f″(p) represents the output feature of sampling point p, r d Represents the normalized direction of the matching ray r, and Normal(p) represents the normal information of the sampling point p; The color of the matching ray is rendered using the opacity density and RGB values of all sampling points to obtain the rendering value of the matching ray: Among them, c(r) represents the rendering value of the matching ray r, represents the transmittance at the sampling point p, σ(p) represents the opacity of the sampling point p, and t represents the displacement vector from any feature point on the ray r to the target camera.
6. An implicit field reconstruction system based on view matching and epipolar geometry constraints, characterized in that: The implicit field reconstruction system based on view matching and epipolar geometry constraints is applied to the implicit field reconstruction method based on view matching and epipolar geometry constraints according to any one of claims 1 to 5, and the implicit field reconstruction system based on view matching and epipolar geometry constraints includes: An implicit field optimization module is configured to obtain an initial image set of a target object and construct an initial three-dimensional implicit field, inputting the photometric error between the initial image set and the multi-view rendering values of the target object into the initial three-dimensional implicit field for training to obtain a three-dimensional implicit field; a feature volume optimization module, configured to extract sampling point features of multiple camera rays, generate target key features based on the multiple sampling point features, input the target key features into the three-dimensional implicit field, input the initial image group into the three-dimensional implicit field, and output a target feature volume; A geometric feature collection module is used to input the target feature body into a geometric network, output a signed distance field value and an output feature, and obtain a geometric result of the target object according to the signed distance field value and the output feature; A rendering module is used to obtain normal information of the target object at each sampling point according to the geometric result, input all the normal information into a texture network, and output rendering values of matching rays among the multiple camera rays.
7. A terminal, characterized in that: The terminal includes: a memory, a processor, and an implicit field reconstruction program based on perspective matching and epipolar geometric constraints stored in the memory and runnable on the processor. When the implicit field reconstruction program based on perspective matching and epipolar geometric constraints is executed by the processor, the steps of the implicit field reconstruction method based on perspective matching and epipolar geometric constraints as described in any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an implicit field reconstruction program based on perspective matching and epipolar geometric constraints. When the implicit field reconstruction program based on perspective matching and epipolar geometric constraints is executed by a processor, the steps of the implicit field reconstruction method based on perspective matching and epipolar geometric constraints as described in any one of claims 1 to 5 are implemented.