Free-form collection method for high-dimensional materials
The free-form acquisition method addresses the challenges of high-dimensional material acquisition by converting appearance scans into a geometric learning problem and using a neural network to optimize illumination and recover high-quality material attributes, enhancing efficiency and accuracy.
Patent Information
- Application Number
- JP2023558118
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-07
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-06-07
AI Technical Summary
Current methods for high-dimensional material acquisition, particularly in computer graphics and vision, face challenges in efficiently scanning object appearances off-plane and achieving high-quality material attribute recovery, especially with portable and lightweight devices.
A free-form acquisition method that converts appearance scans into a geometric learning problem in an unstructured point cloud, using a neural network to aggregate information from unstructured images and optimize illumination patterns for high-quality material recovery.
This method effectively recovers high-quality material attributes from disordered and non-uniformly distributed data, improving the efficiency and accuracy of high-dimensional material acquisition without relying on specific acquisition devices.
Smart Images

Figure 0007692224000136 
Figure 0007692224000137 
Figure 0007692224000138
Abstract
Description
Technical Field
[0001] The present invention relates to a free-form acquisition method for high-dimensional materials and belongs to the fields of computer graphics and computer vision.
Background Art
[0002] The digitization of objects in the real world is one of the core problems in computer graphics and vision. Currently, the digitized real object may be represented by a three-dimensional grid model and a six-dimensional bidirectional reflectance distribution function (SVBRDF) that changes according to space. The digitized real object can clearly reproduce its original state under any viewing angle and lighting conditions, and has important applications in fields such as cultural heritage, e-commerce, computer games, and movie production.
[0003] Although a high-precision geometric model can be easily obtained by a commercial mobile 3D scanner, it is also desired to develop a lightweight device to freely perform appearance scanning for the following reasons. First, if the posture of the video camera can be reliably estimated, objects of different sizes can be scanned. Second, due to the portability of the device, on-site scanning can be performed on objects that do not allow transportation, such as precious cultural properties. In addition, a lightweight device requires a short manufacturing time and low cost, so it is more acceptable to more people. It further provides a user-friendly experience similar to geometric scanning.
[0004] Although the demand is increasing rapidly, effective off-plane appearance scanning remains one of the problems waiting to be solved. On the other hand, almost all conventional movable appearance scanning operations are taken in the case of a single point / parallel light, so the sampling efficiency in the four-dimensional observation and illumination directions is lower, and prior knowledge is required to replace the spatial resolution with angular accuracy (Giljoo Nam, Joo Ho Lee, Diego Gutierrez, and Min H Kim. 2018. Practical SVBRDF acquisition of 3D objects with unstructured flash photography. In SIGGRAPH Asia Technical Papers. 267.). On the other hand, the fixed collection system depends on certain image conditions when the illumination changes, and currently, it is not yet understood how to expand it to mobile devices. Since mobile devices have unstructured and constantly changing images and their external dimensions are smaller, they cannot fully cover the illumination field.
Summary of the Invention
Problems to be Solved by the Invention
[0005] The object of the present invention is to provide a free-form collection method for high-dimensional materials for the drawbacks of the prior art. This method can effectively utilize the collection condition information of each image to recover high-quality object material attributes from the disordered and non-uniformly distributed collection results.
Means for Solving the Problems
[0006] The present invention discloses a free-form acquisition method for high-dimensional materials. The main gist of this method is that a free-form appearance scan can be converted into a geometric learning problem in an unstructured point cloud, and each point in the point cloud may represent an image measurement value and the pose information of the object at the time of image shooting. Based on this gist, the present invention designs a neural network, which can effectively aggregate information in different unstructured images to reconstruct spatially independent reflection attributes, and optimize the illumination pattern used in the acquisition stage to finally obtain high-quality material acquisition results. The present invention does not depend on a specific acquisition device, and may be collected by holding the device by hand using a fixed object, or by placing the object on a rotating table and rotating it for collection by a fixed device, but is not limited to these two methods.
[0007] This method converts the learning of material information into a geometric learning problem in an unstructured point cloud, constructs sampling results in a plurality of different illumination and observation directions into a high-dimensional point cloud, and proposes that each point in the point cloud is a vector composed of an image measurement value and the pose information of the object at the time of image shooting. This method can effectively aggregate the information of unstructured images from a high-dimensional point cloud that is disordered, irregular, unevenly distributed, and has limited accuracy to recover high-quality material attributes. The skeletal representation is F(G(high dimensional point cloud)) = m, The point cloud data feature extraction method G is not limited to a specific network structure, and other methods that can extract features from the point cloud can also be applied. The non-linear mapping network F is not limited to a fully connected network, and the representation of the object material attribute is not limited to the Lumitexel vector m.
[0008] This method includes two stages: a training stage and an acquisition stage.
[0009] The training stage includes: Step (1) of obtaining the parameters of the acquisition device and generating the acquisition results that simulate an actual video camera as training data; Step (2) of training a neural network using the generated training data, and the characteristics of the neural network are The input of the neural network is the Lumitexel vector in k unstructured samplings, where k is the number of samplings, and each value of Lumitexel describes the reflection intensity along a certain observation direction of the sampling point for the incident light from each light source. Lumitexel has a linear relationship with the luminous intensity of the light source and is simulated by a linear fully connected layer (2.1). The first layer of the neural network includes a linear fully connected layer and is used to simulate the illumination pattern used during actual collection and convert the k Lumitexels into camera collection results. These k collection results are combined with the pose information of the corresponding sampling points to form a high-dimensional point cloud (2.2). From the second layer onwards is a feature extraction network, which independently extracts features from each point in the high-dimensional point cloud to obtain feature vectors (2.3). After the feature extraction network, it is a max pooling layer for aggregating the feature vectors extracted from k unstructured images to obtain a global feature vector (2.4). After the max pooling layer, it is a non-linear mapping network for recovering high-dimensional material information based on the global feature vector (2.5), and includes The collection stage is A material collection step, in which the collection device irradiates the target three-dimensional object in sequence according to the illumination pattern, the video camera obtains photos in a set of unstructured images, uses the photos as input, and obtains a geometric model with texture coordinates for the sampling object and the pose of the video camera during photo shooting (1). A material recovery step, which includes obtaining the pose for each vertex of each valid texture coordinate on the sampling object when taking each photo based on the pose of the video camera during photo shooting in the material collection stage, constructing the high-dimensional point cloud as the input of the second-layer feature extraction network of the neural network based on the collected photos and pose information, and calculating and obtaining high-dimensional material information (step (2)).
[0010] Furthermore, the unstructured sampling is free random sampling with a non-fixed viewing angle, the sampling data is disordered, irregular, and unevenly distributed, and it may be collected by holding the collection device by hand using a fixed object, or the object may be placed on a rotating table and rotated and collected by a fixing device.
[0011] Furthermore, in the process of generating training data, when the light source is colored, it is necessary to correct the spectral response relationship among the light source, the sampling object, and the video camera. Regarding the correction method, The spectral distribution curve of the unknown color light source L is Defined as TIFF0007692224000001.tif917, where λ represents the wavelength, and c 1 Represents one of the three channels of RGB, and the spectral distribution curve L(λ) of the light source with light intensity {I R , I G , I B} may be shown in the following formula 2.
Equation
Equation
[0012] Furthermore, in the training stage step (2.1), the relationship between the observed value B in the photograph of the sampling point p on the object surface, the reflection function f r and the light intensity of each light source is given by the following Equation 5,
Equation
[0013] The input of the neural network is the Lumitexel vector in k unstructured samplings, denoted as m(l; P),
Equation
Equation
[0014] Furthermore, in the training stage step (2.3), the formula of the feature extraction network is shown in the following formula 9,
Equation
[0015] Furthermore, in step (2.5) of the training stage, the skeletalized representation of the non-linear mapping network is shown in Equation 11 below,
Equation
[0016] Furthermore, the loss function of the neural network is designed as follows, (1) Assume one Lumitextel space, which is a cube with the center at the spatial position x of the sampling point p of the cube, and the central coordinate system of the cube has the x - axis direction being TIFF0007692224000021.tif97, and the z - axis direction being TIFF0007692224000022.tif87, TIFF0007692224000023.tif87 being the geometric normal vector, TIFF0007692224000024.tif97 being an arbitrary unit vector orthogonal to TIFF0007692224000025.tif87, (2) Assume one video camera, with the observation direction being the positive direction of the z - axis of the cube, (3) In the case of diffuse - reflecting Lumitexel, the resolution of the cube is 6×N d 2 and in the case of specular - reflecting Lumitexel, the resolution of the cube is 6×N s 2 that is, from each face, N d 2 , N s 2 points are uniformly sampled as virtual point light sources with unit light intensity, a. Set the specular reflectivity ρ s of the sampling point to 0, and generate the diffuse - reflection feature vector TIFF0007692224000026.tif79 in this Lumitexel space, b. Set the diffuse reflectivity ρ d to 0, and generate the specular - reflection feature vector TIFF0007692224000027.tif79 in this Lumitexel space, c. Let the output of the neural network be vectors m d , m s , and m d be of the same length as TIFF0007692224000028.tif79, and m s be is the same as the length of TIFF0007692224000029.tif79, and the vector m d、 m s is the diffuse reflection feature vector TIFF0007692224000030.tif79, the specular reflection vector is the prediction of TIFF0007692224000031.tif79, (4) The loss function of the material feature part is shown in the following Equation 12,
Equation
Equation
[0017] Furthermore, in the collection stage, after the material collection is completed, geometric alignment is performed, and then material restoration is performed. Specifically, geometric alignment is to scan the object by a scanner to obtain a geometric model, align it with the geometric model reconstructed in three dimensions, and then replace the geometric model reconstructed in three dimensions.
[0018] Furthermore, for the valid texture coordinates, pixels in the collected photo are sequentially retrieved based on the pose information of the photo and the sampling points, the validity of the pixels is determined, and a high-dimensional point cloud is constructed in combination with the pose for the vertices. For a point p on the surface of the sampled object determined for a certain valid texture coordinate, the criterion for determining that the j-th sampling is valid for the vertex p is as follows: (1) The position x of the vertex p p j is visible to the video camera in this sampling, and x p j is within the sampling space defined when training the network. (2) TIFF0007692224000037.tif1137, TIFF0007692224000038.tif87 is the dot product operation, θ is the lower limit of the valid sampling angle, ω o ′ indicates the direction of the outgoing light in the world coordinate system. TIFF0007692224000039.tif108 indicates the normal vector of the j-th vertex p. (3) The numerical values of each channel of the pixel in the photo are within the interval [a, b], where a and b are the lower and upper limits of the valid sampling luminance, which is expressed as When all three conditions are satisfied, the j-th sampling is considered valid for the vertex p, and the j-th sampling result is added to the high-dimensional point cloud.
[0019] Furthermore, after recovering the material information, the material parameters may be fitted, which is a step of fitting the local coordinate system and roughness. For a point p on the surface of the sampled object determined for a certain valid texture coordinate, based on the single-channel specular reflection vector output from the network, the local coordinate system and roughness in the material parameters are fitted by the L-BFGS-B method in step (1). A step of fitting the reflectance, which obtains the specular reflectance and the diffuse reflectance by a trust region algorithm, keeps the local coordinate system and the roughness obtained in the previous process constant when obtaining them, synthesizes the observed values at the viewing angles used for collection, and makes it as close as possible to the observed values collected and obtained. It is divided into two steps: step (2).
Advantages of the Invention
[0020] The beneficial effects of the present invention are as follows. The method of the present invention converts the learning of material information into a geometric learning problem in an unstructured point cloud, constructs sampling results in a plurality of different illumination and observation directions into a high-dimensional point cloud, and proposes that each point in the point cloud is a vector composed of an image measurement value and the pose information of the object at the time of image shooting. This method can effectively aggregate the information of unstructured images from a high-dimensional point cloud that is disordered, irregular, unevenly distributed, and limited in accuracy, and can recover high-quality material attributes.
Brief Description of the Drawings
[0021]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Embodiment for Carrying Out the Invention
[0022] To make the object, technical solution and advantages of the present invention clearer, the present invention will be described in detail below with reference to the drawings.
[0023] The specific implementation of the method for freely collecting high-dimensional materials according to the present invention may be divided into the following training stage and collection stage.
[0024] I. Training Stage 1. Generate training data, obtain the parameters of the collection device, and the parameters include the distance and angle from the light source to the origin of the sampling space, the characteristic curve of the light source, the distance and angle from the video camera to the origin of the sampling space, the internal parameters and external parameters of the video camera. Use these parameters to generate the collection result that simulates the actual video camera as training data. The rendering model used when generating the training data is the GGX model, and the generation formula is shown in the following formula 1,
Equation
[0025] When the light source used for collection is color, first, it is necessary to correct the spectral response relationships of the light source, the sampling object, and the video camera. Regarding the correction method, the spectral distribution curve of the unknown color light source L is defined as TIFF0007692224000043.tif917, where λ represents the wavelength, c 1 represents one of the three channels of RGB, and the spectral distribution curve L(λ) of the light source with the light intensity {I R , I G , I B} can be shown in Equation 2 below,
Equation
Equation
Equation
[0026] 2. Use the generated training data to train the neural network shown in Figure 4. The characteristics of the neural network include the following (1) to (6).
[0027] (1) Observation value B in the photo of the sampling point p on the object surface, reflection function f rThe relationship between the light intensity of each light source is shown in Equation 5 below: [Equation] Here, I represents the emission information of each light source l, including the spatial position x of the light source l, l the normal vector n of the light source l, l the emission intensity I(l) of the light source l, and P includes the parameter information of the sampling point p, including the spatial position x of the sampling point, p the material parameters n, t, α, x α, y ρ, d ρ, s . TIFF0007692224000050.tif719 explains the light intensity distribution in different incident directions of the light source l, and V represents the binary function of visibility with respect to x l of x p . TIFF0007692224000051.tif711 is the dot product operation of two vectors, and negative values will be cut off to 0. f r (ω i ′; ω o ′, P) is a two-dimensional reflection function with respect to ω o ′ when ω i ′ is constant.
[0028] The input of the neural network is k Lumitexels in unordered and irregular sampling, where k is the number of sampling points, and Lumitexel is a vector denoted as m(l; P), and each value in it explains the reflection intensity along a certain observation direction of the sampling point with respect to the incident light from each light source. [Equation] B in the above equation is the display in a single channel. When the light source is a color light source, B is extended to the following form: [Equation] Here, f r (ω i ′; ωo ′, P, c 2 ) is f r (ω i ′; ω o in the formula of ′, P) It is the result of TIFF0007692224000054.tif858. B has a linear relationship with the luminous intensity of the light source and may be simulated by a linear fully connected layer.
[0029] (2) The first layer of the neural network includes a linear fully connected layer, and the parameter matrix of the linear fully connected layer is obtained by training according to the following formula [Number] Here, W raw is the parameter to be trained, W l is the illumination matrix, which has a size of 1×N in the case of a single-channel light source and 3×N in the case of a color light source. N is the vector length of Lumitexel, and f W is a mapping. By converting W raw , the generated illumination matrix can correspond to the possible luminous intensity of the light source. In this example, the mapping f W selects the Sigmoid function and restricts the initial value of the illumination matrix W l of the first layer network to (0, 1), but f W is not limited to the Sigmoid function.
[0030] Taking Wl as the luminous intensity of the light source, k sampling observation values TIFF0007692224000056.tif970 are calculated and obtained according to the above relationship (1).
[0031] (3) From the second layer, it is a feature extraction network. The k samplings independently extract features to obtain a feature vector, and the formula is shown in Equation 9 below [Number] Here, f is a one-dimensional convolution function, and the size of the convolution kernel is 1×1. TIFF0007692224000058.tif1032 are the spatial positions of the sampling points, the geometric normal vectors, and the geometric tangent vectors when sampling for the j-th time respectively. TIFF0007692224000059.tif87 may be obtained by a geometric model acquired after being reconstructed three-dimensionally or scanned by a scanner. TIFF0007692224000060.tif96 is TIFF0007692224000061.tif87 is an arbitrary unit vector orthogonal to, and by sampling the pose of the video camera for the j-th time, TIFF0007692224000062.tif919 is transformed to TIFF0007692224000063.tif1019 can be obtained, and V feature (j) is the feature vector output from the network when sampling for the j-th time.
[0032] (4) After the feature extraction network is a max-pooling layer. The formula for the max-pooling operation is shown in Equation 10 below.
Number
[0033] (5) After the max-pooling layer is a non-linear mapping network.
Number
[0034] (6) The loss function of the neural network includes the following (6.1) to (6.4). (6.1) Assume one Lumitextel space, which is a cube with the center at the spatial position x p of the sampling point. The central coordinate system of the cube has the x-axis direction as TIFF0007692224000066.tif96, and the z-axis direction as TIFF0007692224000067.tif87, (6.2) Assume one video camera, with the observation direction being the positive direction of the z-axis of the cube, (6.3) In the case of diffuse reflection Lumitexel, the resolution of the cube is 6×N d 2 and in the case of specular reflection Lumitexel, the resolution of the cube is 6×N s 2 That is, from each face, N d 2 , N s 2 points are uniformly sampled as virtual point light sources with unit light intensity. In this embodiment, N d = 8, N s = 32, a. Set the specular reflectivity ρ s of the sampling point to 0, and generate the diffuse reflection feature vector TIFF0007692224000068.tif79 in this Lumitexel space, b. Set the diffuse reflectivity ρ d to 0, and generate the specular reflection feature vector TIFF0007692224000069.tif78 in this Lumitexel space, c. Output of the neural network as vector m d , m s is taken as m d and m is the same length as TIFF0007692224000070.tif79, and m s is the same length as TIFF0007692224000071.tif78. The vector m d、 m s is respectively the diffuse reflection feature vector TIFF0007692224000072.tif79, the specular reflection vector TIFF0007692224000073.tif78 prediction, and (6.4) The loss function of the material feature part is shown in Equation 12 below,
Equation
Equation
[0035] 3. After the training is completed, the parameters W of the linear fully connected layer of the network raw are taken out, and the official W l = f W (W raw ) is used as the illumination pattern after being converted.
[0036] II. Collection stage The collection stage may be further subdivided into a material collection stage, a geometric alignment stage (optionally), and a material recovery stage.
[0037] 1. Material collection stage The collection device irradiates the target three-dimensional object in sequence according to the illumination pattern, the video camera acquires photos in a set of unstructured images, and uses the photos as input to obtain the geometric model of the sampling object and the pose of the video camera at the time of photography by a three-dimensional reconstruction tool disclosed in the industrial field. 2. Geometric alignment stage (optionally) (1) The object is scanned by a high-precision scanner to obtain a geometric model, (2) The geometric model obtained by scanning with the scanner is aligned with the geometric model reconstructed in three dimensions, and the geometric model reconstructed in three dimensions is replaced. The alignment method may use the method CPD disclosed in the field (A. Myronenko and X. Song. 2010. Point Set Registration: Coherent Point Drift. IEEE PAMI 32, 12 (2010), 2262-2275. TIFF0007692224000080.tif873). 3. Material recovery stage (1) The pose of each vertex of the sampling object when taking the j-th photo based on the pose of the video camera at the time of photography in the photo-taking step of the material collection step TIFF0007692224000081.tif1032 is obtained, (2) Process the geometric model of the sampled object obtained by reconstructing in three dimensions or the geometric model of the sampled object obtained by scanning with the scanner after alignment using the tool Iso-charts disclosed in the field to obtain a geometric model with texture coordinates. (3) For the valid texture coordinates, for the set of collected photos r 1 , r 2 , …, r π and the pose information of the sampling points, sequentially extract the pixels in the photos, determine the validity of the pixels, and combine the pose TIFF0007692224000082.tif932 to form the input vector of the feature extraction network of the second layer of the neural network for the high-dimensional point cloud, and calculate and obtain the output vectors m d and m s .
[0038] For a point p on the surface of the sampled object determined for certain valid texture coordinates, the criterion for determining that the j-th sampling is valid for point p is 1) x p j is visible to the video camera in the sampling, and x p j is within the sampling space defined when training the network. 2) TIFF0007692224000083.tif1036, TIFF0007692224000084.tif77 is a dot product operation, θ is the lower limit of the valid sampling angle, and in this embodiment, θ = 0.3. 3) The numerical values of each channel of the pixel in the photo are in the interval [a, b], where a and b are the lower and upper limits of the valid sampling luminance, and in this embodiment, a = 32 and b = 224. When all three conditions are satisfied, the j-th sampling is considered valid for point p, and the j-th sampling result is added to the high-dimensional point cloud. (4) Fit the material parameters, which is divided into two steps of 1) and 2) below. 1) Fitting of local coordinate system and roughness For the point p on the surface of the sampling object determined for a certain effective texture coordinate, based on the single-channel specular reflection vector output from the network, fit the local coordinate system and roughness in the material parameters by the L-BFGS-B method, and the optimization target is shown in the following formula 14.
Equation
[0039] 2) Fitting of reflectance This process obtains the specular reflectance and diffuse reflectance by the trust region algorithm, and the fitting target is shown in the following formula 15.
Equation
Equation
[0040] The following is a system example of a specific collection device. Figure 1 is a three-dimensional display of the system example, Figure 2 is a front view, and Figure 3 is a side view. The collection device is composed of one lamp panel, and one camera for collecting images is fixed on the upper part. A total of 512 LED lamp beads are densely arranged on the lamp panel. The lamp beads are controlled by an FPGA and can adjust the emission luminance and emission time.
[0041] The following is an example of an acquisition system that applies the method of the present invention. The entire system is divided into the following several modules.
[0042] Regarding the preparation module, Provide a data set for network training. When this part uses the GGX model and inputs a set of material parameters, the pose information of k sampling points, and the position of the video camera, a high-dimensional point cloud composed of k reflection situations can be obtained. The network training part uses the Pytorch open-source framework and is trained by an Adam optimization device. The network structure is as shown in Figure 6. Each rectangle represents one layer of neurons, and the number in the rectangle represents the number of neurons in that layer. The leftmost layer is the input layer, and the rightmost layer is the output layer. The solid arrows between layers represent full connections, and the dashed arrows represent convolutions.
[0043] Regarding the collection module, The device is as shown in Figures 1, 2, and 3. The above has described the specific structure. However, the size of the sampling space defined in this system, as well as the spatial positional relationship between the collection device and the sampling space, are shown in Figure 4.
[0044] Regarding the recovery module, A geometric model with texture coordinates is calculated and obtained from the geometric model of the sampled object reconstructed in three dimensions or the geometric model of the sampled object scanned by the scanner after alignment, a trained neural network is loaded, a material feature vector is predicted for each vertex of the geometric model with texture coordinates, and a coordinate system and material parameters for rendering are fitted.
[0045] Figure 5 shows the working process of this embodiment. First, training data is generated, randomly sampled to obtain 200 million Lumitexels, 80% of which is used as the training set and the rest as the validation set. When training the network, the parameters are initialized by the Xavier method and the learning rate is set to 1e-4. The illumination pattern is in color, the size of the illumination matrix is (3, 512), and the three rows of the matrix respectively represent the illumination patterns of the three channels of red, green, and blue. After the training is completed, the illumination matrix is taken out and converted into an illumination pattern. The parameters of each column specify the emission intensity of the light source at that position. Figure 7 shows one illumination pattern of the three channels of red, green, and blue obtained by training the network. The following process is as follows. First, hold the device by hand, make the lamp panel emit light according to the illumination pattern, and simultaneously take pictures of the object with a video camera to obtain a set of sampling results. Second, the geometric model of the sampled object obtained by three-dimensional reconstruction or the geometric model of the sampled object obtained by scanning with the scanner after alignment is processed by Isochart to obtain a geometric model with texture coordinates. Third, for each vertex of the geometric model with texture coordinates, find the corresponding effective actual shooting data based on the pose during sampling and the pixel values of the sampling photos, construct a high-dimensional point cloud and input it into the network to recover the diffuse reflection feature vector and the specular reflection feature vector. Fourth, based on the diffuse reflection feature vector and the specular reflection feature vector output from the network, for each vertex, fit the coordinate system and roughness for rendering by the LBFGS-B method, and obtain the specular reflectance and the diffuse reflectance by the trust region algorithm.
[0046] Figure 8 shows two Lumitexel vectors in the validation set recovered by the above system. The left column is TIFF0007692224000094.tif78, and the right column is the corresponding m s is.
[0047] FIG. 9 shows the material attribute results recovered by performing a material appearance scan on the sampled object by the above system. The first row represents the three components of TIFF0007692224000095.tif840, and the second row represents the three components of TIFF0007692224000096.tif838 of the sampled object. The third row represents the roughness coefficient α x , α y of the sampled object, where the tone value represents the magnitude of the numerical value.
[0048] The above are merely preferred embodiments, and the present invention is not limited to the above embodiments. As long as the technical effects of the present invention can be realized by the same means, all of them should belong to the protection scope of the present invention. Within the protection scope of the present invention, various different modifications and changes can be made to its technical solution and / or embodiments.
Claims
1. A method for freely collecting high-dimensional materials, including a training stage and a collection stage, The training stage includes the following steps (1) to (2), In step (1) of the training stage, the parameters of the collection device are obtained, and a collection result simulating an actual video camera is generated as training data, In step (2) of the training stage, the generated training data is used to train a neural network, The characteristics of the neural network include the following characteristics (2.1) to (2.5), Regarding the characteristic (2.1), the input of the neural network is the Lumitecel vector in k unstructured samplings, k is the number of samplings, and each value of Lumitecel represents the reflection intensity along a certain observation direction of the sampling point for the incident light from each light source. Lumitecel has a linear relationship with the luminous intensity of the light source, and simulation is performed with a linear fully connected layer, Regarding the characteristic (2.2), the first layer of the neural network includes a linear fully connected layer, which is used to simulate the illumination pattern actually used during collection and convert the k Lumitecels into camera collection results. These k collection results are combined with the pose information of the corresponding sampling points to form a high-dimensional point cloud, Regarding the characteristic (2.3), starting from the second layer is a feature extraction network, which independently extracts features from each point in the high-dimensional point cloud to obtain feature vectors. The formula of the feature extraction network is shown in the following formula 9, 【Number 9】 f is a one-dimensional convolution function, the size of the convolution kernel is 1×1, and B(I, P j ) represents the result output from the first-layer network or the measurement value collected and obtained, 【Number】 are respectively the spatial position of the sampling point, the geometric normal vector and the geometric tangent vector of the sampling point when sampling for the jth time, 【Number】 is obtained by a geometric model, 【Number】 is 【Number】 is an arbitrary unit vector orthogonal to, and by sampling the pose of the video camera for the jth time, 【Number】 is transformed to 【Number】 can be obtained, and V feature is the j-th sampled feature vector output from the network, Regarding the characteristic (2.4), after the feature extraction network, it is a max-pooling layer for aggregating the feature vectors extracted from k unstructured images to obtain a global feature vector, Regarding the above feature (2.5), after the max-pooling layer, it is a non-linear mapping network. The non-linear mapping network is skeletonized based on the global feature vector. The skeletonized representation of the non-linear mapping network is shown in Equation 11 below. 【Number 11】 f i+1 is the mapping function of the network of the (i + 1)-th layer, W i+1 is the parameter matrix of the network of the (i + 1)-th layer, b i+1 is the offset vector of the network of the (i + 1)-th layer, y i+1 is the output of the network of the (i + 1)-th layer, d and s respectively represent two branches of diffuse reflection and specular reflection, and the input y 1 d and y 1 s is the global feature vector output from the max pooling layer, The collection stage includes the following steps (1) and (2). In step (1) of the collection stage, it is a material collection step. The collection device irradiates sampling objects that are target three-dimensional objects in order according to the illumination pattern. The video camera acquires photos in a set of unstructured images. Using the photos as input, it acquires a geometric model with texture coordinates on the sampling object and the pose of the video camera during photography. In step (2) of the collection stage, it is a material recovery step. Based on the pose of the video camera during photography in the material collection step, it acquires the pose for each vertex of the effective texture coordinates on the sampling object when each photo is taken. Based on the collected photos and pose information, it constructs the high-dimensional point cloud as the input of the feature extraction network in the second layer of the neural network, and calculates and obtains a specular reflection feature vector and a diffuse reflection feature vector as the output vector of the final layer. The unstructured sampling is free random sampling with a non-fixed viewing angle. The sampling data is disordered, irregular, and unevenly distributed. It may be collected by holding the collection device by hand using a fixed object, or by placing the object on a rotating table and rotating it for collection by a fixing device. This is a free-style collection method for high-dimensional materials.
2. In the training data generation process, when the light source is colored, it is necessary to correct the spectral response relationship among the light source, the sampling object, and the video camera. Regarding the correction method, The spectral distribution curve of the unknown color light source L is 【Number】 is defined as, where λ represents wavelength, and c 1 represents one of the three channels of RGB, and the spectral distribution curve L(λ) of the light source with light intensity {I R , I G , I B} can be expressed by the following Equation 2: 【Number 2】 The reflection spectrum distribution curve p(λ) of any sampling point p is based on three unknown bases with coefficients p R , p G , p B respectively 【Number】 is shown as a linear combination, and c 2 represents one of the three channels of RGB, specifically, as shown in Equation 3 below, [Number 3] The spectral distribution curve of the video camera C is 【Number】 is shown as a linear combination, and under the irradiation of a light source with light intensities {I R , I G , I B}, the measured value at a specific channel c R with sampling points of reflection coefficients {p G , p B} by a video camera is shown in Equation 4 below, 3 【Number 4】 Illumination conditions {I R , I G , I B} = {1, 0, 0} / {0, 1, 0} / {0, 0, 1}, photograph a color test card with known reflection coefficients {p R , p G , p B}, construct a system of linear equations based on the measurement values collected by a video camera, obtain a color correction matrix δ(c 1 , c 2 , c 3 ) of size 3×3×3, and the method for freely collecting a high-dimensional material according to claim 1, characterized in that the color correction matrix shows the spectral response relationship of a light source, a sampling object, and a video camera.
3. In the feature (2.1) in the training stage, the observed value B in the photograph of the sampling point p on the object surface, the reflection function f r and the relationship between the light intensity of each light source are shown in the following formula 5, 【Number 5】 Here, I represents the light emission information of each light source l, and the spatial position x of the light source l l , the normal vector n of the light source l l , includes the light emission intensity I(l) of the light source l, P includes the parameter information of the sampling point p, and the spatial position x of the sampling point p , the material parameters n, t, α x , α y , ρ d , ρ s are included, 【Number】 explains the light intensity distribution in different incident directions of the light source l, and V is a binary function of visibility with respect to x l of x p shows the binary function of visibility with respect to x, 【Number】 is the dot product operation of two vectors, and f r (ω i ′; ω o ′, P) is a two-dimensional reflection function with respect to ω o ′ being constant, and is a function with respect to ω i ′, The input of the neural network is the Lumitexel vector in k unstructured samplings, denoted as m(l; P). 【Number 6】 In the above formula, B is the representation in a single channel. When the light source is a colored light source, B is extended to the following form. 【Number 7】 Here, f r (ω i ′; ω o ′, P, c 2 ) is f r (ω i ′; ω o ′, P) at 【Number】 The free-style collection method for high-dimensional materials according to claim 2, characterized in that it is the result of
4. The loss function of the neural network is designed as follows in steps (1) to (4), In step (1), one Lumitextel space is assumed, and the space is a cube whose center is at the spatial position x of the sampling point p of the cube, and the center coordinate system of the cube has its x-axis direction 【Number】 where the z-axis direction is 【Number】 and 【Number】 is the geometric normal vector, 【Number】 and 【Number】 is an arbitrary unit vector orthogonal to In step (2), one video camera is assumed, and the observation direction is the positive direction of the z-axis of the cube, In step (3), for diffused reflection Lumitecel, the resolution of the cube is 6 × N d 2 and for specular reflection Lumitecel, the resolution of the cube is 6 × N s 2 i.e., N d 2 , N s 2 points are uniformly sampled from each face as virtual point light sources with unit light intensity, and the following steps a to c are included In the step a, the specular reflectance ρ of the sampling point s is set to 0, and the diffuse reflection feature vector in this Lumitecel space 【Number】 is generated, In the step b, the diffuse reflectance ρ d is set to 0, and the specular reflection feature vector in this Lumitecel space 【Number】 is generated, In the step c, the output of the neural network is taken as a vector m d , m s , and m d is 【Number】 is the same as the length of, m s is 【Number】 is the same as the length, and the vector m d、 m s are respectively the diffuse reflection feature vectors 【Number】 , the specular reflection vector 【Number】 is the prediction of, In step (4), the loss function of the material feature part is shown in Equation 12 below, 【Number 12】 Here, λ d and λ s each represent the loss weights of m d , m s respectively, the confidence level β is used to evaluate the loss of specular reflection Lumitexel, and log acts on each dimension of the vector, The determination of the confidence level β is shown in Equation 13 below, 【Number 13】 where 【Number】 The term represents the logarithm of the maximum value of all single light source rendering values sampled at the j-th time, 【Number】 The term represents the logarithm of the maximum value of the single light source rendering value that can be theoretically obtained when sampling at the j-th time, 【Number】 A method for free-form collection of high-dimensional materials according to claim 1, characterized in that is a ratio adjustment coefficient.
5. In the material collection step, after the material collection is completed, geometric alignment is performed, and then material restoration is performed. Specifically, geometric alignment is to scan an object with a scanner to obtain a geometric model, and after aligning it with the geometric model reconstructed in three dimensions, replace the geometric model reconstructed in three dimensions. A method for free-form collection of high-dimensional materials according to claim 1, characterized by the above.
6. For valid texture coordinates, pixels in the photo are sequentially taken out based on the collected photo and the pose information of the sampling points, the validity of the pixels is judged, and a high-dimensional point cloud is constructed in combination with the pose for the vertex. For a point p on the surface of the sampled object determined for a certain valid texture coordinate, the judgment criteria for the j-th sampling to be valid for the vertex p include the following three conditions (1) to (3), In the above condition (1), the position x of the vertex p p j is visible to the video camera in the sampling, and x p j is within the sampling space defined when training the network. In the condition (2), and is a dot product operation, θ is the lower limit of the effective sampling angle, and ω o ′ indicates the direction of the emitted light in the world coordinate system, represents the normal vector of the j-th vertex p, In the condition (3), each channel value of the pixel in the photo is in the interval [a, b], and a and b are the lower and upper limits of the valid sampling luminance, When all of the above three conditions are satisfied, the j-th sampling is regarded as valid for the vertex p, and the j-th sampling result is added to the high-dimensional point cloud. A method for free-form collection of high-dimensional materials according to claim 1, characterized by the above.
7. Fitting the material parameters after recovering the material information includes the steps of fitting the local coordinate system and roughness and the step of fitting the reflectivity. In the step of fitting the local coordinate system and roughness, for a point p on the surface of the sampling object determined for a certain effective texture coordinate, the local coordinate system and roughness in the material parameters are fitted by the L-BFGS-B method based on the single-channel specular reflection vector output from the network. The method for free-form acquisition of a high-dimensional material according to claim 1, wherein in the step of fitting the reflectivity, the specular reflectivity and the diffuse reflectivity are obtained by a trust region algorithm, and the local coordinate system and roughness obtained in the previous process are fixed during the obtaining, and the observed values are synthesized at the viewing angles used for collection so as to be as close as possible to the observed values obtained by collection.
Citation Information
Patent Citations
Three-dimensional object normal vector, geometry and material acquisition method based on neural network
CN110570503A
Method for generating facial skin reflectance model implemented by computer
JP2006277748A
Image processing device, image processing method, and image processing program
JP2019219928A