A sparse-view-oriented local continuous light field construction method and system

CN118115933BActive Publication Date: 2026-09-29BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311870673.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2026-09-29
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

[0005]本发明技术解决问题:克服稀疏视角下光场构建困难、构建结果质量差的问题,提供一种面向稀疏视角的局部连续光场构建方法及系统,在稀疏视角的输入下构建局部光场,并使得构建的结果具有连续性,最终捕获目标场景的局部光场信息,提供不同视角、位置的准确的图像数据,为智能决策提供更多的数据支持

Benefits of technology

[0045](1)本发明相比对空间位置和方向进行编码,光场双平面结构化编码实现的任意图像光线特征参数化描述在构建场景隐式表达模型时只需进行一次查询,避免了多个采样点的查询使得模型消耗显存更少,收敛速度更快;在此基础上,采用极点特征Transformer网络补充任意图像极点特征避免了构建代价体的计算量并且增强了光线之间的几何一致性;设计场景隐式表达模型的整体损失函数综合考虑了颜色一致性与几何一致性,使得优化后的隐式表达模型输出目标光场的子孔径图像集合更加鲁棒;最后,对所述目标光场子孔径图像集合进行角域一致性优化,其中面向角域一致性的损失函数通过惩罚子孔径图像集合的颜色和EPI特征使得提取的目标光场局部连续,最终使得面向稀疏视角构建的局部连续光场可以提供不同视角、位置的准确的图像数据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118115933B_ABST
    Figure CN118115933B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of local continuous light field construction method and system for sparse view angle, the image set of sparse view angle is uniformly coded with light field double plane structure, the parametric description of arbitrary image light characteristic is realized, and the light characteristic of arbitrary image is obtained;With the premise of polar line constraint, the polar point feature of arbitrary image is extracted to the light characteristic of arbitrary image, and the polar point feature of arbitrary image is supplemented using polar point feature Transformer network, and the light enhancement feature of arbitrary image is obtained;The overall loss function of scene implicit expression model is designed, the scene implicit expression model is trained, and the sub-aperture image set of target light field is output;The sub-aperture image set of the target light field is optimized using pre-training angular domain consistency optimization model in angular domain, and the target light field with high angular domain consistency is extracted as local continuous light field, to complete the construction of local continuous light field for sparse view angle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method and system for constructing a local continuous light field for sparse viewpoints, belonging to the field of computer technology. Background Technology

[0002] Video surveillance equipment is a crucial foundation for smart cities. It provides the most basic data support for intelligent decision-making and safeguards public safety, including traffic safety and fire safety. However, video surveillance equipment cannot provide scene information from any angle or location, leading to delays or misjudgments in the intelligent decision-making process. Therefore, how to capture image data from different perspectives and locations with limited heterogeneous video equipment to provide more data support for intelligent decision-making presents both significant opportunities and challenges.

[0003] As a complete representation of the collection of light rays in space, a light field records both the direction and intensity of the light rays, containing rich information about the scene's structure. The regular encoding of geometric information in structured light field data significantly reduces the difficulty of scene information analysis, a key difference between light field data and other multi-view visual data. However, video surveillance equipment exhibits greater sparsity compared to traditional dense array sampling methods for light fields, making it difficult to directly support the dense viewing angle requirements of light field image processing.

[0004] In summary, existing light field construction methods mainly use dense viewpoints as input, which cannot handle the input of sparse viewpoints in real-world scenes, and cannot construct accurate local continuous light fields. Summary of the Invention

[0005] This invention addresses the problem of difficulty in constructing light fields and poor quality of construction results under sparse viewpoints. It provides a method and system for constructing local continuous light fields under sparse viewpoint input, constructing local light fields with continuity in the construction results, and ultimately capturing local light field information of the target scene. This provides accurate image data from different viewpoints and locations, offering more data support for intelligent decision-making.

[0006] Technical solution of the present invention:

[0007] In a first aspect, the present invention provides a method for constructing a locally continuous light field for sparse viewpoints, comprising:

[0008] Step 1: Perform unified optical field dual-plane structured encoding on the image set with sparse viewpoints to achieve parameterized description of the light features of any image and obtain the light features of any image; wherein, the sparse viewpoints are at least 3 and at most 10.

[0009] Step 2: Based on epipolar constraints, extract the epipolar features of any image from the ray features of any image, and use an epipolar feature Transformer network to supplement the epipolar features of any image to obtain the ray enhancement features of any image.

[0010] Step 3: Design the overall loss function of the scene implicit representation model, train the scene implicit representation model, and output a set of sub-aperture images of the target light field;

[0011] Step four: Angular domain consistency optimization is performed on the target light field sub-aperture image set, and the target light field with high angular domain consistency is extracted as a local continuous light field, thereby completing the construction of a local continuous light field for sparse viewpoints.

[0012] Specifically, step one is implemented as follows:

[0013] (1) Based on the center camera C of the light field o Given the camera baseline b and the camera intrinsic parameter matrix K, determine the two light field planes Π and Ω;

[0014] (2) Based on the camera extrinsic parameter matrix [R] in the light field o ,t o With camera intrinsic parameter matrix K o To obtain the structured encoding (s) of the central image in the light field biplane Π of the image set based on sparse viewpoints. * ,t * ,u * ,v * ), that is, the light characteristics of the central image;

[0015] (3) Transformation matrix from the center image to other images in an image set based on a sparse viewpoint Camera extrinsic matrix and camera intrinsic parameter matrix Obtain the light features of any image in the image set.

[0016] Specifically, step two is implemented as follows;

[0017] (1) Perform cross-view projection on the light features of any image, and extract the poles in other images of the image set that match all poles of any image in the image set using epipolar constraints to form the pole features of any image.

[0018] (2) An arbitrary image pole feature is obtained by using a pole feature Transformer network to supplement the pole features of the arbitrary image. The arbitrary image pole feature is then concatenated with the original features of the arbitrary image to obtain the arbitrary image light enhancement features. The pole feature Transformer consists of 8 blocks with an internal feature size of 256. Each block consists of a single-head self-attention layer. Each block is connected by residuals and normalized.

[0019] Specifically, step three is implemented as follows:

[0020] (1) The overall loss function of the designed scene implicit expression model is as follows:

[0021] l=l1+∈l2

[0022]

[0023]

[0024] l is the overall loss function, l1 is the color consistency loss function, and l2 is the geometric consistency loss function; L tar For pixel color true value, The color output by the scene implicitly represents the model; c j,k extreme point Color, β j,k Here, ∈ represents the extreme feature weights, and ∈ represents the weight coefficient, which is 0.5. The l1 loss function ensures the fidelity of the generated light features of any image by penalizing the color; the l2 loss function enhances geometric consistency by directly combining extreme color and extreme feature weights, thereby improving the credibility of the generated light enhancement features of any image.

[0025] (2) A geometric feature Transformer network is used to aggregate arbitrary image light features and arbitrary image light enhancement features, and input them into a multilayer perceptron to construct a scene implicit representation model. The overall loss function of the scene implicit representation model is then used as a penalty for training. The geometric feature Transformer consists of 8 blocks with an internal feature size of 256. Each block consists of a single-head self-attention layer and a multilayer perceptron. The multilayer perceptron uses Gaussian error linear units as activation layers. Each block uses residual connections and is normalized.

[0026] (3) Given the target light field parameters, output the set of sub-aperture images of the target light field through the trained scene implicit expression model.

[0027] Specifically, step four is implemented as follows:

[0028] (1) The loss function designed for corner-domain consistency is as follows:

[0029] l = l r +∈l e

[0030]

[0031]

[0032] l is the overall loss function, l r Let l be the color consistency loss function. e The loss function is the EPI feature function. For the target light field sub-aperture image set, SAIs gt For the set of target light field sub-aperture images, This represents an image slice along the y-axis and v-axis representing the target light field sub-aperture image. E represents an image slice of the target light field sub-aperture image along the x-axis and u-axis. y,v E represents the ground truth of the sub-aperture image of the target light field along the y-axis and v-axis. x,u This represents the ground truth of the image slices along the x-axis and u-axis of the target light field sub-aperture image, where ∈ is the weighting coefficient with a value of 0.3, l r The loss function ensures the spatial consistency of the target light field sub-aperture image set by penalizing the color of the target light field sub-aperture image set. e The loss function improves the angular domain consistency of the target light field sub-aperture image set by penalizing EPI features.

[0033] (2) Perform spatial-angular domain alternating convolution on the target light field sub-aperture image set to obtain continuous light field features that fuse spatial domain and angular domain information. Encode the continuous light field features at their positions and input them into the angular domain Transformer network. Optimize the loss function for angular domain consistency to train a pre-trained angular domain consistency optimization model. The angular domain Transformer consists of 8 blocks with an internal feature size of 512. Each block consists of an eight-head self-attention layer. Every four blocks are connected by residuals and normalized.

[0034] (3) Use a pre-trained angular domain consistency optimization model to extract the target light field with high angular domain consistency as a local continuous light field, thereby completing the construction of a local continuous light field for sparse viewpoints.

[0035] Secondly, the present invention provides a system for constructing a locally continuous light field for sparse viewpoints, comprising:

[0036] The encoding module performs unified optical field dual-plane structured encoding on the image set with sparse viewpoints, realizes parameterized description of ray features of arbitrary images, and obtains ray features of arbitrary images;

[0037] The fusion module, based on epipolar constraints, extracts epipolar features from the ray features of any image and supplements the epipolar features of any image with an epipolar feature Transformer network to obtain the ray enhancement features of any image.

[0038] The scene implicit representation module designs the overall loss function of the scene implicit representation model, trains the scene implicit representation model, and outputs a set of sub-aperture images of the target light field.

[0039] The optimization module performs angular domain consistency optimization on the target light field sub-aperture image set, extracts the target light field with high angular domain consistency as a local continuous light field, thereby completing the construction of a local continuous light field for sparse viewpoints.

[0040] Thirdly, the present invention provides an electronic device, comprising a processor and a memory, wherein:

[0041] Memory, used to store computer programs;

[0042] A processor is used to execute computer programs stored in memory, and in doing so, implements the aforementioned methods or systems.

[0043] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed, it implements the aforementioned method or the aforementioned system.

[0044] The advantages of this invention compared to the prior art are:

[0045] (1) Compared with encoding spatial position and direction, the parameterized description of arbitrary image light features realized by the dual-plane structured encoding of the light field only needs to be queried once when constructing the scene implicit expression model, avoiding the query of multiple sampling points, so that the model consumes less memory and converges faster. On this basis, the pole feature Transformer network is used to supplement the pole features of arbitrary images to avoid the computational burden of constructing the cost volume and enhance the geometric consistency between light rays. The overall loss function of the scene implicit expression model is designed to comprehensively consider color consistency and geometric consistency, so that the sub-aperture image set of the target light field output by the optimized implicit expression model is more robust. Finally, the target light field sub-aperture image set is optimized for angular domain consistency. The loss function for angular domain consistency penalizes the color and EPI features of the sub-aperture image set to make the extracted target light field locally continuous. Finally, the locally continuous light field constructed for sparse viewpoints can provide accurate image data for different viewpoints and positions.

[0046] (2) This invention uses a central camera C of the light field oThe camera baseline b and camera intrinsic parameter matrix K are used to determine the light field biplanes Π and Ω. The four-dimensional light field describes the ray characteristics of any image. The ray parameters are directly mapped to the integral radiation along the ray. The process of querying hundreds of times in the typical scene implicit representation model is reduced to one time, which reduces the memory consumption and improves the convergence speed of the scene implicit representation model, thereby providing more accurate image data.

[0047] (3) Based on epipolar constraints, this invention extracts arbitrary image epipolar features from arbitrary image light features and obtains arbitrary image epipolar supplementary features by supplementing arbitrary image epipolar features through a designed epipolar feature Transformer network. This avoids the computational burden of constructing a cost volume and makes the obtained arbitrary image light enhancement features geometrically consistent.

[0048] (4) The present invention uses a geometric feature Transformer network to aggregate arbitrary image light features and arbitrary image light enhancement features to construct a scene implicit representation model. The overall loss function of the designed scene implicit representation model takes into account both color consistency and geometric consistency, making the target light field sub-aperture image set output by the optimized scene implicit representation model more robust.

[0049] (5) The continuous light field features extracted by the spatial-angular domain alternating convolution of the target light field sub-aperture image set in this invention take into account the spatial resolution and angular resolution information of the sub-aperture image set. After the continuous light field features are positionally encoded, they are input into the angular domain Transformer network for training to obtain a pre-trained angular domain consistency optimization model. The designed loss function for angular domain consistency enhances the angular domain consistency of the light field sub-aperture image set. Attached Figure Description

[0050] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 A flowchart illustrating the implementation of a method for constructing a local continuous light field for sparse viewpoints, as provided in an embodiment of the present invention.

[0052] Figure 2 This is a schematic diagram of the dual-plane structured coding of the optical field in an embodiment of the present invention;

[0053] Figure 3 This is a schematic diagram of cross-view projection based on epipolar constraints in an embodiment of the present invention;

[0054] Figure 4 This is a diagram of the Transformer network structure for pole features in an embodiment of the present invention;

[0055] Figure 5 This is a diagram of the geometric feature Transformer network structure in an embodiment of the present invention;

[0056] Figure 6 This is a diagram of the structure of the pre-trained corner domain uniformity model in an embodiment of the present invention. Detailed Implementation

[0057] To gain a more detailed understanding of the objectives, technical solutions, and advantages of this invention, the implementation of the embodiments of this invention will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration only and are not intended to limit the embodiments of this invention.

[0058] To clearly illustrate the design concept of this invention, the invention will be described in detail below with reference to embodiments.

[0059] like Figure 1 As shown, this embodiment of the invention provides a method for constructing a local continuous light field for sparse viewpoints, including:

[0060] Step 1: Perform unified light field dual-plane structured encoding on the image set with sparse viewpoints to achieve parameterized description of light ray features of any image and obtain the light ray features of any image.

[0061] Optionally, step one above is implemented as follows:

[0062] like Figure 2 As shown, O is the origin of the spherical coordinate system, I o For the centered image, C o As the center camera of the light field, C j1 and C j2 For any two images and The camera, Π(u,v) is the angular domain plane, Ω(s,t) is the spatial plane, (s * ,t * ,u * ,v * ) for I o any pixel x * Structured encoding of Π and Ω. Based on the central camera C of the light field. o The camera baseline b and the camera intrinsic parameter matrix K determine the two light field planes Π and Ω; for the center image I o any pixel x * Combined with the camera extrinsic matrix [R] in the light field o , t o With camera intrinsic parameter matrix K o Obtain the central image Io any pixel x * In the structured coding of optical field biplanes Π and Ω (s * , t * u * v * The central image light characteristics are shown in formula (1):

[0063] (s * , t * u * v * )=Π∪Ω∩(-R o T t o +R o T K o -1 x * (1)

[0064] For any image Pixels Image light characteristics Combine any image To the center image I o Transformation matrix With camera extrinsic matrix and camera intrinsic parameter matrix Solve for any image Pixels Structured encoding of corresponding rays in the two planes Π and Ω of the light field That is, the light characteristics of any image.

[0065] Step 2: Based on the epipolar constraint, extract the epipolar features of any image from the light features of any image, and use the epipolar feature Transformer (transform neural network) to supplement the epipolar features of any image to obtain the light enhancement features of any image.

[0066] Optionally, step two above can be implemented as follows:

[0067] First, for any image light features Perform cross-view projection. For example... Figure 3 As shown, I o For the centered image, C o As the center camera of the light field, C j1 and C j2 For any two images and A light field camera, for any two images and Pick medium pixel epipolar constraints on the image Light characteristics Selecting the pole according to Camera extrinsic matrix and camera intrinsic parameter matrix and For the central image I o Transformation matrix and Projecting poles onto the image Get pixels by Record r j1 The upper pole is at The set of matched pixels is referred to as pixels in this embodiment of the invention. exist The set of poles. Repeat the above steps to extract pixels. The set of poles in any image That is, pixels Arbitrary image pole features based on epipolar constraints.

[0068] Then, an extremum feature Transformer network is used to supplement the extremum features of any image to obtain supplemented extremum features for any image. pixels Define pixel The initial characteristics are:

[0069]

[0070] in, for At pixel position The visual features are extracted by a simple convolutional network. For pixels The color. For pixels. exist The extreme point in Define poles The initial characteristics are:

[0071]

[0072] in, For pixels The light characteristics, for Up projection to World coordinates For pixels Visual features For pixels The color.

[0073] Pixels initial features With pixels exist The extreme point in initial features The result of concatenating the input to the extreme feature Transformer network is:

[0074]

[0075] Where 1≤k≤N,j m ≠j1

[0076] Will With all Separately stitch together and extract pixels With pixels exist The extreme point in Combined weights The higher the weight, the higher the pixel With the extreme point The higher the correlation between the two:

[0077]

[0078] Where W1 is the weight matrix.

[0079] The image is obtained by calculating the weighted average of the output sequence of the Transformer with extreme features. supplementary features

[0080]

[0081] in, The combined weights are calculated using formula (5).

[0082] any image pixels initial features With any image Pole complement features By stitching together, an image is obtained. Arbitrary image ray enhancement features:

[0083]

[0084] like Figure 4 As shown, pixels initial features and pixels exist The extreme point in initial features After concatenation, the data is input into the extreme feature Transformer network to obtain... and Will With all The pixels are obtained by concatenating the inputs separately and feeding them into a multilayer perceptron. With pixels exist The extreme point in Combined weights calculate and The weighted average of the images is obtained supplementary features Finally, the pixels initial features With images Pole complement features By stitching together, an image is obtained. Arbitrary image light enhancement features.

[0085] Step 3: Design the overall loss function of the scene implicit representation model, train the scene implicit representation model, and output a set of sub-aperture images of the target light field;

[0086] Optionally, in step three above:

[0087] First, a geometric feature Transformer network is used to aggregate arbitrary image ray features and arbitrary image ray enhancement features, such as... Figure 5 As shown.

[0088] The result is obtained by concatenating arbitrary image ray features with arbitrary image ray enhancement features:

[0089]

[0090] in, For any image light features, N is the number of samples. Let u be the set of poles that any image matches in any other image. tar v is the index value of any image. tar The index value is any other image.

[0091] The features of the above formula (8) are input into the geometric feature Transformer network to obtain the output sequence:

[0092]

[0093] Where 1 ≤ k ≤ N. The arbitrary ray feature z is obtained by calculating the correlation between the arbitrary image ray feature and the arbitrary image ray enhancement feature. tar:

[0094]

[0095] Where, β j,k The combined weights for light enhancement features are calculated as follows:

[0096]

[0097] in, and is the output sequence obtained from the geometric feature Transformer network. W2 is the weight matrix.

[0098] Then, the light feature z tar The input is converted into pixel color by a multilayer perceptron. Construct an implicit representation model of the scenario.

[0099] Finally, the overall loss function of the scene implicit expression model is used as a penalty for training;

[0100] The overall loss function of the defined scene implicit representation model is as follows:

[0101]

[0102] Where l is the overall loss function, l1 is the color consistency loss function, and l2 is the geometric consistency loss function; L tar For pixel color true value, The color output by the scene implicitly represents the model; c j,k extreme point Color, β j,k ∈ represents the extreme feature weights, where ∈ is the weight coefficient set to 0.5. The l1 loss function ensures the fidelity of the generated light features of arbitrary images by penalizing the color. The l2 loss function strengthens the network's attention to directly related extremes by directly combining extreme colors and extreme feature weights, reduces the entropy of extreme feature weights, enhances geometric consistency, and thus improves the credibility of the generated light enhancement features of arbitrary images.

[0103] Define the angular coordinates of the target light field center as (u,v)=(0,0). For a target structured light field with angular resolution of U×V, given the target light field parameters, the sub-aperture image set L0 of the target light field is output through the trained scene implicit representation model.

[0104] L0=L0(s,t,u,v) (13)

[0105] in,

[0106] Step 4: Use a pre-trained angular domain consistency optimization model to perform angular domain consistency optimization on the target light field sub-aperture image set, and extract the target light field with high angular domain consistency as a local continuous light field, thereby completing the construction of a local continuous light field for sparse viewpoints.

[0107] Optionally, step four above is implemented as follows:

[0108] First, extract the target light field sub-aperture image set to initialize the feature F. (u,v) , H and W are the height and width of the sub-aperture image, respectively, and n1 is the number of convolution kernels.

[0109] The initial features are concatenated along the angular domain dimension to obtain:

[0110]

[0111] in,

[0112] For F cat By performing alternating spatial-angular domain convolutions, a continuous light field feature F that fuses spatial and angular domain information is obtained. spa-ang For the continuous light field characteristic F spa-ang Position encoding is performed to obtain P F =P(F spa-ang The position encoding function P(·) is defined as follows:

[0113]

[0114]

[0115] Where x = {1, 2, ..., N}, L is the position coding function of the reference frequency, and P k (x) is the code for the k-th frequency, w k (α) is the encoding weight function, α={0,1,…,L}.

[0116] The position-encoded continuous light field features are input into the angular domain Transformer network to obtain:

[0117] T k =[H1,H2,…,H k ,…,H N W H (17)

[0118] The corner domain Transformer consists of 8 blocks with an internal feature size of 512. Each block is composed of an eight-head self-attention layer, and every four blocks are connected by residuals and normalized. H For the weight matrix, each attention head Hk The calculation method is as follows:

[0119]

[0120] Where K = {1, 2, ..., N}, N = 8, softmax is the activation function, and W K W Q W V These are the weight parameter matrices for matrices K, Q, and V, respectively. V = LN(F) spa-ang ), For concatenation operations, P F For continuous light field characteristics F spa-ang The positional encoding, where LN is the normalization function;

[0121] Then, the loss function for corner-domain consistency is optimized, and a pre-trained corner-domain consistency optimization model is trained. The loss function for corner-domain consistency is as follows:

[0122]

[0123]

[0124] Where l is the overall loss function, l r Let l be the color consistency loss function. e Here, ||.||1 represents the EPI feature loss function; ||.||1 represents the L1 distance. For the target light field sub-aperture image set, SAIs gt For the set of target light field sub-aperture images, This represents an image slice along the y-axis and v-axis representing the target light field sub-aperture image. E represents an image slice of the target light field sub-aperture image along the x-axis and u-axis. y,v E represents the ground truth of the sub-aperture image of the target light field along the y-axis and v-axis. x,u This represents the ground truth of the image slices along the x-axis and u-axis of the target light field sub-aperture image. For the gradient operator in the x-axis direction, The gradient operator is located in the y-axis direction. For the gradient operator in the u-axis direction, Let be the gradient operator along the v-axis, ∈ be the weight coefficient (0.3), and l be the gradient operator. r The loss function ensures the spatial consistency of the target light field sub-aperture image set by penalizing the color of the target light field sub-aperture image set. e The loss function improves the angular domain consistency of the target light field sub-aperture image set by penalizing EPI features.

[0125] Finally, the pre-trained corner domain consistency optimization model F is used. A Extract the target light field L1(s,t,u,v) = F from the sub-aperture image set L0 of the target light field, which has high angular domain consistency. A (L0) serves as a locally continuous light field, completing the construction of a locally continuous light field oriented towards a sparse viewpoint. and U×V represents the angular domain resolution, (u, v) represents the coordinates within the angular domain plane Π, and (s, t) represents the coordinates within the spatial plane Ω.

[0126] like Figure 6 As shown, using the set of sub-aperture images L0 of the target light field as input, the initialization features F of the target light field sub-aperture image set are obtained after initialization feature extraction. (u,v) , will F (u,v) F is obtained by splicing along the angular domain dimension. cat And for F cat Performing eight alternating spatial-angular domain convolutions yields a continuous light field feature F that fuses spatial and angular domain information. spa-ang ; for continuous light field characteristics F spa-ang Position encoding is performed to obtain P F =P(F spa-ang ), and P F The input is fed into eight angular domain Transformer networks, and finally, after an upsampling operation, a local continuous optical field L1(s,t,u,v) is obtained. Each angular domain Transformer first encodes the position of the normalization operation P. F The input is fed into three linear layers to obtain K, Q, and V matrices. The K and Q matrices are multiplied and scaled, and after passing through the activation layer, they are multiplied with the V matrix. Finally, the input passes through concatenation, linear layers, matrix addition, feedforward network, and matrix addition again.

[0127] Optionally, the present invention also provides an electronic device (computer, server, smartphone, network device, etc.) including a memory and a processor, wherein the memory is used to store a computer program executable by the processor, and when the processor executes the computer program, it implements the above embodiments.

[0128] Optionally, the present invention also provides a program product, such as a computer-readable storage medium, including a program that, when executed by a processor, is used to perform the above embodiments.

[0129] To demonstrate the advantages of this invention, peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and learned perceptual patch similarity (LPIPS) were used to compare it with Neural Radiation Field (NeRF) on the Real Forward Scene Dataset (RFF), the Visual Dependency Challenge Dataset (Shiny), and the Synthetic Dataset (Blender). The RFF dataset includes eight real forward scenes captured using a smartphone; the Shiny dataset includes eight real forward scenes with visual dependency challenge effects; and the Blender dataset includes eight synthetic datasets randomly sampled from the hemisphere surrounding the object. In the RFF, Shiny, and Blender datasets, images were selected as the test set with a stride of 4, and the remaining images were used as the training set. Table 1 shows the comparison of PSNR, SSIM, and LPIPS with NeRF on the publicly available datasets RFF, Shiny, and Blender.

[0130] Compared to NeRF, the scene implicit representation model constructed in this invention improves PSNR by an average of 1.51, SSIM by an average of 2.7%, and LPIPS by an average of 11.5%. This demonstrates that this invention can provide completely accurate image data.

[0131] Table 1

[0132]

[0133] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for constructing a locally continuous light field for sparse viewpoints, characterized in that, include: A unified optical field dual-plane structured encoding is performed on a set of images with sparse viewpoints to achieve parameterized description of ray features in any image, thus obtaining the ray features of any image; Based on epipolar constraints, arbitrary image epipolar features are extracted from arbitrary image light features, and an epipolar feature Transformer network is used to supplement arbitrary image epipolar features to obtain arbitrary image light enhancement features; Design the overall loss function of the scene implicit representation model, train the scene implicit representation model, and output a set of sub-aperture images of the target light field; The target light field sub-aperture image set is optimized for angular domain consistency using a pre-trained angular domain consistency optimization model. The target light field with high angular domain consistency is extracted as a local continuous light field, thereby completing the construction of a local continuous light field for sparse viewpoints. The overall loss function of the implicit representation model for the designed scene, the training of the implicit representation model for the scene, and the output of the sub-aperture image set of the target light field are specifically implemented as follows: The overall loss function of the designed scene implicit representation model is as follows: For the overall loss function, Let the color consistency loss function be... The geometric consistency loss function; For pixel color true value, The color output by the scene implicitly expresses the model; extreme point color, For extreme feature weights, These are the weighting coefficients. The loss function ensures the fidelity of light features in the generated image by penalizing color. The loss function enhances geometric consistency by directly combining extremum color and extremum feature weights, thereby improving the credibility of generating arbitrary image light enhancement features; A geometric feature Transformer network is used to aggregate arbitrary image lighting features and arbitrary image lighting enhancement features, and then input them into a multilayer perceptron to construct a scene implicit representation model. The overall loss function of the scene implicit representation model is then used as a penalty for training. Given the target light field parameters, the trained scene implicit representation model outputs a set of sub-aperture images of the target light field.

2. The method for constructing a locally continuous light field for sparse viewpoints according to claim 1, characterized in that, The method of uniformly encoding the light field in a dual-plane structured manner for the image set with sparse viewpoints enables parameterized description of the light features of any image, and the specific implementation of obtaining the light features of any image is as follows: The two planes of the light field are determined based on the center camera, camera baseline, and camera intrinsic parameter matrix of the light field. Based on the extrinsic and intrinsic parameters of the light field camera, the structured encoding of the central image in the dual planes of the light field in the image set based on sparse viewpoints is obtained. That is, the light characteristics of the central image; Based on the transformation matrix from the center image to other images in the sparse viewpoint image set, the camera extrinsic matrix, and the camera intrinsic matrix, the light characteristics of any image in the image set are obtained.

3. The method for constructing a locally continuous light field for sparse viewpoints according to claim 1, characterized in that, The above-mentioned method, which extracts arbitrary image polar features based on epipolar constraints and supplements arbitrary image polar features with an epipolar feature Transformer network to obtain arbitrary image light enhancement features, is specifically implemented as follows: The light features of any image are projected across the viewpoint, and the poles that match the poles of any image in the image set in other images in the image set are extracted by epipolar constraints to form the pole features of any image. An arbitrary image pole feature supplement is obtained by using a pole feature Transformer network to supplement the pole features of the arbitrary image. The arbitrary image pole feature supplement is then concatenated with the original features of the arbitrary image to obtain the arbitrary image light enhancement feature.

4. The method for constructing a locally continuous light field for sparse viewpoints according to claim 1, characterized in that, The specific implementation of using a pre-trained angular domain consistency optimization model to optimize the angular domain consistency of the target light field sub-aperture image set, and extracting the target light field with high angular domain consistency as a local continuous light field, thereby completing the construction of a local continuous light field for sparse viewpoints is as follows: The loss function designed for corner-domain consistency is as follows: For the overall loss function, Let the color consistency loss function be... The loss function is the EPI feature function. For the target light field sub-aperture image set, For the set of target light field sub-aperture images, This represents an image slice along the y-axis and v-axis representing the target light field sub-aperture image. This represents an image slice along the x-axis and u-axis representing the target light field sub-aperture image. This represents the ground truth of the image slices along the y-axis and v-axis of the target light field sub-aperture image. This represents the ground truth of the image slices along the x-axis and u-axis of the target light field sub-aperture image. These are the weighting coefficients. The loss function ensures the spatial consistency of the target light field sub-aperture image set by penalizing the color of the target light field sub-aperture image set. The loss function improves the angular domain consistency of the target light field sub-aperture image set by penalizing EPI features; For the gradient operator in the x-axis direction, The gradient operator is located in the y-axis direction. For the gradient operator in the u-axis direction, The gradient operator is located in the v-axis direction. The target light field sub-aperture image set is subjected to alternating spatial-angular domain convolution to obtain continuous light field features that fuse spatial and angular domain information. The continuous light field features are then position-encoded and input into an angular domain Transformer network. By optimizing the loss function oriented towards angular domain consistency, a pre-trained angular domain consistency optimization model is obtained. A pre-trained angular domain consistency optimization model is used to extract the target light field with high angular domain consistency as a local continuous light field, thereby completing the construction of a local continuous light field for sparse viewpoints.

5. A system for constructing a locally continuous light field for sparse viewpoints, characterized in that, include: The encoding module performs unified optical field dual-plane structured encoding on the image set with sparse viewpoints, realizes parameterized description of ray features of arbitrary images, and obtains ray features of arbitrary images; The fusion module, based on epipolar constraints, extracts epipolar features from the ray features of any image and supplements the epipolar features of any image with an epipolar feature Transformer network to obtain the ray enhancement features of any image. The scene implicit representation module designs the overall loss function of the scene implicit representation model, trains the scene implicit representation model, and outputs a set of sub-aperture images of the target light field. The optimization module performs angular domain consistency optimization on the target light field sub-aperture image set, extracts the target light field with high angular domain consistency as a local continuous light field, thereby completing the construction of a local continuous light field for sparse viewpoints; The overall loss function of the implicit representation model for the designed scene, the training of the implicit representation model for the scene, and the output of the sub-aperture image set of the target light field are specifically implemented as follows: The overall loss function of the designed scene implicit representation model is as follows: For the overall loss function, Let the color consistency loss function be... The geometric consistency loss function; For pixel color true value, The color output by the scene implicitly expresses the model; extreme point color, For extreme feature weights, These are the weighting coefficients. The loss function ensures the fidelity of light features in the generated image by penalizing color. The loss function enhances geometric consistency by directly combining extremum color and extremum feature weights, thereby improving the credibility of generating arbitrary image light enhancement features; A geometric feature Transformer network is used to aggregate arbitrary image lighting features and arbitrary image lighting enhancement features, and then input them into a multilayer perceptron to construct a scene implicit representation model. The overall loss function of the scene implicit representation model is then used as a penalty for training. Given the target light field parameters, the trained scene implicit representation model outputs a set of sub-aperture images of the target light field.

6. An electronic device, characterized in that, Includes processor and memory, of which: Memory, used to store computer programs; A processor for executing a computer program stored in memory, which, when executed, implements the method described in any one of claims 1-4, or the system described in claim 5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the method described in any one of claims 1-4, or the system described in claim 5.

Citation Information

Patent Citations

  • Maneuvering track front side view synthetic aperture radar tomography method

    CN110146884A

  • Method and apparatus for wide field distortion-compensated imaging

    US5448053A