Unsupervised learning of light field disparity estimation for wetland scenes

By utilizing an unsupervised learning-based light field disparity estimation device, which leverages implicit neural representation and light field geometric consistency, the problem of scarce labeled data and dynamic interference in disparity estimation in wetland scenes is solved, achieving efficient disparity estimation.

CN120852497BActive Publication Date: 2026-03-03BEIJING INFORMATION SCI & TECH UNIV +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Estimating the parallax of light field in wetland scenes faces challenges such as scarce labeled data, dynamic interference, and multimodal light interaction, making it difficult for traditional methods to accurately estimate parallax.

Method used

An unsupervised learning-based light field disparity estimation device includes a disparity estimation network, a view generation network, and an occlusion mask generation module. It utilizes implicit neural expression networks and light field geometric consistency to generate occlusion masks to eliminate occlusion errors. It also calculates feature matching costs and generates disparity maps using a multi-view disparity estimation method.

Benefits of technology

It reduces annotation costs, improves robustness to dynamic scenes, accurately masks occlusion errors, adapts to complex lighting interactions, and provides reliable and robust light field parallax estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852497B_ABST
    Figure CN120852497B_ABST
Patent Text Reader

Abstract

This invention discloses an unsupervised learning light field disparity estimation device and method for wetland scenes, comprising: a disparity estimation network for receiving five-dimensional light field data, extracting multi-view spatiotemporal features, and outputting a disparity map corresponding to the central view; a view generation network employing an implicit neural network to transform the normalized coordinates of the central view based on the disparity map, generating a non-central view; and an occlusion mask generation module for estimating the occlusion relationship between the central view and the non-central view based on the disparity map, generating a final occlusion mask to eliminate occlusion errors. This invention achieves low annotation costs, improves robustness to dynamic scenes, and optimizes the accuracy of ray modeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and machine learning technology, and in particular to an unsupervised learning light field parallax estimation device and method for wetland scenes. Background Technology

[0002] In the research and application of wetland ecosystems, optical field parallax estimation technology is of great significance. However, the unique and complex ecological structure of wetland scenes brings many challenges to optical field parallax estimation.

[0003] Wetland scenes encompass a rich variety of elements, including water bodies, vegetation, and mudflats. Their light field data is influenced by a combination of factors, such as reflection, refraction, and dynamic environmental conditions. In this context, traditional disparity estimation methods, such as supervised methods based on stereo matching or deep learning, face insurmountable challenges.

[0004] (1) Scarcity of labeled data: Labeling the light field data of wetland scenes is not only costly, but also difficult to fully cover these situations due to complex optical effects such as water surface ripples and semi-transparent vegetation in wetlands, resulting in extremely scarce available labeled data.

[0005] (2) Dynamic interference: Dynamic factors such as the continuous flow of water and the swaying of vegetation in the wind can cause inconsistencies in the light field in the time domain. Traditional methods are prone to error accumulation when dealing with such dynamic changes, which in turn seriously affects the accuracy of parallax estimation.

[0006] (3) Multimodal light interaction: The light field in wetland scenes is affected by various factors such as water surface reflection and scattering by suspended particles, resulting in extremely complex light propagation paths. Traditional methods are difficult to effectively model such complex light propagation paths, which greatly reduces the accuracy and reliability of parallax estimation. Summary of the Invention

[0007] The purpose of this invention is to provide an unsupervised learning optical field parallax estimation device and method for wetland scenes, so as to overcome or at least mitigate at least one of the above-mentioned defects of the prior art.

[0008] To achieve the above objectives, the present invention provides an unsupervised learning optical field disparity estimation device, comprising:

[0009] A parallax estimation network is used to receive five-dimensional light field data L∈R. B×C×N×H×W The system extracts multi-view spatiotemporal features and outputs the disparity map D(x, y) corresponding to the central view V. The light field data includes light field views with height H and width W, which are divided into the central view V and non-central views V. iB represents the batch number, C represents the number of channels, N represents the number of viewpoints in the light field data, i represents the sequence number of the non-center view, and the value is a natural number from 0 to N×N. (x, y) represents the pixel coordinates of the center view.

[0010] The view generation network employs an implicit neural network to transform the normalized coordinates of the central view based on the disparity map D(x,y) to generate a non-central view V. i ;

[0011] An occlusion mask generation module is used to estimate the center view V and the non-center view V based on the disparity map D(x,y). i The occlusion relationship is used to generate the final occlusion mask to eliminate occlusion errors;

[0012] Among them, the disparity estimation network, the view generation network, and the occlusion mask generation module are jointly optimized to complete the light field disparity estimation in an unsupervised learning manner.

[0013] Furthermore, the disparity estimation network includes:

[0014] The feature extraction unit, which includes a 3D convolutional layer, a BN layer, and a LeakyReLU activation function, is used to extract features of the light field view in the light field data and perform normalization processing.

[0015] The view selection unit is used to extract features from the normalized light field view, adjust the dimensions of the features, and represent the adjusted light field view features as Feature∈R. B×N×C×H×W This allows us to obtain the attention weights for each viewpoint of the light field data, and adjust the features of each light field view based on the attention weights for each viewpoint.

[0016] The matching cost unit is used to calculate the feature matching cost using a multi-view disparity estimation method.

[0017] The disparity regression unit is used to generate a disparity map D(x, y) based on the matching cost.

[0018] Furthermore, the view generation network includes:

[0019] The image offset calculation unit calculates the horizontal and vertical pixel offsets ΔX(x,y) and ΔY(x,y) of the non-center view relative to the center view, as well as the viewing angle offsets dv and du, based on the disparity map D(x,y).

[0020]

[0021] The coordinate transformation unit is used to convert the normalized coordinates of the center view into the coordinates of the non-center view based on the offset.

[0022] The disparity map generation unit is used to describe the center view as... Using the implicit neural network function C Described as Equation (3):

[0023]

[0024] Furthermore, the view generation network also includes:

[0025] The symmetrical viewpoint projection transformation unit converts the output of the disparity map generation unit into a symmetrical viewpoint projection transformation unit. Mapping to a symmetrical viewpoint (Ns, Nt), and aligning pixels using bilinear interpolation, a non-centered view V is generated. i .

[0026] Furthermore, the occlusion mask generation module includes:

[0027] The initial occlusion mask generation unit has two preset occlusion scenarios: the center view is occluded but not visible, and the center view is visible but not occluded.

[0028] The final occlusion mask generation unit performs an "OR" operation on the two types of initial occlusion masks to generate a final mask containing the two types of occlusion regions;

[0029] The final mask is used to adjust for view generation errors and light field geometry features;

[0030] The final masks for the two types of occlusion are denoted as M. i (x,y) and G i (x s,i ,y s,i ):

[0031]

[0032] Among them, H i To obtain the non-center view V after traversing all pixel coordinates (x, y) of the center view V. i The intersection of the visible area with the central view V is shown in equation (6), C i (x s,i ,y s,i (V) is a non-center view i pixel coordinates (x) s,i ,y s,i The coordinates of distance in the coordinate system of the central view V. max C i (x s,i ,y s,i The maximum distance between any two pixel coordinates in N(Warp) is N(Warp) i (x c ,y c ),γ) is based on Warpi (x c ,y c A square region centered at γ and with sides of length 2γ+1, where γ is used to control the center Warp. i (x c ,y c The number of pixels extending outwards, β is a parameter used to control the sensitivity of ghost detection, (x c ,y c (V) is a non-center view i Pixel coordinates:

[0033]

[0034] Furthermore, the network's loss function Loss total Described as Equation (11):

[0035] Loss total =λLoss L2 +(1-λ)Loss smooth (11)

[0036] In the formula, Loss L2 Mean squared error loss is used to measure the pixel difference between the generated view and the real view. smooth To smooth the loss, the gradient smoothness of the disparity map is optimized, and λ is used to control the loss. L2 and Loss smooth The weight.

[0037] This invention also provides an unsupervised learning method for estimating light field disparity, comprising:

[0038] Step 1: Receive five-dimensional light field data L∈R B×C×N×H×W The system extracts multi-view spatiotemporal features and outputs the disparity map D(x, y) corresponding to the central view V. The light field data includes light field views with height H and width W, which are divided into the central view V and non-central views V. i B represents the batch number, C represents the number of channels, N represents the number of viewpoints in the light field data, i represents the sequence number of the non-center view, and the value is a natural number from 0 to N×N. (x, y) represents the pixel coordinates of the center view.

[0039] Step 2: Using a hidden neural network, the normalized coordinates of the central view are transformed based on the disparity map D(x, y) to generate the non-central view V. i ;

[0040] Step 3: Estimate the central view Vx and non-central view Vy based on the disparity map D(x,y). i The occlusion relationship is used to generate the final occlusion mask to eliminate occlusion errors;

[0041] Step 4: Jointly optimize the outputs of Step 1, Step 2 and Step 3 to complete the light field disparity estimation in an unsupervised learning manner.

[0042] Further, step 1 includes:

[0043] Step 11: Extract the features of the light field view from the light field data and normalize them;

[0044] Step 12: Extract the features of the normalized light field view, adjust the dimensions of the features of the light field view, and represent the adjusted features of the light field view as Feature∈R B×N×C×H×W This allows us to obtain the attention weights for each viewpoint of the light field data, and adjust the features of each light field view based on the attention weights for each viewpoint.

[0045] Step 13: Calculate the feature matching cost using the multi-view parallax estimation method;

[0046] Step 14: Generate disparity map D(x, y) based on matching cost.

[0047] Further, step 2 includes:

[0048] Step 21: Calculate the horizontal and vertical pixel offsets ΔX(x,y) and ΔY(x,y) of the non-center view relative to the center view, as well as the viewing angle offsets dv and du, based on the disparity map D(x,y).

[0049]

[0050] Step 22: Convert the normalized coordinates of the center view to the coordinates of the non-center view based on the offset;

[0051] Step 23, describe the center view as Using the implicit neural network function C Described as Equation (3):

[0052]

[0053] Furthermore, step 3 includes:

[0054] Step 31: Preset two types of occlusion: the center view is occluded and not visible, and the center view is visible but not occluded.

[0055] Step 32: Perform an "OR" operation on the two initial occlusion masks to generate a final mask containing the two types of occlusion regions;

[0056] The final mask is used to adjust for view generation errors and light field geometry features;

[0057] The final masks for the two types of occlusion are denoted as M. i (x,y) and G i (x s,i ,y s,i ):

[0058]

[0059] Among them, H i To obtain the non-center view V after traversing all pixel coordinates (x, y) of the center view V. i The intersection of the visible area with the central view V is shown in equation (6), C i (x s,i ,y s,i (V) is a non-center view i pixel coordinates (x) s,i ,y s,i The coordinates of distance in the coordinate system of the central view V. max C i (x s,i ,y s,i The maximum distance between any two pixel coordinates in N(Warp) is N(Warp) i (x c ,y c ),γ) is based on Warp i (x c ,y c A square region centered at γ and with sides of length 2γ+1, where γ is used to control the center Warp. i (x c ,y c The number of pixels extending outwards, β is a parameter used to control the sensitivity of ghost detection, (x c ,y c (V) is a non-center view i Pixel coordinates:

[0060]

[0061] The present invention has the following advantages due to the adoption of the above technical solutions:

[0062] 1. Since this invention uses optical field geometric consistency (reprojection error, epipolar constraint) to replace manual annotation and solves the problem of data scarcity, it can reduce annotation costs;

[0063] 2. Because this invention uses implicit neural expression (SIREN network) to continuously represent the light field view, it reduces the generation error of traditional interpolation methods, improves the robustness of dynamic scenes, and adapts to dynamic interference;

[0064] 3. Because the present invention uses: preset occlusion rules for the center view and non-center view to accurately shield occlusion area errors and generate a dual-mode occlusion mask, it can cope with complex light interactions.

[0065] The method involved in this invention is based on unsupervised learning technology, which effectively overcomes the problem of difficult data annotation in wetland scenes and provides a reliable and robust basis for light field parallax estimation for multiple related application fields. Attached Figure Description

[0066] Figure 1 This is a flowchart of the network model used in the unsupervised learning light field disparity estimation method for wetland scenes according to an embodiment of the present invention. In the figure, the black arrows represent the forward propagation process, and the yellow arrows represent the backward propagation process. Detailed Implementation

[0067] In the accompanying drawings, the same or similar reference numerals are used to denote the same or similar elements or elements having the same or similar functions. The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0068] In the description of this invention, the terms "center," "longitudinal," "lateral," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting the scope of protection of this invention.

[0069] The unsupervised learning light field parallax estimation device for wetland scenes provided in this embodiment of the invention is used for light field display and reproduction of wetland scenes.

[0070] The unsupervised learning light field disparity estimation device includes a disparity estimation network 1, a view generation network 2, and an occlusion mask generation module 3. Wherein:

[0071] Parallax estimation network 1 is used to receive five-dimensional light field data L∈R B×C×N×H×W This paper extracts multi-view features from light field data and outputs a disparity map corresponding to the central view. Here, B represents the batch size, C represents the number of channels, N represents the number of viewpoints in the light field data, the image corresponding to the central viewpoint is referred to as the central view, and the images corresponding to sub-viewpoints outside the central viewpoint are referred to as non-central views. The central view and non-central views are collectively referred to as the light field view, and H and W represent the height and width of the light field view, respectively.

[0072] In this embodiment, the light field data includes a light field view with an angular resolution of N×N. Figure 1The image shown is a light field view with an angular resolution of 9×9, meaning N=9.

[0073] As a preferred implementation of the disparity estimation network 1, it is implemented using a CNN architecture, specifically including a feature extraction unit 11, a view selection unit 12, a matching cost unit 13, and a disparity regression unit 14.

[0074] The feature extraction unit 11 is used to extract the spatiotemporal features and edge features of the light field view in the light field data and perform normalization processing.

[0075] In one embodiment, the feature extraction unit 11 includes a three-dimensional convolutional layer 111, a BN layer 112, and a LeakyReLU activation function 113, used to extract the spatiotemporal features and edge features of the light field data.

[0076] The three-dimensional convolutional layer 111 is used to extract the spatiotemporal features of the light field view, and while extracting features, it enhances the spatiotemporal feature extraction capability of the three-dimensional convolutional layer 111 by employing cosine position encoding at a fixed position.

[0077] Preferably, the feature extraction unit 111 further includes a two-dimensional convolutional layer initialized with a Sobel operator, which is used to extract edge features of the light field view.

[0078] BN layer 112 is used to normalize the spatiotemporal and edge features of the extracted light field view to accelerate the training process and improve model stability.

[0079] The LeakyReLU activation function 113 is used to introduce nonlinear factors and enhance the expressive power of the feature extraction unit 11.

[0080] The view selection unit 12 is used to receive the spatiotemporal features and edge features of the normalized light field view. The spatiotemporal features and edge features will be referred to as the features of the light field view in the following text.

[0081] The view selection unit 12 is used to calculate a 0-1 attention weight α for each viewpoint of the light field data to adjust the characteristics of each light field view.

[0082] Specifically, the view selection unit 12 includes a Convolutional Block Attention Module (CBAM) 121 and a residual block 122.

[0083] The convolutional block attention module 121 is used to first extract the features of the normalized light field view output by the feature extraction unit 11. For each viewpoint of the light field data, a 0-1 attention weight α is calculated. Specifically, this includes adjusting the dimension of the light field view features to obtain the feature representation of the light field view as Feature∈R B×N×C×H×W This allows for the calculation of a 0-1 attention weight α for each viewpoint of the light field data. Where B represents the batch size, C represents the number of channels, and N, H, and W represent the number of viewpoints for the feature information and the spatial dimension of the feature information, respectively.

[0084] By swapping the viewpoint dimension and the channel dimension, the spatiotemporal features from different viewpoints are adjusted, and channel attention can be used to adjust the weights of different viewpoints. The network designed in this way can pay more attention to the view features that have a greater impact on the parallax results.

[0085] Difference block 122 is used to introduce a channel attention mechanism, multiply the attention weight α by the normalized light field view feature F output by the feature extraction unit 11, that is, to use (1+α)*F to represent the adjustment of the feature of each viewpoint, thereby achieving the purpose of optimizing the feature of the normalized light field view and outputting the optimization result.

[0086] The matching cost unit 13 is used to calculate the matching cost of the feature information output by the view selection unit 12 using the multi-view disparity estimation method, and input the matching cost into the disparity regression unit 14 to obtain the disparity map corresponding to the final center view.

[0087] During the training of the disparity estimation network 1, various methods were employed to augment the light field data. For example, when converting the RGB image of the light field to a grayscale image, random channel perturbation was used to simulate different grayscale conversion methods. Simultaneously, the dynamic range and contrast of the light field data were randomly adjusted to increase its diversity. Furthermore, noise was randomly added to the light field image with a certain probability, and optical simulation strategies such as random flipping or rotation, refocusing, and downsampling were used. In this embodiment, data augmentation is performed before the data is input into the model to enhance the model's generalization ability and prevent overfitting. The trained disparity estimation network receives the original light field data as input, eliminating the need for further data augmentation operations.

[0088] View generation network 2 employs a pre-trained SIREN architecture implicit neural network to receive the normalized coordinates corresponding to the center view of the light field data and the disparity map D(x, y) corresponding to the center view output by disparity estimation network 1. It then transforms the normalized coordinates based on the disparity map D(x, y) corresponding to the center view to obtain the transformed coordinates corresponding to the non-center views. This results in a non-center view.

[0089] As a preferred embodiment of the view generation network 2, the view generation network 2 includes a graph offset calculation unit 21, a coordinate transformation unit 22, and a disparity graph generation unit 23.

[0090] The image offset calculation unit 21 receives the disparity map D(x, y) corresponding to the center view output by the disparity estimation network 1, and calculates the offset of the non-center view relative to the center view using equation (1). The offset includes the horizontal pixel offset ΔX(x, y) and the vertical pixel offset ΔY(x, y) between the non-center view and the center view at pixel position (x, y), as well as the horizontal viewing angle offset dv and the vertical viewing angle offset du between the non-center view and the center view. Among them, the non-center view is the non-center view of the third dimension N.

[0091]

[0092] Where W and H are the width and height of the image, respectively. The non-center view has the same size as the center view, so no additional size processing is required. D(x,y) is the disparity value at pixel position (x,y) obtained by disparity estimation network 1.

[0093] The coordinate transformation unit 22 receives the normalized coordinates corresponding to the center view and, based on the offset of the non-center view relative to the center view, performs coordinate transformation on the normalized coordinates corresponding to the center view using equation (2) to obtain the transformed coordinates corresponding to the non-center view.

[0094]

[0095] The disparity map generation unit 23 is used to describe the center view as And based on the transformed coordinates corresponding to the non-center view Using the implicit neural network function C Described as Equation (3):

[0096]

[0097] This embodiment improves upon existing view generation networks used in unsupervised light field disparity estimation algorithms. By employing implicit neural representation, the pixel-based, discrete view generation process is optimized into a coordinate-based, continuous process, thereby avoiding interpolation and improving the accuracy of the view generation process.

[0098] In one embodiment, to avoid the cumulative error of a single viewpoint generation, the view generation network 2 further includes a symmetric viewpoint projection transformation unit 24, which is used to transform the output of the disparity map generation unit 23. Mapping to symmetrical positions (Ns, Nt), and aligning pixels using bilinear interpolation, generates a non-centered view V. i Where s and t represent the angular coordinates corresponding to the target viewpoint.

[0099] This embodiment reduces the cumulative error of the non-central view obtained by the view generation network 2 by using symmetrical view projection transformation.

[0100] In the above embodiments, in order to save the computing resources occupied by the pre-trained implicit neural expression network and speed up the computing speed, mixed precision is used for training in the pre-training stage and fp16 precision is used in the inference stage.

[0101] The occlusion mask generation module 3 is used to estimate the occlusion relationship between the center view and the non-center view based on the disparity map D(x,y) corresponding to the center view output by the disparity estimation network 1, thereby reducing the impact of occlusion on the disparity results.

[0102] In one embodiment, the occlusion mask generation module 3 includes an initial occlusion mask generation unit 31 and a final occlusion mask generation unit 32, wherein:

[0103] The initial occlusion mask generation unit 31 is used to generate corresponding initial occlusion masks based on two preset occlusion conditions. The first type of occlusion condition is that the mask is occluded in the center view but visible in the off-center view. The second type of occlusion condition is that the mask is visible in the center view but occluded in the off-center view.

[0104] The final occlusion mask generation unit 32 is used to perform an "OR" operation on the initial occlusion mask output by the initial occlusion mask generation unit 31 to obtain a final occlusion mask containing two types of occlusion.

[0105] The final occlusion mask serves two purposes: first, it eliminates view generation errors. For example, when there are differences in the pixels between the view generated by the view generation network 2 and the real view, the final occlusion mask can eliminate the errors caused by occlusion. Second, it adjusts the geometric features of the light field. For example, the final occlusion mask is also applied after the view selection unit 12, multiplying the final occlusion mask with the features of the light field view pixel by pixel. Since the final occlusion mask is a binary image with only two values, 0 and 1, the feature value is directly set to 0 where the final occlusion mask is 0, and the feature value remains unchanged where it is 1. This allows the occlusion mask to adjust the features of the light field view extracted by the view selection unit 12, thereby improving the accuracy of subsequent steps.

[0106] Transform the pixel coordinates (x, y) of the center view V to the non-center view V. iIn the coordinate system, obtain the pixel coordinates (x, y) of the center view V and the non-center view V. i Warp coordinates in the coordinate system i (x c ,y c ):

[0107]

[0108] Among them, (x s,i ,y t,i (i) represents the pixel coordinates of a non-center view. x i y ) represents the angular coordinates of a non-center view.

[0109] With Warp i (x c ,y c Take a neighborhood centered on ) (x c ,y c )∈V i , where γ is the set threshold.

[0110] With Warp i (x c ,y c Centered on ), a square region N(Warp) with a side length of 2γ+1 is defined. i (x c ,y c ), γ), where γ is used to control the number of pixels extending outwards from the center.

[0111] For the first type of occlusion, the final occlusion mask generation unit 32 obtains the final occlusion mask M. i (x,y) specifically includes:

[0112] In the non-visible area of ​​the central view V, in the non-central view V i The center view V may be presented in a visible state. Therefore, the pixel coordinates (x, y) of the center view V are transformed to those of the non-center view V. i When using a coordinate system, some views may not be mapped to non-center views V. i The coordinates of valid pixels, these pixel locations are "holes", which represent areas with missing or invalid data.

[0113] Therefore, obtain each non-center view V i The final occlusion mask M i (x,y), as shown in equation (5):

[0114]

[0115] Among them, Hi To obtain the non-center view V after traversing all pixel coordinates (x, y) of the center view V. i The intersection with the visible area of ​​the central view V is shown in equation (6):

[0116]

[0117] For the second type of occlusion, the final occlusion mask generation unit 32 obtains the final occlusion mask G. i (x s,i ,y s,i Specifically, it includes:

[0118] Each non-center view V i pixel coordinates (x) s,i ,y s,i In the coordinate system of the transformed central view V, as shown in equation (7), the non-central view V is obtained. i pixel coordinates (x) s,i ,y s,i The coordinates C in the coordinate system of the central view V i (x s,i ,y s,i ):

[0119] C i (x s,i ,y s,i )={(x c ,y c )∈V|Warp i (x c ,y c )=(x s,i ,y s,i )} (7)

[0120] Determine C i (x s,i ,y s,i If the number of ) is not greater than 1, then there is no many-to-one mapping, C i (x s,i ,y s,i The maximum distance between any two pixel coordinates in () max The value is 0. Conversely, C is 0. i (x s,i ,y s,i If the number of ) is greater than 1, then the distance is further calculated. max In this embodiment, the Euclidean distance d is taken as an example, as shown in equation (8):

[0121]

[0122] The two adjacent square regions N(Warp) are determined by equation (9). i (x c ,y c Whether ghosting occurs (γ) is determined to determine the final occlusion mask G for the second type of occlusion. i (x s,i ,y s,i ):

[0123]

[0124] In the formula, when distance max When the value is greater than β, it is determined that a second type of occlusion has occurred, and the final occlusion mask G is then determined. i (x s,i ,y s,i The distance is set to 0 to ignore the influence of this region in subsequent processing; when distance... max If the value is not greater than β, it is determined that no second-type occlusion has occurred, and the final occlusion mask G is then determined. i (x s,i ,y s,i The value is 1.

[0125] This embodiment proposes an unsupervised learning method for light field disparity estimation. The core of this method is the use of implicit neural representation to replace the traditional interpolation-based view generation process. By implicitly representing the central view of the light field scene, errors in the view generation stage are reduced. Furthermore, addressing the common occlusion problem in light field disparity estimation, a simple yet effective occlusion mask generation method is proposed. By using two preset occlusion scenarios, a relatively accurate occlusion mask can be obtained.

[0126] The network model designed in this embodiment uses a loss function consisting of two parts. The network is trained by backpropagation through the calculation of the loss function.

[0127] The first part calculates the mean squared error (MSE) between the view generated by the view generation network and the real view, using it as the loss function (10), and obtains the pixel difference loss used to measure the difference between the generated view and the real view. L2 :

[0128]

[0129] Where I1 and I2 represent the view generated by the view generation network and the real view, respectively. M and N are the height and width of the disparity map, and (i,j) represents the index traversal coordinates of I1 and I2. This formula obtains the mean squared error by calculating the squared difference between pixels pixel by pixel and taking the square of the difference.

[0130] The second part uses the smoothing loss function (11) of equation (11) to obtain the smoothing loss. smooth :

[0131] in, and Let represent the gradients of the disparity map D(x, y) corresponding to the central view V obtained by disparity estimation network 1 in the horizontal and vertical directions, respectively. The gradients of the disparity map D(x, y) corresponding to the central view V obtained by disparity estimation network 1 in the x and y directions are given by [the following equations are used to represent the gradients of the disparity map D(x, y) in the x and y directions, respectively]. and Provided.

[0132] Finally, the loss function is given by the following formula (12):

[0133] Loss total =λLoss L2 +(1-λ)Loss smooth (12)

[0134] Here, λ is used to control the weights of the two loss functions, and the parameter used to control the edge weights is set to 0.95 during training.

[0135] This invention also provides an unsupervised learning method for estimating light field disparity, comprising:

[0136] Step 1: Receive five-dimensional light field data L∈R B×C×N×H×W The system extracts multi-view spatiotemporal features and outputs the disparity map D(x, y) corresponding to the central view V. The light field data includes light field views with height H and width W, which are divided into the central view V and non-central views V. i B represents the batch number, C represents the number of channels, N represents the number of viewpoints in the light field data, i represents the sequence number of the non-center view, and the value is a natural number from 0 to N×N. (x, y) represents the pixel coordinates of the center view.

[0137] Step 2: Using a hidden neural network, the normalized coordinates of the central view are transformed based on the disparity map D(x, y) to generate the non-central view V. i ;

[0138] Step 3: Estimate the central view Vx and non-central view Vy based on the disparity map D(x,y). i The occlusion relationship is used to generate the final occlusion mask to eliminate occlusion errors;

[0139] Step 4: Jointly optimize the outputs of Step 1, Step 2 and Step 3 to complete the light field disparity estimation in an unsupervised learning manner.

[0140] In one embodiment, step 1 includes:

[0141] Step 11: Extract the features of the light field view from the light field data and normalize them;

[0142] Step 12: Extract the features of the normalized light field view, adjust the dimensions of the features of the light field view, and represent the adjusted features of the light field view as Feature∈R B×N×C×H×W This allows us to obtain the attention weights for each viewpoint of the light field data, and adjust the features of each light field view based on the attention weights for each viewpoint.

[0143] Step 13: Calculate the feature matching cost using the multi-view parallax estimation method;

[0144] Step 14: Generate disparity map D(x, y) based on matching cost.

[0145] In one embodiment, step 2 includes:

[0146] Step 21: Calculate the horizontal and vertical pixel offsets ΔX(x,y) and ΔY(x,y) of the non-center view relative to the center view, as well as the viewing angle offsets dv and du, based on the disparity map D(x,y).

[0147]

[0148] Step 22: Convert the normalized coordinates of the center view to the coordinates of the non-center view based on the offset;

[0149] Step 23, describe the center view as Using the implicit neural network function C It is described as Equation (3).

[0150] In one embodiment, step 3 includes:

[0151] Step 31: Preset two types of occlusion: the center view is occluded and not visible, and the center view is visible but not occluded.

[0152] Step 32: Perform an OR operation on the two initial occlusion masks to generate a final mask containing the occlusion regions of both types. The final mask is used to adjust for view generation errors and light field geometry features.

[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Those skilled in the art should understand that modifications can be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An unsupervised learning optical field parallax estimation device, characterized in that, include: A parallax estimation network is used to receive five-dimensional light field data L∈R. B×C×N×H×W The system extracts multi-view spatiotemporal features and outputs the disparity map D(x, y) corresponding to the central view V. The light field data includes light field views with height H and width W, which are divided into the central view V and non-central views V. i B represents the batch number, C represents the number of channels, N represents the number of viewpoints in the light field data, i represents the sequence number of the non-center view, and the value is a natural number from 0 to N×N. (x, y) represents the pixel coordinates of the center view. The view generation network employs an implicit neural network to transform the normalized coordinates of the central view based on the disparity map D(x,y) to obtain the transformed coordinates of the non-central views. Generate non-centered view V i ; The occlusion mask generation module is used to estimate the center view V and the non-center view V based on the disparity map D(x,y). i The occlusion relationship is used to generate the final occlusion mask to eliminate occlusion errors; Among them, the disparity estimation network, the view generation network and the occlusion mask generation module are jointly optimized to complete the light field disparity estimation in an unsupervised learning manner; The view generation network includes: The image offset calculation unit calculates the horizontal and vertical pixel offsets ΔX(x,y) and ΔY(x,y) of the non-center view relative to the center view, as well as the viewing angle offsets dv and du, based on the disparity map D(x,y). The coordinate transformation unit is used to convert the normalized coordinates of the center view into the coordinates of the non-center view based on the offset. The disparity map generation unit is used to describe the center view as... Using the implicit neural network function C Described as Equation (3): The view generation network also includes: The symmetrical viewpoint projection transformation unit converts the output of the disparity map generation unit into a symmetrical viewpoint projection transformation unit. Mapping to a symmetrical viewpoint (Ns, Nt), and aligning pixels using bilinear interpolation, a non-centered view V is generated. i .

2. The unsupervised learning optical field disparity estimation device according to claim 1, characterized in that, The disparity estimation network includes: The feature extraction unit, which includes a 3D convolutional layer, a BN layer, and a LeakyReLU activation function, is used to extract features of the light field view in the light field data and perform normalization processing. The view selection unit is used to extract features from the normalized light field view, adjust the dimensions of the features, and represent the adjusted light field view features as Feature∈R. B×N×C×H×W This allows us to obtain the attention weights for each viewpoint of the light field data, and adjust the features of each light field view based on the attention weights for each viewpoint. The matching cost unit is used to calculate the feature matching cost using a multi-view disparity estimation method. The disparity regression unit is used to generate a disparity map D(x, y) based on the matching cost.

3. The unsupervised learning optical field disparity estimation device according to claim 1 or 2, characterized in that, The occlusion mask generation module includes: The initial occlusion mask generation unit has two preset occlusion scenarios: the center view is occluded but not visible, and the center view is visible but not occluded. The final occlusion mask generation unit performs an "OR" operation on the two types of initial occlusion masks to generate a final mask containing the two types of occlusion regions; The final mask is used to adjust for view generation errors and light field geometry features; The final masks for the two types of occlusion are denoted as M. i (x,y) and G i (x s,i ,y s,i ): Among them, H i To obtain the non-center view V after traversing all pixel coordinates (x, y) of the center view V. i The intersection of C with the visible area of ​​the center view V i (x s,i ,y s,i (V) is a non-center view i pixel coordinates (x) s,i ,y s,i The coordinates of distance in the coordinate system of the central view V. max C i (x s,i ,y s,i The maximum distance between any two pixel coordinates in N(Warp) is N(Warp) i (x c ,y c ),γ) is based on Warp i (x c ,y c A square region centered at γ and with sides of length 2γ+1, where γ is used to control the center Warp. i (x c ,y c The number of pixels extending outwards, β is a parameter used to control the sensitivity of ghost detection, (x c ,y c (V) is a non-center view i Pixel coordinates:

4. The unsupervised learning optical field disparity estimation device according to claim 1 or 2, characterized in that, Loss function of the network total Described as Equation (11): Loss total =λLoss L2 +(1-λ)Loss smooth (11) In the formula, Loss L2 Mean squared error loss is used to measure the pixel difference between the generated view and the real view. smooth To smooth the loss, the gradient smoothness of the disparity map is optimized, and λ is used to control the loss. L2 and Loss smooth The weight.

5. An unsupervised learning method for estimating light field disparity, characterized in that, include: Step 1: Receive five-dimensional light field data L∈R B×C×N×H×W Extract multi-view spatiotemporal features and output the disparity map D(x,y) corresponding to the central view V; where the light field data includes light field views with height H and width W, and the light field views are divided into the central view V and non-central view V. i B represents the batch number, C represents the number of channels, N represents the number of viewpoints in the light field data, i represents the sequence number of the non-center view, and the value is a natural number from 0 to N×N. (x, y) represents the pixel coordinates of the center view. Step 2: Using a hidden neural network, the normalized coordinates of the central view are transformed based on the disparity map D(x, y) to obtain the transformed coordinates of the non-central views. Generate non-centered view V i ; Step 3: Estimate the central view Vx and non-central view Vy based on the disparity map D(x,y). i The occlusion relationship is used to generate the final occlusion mask to eliminate occlusion errors; Step 4: Jointly optimize the outputs of Step 1, Step 2 and Step 3 to complete the light field disparity estimation in an unsupervised learning manner; Step 2 includes: Step 21: Calculate the horizontal and vertical pixel offsets ΔX(x,y) and ΔY(x,y) of the non-center view relative to the center view, as well as the viewing angle offsets dv and du, based on the disparity map D(x,y). Step 22: Convert the normalized coordinates of the center view to the coordinates of the non-center view based on the offset; Step 23, describe the center view as Using the implicit neural network function C Described as Equation (3): Step 2 also includes: The symmetrical viewpoint projection transformation unit converts the output of the disparity map generation unit into a symmetrical viewpoint projection transformation unit. Mapping to a symmetrical viewpoint (Ns, Nt), and aligning pixels using bilinear interpolation, a non-centered view V is generated. i .

6. The unsupervised learning light field disparity estimation method according to claim 5, characterized in that, Step 1 includes: Step 11: Extract the features of the light field view from the light field data and normalize them; Step 12: Extract the features of the normalized light field view, adjust the dimensions of the features of the light field view, and represent the adjusted features of the light field view as Feature∈R B×N×C×H×W This allows us to obtain the attention weights for each viewpoint of the light field data, and adjust the features of each light field view based on the attention weights for each viewpoint. Step 13: Calculate the feature matching cost using the multi-view parallax estimation method; Step 14: Generate disparity map D(x, y) based on matching cost.

7. The unsupervised learning light field disparity estimation method according to claim 5 or 6, characterized in that, Step 3 includes: Step 31: Preset two types of occlusion: the center view is occluded and not visible, and the center view is visible but not occluded. Step 32: Perform an "OR" operation on the two initial occlusion masks to generate a final mask containing the two types of occlusion regions; The final mask is used to adjust for view generation errors and light field geometry features; The final masks for the two types of occlusion are denoted as M. i (x,y) and G i (x s,i ,y s,i ): Among them, H i To obtain the non-center view V after traversing all pixel coordinates (x, y) of the center view V. i The intersection of C with the visible area of ​​the center view V i (x s,i ,y s,i (V) is a non-center view i pixel coordinates (x) s,i ,y s,i The coordinates of distance in the coordinate system of the central view V. max C i (x s,i ,y s,i The maximum distance between any two pixel coordinates in N(Warp) is N(Warp) i (x c ,y c ),γ) is based on Warp i (x c ,y c A square region centered at γ and with sides of length 2γ+1, where γ is used to control the center Warp. i (x c ,y c The number of pixels extending outwards, β is a parameter used to control the sensitivity of ghost detection, (x c ,y c (V) is a non-center view i Pixel coordinates:

Citation Information

Patent Citations

  • Light field depth self-supervised learning method based on occlusion region iterative optimization

    CN112288789A

  • Method for synthesizing virtual viewpoint image based on implicit neural scene representation

    CN114666564A