Light field sub-pixel parallax estimation model and training method and system thereof, medium and electronic equipment

By designing a light field subpixel disparity estimation model, including image conversion, feature extraction and subpixel cost construction, and introducing an edge loss function, the problems of small number of disparity labels and occlusion pixels in existing light field disparity estimation are solved, and high-performance and robust disparity estimation are achieved.

CN120070530APending Publication Date: 2025-05-30DONGHUA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510043822.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the existing light field parallax estimation method, the number of parallax tags is small, which affects the result of parallax estimation, and occluding pixels also leads to a decrease in estimation accuracy.

Method used

A light field subpixel disparity estimation model is designed to generate aggregated cost bodies through image conversion, feature extraction, subpixel cost construction and attention enhancement, and a predicted disparity map is output using the parallax regression module. At the same time, an edge loss function is introduced to alleviate the effects of occluding pixels.

Benefits of technology

The number of parallax labels is increased, the accuracy and robustness of parallax estimation is enhanced, and high performance is demonstrated on multiple synthetic and real datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070530A_ABST
    Figure CN120070530A_ABST
Patent Text Reader

Abstract

The invention provides a light field sub-pixel parallax estimation model and a training method and system thereof, a medium and electronic equipment. The light field sub-pixel parallax estimation model comprises an image conversion module used for performing image conversion on a light field sub-aperture image to obtain a light field macro-pixel image; the feature extraction module is used for performing feature extraction on the light field macro-pixel image to obtain a feature map; the cost construction module is used for performing sub-pixel cost construction and attention enhancement on the feature map to obtain an attention enhancement cost body; the cost aggregation module is used for regularizing the attention enhancement cost body to obtain an aggregation cost body; the parallax regression module is used for performing parallax mapping on the aggregation cost body to obtain a predicted parallax map corresponding to the light field sub-aperture image; the predicted disparity map is a disparity map corresponding to a center sub-aperture image in the light field sub-aperture images; by designing sub-pixel cost construction, the problem that the number of parallax labels is small in existing light field parallax estimation is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of physics, in particular to computer vision and machine learning technologies, and particularly to light field parallax estimation. Specifically, it is a light field sub-pixel parallax estimation model and its training method, system, medium, and electronic device. Background Art

[0002] Traditional two-dimensional images record the intensity of light, while light field images record not only the intensity information of light but also the direction information of light. The light field expands our understanding of scene capture and scene description in a more rich and comprehensive way of describing the world we live in. In the field of computer vision, some other applications are mainly realized by analyzing the rich information in the light field, such as parallax estimation, semantic segmentation, refocusing, super-resolution, view synthesis, 3D reconstruction, virtual reality, etc.

[0003] Light field parallax estimation is an important task in the light field and is a basic technology for other subsequent tasks, such as view synthesis and virtual reality. Existing machine learning-based light field parallax estimation algorithms use a cost volume to measure consistency and then obtain a disparity map. However, these methods mostly use a small number of integer disparity labels, which is extremely small compared to the disparity labels in stereo matching and affects the result of parallax estimation. Summary of the Invention

[0004] The purpose of the present invention is to provide a light field sub-pixel parallax estimation model and its training method, system, medium, and electronic device to solve the problems pointed out in the above background art.

[0005] In a first aspect, the present invention provides a light field sub-pixel parallax estimation model. The light field sub-pixel parallax estimation model includes: an image conversion module for performing image conversion on a light field sub-aperture image to obtain a light field macro-pixel image; a feature extraction module for performing feature extraction on the light field macro-pixel image to obtain a feature map; a cost construction module for performing sub-pixel cost construction and attention enhancement on the feature map to obtain an attention-enhanced cost volume; a cost aggregation module for regularizing the attention-enhanced cost volume to obtain an aggregated cost volume; a disparity regression module for performing disparity mapping on the aggregated cost volume to obtain a predicted disparity map corresponding to the light field sub-aperture image; the predicted disparity map is the disparity map corresponding to the central sub-aperture image in the light field sub-aperture image.

[0006] In an implementation of the first aspect, the cost construction module includes: a sub-pixel disparity label establishment unit for establishing a sub-pixel disparity label sequence based on a scaling factor; a sub-pixel disparity cost calculation unit for performing cost calculation on the feature map according to the sub-pixel disparity label sequence to obtain a sub-pixel cost volume; and an attention enhancement unit for enhancing the attention of the sub-pixel cost volume to obtain the attention-enhanced cost volume.

[0007] In an implementation of the first aspect, the sub-pixel disparity cost calculation unit uses dilated convolution for cost calculation.

[0008] In an implementation of the first aspect, the feature extraction module uses dilated convolution with weight sharing for feature extraction; and / or the cost aggregation module uses 3D convolution for regularization.

[0009] In a second aspect, the present invention provides a training method for the light field sub-pixel disparity estimation model described above. The training method includes: obtaining a training data set; the training data set includes multiple light field sub-aperture images; inputting the light field sub-aperture images into the light field sub-pixel disparity estimation model so that the light field sub-pixel disparity estimation model outputs a predicted disparity map; constructing a loss function of the light field sub-pixel disparity estimation model according to the true disparity map and the predicted disparity map, and training the light field sub-pixel disparity estimation model based on the loss function to obtain a trained light field sub-pixel disparity estimation model.

[0010] In an implementation of the second aspect, the obtaining of the training data set includes: obtaining an original light field sub-aperture image; randomly slicing and sampling the original light field sub-aperture image to obtain light field sub-aperture image slices; performing data augmentation on the light field sub-aperture image slices to obtain image slices after data augmentation; and using the image slices after data augmentation as the light field sub-aperture images.

[0011] In an implementation of the second aspect, the loss function includes: a mean absolute loss function and an edge loss function; wherein, the edge loss function is implemented based on the Laplacian operator.

[0012] In this implementation, by introducing an edge loss function into the loss function, the influence brought by occluded pixels in light field disparity estimation is alleviated, and high performance and strong robustness in light field disparity estimation are achieved in the face of multiple synthetic data sets and real data sets.

[0013] In a third aspect, the present invention provides a training system for the above-mentioned light field sub-pixel disparity estimation model. The training system includes: a data acquisition module for acquiring a training data set, where the training data set includes multiple light field sub-aperture images; an input-output module for inputting the light field sub-aperture images into the light field sub-pixel disparity estimation model so that the light field sub-pixel disparity estimation model outputs a predicted disparity map; and a model training module for constructing a loss function of the light field sub-pixel disparity estimation model based on a ground truth disparity map and the predicted disparity map, and training the light field sub-pixel disparity estimation model based on the loss function to obtain a trained light field sub-pixel disparity estimation model.

[0014] In a fourth aspect, the present invention provides an electronic device, which includes: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the above-mentioned training method.

[0015] In a fifth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by an electronic device, the above-mentioned training method is implemented.

[0016] As described above, the light field sub-pixel disparity estimation model and its training method, system, medium and electronic device of the present invention have the following beneficial effects:

[0017] (1) Compared with the prior art, the present invention provides a light field sub-pixel disparity estimation method. Aiming at the target of disparity estimation of the central sub-aperture image of the light field, the problem of few disparity labels in traditional light field disparity estimation is solved by designing sub-pixel cost, and at the same time, the edge loss is used to alleviate the influence of occluded pixels in light field disparity estimation. In the face of multiple synthetic data sets and real data sets, high performance and strong robustness in light field disparity estimation are achieved.

[0018] (2) The present invention adds sub-pixel disparity labels on the basis of the original integer disparity labels to increase the number of disparity labels, then designs corresponding convolution parameters, and finally uses dilated convolution to complete the construction of sub-pixel cost, effectively solving the problem of few disparity labels in existing light field disparity estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It shows a schematic structural diagram of the light field sub-pixel disparity estimation model described in the embodiment of the present invention.

[0020] Figure 2 It shows a schematic structural diagram of the cost construction module described in the embodiment of the present invention.

[0021] Figure 3It shows a flowchart of the training method according to the embodiments of the present invention.

[0022] Figure 4 It shows a flowchart of obtaining a training data set according to the embodiments of the present invention.

[0023] Figure 5 It shows a schematic structural diagram of the training system according to the embodiments of the present invention. Detailed implementation manners

[0024] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0025] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. The diagrams only show the components related to the present invention, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0026] Refer to Figures 1 to 5 The following embodiments of the present invention provide a light field sub-pixel disparity estimation model and its training method, system, medium, and electronic device. Compared with the prior art, the present invention provides a light field sub-pixel disparity estimation method. For the target of light field disparity estimation, by designing sub-pixel disparity labels, constructing a sub-pixel cost volume, and adding attention enhancement, the problem of few disparity labels in traditional light field disparity estimation is solved. At the same time, the edge loss is used to alleviate the influence of occluded pixels in light field disparity estimation. When facing multiple light field synthesis data sets and real data sets, the accuracy and robustness of light field disparity estimation are improved; the present invention adds sub-pixel disparity labels on the basis of the original integer disparity labels, increases the number of disparity labels, then designs corresponding convolution parameters, and finally uses dilated convolution to complete the construction of the sub-pixel cost, effectively solving the problem of few disparity labels in existing light field disparity estimation.

[0027] Next, the technical solutions in the embodiments of the present invention will be described in detail with reference to the accompanying drawings in the embodiments of the present invention.

[0028] In one embodiment, the present invention provides a light field sub-pixel disparity estimation model, which is applied to light field disparity estimation. Specifically, the light field sub-pixel disparity estimation model is used to perform disparity estimation on the input light field sub-aperture image and output a predicted disparity map. Among them, the predicted disparity map is the disparity map corresponding to the central sub-aperture image of the light field sub-aperture image.

[0029] As Figure 1 shown, in one embodiment, the light field sub-pixel disparity estimation model includes:

[0030] An image conversion module 11, configured to perform image conversion on the light field sub-aperture image to obtain a light field macro-pixel image.

[0031] In this embodiment, converting the light field sub-aperture image into a light field macro-pixel image through image conversion adopts conventional technical means in the art, so the working principle thereof will not be elaborated in detail herein.

[0032] A feature extraction module 12, configured to perform feature extraction on the light field macro-pixel image to obtain a feature map.

[0033] In one embodiment, the feature extraction module 12 uses dilated convolution with weight sharing for feature extraction.

[0034] A cost construction module 13, configured to perform sub-pixel cost construction and attention enhancement on the feature map to obtain an attention-enhanced cost volume.

[0035] Specifically, the cost construction module 13 adopts dilated convolution to calculate the cost at the sub-pixel level and splices the costs in the disparity dimension to complete the construction of the sub-pixel cost volume, and adjusts the weights of the channel information based on channel attention to improve the representation ability of the sub-pixel cost volume, and finally obtains an attention-enhanced cost volume.

[0036] As Figure 2 shown, in one embodiment, the cost construction module 13 includes:

[0037] A sub-pixel disparity label establishment unit 131, configured to establish a sub-pixel disparity label sequence based on a scaling factor.

[0038] It should be noted that the sub-pixel disparity label establishment unit 131 determines the disparity range and obtains a sub-pixel disparity label sequence.

[0039] Specifically, the sub-pixel disparity label establishment unit 131 establishes a sub-pixel disparity label sequence by introducing a scaling factor α, which is more accurate than the integer disparity label in the traditional light field disparity estimation task. Among them,

[0040] the scaling factor α takes a positive integer value and satisfies not equal to 0.

[0041] A is the light field angular resolution parameter, indicating that the light field angular resolution is A×A. is the floor function, ensuring that for the disparity labels established for the scaling factor, corresponding pixels can be located in different views of the light field image.

[0042] The sub-pixel disparity label sequence includes multiple sub-pixel disparity labels, where the sub-pixel disparity label satisfies the following formula:

[0043]

[0044] where d(i) is the (i + 1)-th sub-pixel disparity label in the sub-pixel disparity label sequence; d min is the predefined integer minimum disparity label; i is the index, starting from 0, and stopping when d(i) reaches d max ; d max is the predefined integer maximum disparity label.

[0045] It should be noted that the sub-pixel disparity label sequence obtained by the sub-pixel disparity label establishment unit 131 is applicable to the entire data set.

[0046] It should be noted that the working principle of the above corresponding pixel positioning is as follows:

[0047] A light field sub-aperture image is represented as L(u, v, x, y), where u and v represent the angular parameters of the light field sub-aperture image, and x and y represent the spatial parameters of the light field sub-aperture image.

[0048] Assume that L(u c , v c ) is the central sub-aperture image of the light field sub-aperture image, and the disparity of L(u c , v c , x i , y j ) is d(i), (u c and v c represent the angular parameters of the central sub-aperture image, and x i and y j represent the spatial parameters of the central sub-aperture image).

[0049] In the sub-aperture domain, the central sub-aperture image L(u c , v c ) of the light field sub-aperture image and another light field sub-aperture image L(u m , v n )(u m and v nThe horizontal displacement and vertical displacement (dx, dy) of the angular parameter representing the image of another optical field sub-aperture satisfy: where (dx, dy) should be integers. When constructing sub-pixel disparity labels, in order to locate the corresponding pixel L(u m ,v n ,x k ,y l )(x k and y l representing the spatial parameters of the image of another optical field sub-aperture), the following method is adopted: (u m ,v n ,x k ,y l ) = (u c ±αk m ,v c ±αk n ,x i -dx,y j -dy), where k m and k n take values of At this time, the positioning of the corresponding pixel is completed, and then the displacement is transferred to the macro-pixel domain: where dx_m and dy_m respectively represent the horizontal displacement and vertical displacement of the corresponding pixel relative to the pixel of the central sub-aperture image in the macro-pixel domain, and the convolution parameters can be further determined.

[0050] The sub-pixel disparity cost calculation unit 132 is configured to calculate the cost of the feature map according to the sub-pixel disparity label sequence to obtain a sub-pixel cost volume.

[0051] In one embodiment, the sub-pixel disparity cost calculation unit 132 uses dilated convolution for cost calculation.

[0052] Specifically, the sub-pixel disparity cost calculation unit 132 designs different convolution parameters according to the sub-pixel disparity label sequence, and then uses dilated convolution to perform convolution operations to obtain costs, and splices the costs in the disparity dimension to obtain a sub-pixel cost volume.

[0053] In one embodiment, the convolution parameters at least include: a convolution kernel, a dilation rate, and a padding parameter.

[0054] It should be noted that the sub-pixel disparity cost calculation unit 132 is used to design the convolution kernel, dilation rate, and padding parameter, and complete cost calculation and cost volume construction; among them, the convolution kernel, dilation rate, and padding parameter are parameters in the dilated convolution used for cost calculation.

[0055] Specifically, the sub-pixel disparity cost calculation unit 132 calculates the cost for the feature map through dilated convolution, and the convolution kernel size parameter K of the dilated convolution satisfies:

[0056]

[0057] where K is the convolution kernel size parameter, indicating that the convolution kernel size is K×K; the dilation rate D satisfies:

[0058]

[0059] where D is the dilation rate; sign(·) is the sign function; the padding parameter P satisfies:

[0060]

[0061] Design the dilated convolution according to different sub-pixel disparity labels and calculate the sub-pixel cost Cost(d(i)), which can be expressed as:

[0062] Cost(d(i)) = Conv2(F, A, α, d(i));

[0063] where Conv2 is the convolution operation; F is the feature map.

[0064] It should be noted that for a given scaling factor, the design result of the convolution kernel needs to be corrected to some extent. This is because in the sub-pixel disparity label sequence, when reduction can be performed, it is equivalent to using a smaller α. At this time, the convolution kernel size becomes larger, and more angular information in the light field can be added to the calculation process. Therefore, when the above situation occurs, the present invention adopts a larger convolution kernel size. Then, the sub-pixel costs Cost(d(i)) corresponding to each sub-pixel disparity label d(i) are concatenated in the disparity channel to obtain the sub-pixel cost volume.

[0065] The attention enhancement unit 133 is used to enhance the attention of the sub-pixel cost volume to obtain the attention-enhanced cost volume.

[0066] Specifically, the attention enhancement unit 133 performs channel attention operation on the sub-pixel cost volume, calculates the channel weights of the sub-pixel cost volume using channel attention, and obtains the attention-enhanced cost volume using element-wise multiplication.

[0067] It should be noted that the attention enhancement unit 133 adopts conventional technical means in the technical field, so the working principle thereof will not be elaborated in detail herein; through this attention enhancement unit 133, the channel attention weights are adaptively learned, thereby improving the expression ability of the cost volume.

[0068] A cost aggregation module 14 is configured to regularize the attention-enhanced cost volume to obtain an aggregated cost volume.

[0069] In one embodiment, the cost aggregation module 14 uses 3D convolution for regularization.

[0070] A disparity regression module 15 is configured to perform disparity mapping on the aggregated cost volume to obtain a predicted disparity map corresponding to the light field sub-aperture image.

[0071] It should be noted that the predicted disparity map is the disparity map corresponding to the central sub-aperture image in the light field sub-aperture image.

[0072] The present invention provides a light field sub-pixel disparity estimation model. Through this model, light field disparity estimation can be achieved, effectively solving the problem of few disparity labels in the existing light field disparity estimation task.

[0073] As Figure 3 shown, in one embodiment, the present invention provides a training method for applying the above light field sub-pixel disparity estimation model; specifically, the training method includes:

[0074] Step S1: Obtain a training data set.

[0075] In one embodiment, the training data set is sourced from the 4D light field benchmark, which is a data set proposed by four scholars, namely Katrin Honauer et al. (https: / / lightfield-analysis.uni-konstanz.de / ).

[0076] Specifically, the training data set includes multiple light field sub-aperture images.

[0077] As Figure 4 shown, in one embodiment, the obtaining of the training data set includes:

[0078] Step S11: Obtain the original light field sub-aperture image.

[0079] Step S12: Perform random slicing sampling on the original light field sub-aperture image to obtain light field sub-aperture image slices.

[0080] Step S13: Perform data augmentation on the light field sub-aperture image slices to obtain the image slices after data augmentation; the image slices after data augmentation are used as the light field sub-aperture images.

[0081] It should be noted that through the data augmentation operation in step S13, the training data set is extended.

[0082] In one embodiment, the data augmentation includes at least, but is not limited to, any one or two or more of the following processing methods: random flipping and rotation, color channel reassignment, brightness and contrast adjustment, refocusing, and downsampling.

[0083] Step S2: Input the sub-aperture image of the light field into the light field sub-pixel disparity estimation model so that the light field sub-pixel disparity estimation model outputs a predicted disparity map.

[0084] Specifically, input the image slice after data augmentation obtained in step S13 above into the light field sub-pixel disparity estimation model so that the model outputs a corresponding predicted disparity map.

[0085] It should be noted that the working principle of step S2 is the same as that of the light field sub-pixel disparity estimation model introduced above, so it will not be elaborated in detail here.

[0086] Step S3: Construct a loss function of the light field sub-pixel disparity estimation model according to the true disparity map and the predicted disparity map, and train the light field sub-pixel disparity estimation model based on the loss function to obtain a trained light field sub-pixel disparity estimation model.

[0087] It should be noted that for the training method provided by the present invention, the training process of the light field sub-pixel disparity estimation model is end-to-end training, and the parameters of the light field sub-pixel disparity estimation model are updated by performing the gradient descent method based on the loss function to obtain a trained light field sub-pixel disparity estimation model.

[0088] In one embodiment, before step S3, the training method further includes: obtaining a true disparity map.

[0089] It should be noted that this true disparity map corresponds to the original sub-aperture image of the light field in step S11.

[0090] In one embodiment, the training method further includes: performing random slice sampling on the true disparity map to obtain true disparity map slices.

[0091] It should be noted that the random slice sampling of the true disparity map involved in this step is the same as the random slice sampling of the original sub-aperture image of the light field in step S12 above, to ensure that the obtained true disparity map slices correspond one-to-one with the sub-aperture image slices of the light field, so as to realize constructing a loss function according to the predicted disparity map corresponding to the true disparity map slice and the target sub-aperture image slice of the light field (i.e., the sub-aperture image slice corresponding to this true disparity map slice) in step S3.

[0092] In one embodiment, the loss function includes: a mean absolute loss function and an edge loss function; wherein, the edge loss function is implemented based on the Laplacian operator.

[0093] Specifically, the mean absolute loss function is used to calculate the mean absolute loss of the predicted disparity map and the true disparity map slices; the edge loss function is used to perform Laplacian operation on the predicted disparity map and the true disparity map slices, and then calculate the mean absolute loss to obtain the edge loss result; the two are combined to obtain the loss function.

[0094] In one embodiment, the mean absolute loss function and the edge loss function are combined by summation, satisfying:

[0095] L = L 1 +βL edge ;

[0096] where L is the loss function, L 1 is the mean absolute loss function, which is a conventional loss function in the art, so its implementation form will not be elaborated in detail here; L edge is the edge loss function, satisfying:

[0097] L edge = ||Δdp - Δdp gt || 1 ;

[0098] where ||·|| 1 is the mean absolute loss; Δ is the Laplacian operation, which is a common operation in the technical field of the present invention, so its working principle will not be elaborated in detail here; dp and dp gt are the predicted disparity map and the true disparity map slices respectively; β is a tuning parameter used to adjust the importance of the mean absolute loss and the edge loss.

[0099] It should be noted that when dealing with complex scenes, the occlusion phenomenon will seriously affect the estimation accuracy, and existing models often cannot effectively handle these problems; in the present invention, by adding edge loss during the training process of the light field sub-pixel disparity estimation model, the influence of occlusion in the light field disparity estimation task is alleviated, thereby improving the accuracy and robustness of the light field disparity estimation.

[0100] In one embodiment, after training the light field sub-pixel disparity estimation model through the above steps, the trained light field sub-pixel disparity estimation model is obtained, and the trained light field sub-pixel disparity estimation model is tested through this trained light field sub-pixel disparity estimation model to obtain the trained light field sub-pixel disparity estimation model.

[0101] Specifically, the trained light field sub-pixel disparity estimation model is tested using a test data set.

[0102] In one embodiment, the test data set is sourced from the 4D light field benchmark.

[0103] Specifically, the dataset proposed by four scholars, namely Katrin Honauer et al., is divided into a training dataset and a test dataset according to a certain ratio.

[0104] It should be noted that what the specific ratio is does not serve as a condition limiting the present invention. In practical applications, it can be set according to the specific application scenario.

[0105] It should be noted that the testing of the model adopts conventional technical means in the field of machine learning, so it will not be elaborated in detail here.

[0106] After training the light field sub-pixel disparity estimation model by the above training method, a trained light field sub-pixel disparity estimation model is obtained to utilize this trained light field sub-pixel disparity estimation model to achieve light field disparity estimation; when using this trained light field sub-pixel disparity estimation model for light field disparity estimation, that is, by inputting a target light field sub-aperture image into this trained light field sub-pixel disparity estimation model, this trained light field sub-pixel disparity estimation model will output a predicted disparity map corresponding to the target light field sub-aperture image to achieve the light field disparity estimation of the target light field sub-aperture image; the working principle adopted is the same as the working principle of the light field sub-pixel disparity estimation model introduced above, so it will not be elaborated in detail here.

[0107] It should be noted that the present invention designs a scaling factor to establish a sub-pixel disparity label sequence, and at the same time proposes a method that can accurately match pixel points and calculate the cost using dilated convolution. Finally, the costs are concatenated in the disparity dimension to complete the construction of the cost volume; the present invention also designs an edge loss for the training process of the light field sub-pixel disparity estimation model, solves the problem of fewer disparity labels in the traditional light field disparity estimation task, alleviates the influence brought by occluded pixels, realizes the calculation of light field disparity estimation at the sub-pixel level, and improves the accuracy and robustness of light field disparity estimation.

[0108] The protection scope of the training method described in the embodiments of the present invention is not limited to the execution order of the steps listed in this embodiment. Any solution achieved by adding or subtracting steps of the prior art and replacing steps according to the principle of the present invention is included in the protection scope of the present invention.

[0109] The embodiments of the present invention also provide an electronic device, which includes: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the above training method.

[0110] In one embodiment, the present invention provides a development platform (corresponding to the above electronic device) as a server system.

[0111] In one embodiment, the server system is ubuntu18.04, the GPU is NVIDIA3090, and the CPU is Intel(R)Xeon(R)Silver 4210R CPU; the experimental environment is python3.7, the deep learning framework is pytorch 1.12.0, and torchvision0.13.0.

[0112] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by an electronic device, the above training method is implemented.

[0113] Those of ordinary skill in the art can understand that all or part of the steps in the method of the above embodiments can be completed by instructing a processor through a program. The program can be stored in a computer-readable storage medium. The storage medium is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disc, and any combination thereof. The above storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid-state disk (SSD)).

[0114] An embodiment of the present invention also provides a training system. The training system can implement the training method of the present invention. However, the implementation device of the training method of the present invention includes, but is not limited to, the structure of the training system listed in this embodiment. Any structural deformation and replacement of the prior art made according to the principle of the present invention are included in the protection scope of the present invention.

[0115] As Figure 5 shown, in one embodiment, the present invention provides a training system applied to the above light field sub-pixel disparity estimation model; specifically, the training system includes:

[0116] A data acquisition module 51, configured to acquire a training data set; the training data set includes multiple light field sub-aperture images.

[0117] An input / output module 52, configured to input the light field sub-aperture image into the light field sub-pixel disparity estimation model, so that the light field sub-pixel disparity estimation model outputs a predicted disparity map.

[0118] The model training module 53 is configured to construct a loss function of the light field sub-pixel disparity estimation model according to the real disparity map and the predicted disparity map, and train the light field sub-pixel disparity estimation model based on the loss function to obtain a trained light field sub-pixel disparity estimation model.

[0119] It should be noted that the structures and principles of the data acquisition module 51, the input / output module 52, and the model training module 53 correspond one by one to the steps (steps S1 to S3) in the above training method. The specific working principle can also refer to the introduction of the training method in the foregoing embodiments, so it will not be elaborated here.

[0120] In several embodiments provided by the present invention, it should be understood that the disclosed system, device, or method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules / units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of devices or modules or units can be in electrical, mechanical, or other forms.

[0121] The modules / units described as separate components may or may not be physically separated, and the components displayed as modules / units may or may not be physical modules, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules / units can be selected according to actual needs to achieve the purpose of the embodiments of the present invention. For example, in each embodiment of the present invention, the functional modules / units can be integrated in a processing module, or each module / unit can exist physically alone, or two or more modules / units can be integrated in one module / unit.

[0122] Those of ordinary skill in the art should further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0123] The descriptions of the processes or structures corresponding to the above-mentioned various drawings each have their own focuses. For the parts not elaborated in a certain process or structure, reference may be made to the relevant descriptions of other processes or structures.

[0124] The above embodiments are only illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A light field sub-pixel disparity estimation model, characterized in that: The light field sub-pixel disparity estimation model includes: An image conversion module, used for performing image conversion on the light field sub-aperture image to obtain a light field macro-pixel image; A feature extraction module, used to extract features from the light field macro-pixel image to obtain a feature map; A cost construction module, used to perform sub-pixel cost construction and attention enhancement on the feature map to obtain an attention enhancement cost body; A cost aggregation module, used to regularize the attention enhancement cost body to obtain an aggregated cost body; The disparity regression module is used to perform disparity mapping on the aggregated cost volume to obtain a predicted disparity map corresponding to the light field sub-aperture image; the predicted disparity map is a disparity map corresponding to the central sub-aperture image in the light field sub-aperture image.

2. The light field sub-pixel disparity estimation model according to claim 1, characterized in that: The cost construction module includes: a sub-pixel disparity label establishment unit, used to establish a sub-pixel disparity label sequence based on a scaling factor; A sub-pixel disparity cost calculation unit, configured to perform cost calculation on the feature map according to the sub-pixel disparity label sequence to obtain a sub-pixel cost volume; An attention enhancement unit is used to perform attention enhancement on the sub-pixel cost volume to obtain the attention enhanced cost volume.

3. The light field sub-pixel disparity estimation model according to claim 2, characterized in that: The sub-pixel disparity cost calculation unit uses dilated convolution to perform cost calculation.

4. The light field sub-pixel disparity estimation model according to any one of claims 1 to 3, characterized in that: The feature extraction module uses weight-shared dilated convolution to perform feature extraction; and / or The cost aggregation module uses three-dimensional convolution for regularization.

5. A training method for a light field sub-pixel disparity estimation model according to any one of claims 1 to 4, characterized in that: The training method comprises: Acquire a training data set; the training data set includes a plurality of light field sub-aperture images; Inputting the light field sub-aperture image to the light field sub-pixel disparity estimation model so that the light field sub-pixel disparity estimation model outputs a predicted disparity map; A loss function of the light field sub-pixel disparity estimation model is constructed according to the real disparity map and the predicted disparity map, so as to train the light field sub-pixel disparity estimation model based on the loss function and obtain a trained light field sub-pixel disparity estimation model.

6. The training method according to claim 5, characterized in that: The obtaining of the training data set comprises: Obtaining the original light field sub-aperture image; Randomly slicing and sampling the original light field sub-aperture image to obtain light field sub-aperture image slices; Data enhancement is performed on the light field sub-aperture image slice to obtain the image slice after data enhancement; the image slice after data enhancement is used as the light field sub-aperture image.

7. The training method according to claim 5 or 6, characterized in that: The loss function includes: a mean absolute loss function and an edge loss function; wherein the edge loss function is implemented based on a Laplace operator.

8. A training system for a light field sub-pixel disparity estimation model according to any one of claims 1 to 4, characterized in that: The training system comprises: A data acquisition module, used to acquire a training data set; the training data set includes a plurality of light field sub-aperture images; An input-output module, configured to input the light field sub-aperture image into the light field sub-pixel disparity estimation model, so that the light field sub-pixel disparity estimation model outputs a predicted disparity map; The model training module is used to construct a loss function of the light field sub-pixel disparity estimation model according to the real disparity map and the predicted disparity map, so as to train the light field sub-pixel disparity estimation model based on the loss function and obtain a trained light field sub-pixel disparity estimation model.

9. An electronic device, characterized in that: The electronic device comprises: a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory so that the electronic device performs the training method according to any one of claims 5 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by an electronic device, the training method described in any one of claims 5 to 7 is implemented.