Methods, apparatus, electronic devices and storage media for stereoscopic image reconstruction
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2022-07-20
- Publication Date
- 2026-05-26
AI Technical Summary
In the existing technology, due to manufacturing errors of the binocular camera and external impacts, the left and right views move in the y-direction, resulting in inaccurate stereo image reconstruction based on parallax alignment, which affects the image reconstruction effect.
A stereo image reconstruction model is used to perform optical flow pre-alignment on the features of the left and right views, and combined with a fine alignment task, the accuracy of feature alignment is improved.
It improves the accuracy of stereo image reconstruction and achieves better super-resolution results.
Smart Images

Figure CN115147552B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and in particular to a method, apparatus, electronic device, and storage medium for stereoscopic image reconstruction. Background Technology
[0002] Stereo image super-resolution refers to image super-resolution reconstruction based on complementary spatial information in the left and right views of a stereo image.
[0003] In related technologies, stereo image super-resolution schemes are mainly divided into binocular information fusion based on attention mechanisms and binocular information fusion based on disparity (x-direction motion). Among them, binocular information fusion based on disparity alignment has better information fusion effect and is widely used. However, due to manufacturing errors, external impacts, and other reasons, binocular cameras may misalign, causing the left and right views to move in the y-direction as well. This results in inaccurate information fusion using disparity (x-direction motion), low alignment precision, and unsatisfactory image reconstruction. Therefore, achieving high-precision stereo image reconstruction has become particularly important. Summary of the Invention
[0004] This disclosure provides a stereo image reconstruction method, apparatus, electronic device, and storage medium to achieve high-precision super-resolution reconstruction of the left and right views of a stereo image.
[0005] In a first aspect, this disclosure provides a stereoscopic image reconstruction method, the application method including:
[0006] A stereo image reconstruction model is used to perform optical flow pre-alignment on the left and right view features of the stereo image to be processed to obtain pre-aligned features;
[0007] The stereo image reconstruction model is used to perform a fine alignment task on the pre-aligned features;
[0008] Based on the refined alignment task, the left and right views of the stereo image to be processed are reconstructed to obtain the reconstructed stereo image left and right views.
[0009] Secondly, this disclosure also provides a stereoscopic image reconstruction apparatus, the application apparatus comprising:
[0010] The pre-alignment module is used to perform optical flow pre-alignment on the left and right view features of the stereo image to be processed using a stereo image reconstruction model to obtain pre-aligned features;
[0011] A fine alignment module is used to perform a fine alignment task on the pre-aligned features using the stereo image reconstruction model;
[0012] The image reconstruction module is used to reconstruct the left and right views of the stereoscopic image to be processed according to the fine alignment task, so as to obtain the reconstructed stereoscopic image left and right views.
[0013] Thirdly, this disclosure also provides an electronic device, the electronic device comprising:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the stereoscopic image reconstruction method described in any of the above embodiments.
[0017] Fourthly, this disclosure also provides a computer-readable medium storing computer instructions that, when executed by a processor, implement the stereoscopic image reconstruction method described in any of the above embodiments.
[0018] The technical solution of this disclosure employs a stereo image reconstruction model to perform optical flow pre-alignment on the left and right view features of the stereo image to be processed, obtaining pre-aligned features; the stereo image reconstruction model then performs a fine alignment task on the pre-aligned features; and based on the fine alignment task, image reconstruction is performed on the left and right views of the stereo image to be processed to obtain the reconstructed stereo image left and right views. In this solution, during stereo image reconstruction, in addition to using optical flow estimation for optical flow pre-alignment, a fine alignment task is also performed on the pre-aligned features within the stereo image reconstruction model, significantly improving the accuracy of the pre-aligned features between the left and right view features of the stereo image to be processed. This, in turn, enhances the subjective effect of the reconstruction result and achieves better super-resolution of the stereo image.
[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0021] Figure 1 A flowchart of a stereo image reconstruction method provided in this disclosure embodiment;
[0022] Figure 2 This is a schematic diagram illustrating the principle of applying a stereoscopic image reconstruction model according to an embodiment of this disclosure;
[0023] Figure 3 This is a result diagram of the stereoscopic image left view reconstruction process provided in the embodiments of this disclosure;
[0024] Figure 4 This is a result diagram of the stereoscopic image right view reconstruction process provided in the embodiments of this disclosure;
[0025] Figure 5 A flowchart illustrating a pre-trained optical flow estimation network training method in a stereo image reconstruction model provided in this disclosure embodiment;
[0026] Figure 6 This is a schematic diagram of the training sample determination process provided in the embodiments of this disclosure;
[0027] Figure 7 This is a diagram showing the result of the training sample processing procedure provided in the embodiments of this disclosure;
[0028] Figure 8 This is a result diagram of another training sample processing procedure provided in an embodiment of this disclosure;
[0029] Figure 9 A flowchart of a stereo image reconstruction method provided in this disclosure embodiment;
[0030] Figure 10 A structural block diagram of a stereoscopic image reconstruction apparatus provided in this disclosure embodiment;
[0031] Figure 11 A structural block diagram of an electronic device for implementing the stereoscopic image reconstruction method of the present disclosure embodiments. Detailed Implementation
[0032] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0033] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0034] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0035] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0036] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0037] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0038] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0039] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0040] In the following embodiments, each embodiment provides optional features and examples. The various features described in the embodiments can be combined to form multiple optional solutions. Each numbered embodiment should not be regarded as only one technical solution. Furthermore, unless otherwise specified, the embodiments and features in the embodiments of this disclosure can be combined with each other.
[0041] Figure 1 This is a flowchart illustrating a stereoscopic image reconstruction method provided in this embodiment. The technical solution of this embodiment is applicable to achieving high-precision super-resolution reconstruction of the left and right views of a stereoscopic image. This method can be executed by a stereoscopic image reconstruction device, which can be implemented by software and / or hardware, and is generally integrated into any electronic device with network communication capabilities, including but not limited to: computers, personal digital assistants, etc. Figure 1 As shown, the stereoscopic image reconstruction method of this embodiment may include the following steps S110-S130:
[0042] S110. Using a stereo image reconstruction model, perform optical flow pre-alignment on the left and right view features of the stereo image to be processed to obtain pre-aligned features.
[0043] The stereo image reconstruction model refers to a model that performs super-resolution reconstruction of the left and right views of a stereo image based on complementary spatial information in the left and right views. The input of the stereo image reconstruction model is the left and right views of a low-resolution stereo image, and the output is the left and right views of a super-resolution stereo image. The stereo image reconstruction model can perform optical flow pre-alignment on the features of the left and right views of the stereo image to be processed, obtaining pre-aligned features between the features of the left and right views of the stereo image to be processed.
[0044] In this scheme, feature extraction processing can be performed on the stereo image based on image processing technology to obtain the left and right view features of the stereo image to be processed, which are denoted as the left view feature and the right view feature of the stereo image to be processed, respectively. Then, optical flow pre-alignment is performed on the left view feature and the right view feature of the stereo image to be processed using a stereo image reconstruction model.
[0045] In this technical solution, optionally, a stereo image reconstruction model is used to perform an optical flow pre-alignment task on the left and right view features of the stereo image to be processed to obtain pre-aligned features, which may include the following process:
[0046] Based on the pre-trained optical flow estimation network in the stereo image reconstruction model, an optical flow pre-alignment task is performed on the left and right view features of the stereo image to be processed to obtain pre-aligned features.
[0047] In this scheme, the pre-trained optical flow estimation network in the stereo image reconstruction model is used to estimate the optical flow information of the left and right views of the stereo image, so as to improve the accuracy of the reconstructed super-resolution stereo image's left and right views. Specifically, optical flow estimation can be performed on the right view of the stereo image based on the left view, and vice versa.
[0048] In this scheme, the left and right view features of the stereo image to be processed are used as inputs and fed into the pre-trained optical flow estimation network of the stereo image reconstruction model. The optical flow estimation network in the stereo image reconstruction model determines the optical flow information between the left and right view features of the stereo image to be processed. Based on the optical flow information between the left and right view features of the stereo image to be processed, the left view features and the right view features of the stereo image to be processed are pre-aligned, and the pre-aligned features between the left and right view features of the stereo image to be processed are output.
[0049] Optionally, in this embodiment, the pre-trained optical flow estimation network may include the Flownet series optical flow estimation network, the PWC-net series optical flow estimation network, etc. The Flownet series optical flow estimation network uses an end-to-end convolutional neural network to solve the optical flow estimation problem. The PWC-net series optical flow estimation network uses three main components: an image pyramid, mapping, and matching cost capacity calculation. The mapping and matching cost capacity calculation does not require training parameters, which can reduce the number of model parameters.
[0050] In this technical solution, optionally, the optical flow estimation network in the stereo image reconstruction model is used to perform an optical flow pre-alignment task on the left and right view features of the stereo image to be processed to obtain pre-aligned features, which may include steps A1-A2:
[0051] Step A1: Input the left and right views of the stereo image to be processed into the pre-trained optical flow estimation network in the stereo image reconstruction model to obtain the optical flow estimation value between the left and right views of the stereo image to be processed.
[0052] Step A2: Based on the estimated optical flow between the left and right views of the stereo image to be processed, pre-align the features of the left and right views of the stereo image to be processed to obtain the pre-aligned features between the features of the left and right views of the stereo image to be processed.
[0053] The alignment features include features obtained by aligning the features from the left view of the stereo image to the right view of the stereo image, and features obtained by aligning the features from the right view of the stereo image to the left view of the stereo image.
[0054] For example, Figure 2 This is a schematic diagram illustrating the principle of applying a stereoscopic image reconstruction model according to an embodiment of this disclosure, such as... Figure 2 As shown, with I l I represents the left view of the stereoscopic image to be processed. r Taking the right view of the stereoscopic image to be processed as an example. Let I... l and I r The inputs are fed into the pre-trained optical flow estimation network in the stereo image reconstruction model. The pre-trained optical flow estimation network can estimate the optical flow between the left and right views of the stereo image to be processed, and obtain the estimated optical flow value F from the left view to the right view of the stereo image to be processed. lr The estimated optical flow F from the right view of the stereo image to the left view of the stereo image to be processed. rl .
[0055] like Figure 2 As shown, according to F lr and F rl Features f of the left and right views of the stereo image to be processed l and f rPre-alignment can be performed to obtain the pre-aligned features f between the left and right view features of the stereo image to be processed. r2l and f l2r The specific process can be as follows: using F lr and F rl The left view features f of the stereo image to be processed are respectively l Features f of the right view of the stereo image to be processed r Pre-alignment is performed from the left view features of the stereo image to the right view features of the stereo image, and pre-alignment is performed from the right view features of the stereo image to the left view features of the stereo image, respectively, to obtain the pre-aligned features f between the left and right view features of the stereo image to be processed. r2l and f l2r Among them, f r2l This represents the alignment feature obtained by pre-aligning the features from the right view of the stereo image to the left view of the stereo image; f l2r This represents the alignment feature obtained by pre-aligning the features from the left view of the stereo image to the right view of the stereo image.
[0056] For the above optional method, the optical flow estimation value between the left and right views of the stereo image to be processed is used to pre-align the features of the left and right views of the stereo image to be processed, thereby obtaining the pre-aligned features between the features of the left and right views of the stereo image to be processed. This improves the accuracy of the alignment features, and thus enhances the subjective effect of the stereo image reconstruction model construction results when using the alignment features for subsequent image reconstruction.
[0057] S120. A stereo image reconstruction model is used to perform a fine alignment task on the pre-aligned features.
[0058] The fine-grained alignment task is used to perform secondary alignment on the pre-aligned features. That is, the fine-grained alignment task is used in conjunction with a pre-trained optical flow estimation network to improve the accuracy of the pre-aligned features between the left and right view features of the stereo image to be processed, thereby improving the subjective effect of the reconstruction result and obtaining a better super-resolution stereo image.
[0059] S130. Based on the fine alignment task, perform image reconstruction on the left and right views of the stereo image to be processed to obtain the reconstructed stereo image left and right views.
[0060] By performing a fine-grained alignment task, complementary spatial information in the left and right views of a stereo image can be accurately obtained. This complementary spatial information can then be incorporated into the reconstruction process of the left and right views of the stereo image to be processed, thereby improving the subjective effect of the reconstruction results and enabling the stereo image reconstruction model to output a better super-resolution effect.
[0061] In this technical solution, optionally, the process of reconstructing the left and right views of the stereoscopic image to be processed based on the fine alignment task to obtain the reconstructed stereoscopic image can include the following steps:
[0062] The refined alignment features obtained from the refined alignment task and the left and right view features of the stereo image to be processed are input into the fusion reconstruction network of the stereo image reconstruction model to obtain the reconstructed stereo image left and right views after super-resolution reconstruction of the stereo image left and right views.
[0063] refer to Figure 2 , with I l I represents the left view of the stereoscopic image to be processed. r Taking the right view of the stereo image to be processed as an example, the fine-grained alignment feature f is... or2l and f ol2r Features f of the left and right views of the stereo image to be processed l and f r The fusion reconstruction network input into the stereo image reconstruction model can obtain the left and right views S after super-resolution reconstruction of the left and right views of the stereo image to be processed. l and S r Specifically, the refined alignment feature f or2l and the left view features f of the stereo image to be processed l The input is fed into a fusion reconstruction network of the stereo image reconstruction model. Based on the loss function value of the stereo image reconstruction model, the parameters of the stereo image reconstruction model are adjusted, and the super-resolution reconstructed left view S is output. l ; and, refine the alignment feature f ol2r and the right view features f of the stereo image to be processed r The input is a fusion reconstruction network for a stereo image reconstruction model, and the output is a super-resolution reconstructed right view S. r .
[0064] By using a fusion reconstruction network to fuse and reconstruct refined alignment features with the left and right view features of the stereo image to be processed, and by incorporating complementary spatial information from the left and right views of the stereo image into the reconstruction process using refined alignment features, the subjective effect of the reconstruction results can be improved, and the stereo image reconstruction model can output a better super-resolution effect of the stereo image.
[0065] For example, Figure 3 This is a result image of the stereoscopic image left view reconstruction process provided in the embodiments of this disclosure, specifically as follows: Figure 3 As shown in (a), (b), (c), and (d); Figure 4 This is a result image of the stereoscopic image right view reconstruction process provided in the embodiments of this disclosure, specifically as follows: Figure 4 As shown in (a), (b), (c), and (d). Figure 3and Figure 4 In the diagram, (a), (b), (c), and (d) correspond to the results obtained in the following processes: determining the left and right view features of the stereo image to be processed; using a stereo image reconstruction model to perform an optical flow pre-alignment task on the left and right view features of the stereo image to be processed to obtain pre-aligned features; using a stereo image reconstruction model to perform a fine alignment task on the pre-aligned features; and based on the fine alignment task, performing image reconstruction on the left and right views of the stereo image to be processed to obtain the reconstructed stereo image left and right views.
[0066] The technical solution of this disclosure employs a stereo image reconstruction model to perform optical flow pre-alignment on the left and right view features of the stereo image to be processed, obtaining pre-aligned features; then, the stereo image reconstruction model performs a fine alignment task on the pre-aligned features; and finally, based on the fine alignment task, image reconstruction is performed on the left and right views of the stereo image to be processed to obtain the reconstructed stereo image left and right views. In the stereo image reconstruction process, this disclosure, in addition to using optical flow estimation for optical flow pre-alignment, also performs a fine alignment task on the pre-aligned features in the stereo image reconstruction model, greatly improving the accuracy of the pre-aligned features between the left and right view features of the stereo image to be processed, thereby improving the subjective effect of the reconstruction result and achieving better super-resolution of the stereo image.
[0067] Based on the above technical solution in this embodiment, optionally, the stereo image reconstruction model training process may include the following steps B1-B3:
[0068] Step B1: Based on the pre-trained optical flow estimation network in the stereo image reconstruction model, pre-align the left and right view features of the training stereo image to obtain pre-aligned features.
[0069] In this scheme, the left and right views of the training stereo image are input into the pre-trained optical flow estimation network in the stereo image reconstruction model to obtain the optical flow estimation value between the left and right views of the training stereo image; the features of the left and right views of the training stereo image are pre-aligned based on the optical flow estimation value between the left and right views of the training stereo image to obtain the pre-aligned features between the features of the left and right views of the training stereo image; wherein, the alignment features include features obtained by aligning the features of the left view of the stereo image to the features of the right view of the stereo image and features obtained by aligning the features of the right view of the stereo image to the features of the left view of the stereo image.
[0070] For example, see still Figure 2 As shown, with I l I represents the left view of the stereo image used for training. r Taking the right view of the 3D image used for training as an example, let I... l and I rThe inputs are fed into the pre-trained optical flow estimation network in the stereo image reconstruction model. The pre-trained optical flow estimation network can estimate the optical flow between the left and right views of the training stereo image, obtaining the estimated optical flow value F from the left view to the right view of the training stereo image. lr The optical flow estimate F from the right view of the training stereo image to the left view of the training stereo image. rl .
[0071] like Figure 2 As shown, according to F lr and F rl Features f of left and right views of the training stereo image l and f r By performing pre-alignment, we can obtain the pre-aligned features f between the left and right view features of the stereo image used for training. r2l and f l2r The specific process can be as follows: using F lr and F rl The left view features f of the training stereo image are respectively analyzed. l Features f of the right view of the stereo image used for training r Pre-alignment is performed from the left view features of the stereo image to the right view features of the stereo image, and pre-alignment is performed from the right view features of the stereo image to the left view features of the stereo image, respectively, to obtain the pre-aligned features f between the left and right view features of the training stereo image. r2l and f l2r Among them, f r2l This represents the alignment feature obtained by pre-aligning the features from the right view of the stereo image to the left view of the stereo image; f l2r This represents the alignment feature obtained by pre-aligning the features from the left view of the stereo image to the right view of the stereo image.
[0072] Step B2: Control the stereo image reconstruction model to perform a fine alignment training task on the pre-aligned features.
[0073] The refined alignment training task is used to perform secondary alignment on the pre-aligned features. That is, the refined alignment training task is used in conjunction with the pre-trained optical flow estimation network to improve the accuracy of the pre-aligned features between the left and right view features of the training stereo image, thereby improving the subjective effect of the reconstruction result and obtaining a better super-resolution stereo image.
[0074] Step B3: Adjust the parameters of the stereo image reconstruction model according to the fine alignment training task to obtain the trained stereo image reconstruction model.
[0075] In this scheme, the refined alignment features output from the refined alignment training task and the left and right view features of the training stereo image are input into the fusion reconstruction network of the stereo image reconstruction model to obtain the left and right views reconstructed by super-resolution reconstruction of the stereo image's left and right views. Based on the super-resolution reconstructed left and right views and the pre-annotated left and right views of the training stereo image, the loss function value of the stereo image reconstruction model is determined. The parameters of the stereo image reconstruction model are adjusted based on the loss function value to obtain the trained stereo image reconstruction model. The loss function can be set according to the requirements of the stereo image reconstruction model. For example, the loss function can be set to an absolute value loss function, a squared loss function, etc.
[0076] refer to Figure 2 , with I l I represents the left view of the stereo image used for training. r Taking the right view of the 3D image used for training as an example, the refined alignment feature f is... or2l and f ol2r Features f of left and right views of the training stereo image l and f r The fusion reconstruction network input into the stereo image reconstruction model can obtain the left and right views S after super-resolution reconstruction of the left and right views of the stereo image. l and S r Specifically, the refined alignment feature f or2l and training with left view features f of stereo images l The input is fed into a fusion reconstruction network of the stereo image reconstruction model. Based on the loss function value of the stereo image reconstruction model, the parameters of the stereo image reconstruction model are adjusted, and the super-resolution reconstructed left view S is output. l ; and, refine the alignment feature f ol2r and training stereo image right view features f r The input is a fusion reconstruction network for a stereo image reconstruction model, and the output is a super-resolution reconstructed right view S. r .
[0077] By using a fusion reconstruction network to fuse and reconstruct refined alignment features with left and right view features of the training stereo image, the parameters of the stereo image reconstruction model can be adjusted, thereby obtaining a stereo image reconstruction model with better super-resolution output.
[0078] Figure 5 This is a flowchart illustrating a training method for a pre-trained optical flow estimation network in a stereo image reconstruction model, provided by an embodiment of this disclosure. The technical solution of this embodiment further optimizes the determination process of the pre-trained optical flow estimation network based on the above embodiments. This embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 5As shown, the pre-trained optical flow estimation network training method in the stereo image reconstruction model of this embodiment may include the following steps S510-S520:
[0079] S510. Determine the training samples used by the optical flow estimation network; each training sample includes two training images of the real scene and the optical flow pre-labeled values between the two training images.
[0080] In this process, images from the stereo image dataset can be used as the source image set for synthesizing the optical flow dataset, and the source image set can be processed to obtain the training samples used by the optical flow estimation network.
[0081] In this technical solution, optionally, determining the training samples used by the optical flow estimation network may include steps C1-C4:
[0082] Step C1: Determine the optical flow information corresponding to the background image and the foreground image selected from the source image set; wherein the background image and the foreground image are non-repeating images, and the source image set is generated based on real scene images in the stereo image super-resolution dataset.
[0083] Step C2: Select a local region from the foreground image to obtain a local foreground image, and determine the optical flow information corresponding to the local foreground image.
[0084] Step C3: Perform forward transformation on the background image and the foreground local image using their respective optical flow information to obtain the transformed background image and foreground local image.
[0085] Step C4: The background image before conversion and the foreground local image are superimposed in the direction from background to foreground to obtain the first training image. The converted background image and the foreground local image are superimposed in the direction from background to foreground to obtain the second training image. The optical flow information corresponding to the background image and the foreground local image are superimposed to obtain the pre-labeled optical flow information between the two training images, so as to construct the training samples used to train the optical flow estimation network.
[0086] For example, Figure 6 This is a schematic diagram of the training sample determination process provided in the embodiments of this disclosure, as shown below. Figure 6 As shown, images from the stereo image dataset are used as the source image set for synthesizing the optical flow dataset. K unique images are randomly selected from this image set. in, The other K-1 images serve as the foreground images, with one as the background. Perform homography transformation to obtain the corresponding optical flow information F1; Homography transformation is also performed to obtain the corresponding optical flow information {F}. k |k∈[1,2…,K]}. In each foreground image By using a local polygonal region or a local circular region, a foreground local polygonal image or a foreground local circular image can be obtained. The sampled local region has a mask value of 1, while other regions have a mask value of 0. Based on the obtained... F k Through formula Calculate the converted This yields the transformed background image and foreground local image. Here, f represents the optical flow-based forward transformation operation. By layering the image and optical flow from the background to the foreground, the training samples used by the final synthetic optical flow estimation network can be obtained.
[0087] By determining the training samples used by the optical flow estimation network, a real-world optical flow dataset can be constructed, improving the accuracy of the pre-trained optical flow estimation network in predicting optical flow.
[0088] For example, Figure 7 This is a result diagram of the training sample processing procedure provided in the embodiments of this disclosure, such as... Figure 7 As shown, (a) is the original image, and (b) is the image after degradation. To make the generated optical flow dataset more closely resemble the real image, some degradation methods can be added to the image, such as blurring and compression.
[0089] For example, Figure 8 This is a result diagram of another training sample processing procedure provided in this embodiment of the disclosure, such as... Figure 8 As shown, (a) is the background image, (b) is the foreground image, (c) is the foreground local polygon image, and (d) is the training sample used by the optical flow estimation network. By overlaying the image and optical flow layer by layer from the background to the foreground, the training sample used by the final synthesized optical flow estimation network can be obtained.
[0090] S520. Based on the training samples, control the optical flow estimation network to be trained in the stereo image reconstruction model to perform optical flow estimation training task, and obtain the pre-trained optical flow estimation network in the stereo image reconstruction model.
[0091] In this scheme, the ability of the pre-trained optical flow estimation network in the stereo image reconstruction model to output optical flow estimates can be trained based on the loss information when performing the optical flow estimation training task, so that the optical flow estimates output by the pre-trained optical flow estimation network in the stereo image reconstruction model are closer to the true optical flow values.
[0092] In this technical solution, optionally, the optical flow estimation network to be trained in the stereo image reconstruction model is controlled to perform an optical flow estimation training task based on the training samples to obtain the pre-trained optical flow estimation network in the stereo image reconstruction model, which may include steps D1-D3:
[0093] Step D1: Input two training images from the training samples into the optical flow estimation network to be trained in the stereo image reconstruction model to perform optical flow estimation training task and obtain the optical flow estimation value between the two training images in the training samples.
[0094] Step D2: Determine the loss function value for the optical flow estimation training task based on the optical flow estimation value between two training images in the training samples and the optical flow pre-labeling value between two training images.
[0095] Step D3: Adjust the parameters of the optical flow estimation network to be trained based on the loss function value of the optical flow estimation training task to obtain the pre-trained optical flow estimation network.
[0096] The loss function is used to evaluate the degree to which the predicted values of the pre-trained optical flow estimation network differ from the true values. A better loss function indicates better performance of the pre-trained optical flow estimation network. The loss function can be an absolute value loss function or a squared loss function. This embodiment does not impose a specific limitation.
[0097] In this scheme, the optical flow estimation capability of the optical flow estimation network output image in the stereo image reconstruction model can be trained based on two training images in the training samples. At the same time, the loss function value of the optical flow estimation training task is introduced to optimize the optical flow estimation value output by the pre-trained optical flow estimation network in the stereo image reconstruction model to the optical flow pre-labeled value, so that the optical flow estimation value output by the pre-trained optical flow estimation network in the stereo image reconstruction model is closer to the true optical flow value.
[0098] By inputting two training images from the training samples into the optical flow estimation network to be trained in the stereo image reconstruction model, the optical flow estimation network to be trained is carried out. The parameters of the optical flow estimation network to be trained are adjusted to obtain the pre-trained optical flow estimation network, which can improve the accuracy of the pre-trained optical flow estimation network in predicting optical flow.
[0099] The technical solution of this disclosure determines the training samples used by the optical flow estimation network, and controls the optical flow estimation network to be trained in the stereo image reconstruction model to perform an optical flow estimation training task based on the training samples, thereby obtaining a pre-trained optical flow estimation network in the stereo image reconstruction model. By implementing this technical solution, the pre-trained optical flow estimation network can construct an optical flow dataset in a real scene, improving the accuracy of the optical flow estimation network in predicting optical flow. The results of the optical flow alignment are then finely aligned using a fine-grained alignment training task, improving the accuracy of the alignment and thus enhancing the subjective effect of the reconstruction results.
[0100] Figure 9 This is a flowchart of a stereo image reconstruction method provided in this embodiment. The technical solution of this embodiment further optimizes the process of performing fine-grained alignment on pre-aligned features based on the above embodiments. This embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 9 As shown, the stereoscopic image reconstruction method of this embodiment may include the following steps S910-S930:
[0101] S910. A stereo image reconstruction model is used to perform optical flow pre-alignment on the left and right view features of the stereo image to be processed to obtain pre-aligned features.
[0102] S920. The pre-aligned features are refined by using the pre-trained refined alignment network in the stereo image reconstruction model to obtain refined aligned features.
[0103] Specifically, when training the pre-trained refined alignment network, the learning rate of the pre-trained optical flow estimation network is within a preset learning rate range, and the learning rate of the pre-trained optical flow estimation network is lower than the learning rate of the pre-trained refined alignment network. The refined alignment network includes a deformable convolutional network.
[0104] In this embodiment, the preset learning rate range can be set according to the learning requirements of the fine-grained alignment network. Optionally, the learning rate of the pre-trained optical flow estimation network can be set between [5e-7, 1e-5].
[0105] In this technical solution, optionally, the pre-aligned features are finely aligned by using a pre-trained fine-alignment network of the stereo image reconstruction model, which may include steps E1-E2:
[0106] Step E1: Input the pre-aligned features and the left and right view features of the stereo image to be processed into the offset estimation network of the refined alignment network to perform offset estimation and determine the offset of the pre-aligned features.
[0107] refer to Figure 2 , with I l I represents the left view of the stereoscopic image to be processed. r Taking the right view of the stereo image to be processed as an example, the pre-aligned feature f r2l and f l2r Features f of the left and right views of the stereo image to be processed l and f r The offset estimation network of the refined alignment network is used for offset estimation to obtain the offset O of the pre-aligned features. l and O rThis means estimating r×r motion vectors at each pixel location. Specifically, the pre-aligned features f r2l and the left view features f of the stereo image to be processed l The offset estimation network input into the fine alignment network performs offset estimation to obtain the offset O between the left view and the right view of the stereo image to be processed. l ; and, the pre-aligned feature f l2r and the right view features f of the stereo image to be processed r The offset estimation network input into the fine-grained alignment network performs offset estimation to obtain the offset O between the right view and the left view of the stereo image to be processed. r .
[0108] Step E2: Input the offset of the pre-aligned feature and the pre-aligned feature into the deformable convolutional network of the fine alignment network to perform fine alignment of the pre-aligned feature, so as to obtain the fine alignment feature between the left and right view features of the stereo image to be processed.
[0109] Continue to refer to Figure 2 , with I l I represents the left view of the stereoscopic image to be processed. r Taking the right view of the stereo image to be processed as an example, the offset O of the pre-aligned feature is obtained. l and O r Then, set the offset O l and pre-alignment feature f r2l The pre-aligned features are finely aligned by a deformable convolutional network fed into a fine-alignment network, resulting in fine-aligned features f between the left and right view features of the stereo image to be processed. or2l ; and, offset O r and pre-alignment feature f l2r The pre-aligned features are finely aligned by a deformable convolutional network fed into a fine-alignment network, resulting in fine-aligned features f between the left and right view features of the stereo image to be processed. ol2r .
[0110] By inputting the optical flow aligned result into a deformable convolutional network for secondary fine alignment, the alignment accuracy can be improved, thereby enhancing the subjective results of the stereo image reconstruction model.
[0111] S930. Based on the refined alignment task, perform image reconstruction on the left and right views of the stereoscopic image to be processed to obtain the reconstructed stereoscopic image left and right views.
[0112] The technical solution of this disclosure employs a stereo image reconstruction model to perform optical flow pre-alignment on the left and right view features of the stereo image to be processed, obtaining pre-aligned features. Then, a pre-trained fine-alignment network in the stereo image reconstruction model is used to fine-align the pre-aligned features, resulting in fine-aligned features. Based on the fine-alignment task, image reconstruction is performed on the left and right views of the stereo image to be processed, yielding the reconstructed stereo image's left and right views. By implementing this technical solution, the accuracy of alignment is improved by using deformable convolution to fine-align the optical flow alignment results, thereby enhancing the subjective effect of the reconstruction results and achieving better super-resolution of the stereo image.
[0113] Figure 10 This is a structural block diagram of a stereoscopic image reconstruction device provided in an embodiment of this disclosure. The technical solution of this embodiment is applicable to situations where high-precision super-resolution reconstruction of the left and right views of a stereoscopic image is achieved. This device can be implemented by software and / or hardware and is generally integrated into any electronic device with network communication capabilities, including but not limited to: computers, personal digital assistants, etc. Figure 11 As shown, the stereoscopic image reconstruction apparatus of this embodiment may include the following: a pre-alignment module 1010, a fine alignment module 1020, and an image reconstruction module 1030, wherein:
[0114] The pre-alignment module 1010 is used to perform optical flow pre-alignment on the left and right view features of the stereo image to be processed using a stereo image reconstruction model to obtain pre-aligned features;
[0115] The fine alignment module 1020 is used to perform a fine alignment task on the pre-aligned features using the stereo image reconstruction model;
[0116] The image reconstruction module 1030 is used to perform image reconstruction on the left and right views of the stereoscopic image to be processed according to the fine alignment task to obtain the reconstructed stereoscopic image left and right views.
[0117] Based on the above embodiments, optionally, the pre-alignment module 1010 includes:
[0118] The pre-aligned feature acquisition unit is used to perform optical flow pre-alignment task on the left and right view features of the stereo image to be processed based on the pre-trained optical flow estimation network in the stereo image reconstruction model to obtain pre-aligned features.
[0119] Based on the above embodiments, optionally, the pre-alignment feature obtaining unit includes:
[0120] The training sample determination subunit is used to determine the training samples used by the optical flow estimation network; each training sample includes two training images of the real scene and optical flow pre-labeled values between the two training images.
[0121] The pre-trained optical flow estimation network obtains sub-units, which are used to control the optical flow estimation network to be trained in the stereo image reconstruction model to perform optical flow estimation training tasks based on the training samples, thereby obtaining the pre-trained optical flow estimation network in the stereo image reconstruction model.
[0122] Based on the above embodiments, optionally, the training samples determine the sub-units, specifically for:
[0123] Determine the optical flow information corresponding to the background image and the foreground image selected from the source image set; wherein the background image and the foreground image are non-repeating images, and the source image set is generated based on real scene images in the stereo image super-resolution dataset;
[0124] A local foreground image is obtained by selecting a local region from the foreground image, and the optical flow information corresponding to the local foreground image is determined.
[0125] The background image and the foreground local image are forward transformed using their respective optical flow information to obtain the transformed background image and foreground local image;
[0126] The first training image is obtained by superimposing the background image and the foreground local image before conversion in the direction from background to foreground. The second training image is obtained by superimposing the background image and the foreground local image after conversion in the direction from background to foreground. The pre-labeled optical flow information between the two training images is obtained by superimposing the optical flow information corresponding to the background image and the foreground local image respectively, so as to construct the training samples used to train the optical flow estimation network.
[0127] Based on the above embodiments, optionally, the pre-trained optical flow estimation network obtains sub-units, specifically used for:
[0128] Two training images from the training samples are input into the optical flow estimation network to be trained in the stereo image reconstruction model to perform optical flow estimation training task, and the optical flow estimation value between the two training images in the training samples is obtained.
[0129] Based on the optical flow estimation value between two training images in the training samples and the optical flow pre-labeled value between two training images, the loss function value of the optical flow estimation training task is determined;
[0130] The parameters of the optical flow estimation network to be trained are adjusted based on the loss function value of the optical flow estimation training task to obtain the pre-trained optical flow estimation network.
[0131] Based on the above embodiments, optionally, the pre-alignment feature obtaining unit further includes:
[0132] The optical flow estimation value is obtained by sub-unit, which is used to input the left and right views of the stereo image to be processed into the pre-trained optical flow estimation network in the stereo image reconstruction model to obtain the optical flow estimation value between the left and right views of the stereo image to be processed.
[0133] The pre-aligned feature is obtained as a sub-unit, which is used to pre-align the features of the left and right views of the stereo image to be processed based on the optical flow estimation value between the left and right views of the stereo image to be processed, and obtain the pre-aligned features between the features of the left and right views of the stereo image to be processed.
[0134] The alignment features include features obtained by aligning the features from the left view of the stereo image to the right view of the stereo image, and features obtained by aligning the features from the right view of the stereo image to the left view of the stereo image.
[0135] Based on the above embodiments, optionally, the fine alignment module 1020 includes:
[0136] The refined alignment feature acquisition unit is used to refine the pre-aligned features by refining the pre-aligned features through the pre-trained refined alignment network in the stereo image reconstruction model to obtain refined alignment features;
[0137] Specifically, when training the pre-trained refined alignment network, the learning rate of the pre-trained optical flow estimation network is within a preset learning rate range, and the learning rate of the pre-trained optical flow estimation network is lower than the learning rate of the pre-trained refined alignment network. The refined alignment network includes a deformable convolutional network.
[0138] Based on the above embodiments, optionally, the refined alignment feature obtaining unit is specifically used for:
[0139] The pre-aligned features and the left and right view features of the stereo image to be processed are input into the offset estimation network of the refined alignment network for offset estimation to determine the offset of the pre-aligned features.
[0140] The offset of the pre-aligned feature and the pre-aligned feature are input into the deformable convolutional network of the fine alignment network to perform fine alignment of the pre-aligned feature, thereby obtaining the fine alignment feature between the left and right view features of the stereo image to be processed.
[0141] Based on the above embodiments, optionally, the image reconstruction module 1030 includes:
[0142] The refined alignment features obtained from the refined alignment task and the left and right view features of the stereo image to be processed are input into the fusion reconstruction network of the stereo image reconstruction model to obtain the reconstructed stereo image left and right views after super-resolution reconstruction of the stereo image left and right views.
[0143] The stereoscopic image reconstruction apparatus provided in this embodiment can execute a stereoscopic image reconstruction method provided in any of the embodiments of this disclosure, and has the corresponding functions and beneficial effects of executing the stereoscopic image reconstruction method. For details, please refer to the relevant operations of a stereoscopic image reconstruction method in the foregoing embodiments.
[0144] The following is for reference. Figure 11 The diagram illustrates a structural schematic of an electronic device 1100 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 11 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0145] like Figure 11 As shown, electronic device 1100 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1101, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1102 or a program loaded from storage device 1106 into random access memory (RAM) 1103. The RAM 1103 also stores various programs and data required for the operation of electronic device 1100. The processing device 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.
[0146] Typically, the following devices can be connected to I / O interface 1105: input devices 1106 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 1107 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1106 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1109. Communication device 1109 allows electronic device 1100 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 11 An electronic device 1100 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0147] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the stereoscopic image reconstruction method shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1109, or installed from storage device 1106, or installed from ROM 1102. When the computer program is executed by processing device 1101, it performs the functions defined above in the stereoscopic image reconstruction method of embodiments of this disclosure.
[0148] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0149] The electronic device provided in this embodiment and the stereoscopic image reconstruction method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0150] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the stereoscopic image reconstruction method provided in the above embodiments.
[0151] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0152] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0153] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0154] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: perform an optical flow pre-alignment task on the left and right view features of the stereo image to be processed using a stereo image reconstruction model to obtain pre-aligned features; perform a fine alignment task on the pre-aligned features using the stereo image reconstruction model; and perform image reconstruction on the left and right views of the stereo image to be processed based on the fine alignment task to obtain the reconstructed stereo image left and right views.
[0155] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0157] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the units are not necessarily limiting in certain circumstances; for example, a pre-alignment module can also be described as "using a stereo image reconstruction model to perform optical flow pre-alignment on the left and right view features of the stereo image to be processed to obtain pre-aligned features."
[0158] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0159] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0160] According to one or more embodiments of this disclosure, Example 1 provides a stereo image reconstruction method, the application method comprising:
[0161] A stereo image reconstruction model is used to perform optical flow pre-alignment on the left and right view features of the stereo image to be processed to obtain pre-aligned features;
[0162] The stereo image reconstruction model is used to perform a fine alignment task on the pre-aligned features;
[0163] Based on the refined alignment task, the left and right views of the stereo image to be processed are reconstructed to obtain the reconstructed stereo image left and right views.
[0164] Example 2, based on the method described in Example 1, uses a stereo image reconstruction model to perform an optical flow pre-alignment task on the left and right view features of the stereo image to be processed to obtain pre-aligned features, including:
[0165] Based on the pre-trained optical flow estimation network in the stereo image reconstruction model, an optical flow pre-alignment task is performed on the left and right view features of the stereo image to be processed to obtain pre-aligned features.
[0166] Example 3: According to the method described in Example 2, the process of determining the pre-trained optical flow estimation network includes:
[0167] Determine the training samples used by the optical flow estimation network; each training sample includes two training images of the real scene and the optical flow pre-labeled values between the two training images.
[0168] Based on the training samples, the optical flow estimation network to be trained in the stereo image reconstruction model is controlled to perform an optical flow estimation training task, thereby obtaining the pre-trained optical flow estimation network in the stereo image reconstruction model.
[0169] Example 4, based on the method described in Example 3, determines the training samples used by the optical flow estimation network, including:
[0170] Determine the optical flow information corresponding to the background image and the foreground image selected from the source image set; wherein the background image and the foreground image are non-repeating images, and the source image set is generated based on real scene images in the stereo image super-resolution dataset;
[0171] A local foreground image is obtained by selecting a local region from the foreground image, and the optical flow information corresponding to the local foreground image is determined.
[0172] The background image and the foreground local image are forward transformed using their respective optical flow information to obtain the transformed background image and foreground local image;
[0173] The first training image is obtained by superimposing the background image and the foreground local image before conversion in the direction from background to foreground. The second training image is obtained by superimposing the background image and the foreground local image after conversion in the direction from background to foreground. The pre-labeled optical flow information between the two training images is obtained by superimposing the optical flow information corresponding to the background image and the foreground local image respectively, so as to construct the training samples used to train the optical flow estimation network.
[0174] Example 5, according to the method described in Example 3, controls the optical flow estimation network to be trained in the stereo image reconstruction model to perform an optical flow estimation training task based on the training samples, to obtain the pre-trained optical flow estimation network in the stereo image reconstruction model, including:
[0175] Two training images from the training samples are input into the optical flow estimation network to be trained in the stereo image reconstruction model to perform optical flow estimation training task, and the optical flow estimation value between the two training images in the training samples is obtained.
[0176] Based on the optical flow estimation value between two training images in the training samples and the optical flow pre-labeled value between two training images, the loss function value of the optical flow estimation training task is determined;
[0177] The parameters of the optical flow estimation network to be trained are adjusted based on the loss function value of the optical flow estimation training task to obtain the pre-trained optical flow estimation network.
[0178] Example 6, according to the method described in Example 2, performs an optical flow pre-alignment task on the left and right view features of the stereo image to be processed, based on the pre-trained optical flow estimation network in the stereo image reconstruction model, to obtain pre-aligned features, including:
[0179] The left and right views of the stereo image to be processed are input into the pre-trained optical flow estimation network in the stereo image reconstruction model to obtain the optical flow estimation value between the left and right views of the stereo image to be processed.
[0180] Based on the estimated optical flow between the left and right views of the stereo image to be processed, the features of the left and right views of the stereo image to be processed are pre-aligned to obtain the pre-aligned features between the features of the left and right views of the stereo image to be processed.
[0181] The alignment features include features obtained by aligning the features from the left view of the stereo image to the right view of the stereo image, and features obtained by aligning the features from the right view of the stereo image to the left view of the stereo image.
[0182] Example 7, according to the method described in Example 1, uses the stereo image reconstruction model to perform a fine alignment task on the pre-aligned features, including:
[0183] The pre-aligned features are refined by aligning the pre-aligned features using a pre-trained refined alignment network in the stereo image reconstruction model to obtain refined aligned features.
[0184] Specifically, when training the pre-trained refined alignment network, the learning rate of the pre-trained optical flow estimation network is within a preset learning rate range, and the learning rate of the pre-trained optical flow estimation network is lower than the learning rate of the pre-trained refined alignment network. The refined alignment network includes a deformable convolutional network.
[0185] Example 8, according to the method described in Example 7, involves refining the pre-aligned features using a pre-trained fine-tuning alignment network of the stereo image reconstruction model to obtain fine-tuned aligned features, including:
[0186] The pre-aligned features and the left and right view features of the stereo image to be processed are input into the offset estimation network of the refined alignment network for offset estimation to determine the offset of the pre-aligned features.
[0187] The offset of the pre-aligned feature and the pre-aligned feature are input into the deformable convolutional network of the fine alignment network to perform fine alignment of the pre-aligned feature, thereby obtaining the fine alignment feature between the left and right view features of the stereo image to be processed.
[0188] Example 9, according to the method described in Example 1, involves reconstructing the left and right views of the stereoscopic image to be processed based on the refined alignment task to obtain the reconstructed stereoscopic image's left and right views, including:
[0189] The refined alignment features obtained from the refined alignment task and the left and right view features of the stereo image to be processed are input into the fusion reconstruction network of the stereo image reconstruction model to obtain the reconstructed stereo image left and right views after super-resolution reconstruction of the stereo image left and right views.
[0190] According to one or more embodiments of this disclosure, Example 10 provides a stereoscopic image reconstruction apparatus, the application apparatus comprising:
[0191] The pre-alignment module is used to perform optical flow pre-alignment on the left and right view features of the stereo image to be processed using a stereo image reconstruction model to obtain pre-aligned features;
[0192] A fine alignment module is used to perform a fine alignment task on the pre-aligned features using the stereo image reconstruction model;
[0193] The image reconstruction module is used to reconstruct the left and right views of the stereoscopic image to be processed based on the fine alignment task to obtain the reconstructed stereoscopic image left and right views.
[0194] According to one or more embodiments of this disclosure, Example 11 provides an electronic device, the electronic device comprising:
[0195] At least one processor; and
[0196] A memory communicatively connected to the at least one processor; wherein,
[0197] The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the stereo image reconstruction method described in any one of Examples 1-9.
[0198] According to one or more embodiments of the present disclosure, Example 12 provides a computer-readable medium storing computer instructions for causing a processor to execute the stereoscopic image reconstruction method described in any one of Examples 1-9.
[0199] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0200] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0201] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A stereoscopic image reconstruction method characterized by, The application method includes: A stereo image reconstruction model is used to perform optical flow pre-alignment on the left and right view features of the stereo image to be processed to obtain pre-aligned features; wherein, the alignment features include features obtained by aligning the features from the left view features of the stereo image to the right view features of the stereo image, and features obtained by aligning the features from the right view features of the stereo image to the left view features of the stereo image; The stereo image reconstruction model is used to perform a fine alignment task on the pre-aligned features; Based on the refined alignment task, the left and right views of the stereo image to be processed are reconstructed to obtain the reconstructed stereo image left and right views. Specifically, a stereo image reconstruction model is used to perform optical flow pre-alignment on the left and right view features of the stereo image to be processed to obtain pre-aligned features, including: Based on the pre-trained optical flow estimation network in the stereo image reconstruction model, an optical flow pre-alignment task is performed on the left and right view features of the stereo image to be processed to obtain pre-aligned features; The process of determining the pre-trained optical flow estimation network includes: Determine the training samples used by the optical flow estimation network; each training sample includes two training images of the real scene and the optical flow pre-labeled values between the two training images. Based on the training samples, the optical flow estimation network to be trained in the stereo image reconstruction model is controlled to perform an optical flow estimation training task, thereby obtaining the pre-trained optical flow estimation network in the stereo image reconstruction model. Among them, determining the training samples used by the optical flow estimation network includes: Determine the optical flow information corresponding to the background image and the foreground image selected from the source image set; wherein the background image and the foreground image are non-repeating images, and the source image set is generated based on real scene images in the stereo image super-resolution dataset; A local foreground image is obtained by selecting a local region from the foreground image, and the optical flow information corresponding to the local foreground image is determined. The background image and the foreground local image are forward transformed using their respective optical flow information to obtain the transformed background image and foreground local image; The first training image is obtained by superimposing the background image and the foreground local image before conversion in the direction from background to foreground. The second training image is obtained by superimposing the background image and the foreground local image after conversion in the direction from background to foreground. The pre-labeled optical flow information between the two training images is obtained by superimposing the optical flow information corresponding to the background image and the foreground local image respectively, so as to construct the training samples used to train the optical flow estimation network.
2. The method of claim 1, wherein, Based on the training samples, the optical flow estimation network to be trained in the stereo image reconstruction model is controlled to perform an optical flow estimation training task, thereby obtaining the pre-trained optical flow estimation network in the stereo image reconstruction model, including: Two training images from the training samples are input into the optical flow estimation network to be trained in the stereo image reconstruction model to perform optical flow estimation training task, and the optical flow estimation value between the two training images in the training samples is obtained. Based on the optical flow estimation value between two training images in the training samples and the optical flow pre-labeled value between two training images, the loss function value of the optical flow estimation training task is determined; The parameters of the optical flow estimation network to be trained are adjusted based on the loss function value of the optical flow estimation training task to obtain the pre-trained optical flow estimation network.
3. The method of claim 1, wherein, Based on the pre-trained optical flow estimation network in the stereo image reconstruction model, an optical flow pre-alignment task is performed on the left and right view features of the stereo image to be processed to obtain pre-aligned features, including: The left and right views of the stereo image to be processed are input into the pre-trained optical flow estimation network in the stereo image reconstruction model to obtain the optical flow estimation value between the left and right views of the stereo image to be processed. Based on the optical flow estimation values between the left and right views of the stereo image to be processed, the features of the left and right views of the stereo image to be processed are pre-aligned to obtain the pre-aligned features between the features of the left and right views of the stereo image to be processed.
4. The method of claim 1, wherein, The stereo image reconstruction model is used to perform a fine-grained alignment task on the pre-aligned features, including: The pre-aligned features are refined by aligning the pre-aligned features using a pre-trained refined alignment network in the stereo image reconstruction model to obtain refined aligned features. Specifically, when training the pre-trained refined alignment network, the learning rate of the pre-trained optical flow estimation network is within a preset learning rate range, and the learning rate of the pre-trained optical flow estimation network is lower than the learning rate of the pre-trained refined alignment network. The refined alignment network includes a deformable convolutional network.
5. The method of claim 4, wherein, The pre-aligned features are refined by using a pre-trained fine-alignment network of the stereo image reconstruction model to obtain fine-aligned features, including: The pre-aligned features and the left and right view features of the stereo image to be processed are input into the offset estimation network of the refined alignment network for offset estimation to determine the offset of the pre-aligned features. The offset of the pre-aligned feature and the pre-aligned feature are input into the deformable convolutional network of the fine alignment network to perform fine alignment of the pre-aligned feature, thereby obtaining the fine alignment feature between the left and right view features of the stereo image to be processed.
6. The method of claim 1, wherein, Based on the refined alignment task, image reconstruction is performed on the left and right views of the stereo image to be processed to obtain the reconstructed stereo image left and right views, including: The refined alignment features obtained from the refined alignment task and the left and right view features of the stereo image to be processed are input into the fusion reconstruction network of the stereo image reconstruction model to obtain the reconstructed stereo image left and right views after super-resolution reconstruction of the stereo image left and right views.
7. A stereoscopic image reconstruction apparatus, characterized by comprising: The application device includes: The pre-alignment module is used to perform optical flow pre-alignment on the left and right view features of the stereo image to be processed using a stereo image reconstruction model to obtain pre-aligned features; wherein, the alignment features include features obtained by aligning the left view features of the stereo image to the right view features of the stereo image and features obtained by aligning the right view features of the stereo image to the left view features of the stereo image. A fine alignment module is used to perform a fine alignment task on the pre-aligned features using the stereo image reconstruction model; The image reconstruction module is used to reconstruct the left and right views of the stereo image to be processed according to the fine alignment task, so as to obtain the reconstructed stereo image left and right views. The pre-alignment module includes: The pre-aligned feature acquisition unit is used to perform optical flow pre-alignment task on the left and right view features of the stereo image to be processed based on the pre-trained optical flow estimation network in the stereo image reconstruction model to obtain pre-aligned features. The pre-aligned features yield units, including: The training sample determination subunit is used to determine the training samples used by the optical flow estimation network; each training sample includes two training images of the real scene and optical flow pre-labeled values between the two training images. The pre-trained optical flow estimation network obtains sub-units, which are used to control the optical flow estimation network to be trained in the stereo image reconstruction model to perform optical flow estimation training tasks based on the training samples, thereby obtaining the pre-trained optical flow estimation network in the stereo image reconstruction model. Among them, the training samples determine the sub-units, specifically used for: Determine the optical flow information corresponding to the background image and the foreground image selected from the source image set; wherein the background image and the foreground image are non-repeating images, and the source image set is generated based on real scene images in the stereo image super-resolution dataset; A local foreground image is obtained by selecting a local region from the foreground image, and the optical flow information corresponding to the local foreground image is determined. The background image and the foreground local image are forward transformed using their respective optical flow information to obtain the transformed background image and foreground local image; The first training image is obtained by superimposing the background image and the foreground local image before conversion in the direction from background to foreground. The second training image is obtained by superimposing the background image and the foreground local image after conversion in the direction from background to foreground. The pre-labeled optical flow information between the two training images is obtained by superimposing the optical flow information corresponding to the background image and the foreground local image respectively, so as to construct the training samples used to train the optical flow estimation network.
8. An electronic device, comprising: The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the stereoscopic image reconstruction method according to any one of claims 1-6.
9. A computer readable medium characterized by The computer-readable medium stores computer instructions that cause a processor to execute the stereoscopic image reconstruction method according to any one of claims 1-6.