Monocular Three-Dimensional Plane Recovery Method, Device and Storage Medium
Through multi-scale feature extraction and the encoder-decoder architecture of the Transformer module, the problem of insufficient recognition of small plane areas in monocular three-dimensional plane recovery is solved, and the recovery accuracy and robustness are improved.
Patent Information
- Application Number
- CN202210739676.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-06-28
AI Technical Summary
In the prior art, the monocular three-dimensional plane recovery method lacks the ability to identify small plane areas, resulting in loss of image information and affecting recovery accuracy and robustness.
The encoder-decoder architecture of multi-scale feature extraction and Transformer modules is adopted to extract the internal and associated features of the image blocks respectively, and the decoder weight is updated through the loss function to improve the comprehensiveness and robustness of feature extraction.
It effectively reduces image information loss, improves the accuracy and robustness of monocular three-dimensional plane recovery, and enhances the detection accuracy of scene changes.
Smart Images

Figure CN115115691B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image data processing, and particularly to a monocular three-dimensional plane recovery method, device, and storage medium. Background Art
[0002] Three-dimensional plane recovery requires segmenting the plane region of the scene from the image dimension, and at the same time estimating the plane parameters of the corresponding region. Based on the plane region and plane parameters, three-dimensional plane recovery can be realized, and the predicted three-dimensional plane can be reconstructed.
[0003] In the related art, monocular three-dimensional plane recovery focuses on reconstruction accuracy, and strengthens the accuracy of the model structure by analyzing the edges of the plane structure and the embedding of the scene, but lacks the ability to identify small plane regions, and is prone to losing pixel regions with a small proportion during the plane detection process, affecting the accuracy of monocular three-dimensional plane recovery. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention provides a monocular three-dimensional plane recovery method, device, and storage medium, which can extract the internal features of the feature map, effectively improve the comprehensiveness of feature extraction, and further improve the accuracy of monocular three-dimensional plane recovery.
[0005] The first aspect embodiment of the present invention provides a monocular three-dimensional plane recovery method, including:
[0006] Performing multi-scale feature extraction on the input image to obtain a first feature map and a second feature map at two scales;
[0007] Inputting the first feature map into a first inner encoder and a first outer encoder respectively, and extracting the first internal feature of the first image block in the first feature map and the first correlation feature between each first image block;
[0008] Inputting the second feature map into a second inner encoder and a second outer encoder respectively, and extracting the second internal feature of the second image block in the second feature map and the second correlation feature between each second image block;
[0009] Fusing the first internal feature and the first correlation feature and inputting them into a first decoder for decoding to obtain predicted plane parameters and a predicted plane region;
[0010] Fusing the second internal feature and the second correlation feature and inputting them into a second decoder for decoding to obtain a predicted non-planar region, where the predicted non-planar region is used to verify the predicted plane region;
[0011] Performing three-dimensional recovery according to the plane parameters and the plane region to obtain a predicted three-dimensional plane.
[0012] According to the above embodiments of the present invention, there are at least the following beneficial effects: By setting an inner encoder and an outer encoder, the internal features of the image blocks in the corresponding feature maps and the correlation features between the image blocks are respectively extracted, and then the internal features and the correlation features are fused and input into the decoder for decoding, which can effectively improve the comprehensiveness of feature extraction, reduce the probability of image information loss, and further improve the accuracy of monocular three-dimensional plane recovery. In addition, the predicted plane region can be verified by predicting the non-plane region, which can further improve the robustness of monocular three-dimensional plane recovery.
[0013] According to some embodiments of the first aspect of the present invention, multi-scale feature extraction is performed on the input image to obtain a first feature map and a second feature map at two scales, including:
[0014] Perform multi-scale feature extraction on the input image to obtain a first extraction map and a second extraction map at two scales;
[0015] Embed the corresponding position information into the first extraction map and the second extraction map respectively to obtain a first feature map and a second feature map at two scales.
[0016] According to some embodiments of the first aspect of the present invention, the first feature map is respectively input into a first inner encoder and a first outer encoder to respectively extract a first internal feature of a first image block in the first feature map and a first correlation feature between each first image block, including:
[0017] Slice the first feature map into a plurality of first image blocks;
[0018] Input each first image block into the first inner encoder to extract the first internal feature of each first image block;
[0019] Input each first image block into the first outer encoder to extract the first correlation feature between each first image block.
[0020] According to some embodiments of the first aspect of the present invention, the second feature map is respectively input into a second inner encoder and a second outer encoder to respectively extract a second internal feature of a second image block in the second feature map and a second correlation feature between each second image block, including:
[0021] Slice the second feature map into a plurality of second image blocks;
[0022] Input each second image block into the second inner encoder to extract the second internal feature of each second image block;
[0023] Input each second image block into the second outer encoder to extract the second correlation feature between each second image block.
[0024] According to some embodiments of the first aspect of the present invention, after fusing the first internal feature and the first associated feature, the result is input into a first decoder for decoding to obtain predicted plane parameters and a predicted plane region, including:
[0025] Element-wise add the first internal feature and the first associated feature to obtain a first fused feature;
[0026] Input the first fused feature into the first decoder for decoding and classification with the plane region and plane parameters as labels to obtain predicted plane parameters and a predicted plane region.
[0027] According to some embodiments of the first aspect of the present invention, after fusing the second internal feature and the second associated feature, the result is input into a second decoder for decoding to obtain a predicted non-planar region, including:
[0028] Element-wise add the second internal feature and the second associated feature to obtain a second fused feature;
[0029] Input the second fused feature into the second decoder for decoding and classification with the non-planar region as a label to obtain a predicted non-planar region.
[0030] According to some embodiments of the first aspect of the present invention, after fusing the second internal feature and the second associated feature and inputting the result into the second decoder for decoding to obtain a predicted non-planar region, it further includes:
[0031] Update the weights of the first decoder according to the predicted plane region, the predicted non-planar region, and the loss function.
[0032] According to some embodiments of the first aspect of the present invention, updating the weights of the first decoder according to the predicted non-planar region and the loss function includes:
[0033] Update the weights of the first decoder according to the predicted plane region, the predicted non-planar region, and the cross-entropy loss function, where the cross-entropy loss function is:
[0034]
[0035] Y + and Y - respectively represent the labeled pixels of the plane region and the non-planar region, and P i represents the probability that the i-th pixel belongs to the plane region, represents the probability that the i-th pixel belongs to the non-planar region, and w is the ratio of the labeled pixels of the plane region to the labeled pixels of the non-planar region.
[0036] Embodiments of the second aspect of the present invention provide an electronic device, including:
[0037] A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, a monocular three-dimensional plane recovery method according to any one of the first aspect is implemented.
[0038] Since the electronic device in the embodiment of the second aspect applies the monocular three-dimensional plane recovery method according to any one of the first aspect, it has all the beneficial effects of the first aspect of the present invention.
[0039] A computer storage medium according to an embodiment of the third aspect of the present invention stores computer-executable instructions for executing the monocular three-dimensional plane recovery method according to any one of the first aspect.
[0040] Since the computer storage medium in the embodiment of the third aspect can execute the monocular three-dimensional plane recovery method according to any one of the first aspect, it has all the beneficial effects of the first aspect of the present invention.
[0041] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings
[0042] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:
[0043] Figure 1 is a main step diagram of the monocular three-dimensional plane recovery method according to an embodiment of the present invention;
[0044] Figure 2 is a schematic diagram of the steps of S100 in the monocular three-dimensional plane recovery method according to an embodiment of the present invention;
[0045] Figure 3 is a schematic diagram of the steps of S200 in the monocular three-dimensional plane recovery method according to an embodiment of the present invention;
[0046] Figure 4 is a schematic diagram of the steps of S300 in the monocular three-dimensional plane recovery method according to an embodiment of the present invention;
[0047] Figure 5 is a schematic diagram of the steps of S400 in the monocular three-dimensional plane recovery method according to an embodiment of the present invention;
[0048] Figure 6 is a schematic diagram of the steps of S500 in the monocular three-dimensional plane recovery method according to an embodiment of the present invention;
[0049] Figure 7 is a framework diagram of the network applied in the monocular three-dimensional plane recovery method according to an embodiment of the present invention. Detailed Embodiments
[0050] In the description of the present invention, unless otherwise clearly defined, terms such as "setting", "installing", "connecting", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in the present invention in combination with the specific content of the technical solution. In the description of the present invention, the meaning of "several" is one or more, the meaning of "multiple" is more than two, "greater than", "less than", "exceeding", etc. are understood as not including the recited number, and "above", "below", "within", etc. are understood as including the recited number. In addition, features defined with "first", "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise stated, the meaning of "multiple" is two or more.
[0051] With the development of deep learning, the field of computer vision has received more and more attention from researchers. The three-dimensional plane recovery and reconstruction technology is one of the mainstream research tasks in the current field of computer vision. The three-dimensional plane recovery of a single image requires segmenting the plane instance region of the scene from the image dimension and simultaneously estimating the plane parameters of each instance region. Generally, non-planar regions will be represented by the depth estimated by the network model. This technology has broad application prospects in the fields of virtual reality, augmented reality, robotics, etc.
[0052] The plane detection and recovery method for a single image needs to simultaneously conduct research on image depth, plane normal, plane segmentation, etc. The traditional three-dimensional plane recovery and reconstruction method based on manually extracting features only extracts the shallow texture information of the image and depends on the prior conditions of plane geometry, having the drawback of weak generalization ability. And the real indoor scene is very complex, and the multiple shadows generated by complex light and various folded occluders will affect the quality of plane recovery and reconstruction, making it difficult for traditional methods to handle the plane reconstruction task of complex indoor scenes. Plane recovery and reconstruction is an important research direction in three-dimensional reconstruction. Currently, most three-dimensional reconstruction methods first generate point cloud data through three-dimensional vision methods, then generate a non-linear scene surface by fitting relevant points, and then optimize the overall reconstruction model through global reasoning. The segmented plane recovery and reconstruction combines the visual instance segmentation method to identify the plane region of the scene, representing the plane with three parameters and a segmentation mask in the Cartesian coordinate system, having better reconstruction accuracy and effect. The segmented plane recovery and reconstruction is a multi-stage reconstruction method, and the accuracy of plane recognition and parameter estimation will both affect the result of the final model.
[0053] The three-dimensional plane recovery requires segmenting the plane region of the scene from the image dimension and simultaneously estimating the plane parameters of the corresponding region. Based on the plane region and plane parameters, the three-dimensional plane recovery can be realized, and the predicted three-dimensional plane can be reconstructed.
[0054] There are the following several plane recovery methods: The end-to-end convolutional neural network architecture Planenet can infer a fixed number of plane instance masks and plane parameters from a single RGB image; learn directly from the depth modality of the loss induced by the plane structure by predicting a fixed number of planes; improve the two-stage Mask R-CNN method, replace object category classification with plane geometry prediction, and then refine the plane segmentation mask with a convolutional neural network; adopt the associated embedding method by predicting per-pixel plane parameters, train the network parameters to map each pixel to the embedding space, and then cluster the embedded pixels to generate plane instances; a plane refinement method constrained by the Manhattan world hypothesis, which strengthens the refinement of plane parameters by restricting the geometric relationship between plane instances; a divide-and-conquer method for panoramic plane segmentation in the horizontal and vertical directions, which can effectively recover distorted plane instances for the pixel distribution difference between panoramic images and ordinary images; the method PlaneTR based on the Transformer module, which effectively improves the efficiency of plane detection by adding the center and edge features of plane instances.
[0055] In related technologies, monocular three-dimensional plane recovery focuses on reconstruction accuracy, strengthens the accuracy of the model structure by analyzing the edges of the plane structure and the embedding of the scene, but lacks the ability to recognize small plane regions, and is prone to losing small pixel regions during the plane detection process, affecting the accuracy of monocular three-dimensional plane recovery.
[0056] Based on this, in order to obtain better results and use fewer computing resources, applying the encoder part of the Transformer module to the image patch sequence and applying it to the image classification task can obtain better results and fewer computing resources than the state-of-the-art convolutional network. If the object detection problem is described as a sequence-to-sequence prediction problem, directly predict a set of objects that interact with the context feature sequence from the learned object queries. A new and simple object detection paradigm is proposed, which is based on the standard Transformer encoder-decoder architecture, which gets rid of many manually designed components such as anchor generation and non-maximum suppression. To solve the suboptimal representation learning caused by the lack of learning ability of the convolutional network for low-level feature tensors, semantic segmentation is redefined as a sequence-to-sequence prediction task, and a pure encoder based on the self-attention mechanism is proposed, eliminating the dependence on convolutional operations and solving the problem of limited receptive fields.
[0057] The following refers to Figures 1 to 7 Describe a monocular three-dimensional plane recovery method, device and storage medium of the present invention, which can extract the internal features of the feature map, effectively improve the comprehensiveness of feature extraction, and thus improve the accuracy of monocular three-dimensional plane recovery.
[0058] Reference Figure 1 As shown, a monocular three-dimensional plane recovery method according to an embodiment of the first aspect of the present invention includes at least the following steps:
[0059] S100. Perform multi-scale feature extraction on the input image to obtain a first feature map and a second feature map at two scales;
[0060] S200. Input the first feature map into a first inner encoder and a first outer encoder respectively, and extract the first internal feature of the first image block in the first feature map and the first correlation feature between each first image block;
[0061] S300. Input the second feature map into a second inner encoder and a second outer encoder respectively, and extract the second internal feature of the second image block in the second feature map and the second correlation feature between each second image block;
[0062] S400. After fusing the first internal feature and the first correlation feature, input them into a first decoder for decoding to obtain predicted plane parameters and a predicted plane region;
[0063] S500. After fusing the second internal feature and the second correlation feature, input them into a second decoder for decoding to obtain a predicted non-planar region, where the predicted non-planar region is used to verify the predicted plane region;
[0064] S600. Perform three-dimensional recovery according to the plane parameters and the plane region to obtain a predicted three-dimensional plane.
[0065] By performing multi-scale feature extraction on the input image, the comprehensiveness of the obtained information can be improved. By setting an inner encoder and an outer encoder to extract the internal features of the image blocks in the corresponding feature maps and the correlation features between the image blocks respectively, and then fusing the internal features and the correlation features and inputting them into the decoder for decoding, the comprehensiveness of feature extraction can be effectively improved, the probability of image information loss can be reduced, and thus the accuracy of monocular three-dimensional plane recovery can be improved. In addition, the predicted plane region can be verified by the predicted non-planar region, which can further improve the robustness of monocular three-dimensional plane recovery.
[0066] It can be understood that, as shown in the reference Figure 2 Step S100, performing multi-scale feature extraction on the input image to obtain a first feature map and a second feature map at two scales, includes:
[0067] S110. Perform multi-scale feature extraction on the input image to obtain a first extraction map and a second extraction map at two scales;
[0068] S120. Embed the corresponding position information into the first extraction map and the second extraction map respectively to obtain the first feature map and the second feature map at two scales.
[0069] Specifically, in step S110, perform multi-scale feature extraction on the input image through the HRNet convolutional network to obtain the first extraction map and the second extraction map at two scales.
[0070] In step S120, embed the corresponding position information into the first extraction map and the second extraction map respectively through position embedding, and after converting them into tokens respectively, obtain the first feature map and the second feature map at two scales.
[0071] It should be noted that the scale corresponding to the first feature map is the scale of HW / 16, and the scale corresponding to the second feature map is HW / 32, where H and W respectively represent the height and width of the input image.
[0072] To obtain more details, the input data is further encoded into sub-patches (i.e., subdivided image blocks) through the attention mechanism. By dividing the feature map into multiple non-overlapping regions and performing Windows Multi-Head Self-Attention (W-MSA) on the patch embeddings of different feature maps, the computational complexity can be effectively reduced. The tokens from different stages of the vision transformer are combined into image-like representations of different resolutions, and a convolutional decoder is used to gradually combine them into full-resolution predictions. Compared with the fully convolutional network, the multi-scale dense vision transformer avoids the feature loss caused by the downsampling operation after the image block embedding calculation and provides more refined and globally consistent predictions.
[0073] It can be understood that, as shown in Figure 3 , in step S200, input the first feature map into the first inner encoder and the first outer encoder respectively, and extract the first internal feature of the first image block in the first feature map and the first correlation feature between each first image block respectively, including:
[0074] S210. Cut the first feature map into multiple first image blocks.
[0075] S220. Input each first image block into the first inner encoder to extract the first internal feature of each first image block.
[0076] S230. Input each first image block into the first outer encoder to extract the first correlation feature between each first image block, where the first correlation feature is used to characterize the relationship between each image block.
[0077] By slicing the first feature map into multiple first image patches, it is possible to effectively avoid the loss of small - proportion pixel plane regions.
[0078] It can be understood that, referring to Figure 4 as shown, in step S300, the second feature map is input into the second inner encoder and the second outer encoder respectively, and the second internal features of the second image patches in the second feature map and the second correlation features between the respective second image patches are extracted respectively, including:
[0079] S310: Slice the second feature map into multiple second image patches;
[0080] S320: Input each second image patch into the second inner encoder to extract the second internal features of each second image patch;
[0081] S330: Input each second image patch into the second outer encoder to extract the second correlation features between the respective second image patches, where the second correlation features are used to characterize the relationship between the image patches.
[0082] By slicing the first feature map into multiple first image patches, it is possible to effectively avoid the loss of small - proportion pixel plane regions.
[0083] It can be understood that, referring to Figure 5 as shown, in step S400, after fusing the first internal features and the first correlation features, input them into the first decoder for decoding to obtain the predicted plane parameters and the predicted plane region, including:
[0084] S410: Perform element - wise addition on the first internal features and the first correlation features to obtain the first fusion feature;
[0085] S420: Input the first fusion feature into the first decoder for decoding and classification with the plane region and plane parameters as labels to obtain the predicted plane parameters and the predicted plane region.
[0086] By fusing the first internal features and the first correlation features, it is possible to effectively improve the comprehensiveness of feature extraction, and thus improve the accuracy of the final 3D plane recovery.
[0087] It can be understood that, referring to Figure 6 as shown, in step S500, after fusing the second internal features and the second correlation features, input them into the second decoder for decoding to obtain the predicted non - plane region, including:
[0088] S510: Perform element - wise addition on the second internal features and the second correlation features to obtain the second fusion feature;
[0089] S520: Input the second fusion feature into the second decoder for decoding and classification with the non - plane region as the label to obtain the predicted non - plane region.
[0090] By fusing the second internal feature and the second related feature, the comprehensiveness of feature extraction can be effectively improved, and thus the accuracy of the final 3D plane recovery can be improved.
[0091] When the scene changes, the detection accuracy during 3D plane recovery in related technologies is significantly insufficient and the robustness is low.
[0092] Based on this, to improve the detection accuracy and robustness to scene changes, it can be understood that in step S500, after fusing the second internal feature and the second associated feature and inputting them into the second decoder for decoding to obtain the predicted non-planar region, it further includes:
[0093] Updating the weights of the first decoder according to the predicted planar region, the predicted non-planar region, and the loss function.
[0094] During the 3D plane recovery process, by iteratively updating the first decoder through the loss function, the accuracy of predicting the planar region during 3D plane recovery can be effectively improved, the performance of the overall network is dynamically updated, and the detection accuracy and robustness to scene changes can be improved.
[0095] It can be understood that updating the weights of the first decoder according to the predicted planar region, the predicted non-planar region, and the loss function is specifically:
[0096] Updating the weights of the first decoder according to the predicted planar region, the predicted non-planar region, and the cross-entropy loss function, where the cross-entropy loss function is:
[0097]
[0098] Y + and Y - respectively represent the labeled pixels of the planar region and the non-planar region, and P i represents the probability that the i-th pixel belongs to the planar region, represents the probability that the i-th pixel belongs to the non-planar region, and w is the ratio of the labeled pixels of the planar region to the labeled pixels of the non-planar region. Due to the differences in the scales of the first decoder and the second decoder and the definitions of positive and negative labels, the planar region is finally optimized through variational information, and the positive and negative labels are the planar region label and the non-planar region label.
[0099] Mutual information is a measure of the degree of dependence between two random variables based on Shannon entropy, which can capture the non-linear statistical correlation between variables. The mutual information between X and Z can be understood as the reduction in the uncertainty in X given Z:
[0100]
[0101] Among them, \(H(X)\) is the Shannon entropy, \(H(X|Z)\) is the conditional entropy of \(Z\) given \(X\), \(P XZ is the joint probability distribution of two variables, \(P X and \(P Z are their respective marginal probability distributions. At the same time, the mutual information is equivalent to the KL divergence (Kullback-Leibler) of \(P XZ and \(P X and \(P Z product:
[0102]
[0103] When the divergence between the joint probability \(P XZ and the marginal product is greater, the dependence between \(X\) and \(Z\) is stronger. Therefore, for two completely independent variables, there is no mutual information. Mutual information is usually used in unsupervised representation learning networks, but mutual information estimation is difficult to estimate as a bijective function and may lead to suboptimal representations that are irrelevant to downstream tasks. And a highly non-linear evaluation framework may bring better downstream performance but violates the purpose of learning effective and transferable data representations. Based on the mutual information-based knowledge distillation framework, the mutual information is defined as the difference between the entropy value of the teacher model and the entropy value of the teacher model under the condition of the student model. By maximizing the mutual information between the teacher-student networks, the student model learns the feature distribution of the teacher model.
[0104] Based on this, the present invention enhances the feature expression by maximizing the planar feature mutual information of two network branches. In the PlaneMT network model framework, two network branches of different scales respectively correspond to the first decoder and the second decoder, which are respectively used to detect and obtain the predicted planar region \(S P and the predicted non-planar region \(S'\) N-P , where, in the most ideal case, the predicted planar region and the predicted non-planar region are inverted:
[0105] S' P ∶=S' N-P
[0106] Thus, the output predicted planar region variables \(S P and \(S'\) P of the two network branches are used as variational information measures for information maximization:
[0107]
[0108] Since the mutual information is difficult to calculate, a variational lower bound is proposed for each mutual information term \(I(X;Z)\), and a variable Gaussian \(q(x|z)\) is used to simulate \(p(x|z)\):
[0109] I(SP ; S' P ) = H(S P ) - H(S P |S' P )
[0110]
[0111] The last inequality represents the non - negativity of the KL divergence D KL .
[0112] The framework diagram of the network applied by the monocular three - dimensional plane recovery method according to the embodiment of the present invention is shown in Figure 7 . After the backbone network extracts features to obtain feature maps with sizes of 12×16 and 6×8, the feature map with a size of 12×16 is input into the first inner and outer encoders through POS (Position Embedding), and the feature map with a size of 6×8 is input into the second inner and outer encoders through POS. Among them, the loss function adopts the mutual information loss function.
[0113] In addition, an embodiment of the second aspect of the present invention further provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor.
[0114] The processor and the memory can be connected through a bus or other means.
[0115] As a non - transient computer - readable storage medium, the memory can be used to store non - transient software programs and non - transient computer - executable programs. In addition, the memory may include high - speed random - access memory, and may also include non - transient memory, such as at least one magnetic disk storage device, a flash memory device, or other non - transient solid - state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above - mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0116] The non - transient software programs and instructions required to implement the monocular three - dimensional plane recovery method of the above - mentioned first - aspect embodiment are stored in the memory. When executed by the processor, the monocular three - dimensional plane recovery method in the above - mentioned embodiment is executed. For example, the method steps S100 to S600, the method steps S110 to S120, the method steps S210 to S230, the method steps S310 to S330, the method steps S410 to S420, and the method steps S510 to S520 described above are executed.
[0117] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0118] In addition, the third aspect embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions, which are executed by a processor or a controller, for example, executed by a processor in the above device embodiment, enabling the processor to execute the monocular three-dimensional plane recovery method in the above embodiment, for example, execute the method steps S100 to S600, method steps S110 to S120, method steps S210 to S230, method steps S310 to S330, method steps S410 to S420, method steps S510 to S520 described above.
[0119] Those of ordinary skill in the art can understand that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium generally contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0120] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0121] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A monocular three-dimensional plane recovery method, characterized in that, Including: Performing multi-scale feature extraction on the input image to obtain a first feature map and a second feature map at two scales; Inputting the first feature map into a first inner encoder and a first outer encoder respectively, and extracting a first internal feature of a first image block in the first feature map and a first correlation feature between each of the first image blocks; Inputting the second feature map into a second inner encoder and a second outer encoder respectively, and extracting a second internal feature of a second image block in the second feature map and a second correlation feature between each of the second image blocks; Fusing the first internal feature and the first correlation feature and inputting the result into a first decoder for decoding to obtain predicted plane parameters and a predicted plane region; Fusing the second internal feature and the second correlation feature and inputting the result into a second decoder for decoding to obtain a predicted non-planar region, where the predicted non-planar region is used to verify the predicted plane region; Performing three-dimensional restoration based on the plane parameters and the plane region to obtain a predicted three-dimensional plane.
2. The monocular three-dimensional plane recovery method according to claim 1, characterized in that The performing multi-scale feature extraction on the input image to obtain a first feature map and a second feature map at two scales includes: Performing multi-scale feature extraction on the input image to obtain a first extraction map and a second extraction map at two scales; Embedding corresponding position information into the first extraction map and the second extraction map respectively to obtain a first feature map and a second feature map at two scales.
3. The monocular three-dimensional plane recovery method according to claim 1, wherein The inputting the first feature map into a first inner encoder and a first outer encoder respectively, and extracting a first internal feature of a first image block in the first feature map and a first correlation feature between each of the first image blocks includes: Cutting the first feature map into a plurality of first image blocks; Inputting each of the first image blocks into the first inner encoder to extract the first internal feature of each of the first image blocks; Inputting each of the first image blocks into the first outer encoder to extract the first correlation feature between each of the first image blocks.
4. The monocular three-dimensional plane recovery method according to claim 1, characterized in that The inputting the second feature map into a second inner encoder and a second outer encoder respectively, and extracting a second internal feature of a second image block in the second feature map and a second correlation feature between each of the second image blocks includes: Cutting the second feature map into a plurality of second image blocks; Inputting each of the second image blocks into the second inner encoder to extract the second internal feature of each of the second image blocks; Inputting each of the second image blocks into the second outer encoder to extract the second correlation feature between each of the second image blocks.
5. The monocular three-dimensional plane recovery method according to claim 1, characterized in that The fusing the first internal feature and the first correlation feature and inputting the result into a first decoder for decoding to obtain predicted plane parameters and a predicted plane region includes: Performing element-wise addition on the first internal feature and the first correlation feature to obtain a first fusion feature; Inputting the first fusion feature into the first decoder for decoding and classification with the plane region and the plane parameters as labels to obtain the predicted plane parameters and the predicted plane region.
6. The monocular three-dimensional plane recovery method according to claim 1, characterized in that Fusing the second internal feature and the second associated feature and then inputting the fused result into a second decoder for decoding to obtain a predicted non-planar region, including: Performing element-wise addition on the second internal feature and the second associated feature to obtain a second fused feature; Inputting the second fused feature into the second decoder for decoding and classification with the non-planar region as the label to obtain the predicted non-planar region.
7. The monocular three-dimensional plane recovery method according to claim 1, characterized in that After fusing the second internal feature and the second associated feature and then inputting the fused result into a second decoder for decoding to obtain a predicted non-planar region, further including: Updating the weights of the first decoder according to the predicted planar region, the predicted non-planar region, and a loss function.
8. The monocular three-dimensional plane recovery method according to claim 7, characterized in that The updating the weights of the first decoder according to the predicted non-planar region and the loss function includes: Updating the weights of the first decoder according to the predicted planar region, the predicted non-planar region, and a cross-entropy loss function, where the cross-entropy loss function is: Y + and Y - represent the marked pixels of the planar region and the non-planar region respectively, and P i represents the probability that the i-th pixel belongs to the planar region, represents the probability that the i-th pixel belongs to the non-planar region, and w is the ratio of the marked pixels of the planar region to the marked pixels of the non-planar region.
9. An electronic device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the monocular three-dimensional plane recovery method according to any one of claims 1 to 8 is implemented.
10. A computer storage medium, characterized in that, Stored with computer-executable instructions for executing the monocular three-dimensional plane recovery method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Scene reconstruction in three-dimensions from two-dimensional images
CN114026599A
Image depth estimation method and device, electronic equipment and storage medium
CN114266814A