Panoramic room layout estimation method and system and computer device
Through the panoramic room layout estimation method with a dual-branch architecture, the feature extraction and fusion of the ERP branch and the Cube branch are used to solve the panoramic distortion problem and improve the accuracy of room layout estimation, especially in complex room structures and non-Manhattan scenarios.
Patent Information
- Application Number
- CN202510818413.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-10-10
AI Technical Summary
Existing panoramic room layout estimation methods fail to fully utilize the complementary advantages of different projection methods, resulting in insufficient performance when dealing with complex room structures and non-Manhattan scenes.
A dual-branch architecture, ERP branch and Cube branch, is adopted. The ERP branch processes panoramic images through saliency maps and deformable convolutional feature maps, while the Cube branch extracts local features by multi-view stitching of cube maps and fuses the feature sequences of the two to predict room layout parameters.
The accuracy of room layout estimation is improved, the panoramic distortion problem is solved, and the ability to handle complex room structures and non-Manhattan scenes is enhanced.
Smart Images

Figure CN120765873A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision room layout estimation, and in particular to a panoramic room layout estimation method, system and computer device. Background Art
[0002] Room layout estimation is a key task in computer vision and 3D scene understanding. It aims to recover the structured geometric information of indoor scenes from a single or multiple images, including the location and reconstruction of major planar boundaries such as walls, ceilings, and floors. This technology has broad application value in augmented reality (AR), virtual reality (VR), robot navigation, smart homes, and other fields. With the rapid development of virtual reality technology and the growing demand for immersive indoor experiences, room layout estimation technology using panoramic images has received increasing attention. 360° panoramic cameras can capture complete indoor environment information, providing richer spatial cues for layout estimation, but also bringing new technical challenges.
[0003] Compared to traditional planar images, panoramic images represent a full 360° × 180° field of view through equirectangular or cubic projections. A typical indoor environment typically consists of a floor, a ceiling, and several vertical walls. This Manhattan world assumption provides useful prior knowledge for layout estimation. Current methods often combine geometric reasoning and deep learning to predict room layout parameters through end-to-end training, but their performance in handling complex room structures and non-Manhattan scenes remains to be improved. Existing panoramic room layout estimation methods are mostly based on a single projection design and propose improvements to address their respective distortion issues. However, these methods typically rely solely on a single projection and fail to fully utilize the complementary advantages of different projection methods. Therefore, exploring panoramic room layout estimation methods that combine dual projections is crucial. Summary of the Invention
[0004] In view of this, the present invention proposes a panoramic room layout estimation method, system and computer device, aiming to solve the problems existing in the current technology.
[0005] The present invention proposes a panoramic room layout estimation method, comprising:
[0006] S1. Extracting global features from a panoramic image of a room through an ERP branch and processing the extracted features to obtain a global one-dimensional feature sequence of the panoramic image.
[0007] S2. Convert the panoramic image in step S1 into a cube map, extract local features from the cube map based on the Cube branch, and process the extracted features to obtain a local one-dimensional feature sequence of the cube map.
[0008] S3. Fusing the global one-dimensional feature sequence in step S1 and the local one-dimensional feature sequence in step S2 to obtain a predicted horizontal line depth and a predicted room height.
[0009] In some embodiments of the present application, step S1 specifically includes:
[0010] S11, extracting global features from the panoramic image through the ERP branch and outputting an original feature map;
[0011] S12, after processing the panoramic image through the deformable convolution DCNv3 to output a deformable convolution feature map, the slice similarity, normalization layer and standard deviation layer processing are performed in sequence to form a saliency map;
[0012] S13, adding the saliency map to the original feature map and performing high compression through a high compression module to obtain the global one-dimensional feature sequence.
[0013] In some embodiments of the present application, the ResNet-34 model is used as the ERP branch in step S11.
[0014] In some embodiments of the present application, step S2 specifically includes:
[0015] S21, converting the panoramic image into a cube map;
[0016] S22, splitting the six faces of the cube map in step S21, horizontally splicing the top, bottom, left, and right views to obtain a horizontal image, and vertically splicing the top view and the bottom view to obtain a vertical image;
[0017] S23, extracting local features from the horizontal image in step S22 through the Cube branch to obtain local features Figure 1 , the local features Figure 1 A local one-dimensional feature sequence 1 is obtained by performing high compression through a high compression module;
[0018] S24, extract local features from the two views of the vertical image in step S22 through the Cube branch to obtain local features Figure 2 And local feature diagram 3, the local feature Figure 2 The height of local feature map 3 and local feature map 3 are compressed by the height compression module to obtain local one-dimensional feature sequence 2 and local one-dimensional feature sequence 3;
[0019] S25, linearize the local one-dimensional feature sequence 2 and the local one-dimensional feature sequence 3 through the linear layer to obtain the cube height h Cube .
[0020] In some embodiments of the present application, S21, when converting the panoramic picture into a cube map, specifically comprises the following steps:
[0021] Given that a pixel p is located on the i-th face of the cube, the coordinates of p in the camera system are (x, y, z), that is, the pixel p=(p x , p y , p z ), the pixel p is equirectangularly projected, and the specific formula is as follows:
[0022]
[0023] where w represents the pixel width of the cube map; K represents the camera intrinsic parameter of the perspective projection; R i is the rotation matrix of the i-th face of any camera coordinate system external matrix; q represents the direction vector of the pixel converted into the global coordinate system; θ represents the longitude in the equirectangular projection; φ represents the latitude in the equirectangular projection; q x , q z , q y are the components of q in x, y, and z respectively.
[0024] In some embodiments of the present application, the ResNet-34 model is used as the Cube branch of steps S23 and S24.
[0025] In some embodiments of the present application, step S3 specifically comprises:
[0026] The global one-dimensional feature sequence in step S1 is processed through a self-attention mechanism, and the specific formula is as follows:
[0027]
[0028] where Q is a query matrix; K is a key matrix; V is a value matrix; K T is the transpose matrix of the key matrix K; d k represents the dimension of the feature vector;
[0029] Q, K, and V are learned from the equirectangular image features through a full connection layer, and Q, K, and V are input into the graph feature interaction mechanism for processing, and the specific formula is as follows:
[0030] G'=(FxF T )×G;
[0031] F g =F+G'×F;
[0032] where F is the output of the self-attention mechanism; F Trepresents the transposed matrix of F; G is a learnable weight matrix and is initialized to zero; Fg is the final output of the graph feature interaction mechanism;
[0033] Take Fg as the query, local one-dimensional feature sequence 1, local one-dimensional feature sequence 1 and local one-dimensional feature sequence 3 as the key and value, pass them into the cross attention mechanism for processing, and then linearize them through the linear layer to obtain the predicted horizontal depth and predicted room height.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] The present invention extracts features from panoramic images and cube images by designing a dual-branch architecture, one of which is the ERP branch and the other is the Cube branch. In the ERP branch, a saliency map is introduced to address the problem of panoramic distortion. Given a deformable convolution feature map, the cosine similarity between the center pixel and all pixels in the kernel is calculated, and then normalized to indicate the importance of the center pixel. In the Cube branch, the four views of the cube map are spliced together to obtain a feature sequence, and the top and bottom views are used to obtain the height of the cube room. The feature sequences of the two branches are combined, and the combined features are used to predict the layout boundary parameters through a linear layer to obtain the predicted depth of the horizontal line and the predicted height of the room, thereby improving the accuracy of the room layout estimation.
[0036] On the other hand, the present application provides a panoramic room layout estimation system, which applies the panoramic room layout estimation method, including:
[0037] An ERP branch module is used to extract global features from the panoramic picture of the room through the ERP branch and then process the extracted features to obtain a global one-dimensional feature sequence of the panoramic picture;
[0038] A Cube branch module is used to convert the panoramic image into a cube map, extract local features from the cube map based on the Cube branch, and process the extracted local features to obtain a local one-dimensional feature sequence of the cube map;
[0039] The fusion module is used to fuse the global one-dimensional feature sequence and the local one-dimensional feature sequence to obtain the horizontal line prediction depth and the room prediction height.
[0040] In another aspect, the present application further provides a computer device, comprising:
[0041] at least one processor;
[0042] at least one memory for storing at least one program;
[0043] When the at least one program is executed by the at least one processor, the at least one processor is enabled to implement the panoramic room layout estimation method.
[0044] It can be understood that the panoramic room layout estimation system and the computer device in this embodiment have the same beneficial effects as the panoramic room layout estimation method, and will not be described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0046] Figure 1 A flowchart of a panoramic room layout estimation method provided by an embodiment of the present invention;
[0047] Figure 2 This is a structural block diagram of a panoramic room layout estimation system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0048] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, unless there is a conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0049] See Figure 1 As shown, this embodiment provides a panoramic room layout estimation method, including:
[0050] S1. Extracting global features from a panoramic image of a room through an ERP branch and processing the extracted features to obtain a global one-dimensional feature sequence of the panoramic image.
[0051] S2. Convert the panoramic image in step S1 into a cube map, extract local features from the cube map based on the Cube branch, and process the extracted features to obtain a local one-dimensional feature sequence of the cube map.
[0052] S3. Fusing the global one-dimensional feature sequence in step S1 and the local one-dimensional feature sequence in step S2 to obtain a predicted horizontal line depth and a predicted room height.
[0053] It can be understood that a dual-branch architecture is adopted in this embodiment. The ERP branch calculates and normalizes the cosine similarity of the saliency map and the deformable convolution feature map, effectively solving the panoramic distortion problem and ensuring the accuracy of the global features. The Cube branch comprehensively captures the local features of the room by splicing multiple views of the cube map. Finally, the features of the two branches are combined, and the key parameters of the room layout are predicted through the linear layer, which not only solves the problem of panoramic distortion, but also improves the accuracy of the overall estimation through feature fusion.
[0054] Furthermore, step S1 specifically includes:
[0055] S11, extracting global features from the panoramic image through the ERP branch and outputting an original feature map;
[0056] S12, after processing the panoramic image through the deformable convolution DCNv3 to output a deformable convolution feature map, the slice similarity, normalization layer and standard deviation layer processing are performed in sequence to form a saliency map;
[0057] S13, adding the saliency map to the original feature map and performing high compression through a high compression module to obtain the global one-dimensional feature sequence.
[0058] It is understandable that in this embodiment, a saliency map is introduced to address the panoramic distortion problem present in the ERP branch. The saliency map is used to identify important areas in the image, and then the saliency map is added to the original feature map containing global feature information to enhance the feature expression of the key areas. Next, the features are reduced in dimensionality by a high compression module to generate a concise and effective feature representation of a global one-dimensional feature sequence for subsequent analysis and application. The addition of the original feature map ensures the comprehensiveness of the features and avoids information loss. By reducing the feature dimension, the high compression module reduces computational complexity and storage requirements while maintaining key information.
[0059] Furthermore, the ResNet-34 model is used as the ERP branch in step S11.
[0060] Furthermore, step S2 specifically includes:
[0061] S21, converting the panoramic image into a cube map;
[0062] S22, splitting the six faces of the cube map in step S21, horizontally splicing the top, bottom, left, and right views to obtain a horizontal image, and vertically splicing the top view and the bottom view to obtain a vertical image;
[0063] S23, extracting local features from the horizontal image in step S22 through the Cube branch to obtain local features Figure 1, the local features Figure 1 A local one-dimensional feature sequence 1 is obtained by performing high compression through a high compression module;
[0064] S24, extract local features from the two views of the vertical image in step S22 through the Cube branch to obtain local features Figure 2 And local feature diagram 3, the local feature Figure 2 The height of local feature map 3 and local feature map 3 are compressed by the height compression module to obtain local one-dimensional feature sequence 2 and local one-dimensional feature sequence 3;
[0065] S25, linearize the local one-dimensional feature sequence 2 and the local one-dimensional feature sequence 3 through the linear layer to obtain the cube height h Cube .
[0066] Furthermore, S21, converting the panoramic image into a cube map, specifically includes the following steps:
[0067] Given a pixel p located on the i-th face of the cube, the coordinates of p in the camera system are (x, y, z), that is, pixel p = (p x , p y , p z ), perform equirectangular projection on pixel p. The specific formula is as follows:
[0068]
[0069] Where w represents the pixel width of the cube map; K represents the camera intrinsic parameter of the perspective projection; R i is the rotation matrix of the i-th face of the external matrix of any camera coordinate system; q represents the pixel The direction vector converted to the global coordinate system; θ represents the longitude in the equirectangular projection; φ represents the latitude in the equirectangular projection; q x ,q z ,q y are the x, y, and z components of q respectively.
[0070] It can be understood that in this embodiment, the panoramic image is first converted into a cube map, which simplifies the image representation. Then, the six faces of the cube map are split into four views of the top, bottom, left, and right, and a top view and a bottom view. The top, bottom, left, and right views are horizontally spliced to form a horizontal image, and the top and bottom views are vertically spliced to form a vertical image. Local features of the horizontal and vertical images are extracted through the Cube branch, and these features are compressed into a one-dimensional feature sequence using a height compression module, thereby simplifying the data structure and improving processing efficiency. Finally, these one-dimensional feature sequences are processed through a linear layer to ensure the linear processability of the data, and the cube height h is obtained. Cube , which helps in further data analysis.
[0071] Furthermore, the ResNet-34 model is used as the Cube branch of step S23 and step S24.
[0072] Furthermore, step S3 specifically includes:
[0073] The global one-dimensional feature sequence in step S1 is processed by the self-attention mechanism. The specific formula is as follows:
[0074]
[0075] Among them, Q is the query matrix; K is the key matrix; V is the value matrix; K T is the transposed matrix of the key matrix K; d k represents the dimension of the feature vector;
[0076] Q, K, and V are all learned from the equidistant rectangular image features through the fully connected layer, and Q, K, and V are passed into the image feature interaction mechanism for processing. The specific formula is as follows:
[0077] G'=(F×F T )×G;
[0078] F g =F+G'×F;
[0079] Among them, F is the output of the self-attention mechanism; F T represents the transposed matrix of F; G is a learnable weight matrix and is initialized to zero; Fg is the final output of the graph feature interaction mechanism;
[0080] Take Fg as the query, local one-dimensional feature sequence 1, local one-dimensional feature sequence 1 and local one-dimensional feature sequence 3 as the key and value, pass them into the cross attention mechanism for processing, and then linearize them through the linear layer to obtain the predicted horizontal depth and predicted room height.
[0081] It can be understood that in this embodiment, the self-attention mechanism is used to assign different weights to different positions by calculating the similarity between the query Q (Query), key K (Key) and value V (Value), thereby highlighting important information. The graph feature interaction mechanism uses the output of the self-attention mechanism to further process the features, and performs weighted summation through the learnable weight matrix G (initialized to zero) to achieve complex interactions between features, and obtain the final graph feature interaction output Fg, which further enhances the expressiveness of the features. The cross-attention mechanism establishes connections between different feature sequences and can comprehensively consider the influence of multiple features. Finally, the linear layer is used to perform a linear transformation on the output of the cross-attention mechanism to ensure the interpretability and stability of the output.
[0082] Further, step S3 further comprises:
[0083] S4, input evaluation data set, obtain evaluation result. The test image of the four panoramic image data sets of Zind, MatterportLayout, PanoContext and Standford2D3D is used as the evaluation data set, and the 3DIou and 2DIou indexes used for panoramic room estimation are used for evaluation.
[0084] On the other hand, in combination Figure 2 As shown in the accompanying drawings, the present application provides a panoramic room layout estimation system, which applies the panoramic room layout estimation method, comprising:
[0085] An ERP branch module is configured to process the global features extracted from the panoramic picture of the room by the ERP branch to obtain a global one-dimensional feature sequence of the panoramic picture;
[0086] A Cube branch module is configured to convert the panoramic picture into a cube map, process the local features extracted from the cube map based on the Cube branch to obtain a local one-dimensional feature sequence of the cube map;
[0087] A fusion module is configured to fuse and process the global one-dimensional feature sequence and the local one-dimensional feature sequence to obtain the horizontal line predicted depth and the room predicted height.
[0088] In another aspect, the present application further provides a computer device, comprising:
[0089] At least one processor;
[0090] At least one memory for storing at least one program;
[0091] When the at least one program is executed by the at least one processor, the at least one processor implements the panoramic room layout estimation method.
[0092] It can be understood that the panoramic room layout estimation system and the computer device in the embodiment have the same beneficial effects as the panoramic room layout estimation method, and will not be described here.
[0093] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.
[0094] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0095] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0096] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A panoramic room layout estimation method, characterized in that: include: S1. Extracting global features from a panoramic image of a room through an ERP branch and processing the extracted features to obtain a global one-dimensional feature sequence of the panoramic image. S2. Convert the panoramic image in step S1 into a cube map, extract local features from the cube map based on the Cube branch, and process the extracted features to obtain a local one-dimensional feature sequence of the cube map. S3. Fusing the global one-dimensional feature sequence in step S1 and the local one-dimensional feature sequence in step S2 to obtain a predicted horizontal line depth and a predicted room height.
2. A panoramic room layout estimation method according to claim 1, characterized in that: Step S1 specifically includes: S11, extracting global features from the panoramic image through an ERP branch and outputting an original feature map; S12, after processing the panoramic image through the deformable convolution DCNv3 to output a deformable convolution feature map, the slice similarity, normalization layer and standard deviation layer processing are performed in sequence to form a saliency map; S13, adding the saliency map to the original feature map and performing high compression through a high compression module to obtain the global one-dimensional feature sequence.
3. The method for estimating panoramic room layout according to claim 2, wherein: The ResNet-34 model is used as the ERP branch in step S11.
4. The method for estimating panoramic room layout according to claim 1, wherein: Step S2 specifically includes: S21, converting the panoramic image into a cube map; S22, splitting the six faces of the cube map in step S21, horizontally splicing the top, bottom, left, and right views to obtain a horizontal image, and vertically splicing the top view and the bottom view to obtain a vertical image; S23, extracting local features from the horizontal map in step S22 through a Cube branch to obtain a local feature map 1, and highly compressing the local feature map 1 through a high compression module to obtain a local one-dimensional feature sequence 1; S24, extracting local features from the two views of the vertical image in step S22 through the Cube branch to obtain a local feature map 2 and a local feature map 3, and compressing the heights of the local feature map 2 and the local feature map 3 through a height compression module to obtain a local one-dimensional feature sequence 2 and a local one-dimensional feature sequence 3; S25, linearize the local one-dimensional feature sequence 2 and the local one-dimensional feature sequence 3 through the linear layer to obtain the cube height h Cube .
5. The method for estimating panoramic room layout according to claim 4, wherein: S21, converting the panoramic image into a cube map, specifically includes the following steps: Given a pixel p located on the i-th face of the cube, the coordinates of p in the camera system are (x, y, z), that is, pixel p = (p x , p y , p z ), perform equirectangular projection on pixel p. The specific formula is as follows: Where w represents the pixel width of the cube map; K represents the camera intrinsic parameter of the perspective projection; R i is the rotation matrix of the i-th face of the external matrix of any camera coordinate system; q represents the pixel The direction vector converted to the global coordinate system; θ represents the longitude in the equirectangular projection; φ represents the latitude in the equirectangular projection; q x ,q z ,q y are the x, y, and z components of q respectively.
6. The method for estimating panoramic room layout according to claim 4, wherein: Use the ResNet-34 model as the Cube branch of step S23 and step S24.
7. The method for estimating panoramic room layout according to claim 1, wherein: Step S3 specifically includes: The global one-dimensional feature sequence in step S1 is processed by the self-attention mechanism. The specific formula is as follows: Among them, Q is the query matrix; K is the key matrix; V is the value matrix; K T is the transposed matrix of the key matrix K; d k represents the dimension of the feature vector; Q, K, and V are all learned from the equidistant rectangular image features through the fully connected layer, and Q, K, and V are passed into the image feature interaction mechanism for processing. The specific formula is as follows: G'=(F×F T )×G; F g =F+G'×F; Among them, F is the output of the self-attention mechanism; F T represents the transposed matrix of F; G is a learnable weight matrix and is initialized to zero; Fg is the final output of the graph feature interaction mechanism; Take Fg as the query, local one-dimensional feature sequence 1, local one-dimensional feature sequence 1 and local one-dimensional feature sequence 3 as the key and value, pass them into the cross attention mechanism for processing, and then linearize them through the linear layer to obtain the predicted horizontal depth and predicted room height.
8. A panoramic room layout estimation system, characterized in that: Applying the panoramic room layout estimation method according to any one of claims 1 to 7, comprising: An ERP branch module is used to extract global features from the panoramic picture of the room through the ERP branch and then process the extracted features to obtain a global one-dimensional feature sequence of the panoramic picture; A Cube branch module is used to convert the panoramic image into a cube map, extract local features from the cube map based on the Cube branch, and process the extracted local features to obtain a local one-dimensional feature sequence of the cube map; The fusion module is used to fuse the global one-dimensional feature sequence and the local one-dimensional feature sequence to obtain the horizontal line prediction depth and the room prediction height.
9. A computer device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor is enabled to implement the panoramic room layout estimation method according to any one of claims 1 to 7.