Panel representation and distortion reduction in 360 panoramic images
The panoramic images are processed through neural networks, and the panel geometry embedding network and local to global transformer network are used to solve the problem of distortion in panoramic images, and the continuity of depth maps, layouts and semantic maps is improved, and the understanding of the indoor environment is enhanced.
Patent Information
- Application Number
- CN202480004412.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-16
- Filing Date
- 2024-05-17
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to effectively deal with the distortion problem in panoramic images, resulting in the continuity of depth maps, layouts and semantic maps.
The neural network is used to process panoramic images, embed the network and local to global transformer network through panel geometry, encode the local and global geometric features of the panel, reduce the impact of distortion and improve the continuity of the image.
It effectively reduces distortion in panoramic images, improves the continuity of depth maps, layouts and semantic maps, and enhances the ability to understand the indoor environment.
Smart Images

Figure CN120051796A_ABST
Abstract
Description
Incorporation by Reference
[0001] This application is based on and claims priority to U.S. non-provisional patent application No. 18 / 666,438, filed on May 16, 2024, entitled “Panel Representation and Distortion Reduction in 360 Panorama,” and U.S. Provisional Patent Application No. 63 / 467,569, filed on May 18, 2023, entitled “Methods of understanding Indoor Environments using Panel Representation in 360 Panorama Images,” both of which are incorporated herein by reference in their entirety. Technical Field
[0002] The present disclosure relates generally to image processing, and in particular to processing panoramic images using neural networks to generate depth maps, layouts, semantic maps, etc. with reduced distortion and improved continuity. Background Art
[0003] Extracting / predicting semantic content and identifying objects from digital images using computer vision techniques is essential for many autonomous applications. Panoramic images generated in various formats can differ from typical perspective 2D images in terms of various geometric and other properties. Computer vision techniques and architectures for processing panoramic images can be designed to explore such properties in order to improve prediction accuracy and reduce the negative effects of panoramic distortion. Summary of the invention
[0004] The present disclosure relates generally to image processing, and in particular to processing panoramic images using neural networks to generate depth maps, layouts, semantic maps, etc. with reduced distortion and improved continuity. Methods and systems are described for generating such maps by exploiting several basic properties of these panoramic images and by using a panoramic panel representation and a neural network framework. A panel geometry embedding network is incorporated for encoding local and global geometric features of the panel to reduce the negative effects of panoramic distortion. A local-to-global transformer network is also incorporated for capturing geometric context and aggregating local information within a panel and global context on a panel-by-panel basis.
[0005] In some example embodiments, a method for processing a panoramic image dataset by a computing circuit is disclosed. The method may include generating a plurality of data panels from the panoramic image dataset; executing a first neural network to process the plurality of data panels to generate an embedding set representing geometric features of the plurality of data panels; executing a second neural network to process the plurality of data panels and the embedding set to generate a plurality of mapping panels; and fusing the plurality of mapping panels into a mapping dataset of the panoramic image dataset.
[0006] In the above example implementation, the mapping dataset includes one of a depth map, a layout map, or a semantic map corresponding to the panoramic image dataset.
[0007] In any of the above-mentioned example embodiments, the panoramic image data set includes a two-dimensional data array; and each of the multiple data panels includes a sub-array of the data array, which is in the entirety of a first dimension of the two dimensions and in a segment of a second dimension of the two dimensions.
[0008] In any of the above example implementations, the first dimension represents a gravity direction of the panoramic image dataset, and the second dimension represents a horizontal direction of the panoramic image dataset.
[0009] In any of the above-mentioned example embodiments, the multiple data panels are continuously generated from the panoramic image data set using a window, the window having a length of the entirety of the first dimension in the first dimension and a predetermined width in the second dimension, and the window slides a predetermined step along the second dimension.
[0010] In any of the above example implementations, the window slides continuously from one edge of the panoramic image dataset in the second dimension to another edge of the panoramic image dataset in the second dimension.
[0011] In any of the above example embodiments, the first neural network is configured to encode local and global geometric features of multiple data panels to reduce the effects of geometric distortion in the panoramic image dataset and enhance the preservation of geometric continuity across multiple mapping panels.
[0012] In any of the above example embodiments, the first neural network includes a multilayer perceptron (MLP) network LP, which is used to process geometric information extracted from multiple data panels to generate an embedding set including a global geometric feature set and a local geometric feature set of the multiple data panels.
[0013] In any of the above example implementations, the second neural network is configured to process the plurality of data panels based on the embedding set and reduce geometric distortion in the panoramic image dataset.
[0014] In any of the above example implementations, the second neural network includes: a downsampling network; a transformer network; and an upsampling network.
[0015] In the above-mentioned example implementation, the downsampling network is configured to process multiple data panels and embedding sets to generate a series of downsampled features with decreasing resolution; the transformer network is configured to process the downsampled features with the lowest resolution to generate transformed low-resolution features; and the upsampling network is configured to process the transformed low-resolution features and a series of downsampled features to generate the multiple mapping panels.
[0016] In any of the above example embodiments, the transformer network includes a feature processor.
[0017] In any of the above example embodiments, the feature processor is configured to increase continuity of the geometric feature.
[0018] In any of the above example implementations, the feature processor is configured to aggregate local information within each of a plurality of data planes to capture per-panel context.
[0019] Aspects of the present disclosure also provide an electronic device or apparatus, which includes a circuit or processor configured to execute any one of the above method embodiments.
[0020] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions, which, when executed by an electronic device, cause the electronic device to perform any one of the above method embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:
[0022] Figure 1 An example 360° panoramic image of an indoor scene is shown in equirectangular representation.
[0023] Figure 2 Shows Figure 1 Polar coordinate view of 360 panoramic images for a more realistic perspective.
[0024] Figure 3a Shows from Figure 1 Example depth map extracted from a 360° panoramic image.
[0025] Figure 3b Shows from Figure 1 Example semantic graph extracted from a 360° panoramic image.
[0026] Figure 3c Shows from Figure 1 Example layout for 360° panoramic image extraction.
[0027] Figure 4 An example block diagram of a system including PanelNet for processing panoramic images to generate depth maps, semantic maps, and layouts is shown.
[0028] Figure 5 Shows Figure 4 An example implementation of PanelNet.
[0029] Figure 6 Shows Figure 4 Another example implementation of PanelNet.
[0030] Figure 7 Shows Figure 6 An example processing component of PanelNet.
[0031] Figure 8 Shows Figure 7 An example processing component of the PanelNet local-to-global transformer network.
[0032] Fig. 9 An example processing flow for panel fusion to generate a mapped dataset is shown.
[0033] Fig.10 A schematic diagram of a computer system according to an example embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0034] Throughout the specification and claims, terms may have subtle meanings that are suggested or implied by the context beyond the explicitly stated meaning. The phrases "in one embodiment / implementation" or "in some embodiments / implementations" as used herein do not necessarily refer to the same embodiment / implementation, and the phrases "in another embodiment / implementation" or "in other embodiments" as used herein do not necessarily refer to different embodiments. For example, the claimed subject matter is intended to include all or part of the combination of the exemplary embodiments / implementations.
[0035] In general, terms can be understood at least in part from their use in context. For example, terms such as "and", "or", or "and / or" as used herein can include various context-dependent meanings. Typically, "or", if used in an association list, such as A, B, or C, is intended to mean A, B, and C, which are used herein in an inclusive sense, and A, B, or C, which are used herein in an exclusive sense. In addition, the terms "one or more", "at least one", "one", "an", or "the" used herein depend at least in part on the context and can be used in a singular sense or in a plural sense. In addition, the term "based on" or "determined by..." can be understood as not necessarily intended to express an exclusive set of factors, but can allow for the presence of additional factors that are not necessarily explicitly described again, which depends at least in part on the context.
[0036] The following disclosure relates generally to image processing, and specifically to processing panoramic images using neural networks to generate depth maps, layouts, semantic maps, etc. with reduced distortion and improved continuity. Methods and systems are described for generating such maps by exploiting several basic properties of these panoramic images and by using a panoramic panel representation and a neural network framework. A panel geometry embedding network is incorporated to encode local and global geometric features of the panel to reduce the negative effects of panoramic distortion. A local-to-global transformer network is also incorporated to capture geometric context and aggregate local information within a panel and global context on a panel-by-panel basis. Panoramic image format
[0037] Compared to conventional perspective 2D images, such as those taken from ordinary cameras with a fixed, limited viewing angle, panoramic images provide a wide (e.g., 360 degree) field of view (FoV) around a particular viewing axis. Panoramic images can be taken by a dedicated camera with a wide / multi-lens optical system, or can be created by stitching together multiple overlapping conventional images taken from multiple viewing angles around the viewing axis. A complete panoramic image can be referred to as a 360 panorama and can be represented in digital form using one of several example data formats. One example format can be based on an equirectangular projection (ERP) representation. Similar to conventional images, a panoramic image in an ERP format can be represented by a data set in a 2D pixel array, but can include complete 360 panoramic information about the viewing axis. Each pixel of a panoramic image of a 360 scene in the ERP formatted data set contains perspective imaging information (e.g., RGB or YUV values) corresponding to a perspective viewing stereo angle unit in the 360 scene.
[0038] Figure 1An ERP representation of an example 360 panoramic image is illustrated in FIG. The example 360 panoramic image is represented by a 2D data array 100 along a horizontal direction 102 and a vertical direction 104. The example 360 scene representation is generated around a viewing axis along the vertical axis 104. Thus, the data array 100 contains image content having a continuously varying viewing angle horizontally over a full 360 degree range and having a fixed perspective FoV within a specific vertical viewing angle range. Thus, as Figure 1 The indicated horizontal direction represents the panning direction of the 360-degree panoramic image. Figure 1 In the example of , the 2D pixels are evenly distributed in the horizontal direction 102 and the vertical direction 104 .
[0039] Figure 2 Shows Figure 1 An example ERP formatted 360° panoramic image 3D view 200 showing the Figure 1 A 2D pixel array is more realistic than a perspective image. Figure 2 The 3D FoV in is basically characterized by two perspective views, one with a FoV range in the vertical direction and the other spanning all 360 degrees around the vertical axis (i.e., the translation axis), as further described below. Figure 1 The 2D data array of inherently includes large perspective distortions toward the upper edge 110 and the lower edge 112 because these edges correspond in their entirety to Figure 2 The 3D perspective view of FIG. 2 is at its poles 202 (zenith) and 204 (nadir). In other words, Figure 2 The reduced 3D perspective viewing range in the translation direction towards the pole is stretched to Figure 1 The ERP formatted pixel array has a complete pixel range in the horizontal direction 102.
[0040] The above 2D ERP representation for 360° panoramic images is essentially continuous in the translation direction. Figure 1 The vertical edges 106 and 108 of the 2D data array represented by the ERP maintain continuity.
[0041] The above panoramic images can be generated from any 360-degree scene. For example, such a 360-degree scene can be in an indoor or outdoor environment. A real or natural scene in any of such environments can be characterized by the alignment of objects determined by gravity. In this disclosure, for the purpose of clarity and illustration, the 360-degree scene can be generated from any 360-degree scene. For example, such a 360-degree scene can be in an indoor or outdoor environment. A real or natural scene in any of such environments can be characterized by the alignment of objects determined by gravity. Figure 1 and Figure 2 In FIG. 1 , gravity is considered as the vertical direction. Therefore, a typical panoramic image can be generated by translating around the gravity axis (along the vertical direction) and in the horizontal plane. However, the various embodiments below are not limited to this panoramic geometric arrangement. Panoramic Image Information Extraction Based on Computer Vision
[0042] For some applications and image analysis tasks, additional information can be extracted or generated from the input image dataset based on computer vision techniques and modeling. Such information extraction or generation may include, but is not limited to, image classification, image segmentation, object identification and recognition, depth estimation, object layout generation, etc. Just as with conventional 2D image imaging, such information can also be extracted or generated from panoramic images using computer vision modeling. For example, the input for such information extraction or generation can be the above-mentioned Figure 1 Described is an ERP dataset of panoramic images.
[0043] The information extracted or generated from the panoramic images described above can be of various types and complexities. For example, the output of a classification model can be simple and compact. On the other hand, the output dataset representing the extracted or generated information (such as depth estimation information at each pixel, object layout, and semantic information of the pixels of the panoramic image) may be more complex. Such information can be used to construct, for example, depth maps, semantic maps, and layouts.
[0044] As an example, Figure 3a , Figure 3b and Figure 3c The diagrams show the Figure 1 The depth map, semantic map, and layout map extracted / generated by the ERP representation of the example panoramic image 100. These example depth maps, semantic maps, and layout maps are generated in the same ERP representation. In other words, each pixel in these maps can correspond to Figure 1 However, the contents of these images represent the extracted depth, semantic and layout information at the pixel level, rather than Figure 1 Optical information directly measured in the input panoramic image 100. Such information in a 2D array of similar or different size to the original data array of the panoramic image can be referred to as a mapping dataset. Such a mapping dataset can be extracted or generated based on computer vision techniques and modeling. Therefore, depending on the details of the type of mapping dataset extracted or generated for a specific target task, such modeling can involve functions including but not limited to object recognition / identification, depth estimation, semantic extraction, etc. The mapping dataset so generated can then be used to predict, for example, a 3D model of the spatial content of the panoramic image.
[0045] For example, panoramic depth prediction can be generated via computer vision modeling to determine the 3D position of objects identified in a panoramic image. For another example, panoramic layout prediction can be generated via computer vision modeling to obtain the layout structure of a captured layout embedded in a panoramic image of a 3D scene. For another example, panoramic semantic segmentation can represent another important dense prediction task to generate pixel semantic information for understanding the content of a panoramic image.
[0046] Example computer vision models for generating the above-mentioned mapping data sets may include, for example, various neural networks (e.g., convolutional neural networks, CNNs) and / or other data analysis algorithms or components. The above-mentioned various example graphs may be crucial for many practical applications. In indoor environments, such practical applications may include, but are not limited to, room reconstruction, robot navigation, and virtual reality applications. Early methods focused on modeling indoor scenes using perspective images. With the development of CNNs and omnidirectional photography, panoramic images have become candidates for the above-mentioned mapping data sets. Compared with the use of traditional perspective images, panoramic images have a larger FoV and provide geometric context, especially the geometric context of the indoor environment, which can be learned in a continuous manner via training.
[0047] Although the ERP format provides a convenient representation of panoramic images, modeling the overall panoramic scene in the ERP format or representation through computer vision can be challenging. For example, as mentioned above, ERP distortion increases when ERP pixels are close to the zenith or nadir of a panoramic image, which may reduce the ability of a convolutional network structure that can be included in a computer vision model designed for undistorted perspective images.
[0048] In some example implementations for eliminating the effects of the above-mentioned ERP distortion, the panorama can be first decomposed into perspective patches, such as slice images, so that the computer vision model can be configured to extract image information and features at the patch level, where the relative distortion across patches (e.g., the relative distortion within each patch) is small. However, partitioning a typical gravity-aligned panoramic image into discontinuous patches may destroy the local continuity of gravity-aligned scenes and objects, thereby still limiting the performance of typical distortion-free modeling.
[0049] The example embodiments disclosed below further provide a computer vision system that relates to a partitioning method and a neural network architecture for processing panoramic images to extract or generate imaging information for understanding content included in the panoramic image, which imaging information can eliminate the effects of ERP distortion and simultaneously maintain continuity between image partitions. While these embodiments are particularly suitable for information extraction and understanding of indoor panoramic scenes (including gravity-aligned indoor objects and generated in ERP representations), they can also be applied to analyzing panoramic images in other environments and in representations other than ERP formats. The various neural networks in these embodiments can be pre-trained using training panoramic images and can be retrained and updated as more training data sets become available.
[0050] Specifically, the input ERP panoramic image can be partitioned in a continuous manner, and these partitions can be processed by a neural network structure called PanelNet, which can be designed, configured and trained to handle major panoramic understanding tasks such as depth estimation, semantic segmentation and layout prediction. In some example embodiments, only the last layer or layers of the decoder in PanelNet may need to be slightly modified to accommodate the different extraction / generation tasks described above.
[0051] In some example embodiments further described below, PanelNet can be based on at least two basic properties of the ERP representation of the panoramic image: (1) the ERP representation of the panoramic image is continuous and seamless in the horizontal direction, as described above with respect to Figure 1 described; and (2) gravity plays an important role in object alignment in typical panoramic scenes, especially in indoor environment design, which makes it crucial to design and tune PanelNet for extracting gravity-aligned features. As described in further detail below, the example panel representation of ERP can be adapted to these basic properties for processing by PanelNet, thereby improving the performance of extracting / generating / predicting the above-mentioned mapping datasets. For example, an ERP dataset of a panoramic image can be partitioned into continuous panels with corresponding global and local 3D geometries, which preserves gravity-aligned features within the panel and maintains global continuity across panels. PanelNet can include a geometric embedding network for panel representation that encodes local and global features of the panel for processing by an encoder within PanelNet to reduce the negative impact of ERP distortion without adding further explicit distortion processing networks. In some example embodiments, a transformer network can be included as a feature processor that can perform local-to-global feature processing and extraction by using local information aggregation of window blocks and accurate panel-by-panel context capture using panel blocks to further enhance continuity. Example overall architecture:
[0052] Figure 4 An example overall implementation 400 for processing an ERP representation 402 of a panoramic image is illustrated. The example implementation 400 includes generating a panel representation 404 including a continuous panel 406 of the ERP and a geometric representation 408 of the panel; the panel representation 404 is processed by PanelNet 410 to generate one or more mapping data sets 420 related to at least one of, for example, a depth map 422, a semantic map 424, and a layout 426.
[0053] Figure 5 Shows Figure 4Further details of an example implementation 500 of PanelNet 410 are provided below. Figure 5 As shown, the panel representation 404 can be processed by PanelNet 410. PanelNet 410 can be implemented in an encoder-decoder manner, including a panel encoder 502, a panel decoder 504, and a fusion network 506 that sequentially processes the panel representation 404 and is used to generate one or more mapping data sets 420. The fusion network 506 can be configured to combine the processed and mapped panels into one or more mapping data sets 420.
[0054] Figure 6 Shows Figure 4 Further details of another example implementation 600 of PanelNet 410 are provided. Figure 5 Compared to the example implementation 500, Figure 6 The example implementation 600 further includes a transformer 608 in PanelNet 410. Thus, the panel representation 404 may be processed sequentially by the panel encoder 602, the transformer 608, the panel decoder 604, and the fusion network 606 to generate one or more mapping data sets 420. There may be Figure 5 and Figure 6 There are other inter-block data dependencies between the various blocks of PanelNet 410 , but they are not explicitly shown for simplicity.
[0055] Figure 7 Further details are shown following Figure 4 and Figure 6 Example implementation 700. Example implementations may include panel representation generation 404, encoder network 602, transformer network 608, decoder network 604, fusion network 608, further details are provided below. Vertical panel partitioning and geometric embedding
[0056] exist Figure 7 In some example embodiments, an input ERP representation of a panoramic image in, for example, RGB or YUV may include H e ×W e The input ERP representation may be partitioned to generate a plurality of ERP panels 408 using a masking window that continuously slides through the 2D array of the ERP representation 402 in the horizontal direction (the panning direction of the panoramic image). The masking window may have a height H that spans the entire vertical direction of the ERP representation 402. e , and can have a width or interval I (in pixels). The masking window can slide with a step size S. Since the two vertical edges of the input ERP ( Figure 1106 and 108) are continuous, the masking window will slide horizontally from one end of the ERP representation to the other by crossing one vertical edge into the other vertical edge. The number of continuous panels thus generated will be N = W e / S. In some example embodiments, the stride S may be smaller than the window width or interval I, and thus, adjacent panels may overlap to enhance the horizontal continuity between panels. Thus, in the example embodiments of the above panels, a continuous and seamless vertical panel is generated. Because in such example embodiments the ERP representation 402 is not partitioned in the vertical direction, the ERP panel 408 thus generated will therefore be continuous in the vertical direction without additional tiles.
[0057] like Figure 7 As shown, a geometric representation of the ERP panel 408 may be further generated. This geometric representation ( Figure 4 406) may include a geometric embedding 705 generated by a geometric parameterization process 702 of the vertical ERP panel 408, and then generated by a geometric embedding network 704 (which may be implemented as a multi-layer perceptron (MLP) network as described in further detail). In conjunction with the geometric embedding network 704, the geometric representation of the panel generated as the geometric embedding 705 may help reduce the negative effects of panoptic distortion in the ERP representation 402. The geometric embedding may include a trained multi-dimensional embedding space, and each of the geometric embeddings may be a vector in the trained multi-dimensional embedding space and may represent each set of geometric parameters parameterized via the geometric parameterization process 702.
[0058] In some example embodiments, the geometric embedding generation process may be configured to combine the geometric features of the EPR panel with the image features and thereby reduce the negative impact of ERP distortion. e (x e ,y e )(where x e and e , respectively, representing the horizontal and vertical coordinates of a pixel in the ERP representation, will correspond to the azimuth and polar angles representing the corresponding direction in the FoV. and θ. Thus, the EPR has a pixel position x e ,y e The pixels can be mapped to the angular direction The angular direction can be further converted to the unit sphere P in the FOV s The absolute 3D world coordinate P on s (x s ,y s ,zs ), which has the following conversion relationship:
[0059] Then, the transformed 3D world coordinates P of all pixels in all panels are s (x s ,y s ,z s ) can be used to generate global geometric features. Since each ERP panel above has the same distortion profile in the vertical direction as any other ERP panel (as determined by the way the ERP is partitioned into vertical panels), the relative position of each pixel with respect to the panel it is in is also important. In some example embodiments, a relative 3D local position P(x ′ ,y ′ ,z ′ ). The global 3D world coordinates of the randomly selected ERP panels can be selected to represent the relative 3D positions of all ERP panels. Due to the vertical partitioning method used to generate the ERP panels, z s will be equal to z ′ Thus, the final output parameter set of a point on the ERP panel from the geometric parameterization process 702 may be its local coordinates and global coordinates (x s ,y s ,z s ,x ′ ,y ′ This set of geometric parameters for each of the pixels of the ERP panel can then be input into a geometric embedding network 704 to generate a geometric embedding 705.
[0060] like Figure 7 As further shown in , a geometric embedding representing both global and local geometric features may be generated by an MLP network 704, which may be implemented, for example, as a two-layer MLP network.
[0061] Various transformations between pixel positions and ERP pixel positions and 3D world coordinates will be determined by the partitioning of the vertical ERP panel. As described above, the vertical ERP panel partitioning is determined by the width I and stride S of the sliding window. Therefore, given I and S, local and global geometric features are determined and generated together as part of the geometric embedding 705.
[0062] Therefore, global geometric features, called global geometric embeddings, can be extracted across ERP panels to record the location of segmented panels in the panorama. The global geometric information can include, for example, panel location information in the panorama ERP image, such as the panel center pixel location in the ERP image and the boundary range of each panel in the ERP image, the spherical geometry in the corresponding spherical coordinates, etc. These global features can be used across ERP panels by the decoder network (described in further detail below) when processing each panel, such as Fig. 9 , where the decoder network is shown as being used to process each panel (in Fig. 9 ) to generate a mapping panel for fusion. Encoder Network
[0063] like Figure 7 As further shown in , the encoder network 602 can be configured to process the panel representation 404 with the help of the geometric embedding 705. For example, the encoder network 602 can include a multi-layer downsampling neural network to process the vertical ERP panel 408 to generate feature maps with decreasing resolutions. For example, the multi-layer downsampling neural network of the decoder 602 can be based on a ResNet-34-based neural network architecture as a feature extractor. As an example only, the multi-layer downsampling neural network can be configured to generate feature maps at 4 different scales (or resolutions). Therefore, the multi-layer downsampling neural network of the decoder 602 can correspondingly include 4 stages of downsampling network layers, such as Figure 7 The highest resolution layer (input layer) and the lowest resolution layer (output layer) are indicated by 706 and 708, respectively.
[0064] In some example embodiments, a 1×1 convolutional layer may be applied to reduce the dimension of the final feature map for each EPR panel to f b ∈R Cb×Hb×Wb , where, for example, for any interval and stride, H b =H e / 32,W b =I / 32, C b =D / (H eb ×W eb ), and D=512.
[0065] like Figure 7As further shown in FIG. 6 , the geometric features included in the geometric embedding 705 can be added to the first layer 706 (highest resolution layer) of the encoder network 602 to make the encoder network 602 aware of the ERP distortion. For example, a geometric embedding vector or feature associated with the position of each pixel position can be generated as described above and combined with the image content of the corresponding pixel of the vertical panel at the highest resolution layer 706 of the encoder 602, so that both global and local geometric features are incorporated into the encoder network 602. Converter Network
[0066] like Figure 7 As further shown in the example implementation of , the final feature map 708 (eg, having the lowest resolution) can then be used as input to the transformer network 608 .
[0067] In some examples, transformer network 608 may be implemented as a local-to-global transformer network for performing information aggregation, as described in further detail below, to specifically extract remote dependencies in remote ERP dashboards.
[0068] Specifically, although partitioning the ERP into continuous vertical panels via sliding windows as described above can help maintain the continuity of structures or objects in the panoramic scene, capturing long-range dependencies is still crucial. Since the ERP representation is seamless in the horizontal direction, two vertical ERP panels that are far apart on the panorama may have a closer real-world distance and therefore be correlated. Such correlation may not be easily captured. To address this issue and further improve local information aggregation, the local-to-global transformer network 608 can be designed and configured to include at least two major important components: (1) a window block for enhancing geometric relationships within a local panel, and (2) a panel block for capturing long-range context between panels. Example local-to-global transformer network in Figure 8 Shown as 800.
[0069] In some example embodiments, for each ERP panel, the input feature map f from the last layer of decoder 602 is b ∈R Cb×Hb×Wb Can be shaped into flattened 2D feature tiles where (P×P) represents the size of the feature block, and N w =H b W b / P 2 is the number of feature patches in the current window block. In some example embodiments, P can be 1, 2, or 4 for window blocks of different resolutions. A learnable position embedding Used to maintain the location information of feature tiles.
[0070] In the panel block, global information can be aggregated via per-panel multi-head self-attention. The feature maps of all panels can be compressed into N 1-D feature vectors f p ∈R N×D , and is then used as a token in the panel block. Similar to the window block, the learnable position embedding E p ∈R N×D Added to the token to preserve tile-by-tile position information.
[0071] In some example embodiments, and as Figure 8 As shown in the example local-to-global transformer network 800, a multi-head self-attention module (MSA) 802 and a feed-forward network (FFN) 804 can be stacked together. As shown in 806 and 808, a LayerNorm (LN) operation can be performed before each MSA and FFN. Figure 8 As further shown, the local to global transformer block can be calculated as where l is the block number at each stage, and and z l Represents the output feature map of window / panel-MSA and FFN. In order to aggregate features from local to global, window blocks can be stacked in order from small to large according to the window size. Panel blocks can be stacked after window blocks, such as Figure 8 As shown. In some example embodiments, multiple (e.g., 12) transformer blocks may be used, and they may be placed in the following example order: low-resolution window block (2), medium-resolution window block (2), high-resolution window block (2), panel block (6). Since the compression operation in the panel block reduces the impact of local information aggregation performed by the window block, the performance of the transformer network may be degraded when the order is disrupted. Decoder network and fusion network
[0072] like Figure 7 As illustrated, the output from the local-to-global transformer network 608 may be provided to a decoder network 604. The decoder network 604 may be implemented as a multi-layer or multi-stage upsampling network for restoring the original resolution of the ERP panel.
[0073] In some example embodiments, for each decoder layer, its feature map may be concatenated with the feature map generated by the corresponding layer or stage in encoder 602, such as Figure 7As indicated by the vertical arrow from encoder 602 to decoder 604 in FIG. 6 . In this way, the up-conversion stage decoder 604 can be configured to correspond inversely to the down-conversion stage of encoder 602. Therefore, the multi-layer upsampling network can be implemented by multi-stage up-convolution to gradually restore the feature map to the input resolution of the ERP panel.
[0074] The output from the decoder 604 may represent a panel-by-panel mapping data set and may be referred to as a mapping panel. The mapping panel at each pixel may contain prediction information, such as depth information, layout information, or semantic information, rather than the original RBG or YUV image information. The mapping panels may then be fused or merged together to form one or more overall mapping data sets corresponding to the input ERP.
[0075] In some example embodiments, a learnable confidence map may be predicted by the fusion network 608 to improve the final merge or fusion result. For the final merge, the predictions of all mapping panels may be averaged. In some example embodiments, a network for predicting one type of mapping dataset (e.g., layout) may be trained by slightly modifying the network structure. Figure 7 The model can be used to predict the mapping dataset for another 360-degree dense prediction task (e.g., semantic segmentation). Figure 7 In an example of applying the model to layout prediction in an indoor environment, the LGT-Net representation can be used to represent the room layout as floor boundaries and room heights. In some example embodiments, one linear layer for generating floor boundaries and two linear layers for generating room heights can be added after the last decoder layer. Figure 7 In some example embodiments, a default length of the output 1-D floor border may be 1024. Loss Function
[0076] Can be trained jointly or in stages Figure 7 A model of (by iteratively fixing some subnetworks and training other subnetworks). The loss function can be designed according to the specific prediction task. As an example, for depth estimation, the loss function can be based on the reverse Huber loss (BerHu). In some example embodiments, training can be performed in a fully supervised manner. An example BerHu loss function can be expressed in the following form: Where e represents the error term, and the error threshold c is used to determine where the switch from L1 loss (e.g., minimum absolute deviation) to L2 loss (e.g., minimum squared error) occurs. A combination of L1 loss, normal loss, and normal gradient loss for horizontal depth and room height can be optimized to train a model for indoor environments. Figure 7 Model.
[0077] For semantic segmentation prediction, in some example embodiments, a loss function based on a cross entropy loss with per-class weights may be used. Example PanelNet training and testing
[0078] To test the above PanelNet implementation, a real-world dataset consisting of 1,413 panoramas collected in 6 large-scale indoor regions, called Stanford2D3D, was used. For depth estimation, the dataset was split into region 1, region 2, region 3, region 4, region 6 for training, and region 5 for testing. For semantic segmentation, a 3-fold split of the dataset was used for training, evaluation, and testing. The example resolutions for depth estimation and semantic segmentation were 512×1024 and 256×512, respectively.
[0079] In addition, datasets called PanoContext and extended Stanford2D3D are also used for training and testing of the above PanelNet implementation. These two datasets include two cuboid room layout datasets. For example, PanoContext contains 514 annotated cuboid room layouts collected from the SunCG dataset. Specifically, 571 panoramas were collected from Stanford2D3D and annotated with room layouts. The input resolution of both datasets is 512×1024. For these datasets, the same example segmentation used for training and testing above was adopted.
[0080] In addition, a dataset called Matterport3D can also be used. This dataset includes a large-scale RGB-D dataset containing 10,800 panoramic images collected in 90 indoor scenes. This dataset can be used in particular for our depth estimation training evaluation and testing. The dataset can be split into 7829 panoramas from 61 houses for training, and the rest for testing. A resolution of 512×1024 can be used for training and testing.
[0081] In addition, for depth estimation, the performance of PanelNet implementations can be evaluated using standard depth estimation metrics, including mean relative error (MRE), mean absolute error (MAE), root mean square error (RMSE), log-based root mean square error (RMSE(log)), and threshold-based accuracy, such as δ 1 , δ 2 and δ 3For semantic segmentation, the performance of PanelNet implementations is evaluated using standard semantic segmentation metrics including class-wise mIoU and class-wise mAc. For layout prediction, the performance of PanelNet implementations is evaluated using 3D Intersection over Union (3DIoU).
[0082] One or more PanelNet models can be implemented in PyTorch and trained on, for example, eight NVIDIA GTX1080Ti GPUs with a batch size of 16. The network is trained using the Adam optimizer and the initial learning rate can be set to 0.0001. For depth estimation, the network / model can be trained on the above-mentioned Stanford2D3D dataset for, for example, 100 rounds and on the above-mentioned Matterport3D dataset for, for example, 60 rounds. The network / model can be trained on the above-mentioned semantic segmentation dataset for 200 rounds and on the above-mentioned layout prediction dataset for 1000 rounds. Random flipping, random horizontal rotation, and random gamma enhancement can be further employed for data augmentation. Example default strides and intervals for depth estimation of 32 and 128, respectively, can be used, while the stride can be set to, for example, 16 for semantic segmentation.
[0083] The above methods for PanelNet can be evaluated against the state-of-the-art panoramic depth estimation algorithms in Table 1 below. The results can be the average of the best results of three training sessions. The results of SliceNet on Stanford2D3D are reproduced by fixing the metric and retraining and re-evaluating the Omnifusion model for 2 iterations on the corresponding Matterport3D dataset. Table 1 shows that the PanelNet model implementation outperforms the existing models on all metrics of both datasets. Table 1
[0084] Compared to the PanelNet implementation, methods that directly operate on the panorama predict a continuous background but lack object details. Fusion-based methods generate clear depth boundaries, while artifacts caused by patch-by-patch differences lead to inconsistent depth predictions, which cannot be removed with their patch fusion modules or iterative mechanisms. However, with the help of the above-mentioned local-to-global transformer network, the PanelNet implementation maintains the geometric continuity of the room structure and shows excellent performance even in some challenging scenes. The PanelNet model is also able to generate clear object depth edges.
[0085] PanelNet is further evaluated against state-of-the-art panoptic semantic segmentation methods, as shown in Table 2 below. For example, the PanelNet model improves the mIoU metric by 6.9% and the mAcc metric by 8.9% compared to the existing Ho-HoNet implementation. PanelNet provides a strong ability to segment out objects with smooth surfaces. The segmentation edges generated by PanelNet appear natural and continuous. This is because the local-to-global transformer network is able to successfully capture the geometric context of the object. The PanelNet model is also able to segment small objects from the background. The segmentation boundaries of the ceiling and walls generated by the PanelNet model are highly smooth, indicating the ability of the panel geometry embedding network to learn ERP distortion. Table 2 method enter QUR mAcc TangentImg RGB-D 41.8 54.9 HoHoNet RGB-D 43.3 53.9 PanelNet RGB 46.3 58.7
[0086] The PanelNet model approach is further evaluated against state-of-the-art panoramic layout estimation methods, as shown in Table 3 below. By adding a linear layer at the end of the depth estimation network as described above, the PanelNet model achieves competitive performance against existing technology methods specifically designed for layout estimation. Since the PanelNet model was originally designed for dense prediction, it suffers from information loss during upsampling and channel compression. The PanelNet-based layout model shares the same structure as the depth estimation model before the linear layer. The PanelNet model can be activated with weights pre-trained on the depth estimation dataset to reduce training overhead. The PanelNet-based layout prediction model performs best when the stride is set to 64 and the interval is set to 128. Table 3 method PanoContext Stanford2D3D LayoutNet v2 85.02 82.66 DuLa-Net v2 83.77 86.60 HorizonNet 84.23 83.51 AtlantaNet - 83.94 LGT-Net 85.16 85.76 PanelNet 84.52 85.91
[0087] Ablation studies can be further performed to evaluate the impact of various elements and hyperparameters of PanelNet on, for example, the Stanford2D3D dataset for depth estimation, as shown in Table 4 below. For all networks, the stride can be set to an example value of 32 and the interval can be set to an example value of 128. A baseline model with a ResNet-34 encoder and a deep decoder as illustrated above can be used. Since partitioning the entire panorama into overlapping vertical panels greatly increases the computational complexity, ResNet-34 instead of a visual transformer can be used as the backbone network (encoder and decoder). As shown in Table 4, the performance improvement of adding a panel geometry embedding network to the pure CNN structure of PanelNet may be small because the network has a low ability to aggregate distortion information with image features. However, by applying a local-to-global transformer network as a feature processor, the baseline network achieves significant performance improvements on all evaluation metrics. Benefiting from the information aggregation ability of the local-to-global transformer network, the panel geometry embedding network more fully exploits its ability to perceive distortion and improves performance both quantitatively and qualitatively. The combination of the local-to-global transformer network and the panel geometry embedding network leads to the clearest object edges in depth estimation. For panel patches, we further evaluate the effect of per-panel relative position embedding similar to LGT-Net. However, it seems to bring minimal performance improvement of depth estimation while increasing computational complexity. Table 4 method Training Memory MRE MAE RMSE <![CDATA[δ 1 ]]> <![CDATA[δ 2 ]]> Baseline 10231 0.1033 0.1859 0.3212 0.8976 0.9741 Baseline+Geo(G) 10371 0.1029 0.1861 0.3205 0.8980 0.9765 Baseline+Geo(G+L) 10509 0.1000 0.1815 0.3149 0.9012 0.9775 Baseline + Transformer (P) 10359 0.0904 0.1652 0.3058 0.9123 0.9776 Baseline + Transformer (P+W) 10379 0.0854 0.1610 0.3016 0.9164 0.9785 Baseline+Geo(G+L)+Transformer(P) 10639 0.0851 0.1572 0.2954 0.9218 0.9789 Baseline+Geo(G+L)+Transformer(P+W) 10659 0.0829 0.1520 0.2933 0.9242 0.9796
[0088] Ablation studies can be further conducted to verify the usefulness of panel representation for, e.g., slice image partitioning. The Omnifusion implementation can be used as a comparison because it has a similar input format and can be trained via the same encoder-decoder CNN architecture as the PanelNet model. The comparison is shown in Table 5. As shown in Table 5, the panel representation with a pure CNN architecture outperforms the original Omnifusion, which demonstrates the superiority of the panel representation. The default transformer of Omnifusion can be replaced with a local-to-global transformer network. However, the local-to-global transformer network does not seem to bring significant performance improvements for slice images because the discontinuous slice tiles reduce the ability of the window tiles to aggregate local information in the vertical direction, which reduces the continuity of depth estimation for gravity-aligned objects and scenes. In contrast, the vertical continuity is maintained within the vertical panels of PanelNet. With the panel representation, the local-to-global transformer exerts the greatest information aggregation capability. Table 5
[0089] The effects of the panel size and stride of PanelNet on the model performance and speed are further evaluated, as shown in Table 6. For Table 6, the FPS is obtained by measuring the average inference time on a single NVIDIA GTX 1080Ti GPU. It is observed that for PanelNet models with the same panel spacing (i.e., sliding window width), a smaller stride improves performance. For the same stride, PanelNet models with larger panels have better performance. In theory, a smaller stride improves performance because more overlapping areas of consecutive panels maintain horizontal consistency. Larger panels also lead to better performance because larger panels provide larger FoVs that contain more geometric context within the panels. However, it is observed that constantly increasing the spacing may have a negative impact on performance. Specifically, larger panels bring higher computational complexity, which forces the stride to increase to reduce computational overhead. This causes the performance gain brought by the larger FoV to be eliminated by the consistency loss caused by less overlap. For optimal performance, the spacing can be set to, for example, 128, and the stride can be set to, for example, 32. Table 6 I S #panel FPS MRE RMSE <![CDATA[δ 1 ]]> 64 16 128 6.4 0.0866 0.3040 0.9181 64 32 64 12.4 0.0909 0.3207 0.9102 64 64 32 24.4 0.0952 0.3319 0.9041 128 32 32 6.9 0.0829 0.2933 0.9242 128 64 16 13.5 0.0892 0.3109 0.9172 128 128 8 25.7 0.0920 0.3181 0.9103 256 64 16 7.5 0.0894 0.3047 0.9132 256 128 8 13.9 0.0908 0.3069 0.9182 256 256 4 26.4 0.0986 0.3248 0.8991
[0090] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.10 A computer system (1000) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0091] Computer software may be encoded using any suitable machine code or computer language that may be assembled, compiled, linked, or similar mechanisms to create code comprising instructions that may be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., either directly or through interpretation, microcode execution, etc.
[0092] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0093] Fig.10 The components shown for the computer system (1000) are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement on any one component or combination of components illustrated in the exemplary embodiment of the computer system (1000).
[0094] The computer system (1000) may include certain human interface input devices. The input human interface devices may include one or more of the following (only one of each is depicted): keyboard (1001), mouse (1002), trackpad (1003), touch screen (1010), data gloves (not shown), joystick (1005), microphone (1006), scanner (1007), camera (1008).
[0095] The computer system (1000) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback of a touch screen (1010), a data glove (not shown), or a joystick (1005), but there may also be tactile feedback devices that are not used as input devices), audio output devices (such as: speakers (1009), headphones (not depicted)), visual output devices (such as screens (1010), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which are capable of outputting two-dimensional visual output or more than three-dimensional output through means such as stereo output; virtual reality glasses (not depicted), holographic displays, and smoke canisters (not depicted)), and printers (not depicted).
[0096] The computer system (1000) may also include human-accessible storage devices and their associated media, such as optical media including media (1021) such as CD / DVD ROM / RW (1020) with CD / DVD, thumb drives (1022), removable hard drives or solid-state drives (1023), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), and the like.
[0097] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other volatile signals.
[0098] The computer system (1000) may also include an interface (1054) to one or more communication networks (1055). The network may be, for example, wireless, wired, optical. The network may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicular and industrial networks including CAN buses, etc.
[0099] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1040) of the computer system (1000).
[0100] The core (1040) may include one or more central processing units (CPUs) (1041), graphics processing units (GPUs) (1042), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (1043), hardware accelerators (1044) for certain tasks, graphics adapters (1050), etc. These devices, along with read-only memory (ROM) (1045), random access memory (1046), internal mass storage devices (1047) such as internal non-user accessible hard drives, SSDs, etc., may be connected via a system bus (1048). In some computer systems, the system bus (1048) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (1048), or via a peripheral bus (1049). In an example, a screen (1010) may be connected to a graphics adapter (1050). Architectures for peripheral buses include PCI, USB, etc.
[0101] The computer readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of a type well known and available to those skilled in the art of computer software.
[0102] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various substitute equivalents that fall within the scope of the present disclosure. It should therefore be appreciated that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.
Claims
1. A method for processing a panoramic image data set by a computing circuit, characterized in that: include: generating a plurality of data panels from the panoramic image dataset; executing a first neural network to process the plurality of data panels to generate a set of embeddings representing geometric features of the plurality of data panels; executing a second neural network to process the plurality of data panels and the set of embeddings to generate a plurality of mapping panels; as well as The plurality of mapping panels are fused into a mapping dataset of the panoramic image dataset.
2. The method according to claim 1, characterized in that: The mapping dataset includes one of a depth map, a layout map, or a semantic map corresponding to the panoramic image dataset.
3. The method according to claim 1, characterized in that: The panoramic image data set comprises a two-dimensional data array; and Each of the plurality of data panels includes a sub-array of the data array, the sub-array being in its entirety in a first dimension of the two dimensions and in a section in a second dimension of the two dimensions.
4. The method according to claim 3, characterized in that The first dimension represents a gravity direction of the panoramic image dataset, and the second dimension represents a horizontal direction of the panoramic image dataset.
5. The method according to claim 3, characterized in that: in, The plurality of data panels are continuously generated from the panoramic image data set using a window having a length of the entirety of the first dimension in the first dimension and a predetermined width in the second dimension, the window sliding by a predetermined step along the second dimension.
6. The method according to claim 5, characterized in that The window slides continuously from one edge of the panoramic image dataset in the second dimension to another edge of the panoramic image dataset in the second dimension.
7. The method according to any one of claims 3 to 6, characterized in that The first neural network is configured to encode local and global geometric features of the plurality of data panels to reduce effects of geometric distortion in the panoramic image dataset and enhance preservation of geometric continuity across the plurality of mapped panels.
8. The method according to any one of claims 3 to 6, characterized in that The first neural network includes a multi-layer perceptron (MLP) network LP, which is used to process the geometric information extracted from the multiple data panels to generate the embedding set including the global geometric feature set and the local geometric feature set of the multiple data panels.
9. The method according to claim 8, characterized in that The second neural network is configured to process the plurality of data panels and reduce geometric distortion in the panoramic image dataset based on the set of embeddings.
10. The method according to claim 8, characterized in that The second neural network comprises: Downsampling network; a converter network; and Upsampling network.
11. The method according to claim 10, characterized in that: The downsampling network is configured to process the plurality of data panels and the set of embeddings to generate a series of downsampled features of decreasing resolution; The transformer network is configured to process the lowest resolution downsampled features to generate transformed low resolution features; and The upsampling network is configured to process the transformed low-resolution features and the series of downsampled features to generate the plurality of mapping panels.
12. The method according to claim 10, characterized in that The transformer network includes a feature processor.
13. The method according to claim 12, characterized in that The feature processor is configured to increase continuity of the geometric feature.
14. The method according to claim 12, characterized in that The feature processor is configured to aggregate local information of each of the plurality of data planes to capture per-panel context.
15. A device for processing a panoramic image data set, characterized in that: The apparatus comprises a memory for storing computer instructions and at least one processor for executing the computer instructions to: generating a plurality of data panels from the panoramic image dataset; executing a first neural network to process the plurality of data panels to generate a set of embeddings representing geometric features of the plurality of data panels; executing a second neural network to process the plurality of data panels and the set of embeddings to generate a plurality of mapping panels; as well as The plurality of mapping panels are fused into a mapping dataset of the panoramic image dataset.
16. The device according to claim 15, characterized in that: The panoramic image data set comprises a two-dimensional data array; and Each of the plurality of data panels includes a sub-array of the data array, the sub-array being in its entirety in a first dimension of the two dimensions and in a section in a second dimension of the two dimensions.
17. The device according to claim 16, characterized in that: The plurality of data panels are continuously generated from the panoramic image dataset using a window, the window having a length of the entirety of the first dimension in the first dimension and a predetermined width in the second dimension, the window sliding along the second dimension by a predetermined step and sliding from one edge of the panoramic image dataset in the second dimension to another edge of the panoramic image dataset in the second dimension; and The first neural network is configured to encode local and global geometric features of the plurality of data panels to reduce effects of geometric distortion in the panoramic image dataset and enhance preservation of geometric continuity across the plurality of mapped panels.
18. The device according to any one of claims 16 to 17, characterized in that: The first neural network comprises a multi-layer perceptron (MLP) network LP for processing geometric information extracted from the plurality of data panels to generate the embedding set comprising a global geometric feature set and a local geometric feature set of the plurality of data panels; The second neural network includes a downsampling network, a transformer network and an upsampling network; The downsampling network is configured to process the plurality of data panels and the set of embeddings to generate a series of downsampled features of decreasing resolution; The transformer network is configured to process the lowest resolution downsampled features to generate transformed low resolution features; and The upsampling network is configured to process the transformed low-resolution features and the series of downsampled features to generate the plurality of mapping panels.
19. The device according to claim 18, characterized in that: The transformer network includes a feature processor; and The feature processor is configured to increase continuity of the geometric features and aggregate local information within each of the plurality of data panels to capture panel-by-panel context.
20. A non-volatile computer-readable medium for storing instructions, characterized in that: When executed by at least one processor, the instructions are configured to cause the processor to process the panoramic image dataset by: generating a plurality of data panels from the panoramic image dataset; executing a first neural network to process the plurality of data panels to generate a set of embeddings representing geometric features of the plurality of data panels; executing a second neural network to process the plurality of data panels and the set of embeddings to generate a plurality of mapping panels; as well as The plurality of mapping panels are fused into a mapping dataset of the panoramic image dataset.