Method, apparatus, and computer program for estimating a Manhattan layout associated with a scene

By combining geometric and semantic information from multiple two-dimensional images, the method addresses the challenge of accurately reconstructing Manhattan layouts from panoramic images, facilitating precise three-dimensional scene reconstruction for virtual and augmented reality applications.

JP7706566B2Active Publication Date: 2025-07-11TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023560170
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-11-04
Filing Date
2022-11-08
Publication Date
2025-07-11
Estimated Expiration
2042-11-08

AI Technical Summary

Technical Problem

Existing methods for reconstructing three-dimensional real-world scenes from two-dimensional images face challenges in accurately estimating Manhattan layouts due to incomplete geometric and semantic information, particularly when multiple panoramic images are used, which can be obstructed by objects and require sophisticated data processing.

Method used

A method that combines geometric and semantic information from multiple two-dimensional images to estimate a Manhattan layout, involving line detection, alignment, and polygon generation, followed by triangulation or voxelization to create a three-dimensional shape, using techniques like ray-casting and deep learning for semantic segmentation.

Benefits of technology

This approach enhances the accuracy of estimating Manhattan layouts by integrating geometric and semantic data, enabling precise reconstruction of three-dimensional scenes from multiple panoramic images, suitable for applications in virtual and augmented reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007706566000016
    Figure 0007706566000016
  • Figure 0007706566000017
    Figure 0007706566000017
  • Figure 0007706566000018
    Figure 0007706566000018
Patent Text Reader

Abstract

A plurality of two-dimensional (2D) images of a scene are received. Geometry information and semantic information for each of the plurality of 2D images are determined. The geometry information is indicative of detected lines and reference directions in each of the 2D images. The semantic information includes classification information of pixels in each of the 2D images. A layout estimation associated with each of the 2D images of the scene is determined based on the geometry information and the semantic information of each of the 2D images. A joint layout estimation associated with the scene is determined based on the plurality of determined layout estimates associated with the plurality of 2D images of the scene. A Manhattan layout associated with the scene is generated based on the joint layout estimation. The Manhattan layout includes at least a three-dimensional (3D) shape of the scene including mutually orthogonal walls.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority to U.S. Patent Application No. 17 / 981,156, filed Nov. 4, 2022, entitled “Manhattan Layout Estimation Using Geometry and Semantic Information,” which in turn claims the benefit of priority to U.S. Provisional Application No. 63 / 306,001, filed Feb. 2, 2022, entitled “Method for Manhattan Layout Estimation from Multiple Panoramic Images Using Geometry and Semantic Segmentation Information.” The disclosure of the prior applications is hereby incorporated by reference in its entirety.

[0002] This disclosure generally describes embodiments related to image coding.

Background Art

[0003] The description of the background art provided herein is for the purpose of generally presenting the context of the disclosure. Aspects of the description that are not within the scope of this background art section and that may not be considered prior art at the time of filing are not admitted to be prior art to the disclosure, either expressly or by implication.

[0004] Various techniques have been developed for capturing and representing the world, such as objects in a three - dimensional (3D) space or shape, the environment of the world, etc. A 3D representation of the world may enable more immersive interactions and more immersive communications. In some examples, 3D shapes are widely used to represent such immersive content.

[0005] Conventionally, a method of reconstructing a three-dimensional (3D) real-world scene from a single two-dimensional (2D) image by identifying junction points that satisfy the geometric constraints of the scene based on intersecting lines, vanishing points, and vanishing lines that are orthogonal to each other. By sampling the 2D image according to the junction points, possible layouts of the scene are generated. Then, an energy function is maximized to select the optimal layout from the possible layouts. The energy function is known to evaluate the possible layouts using a conditional random field (CRF) model (see Patent Document 1).

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Summary of the Invention

Means for Solving the Problems

[0007] Aspects of the present disclosure provide methods and apparatuses for image coding (e.g., compression and decompression). In some examples, an apparatus for image coding includes a processing circuit.

[0008] According to one aspect of the present disclosure, a method for estimating a Manhattan layout associated with a scene is provided. In this method, a plurality of two-dimensional (2D) images of the scene can be received. The geometric information and semantic information of each of the plurality of 2D images can be determined. The geometric information can indicate lines and reference directions detected in each 2D image. The semantic information can include pixel classification information in each 2D image. The layout estimation associated with each 2D image of the scene can be determined based on the geometric information and semantic information of each 2D image. The combined layout estimation associated with the scene can be determined based on a plurality of determined layout estimations associated with the plurality of 2D images of the scene. The Manhattan layout associated with the scene can be generated based on the combined layout estimation. The Manhattan layout can include at least a three-dimensional (3D) shape of the scene including walls orthogonal to each other.

[0009] To determine geometric information and semantic information, the first geometric information of the first 2D image among a plurality of 2D images can be extracted. The first geometric information can include at least one of a detected line, a reference direction of the first 2D image, a ratio of a first distance from the ceiling to the ground to a second distance from the camera to the ground, or a relative pose (e.g., an angle or a distance) between the first 2D image and a second 2D image among the plurality of 2D images. The pixels of the first 2D image can be labeled to generate first semantic information, and the first semantic information can indicate first structural information of the pixels in the first 2D image.

[0010] To determine a layout estimation associated with each 2D image of the scene, a first layout estimation of a plurality of layout estimations associated with the scene can be determined based on the first geometric information and the first semantic information of the first 2D image. To determine the first layout estimation, it can be determined whether each of the detected lines is a boundary line corresponding to a wall boundary in the scene. The boundary lines of the detected lines can be aligned with the reference direction of the first 2D image. A first polygon indicating the first layout estimation can be generated based on the boundary lines aligned using one of 2D polygon noise removal and staircase removal.

[0011] To generate the first polygon, a plurality of incomplete boundary lines of the boundary lines can be completed based on one of (i) estimating a plurality of incomplete boundary lines based on a combination of a ceiling boundary line and a floor boundary line of the boundary lines, and (ii) connecting a pair of incomplete boundary lines of the plurality of incomplete boundary lines. A pair of incomplete boundary lines can be connected based on one of (i) adding a perpendicular line to the pair of incomplete boundary lines in response to the pair of incomplete boundary lines being parallel, and (ii) extending at least one of the pair of incomplete boundary lines such that an intersection of the pair of incomplete boundary lines is located on the extended pair of incomplete boundary lines.

[0012] To determine a combined layout estimate associated with a scene, a reference polygon can be determined by combining a plurality of polygons via a polygon summation algorithm. Each of the plurality of polygons can correspond to a respective layout estimate of a plurality of determined layout estimates. A shrunk polygon can be determined based on the reference polygon. The shrunk polygon can include updated edges updated from the edges of the reference polygon. A final polygon can be determined based on the shrunk polygon using one of 2D polygon noise removal and staircase removal. The final polygon can correspond to the combined layout estimate associated with the scene.

[0013] To determine the shrunk polygon, a plurality of candidate edges can be determined from a plurality of polygons of the edges of the reference polygon. Each of the plurality of candidate edges can correspond to a respective edge of the reference polygon. The updated edges of the shrunk polygon can be generated by replacing one or more edges of the reference polygon with one or more candidate edges that are closer to the original view positions in a plurality of images than the corresponding one or more edges of the reference polygon in response to the one or more candidate edges.

[0014] In some embodiments, each of the plurality of candidate edges can be parallel to the corresponding edge of the reference polygon. The projected overlap between each candidate edge and the corresponding edge of the reference polygon can be made larger than a threshold.

[0015] To determine a combined layout estimate associated with a scene, an edge set including the edges of the final polygon can be determined. Based on the edge set, a plurality of edge groups can be generated. A plurality of inner edges of the final polygon can be generated. The plurality of inner edges can be indicated by a plurality of average edges of one or more edge groups of the edge set. Each of one or more of the plurality of edge groups can include a respective number of edges greater than a target value. Each of the plurality of average edges can be obtained by averaging each one edge of one or more of the edge groups.

[0016] In some embodiments, the plurality of edge groups can include a first edge group. The first edge group can further include a first edge and a second edge. The first edge and the second edge can be parallel. The distance between the first edge and the second edge can be less than a first threshold value. The projected overlapping region between the first edge and the second edge can be made greater than a second threshold value.

[0017] To generate a Manhattan layout associated with a scene, the Manhattan layout associated with the scene can be generated based on a triangular mesh triangulated from the combined layout estimate, a quadrilateral mesh quadrilateralized from the combined layout estimate, a sampling point sampled from one of the triangular mesh and the quadrilateral mesh, or a discrete grid generated from one of the triangular mesh and the quadrilateral mesh via voxelization.

[0018] In some embodiments, the Manhattan layout associated with a scene can be generated based on a triangular mesh that is triangulated from a combined layout estimate. Thus, in order to generate the Manhattan layout associated with a scene, the ceiling and floor surfaces in the scene can be generated by triangulating the combined layout estimate. The wall surfaces in the scene can be generated by triangulating a rectangle that encloses the ceiling boundary line and the floor boundary line in the scene. The texture of the Manhattan layout associated with the scene can be generated via a ray-casting based process.

[0019] According to another aspect of the present disclosure, an apparatus is provided. The apparatus includes a processing circuit. The processing circuit can be configured to execute any of the methods for estimating a Manhattan layout associated with a scene.

[0020] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to execute any of the methods for estimating a Manhattan layout associated with a scene.

[0021] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

Brief Description of the Drawings

[0022]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 9C

Figure 10A

Figure 10B

Figure 10C

Figure 11

Figure 12

Figure 13

Figure 14

Modes for Carrying Out the Invention

[0023] Aspects of the present disclosure include technologies in the field of three-dimensional (3D) media processing.

[0024] Technological developments in 3D media processing, such as advancements in 3D capture, 3D modeling, 3D rendering, etc., have facilitated the ubiquitous presence of 3D media content across some platforms and devices. In one example, the first steps of a baby can be captured on one continent, and media technology can enable grandparents on another continent to view (and in some cases interact with) the baby and enjoy an immersive experience with the baby. According to one aspect of the present disclosure, to improve the immersive experience, 3D models are becoming increasingly sophisticated, and the creation and consumption of 3D models occupy a significant amount of data resources such as data storage and data transmission resources.

[0025] In some embodiments, point clouds and meshes can be used as 3D models to represent immersive content.

[0026] A point cloud can generally refer to a set of points in 3D space, and each point has associated attributes such as color, material properties, texture information, intensity attributes, reflectivity attributes, motion-related attributes, modality attributes, and various other attributes. The point cloud can be used to reconstruct an object or scene as a synthesis of such points.

[0027] A mesh of an object (also referred to as a mesh model) can include polygons that describe the surface of the object. Each polygon can be defined by the vertices of the polygon in 3D space and information on how the vertices are connected to the polygon. The information on how the vertices are connected is referred to as connectivity information. In some examples, the mesh can also include attributes such as color, normal, etc., associated with the vertices.

[0028] In some embodiments, some coding tools for point cloud compression (PCC) can be used for mesh compression. For example, the mesh can be remeshed to generate a new mesh, and the connectivity information of the vertices of the new mesh can be inferred (or predefined). The vertices of the new mesh and the attributes associated with the vertices of the new mesh can be regarded as points in a point cloud and compressed using a PCC codec.

[0029] Point clouds can be used to reconstruct an object or scene as a synthesis of such points. Points can be captured using multiple cameras, depth sensors, or lidars in various settings and can be composed of thousands to up to billions of points to realistically represent the reconstructed scene or object. Patches can generally refer to a continuous subset of the surface described by the point cloud. In one example, a patch includes points having surface normal vectors whose deviation from each other is less than a threshold amount.

[0030] PCC can be performed according to various methods, such as a geometry-based method called G-PCC and a video coding-based method called V-PCC. According to some aspects of the present disclosure, G-PCC is a purely geometry-based approach that directly encodes 3D geometry and shares little with video coding, and V-PCC is highly based on video coding. For example, V-PCC can map the points of a 3D cloud to the pixels of a 2D grid (image). The V-PCC method can utilize a general-purpose video codec for point cloud compression. The PCC codec (encoder / decoder) in the present disclosure can be a G-PCC codec (encoder / decoder) or a V-PCC codec.

[0031] According to one aspect of the present disclosure, the V-PCC method can compress the geometry, occupancy, and texture of a point cloud as three separate video sequences using a video codec. The additional metadata required to interpret the three video sequences is compressed separately. A small portion of the entire bitstream is metadata, which can be efficiently encoded / decoded using, for example, a software implementation form. Most of the information is processed by the video codec.

[0032] FIG. 1 shows a block diagram of a communication system (100) in some examples. The communication system (100) includes a plurality of terminal devices that can communicate with each other via, for example, a network (150). For example, the communication system (100) includes a pair of terminal devices (110) and (120) interconnected via a network (150). In the example of FIG. 1, the first pair of terminal devices (110) and (120) can perform unidirectional transmission of point cloud data. For example, the terminal device (110) can compress a point cloud (e.g., points representing a structure) captured by a sensor (105) connected to the terminal device (110). The compressed point cloud can be transmitted to another terminal device (120) via the network (150) in the form of, for example, a bitstream. The terminal device (120) can receive the compressed point cloud from the network (150), decompress the bitstream to reconstruct the point cloud, and appropriately display the reconstructed point cloud. Unidirectional data transmission can be common in media serving applications and the like.

[0033] In the example of FIG. 1, the terminal devices (110) and (120) can be shown as a server and a personal computer, but the principles of the present disclosure may not be so limited. Embodiments of the present disclosure find applications in laptop computers, tablet computers, smartphones, game terminals, media players, and / or dedicated three-dimensional (3D) devices. The network (150) represents any number of networks that transmit compressed point clouds between the terminal devices (110) and (120). The network (150) can include, for example, wireline (wired) and / or wireless communication networks. The network (150) can exchange data over circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, the Internet, and the like.

[0034] FIG. 2 shows a block diagram of a streaming system (200) in some examples. The streaming system (200) is an application for using point clouds. The disclosed subject matter may be equally applicable to other point cloud-enabled applications such as 3D telepresence applications, virtual reality applications, and the like.

[0035] The streaming system (200) may include a capture subsystem (213). The capture subsystem (213) may include a point cloud source (201), such as a light detection and ranging (LIDAR) system, a 3D camera, a 3D scanner, a graphic generation component that generates non-compressed point clouds in software, and the like that generate, for example, non-compressed point clouds (202). In one example, the point cloud (202) includes points captured by a 3D camera. The point cloud (202) is shown in bold lines to emphasize the high data volume as compared to the compressed point cloud (204) (bit stream of the compressed point cloud). The compressed point cloud (204) can be generated by an electronic device (220) that includes an encoder (203) coupled to the point cloud source (201). The encoder (203) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The compressed point cloud (204) (or the bit stream of the compressed point cloud (204)) is shown in thin lines to emphasize that it has a lower data volume as compared to the stream of the point cloud (202) and can be stored in the streaming server (205) for future use. One or more streaming client subsystems, such as the client subsystems (206) and (208) of FIG. 2, can access the streaming server (205) to retrieve copies (207) and (209) of the compressed point cloud (204). The client subsystem (206) may include, for example, a decoder (210) within an electronic device (230). The decoder (210) decodes an input copy (207) of the compressed point cloud and creates an output stream of a reconstructed point cloud (211) that can be rendered on a rendering device (212).

[0036] Note that the electronic devices (220) and (230) can include other components (not shown). For example, the electronic device (220) can include a decoder (not shown), and the electronic device (230) can also include an encoder (not shown).

[0037] In some streaming systems, compression point clouds (204), (207), and (209) (e.g., bitstreams of compression point clouds) can be compressed according to some standards. In some examples, video coding standards are used in the compression of point clouds. Examples of those standards include High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), and the like.

[0038] Figure 3 shows a block diagram of a V-PCC encoder (300) for encoding a point cloud frame according to some embodiments. In some embodiments, the V-PCC encoder (300) can be used in a communication system (100) and a streaming system (200). For example, the encoder (203) can be configured and operate in the same manner as the V-PCC encoder (300).

[0039] The V-PCC encoder (300) receives a point cloud frame as an uncompressed input and generates a bitstream corresponding to the compressed point cloud frame. In some embodiments, the V-PCC encoder (300) can receive a point cloud frame from a point cloud source such as the point cloud source (201).

[0040] In the example of Figure 3, the V-PCC encoder (300) includes a patch generation module (306), a patch packing module (308), a geometry image generation module (310), a texture image generation module (312), a patch information module (304), an occupancy map module (314), a smoothing module (336), image padding modules (316) and (318), a group extension module (320), video compression modules (322), (323), and (332), an auxiliary patch information compression module (338), an entropy compression module (334), and a multiplexer (324).

[0041] According to one aspect of the present disclosure, a V-PCC encoder (300) converts a 3D point cloud frame into an image-based representation, along with some metadata (e.g., occupancy map and patch information) used to convert and revert the compressed point cloud to the expanded point cloud. In some examples, the V-PCC encoder (300) can convert a 3D point cloud frame into a geometry image, a texture image, and an occupancy map, and then use video coding techniques to encode the geometry image, the texture image, and the occupancy map into a bitstream. Generally, a geometry image is a 2D image having pixels filled with geometry values associated with points projected onto the pixels, and the pixels filled with geometry values may be referred to as geometry samples. A texture image is a 2D image having pixels filled with texture values associated with points projected onto the pixels, and the pixels filled with texture values may be referred to as texture samples. An occupancy map is a 2D image having pixels filled with values indicating whether a patch is occupied or not.

[0042] The patch generation module (306) segments the point cloud into a set of patches (e.g., the patches are defined as continuous subsets of the surface described by the point cloud), which may or may not overlap such that each patch can be described by a depth field with respect to a plane in 2D space. In some embodiments, the patch generation module (306) aims to decompose the point cloud into the minimum number of patches with smooth boundaries and minimize the reconstruction error.

[0043] In some examples, the patch information module (304) can collect patch information indicating the size and shape of the patches. In some examples, the patch information can be packed into an image frame and then encoded by an auxiliary patch information compression module (338) to generate compressed auxiliary patch information.

[0044] In some examples, the patch packing module (308) maps the extracted patches onto a two-dimensional (2D) grid while minimizing unused space and ensuring that a unique patch is associated with each M×M (e.g., 16×16) block of the grid. Efficient patch packing can directly affect compression efficiency either by minimizing unused space or by ensuring temporal consistency.

[0045] The geometry image generation module (310) can generate a 2D geometry image associated with the geometry of the point cloud at a given patch location. The texture image generation module (312) can generate a 2D texture image associated with the texture of the point cloud at a given patch location. The geometry image generation module (310) and the texture image generation module (312) utilize the 3D-2D mapping calculated during the packing process to store the geometry and texture of the point cloud as images. To better handle the case where multiple points are projected onto the same sample, each patch is projected onto two images referred to as layers. In one example, the geometry image is represented by a W×H monochrome frame in the YUV420-8-bit format. To generate the texture image, the texture generation procedure utilizes the reconstructed / smoothed geometry to calculate the color associated with the resampled points.

[0046] The occupancy map module (314) can generate an occupancy map that describes the padding information in each unit. For example, the occupancy image includes a binary map indicating whether each cell of the grid belongs to free space or to the point cloud. In one example, the occupancy map uses binary information that describes for each pixel whether the pixel is padded or not. In another example, the occupancy map uses binary information that describes for each block of pixels whether the block of pixels is padded or not.

[0047] The occupancy map generated by the occupancy map module (314) can be compressed using reversible coding or irreversible coding. When reversible coding is used, the entropy compression module (334) is used to compress the occupancy map. When irreversible coding is used, the video compression module (332) is used to compress the occupancy map.

[0048] Note that the patch packing module (308) can leave some empty space between the 2D patches packed within the image frame. The image padding modules (316) and (318) can fill (referred to as padding) the empty space to generate an image frame suitable for 2D video and image codecs. Image padding, also referred to as background filling, can fill the unused space with redundant information. In some examples, good background filling minimally increases the bitrate but does not introduce significant coding distortion around the patch boundaries.

[0049] The video compression modules (322), (323), and (332) can encode 2D images such as the padded geometry image, the padded texture image, and the occupancy map based on suitable video coding standards such as HEVC, VVC. In one example, the video compression modules (322), (323), and (332) are individual components that operate separately. Note that the video compression modules (322), (323), and (332) can be implemented as a single component in another example.

[0050] In some examples, the smoothing module (336) is configured to generate a smoothed image of the reconstructed geometry image. The smoothed image can be provided to the texture image generation module (312). The texture image generation module (312) can then adjust the generation of the texture image based on the reconstructed geometry image. For example, when the patch shape (e.g., geometry) is slightly distorted during encoding and decoding, distortion can be taken into account when generating the texture image to correct the distortion in the patch shape.

[0051] In some embodiments, the group expansion module (320) is configured to pad the pixels around the object boundary with redundant low-frequency content to improve the coding gain and visual quality of the reconstructed point cloud.

[0052] The multiplexer (324) can multiplex the compressed geometry image, the compressed texture image, the compressed occupancy map, and the compressed auxiliary patch information into a compressed bitstream.

[0053] FIG. 4 shows a block diagram of a V-PCC decoder (400) for decoding a compressed bitstream corresponding to a point cloud frame in some examples. In some examples, the V-PCC decoder (400) can be used in a communication system (100) and a streaming system (200). For example, the decoder (210) can be configured to operate in the same manner as the V-PCC decoder (400). The V-PCC decoder (400) receives the compressed bitstream and generates a reconstructed point cloud based on the compressed bitstream.

[0054] In the example of FIG. 4, the V-PCC decoder (400) includes a demultiplexer (432), video decompression modules (434) and (436), an occupancy map decompression module (438), an auxiliary patch information decompression module (442), a geometry reconstruction module (444), a smoothing module (446), a texture reconstruction module (448), and a color smoothing module (452).

[0055] The demultiplexer (432) can receive the compressed bitstream and separate it into a compressed texture image, a compressed geometry image, a compressed occupancy map, and compressed auxiliary patch information.

[0056] The video decompression modules (434) and (436) can decode the compressed image according to an appropriate standard (e.g., HEVC, VVC, etc.) and output the decompressed image. For example, the video decompression module (434) decodes the compressed texture image and outputs the decompressed texture image, and the video decompression module (436) decodes the compressed geometry image and outputs the decompressed geometry image.

[0057] The occupancy map decompression module (438) can decode the compressed occupancy map according to an appropriate standard (e.g., HEVC, VVC, etc.) and output the decompressed occupancy map.

[0058] The auxiliary patch information decompression module (442) can decode the compressed auxiliary patch information according to an appropriate standard (e.g., HEVC, VVC, etc.) and output the decompressed auxiliary patch information.

[0059] The geometry reconstruction module (444) can receive the decompressed geometry image and generate a reconstructed point cloud geometry based on the decompressed occupancy map and the decompressed auxiliary patch information.

[0060] The smoothing module (446) can smooth out the mismatches at the edges of the patches. The smoothing procedure aims to reduce potential discontinuities that may occur at the patch boundaries due to compression artifacts. In some embodiments, a smoothing filter can be applied to the pixels located on the patch boundaries to reduce the distortion that can be caused by compression / decompression.

[0061] The texture reconstruction module (448) can determine the texture information of the points within the point cloud based on the unfolded texture image and the smoothed geometry.

[0062] The color smoothing module (452) can smooth out the coloring mismatches. Non-neighboring patches in 3D space are often packed adjacent to each other in 2D video. In some examples, the pixel values from non-neighboring patches can be mixed by a block-based video codec. The goal of color smoothing is to reduce the visible artifacts that appear at the patch boundaries.

[0063] FIG. 5 shows a block diagram of a video decoder (510) in some examples. The video decoder (510) can be used in the V-PCC decoder (400). For example, the video unfolding modules (434) and (436), and the occupancy map unfolding module (438) can be configured similarly to the video decoder (510).

[0064] The video decoder (510) may include a parser (520) for reconstructing symbols (521) from a compressed image such as a coded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (510). The parser (520) can perform syntax analysis / entropy decoding on the received coded video sequence. The coding of the coded video sequence can follow a video coding technology or standard and can follow various principles including variable length coding with or without context dependence, Huffman coding, arithmetic coding, etc. The parser (520) can extract a set of subgroup parameters regarding at least one of the subgroups of pixels within the video decoder from the coded video sequence based on at least one parameter corresponding to a group. The subgroups can include picture groups (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (520) may also extract from coded video sequence information such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0065] The parser (520) can perform an entropy decoding / syntax analysis operation on the video sequence received from the buffer memory to create symbols (521).

[0066] The restoration of the symbols (521) can include a plurality of different units depending on the type of the coded video picture or a part thereof (such as between pictures and within pictures, between blocks and within blocks), and other factors. How each unit is involved can be controlled by subgroup control information parsed from the coded video sequence by the parser (520). Such a flow of subgroup control information between the parser (520) and the following plurality of units is not shown for clarity.

[0067] In addition to the function blocks already described, the video decoder (510) can conceptually be subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units interact closely with each other and can at least partially be integrated with each other. However, for the purpose of explaining the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0068] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives from the parser (520) the quantized transform coefficients and control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. as symbols (s) (521). The scaler / inverse transform unit (551) can output a block containing sample values that can be input to the aggregator (555).

[0069] In some cases, the output samples of the scaler / inverse transform unit (551) may be related to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed part of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses the surrounding already reconstructed information fetched from the current picture buffer (558) to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, the partially reconstructed current picture and / or the fully reconstructed current picture. The aggregator (555) may, in some cases, add the prediction information generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a sample-by-sample basis.

[0070] In other cases, the output samples of the scaler / inverse transform unit (551) can be associated with inter-coded and potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (553) can access the reference picture memory (557) to fetch the samples used for prediction. After motion-compensating the samples fetched according to the symbols (521) related to the block, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (in this case, called the residual samples or residual signal) in order to generate output sample information. The address in the reference picture memory (557) from which the motion compensation prediction unit (553) fetches the prediction samples can be controlled, for example, by the motion vectors available to the motion compensation prediction unit (553) in the form of symbols (521) that may have X, Y, and reference picture components. Motion compensation can also include interpolation of the fetched sample values, a motion vector prediction mechanism, etc. from the reference picture memory (557) when exact sub-sample motion vectors are used.

[0071] The output samples of the aggregator (555) can undergo various loop filtering techniques in the loop filter unit (556). The video compression technology is controlled by the parameters included in the coded video sequence (also called the coded video bitstream), and can include in-loop filter techniques made available to the loop filter unit (556) as symbols (521) from the parser (520), but can respond not only to the meta information obtained during the decoding of the previous part of the coded picture or coded video sequence (in decoding order), but also to the previously reconstructed and loop-filtered sample values.

[0072] The output of the loop filter unit (556) can be output to the rendering device and can be a sample stream that can be stored in the reference picture memory (557) for use in future inter-picture prediction.

[0073] Once fully reconstructed, a particular coded picture can be used as a reference picture for future prediction. For example, when the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be repositioned before starting the reconstruction of subsequent coded pictures.

[0074] The video decoder (510) may perform a decoding operation according to a predetermined video compression technique in a standard such as ITU-T Rec.H.265. The coded video sequence can conform to the syntax specified by the video compression technique or standard in the sense that it conforms to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile can select specific tools from all the tools available in the video compression technique or standard as the tools that are only available under that profile. Also, for compliance, it may be necessary that the complexity of the coded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (measured, for example, in megasamples per second), the maximum reference picture size, etc. The limitations set by the level may, in some cases, be further restricted by the specifications of the virtual reference decoder (HRD) and the metadata for HRD buffer management signaled within the coded video sequence.

[0075] FIG. 6 shows a block diagram of a video encoder (603) according to an embodiment of the present disclosure. The video encoder (603) can be used in a V-PCC encoder (300) that compresses point clouds. In one example, the video compression modules (322) and (323) and the video compression module (332) are configured in the same manner as the encoder (603).

[0076] The video encoder (603) may receive images such as a padded geometry image, a padded texture image, etc., and generate a compressed image.

[0077] According to one embodiment, a video encoder (603) can code and compress pictures (images) of a source video sequence into a coded video sequence (compressed images) in real time or under any other temporal constraints required by an application. Implementing an appropriate coding speed is one function of a controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units described below. The coupling is not shown for clarity. Parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, …), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions related to the video encoder (603) optimized for a particular system design.

[0078] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop may include a source coder (630) (which is responsible for creating symbols such as a symbol stream based on, for example, an input picture to be coded and reference picture(s)) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in the same way as a (remote) decoder also does (any compression between the symbols and the coded video bitstream is reversible in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream results in bit-exact results regardless of the position of the decoder (local or remote), the content of the reference picture memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as the reference picture samples that the decoder will "see" when using prediction during decoding. This basic principle of the synchronization of the reference pictures (and the resulting drift if the synchronization cannot be maintained, for example, due to channel errors) is also used in some related technologies.

[0079] The operation of the "local" decoder (633) may be the same as that of a "remote" decoder such as the video decoder (510), which has already been described in detail above in conjunction with FIG. 5. However, referring briefly to FIG. 5 as well, since the symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy coder (645) and the parser (520) may be reversible, the entropy decoding part of the video decoder (510) including the parser (520) may not be fully implemented in the local decoder (633).

[0080] During operation, in some examples, source coder (630) may perform motion compensated predictive coding that predictively codes an input picture by referring to one or more previously coded pictures from a video sequence designated as a "reference picture". In this way, coding engine (632) codes the difference between a pixel block of the input picture and a pixel block of a reference picture (s) that may be selected as a prediction reference (s) for the input picture.

[0081] Local video decoder (633) may decode the coded video data of a picture that may be designated as a reference picture, based on the symbols created by source coder (630). The operation of coding engine (632) may advantageously be a non - reversible process. When the coded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. Local video decoder (633) can replicate the decoding process that may be performed by the video decoder for a reference picture and store the reconstructed reference picture in reference picture cache (634). In this way, video encoder (603) may locally store a copy of the reconstructed reference picture having common content as the reconstructed reference picture that would be obtained by a remote video decoder (without transmission errors).

[0082] The predictor (635) can perform predictive search for the coding engine (632). That is, in the case of a new picture to be coded, the predictor (635) can find sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors and block shapes that can serve as appropriate prediction criteria for the new picture, and search the reference picture memory (634). The predictor (635) can operate on sample blocks for each pixel block to find appropriate prediction criteria. In some cases, the input picture can have prediction criteria derived from a plurality of reference pictures stored in the reference picture memory (634), as determined by the search results obtained by the predictor (635).

[0083] The controller (650) can manage the coding operations of the source coder (630), including, for example, setting parameters and subgroup parameters used for encoding video data.

[0084] The outputs of all the aforementioned functional units can undergo entropy coding within the entropy coder (645). The entropy coder (645) converts the symbols generated by various functional units into a coded video sequence by reversibly compressing the symbols according to techniques such as Huffman coding, variable-length coding, and arithmetic coding.

[0085] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) can assign a specific coded picture type to each coded picture, which may affect the coding technique applicable to each picture. For example, a picture may often be assigned as one of the following picture types.

[0086] An intracoded picture (I picture) can be a picture that can be coded and decoded without using any other picture in the sequence as a prediction source. Some video coders allow for different types of intracoded pictures, including for example, an Instantaneous Decoder Refresh (IDR) picture. Those skilled in the art are aware of those variants of I pictures, as well as their respective uses and characteristics.

[0087] A predicted picture (P picture) can be a picture that can be coded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.

[0088] A bi - directionally predicted picture (B picture) can be a picture that can be coded and decoded using intra prediction or inter prediction that uses at most two motion vectors and a reference index to predict the sample values of each block. Similarly, multiple predicted pictures can use three or more reference pictures and associated metadata for the restoration of a single block.

[0089] The source picture is typically spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and may be coded block by block. The blocks may be predictively coded with reference to other (already coded) blocks determined by the coding assignment applied to each picture of the block. For example, blocks of an I picture may be coded non - predictively or may be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra - prediction). Pixel blocks of a P picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0090] The video encoder (603) can perform coding operations in accordance with a predetermined video coding technology or standard such as ITU - T Rec.H.265. In its operation, the video encoder (603) can perform various compression operations including predictive coding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the coded video data can conform to the syntax specified by the video coding technology or standard being used.

[0091] Video can be in the form of multiple source pictures (images) in a time sequence. Intra-picture prediction (often abbreviated as intra prediction) utilizes the spatial correlation within a given picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. When a block within the current picture is similar to a reference block within a reference picture that has been previously coded and is still buffered within the video, the block within the current picture can be coded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension to identify the reference picture when multiple reference pictures are being used.

[0092] In some embodiments, dual prediction techniques can be used in inter-picture prediction. According to the dual prediction technique, two reference pictures such as a first reference picture and a second reference picture are used, both of which are prior to the decoding order of the current picture within the video (however, the display order may be past and future respectively). A block within the current picture can be coded by a first motion vector pointing to a first reference block within the first reference picture and by a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0093] Furthermore, to improve coding efficiency, merge mode techniques can be used in inter-picture prediction.

[0094] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in block units. For example, according to the HEVC standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU contains three coding tree blocks (CTBs), which are one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree divided into one or more coding units (CUs). For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of the CU, such as an inter prediction type or an intra prediction type. The CU is divided into one or more prediction units (PUs) according to the temporal predictability and / or spatial predictability. Generally, each PU contains one luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block contains a matrix of values (e.g., luma values) for pixels such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels.

[0095] FIG. 7 shows a block diagram of a G-PCC encoder (700) in some examples. The G-PCC encoder (700) can be configured to receive point cloud data and compress the point cloud data to generate a bitstream that conveys the compressed point cloud data. In one embodiment, the G-PCC encoder (700) may include a position quantization module (710), a duplicate point removal module (712), an octree encoding module (730), an attribute transfer module (720), a level of detail (LOD) generation module (740), an attribute prediction module (750), a residual quantization module (760), an arithmetic coding module (770), an inverse residual quantization module (780), an addition module (781), and a memory (790) for storing the reconstructed attribute values.

[0096] As shown, an input point cloud (701) can be received by the G-PCC encoder (700). The positions (e.g., 3D coordinates) of the point cloud (701) are provided to the quantization module (710). The quantization module (710) is configured to quantize the coordinates to generate quantized positions. The duplicate point removal module (712) is configured to receive the quantized positions and perform a filtering process to identify and remove duplicate points. The octree encoding module (730) is configured to receive the filtered positions from the duplicate point removal module (712) and perform an octree-based encoding process to generate a series of occupancy codes that describe a 3D grid of voxels. The occupancy codes are provided to the arithmetic coding module (770).

[0097] The attribute transfer module (720) is configured to receive the attributes of the input point cloud and execute an attribute transfer process to determine the attribute values of each voxel when multiple attribute values are associated with respective voxels. The attribute transfer process can be executed on the sorted points output from the octree encoding module (730). The attributes after the transfer operation are provided to the attribute prediction module (750). The LOD generation module (740) operates on the sorted points output from the octree encoding module (730) and is configured to reorganize the points into different LODs. The LOD information is supplied to the attribute prediction module (750).

[0098] The attribute prediction module (750) processes the points according to the LOD-based order indicated by the LOD information from the LOD generation module (740). The attribute prediction module (750) generates an attribute prediction for the current point based on the reconstructed attributes of the set of neighboring points of the current point stored in the memory (790). Thereafter, a prediction residual can be obtained based on the original attribute values received from the attribute transfer module (720) and the locally generated attribute prediction. When a candidate index is used in each attribute prediction process, the index corresponding to the selected prediction candidate can be provided to the arithmetic coding module (770).

[0099] The residual quantization module (760) is configured to receive the prediction residual from the attribute prediction module (750) and perform quantization to generate a quantized residual. The quantized residual is provided to the arithmetic coding module (770).

[0100] The inverse residual quantization module (780) is configured to receive the quantized residual from the residual quantization module (760) and generate a reconstructed prediction residual by performing the inverse of the quantization operation performed in the residual quantization module (760). The addition module (781) is configured to receive the reconstructed prediction residual from the inverse residual quantization module (780) and the respective attribute predictions from the attribute prediction module (750). By combining the reconstructed prediction residual and the attribute prediction, a reconstructed attribute value is generated and stored in the memory (790).

[0101] The arithmetic coding module (770) is configured to receive an occupancy code, a candidate index (if used), a quantized residual (if generated), and other information, and perform entropy encoding to further compress the received values or information. As a result, a compressed bitstream (702) carrying the compressed information can be generated. The bitstream (702) may be transmitted to a decoder that decodes the compressed bitstream, provided in another way, or stored in a storage device.

[0102] FIG. 8 shows a block diagram of a G-PCC decoder (800) according to an embodiment. The G-PCC decoder (800) can be configured to receive a compressed bitstream, perform point cloud data expansion to expand the bitstream, and generate decoded point cloud data. In one embodiment, the G-PCC decoder (800) may include an arithmetic decoding module (810), an inverse residual quantization module (820), an octree decoding module (830), a LOD generation module (840), an attribute prediction module (850), and a memory (860) for storing the reconstructed attribute values.

[0103] As shown, the compressed bitstream (801) can be received in the arithmetic decoding module (810). The arithmetic decoding module (810) is configured to decode the compressed bitstream (801) to obtain the quantized residuals (if generated) and the occupancy codes of the point cloud. The octree decoding module (830) is configured to determine the reconstructed positions of the points within the point cloud according to the occupancy codes. The LOD generation module (840) is configured to reorganize the points into different LODs based on the reconstructed positions and determine the LOD-based order. The inverse residual quantization module (820) is configured to generate the reconstructed residuals based on the quantized residuals received from the arithmetic decoding module (810).

[0104] The attribute prediction module (850) is configured to execute an attribute prediction process to determine the attribute prediction of the points according to the LOD-based order. For example, the attribute prediction of the current point can be determined based on the reconstructed attribute values of the neighboring points of the current point stored in the memory (860). In some examples, the attribute prediction can be combined with the respective reconstructed residuals to generate the reconstructed attributes for the current point.

[0105] The sequence of the reconstructed attributes generated from the attribute prediction module (850), together with the reconstructed positions generated from the octree decoding module (830), corresponds to the decoded point cloud (802) output from the G-PCC decoder (800) in one example. In addition, the reconstructed attributes are also stored in the memory (860) and can then be used to derive the attribute prediction of subsequent points.

[0106] In various embodiments, the encoder (300), decoder (400), encoder (700), and / or decoder (800) can be implemented using hardware, software, or a combination thereof. For example, the encoder (300), decoder (400), encoder (700), and / or decoder (800) can be implemented using a processing circuit such as one or more integrated circuits (ICs) that operate with or without software, such as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs). In another example, the encoder (300), decoder (400), encoder (700), and / or decoder (800) can be implemented as software or firmware that includes instructions stored in a non-volatile (or non-transitory) computer-readable storage medium. When the instructions are executed by a processing circuit such as one or more processors, the processing circuit is caused to perform the functions of the encoder (300), decoder (400), encoder (700), and / or decoder (800).

[0107] Note that the attribute prediction modules (750) and (850) configured to implement the attribute prediction techniques disclosed herein can be included in other decoders or encoders that may have the same or different structures as those shown in FIGS. 7 and 8. Additionally, the encoder (700) and decoder (800) can be included in the same device or separate devices in various examples.

[0108] This disclosure includes embodiments related to the estimation of Manhattan layouts, including the estimation of the Manhattan layout of a scene from various panoramas of the scene. The embodiments can be used to create virtual reality and augmented reality applications such as virtual tours. For example, the Manhattan layout can be estimated from a number of panoramic images using geometric shapes and segmentation information.

[0109] In applications such as robotics, virtual reality, and augmented reality, it is a common technique to estimate the layout of a room from an image. The layout of a room can include the position, orientation, and height of the walls of the room relative to a specific reference point. For example, wall intersections, 3D meshes, or point clouds can be employed to depict the layout of a room. In the Manhattan layout of a room, the walls of the room are perpendicular to each other. A panoramic image can be generated via a camera such as a panoramic camera. A panoramic image can be applied to depict the layout of a room. However, it can be difficult to estimate the Manhattan layout of a room by analyzing multiple panoramic images. A panorama (or panoramic image) can encapsulate 360-degree information in a scene, and the 360-degree information can contain much more data than a perspective image.

[0110] The Manhattan layout of a room can be estimated using geometric information and semantic segmentation of information (e.g., pixels) from multiple panoramic views (or panoramic images). Semantic segmentation can be performed using deep learning algorithms that associate labels or categories with all pixels in an image. Using semantic segmentation, a set of pixels forming distinct categories can be recognized.

[0111] Since a single panorama may not provide an accurate representation of the room layout, multiple panoramas (or panorama images) can be used to estimate the Manhattan layout of the room. For example, objects in the room may block the boundaries of the room walls, or the room may be very large and a single panorama image may not fully capture the room. The geometry information can include two main directions of the Manhattan layout of the room (e.g., the X direction and the Z direction), as well as line segment information extracted from the panorama. However, since the geometry information focuses on the geometric content of the room, the geometry information may lack semantic information. Therefore, semantic segmentation can be used to provide semantic information for the pixels of the panorama. Semantic segmentation can determine the respective categories of each pixel of the panorama by referring to the labeling of the panorama (or panorama image). For example, a pixel can be labeled as one of the floor, wall, etc. inside the room based on semantic segmentation.

[0112] Figures 9A - 9C show an exemplary geometric and semantic representation of a panorama. As shown in Figure 9A, a panorama (900) is provided, and the panorama (900) can include a scene of a hotel room. In Figure 9B, the panorama (900) can be marked with information on geometric line segments. In Figure 9C, semantic information for the pixels of the room can be provided based on semantic segmentation. For example, the semantic information of the pixels of the room can indicate the ceiling (902), floor (904), sofa (906), wall (908), etc.

[0113] The layout of the room can be represented in various ways, including 3D meshes, boundary lines, and point clouds. In the present disclosure, the 3D mesh and boundary lines of the room can be used to describe the layout of the room. The boundary lines can be represented by one or more polygons, and the 3D mesh of the room can be created from the polygons.

[0114] Figures 10A to 10C show an exemplary polygonal representation and an exemplary mesh representation of a room. As shown in Figure 10A, an estimated room layout (1000) can be provided. In Figure 10B, a polygonal representation (1002) of the estimated room layout can be provided, and the corners of the room layout can be labeled with numbers such as 0 to 7. In Figure 10C, a 3D mesh (1004) can be generated from the polygon (1002). According to Figures 10A to 10C, when a room layout (e.g., (1000)) is established, the wall surface can be projected onto the floor surface to obtain a polygon (e.g., (1002)). Next, when the polygon and the wall height are obtained, the layout and the 3D mesh (e.g., (1004)) can be derived.

[0115] In the present disclosure, the estimation of the Manhattan layout of a room can be performed by estimating a polygon (e.g., polygon (1002)) based on the geometric information and semantic information of a scene (e.g., room layout (1000)). The polygon can be represented by a directed graph G = (v, e), where v is the set of corners of the polygon and e is the set of edges connecting the corners. The position of each corner joint can be defined as 2D coordinates (x, y) in the 2D space with respect to the camera position. Each edge can be represented as a directed line segment (p s , p e ), where p s ∈ v and p e ∈ v are the start point and the end point of the edge, respectively. The line function of each edge can be represented as ax + by + c = 0, where n = (a, b) is

Number

[0116] FIG. 11 shows an overview of a system (or process) (1100) for estimating the Manhattan layout of a scene (e.g., a room). As shown in FIG. 11, at (1110), an input image can be provided. The input image can include a set of panoramas (or multiple panorama images) of the scene. The set of panoramas can capture the scene from different view positions to represent the scene more accurately. In steps (1120) and (1130) of the process (1100), the geometric information of each panorama image can be extracted, and the semantic information of each panorama can be determined based on the semantic segmentation from the corresponding panorama image. The geometric information (or geometric factors) can include (1) lines detected in each panorama image, (2) the main directions (e.g., the X direction and the Z direction) of each panorama image, (3) the ratio between the distance from the ceiling of the room to the floor of the room and the distance from the camera to the floor, and (4) the relative pose (e.g., relative position, angle, or distance) between panorama images such as between two respective panorama images. The semantic information can be obtained through semantic segmentation. Semantic segmentation can assign a semantic meaning (e.g., floor, door, etc.) to each pixel of the panorama image. At step (1140), the respective layout of the scene can be estimated based on the geometric information and semantic segmentation of each panorama image. At step (1150), the layouts estimated from each panorama image (or images of the panorama) can be combined to generate a final estimate of the room layout. At step (1160), based on the estimated layout (or the final estimate of the room layout), a 3D mesh (or Manhattan layout) of the room can be generated.

[0117] As in the example of step (1140) of FIG. 11, to determine a layout estimation associated with a scene (e.g., a room) based on the geometric information and semantic information of each panoramic image, the result of semantic segmentation can be used to determine whether each line segment of the panorama (or panoramic image) represents the boundary of a room wall. Considering a line l represented as a sequence of points p0, p1,..., p n in the panoramic image, it is possible to check whether the neighboring pixels of each point p i include a wall. If the neighboring pixels of point p i include only walls or have no wall pixels at all, point p i may not be considered a boundary point. If the number of boundary points exceeds a specific threshold, for example, if 80% of the points are boundaries, the dotted line (or line) l can be specified as a boundary (or boundary line).

[0118] The boundary line can be aligned with the main direction (or principal direction) of the panorama. To align the boundary line with the main direction, each boundary line can be projected onto a horizontal plane (e.g., the X-Z plane). Since the two main directions (e.g., X and Z) are perpendicular, the angles between the projected boundary line and the two main directions can be calculated. Then, the projected boundary line can be rotated around the center of the projected boundary line to make the projected boundary line parallel to the main direction.

[0119] However, objects within the image (or panoramic image) may block the boundary line. Therefore, one or more boundary lines may be incomplete. In the first approach, a combination of the ceiling boundary line and the floor boundary line can be used to estimate the incomplete boundary line (or complete one or more incomplete boundary lines). For example, each floor line (or floor boundary line) can correspond to a ceiling line (or ceiling boundary line), and the distance between the floor line and the ceiling line can be fixed within the scene. The incomplete boundary line can be projected onto the ceiling line and the floor line respectively. The coordinates of the points of the incomplete boundary line on the corresponding projected boundaries (or projected boundary lines) of the ceiling and the floor can be as follows. p c,1 , p c,2 , p c,3 , …, p c,n p f,1 , p f,2 , p f,3 , …, p f,n Based on the ratio r between the first distance from the ceiling to the ground and the second distance from the camera to the ground subsequently, scale (or estimate) the points of the unfinished boundary line and combine the projection points on the ceiling boundary line and the floor boundary line of the following formula (1). can be combined.

Equation

[0120] In the second method, the projected line segments (e.g., the boundary lines projected onto the horizontal plane) can be connected using the Manhattan layout hypothesis. As defined by the Manhattan layout hypothesis, each pair of two connected boundaries can be parallel or perpendicular. Therefore, the projected boundary lines can be sorted according to the original spatial coordinates of the boundary lines in the image space of the scene (e.g., the X - Y - Z space). When two lines (or boundary lines) are parallel, a vertical line can be added to connect the two lines together. When two lines are perpendicular, it is possible to determine whether the intersection point of the two lines is on the two line segments (or the two lines). In response to the intersection point not being on the two lines, the two line segments (or the two lines) can be extended so that the intersection point can be placed on the two line segments.

[0121] The polygon can be obtained based on one or a combination of the first and second methods. In some embodiments, geometric processing methods such as 2D polygon noise removal and staircase removal can be used to refine the polygon. Thereby, a plurality of polygons can be obtained based on the panoramic image. Each polygon is derived from its respective panoramic image and can represent the respective layout estimation of the scene (e.g., a room).

[0122] For noise removal of existing curves (or line segments), various techniques can be applied. In one example, the boundaries of the curve (or line segment) can be adapted to regions with noisy points, and then thinning of the regions can be applied. In one example, multi-scale analysis such as a Gaussian kernel can be applied. The multi-scale analysis can preserve sharp points with an impact detector and output the curve as a set of smooth arcs and corners. In another example, Gaussian smoothing can be applied to the noise estimated by local analysis using a fixed number n of neighboring numbers such as n = 30.

[0123] Staircase removal can reduce staircase artifacts. Staircase artifacts can be a common artifact that can be observed in many noise removal tasks such as one-dimensional signal noise removal, two-dimensional image noise removal, and video noise removal. Image noise removal techniques can flatten one or more regions of an image signal, thereby generating staircase artifacts in the image signal. As a result, staircase artifacts can appear as unwanted false steps or unwanted flat regions in one or more regions of the image signal, where the image signal would otherwise vary smoothly.

[0124] As in the example of step (1150) of FIG. 11, in order to generate a layout estimation of a scene based on each layout estimation from each panoramic image, the polygon derived in the previous step (e.g., step (1140)) can be transformed into the same coordinate system based on the estimated relative positions between the panoramic images. The transformed (or derived) polygon from step (1140) can be represented as poly0, poly1,... poly n Two separate processes, namely contour estimation and inner edge estimation, can be employed to obtain the layout in order to generate a layout estimation of the scene.

[0125] In contour estimation, the contour of the layout estimation can be determined. Contour estimation can utilize the polygon sum algorithm to combine the transformed polygons poly0, poly1,... poly n together as the baseline polygon poly base .

[0126] Next, the baseline polygon poly base can be shrunk to form a shrunk polygon poly shurink . To shrink the baseline polygon poly base , for each edge e base of the baseline polygon poly i , candidate edges

Number

Number

Number

Number

Number

[0127] All candidate edges [Number] Among them, the candidate edge closer to the origin view position (e.g., the position of the edge in the original panoramic image) [Number] can be used to replace e i Accordingly, one or more candidate edges [Number] When one or more are close to the origin view position, one or more edges e i can be replaced with the corresponding one or more candidate edges [Number]

[0128] In the formation of the baseline polygon poly base all edges of the transformed polygons poly0, poly1,... poly n are merged together in the baseline polygon poly base and the coincidence between the edge e i of the reference polygon and the edges of the transformed polygon need not be considered. By projecting the reference polygon poly base onto each transformed polygon, one or more candidate edges [Number] When close to the origin view position, one or more e i can be replaced with the corresponding one or more candidate edges [Number] Accordingly, the edge e i of the reference polygon and the edge of the transformed polygon can be made to coincide in terms of direction, size, position, etc. ​

[0129] Next, each edge e base of the baseline polygon poly i can be compared with the corresponding candidate edge [Number] Based on whether the corresponding candidate edge [Number] is close to the origin view position, each edge e i can be retained or replaced. The shrunk polygon poly shurink can be formed by replacing one or more edges e base of the baseline polygon poly i .

[0130] The final polygon poly final can be further obtained based on the shrunk polygon by using geometric processing methods such as 2D polygon noise removal and staircase removal.

[0131] Inner edge estimation can be configured to restore the inner edge of the final polygon poly final . To restore the inner edge, all edges of the transformed polygons poly0, poly1,... poly final inside the final polygon poly n can be put into the set E'. Next, an edge can be clustered (or grouped) using a space-based edge voting strategy. For example, two edges in the final polygon poly final can be grouped when the two edges satisfy at least one of the following conditions. (1) The two edges are parallel. (2) The distance between the two edges is small enough, such as less than a first threshold. (3) The projected overlap between the two edges is large enough, such as greater than a second threshold.

[0132] Furthermore, if a group of edges contains more than a specific number of edges, calculate the average edge of the group of edges to obtain the restored inner edge of the final polygon poly final which can represent the restored inner edge of the final polygon poly final and can be added to the final polygon poly as the restored inner edge.

[0133] The 3D shape (or Manhattan layout) of the room can be generated based on the estimated polygon (e.g., the final polygon poly final ) using various representations.

[0134] In one embodiment, the 3D shape of the room can be generated using a triangular mesh. For example, the ceiling and floor surfaces of the room can be generated by triangulating the final polygon poly final . The wall surfaces of the room can be generated by triangulating a rectangle surrounded by the ceiling boundary line and the floor boundary line. To generate the texture of the 3D mesh (or 3D shape), a ray-casting-based method can be further applied.

[0135] In one embodiment, a quadrilateral can be used to represent the 3D shape of the room by quadrilateralizing the final polygon poly final .

[0136] In one embodiment, a point cloud can be used to represent the 3D shape of the room by sampling points from a triangular mesh or a quadrilateral mesh. The triangular mesh can be obtained by triangulating the final polygon poly final . The quadrilateral mesh can be obtained by quadrilateralizing the final polygon poly final .

[0137] In one embodiment, the 3D shape of the room can be generated by voxelizing the final polygon. Thus, the 3D model (e.g., the final polygon polyfinal By converting (1), a voxel (or 3D shape) can be created from volumetric data (e.g., the 3D shape of a room).

[0138] In the present disclosure, a method for estimating the Manhattan layout of a scene (e.g., a room) can be provided. The Manhattan layout of a scene can be estimated from a plurality of panoramic images of the scene using geometric information and semantic segmentation information associated with the scene.

[0139] In one embodiment, the main direction, line segments, and semantic segmentation can be used together to estimate the layout of a scene from a single panorama (or panoramic image) of a plurality of panoramic images.

[0140] In one embodiment, the pose information (e.g., angle or distance) of a panoramic image can be used to combine the layout of each panorama into the final layout estimation.

[0141] In one embodiment, triangulating the final polygon (e.g., poly final ), quadrilaterizing the final polygon, generating a point cloud based on the final polygon, or voxelizing the model (e.g., final layout estimation or final polygon poly final ) can be used to generate a 3D shape (or Manhattan layout) from the final room layout.

[0142] In one embodiment, the line segments in a plurality of panoramas can be detected by a line detection method. For example, as shown in step (1140) of FIG. 11, semantic segmentation can assign a semantic meaning (e.g., floor, door, etc.) to each pixel of a panoramic image. Using the results of semantic segmentation, it can be determined whether each line segment of the panorama (or panoramic image) represents the boundary of a wall of the room.

[0143] In one embodiment, the main directions (e.g., the X and Z directions) of the panoramic image can be obtained by analyzing the statistical information of the line segments in the panoramic image.

[0144] In one embodiment, the semantic segmentation of the panoramic image can be achieved using deep learning-based semantic segmentation technology. For example, the semantic segmentation can be a deep learning algorithm that associates labels or categories with each pixel of the image (e.g., the panoramic image of the scene).

[0145] In one embodiment, the panoramic pose estimation (e.g., angle or distance) of the panoramic image can be achieved using image alignment technology. Based on the alignment of two panoramic images, the relative angle or relative distance between the two panoramic images can be determined.

[0146] In one embodiment, the ratio between the distance from the ceiling to the ground of the scene and the distance from the camera to the ground can be calculated using the segmentation information.

[0147] FIG. 12 shows a diagram of a framework (1200) for image processing according to some embodiments of the present disclosure. The framework (1200) includes a video encoder (1210) and a video decoder (1250). The video encoder (1210) encodes an input (1205), such as a plurality of panoramic images of a scene (e.g., a room), into a bitstream (1245), and the video decoder (1250) decodes the bitstream (1245) to generate a reconstructed 3D shape (1295), such as the Manhattan layout of the scene.

[0148] The video encoder (1210) can be any suitable device such as, for example, a computer, a server computer, a desktop computer, a laptop computer, a tablet computer, a smartphone, a gaming device, an AR device, a VR device, etc. The video decoder (1250) can be any suitable device such as, for example, a computer, a client computer, a desktop computer, a laptop computer, a tablet computer, a smartphone, a gaming device, an AR device, a VR device, etc. The bitstream (1245) can be transmitted from the video encoder (1210) to the video decoder (1250) via any suitable communication network (not shown).

[0149] In the example of FIG. 12, the video encoder (1210) includes a segmentation module (1220), an encoder (1230), and an extraction module (1240) that are combined together. The segmentation module (1220) is configured to assign semantic meanings (e.g., floor, door, etc.) to each pixel of the panoramic image associated with the scene. The semantic information of each panoramic image can be transmitted to the encoder (1230) through the bitstream (1225) to the encoder (1230). The extraction module (1240) is configured to extract the geometric information of each panoramic image. The geometric information can be transmitted to the encoder (1230) through the bitstream (1227). The encoder (1230) is configured to generate the 3D shape (or Manhattan layout) of the scene based on the geometric information and semantic information of each panoramic image. For example, the encoder (1230) can generate a respective layout estimation (or polygon) of the scene based on each panoramic image. The layout estimations of the panoramic images can be fused to form a final layout estimation (or final polygon). The 3D shape of the scene can be generated by triangulating the final polygon, quadrilaterating the final polygon, generating a point cloud based on the final polygon, or voxelizing the final polygon.

[0150] In the example of FIG. 12, a bitstream (1245) is provided to a video decoder (1250). The video decoder (1250) includes a decoder (1260) and a reconstruction module (1290) coupled together as shown in FIG. 12. In one example, the decoder (1260) corresponds to the encoder (1230), can decode the bitstream (1245) encoded by the encoder (1230), and generate decoded information (1265). The decoded information (1265) can be further provided to the reconstruction module (1290). Accordingly, the reconstruction module (1290) can reconstruct the 3D shape (or Manhattan layout) (1295) of the scene based on the decoded information (1265).

[0151] FIG. 13 shows a flowchart illustrating an overview of a process (1300) according to an embodiment of the present disclosure. In various embodiments, the process (1300) is executed by a processing circuit. In some embodiments, the process (1300) is implemented with software instructions, and thus when the processing circuit executes the software instructions, the processing circuit executes the process (1300). The process starts at (S1301) and proceeds to (S1310).

[0152] (S1310), a plurality of two-dimensional (2D) images of the scene are received.

[0153] (S1320), the geometric information and semantic information of each of the plurality of 2D images are determined. The geometric information indicates the detected lines and reference directions in each 2D image. The semantic information includes the classification information of the pixels in each 2D image.

[0154] (S1330), a layout estimation associated with each 2D image of the scene is determined based on the geometric information and semantic information of each 2D image.

[0155] In (S1340), the combined layout estimation associated with the scene is determined based on a plurality of determined layout estimations associated with a plurality of 2D images of the scene.

[0156] In (S1350), the Manhattan layout associated with the scene is generated based on the combined layout estimation. The Manhattan layout includes at least a three-dimensional (3D) shape of the scene including walls orthogonal to each other.

[0157] To determine geometric information and semantic information, the first geometric information of the first 2D image among the plurality of 2D images can be extracted. The first geometric information can include at least one of a detected line, a reference direction of the first 2D image, a ratio of a first distance from the ceiling to the ground to a second distance from the camera to the ground, or a relative pose (e.g., an angle or a distance) between the first 2D image and the second 2D image among the plurality of 2D images. The pixels of the first 2D image can be labeled to generate first semantic information, and the first semantic information can indicate first structural information of the pixels in the first 2D image.

[0158] To determine the layout estimation associated with each 2D image of the scene, the first layout estimation of the plurality of determined layout estimations associated with the scene can be determined based on the first geometric information and the first semantic information of the first 2D image. To determine the first layout estimation, it can be determined whether each of the detected lines is a boundary line corresponding to a wall boundary in the scene. The boundary lines of the detected lines can be aligned with the reference direction of the first 2D image. The first polygon indicating the first layout estimation can be generated based on the boundary lines aligned using one of 2D polygon noise removal and staircase removal.

[0159] To generate the first polygon, a plurality of incomplete boundary lines of the boundary line can be completed based on one of (i) estimating a plurality of incomplete boundary lines based on a combination of the ceiling boundary line and the floor boundary line of the boundary line, and (ii) connecting a pair of incomplete boundary lines among the plurality of incomplete boundary lines. A pair of incomplete boundary lines can be connected based on one of (i) adding a perpendicular line to the pair of incomplete boundary lines in response to the pair of incomplete boundary lines being parallel, and (ii) extending at least one of the pair of incomplete boundary lines, such that an intersection point of the pair of incomplete boundary lines is located on the extended pair of incomplete boundary lines.

[0160] To determine a combined layout estimate associated with the scene, the reference polygon can be determined by combining a plurality of polygons via a polygon summation algorithm. Each of the plurality of polygons can correspond to a respective layout estimate of the plurality of determined layout estimates. The shrunk polygon can be determined based on the reference polygon. The shrunk polygon can include updated edges updated from the edges of the reference polygon. The final polygon can be determined based on the shrunk polygon using one of 2D polygon noise removal and staircase removal. The final polygon can correspond to the combined layout estimate associated with the scene.

[0161] To determine the shrunk polygon, a plurality of candidate edges can be determined from a plurality of polygons of the edges of the reference polygon. Each of the plurality of candidate edges can correspond to a respective edge of the reference polygon. The updated edges of the shrunk polygon can be generated by replacing one or more edges of the reference polygon with one or more candidate edges that are closer to the original view position in a plurality of 2D images than the corresponding one or more edges of the reference polygon.

[0162] In some embodiments, each of the plurality of candidate edges can be parallel to a corresponding edge of the reference polygon. The projected overlap between each candidate edge and the corresponding edge of the reference polygon can be made larger than a threshold value.

[0163] To determine a combined layout estimate associated with the scene, a set of edges including the edges of the final polygon can be determined. A plurality of edge groups can be generated based on the set of edges. A plurality of inner edges of the final polygon can be generated. The plurality of inner edges can be indicated by a plurality of average edges of one or more edge groups of the set of edges. Each of one or more of the plurality of edge groups can include a respective number of edges greater than a target value. Each of the plurality of average edges can be obtained by averaging one edge of each of one or more of the plurality of edge groups.

[0164] In some embodiments, the plurality of edge groups can include a first edge group. The first edge group can further include a first edge and a second edge. The first edge and the second edge can be parallel. The distance between the first edge and the second edge can be less than a first threshold value. The projected overlap region between the first edge and the second edge can be made larger than a second threshold value.

[0165] To generate a Manhattan layout associated with the scene, the Manhattan layout associated with the scene can be generated based on a triangular mesh triangulated from the combined layout estimate, a quadrilateral mesh quadrilateralized from the combined layout estimate, a sampling point sampled from one of the triangular mesh and the quadrilateral mesh, or a discrete grid generated from one of the triangular mesh and the quadrilateral mesh via voxelization.

[0166] In some embodiments, the Manhattan layout associated with a scene can be generated based on a triangular mesh triangulated from a combined layout estimation. Thus, in order to generate the Manhattan layout associated with a scene, the ceiling and floor surfaces in the scene can be generated by triangulating the combined layout estimation. The wall surfaces in the scene can be generated by triangulating a rectangle surrounding the ceiling boundary line and the floor boundary line in the scene. The texture of the Manhattan layout associated with the scene can be generated via a ray-casting-based process.

[0167] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 14 shows a computer system (1400) suitable for implementing certain embodiments of the disclosed subject matter.

[0168] The computer software can be coded using any suitable machine code or computer language that can undergo mechanisms such as assembly, compilation, linking, etc., and can create code containing instructions that can be executed directly, or via interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.

[0169] The instructions can be executed on various types of computers or computer components, including, for example, personal computers, tablet computers, servers, smartphones, game consoles, Internet of Things devices, etc.

[0170] The components shown in FIG. 14 with respect to the computer system (1400) are essentially illustrative and are not intended to imply any limitation regarding the use or functionality of the computer software implementing the embodiments of the present disclosure. Also, the configuration of the components should not be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiments of the computer system (1400).

[0171] The computer system (1400) may include a specific human interface input device. Such a human interface input device can respond to input by one or more human users via, for example, tactile input (such as keystrokes, swipes, movements of a data glove), audio input (such as voice, clapping), visual input (such as gestures), and olfactory input (not depicted). The human interface device can also be used to capture specific media that is not necessarily directly related to conscious input by humans, such as audio (such as voice, music, ambient sound), images (such as scanned images, photographic images obtained from a still camera), and video (such as two-dimensional video, three-dimensional video including stereoscopic video).

[0172] The input human interface device may include one or more of a keyboard (1401), a mouse (1402), a trackpad (1403), a touch screen (1410), a data glove (not shown), a joystick (1405), a microphone (1406), a scanner (1407), and a camera (1408) (only one of each is shown).

[0173] The computer system (1400) may also include certain human interface output devices. Such human interface output devices may, for example, stimulate the senses of one or more human users via tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by a touch screen (1410), a data glove (not shown), or a joystick (1405), although there may also be tactile feedback devices that do not function as input devices), audio output devices (such as a speaker (1409), headphones (not shown), etc.), visual output devices (such as screens (1410) including CRT screens, LCD screens, plasma screens, OLED screens, regardless of whether each has a touch screen input function and regardless of whether each has a tactile feedback function, some of which may be capable of outputting two-dimensional visual output or three-dimensional or higher output via means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and a printer (not shown) may also be included.

[0174] The computer system (1400) may also include a storage device accessible to humans, as well as optical media including a CD / DVD ROM / RW (1420) having a CD / DVD or similar medium (1421), a thumb drive (1422), a removable hard drive or solid state drive (1423), legacy magnetic media such as tapes and floppy disks (not shown), and dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), and their associated media can also be included.

[0175] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include a transmission medium, a carrier wave, or other transient signals.

[0176] The computer system (1400) can also include an interface (1454) to one or more communication networks (1455). The network can be, for example, wireless, wired, optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, terrestrial broadcast TV, vehicle and industrial including CANBus, etc. Certain networks generally require an external network interface adapter connected to a specific general-purpose data port or peripheral bus (1449) (such as a USB port of the computer system (1400)), and others are generally integrated into the core of the computer system (1400) by connection to the system bus as described below (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1400) can communicate with other entities. Such communication can be only unidirectional reception (e.g., TV broadcast), only unidirectional transmission (e.g., CANbus to a specific CANbus device), or bidirectional with other computer systems using, for example, local or wide area digital networks. Specific protocols and protocol stacks can be used for each of those networks and network interfaces as described above.

[0177] The aforementioned human interface devices, human-accessible storage devices, and network interfaces can be connected to the core (1440) of the computer system (1400).

[0178] The core (1440) can include one or more central processing units (CPUs) (1441), a graphics processing unit (GPU) (1442), a dedicated programmable processing device in the form of a field programmable gate array (FPGA) (1443), a hardware accelerator for specific tasks (1444), a graphics adapter (1450), etc. These devices can be connected via a system bus (1448) together with a read-only memory (ROM) (1445), a random access memory (1446), an internal hard drive that is not accessible to the user, an internal mass storage such as an SSD (1447). In some computer systems, the system bus (1448) can be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be directly connected to the system bus (1448) of the core or can be connected via a peripheral bus (1449). In one example, a screen (1410) can be connected to the graphics adapter (1450). Architectures for peripheral buses include PCI, USB, etc.

[0179] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) can execute specific instructions that can, in combination, constitute the aforementioned computer code. That computer code can be stored in the ROM (1445) or RAM (1446). Migration data can also be stored in the RAM (1446), while permanent data can be stored, for example, in the internal mass storage (1447). Fast storage to and retrieval from any of the memory devices can be enabled through the use of a cache memory that can be closely associated with one or more CPUs (1441), GPUs (1442), mass storage (1447), ROM (1445), RAM (1446), etc.

[0180] A computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or may be of the kind well known and available to those of ordinary skill in the computer software arts.

[0181] For purposes of illustration and not limitation, a computer system (1400) having an architecture, specifically a core (1440), can provide functionality as a result of software embodied on one or more tangible computer-readable media being executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.). Such computer-readable media can be associated with media such as the user-accessible mass storage introduced above, and specific storage of the core (1440) of a non-transitory nature such as core internal mass storage (1447) or ROM (1445). The software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (1440). The computer-readable media can include one or more memory devices or chips, depending on the specific requirements. The software can cause the core (1440), and specifically the processors (including CPUs, GPUs, FPGAs, etc.) therein, to define data structures stored in RAM (1446) and modify such data structures according to processes defined by the software, thereby executing specific processes described herein, or specific portions of specific processes. Additionally, or alternatively, the computer system can provide functionality as a result of circuitry (e.g., accelerator (1444)) wired or otherwise embodied to operate instead of or in conjunction with software to execute specific processes described herein, or specific portions of specific processes. References to software can, as necessary, include logic, and vice versa. References to computer-readable media can, as necessary, include circuitry (such as integrated circuits (ICs)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0182] Although the present disclosure describes several exemplary embodiments, there are changes, substitutions, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be understood by those skilled in the art that many systems and methods, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure, can be devised.

Description of Reference Numerals

[0183] 100 Communication system, 105 Sensor, 110 Terminal device, 120 Terminal device, 150 Network, 200 Streaming system, 201 Point cloud source, 202 Point cloud, 203 Encoder, 204 Compressed point cloud, 205 Streaming server, 206 Client subsystem, 207 Copy of compressed point cloud, 208 Client subsystem, 209 Copy of compressed point cloud, 210 Decoder, 211 Reconstructed point cloud that can be rendered, 212 Rendering device, 213 Capture subsystem, 220 Electronic device, 230 Electronic device, 300 V-PCC encoder, 304 Patch information module, 306 Patch generation module, 308 Patch packing module, 310 Geometry image generation module, 312 Texture image generation module, 314 Occupancy map module, 316 Image padding module, 318 Image padding module, 320 Group expansion module, 322 Video compression module, 323 Video compression module, 324 Multiplexer, 332 Video compression module, 334 Entropy compression module, 336 Smoothing module, 338 Auxiliary patch information compression module, 400 V-PCC decoder, 432 Demultiplexer, 434 Video decompression module, 436 Video decompression module, 438 Occupancy map decompression module, 442 Auxiliary patch information decompression module, 444 Geometry reconstruction module, 446 Smoothing module, 448 Texture reconstruction module, 452 Color smoothing module, 510 Video decoder, 520 Parser, 521 Symbol, 551 Scaler / Inverse transform unit, 552 Intra prediction unit, 553 Motion compensation prediction unit, 555 Aggregator, 556 Loop filter unit, 557 Reference picture memory, 558 Current picture buffer, 603 Video encoder, 630 Source coder, 632 Coding engine, 633 (Local) decoder, 634 Reference picture memory, 635 Predictor, 645 Entropy coder, 650 Controller, 700 G-PCC encoder, 701 Input point cloud, 702 Compressed bitstream, 710 Position quantization module, 712Duplicate Point Removal Module, 720 Attribute Transfer Module, 730 Octree Encoding Module, 740 Detail Level Generation Module, 750 Attribute Prediction Module, 760 Residual Quantization Module, 770 Arithmetic Coding Module, 780 Inverse Residual Quantization Module, 781 Addition Module, 790 Memory for Storing Reconstructed Attribute Values, 800 G-PCC Decoder, 801 Compressed Bitstream, 802 Decoded Point Cloud, 810 Arithmetic Decoding Module, 820 Inverse Residual Quantization Module, 830 Octree Decoding Module, 840 LOD Generation Module, 850 Attribute Prediction Module, 860 Memory for Storing Reconstructed Attribute Values, 900 Panorama, 902 Ceiling, 904 Floor, 906 Sofa, 908 Wall, 1000 Room Layout, 1002 Polygon, 1004 3D Mesh, 1100 Process, 1110 Input Panorama, 1120 Extraction of Geometric Information, 1130 Semantic Segmentation, 1140 Estimation of a Single Panorama Layout, 1150 Fusion of Multiple Panorama Layouts, 1160 Generation of Layout & Mesh, 1200 Framework, 1205 Input, 1210 Video Encoder, 1220 Segmentation Module, 1225 Bitstream to Encoder, 1227 Bitstream, 1230 Encoder, 1240 Extraction Module, 1245 Bitstream, 1250 Video Decoder, 1260 Decoder, 1265 Decoded Information, 1290 Reconstruction Module, 1295 Reconstructed 3D Shape, 1300 Process, 1400 Computer System, 1401 Keyboard, 1402 Mouse, 1403 Trackpad, 1405 Joystick, 1406 Microphone, 1407 Scanner, 1408 Camera, 1409 Speaker, 1410 Touch Screen, 1420 CD / DVD ROM / RW, 1421 CD / DVD or Similar Medium, 1422 Thumb Drive, 1423 Removable Hard Drive or Solid State Drive, 1440 Core, 1441 Central Processing Unit (CPU), 1442 Graphics Processing Unit (GPU), 1443Field Programmable Gate Array (FPGA), 1444 Hardware Accelerator for Specific Tasks, 1445 Read-Only Memory (ROM), 1446 Random Access Memory, 1447 Internal Mass Storage, 1448 System Bus, 1449 Peripheral Bus, 1450 Graphics Adapter, 1454 Network Interface, 1455 Communication Network

Claims

Claim 1 A method for estimating a Manhattan layout associated with a scene, the method comprising: receiving a plurality of two-dimensional (2D) images of the scene; determining geometric information and semantic information for each of the plurality of 2D images, wherein the geometric information indicates lines and principal directions detected in each of the 2D images, and the semantic information includes classification information of pixels in each of the 2D images; determining a layout estimation associated with each of the 2D images of the scene based on the geometric information and the semantic information of each of the 2D images; determining a combined layout estimation associated with the scene based on the plurality of determined layout estimations associated with the plurality of 2D images of the scene; generating the Manhattan layout associated with the scene based on the combined layout estimation, the Manhattan layout including at least a three-dimensional (3D) shape of the scene including walls orthogonal to each other; the step of determining the geometric information and the semantic information comprises: extracting first geometric information of a first 2D image among the plurality of 2D images, the first geometric information including at least one of detected lines, a principal direction of the first 2D image, a ratio of a first distance from a ceiling to a ground to a second distance from a camera to the ground, or a relative pose between the first 2D image and a second 2D image among the plurality of 2D images; further comprising labeling pixels of the first 2D image to generate first semantic information, the first semantic information indicating first structural information of the pixels in the first 2D image; the step of determining the layout estimation associated with each of the 2D images of the scene further comprises: determining a first layout estimation of the plurality of determined layout estimations associated with the scene based on the first geometric information and the first semantic information of the first 2D image; the step of determining the first layout estimation comprises: Determining whether each of the detected lines is a boundary line corresponding to a wall boundary in the scene; Aligning the boundary lines of the detected lines with the main direction of the first 2D image; Further including generating a first polygon representing the first layout estimation based on the aligned boundary lines; The step of determining the combined layout estimation associated with the scene; Determining a reference polygon by combining a plurality of polygons via a polygon sum algorithm, each of the plurality of polygons corresponding to each layout estimation of the plurality of determined layout estimations; Determining a shrunk polygon based on the reference polygon, the shrunk polygon including updated edges updated from the edges of the reference polygon; Further including determining a final polygon based on the shrunk polygon, the final polygon corresponding to the combined layout estimation associated with the scene. A method. **Claim 2**: A method for estimating a Manhattan layout associated with a scene, the method comprising: Receiving a plurality of two-dimensional (2D) images of the scene; Determining geometric information and semantic information for each of the plurality of 2D images, the geometric information indicating lines and main directions detected in each of the 2D images, and the semantic information including classification information of pixels in each of the 2D images; Determining a layout estimation associated with each of the 2D images of the scene based on the geometric information and the semantic information of each of the 2D images; Determining a combined layout estimation associated with the scene based on the plurality of determined layout estimations associated with the plurality of 2D images of the scene; Generating the Manhattan layout associated with the scene based on the combined layout estimation, the Manhattan layout including at least a three-dimensional (3D) shape of the scene including wall surfaces orthogonal to each other; The step of determining the geometric information and the semantic information; A step of extracting first geometry information of a first 2D image among the plurality of 2D images, wherein the first geometry information includes at least one of a detected line, a main direction of the first 2D image, a ratio of a first distance from a ceiling to a ground and a second distance from a camera to the ground, or a relative pose between the first 2D image and a second 2D image among the plurality of 2D images. A step of labeling pixels of the first 2D image to generate first semantic information, wherein the first semantic information indicates first structure information of the pixels in the first 2D image, further including the step. The step of determining the layout estimation associated with each 2D image of the scene. A step of determining a first layout estimation of the plurality of determined layout estimations associated with the scene based on the first geometry information and the first semantic information of the first 2D image. The step of determining the first layout estimation. A step of determining whether each of the detected lines is a boundary line corresponding to a wall boundary in the scene. A step of aligning the boundary line of the detected line with the main direction of the first 2D image. A step of generating a first polygon indicating the first layout estimation based on the aligned boundary line. The step of generating the first polygon. A step of estimating a plurality of unfinished boundary lines based on a combination of a ceiling boundary line and a floor boundary line of the boundary line. A method further including a step of completing the plurality of unfinished boundary lines of the boundary line based on one of the following steps: (i) adding a perpendicular line to a pair of unfinished boundary lines in response to the pair of unfinished boundary lines being parallel, and (ii) extending at least one of the pair of unfinished boundary lines so that an intersection point of the pair of unfinished boundary lines is located on the extended pair of unfinished boundary lines, and connecting based on one of the steps.

3. The step of determining the shrunk polygon. A step of determining a plurality of candidate edges from the plurality of polygons for the edges of the reference polygon, each of the plurality of candidate edges corresponding to each edge of the reference polygon, the step; The method according to claim 1, further comprising: generating the updated edge of the shrunk polygon by replacing one or more edges of the reference polygon with the corresponding one or more candidate edges in response to the one or more candidate edges that are closer to the original view position in the plurality of 2D images than the corresponding one or more edges of the reference polygon.

4. Each of the plurality of candidate edges is parallel to the corresponding edge of the reference polygon, The projected overlapping portion between each of the candidate edges and the corresponding edge of the reference polygon is larger than a threshold value, The method according to claim 3.

5. The step of determining the combined layout estimation associated with the scene, A step of determining a set of edges including the edges of the final polygon; A step of generating a plurality of edge groups based on the set of edges; A step of generating a plurality of inner edges of the final polygon represented by a plurality of average edges of one or more edge groups of the set of edges, each of the one or more edge groups of the plurality of edge groups including a respective number of edges greater than a target value, and each of the plurality of average edges being obtained by averaging each of the respective one edge of the one or more edge groups, the method according to claim 1, further comprising the step.

6. The plurality of edge groups includes a first edge group, The first edge group further includes a first edge and a second edge, the first edge and the second edge are parallel, the distance between the first edge and the second edge is less than a first threshold value, and the projected overlapping area between the first edge and the second edge is larger than a second threshold value, The method according to claim 5.

7. The step of generating the Manhattan layout associated with the scene is The method according to claim 1, further comprising the step of generating the Manhattan layout associated with the scene based on the triangular mesh obtained by triangular division from the combined layout estimation, the quadrilateral mesh obtained by quadrilateral division from the combined layout estimation, a sampling point sampled from one of the triangular mesh and the quadrilateral mesh, or one of the discrete grids generated from one of the triangular mesh and the quadrilateral mesh via voxelization.

8. The Manhattan layout associated with the scene is generated based on the triangular mesh obtained by triangular division from the combined layout estimation. The step of generating the Manhattan layout associated with the scene comprises further the step of generating a ceiling surface and a floor surface in the scene by triangularly dividing the combined layout estimation, the step of generating the wall surfaces in the scene by triangularly dividing a rectangle surrounding the ceiling boundary line and the floor boundary line in the scene, and the step of generating a texture of the Manhattan layout associated with the scene via a ray-casting based process. The method according to claim 7.

9. An apparatus for estimating a Manhattan layout associated with a scene, the apparatus performing the method according to any one of claims 1 to 8.

10. A computer program comprising one or more instructions which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method of reconstituting scene from single two-dimensional image

    JP2014220804A

  • Photography-based 3D modeling system and method, automatic 3D modeling device and method

    JP2022501684A