Method and apparatus for reducing reality in panoramic images
Patent Information
- Application Number
- CN202310603253.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-26
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2043-05-26
AI Technical Summary
[0004]然而,当该方法应用于三维图像的修复时,修复后的图像的真实感相对较差
[0039]The panoramic image reduction reality method and apparatus disclosed herein, applied to indoor scenes, include: generating layout features based on an acquired masked layout boundary image, a mask image, and a masked panoramic image, wherein the layout features characterize the structural features of the original panoramic image at the layout level; generating a style matrix corresponding to the structured regions of the indoor scene based on the acquired masked panoramic image and the original panoramic image, wherein the style matrix characterizes the structural semantic information corresponding to the structured regions; performing a fill process on a preset structured mask according to the style matrix to obtain structured region texture features; and performing panoramic image inpainting processing based on the layout features and the structured region texture features to obtain a reduced reality predicted image corresponding to the masked panoramic image. In this embodiment, by generating layout features that characterize the structural features of the original panoramic image at the layout level, and a style matrix that characterizes the structural semantic information corresponding to the structured regions, the structured region texture features are obtained by filling based on the style matrix. Based on the layout features and the structured region texture features, the technical features of the predicted image are obtained. This allows the realistic restoration capability to be combined with the preservation of boundary structure, better restoring the scene structure. It can also explore the complementarity between structure and texture, thereby preserving the structure of the indoor scene while generating a background image that includes realistic background texture. This can improve the accuracy and reliability of the predicted image, so that the background image of the object to be removed in the predicted image can highly restore the image of the real scene.
Smart Images

Figure CN116797768B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and in particular to a method and apparatus for reducing the reality of panoramic images. Background Technology
[0002] Reducing reality is achieved by drawing a mask area of the object to be removed on a panoramic image and then rendering the truth of the scene behind the object in the mask area. This rendering operation is called image inpainting in image processing terminology.
[0003] In existing technologies, image inpainting is mainly two-dimensional image inpainting, and it mainly generates realistic textures by searching for nearest neighbors or copying related blocks.
[0004] However, when this method is applied to the restoration of 3D images, the restored images have a relatively poor sense of realism. Summary of the Invention
[0005] This disclosure provides a method and apparatus for reducing reality using panoramic images, applied to indoor scenes, to improve the effectiveness and reliability of reducing reality.
[0006] In a first aspect, this disclosure provides a method for reducing the realism of panoramic images, applied to indoor scenes, the method comprising: Based on the acquired masked layout boundary image, mask image, and masked panoramic image, layout features are generated, wherein the layout features characterize the structural features of the original panoramic image at the layout level. Based on the acquired masked panoramic image and the original panoramic image, a style matrix corresponding to the structured regions of the indoor scene is generated, wherein the style matrix represents the structural semantic information corresponding to the structured regions. The preset structured mask is filled according to the style matrix to obtain structured region texture features; Based on the layout features and the structured region texture features, panoramic image inpainting is performed to obtain a reduced-reality prediction image corresponding to the masked panoramic image.
[0007] In some embodiments, based on the acquired masked layout boundary image, mask image, and masked panoramic image, layout features are generated, including: Based on the masked layout boundary image, the mask image, and the masked panoramic image, layout boundary prediction is performed to obtain a boundary layout map. The boundary layout diagram is subjected to structural feature extraction processing to obtain the layout boundary features; The layout features are generated based on the layout boundary features, the mask image, and the masked panoramic image.
[0008] In some embodiments, the boundary layout map is obtained based on a pre-trained layout boundary prediction model; the layout boundary prediction model includes a downsampling convolutional layer, a transformer block, and a transposed upsampling convolutional layer connected in sequence.
[0009] In some embodiments, the layout boundary features are obtained based on a layout feature extraction model; the layout feature extraction model includes a downsampling gated convolutional layer, a dilated convolutional residual block, and an upsampling gated convolutional layer connected in sequence.
[0010] In some embodiments, the masked layout boundary image is obtained by predicting the Manhattan layout boundary of the target object in the original panoramic image and then masking the Manhattan layout boundary. The target objects include walls, ceilings, and floors.
[0011] In some embodiments, the Manhattan layout boundary is determined based on a pre-trained layout structure image generation model, which includes an encoder and a decoder connected in sequence. The encoder takes the original panoramic image as input and the decoder outputs the Manhattan layout boundary.
[0012] In some embodiments, the encoder includes a convolutional layer, and a noise linear rectified function and a pooling layer respectively connected to the output of the convolutional layer; The decoder includes an upsampling layer, and a convolutional layer and an activation layer that are sequentially connected to the output of the upsampling layer.
[0013] In some embodiments, based on the acquired masked panoramic image and the original panoramic image, a style matrix corresponding to the structured regions of the indoor scene is generated, including: Based on the target object, the masked panoramic image is subjected to structured segmentation processing to obtain a structured region map including the structured region corresponding to the target object; The style matrix is constructed based on the structural semantic information of the structured region map.
[0014] In some embodiments, the structured region map is obtained by processing the masked panoramic image based on a pre-trained structured encoder; the structured encoder includes a skip-connected downsampling convolutional layer and an upsampling convolutional layer.
[0015] In some embodiments, the style matrix is obtained by processing the structured region map and the original panoramic image based on a pre-trained semantic prior encoder; the semantic prior encoder includes a convolutional layer, a transposed convolutional layer, and an average pooling layer connected in sequence.
[0016] In some embodiments, a preset structured mask is filled according to the style matrix to obtain structured region texture features, including: Based on the style matrix, preset Gaussian noise, layout features, and structured mask, local feature extraction processing is performed to obtain an initial local texture; The initial local texture of the repaired region corresponding to the mask image is repaired according to the style matrix to obtain the structured region texture features.
[0017] In some embodiments, the structured region texture features are generated based on a pre-trained residual network model. The inputs of the residual network model are the style matrix, preset Gaussian noise, the layout features, and the structured mask. The residual network model includes convolutional layers, and each convolutional layer of the residual network model includes a building block, a noise linear rectified function, and a convolutional kernel connected in sequence.
[0018] In some embodiments, panoramic image inpainting is performed based on the layout features and the structured region texture features to obtain a reduced-reality predicted image corresponding to the masked panoramic image, including: The layout features are convolved to obtain the first convolutional layout features; The first convolutional layout feature and the structured region texture feature are fused to obtain a combined feature; The layout features are convolved to obtain second convolutional layout features, and global feature extraction is performed on the layout features to obtain global features. The second convolutional layout feature and the global feature are fused to obtain the frequency domain layout feature; The combined features and the frequency domain layout features are fused to obtain the predicted image.
[0019] In some embodiments, the predicted image is obtained by processing the layout features and the structured region texture features based on a pre-trained Fourier convolutional fusion model; the Fourier convolutional fusion model includes a downsampling convolutional layer, a Fourier convolutional fusion layer, an upsampling convolutional layer, a spectral transform block, and a fusion module.
[0020] In some embodiments, the predicted image is generated based on a pre-trained insulation network model, the input of which is the boundary layout map, the mask image, and the masked panoramic image; The repair network model is trained based on a fusion loss function, which is obtained by fusing the absolute error loss function, the adversarial loss function, and the advanced synthetic perception loss function.
[0021] Secondly, this disclosure provides a panoramic image reduction reality device for use in indoor scenes, the device comprising: The first generation unit is used to generate layout features based on the acquired masked layout boundary image, mask image, and masked panoramic image, wherein the layout features characterize the structural features of the original panoramic image at the layout level. The second generation unit is used to generate a style matrix corresponding to the structured region of the indoor scene based on the acquired masked panoramic image and the original panoramic image, wherein the style matrix represents the structural semantic information corresponding to the structured region. The filling unit is used to fill the preset structured mask according to the style matrix to obtain the structured region texture features; The repair unit is used to perform panoramic image repair processing based on the layout features and the structured region texture features to obtain a reduced-reality prediction image corresponding to the masked panoramic image.
[0022] In some embodiments, the first generating unit includes: The prediction subunit is used to predict the layout boundary based on the masked layout boundary image, the mask image, and the masked panoramic image to obtain a boundary layout map. Extract sub-units to perform structural feature extraction processing on the boundary layout diagram to obtain layout boundary features; A generation subunit is used to generate the layout features based on the layout boundary features, the mask image, and the masked panoramic image.
[0023] In some embodiments, the boundary layout map is obtained based on a pre-trained layout boundary prediction model; the layout boundary prediction model includes a downsampling convolutional layer, a transformer block, and a transposed upsampling convolutional layer connected in sequence.
[0024] In some embodiments, the layout boundary features are obtained based on a layout feature extraction model; the layout feature extraction model includes a downsampling gated convolutional layer, a dilated convolutional residual block, and an upsampling gated convolutional layer connected in sequence.
[0025] In some embodiments, the masked layout boundary image is obtained by predicting the Manhattan layout boundary of the target object in the original panoramic image and then masking the Manhattan layout boundary. The target objects include walls, ceilings, and floors.
[0026] In some embodiments, the Manhattan layout boundary is determined based on a pre-trained layout structure image generation model, which includes an encoder and a decoder connected in sequence. The encoder takes the original panoramic image as input and the decoder outputs the Manhattan layout boundary.
[0027] In some embodiments, the encoder includes a convolutional layer, and a noise linear rectified function and a pooling layer respectively connected to the output of the convolutional layer; The decoder includes an upsampling layer, and a convolutional layer and an activation layer that are sequentially connected to the output of the upsampling layer.
[0028] In some embodiments, the second generating unit includes: The segmentation subunit is used to perform structured segmentation processing on the masked panoramic image according to the target object, so as to obtain a structured region map including the structured region corresponding to the target object; A sub-unit is constructed to build the style matrix based on the structural semantic information of the structured region map.
[0029] In some embodiments, the structured region map is obtained by processing the masked panoramic image based on a pre-trained structured encoder; the structured encoder includes a skip-connected downsampling convolutional layer and an upsampling convolutional layer.
[0030] In some embodiments, the style matrix is obtained by processing the structured region map and the original panoramic image based on a pre-trained semantic prior encoder; the semantic prior encoder includes a convolutional layer, a transposed convolutional layer, and an average pooling layer connected in sequence.
[0031] In some embodiments, the filling unit includes: The first processing subunit is used to perform local feature extraction processing based on the style matrix, preset Gaussian noise, the layout features, and the structured mask to obtain an initial local texture; The repair subunit is used to repair the initial local texture of the repair region corresponding to the mask image according to the style matrix, so as to obtain the structured region texture features.
[0032] In some embodiments, the structured region texture features are generated based on a pre-trained residual network model. The inputs of the residual network model are the style matrix, preset Gaussian noise, the layout features, and the structured mask. The residual network model includes convolutional layers, and each convolutional layer of the residual network model includes a building block, a noise linear rectified function, and a convolutional kernel connected in sequence.
[0033] In some embodiments, the repair unit includes: A convolutional subunit is used to perform convolution processing on the layout features to obtain a first convolutional layout feature; The first fusion subunit is used to fuse the first convolutional layout feature and the structured region texture feature to obtain a combined feature; The second processing subunit is used to perform convolution processing on the layout features to obtain second convolution layout features, and to perform global feature extraction processing on the layout features to obtain global features. The second fusion subunit is used to fuse the second convolutional layout feature and the global feature to obtain the frequency domain layout feature; The third fusion subunit is used to fuse the combined features and the frequency domain layout features to obtain the predicted image.
[0034] In some embodiments, the predicted image is obtained by processing the layout features and the structured region texture features based on a pre-trained Fourier convolutional fusion model; the Fourier convolutional fusion model includes a downsampling convolutional layer, a Fourier convolutional fusion layer, an upsampling convolutional layer, a spectral transform block, and a fusion module.
[0035] In some embodiments, the predicted image is generated based on a pre-trained insulation network model, the input of which is the boundary layout map, the mask image, and the masked panoramic image; The repair network model is trained based on a fusion loss function, which is obtained by fusing the absolute error loss function, the adversarial loss function, and the advanced synthetic perception loss function.
[0036] Thirdly, this disclosure provides a processor-readable storage medium storing a computer program for causing the processor to perform the method described in the first aspect.
[0037] Fourthly, this disclosure provides an electronic device, including: a processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in the first aspect.
[0038] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0039] The panoramic image reduction reality method and apparatus disclosed herein, applied to indoor scenes, include: generating layout features based on an acquired masked layout boundary image, a mask image, and a masked panoramic image, wherein the layout features characterize the structural features of the original panoramic image at the layout level; generating a style matrix corresponding to the structured regions of the indoor scene based on the acquired masked panoramic image and the original panoramic image, wherein the style matrix characterizes the structural semantic information corresponding to the structured regions; performing a fill process on a preset structured mask according to the style matrix to obtain structured region texture features; and performing panoramic image inpainting processing based on the layout features and the structured region texture features to obtain a reduced reality predicted image corresponding to the masked panoramic image. In this embodiment, by generating layout features that characterize the structural features of the original panoramic image at the layout level, and a style matrix that characterizes the structural semantic information corresponding to the structured regions, the structured region texture features are obtained by filling based on the style matrix. Based on the layout features and the structured region texture features, the technical features of the predicted image are obtained. This allows the realistic restoration capability to be combined with the preservation of boundary structure, better restoring the scene structure. It can also explore the complementarity between structure and texture, thereby preserving the structure of the indoor scene while generating a background image that includes realistic background texture. This can improve the accuracy and reliability of the predicted image, so that the background image of the object to be removed in the predicted image can highly restore the image of the real scene. Attached Figure Description
[0040] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0041] Figure 1 This explains the implementation principle of panoramic image DR for indoor scenes in related technologies; Figure 2 This is a schematic diagram of a panoramic image reduction method according to an embodiment of the present disclosure; Figure 3 A schematic diagram of a panoramic image reduction method according to another embodiment of this disclosure; Figure 4 A schematic diagram illustrating the principle of a panoramic image reduction method according to an embodiment of this disclosure; Figure 5 This is a schematic diagram illustrating the principle of the structured region texture extraction model according to an embodiment of the present disclosure; Figure 6 This is a schematic diagram illustrating the principle of the Fourier convolution fusion model according to an embodiment of this disclosure; Figure 7 This is a schematic diagram showing the effect comparison between the technical solutions of the embodiments of this disclosure and the technical solutions in related technologies; Figure 8This is a schematic diagram showing the comparison results of the technical solutions of the embodiments of this disclosure and the technical solutions in related technologies. Figure 9 This is a schematic diagram showing the effect comparison between the technical solutions of the embodiments of this disclosure and the technical solutions in related technologies; Figure 10 A schematic diagram of a device for reducing reality from panoramic images according to an embodiment of the present disclosure; Figure 11 This is a schematic diagram of an electronic device used to implement the panoramic image reduction reality method according to embodiments of the present disclosure.
[0042] The accompanying drawings have illustrated specific embodiments of this disclosure, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this disclosure to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0043] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0044] It should be understood that the terms “comprising” and “having” and any variations thereof in the embodiments of this disclosure are intended to cover but not exclude inclusion. For example, a product or device that includes a series of components is not necessarily limited to those components that are explicitly listed, but may include other components that are not explicitly listed or that are inherent to such product or device.
[0045] In this disclosure, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0046] In this disclosure, the term "multiple" refers to two or more, and other quantifiers are similar.
[0047] The terms “first,” “second,” “third,” etc., used in this disclosure are used to distinguish similar or related objects or entities and do not necessarily imply a specific order or sequence, unless otherwise indicated. It should be understood that such terms can be used interchangeably where appropriate, for example, in situations where implementation can proceed in an order other than those given in the illustrations or descriptions of embodiments of this disclosure.
[0048] As used in this disclosure, the term "unit / module" means any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code capable of performing the functions associated with that element.
[0049] To facilitate the reader's understanding of this disclosure, at least some of the terms used in this disclosure are explained below: Augmented Reality (AR) is a real-time interactive visualization method that uses computer technology and graphics methods to add virtual objects to the real world, making them exist in the same image or space.
[0050] Mixed Reality (MR) is a further development of virtual reality technology. This technology enhances the realism of the user experience by introducing real-world scene information into the virtual environment and establishing an interactive feedback loop between the virtual world, the real world, and the user.
[0051] Deep learning (DL) is a subfield of machine learning (ML). It involves learning the inherent patterns and hierarchical representations of sample data. The information gained during this learning process is of great help in interpreting data such as text, images, and sound.
[0052] Diminished Reality (DR) refers to the process of removing physical objects from the user's visual perception.
[0053] Image inpainting refers to a rendering operation that involves drawing a mask area of the object to be removed on a panoramic image and then rendering the real values of the scene behind the object in the mask area.
[0054] AR, as an important method for showcasing interior design, can help users intuitively understand the spatial relationships and size of an interior space. Users can place designed interior scene models into the corresponding real world or into panoramic images taken based on the designed interior scene space, and can move corresponding virtual objects (such as furniture) to appropriate positions in the real scene, thereby experiencing the effects of virtual design in the real world.
[0055] Typically, indoor scenes are pre-arranged, and some existing real-world objects are replaced during the design process. In this case, virtual objects partially overlap with real-world objects, and real-world objects cannot be completely covered by virtual objects. Therefore, the effect of AR is greatly reduced, and it is also impractical to remove all real-world objects from the indoor scene.
[0056] Therefore, in addition to adding virtual objects to real-world scenes, virtually removing real-world objects is also crucial; this process is known as Dynamic Removal (DR). DR applications in real-world scenes can hide, eliminate, and reveal objects while simultaneously perceiving the environment. Compared to AR and MR, which add virtual objects to real-world scenes, DR requires detecting unwanted real-world objects and replacing them with hidden backgrounds in the generated image. In indoor scenes, the most basic operation involves removing clutter (such as furniture and other non-permanent objects), which can be defined using interactive masks or semantic and instance segmentation.
[0057] In related technologies, hidden background images can be synthesized through reprojection. However, this method uses multiple cameras observing from different viewpoints of the same scene to generate the background image, but for indoor scenes, the background behind the object to be removed is unknown. For example, the object to be removed—furniture—is usually placed against a wall, and its background is obscured from any angle. Therefore, the multi-camera approach cannot achieve background image reconstruction in indoor scenes.
[0058] Alternatively, a reasonable method can be used to generate a background image, rather than restoring the actual background image. For example, the surrounding area of the object to be removed can be analyzed to recover the background image of the object from the image of the surrounding area. However, this type of method is usually limited to small removal areas and regular scenes.
[0059] In contrast, structural reasoning in indoor scenes is crucial for DR, as it not only improves texture reprojection and parallax effects but also provides a foundation for image editing operations.
[0060] For example, Figure 1This describes the implementation principle of panoramic image DR (Digital Removal) for indoor scenes in related technologies. The target image is the image corresponding to the object to be removed; the target mask can be appended to the original panoramic image to transform the target image into a target mask; the source area is multiple regions obtained by structurally dividing the original panoramic image, such as the area corresponding to the walls, the area corresponding to the ceiling, and the area corresponding to the floor; the source mask aims to use the texture features of the source area to fill the structural area corresponding to the source mask, thereby obtaining the DR predicted image.
[0061] Based on the above analysis, it can be seen that related technologies can use two-dimensional image inpainting methods to obtain the background image, and mainly generate realistic textures through nearest neighbor search or copying related blocks. In cases of large-area texture repetition, damaged images can be realistically repaired. With the development of deep learning, this inpainting task has been modeled as conditional generation that learns a function mapping between the damaged image and the original undamaged input image. Semantic and structural conditional information, such as lines, edges, and approximate images, can be used to assist the inpainting task.
[0062] For example, deep learning-based image inpainting methods use edge (canny) information as important prior information. Another example is EdgeConnect, a generative image inpainting method based on adversarial edge learning, which, considering the importance of edge-preserving structure generation, divides the image inpainting problem into structure prediction and image completion, predicting the image structure of missing areas in the form of edge mapping. Yet another example is the Incremental Structure Enhancement Inpainting Model (ZITS), which incrementally adds auxiliary information to the trained inpainting model without requiring further training. Furthermore, the GateConv method automatically learns masks from a large number of examples, allowing users to use free-form masks as input to guide inpainting. Markov adversarial networks propose a method for training effective texture synthesis, reflecting the importance of feature fusion at different scales. Finally, LaMa, a robust high-resolution large-mask inpainting method based on Fourier convolution, proposes a high-resolution robust large-mask inpainting method that increases the receptive field of the inpainting network and loss function, enabling image inpainting in large blank areas.
[0063] However, the above method is a two-dimensional image restoration method, while panoramic images are equidistant projections (ERPs). Therefore, when the above method is applied to the restoration of panoramic images, it will cause two-level distortion problems due to equirectangular projection. That is, the above method cannot be directly applied to the restoration task of panoramic images.
[0064] In panoramic image-based restoration tasks, PanoDR (PanoReality Reduction) of indoor scenes guides the generation of background images in the same scene by predicting the structure of the interior, thereby achieving the purpose of background image reconstruction.
[0065] For example, the Instant method for automatic clearing of indoor scenes uses an end-to-end approach to compute an attention mask for clutter in the image based on the geometric differences between the full scene and the empty scene. This attention mask is propagated through gated convolution, which drives the generation of the output image and its depth. Another example is the 360-degree panoramic image inpainting network based on cube maps (PIINET), which applies two-dimensional image inpainting methods to panoramic images through the transformation between cube mapping and isometric projection.
[0066] With the widespread adoption of consumer-grade 360-degree cameras, low-cost, high-quality scene capture can be achieved with a single lens, driving advancements in the field of indoor scene understanding. PanoContext, a whole-house 3D context model for panoramic scene understanding, uses spherical panoramic images to estimate the layout of indoor scenes (such as rooms), enabling reconstruction of indoor scenes from a single viewpoint. In addition to structural understanding of the indoor scene, panoramic images also provide an understanding of the semantic content of the overall scene, such as semantic segmentation. Considering the importance of scene understanding for scene reconstruction, PanoDR applies prior semantic information from panoramic images to image inpainting, helping to restore complete Manhattan boundaries.
[0067] DR can also be viewed as an image transfer task, as it maps the textured portions of an indoor scene in a panoramic image to masked regions. In this context, preserving visual content and style is crucial. For example, Conditional Generative Adversarial Networks (GANs) can be used as a general solution to the transfer problem, where semantic synthesis uses semantic tags to reconstruct the image based on the semantic mapping, preserving class boundaries. However, preserving semantic information in deep layers constructed from stacked convolutions, normalization, and nonlinear layers is difficult because normalization layers tend to obscure semantic inputs. Therefore, spatially adaptive normalization can be introduced, where the input mapping modulates activation in the normalization layer through spatially adaptive learning transformation. Furthermore, a region-by-region style matrix can be introduced, allowing users to select different styles of input images for each semantic region. PanoDR applies these methods to panoramic images of indoor scenes, mapping each pixel to types such as ceiling, walls, and floor based on pixel-level semantic priors, while simultaneously utilizing building blocks (SEANs) for inpainting.
[0068] However, while the aforementioned methods can learn meaningful semantics in 2D images and generate coherent structures and textures for missing regions, the resulting background images lack realism. The isometric projection algorithm for panoramic images suffers from structural distortion, making structure recovery difficult. Furthermore, due to the omnidirectional nature of panoramic images—that is, the continuity between different directions—this leads to smooth boundaries and texture artifacts when converting panoramic images to 2D images for inpainting.
[0069] To avoid the aforementioned technical problems, this disclosure proposes a technical concept developed through inventive effort: Based on a two-dimensional image inpainting network, to learn the overall structure of a panoramic image of an indoor scene, a pre-trained structure restoration module extracts structural layout features to combine realistic restoration capabilities with boundary structure preservation, thereby better restoring the scene structure; then, to maintain consistency between the generated region texture and the textures of other regions in the panoramic image, a structured region texture extraction module aggregates local texture features to restore the removed object region; and a Fourier convolution fusion module fuses local texture features and structural layout features to explore the complementarity between structure and texture, thereby preserving the structure of the indoor scene while generating realistic background textures.
[0070] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this disclosure.
[0071] Based on the above technical concept, this disclosure provides a method for reducing reality from panoramic images, which can be applied to indoor scenes.
[0072] Please see Figure 2 , Figure 2 This is a schematic diagram of a panoramic image reduction reality method according to an embodiment of the present disclosure, as shown below. Figure 2 As shown, the method includes: S201: Based on the acquired masked layout boundary image, mask image, and masked panoramic image, generate layout features, where the layout features characterize the structural features of the original panoramic image at the layout level.
[0073] For example, the execution subject of this embodiment can be a device for panoramic image reduction to reality. This device can be a server, a terminal device, a processor, a chip, etc., and will not be listed here. This embodiment uses a server as an example for illustrative explanation.
[0074] For example, if the device is a server, it can be a cloud server or a local server; it can be a standalone server or a server cluster, and this embodiment does not limit it.
[0075] Among them, the masked layout boundary image is the image obtained by masking the Manhattan layout boundary. The Manhattan layout boundary is the layout boundary used to represent the original panoramic image. In an indoor scene, the Manhattan layout boundary can be understood as the boundary between walls, the boundary between ceiling and walls, and the boundary between walls and floors.
[0076] The masked image is an image that includes a mask. The masked panoramic image is an image obtained by masking the original panoramic image based on the masked image. The original panoramic image is an unmasked image, which is a 3D image of the captured indoor scene.
[0077] Layout features characterize structural features, specifically the structural features of the original panoramic image at the layout level, such as the structural features of the walls, ceiling, and floor in the original panoramic image. These structural features can be understood as the boundary features conforming to the Manhattan layout.
[0078] It should be understood that this embodiment does not limit the method of obtaining the masked layout boundary image, the mask image, and the masked panoramic image, for example: In one example, the server can connect to the acquisition device and receive the masked layout boundary image, the masked image, and the masked panoramic image sent by the acquisition device.
[0079] In another example, the server can provide a tool for loading images, which users can use to transfer the masked layout boundary image, the mask image, and the masked panoramic image to the server.
[0080] The tool for loading images can be an interface for connecting to external devices, such as an interface for connecting to other storage devices. Through this interface, the masked layout boundary image, mask image, and masked panoramic image transmitted by the external device can be obtained. Alternatively, the tool for loading images can be a display device. For example, the server can input an interface for loading images on the display device, and the user can import the masked layout boundary image, mask image, and masked panoramic image to the server through this interface.
[0081] S202: Based on the acquired masked panoramic image and the original panoramic image, generate a style matrix corresponding to the structured regions of the indoor scene, wherein the style matrix represents the structural semantic information corresponding to the structured regions.
[0082] Based on the above analysis, a structured region can be understood as the area corresponding to different target objects in an indoor scene. Target objects include walls, ceilings, and floors. In other words, a structured region can include the area corresponding to the walls, the area corresponding to the ceiling, and the area corresponding to the floor.
[0083] Correspondingly, structural semantic information can be understood as the semantic information of a structured region in relation to the type of the target object. The style matrix can be further understood as the style code of a structured region in relation to the type of the target object. For example, the style matrix corresponding to a wall represents the structural style code of the wall, the style matrix corresponding to a ceiling represents the structural style code of the ceiling, and the style matrix corresponding to a floor represents the structural style code of the floor.
[0084] Similarly, for the methods of obtaining the masked panoramic image and the original panoramic image, please refer to the implementation principle of the above example, which will not be repeated here.
[0085] S203: Fill the preset structured mask according to the style matrix to obtain the structured region texture features.
[0086] This embodiment does not limit the content of the structured mask; for example, it can be determined based on requirements, historical records, and experiments.
[0087] Based on the above analysis, it can be seen that the style matrix can be used to represent the structural semantic information corresponding to each of the multiple structured regions. Therefore, in this step, the structured mask of each structured region in the structured mask can be filled based on the style matrix to obtain the texture features of the structured region.
[0088] Taking the style matrix including the style matrix corresponding to the wall as an example, the mask corresponding to the wall in the structured mask can be filled based on the structural semantic information corresponding to the wall (i.e. the style matrix corresponding to the wall, which can be further described as the style code corresponding to the wall), thereby obtaining the structured region texture features corresponding to the wall in the structured region texture features.
[0089] S204: Perform panoramic image inpainting based on layout features and structured region texture features to obtain a reduced-reality prediction image corresponding to the masked panoramic image.
[0090] For example, after obtaining the layout features and structured region texture features, the background image of the object to be removed can be repaired and predicted based on the layout features and structured region texture features, so as to obtain a predicted image that includes the background image of the object to be removed, that is, the predicted image includes the background image of the object to be removed, as well as the background image of the object not to be removed.
[0091] Based on the above analysis, this disclosure provides a method for reducing reality in panoramic images. This method can be applied to indoor scenes. The method includes: generating layout features based on the acquired masked layout boundary image, mask image, and masked panoramic image, wherein the layout features characterize the structural features of the original panoramic image at the layout level; generating a style matrix corresponding to the structured regions of the indoor scene based on the acquired masked panoramic image and the original panoramic image, wherein the style matrix characterizes the structural semantic information corresponding to the structured regions; performing filling processing on a preset structured mask according to the style matrix to obtain the texture features of the structured regions; and performing panoramic image inpainting processing based on the layout features and the texture features of the structured regions to obtain the reduced reality corresponding to the masked panoramic image. In this embodiment, the predicted image is less realistic. By generating layout features that represent the structural features of the original panoramic image at the layout level, and a style matrix that represents the structural semantic information corresponding to the structured regions, the texture features of the structured regions are obtained by filling based on the style matrix. The technical features of the predicted image are obtained by repairing based on the layout features and the texture features of the structured regions. This allows the realistic restoration capability to be combined with the preservation of boundary structure, better restoring the scene structure. It can also explore the complementarity between structure and texture, so that the structure of the indoor scene is preserved while generating a background image including realistic background texture. This can improve the accuracy and reliability of the predicted image, so that the background image of the object to be removed in the predicted image can highly restore the image of the real scene.
[0092] Based on the above analysis, it can be seen that the background image of the removed object can be repaired using deep learning. This disclosure can also use deep learning to obtain the predicted image. To facilitate a deeper understanding of the implementation principle of this disclosure, the following is a combination of... Figures 3 to 9 The method for reducing reality in panoramic images disclosed herein is described in detail.
[0093] in, Figure 3 This is a schematic diagram of a panoramic image reduction reality method according to another embodiment of the present disclosure, which can be applied to indoor scenes, such as... Figure 3 As shown, the method includes: S301: From the acquired original panoramic image, predict the Manhattan layout boundary and perform masking processing on the Manhattan layout boundary to obtain the masked layout boundary image.
[0094] It should be understood that, in order to avoid tedious descriptions, the technical features that are the same as or similar to those in the above embodiments will not be repeated in this embodiment.
[0095] For example, regarding the implementation method for acquiring the original panoramic image, please refer to the examples above. Similarly, regarding the execution entity of this embodiment, please refer to the examples above.
[0096] For example, the masked layout boundary image is obtained by predicting the Manhattan layout boundary from the target objects in the original panoramic image and then masking the Manhattan layout boundary. The target objects include walls, ceilings, and floors.
[0097] For example, if the original panoramic image includes a room, the Manhattan layout boundaries of the room can be predicted from the original panoramic image.
[0098] In some embodiments, the Manhattan layout boundary is determined based on a pre-trained layout structure image generation model, which includes an encoder and a decoder connected in sequence. The encoder takes the original panoramic image as input and the decoder outputs the Manhattan layout boundary.
[0099] This embodiment does not limit the training method of the layout structure image generation model. For example, sample data can be obtained to train the basic network model based on the sample data, so as to train the basic network model to learn the ability to predict the Manhattan layout boundary, thereby obtaining the layout structure image generation model.
[0100] In some embodiments, the encoder includes a convolutional layer, and a noise rectified linear function (ReLU) and a pooling layer respectively connected to the output of the convolutional layer. The decoder includes an upsampling layer, and a convolutional layer and an activation layer (Sigmoid) sequentially connected to the output of the upsampling layer.
[0101] For example, the layout structure image generation model can employ a layout network (LayoutNet), which includes an encoder and a decoder. The encoder takes the original panoramic image as input and uses an alignment method to stitch together the original panoramic image with a resolution of 512x1024 (perspective view of 512x512) with Manhattan line segment feature maps located in three orthogonal vanishing directions.
[0102] The encoder consists of seven convolutional layers with 3x3 kernels, each followed by a ReLU operation and a max-pooling layer with a downsampling factor of 2. The first convolutional layer can contain 32 image features from the original panoramic image, and the size is doubled after each convolutional layer to ensure better learning of image features from the high-resolution original panoramic image.
[0103] The decoder can be a layout boundary map predictor. The input of the layout boundary map predictor is the image features output by the encoder, and the output is the Manhattan layout boundary. The Manhattan layout boundary can include a three-channel probability prediction of the wall-to-wall, ceiling-to-wall, and wall-to-floor boundaries in the original panoramic image, including visible boundaries and occlusion boundaries.
[0104] The decoder can include 7 nearest neighbor upsampling layers. The output of each nearest neighbor upsampling layer can be connected to a convolutional layer with a kernel size of 3x3. The last layer is a sigmoid, which can add skip connections to each convolutional layer to prevent the prediction results of the upsampling operation of the nearest neighbor upsampling layer from being shifted.
[0105] S302: Based on the masked layout boundary image, the mask image, and the masked panoramic image, perform layout boundary prediction to obtain the boundary layout map.
[0106] In some embodiments, the boundary layout map is obtained based on a pre-trained layout boundary prediction model; the layout boundary prediction model includes a downsampling convolutional layer, a transformer block, and a transposed upsampling convolutional layer connected in sequence. The input to the layout boundary prediction model is a masked layout boundary image, a mask image, and a masked panoramic image, and the output is the boundary layout map.
[0107] For example, the layout boundary prediction model can use Transformer as the backbone network so that even when the input image is at a low resolution, the backbone network Transformer can recover the occluded boundary layout map.
[0108] like Figure 4 As shown, the input to the layout boundary prediction model is the masked layout boundary image (Masked layoutLm), the mask image (MskM), and the masked panoramic image (Masked imageIm). The layout boundary prediction model consists of 3 layers (e.g., Figure 3 The diagram shows 3 (×3, similar descriptions will not be repeated below) downsampling convolutional layers (Conv-layers), 8 Transformer blocks, and 3 transposed convolutional upsampling convolutional layers (TConv-layers). To reduce the computational burden of attention learning, the mappings of the stitched masked layout boundary image, the mask image, and the masked panoramic image are fed into 3 convolutional downsampling layers for downsampling, and then fed into Transformer blocks to recover the downsampled features. Finally, 3 transposed convolutional upsampling convolutional layers are used to upsampling the recovered features to obtain the restored layout map (Restored layout Rm).
[0109] Within the Transformer block, axial attention and standard attention mechanisms can be used alternately to overcome the quadratic complexity problem of standard attention, and positional encoding is used in each axial attention block. Three transposed convolutional upsampling layers can upsample to a resolution of 512x256 to generate a complete boundary layout map.
[0110] S303: Perform structural feature extraction processing on the boundary layout map to obtain the layout boundary features.
[0111] In some embodiments, the layout boundary features are obtained based on a layout feature extraction model; the layout feature extraction model includes a downsampling gated convolutional layer, a dilated convolutional residual block, and an upsampling gated convolutional layer connected in sequence. The input to the layout feature extraction model is a boundary layout map, and the output is the layout boundary features.
[0112] like Figure 4 As shown, the layout feature extraction model consists of 3 downsampling gated convolutional layers (GateConvDownsample), 3 dilated convolutional residual blocks, and 3 upsampling gated convolutional layers (GateConvUpsample). The downsampling gated convolutional layers can be understood as encoders, and the upsampling gated convolutional layers as decoders. The gated convolutional layers can selectively pass useful features.
[0113] In particular, by combining the layout structure image generation model, the layout boundary prediction model, and the layout feature extraction model, the problem of layout structure distortion caused by two-level distortion of panoramic images of indoor scenes in related technologies can be solved. The layout boundary features of indoor scenes can be extracted for subsequent reconstruction of background images after removing objects in the same indoor scene, ensuring the generation of realistic layout structures of indoor scenes.
[0114] Accordingly, in some embodiments, such as Figure 4 As shown, we can call the model that has the corresponding functions of layout structure image generation model, layout boundary prediction model, and layout feature extraction model a structure restoration module (SRM). That is, the structure restoration module includes the layout structure image generation model, layout boundary prediction model, and layout feature extraction model connected in sequence.
[0115] When training the structural repair model, either a "whole-system training" approach or a "segmented training" approach can be used; this embodiment does not impose a limitation. "Whole-system training" can be understood as training the layout structure image generation model, layout boundary prediction model, and layout feature extraction model as a whole. "Segmented training" can be understood as training the layout structure image generation model, layout boundary prediction model, and layout feature extraction model separately.
[0116] S304: Generate layout features based on layout boundary features, mask image, and the masked panoramic image.
[0117] like Figure 4As shown, the masked image and the masked panoramic image can be concatenated, then fused with layout boundary features (such as through addition), and then downsampled to obtain the layout features. The downsampling can specifically employ methods such as... Figure 4 The three-layer downsampling layer shown is implemented.
[0118] S305: Based on the target object in the indoor scene, perform structured segmentation processing on the masked panoramic image to obtain a structured region map including the structured region corresponding to the target object.
[0119] In some embodiments, the structured region map is obtained by processing the masked panoramic image based on a pre-trained structured encoder; the structured encoder includes a skip-connected downsampling convolutional layer and an upsampling convolutional layer.
[0120] For example, if the target objects include walls, ceilings, and floors, the structure encoder can divide the interior scene into a structured region map that includes three structured regions (the structured region corresponding to the walls, the structured region corresponding to the ceiling, and the structured region corresponding to the floor).
[0121] like Figure 5 As shown, the input to the structure encoder is the masked panoramic image, and the output is a structure region map (S). The structure encoder consists of four downsampled convolutional layers with skip connections and four upsampled convolutional layers, which can perform activation operations using normalization and ReLU.
[0122] S306: Construct a style matrix based on the structural semantic information of the structured region map.
[0123] In some embodiments, the style matrix is obtained by processing a structured region map and a raw panoramic image based on a pre-trained semantic prior encoder; the semantic prior encoder includes a convolutional layer, a transposed convolutional layer, and an average pooling layer connected in sequence.
[0124] like Figure 5 As shown, the semantic prior encoder takes a structured region map and the original panoramic image as input, and outputs a 512x3 style matrix, where 3 represents the number of structured regions. Each column of the style matrix corresponds to the style code of the structural semantic information of a structured region. Specifically, as... Figure 5 As shown, the semantic prior encoder can include 4 convolutional layers, 4 transposed convolutional layers, and 1 region-wise average pooling layer. The average pooling layer can exclude irrelevant texture information from the original panoramic image.
[0125] S307: Fill the preset structured mask according to the style matrix to obtain the structured region texture features.
[0126] In some embodiments, S307 may include the following steps: The first step is to perform local feature extraction based on the style matrix, preset Gaussian noise, layout features, and structured mask to obtain the initial local texture.
[0127] The second step is to repair the initial local texture of the repaired area corresponding to the mask image based on the style matrix, and obtain the structured region texture features.
[0128] In some embodiments, the structured region texture features are generated based on a pre-trained residual network (SEANResNet) model, such as... Figure 5 As shown, the input to the residual network model is a style matrix, preset Gaussian noise, layout features, and a structured mask. The residual network model includes convolutional layers, which consist of a SEAN module, ReLU, and convolutional kernel connected in sequence.
[0129] For example, the input to the residual network model is a style matrix transformed by a 1x1 convolutional layer, pre-defined Gaussian noise, layout features, and a structured mask; the output is structured region texture features. The residual network model includes three convolutional layers, each containing a SEAN module (such as...). Figure 5 The SEAN shown), ReLU, and a 3x3 convolutional kernel (as shown) Figure 5 (As shown in the 3x3Conv).
[0130] Similarly, in this embodiment, by combining the structural encoder, semantic prior encoder, and residual network model, the problem of generating unrealistic texture features of structured regions in indoor scenes can be solved. This is achieved by combining structural semantic information to extract highly realistic texture features of structured regions. Correspondingly, as... Figure 4 As shown, we can call a model with the functions of a structural encoder, a semantic prior encoder, and a residual network model a structured region texture extraction module (SRTE-M).
[0131] S308: Perform panoramic image inpainting based on layout features and structured region texture features to obtain a reduced-reality prediction image corresponding to the masked panoramic image.
[0132] In some embodiments, S308 may include the following steps: The first step is to perform convolution processing on the layout features to obtain the first convolutional layout features.
[0133] The second step is to fuse the first convolutional layout features and the structured region texture features to obtain combined features.
[0134] The third step is to perform convolution processing on the layout features to obtain the second convolution layout features, and then perform global feature extraction processing on the layout features to obtain the global features.
[0135] Fourth step: Fuse the second convolutional layout features and global features to obtain the frequency domain layout features.
[0136] Step 5: Fuse the combined features and frequency domain layout features to obtain the predicted image.
[0137] In some embodiments, such as Figure 4 As shown, the predicted image is obtained by processing layout features and structured region texture features based on a pre-trained Fourier convolutional fusion model. The Fourier convolutional fusion model can include downsampling convolutional layers, Fourier convolutional fusion layers, upsampling convolutional layers, spectral transform blocks, and fusion modules.
[0138] In other embodiments, the downsampling convolutional layer is a layer connected to the output of a Fourier convolutional fusion model. For example, as... Figure 4 As shown, the output of the Fourier convolutional fusion model is connected to the downsampling convolutional layer, which has three layers. The output of the downsampling convolutional layer is the predicted image.
[0139] For example, the Fourier convolution fusion model can be the fast fourier convolution fusion (FFCF) model, which includes 3 downsampling convolutional layers, 9 (fast) Fourier convolution fusion layers, and 3 upsampling convolutional layers.
[0140] In some embodiments, such as Figure 6 As shown, the Fourier convolution fusion model includes a convolutional layer, a Fourier convolution fusion layer, a spectral transform block, a normalization and ReLU layer, and a fusion module.
[0141] Combining the first step above and Figure 6 The layout features can be input into the convolutional layer to obtain the first convolutional layout features.
[0142] Combining the second step above and Figure 6 It can fuse the first convolutional layout features and structured region texture features based on the Fourier convolutional fusion layer, and on this basis, activation can be performed in the normalization and ReLU layers to obtain combined features.
[0143] Combining the third step above and Figure 6 The layout features can be input into the convolutional layer and the spectral transform block respectively, and the outputs of the convolutional layer and the spectral transform block can be fused by the Fourier convolutional fusion layer. Furthermore, activation can be performed in the normalization and ReLU layers to obtain the frequency domain layout features.
[0144] Combining the fifth step above and Figure 6 The combined features and frequency domain layout features can be input into the fusion module, which includes one cascaded layer and three convolutional layers. The output of the fusion module is upsampled to generate the restored output of the masked panoramic image, thus obtaining... Figure 4 The predicted image shown.
[0145] Texture restoration is achieved through the interaction of a Fourier convolution fusion model and a structured region texture extraction model. Therefore, as... Figure 4 As shown, a model that has the functions of both Fourier convolution fusion model and structured region texture extraction model can be called an inpainting network model.
[0146] like Figure 7 As shown, the first column contains two different panoramic images of an indoor scene, the second column contains a perspective image corresponding to the panoramic image in the first column, the third column contains a predicted image obtained using a solution in related technologies, the fourth column contains a predicted image obtained using a solution in this embodiment, and the fifth column contains the ground truth value corresponding to the panoramic image in the first column.
[0147] Combination Figure 7 It can be seen that, relatively speaking, the predicted image obtained by adopting the scheme provided in the embodiments of this disclosure can better match the true value and has a higher ability to restore the true value, that is, it has higher accuracy and reliability.
[0148] Based on the above analysis, predicted images can be obtained using deep learning. However, deep learning requires pre-built models, such as the structure restoration model, structured region texture extraction model, and Fourier convolution fusion model described in the previous embodiment. Furthermore, model building requires sample data. In this embodiment, the sample data can be a pre-built structured panoramic image dataset. For example, we can construct a structured panoramic image dataset (SD) based on the structured 3D dataset in related technologies.
[0149] For example, the structured panoramic image dataset includes multiple sets (e.g., 14,528 sets) of panoramic images of indoor scenes. Each set of panoramic images includes panoramic images before and after the object to be removed, a mask image of the object to be removed, a boundary layout map of the panoramic image, and a Manhattan layout boundary, with a resolution of 1024x512.
[0150] In order to determine the object to be removed, the semantic labels of each set of indoor scenes can be used to randomly select the target edge composed of the largest connected component of the available foreground and fill it to represent the object to be removed, which is called a mask.
[0151] The Structured3D dataset includes multiple sets of panoramic images corresponding to empty and full indoor scenes. For each set of panoramic images, the region of the object to be removed in the panoramic image of the full indoor scene can be replaced with the background image of the corresponding location in the empty indoor scene to construct a ground truth (GT) image for panoramic image reduction. An empty indoor scene refers to an indoor scene without the object to be removed, while a full indoor scene refers to an indoor scene including the object to be removed.
[0152] In the Structured3D dataset, the panoramic image of an empty indoor scene is generated by extracting the foreground objects from the panoramic image of a full indoor scene and subtracting the objects to be removed to generate a foreground image without the objects to be removed. The foreground image is then added to the panoramic image of the empty indoor scene to avoid the drawback that there will be lighting differences between the replacement area corresponding to the object to be removed and the original area because the image is rendered based on physical lighting.
[0153] The Structured3D dataset includes the locations of connection points in a structured layout. The boundaries of the structured layout are reconstructed on the panoramic image based on these connection points to update the aforementioned baseline map. Furthermore, different regions of the structured layout can be filled with different colors (e.g., red for the ceiling, blue for the walls, and green for the floor) to obtain the final baseline map.
[0154] Correspondingly, taking the training of a layout structure image generation model as an example, the panoramic image in front of the object to be removed in the structured panoramic image dataset can be used as the image for predicting the Manhattan layout boundary, and the Manhattan layout boundary in the structured panoramic image dataset can be used as the ground truth. By calculating the binary cross-entropy error between the predicted value (predicted pixel probability) and the ground truth (pixel probability in the Manhattan layout boundary), the binary cross-entropy error can be used as the loss to achieve better training results.
[0155] Taking the training of a layout boundary prediction model as an example, the mask image of the object to be removed, the panoramic image of the masked image in front of the object to be removed, and the Manhattan layout boundary of the masked image can be used as images for predicting the boundary layout map in the structured panoramic image dataset. The boundary layout map in the structured panoramic image dataset can be used as the ground truth to train the layout boundary prediction model.
[0156] Taking the training of the Fourier convolutional fusion model as an example, the output of the layout boundary prediction model can be used as part of the input of the Fourier convolutional fusion model. The panoramic image before the object to be removed and the mask image in the structured panoramic image dataset are combined to obtain the prediction image. The panoramic image after removing the object to be removed is used as the ground truth of the prediction image to train the Fourier convolutional fusion model.
[0157] Based on the above analysis, it can be seen that the restoration network model can include a Fourier convolution fusion model and a structured region texture extraction model, and the restoration network model can be trained on a structured panoramic image dataset.
[0158] For example, the predicted image is generated based on a pre-trained insulation network model, whose inputs are a boundary layout map, a mask image, and a masked panoramic image.
[0159] In some embodiments, the repair network model is trained based on a fusion loss function, which is obtained by fusing the absolute error (L1) loss function, the adversarial loss function, and the advanced synthetic perception loss function.
[0160] The L1 loss function represents the difference between the predicted image and the ground truth. The adversarial loss function is obtained by inputting the predicted image and the ground truth into a predefined generator and discriminator, respectively, and is used to represent the difference between the predicted image and the ground truth.
[0161] For example, the L1 loss function It can be determined based on Equation 1, Equation 1: in, A mask between 0 and 1. To predict the image, This is to predict the ground truth image corresponding to the image (such as the ground truth image corresponding to the predicted image in the baseline image of the structured panoramic image dataset mentioned above). For multiplication operations, The 1 in the diagram represents the obscured area.
[0162] Advanced Synthesis-Aware Loss Function It can be determined based on Equation 2, Equation 2: in, , , For average operation, For network activation layer, For structured region features, N represents the total number of feature elements in the feature map. For matrix functions, and It is a preset feature set. and These are preset coefficients, which can be determined based on demand, historical records, and experiments, such as... It can be 0.12. It can be 40.0.
[0163] Adversarial loss function It can be determined based on Equation 3, Equation 3: in, , , For discriminator loss, For generator loss, Let G be the feature matching loss, G be the generator, and D be the discriminator.
[0164] Fusion loss function It can be determined based on Equation 4, Equation 4: Similarly, , , These are preset coefficients, which can be determined based on demand, historical records, and experiments, such as... It is 10.0. It is 10.0. It is 30.0.
[0165] Based on the above analysis, we can construct a structured panoramic image dataset on the basis of the Structured3D dataset in related technologies. We can divide the structured panoramic image dataset into three parts: training set, validation set, and test set. We train the repair network model on the training set, evaluate the repair network model on the validation set, and once the optimal parameters are found, we test it once on the test set. The error on the test set is used as an approximation of the generalization error.
[0166] When evaluating the repair network model on the validation set, one or more metrics can be used, such as Mean Absolute Error (MAE), Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS).
[0167] Figure 8 This document presents the evaluation and comparison results of the restoration network model disclosed herein with five image restoration models in related technologies under the aforementioned indicators. Figure 8 As shown, the five image restoration models in the related technologies include two-dimensional image restoration models and three-dimensional image restoration models, such as PanoDR. The two-dimensional image restoration models specifically include causal-based time series domain generalization (CTSDG) models, ZITS, Latent Diffusion Models (LDMs), and LaMa.
[0168] Combination Figure 8 It can be seen that LaMa outperforms PanoDR compared to 2D image inpainting models, possibly because PanoDR does not smooth the boundaries of the filled regions. In terms of PSNR, LaMa performs best. Regarding other metrics, the inpainting network model of this disclosure achieves a higher LPIPS score, indicating that the restored panoramic image (i.e., the predicted image) is closer to the true value. Furthermore, the inpainting network model of this disclosure achieves better performance on SSIM and MAE metrics. This demonstrates that the method of this disclosure can better recover the structural and textural information of the removed regions in panoramic images.
[0169] Based on the above analysis, the method disclosed herein can be implemented using SRM and SRTE-M. Figure 8 The study also included the results of ablation experiments on SRM and SRTE-M to evaluate the effectiveness of each.
[0170] Combination Figure 8 It can be seen that, relatively speaking, SRM performs better than SRTE-M. This is because SRTE-M takes structured regions as input, providing both local texture information and local structural information, thereby recovering the local information of the masked areas in the panoramic image.
[0171] The disclosed repair network model can be run on a graphics processing unit (GPU) using an optimizer with a learning rate of 6e-4, a 1000-step warmup, and cosine decay. The structural encoder used in SRTE-M is optimized using the optimizer's default parameters with a learning rate of 0.0001 and a batch size of 4. The Fourier convolutional fusion model is trained using the optimizer with a learning rate of 1e-3 for the generator and 1e-4 for the discriminator. The input panoramic image resolution is 512×256. The Fourier convolutional fusion model uses the same weights to initialize the structural encoder's weights, while the weights of other models are initialized to 0 and normally distributed with a value of 0.02.
[0172] Furthermore, to better demonstrate the advantages of the repair network model of this disclosure, we have conducted a qualitative comparison of the method of this disclosure with other methods, and the comparison results can be found in [reference needed]. Figure 9 .
[0173] like Figure 9 As shown, the method disclosed in this disclosure recovers the Manhattan structure of panoramic images more accurately than other methods, and the generated textures are more consistent with the true ground values. Although PanoDR has obvious texture stitching artifacts, it can recover the structural information of panoramic images relatively well. In addition, the texture of the panoramic image recovered by LaMa is smoother, but there are still some deviations in the recovery of indoor structures. Compared with other methods, the method disclosed in this disclosure combines the layout boundary information of panoramic images and extracts structured region information, which is conducive to more realistic and accurate image restoration.
[0174] Based on the above technical concept, this disclosure provides a device for panoramic image reduction of reality, which can be applied to indoor scenes.
[0175] Please see Figure 10 , Figure 10 A schematic diagram of a panoramic image reduction reality apparatus according to an embodiment of this disclosure, as shown below. Figure 10 As shown, the device 1000 includes: The first generation unit 1001 is used to generate layout features based on the acquired masked layout boundary image, mask image, and masked panoramic image, wherein the layout features characterize the structural features of the original panoramic image at the layout level.
[0176] The second generation unit 1002 is used to generate a style matrix corresponding to the structured region of the indoor scene based on the acquired masked panoramic image and the original panoramic image, wherein the style matrix represents the structural semantic information corresponding to the structured region.
[0177] The filling unit 1003 is used to fill the preset structured mask according to the style matrix to obtain the structured region texture features.
[0178] Repair unit 1004 is used to perform panoramic image repair processing based on the layout features and the structured region texture features to obtain a reduced-reality prediction image corresponding to the masked panoramic image.
[0179] In some embodiments, the first generating unit 1001 includes: The prediction subunit is used to predict the layout boundary based on the masked layout boundary image, the mask image, and the masked panoramic image to obtain a boundary layout map. Extract sub-units to perform structural feature extraction processing on the boundary layout diagram to obtain layout boundary features; A generation subunit is used to generate the layout features based on the layout boundary features, the masked layout boundary image, and the masked panoramic image.
[0180] In some embodiments, the boundary layout map is obtained based on a pre-trained layout boundary prediction model; the layout boundary prediction model includes a downsampling convolutional layer, a transformer block, and a transposed upsampling convolutional layer connected in sequence.
[0181] In some embodiments, the layout boundary features are obtained based on a layout feature extraction model; the layout feature extraction model includes a downsampling gated convolutional layer, a dilated convolutional residual block, and an upsampling gated convolutional layer connected in sequence.
[0182] In some embodiments, the masked layout boundary image is obtained by predicting the Manhattan layout boundary of the target object in the original panoramic image and then masking the Manhattan layout boundary. The target objects include walls, ceilings, and floors.
[0183] In some embodiments, the Manhattan layout boundary is determined based on a pre-trained layout structure image generation model, which includes an encoder and a decoder connected in sequence. The encoder takes the original panoramic image as input and the decoder outputs the Manhattan layout boundary.
[0184] In some embodiments, the encoder includes a convolutional layer, and a noise linear rectified function and a pooling layer respectively connected to the output of the convolutional layer; The decoder includes an upsampling layer, and a convolutional layer and an activation layer that are sequentially connected to the output of the upsampling layer.
[0185] In some embodiments, the second generating unit 1002 includes: The segmentation subunit is used to perform structured segmentation processing on the masked panoramic image according to the target object, so as to obtain a structured region map including the structured region corresponding to the target object; A sub-unit is constructed to build the style matrix based on the structural semantic information of the structured region map.
[0186] In some embodiments, the structured region map is obtained by processing the masked panoramic image based on a pre-trained structured encoder; the structured encoder includes a skip-connected downsampling convolutional layer and an upsampling convolutional layer.
[0187] In some embodiments, the style matrix is obtained by processing the structured region map and the original panoramic image based on a pre-trained semantic prior encoder; the semantic prior encoder includes a convolutional layer, a transposed convolutional layer, and an average pooling layer connected in sequence.
[0188] In some embodiments, the filling unit 1003 includes: The first processing subunit is used to perform local feature extraction processing based on the style matrix, preset Gaussian noise, the layout features, and the structured mask to obtain an initial local texture; The repair subunit is used to repair the initial local texture of the repair region corresponding to the mask image according to the style matrix, so as to obtain the structured region texture features.
[0189] In some embodiments, the structured region texture features are generated based on a pre-trained residual network model. The inputs of the residual network model are the style matrix, preset Gaussian noise, the layout features, and the structured mask. The residual network model includes convolutional layers, and each convolutional layer of the residual network model includes a building block, a noise linear rectified function, and a convolutional kernel connected in sequence.
[0190] In some embodiments, the repair unit 1004 includes: A convolutional subunit is used to perform convolution processing on the layout features to obtain a first convolutional layout feature; The first fusion subunit is used to fuse the first convolutional layout feature and the structured region texture feature to obtain a combined feature; The second processing subunit is used to perform convolution processing on the layout features to obtain second convolution layout features, and to perform global feature extraction processing on the layout features to obtain global features. The second fusion subunit is used to fuse the second convolutional layout feature and the global feature to obtain the frequency domain layout feature; The third fusion subunit is used to fuse the combined features and the frequency domain layout features to obtain the predicted image.
[0191] In some embodiments, the predicted image is obtained by processing the layout features and the structured region texture features based on a pre-trained Fourier convolutional fusion model; the Fourier convolutional fusion model includes a downsampling convolutional layer, a Fourier convolutional fusion layer, an upsampling convolutional layer, a spectral transform block, and a fusion module.
[0192] In some embodiments, the predicted image is generated based on a pre-trained insulation network model, the input of which is the boundary layout map, the mask image, and the masked panoramic image; The repair network model is trained based on a fusion loss function, which is obtained by fusing the absolute error loss function, the adversarial loss function, and the advanced synthetic perception loss function.
[0193] It should be noted that the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0194] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0195] According to embodiments of this disclosure, this disclosure also provides a computer program product comprising: a computer program stored in a readable storage medium, at least one processor of an electronic device being able to read the computer program from the readable storage medium, and the at least one processor executing the computer program causing the electronic device to perform the scheme provided in any of the above embodiments.
[0196] Figure 11 A schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0197] like Figure 11As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1102 or a computer program loaded into random access memory (RAM) 1103 from storage unit 1108. The RAM 1103 may also store various programs and data required for the operation of device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.
[0198] Multiple components in device 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of monitors, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0199] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as panoramic image reduction of reality methods. For example, in some embodiments, the panoramic image reduction of reality method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by computing unit 1101, one or more steps of the panoramic image reduction of reality method described above can be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured by any other suitable means (e.g., by means of firmware) to perform a method of panoramic image reduction of reality.
[0200] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0201] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0202] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0203] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0204] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0205] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0206] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0207] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-executable instructions. These computer-executable instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0208] These processor-executable instructions may also be stored in a processor-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the processor-readable memory produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0209] These processors can execute instructions that can also be loaded onto a computer or other programmable data processing device, causing a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0210] Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from its spirit and scope. Therefore, if such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, this disclosure is also intended to include such modifications and variations.
Claims
1. A method for reducing reality from panoramic images, applied to indoor scenes, characterized in that, The method includes: Based on the obtained masked layout boundary image, mask image, and masked panoramic image, layout features are generated, wherein the layout features characterize the structural features of the original panoramic image at the layout level. Based on the acquired masked panoramic image and the original panoramic image, a style matrix corresponding to the structured regions of the indoor scene is generated, wherein the style matrix represents the structural semantic information corresponding to the structured regions. The preset structured mask is filled according to the style matrix to obtain structured region texture features; Based on the layout features and the structured region texture features, panoramic image restoration processing is performed to obtain a reduced-reality prediction image corresponding to the masked panoramic image. Based on the layout features and the structured region texture features, panoramic image inpainting is performed to obtain a reduced-reality predicted image corresponding to the masked panoramic image, including: The layout features are convolved to obtain the first convolutional layout features; The first convolutional layout feature and the structured region texture feature are fused to obtain a combined feature; The layout features are convolved to obtain second convolutional layout features, and global feature extraction is performed on the layout features to obtain global features. The second convolutional layout feature and the global feature are fused to obtain the frequency domain layout feature; The combined features and the frequency domain layout features are fused to obtain the predicted image.
2. The method according to claim 1, characterized in that, Based on the acquired masked layout boundary image, mask image, and masked panoramic image, layout features are generated, including: Based on the masked layout boundary image, the mask image, and the masked panoramic image, layout boundary prediction is performed to obtain a boundary layout map. The boundary layout diagram is subjected to structural feature extraction processing to obtain the layout boundary features; The layout features are generated based on the layout boundary features, the mask image, and the masked panoramic image.
3. The method according to claim 1, characterized in that, The masked layout boundary image is obtained by predicting the Manhattan layout boundary of the target object in the original panoramic image and then masking the Manhattan layout boundary. The target objects include walls, ceilings, and floors.
4. The method according to claim 3, characterized in that, Based on the acquired masked panoramic image and the original panoramic image, a style matrix corresponding to the structured regions of the indoor scene is generated, including: Based on the target object, the masked panoramic image is subjected to structured segmentation processing to obtain a structured region map including the structured region corresponding to the target object; The style matrix is constructed based on the structural semantic information of the structured region map.
5. The method according to claim 1, characterized in that, The preset structured mask is filled according to the style matrix to obtain structured region texture features, including: Based on the style matrix, preset Gaussian noise, layout features, and structured mask, local feature extraction is performed to obtain an initial local texture; The initial local texture of the repaired region corresponding to the mask image is repaired according to the style matrix to obtain the structured region texture features.
6. The method according to claim 2, characterized in that, The predicted image is generated based on a pre-trained repair network model, the input of which is the boundary layout map, the mask image, and the masked panoramic image; The repair network model is trained based on a fusion loss function, which is obtained by fusing the absolute error loss function, the adversarial loss function, and the advanced synthetic perception loss function.
7. A panoramic image reduction reality device, applied to an indoor scene, characterized in that, The method for reducing reality from panoramic images as described in claim 1 includes: The first generation unit is used to generate layout features based on the acquired masked layout boundary image, mask image, and masked panoramic image, wherein the layout features characterize the structural features of the original panoramic image at the layout level. The second generation unit is used to generate a style matrix corresponding to the structured region of the indoor scene based on the acquired masked panoramic image and the original panoramic image, wherein the style matrix represents the structural semantic information corresponding to the structured region. A filling unit is used to fill a preset structured mask according to the style matrix to obtain structured region texture features; The repair unit is used to perform panoramic image repair processing based on the layout features and the structured region texture features to obtain a reduced-reality prediction image corresponding to the masked panoramic image.
8. The apparatus according to claim 7, characterized in that, The masked layout boundary image is obtained by predicting the Manhattan layout boundary from the target object in the original panoramic image and then masking the Manhattan layout boundary. The target object includes walls, ceilings, and floors.
9. A processor-readable storage medium, characterized in that, The processor-readable storage medium stores a computer program for causing the processor to perform the method of any one of claims 1 to 6.
10. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Diminished and mediated reality effects from reconstruction
US20140321702A1
Image processing method, image processing device, and electronic apparatus
US20150245007A1