Method for generating 3-dimensional mesh based on ERP image and computing device using same

The method addresses the challenges of generating accurate 3D meshes from 360-degree videos by converting ERP images into perspective images, removing noise, and using a mesh generation model to create high-quality 3D meshes.

WO2025110438A1PCT designated stage expired Publication Date: 2025-05-30NAVER CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/013621
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-09-09
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing methods for generating 3D meshes from 360-degree videos face challenges due to severe distortion and the inability to distinguish mirrors or glass, leading to inaccurate and noisy 3D models.

Method used

A method that converts an ERP image into perspective images, removes noise using inpainting techniques for reflective surfaces, and generates accurate depth and normal maps to create a high-quality 3D mesh, utilizing a mesh generation model that learns optimal hyperparameters.

Benefits of technology

The method effectively generates accurate and noise-free 3D meshes from ERP images, improving the quality of 3D modeling for indoor spaces by addressing distortion and reflective surface issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024013621_30052025_PF_FP_ABST
    Figure KR2024013621_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method for generating a 3-dimensional mesh based on an equirectangular project (ERP) image and a computing device using same, and the method for generating a 3-dimensional mesh based on an ERP image, using a computing device, according to an embodiment of the present invention, may include the steps of: generating a plurality of perspective images by converting an ERP image for a target environment; if a perspective image including a specific material area exists among the plurality of perspective images, removing the specific material area via an inpainting technique; generating a depth map and a normal map corresponding to the ERP image by using the plurality of perspective images; and generating a 3-dimensional mesh corresponding to the depth map and the normal map by using a mesh generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Method for generating 3D mesh based on ERP image and computing device using the same

[0001] The present invention relates to a method for generating a three-dimensional mesh from an ERP image captured using a 360-degree monocular camera and a computing device using the same.

[0002] 360-degree video can refer to video or image content that is captured or played back simultaneously in all directions (360 degrees). For example, 360-degree video can be displayed on a three-dimensional spherical surface. 360-degree video can be created by capturing images or videos from multiple viewpoints through one or more cameras, connecting the captured images to create a single panoramic or spherical image, and projecting it onto a 2D picture.

[0003] Here, there may be a case where you want to capture an indoor space in 360-degree video and create a 3D model of the space based on that video. However, when performing 3D modeling based on 360-degree video, the distortion within the 360-degree video is severe, making it difficult to apply deep learning models for typical depth estimation.

[0004] Additionally, indoor spaces may contain mirrors or glass, but because mirrors or glass within 360-degree video cannot be distinguished, shapes appearing on them may be reflected within the 3D model. This can lead to issues such as creating 3D models of spaces that do not actually exist.

[0005] The present invention aims to provide a method for generating a 3D mesh based on an ERP image, which can generate an accurate 3D mesh from an ERP image captured using a 360-degree monocular camera, and a computing device using the same.

[0006] The present invention aims to provide a method for generating a 3D mesh based on an ERP image, which can remove noise or errors on a 3D mesh caused by a mirror or glass included in an ERP image, and a computing device using the same.

[0007] The present invention provides a 3D mesh generation method based on an ERP image, which utilizes a mesh generation model that performs learning by scheduling each hyperparameter, and a computing device that uses the same.

[0008] The present invention aims to provide a method for generating a 3D mesh based on an ERP image, which can remove noise by simplifying a plane within the generated 3D mesh, and a computing device using the same.

[0009] According to one embodiment of the present invention, a method for generating a three-dimensional mesh based on an ERP (Equirectangular Project) image using a computing device may include the steps of: generating a plurality of perspective images by converting an ERP image of a target environment; removing a perspective image including a specific material area from among the plurality of perspective images using an inpainting technique, if the perspective image includes a specific material area among the plurality of perspective images; generating a depth map and a normal map corresponding to the ERP image using the plurality of perspective images; and generating a three-dimensional mesh corresponding to the depth map and the normal map using a mesh generation model.

[0010] A computing device for generating a three-dimensional mesh based on an ERP (Equirectangular Project) image according to one embodiment of the present invention includes a processor, wherein the processor may perform the following operations: generating a plurality of perspective images by converting an ERP image for a target environment; removing a perspective image including a specific material area from among the plurality of perspective images using an inpainting technique, if such a perspective image exists; generating a depth map and a normal map corresponding to the ERP image using the plurality of perspective images; and generating a three-dimensional mesh corresponding to the depth map and the normal map using a mesh generation model.

[0011] Additionally, the solutions to the aforementioned problems do not enumerate all features of the present invention. The various features of the present invention, along with their corresponding advantages and effects, can be understood in more detail by referring to the specific embodiments below.

[0012] According to a method for generating a 3D mesh based on an ERP image and a computing device using the same according to one embodiment of the present invention, noise or errors on a 3D mesh caused by mirrors or glass included in an ERP image can be removed, so that an accurate 3D mesh for a target environment can be generated.

[0013] According to a 3D mesh generation method based on an ERP image according to one embodiment of the present invention and a computing device using the same, learning can be performed by scheduling each hyperparameter when learning a mesh generation model, so it is possible to implement a mesh generation model capable of generating a high-performance 3D mesh without having to find optimal hyperparameters.

[0014] According to a method for generating a 3D mesh based on an ERP image according to one embodiment of the present invention and a computing device using the same, it is possible to remove noise within the 3D mesh and simplify the plane through post-processing of the generated 3D mesh, thereby providing a high-quality 3D mesh.

[0015] However, the effects that can be achieved by the ERP image-based 3D mesh generation method and the computing device using the same according to embodiments of the present invention are not limited to those mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary skill in the art to which the present invention pertains from the description below.

[0016] Figure 1 is an exemplary diagram showing ERP image generation using a 360-degree monocular camera according to one embodiment of the present invention.

[0017] Figure 2 is a schematic diagram showing a three-dimensional mesh generation device according to one embodiment of the present invention.

[0018] Figure 3 is an exemplary diagram showing an ERP image and a cube map image according to one embodiment of the present invention.

[0019] Figure 4 is an exemplary diagram showing a perspective image and a masking image according to one embodiment of the present invention.

[0020] Figure 5 is a schematic diagram showing inpainting for a masking area according to one embodiment of the present invention.

[0021] Figure 6 is an exemplary diagram showing projection of points located within a set range from a plane onto the plane according to one embodiment of the present invention.

[0022] Figure 7 is a block diagram showing a computing device according to one embodiment of the present invention.

[0023] Figures 8 and 9 are flowcharts showing a method for generating a 3D mesh based on an ERP image according to one embodiment of the present invention.

[0024] Hereinafter, embodiments disclosed in the present specification will be described in detail with reference to the attached drawings. Regardless of the drawing numbers, identical or similar components will be given the same reference numbers, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are given or used interchangeably only for the convenience of writing the specification, and do not have distinct meanings or roles in themselves. That is, the term "part" used in the present invention means a hardware component such as software, FPGA, or ASIC, and the "part" performs certain roles. However, the "part" is not limited to software or hardware. The "part" may be configured to be on an addressable storage medium, or may be configured to reproduce one or more processors. Thus, as an example, a 'part' may include components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. The functionality provided within the components and 'parts' may be combined into a smaller number of components and 'parts' or further separated into additional components and 'parts'.

[0025] In addition, when describing the embodiments disclosed in this specification, if it is determined that a detailed description of a related known technology may obscure the gist of the embodiments disclosed in this specification, the detailed description thereof will be omitted. In addition, the attached drawings are only intended to facilitate easy understanding of the embodiments disclosed in this specification, and the technical ideas disclosed in this specification are not limited by the attached drawings, and should be understood to include all modifications, equivalents, and substitutes included in the spirit and technical scope of the present invention.

[0026] Figure 1 is an exemplary diagram showing the creation of an ERP (Equirectangular Projection) image using a 360-degree monocular camera (Omnidirectional Camera) according to one embodiment of the present invention.

[0027] Referring to Fig. 1, an ERP image may be a photograph taken using a 360-degree monocular camera (1), and, depending on the embodiment, an image captured from a video taken using the 360-degree monocular camera (1) may also be utilized as an ERP image. In this case, the 360-degree monocular camera (1) may also generate and provide information regarding the camera height, etc., when taking the ERP image.

[0028] Here, the target environment (T) to be photographed using a 360-degree monocular camera (1) may be a specific object or space, and in some embodiments, the target environment (T) may be an indoor space of a building for use in real estate transactions, etc.

[0029] The 360-degree monocular camera (1) can be configured independently, but depending on the embodiment, it can also be implemented by being combined with various terminal devices such as a smartphone, tablet PC, PDA (Personal Digital Assistant), laptop computer, wearable device, etc. In other words, ERP images of the target environment (T) can be created using the 360-degree monocular camera (1) equipped on one's own smartphone, etc.

[0030] Thereafter, the user can request the generation of a 3D mesh corresponding to the captured ERP image by transmitting the captured ERP image to a 3D mesh generation device via a wired or wireless network. In other words, the user can request the generation of a 3D mesh from the 3D mesh generation device in order to generate a 3D model corresponding to the target environment based on the 3D mesh.

[0031] At this time, the user can connect to the network using his / her terminal device and communicate with the 3D mesh generation device through the network. That is, the user can use the terminal device to transmit ERP images, etc. captured by a 360-degree monocular camera (1) to the 3D mesh generation device, and the 3D mesh generation device can generate a 3D mesh for the target environment based on the received ERP image. Thereafter, based on the 3D mesh, it is also possible to generate a 3D model, etc. for the target environment.

[0032] Here, the communication method between the terminal device and the 3D mesh generation device is not limited, and may include not only a communication method utilizing a communication network that the network may include (for example, a mobile communication network, wired Internet, wireless Internet, broadcasting network, satellite network, etc.), but also short-range wireless communication between devices. For example, the network may include any one or more of a personal area network (PAN), a local area network (LAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a broadband network (BBN), and the Internet. In addition, the network may include any one or more of a network topology including, but not limited to, a bus network, a star network, a ring network, a mesh network, a star-bus network, a tree, or a hierarchical network.

[0033] Additionally, although it is illustrated here that an ERP image is created using a 360-degree monocular camera (1), it is also possible to create an ERP image using a panorama mode of a general camera, etc., depending on the embodiment.

[0034] Meanwhile, depending on the embodiment, there may be a case where a three-dimensional model of an indoor space is desired to be created from an ERP image captured by a 360-degree monocular camera (1). Here, in order to perform a three-dimensional model based on the ERP image, it is necessary to estimate the depth or normal within the ERP image. However, since the ERP image is severely distorted compared to a general perspective image, it may be difficult to obtain accurate depth or normal estimation results using a general depth estimation model based on a perspective image.

[0035] Furthermore, indoor spaces may contain mirrors and glass, but 3D mesh generators cannot distinguish between mirrors and glass within ERP images. Therefore, depth and normal maps may be generated by reflecting the shapes appearing on the mirrors and glass. This creates depth and normal maps for spaces that do not actually exist, potentially leading to issues such as incorrect 3D mesh generation.

[0036] In addition, the 3D mesh generation device may further include a mesh generation model for generating a 3D mesh, and the mesh generation model may be trained to generate a 3D mesh using a loss function including depth loss and normal loss corresponding to the difference between the depth map and normal map of the generated 3D mesh and the depth map and normal map of the ERP image. At this time, the performance of the 3D mesh generated by the mesh generation model may vary depending on each hyperparameter set for the depth loss and normal loss. That is, when training the mesh generation model, the setting of each hyperparameter is important, but there is a problem that it is difficult to set the optimized hyperparameters each time according to the characteristics of each ERP image, etc.

[0037] Accordingly, a 3D mesh generation device according to an embodiment of the present invention can convert an ERP image into a perspective image, generate a depth map and a normal map corresponding to the perspective image, and generate a 3D mesh based on the same. In addition, by preprocessing the ERP image or the perspective image, it is possible to prevent the creation of a space that does not actually exist by inpainting in advance for a specific material area such as a mirror or glass. In the case of hyperparameters, the 3D mesh generation device according to an embodiment of the present invention can learn a mesh generation model capable of generating a high-performance 3D mesh without having to find the optimal hyperparameter by performing learning while scheduling each hyperparameter.

[0038] Hereinafter, a three-dimensional mesh generation device according to one embodiment of the present invention will be described with reference to FIG. 2.

[0039] Figure 2 is a schematic diagram showing a three-dimensional mesh generation device according to one embodiment of the present invention.

[0040] Referring to FIG. 2, a 3D mesh generation device (100) according to one embodiment of the present invention may include a receiving unit (110), an image conversion unit (120), a preprocessing unit (130), a depth and normal estimation unit (140), a mesh generation unit (150), and a postprocessing unit (160).

[0041] The receiving unit (110) can receive ERP images via a wired or wireless communication network. That is, the receiving unit (110) can receive ERP images directly from a 360-degree monocular camera (1) that generates ERP images or a user's terminal device, and depending on the embodiment, it is also possible to receive ERP images via a separate server or storage device that stores ERP images generated by the 360-degree monocular camera.

[0042] The image conversion unit (120) can convert an ERP image for a target environment to generate a plurality of perspective images. That is, as illustrated in FIG. 3(a), an ERP image on a spherical domain can be input, and the image conversion unit (120) can convert the ERP image to generate a cubemap image on a corresponding cubemap domain, as illustrated in FIG. 3(b). Referring to FIG. 3(b), the cubemap image can be generated in a form corresponding to a development diagram of a rectangular parallelepiped, and each of the front, back, left side, right side, top side, and bottom side corresponding to the target environment (T) can be displayed in a perspective form. That is, the areas corresponding to the front, back, left side, right side, top side, and bottom side in the ERP image can be distinguished, and the areas can be converted into a perspective form to generate a cubemap image. Here, the image conversion unit (120) can generate the cubemap image by utilizing various algorithms for converting from an ERP image to a cubemap image. Thereafter, the image conversion unit (120) can generate a perspective image by dividing the six perspective images corresponding to the six faces included in the cube map image. That is, six perspective images corresponding to the front, back, left side, right side, top side, and bottom side can be divided and generated, respectively.

[0043] The preprocessing unit (130) can remove the specific material area using an inpainting technique if a perspective image including a specific material area exists among a plurality of perspective images. Here, the specific material area includes a surface that is transparent or reflective, such as a mirror, glass, or a metal surface. However, the present invention is not limited thereto, and depending on the embodiment, a user or the like can define a required area as a specific material area and utilize it. In addition, although the preprocessing unit (130) is shown here to remove the specific material area included in the perspective image, depending on the embodiment, the preprocessing unit (130) can also remove the specific material area in advance from the ERP image.

[0044] The preprocessing unit (130) can regenerate pixel values ​​of a specific material area using the inpainting technique, and at this time, the preprocessing unit (130) can regenerate the specific material area so that it appears to be made of the same material as the adjacent area. Since the specific material area reflects or projects light to appear as if another space exists inside, it is necessary to remove characteristics such as light reflection or projection of the specific material area before estimating a depth map or normal map based on a perspective image. Accordingly, the preprocessing unit (130) can redraw the specific material area with the same material as the adjacent area of ​​the specific material area. For example, if a mirror is included on a wall in an indoor space, the mirror can be redrawn with the material of the wallpaper appearing next to the mirror to remove the reflected area inside the mirror. That is, in order to prevent normals from being estimated differently due to pixel values ​​within the new specific material area created through inpainting, the preprocessing unit (130) can inpaint the specific material area so that it is recognized as a flat plane.

[0045] First, the preprocessing unit (130) can receive a masking image including a masking area that masks a specific material area within a perspective image. Referring to Fig. 4(a), a mirror (A) may be included within the perspective image, and since the space opposite to the mirror (A) may be reflected and appear, it corresponds to a specific material area. Accordingly, as illustrated in Fig. 4(b), a masking image including a masking area (M) that masks the mirror (A) can be generated.

[0046] Here, the masking image can be generated within the preprocessing unit (130). That is, the preprocessing unit (130) can generate a masking image by applying segmentation or object recognition techniques to each perspective image or ERP image to recognize a specific material area and automatically masking the recognized specific material area.

[0047] Additionally, depending on the embodiment, it is also possible to create a masking image by having a worker or the like visually confirm a specific material area, such as a mirror or glass, included in a perspective image or ERP image and masking the specific material area. In other words, since it may be difficult to automatically recognize a mirror or glass included in a perspective image or ERP image through image processing, it is also possible to request a masking task from a worker or the like, and then receive and process each masking image.

[0048] Upon receiving a masking image, the preprocessing unit (130) can inpaint the interior of the masking area based on the pixel values ​​of adjacent pixels located within the adjacent area bordering the masking area. That is, in order to inpaint with the same material as the material of the adjacent area, the pixel values ​​of the adjacent pixels can be checked and inpainting can be performed with the corresponding pixel values.

[0049] Specifically, the preprocessing unit (130) can extract pixel values ​​of adjacent pixels that are in horizontal contact with the masking area among adjacent pixels, and reset the pixel values ​​within the masking area with the interpolated value generated by interpolating the pixel values. In the case of an indoor space, it can be divided into a ceiling surface, a floor surface, and a side surface. Generally, since mirrors, glass, etc. are located on the side surface within an indoor space, the preprocessing unit (130) can perform inpainting based on the pixel values ​​of each adjacent pixel that is in horizontal contact with each specific material area. Through this, it is possible to make it difficult to distinguish the inpainted specific material area from the ceiling surface or the floor surface and from the side surface where the specific material area is located.

[0050] Referring to Fig. 5, there may be cases where adjacent areas exist on the left and right sides of the masking area (M) in the masking image, as in Fig. 5(a). In this case, linear interpolation may be performed on each pixel value bordering the left and right sides of the masking area (M) to generate an interpolation value, and each pixel value within the masking area (M) may be filled with the interpolation value. That is, each pixel included in the masking area (M) of Fig. 5(a) may be divided into rows, and inpainting may be performed on each row as a unit. Specifically, for each row included in the masking area (M), each adjacent pixel (a, b) bordering the left and right sides of the row may be specified. Thereafter, the pixel values ​​of the adjacent pixels (a, b) may be linearly interpolated to generate an interpolation value, and the interpolation value may be reset to the pixel values ​​for all pixels included in the row. In the same way, the masking area (M) can be inpainted by resetting the pixel values ​​of the entire row included within the masking area (M).

[0051] In addition, there may be a case where the adjacent area exists only on one of the left and right sides of the masking area (M), as in 5(b). In this case, for each row included in the masking area (M), each adjacent pixel (a) included in the adjacent area on the left side can be specified, and the pixel values ​​of all pixels in the corresponding row can be reset to the pixel value of the adjacent pixel (a). In the same manner, the masking area (M) can be inpainted by resetting the pixel values ​​of all rows included in the masking area (M).

[0052] Meanwhile, as shown in Fig. 5(c), there may be cases where there is no adjacent area horizontally adjacent to the masking area (M). In this case, the pixel values ​​of each row included in the masking area (M) can be reset to a preset single color. For example, the pixel values ​​of the corresponding pixels can be set to (r, g, b) = (128, 128, 128) for inpainting.

[0053] Once preprocessing is complete, the depth and normal estimation unit (140) can generate a depth map and a normal map corresponding to the ERP image using multiple perspective images. Here, the depth and normal estimation unit (140) can include a depth estimation model that generates each depth map, and a normal estimation model that generates a normal map.

[0054] The depth and normal estimation unit (140) can generate respective depth maps corresponding to a plurality of perspective images using a depth estimation model. Here, the depth estimation model is learned to set depth values ​​corresponding to each pixel in the perspective image when the perspective image is input, and may be implemented in various ways such as a deep learning model or a neural network model. The depth and normal estimation unit (140) can utilize various types of depth estimation models depending on the embodiment, and any model that sets a depth value for the input perspective image and generates a depth map corresponding to the perspective image can be used as a depth estimation model.

[0055] In addition, the depth and normal estimation unit (140) can generate normal maps corresponding to multiple perspective images using a normal estimation model. In order to generate a corresponding 3D mesh from an ERP image, a normal map is required along with a depth map. Therefore, it is also possible to generate normal maps together with a depth map by further including a normal estimation model. When multiple perspective images are input, the normal estimation model generates a normal map indicating normal information for each plane included in the perspective images, and can be implemented in various ways based on a deep learning model or a neural network model. In other words, anything that can generate a normal map corresponding to the perspective image by setting a normal vector for the input perspective image can be used as a normal estimation model.

[0056] In some embodiments, the depth and normal estimation unit (140) may convert each of the depth maps and normal maps into a spherical domain corresponding to the ERP image, thereby generating the ERP depth map and the ERP normal map, respectively. That is, the depth maps and normal maps generated on the cubemap domain may be combined, and converted back into a spherical domain, thereby generating the ERP depth map and the ERP normal map. At this time, the depth and normal estimation unit (140) may parameterize the depth maps and the normal maps, and update the depth values ​​of the depth maps and the normal vectors of the normal maps, thereby maintaining consistency between the depth maps and the normal maps while ensuring that the scales are matched. Thereafter, the ERP depth map and the ERP normal map may be generated by converting based on the updated depth maps and normal maps.

[0057] The mesh generation unit (150) can generate a 3D mesh corresponding to the depth map and normal map using the mesh generation model. That is, the mesh generation unit (150) can input a plurality of depth maps and normal maps corresponding to each perspective image into the mesh generation model, and generate a 3D mesh corresponding to the corresponding ERP image. However, depending on the embodiment, it is also possible to input each ERP depth map and ERP normal map into the mesh generation model to generate a 3D mesh.

[0058] Here, the mesh generation model may be one that generates a 3D mesh based on neural rendering, such as mono-SDF (Signed Distance Function). The mesh generation model may be generated by learning based on a loss function that includes depth loss and normal loss, which are the differences between the depth map and normal map of the generated 3D mesh and the depth map and normal map generated from the ERP image. For example, the loss function for learning the mesh generation model may be set as follows.

[0059] L = Lrgb + λ1L eikonal + λ2L depth + λ3L normal

[0060] Here, L rgb is color loss, L eikonal is an iconic loss, L depth is the depth loss, L normal is the normal loss, λ1 corresponds to the iconic weight, λ2 corresponds to the depth weight, and λ3 corresponds to the normal weight. λ1, λ2, and λ3 correspond to hyperparameters used to learn the mesh generation model.

[0061] That is, the loss function for learning the mesh generation model may include depth loss, normal loss, etc., and depth weights and normal weights may be applied to the depth loss and normal loss, respectively. Here, when using fixed hyperparameters, it is necessary to set the optimal hyperparameters according to the characteristics of each data during learning. For example, if the normal weight is set higher than the depth weight, learning is performed centered on the normal rather than the depth, so each distance in the generated 3D mesh may appear different, but each plane can be generated straight. That is, in the case of a thin wall, the plane of the thin wall may appear flat in the 3D mesh, but it may appear thick or be incorrectly generated as two different planes.

[0062] Additionally, when the depth weight is set higher than the normal weight, the depth is learned more importantly, so each distance within the 3D mesh may appear consistent, but problems such as each plane not being generated in a straight shape may occur. In other words, in the case of a thin wall, the distance between the thin wall and another wall may be well expressed, but problems such as a hole being created in the thin wall may occur.

[0063] Accordingly, the mesh generation unit (150) can learn by changing each weight applied to the depth loss and the normal loss during learning of the mesh generation model according to a preset scheduling. That is, the depth weight applied to the depth loss can be scheduled to decrease from the initial depth weight at each epoch to reach the target depth weight, and the normal weight applied to the normal loss can be scheduled to increase from the initial normal weight at each epoch to reach the target normal weight. At this time, the initial depth weight can be set to be larger than the initial normal weight, and the target depth weight can be set to be smaller than the target normal weight. For example, in the case of the depth weight, the initial depth weight can be set to be 0.20, the target depth weight can be set to be 0.14, and the weight can be set to decrease by 0.01 every 20 epochs, and in the case of the normal weight, the initial normal weight can be set to be 0.02, the target normal weight can be set to be 0.14, and the weight can be set to increase by 0.02 every 20 epochs.

[0064] In this case, since the depth weight is high and the normal weight is low at the beginning, structures such as walls within the 3D mesh can be generated to some extent according to the distance at the beginning, and as learning progresses, the depth weight becomes small and the normal weight becomes high, so that the noises on the plane of structures such as walls within the 3D mesh can be gradually removed and generated neatly.

[0065] That is, by learning a mesh generation model by changing the hyperparameters at each epoch rather than using fixed values ​​for the hyperparameters, it is possible to learn a mesh generation model that can generate high-performance 3D meshes without having to find optimal hyperparameters.

[0066] Meanwhile, 3D meshes generated by mesh generation models can generally provide information about relative distances and ratios between internal points, but cannot provide information about the actual distances between those points. In other words, since ERP images and the like do not contain information about actual distances, 3D meshes alone may not be able to provide information about actual distances.

[0067] Accordingly, the mesh generation unit (150) may learn by including an additional item in the depth loss when learning the mesh generation model so that the actual distance value in the 3D mesh can be reflected. That is, when the target environment (T) is an indoor space, an additional item corresponding to the difference between the depth value of the center point in the depth map of the 3D mesh and the camera height value of the camera that captured the ERP image may be additionally added to the depth loss and learned. At this time, the mesh generation unit (150) may learn by reflecting the additional item only for the perspective image corresponding to the floor surface of the target environment among a plurality of perspective images.

[0068] In the case of an ERP image generated by a 360-degree monocular camera (1), the depth value of the center point within the floor plane corresponds to the camera height of the 360-degree monocular camera (1). Therefore, if a loss function is set so that the depth value of the center point within the floor plane approaches the actual camera height, the depth value of the center point within the floor plane can appear as the actual camera height value in the generated 3D mesh. In this case, since the remaining depth values ​​within the 3D mesh also appear to correspond to actual distance values, it is possible to generate a 3D mesh that reflects a metric scale that can measure distances, etc., of an actual target environment.

[0069] The post-processing unit (160) can detect planes included in a 3D mesh and perform post-processing to remove noise by simplifying the planes when the target environment is an indoor space. Specifically, the post-processing unit (160) can first convert the 3D mesh into a point cloud and distinguish each plane included in the point cloud using a segmentation technique such as RANSAC (Random Sample Consensus).

[0070] Afterwards, the ceiling and floor planes can be distinguished based on the normal direction of each plane. Specifically, by comparing the normal direction of each plane with a predefined vertical direction, candidate planes that match the vertical direction within a margin of error can be found. Among the candidate planes, those that contain the largest number of points above and below the center of the global coordinate system can be identified and set as the ceiling and floor planes, respectively.

[0071] In addition, points located within a set range in the normal direction of the ceiling and floor planes can be projected onto the respective ceiling or floor planes and included in the ceiling or floor planes. That is, as illustrated in Fig. 6, each point located within a set range (threshold) in the normal direction of each ceiling or floor plane can be found, and the points can be projected onto a plane. Through this, noises located around the ceiling or floor plane can be simplified by organizing them into the ceiling or floor plane.

[0072] Thereafter, the post-processing unit (160) can set the remaining points, excluding the ceiling and floor surfaces, as sides based on their respective normals. Specifically, points having normals perpendicular to the normal vector of the ceiling or floor surface can be clustered, and at this time, points adjacent to the points can be further included and clustered based on the normal direction. Here, each clustered surface can be distinguished into its respective side surface. That is, points having normals perpendicular to the normal vector of the ceiling or floor surface correspond to sides within the indoor space, and thus can be clustered.

[0073] Meanwhile, the Manhattan-world assumption states that each surface within an indoor space is created to be orthogonal to each other. When the Manhattan assumption is applied, if the normal direction of the ceiling or floor surface is set to the z-axis in the Cartesian coordinate system, the remaining surfaces can be defined as having normal directions corresponding to the x-axis and y-axis, respectively. Therefore, points having normal directions corresponding to the x-axis and y-axis can be clustered and distinguished as surfaces.

[0074] On the other hand, if the Manhattan assumption does not apply, the side faces may be perpendicular to the normals of the ceiling and floor, but may not form perpendicular angles between the side faces. Therefore, the side faces can be identified by clustering points whose normals are perpendicular to the normal vectors of the ceiling or floor.

[0075] Afterwards, for each side, points located within the set range in the normal direction of each side can be projected onto the corresponding side and included in each side.

[0076] Finally, once the ceiling, floor, and side surfaces are identified, the tangent lines connecting the ceiling, floor, and side walls can be detected, and the presence of areas extending beyond these lines can be determined. If areas extending beyond these lines are present, these areas can be projected within the tangent lines and eliminated. In other words, since ERP images capture indoor spaces, if the Manhattan assumption is satisfied, the ceiling, floor, and side walls can be combined to form a rectangular parallelepiped. Areas extending beyond the tangent lines are considered to be areas of error, as they protrude from the corresponding rectangular parallelepiped. Therefore, by projecting these outliers onto their respective planes and eliminating them, it is possible to simplify and generate the entire 3D mesh.

[0077] Figure 7 is a block diagram illustrating a computing environment (10) suitable for use in exemplary embodiments. In the illustrated embodiment, each component may have different functions and capabilities other than those described below, and may include additional components other than those described below.

[0078] The illustrated computing environment (10) includes a computing device (12). In one embodiment, the computing device (12) may be a device that generates a three-dimensional mesh based on an ERP image (e.g., a three-dimensional mesh generating device (100)).

[0079] A computing device (12) includes at least one processor (14), a computer-readable storage medium (16), and a communication bus (18). The processor (14) may cause the computing device (12) to operate according to the exemplary embodiments mentioned above. For example, the processor (14) may execute one or more programs stored in the computer-readable storage medium (16). The one or more programs may include one or more computer-executable instructions, which, when executed by the processor (14), may be configured to cause the computing device (12) to perform operations according to the exemplary embodiments.

[0080] A computer-readable storage medium (16) is configured to store computer-executable instructions or program code, program data, and / or other suitable forms of information. A program (20) stored in the computer-readable storage medium (16) includes a set of instructions executable by the processor (14). In one embodiment, the computer-readable storage medium (16) may be a memory (volatile memory such as random access memory, non-volatile memory, or a suitable combination thereof), one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, any other form of storage medium that is accessible by the computing device (12) and capable of storing desired information, or a suitable combination thereof.

[0081] A communication bus (18) interconnects various other components of the computing device (12), including the processor (14) and computer-readable storage media (16).

[0082] The computing device (12) may also include one or more input / output interfaces (22) that provide interfaces for one or more input / output devices (24) and one or more network communication interfaces (26). The input / output interfaces (22) and the network communication interfaces (26) are connected to the communication bus (18). The input / output devices (24) may be connected to other components of the computing device (12) via the input / output interfaces (22). Exemplary input / output devices (24) may include input devices such as pointing devices (such as a mouse or a trackpad), a keyboard, a touch input device (such as a touchpad or a touchscreen), a voice or sound input device, various types of sensor devices and / or photographing devices, and / or output devices such as display devices, printers, speakers and / or network cards. The exemplary input / output devices (24) may be included within the computing device (12) as a component constituting the computing device (12), or may be connected to the computing device (12) as a separate device distinct from the computing device (12).

[0083] Figures 8 and 9 are flowcharts illustrating a method for generating a 3D mesh based on an ERP image according to one embodiment of the present invention. Here, each step of Figures 8 and 9 can be performed by a computing device according to one embodiment of the present invention.

[0084] Referring to FIG. 8, the computing device can receive an ERP image for the target environment (S110). That is, the computing device can receive the ERP image directly from a 360-degree monocular camera or a user's terminal device that generates the ERP image via a wired or wireless communication network, or from a separate server or storage device that stores the ERP image.

[0085] Thereafter, the computing device can convert the ERP image for the target environment to generate multiple perspective images (S120). That is, when an ERP image on a sphere domain is input, the computing device can convert the ERP image to generate a cubemap image on a corresponding cubemap domain. The cubemap image can be generated in a form corresponding to a development diagram of a rectangular parallelepiped, and each of the front, back, left side, right side, top side, and bottom side corresponding to the target environment can be displayed in a perspective form. In addition, the computing device can generate the perspective image by dividing each of the six perspective images corresponding to the six faces included in the cubemap image. That is, each of the six perspective images corresponding to the front, back, left side, right side, top side, and bottom side can be divided and generated.

[0086] If a perspective image containing a specific material region exists among multiple perspective images, the computing device can remove the specific material region using an inpainting technique (S130). Here, the specific material region includes a transparent or reflective surface, such as a mirror, glass, or metal surface.

[0087] A computing device can use inpainting techniques to re-create specific material regions, reproducing them so that they appear to be made of the same material as adjacent areas. Since specific material regions reflect or project light to create the illusion of another space within them, it is necessary to remove the light reflection and projection characteristics of specific material regions before estimating depth maps or normal maps based on perspective images. Therefore, the computing device can be configured to draw the specific material region with the same material as adjacent areas.

[0088] Specifically, the computing device can receive a masking image including a masking area in which a specific material area within a perspective image is masked. Here, the masking image may be generated within the preprocessing unit (130), and in some embodiments, a masking image generated by a worker or the like may be provided externally.

[0089] When receiving a masking image, the computing device can inpaint the inside of the masking area based on the pixel values ​​of adjacent pixels located within the adjacent area that borders the masking area. In some embodiments, among the adjacent pixels, the pixel values ​​of adjacent pixels that border the masking area horizontally can be extracted, and the pixel values ​​within the masking area can be reset using the interpolated values ​​generated by interpolating the pixel values. Here, by performing inpainting based on the pixel values ​​of each adjacent pixel that borders the horizontal direction, the inpainted specific material area can be implemented so that it is distinguishable from the ceiling or the floor, but difficult to distinguish from the side surface where the specific material area is located.

[0090] Thereafter, the computing device can generate a depth map and a normal map corresponding to the ERP image using a plurality of perspective images (S140). Here, the computing device can generate the depth map and the normal map, respectively, using a depth estimation model that generates each depth map and a normal estimation model that generates each normal map. The depth estimation model and the normal estimation model can be implemented in various ways, such as a deep learning model or a neural network model, and any model can be used as long as it sets a depth value and a normal vector for an input perspective image to generate a corresponding depth map and normal map.

[0091] In some embodiments, the computing device may convert each of the depth maps and normal maps into a spherical domain corresponding to the ERP image, thereby generating the ERP depth map and the ERP normal map, respectively. That is, the depth maps and normal maps generated in the cubemap domain may be combined, and then converted back into the spherical domain, thereby generating the ERP depth map and the ERP normal map. At this time, the computing device may parameterize the depth maps and the normal maps, and update the depth values ​​of the depth maps and the normal vectors of the normal maps, thereby maintaining consistency between the depth maps and the normal maps while ensuring that the scales are matched. Thereafter, the ERP depth map and the ERP normal map may be generated by converting based on the updated depth maps and normal maps.

[0092] The computing device can generate a 3D mesh corresponding to the depth map and the normal map using the mesh generation model (S150). That is, the computing device can input a plurality of depth maps and normal maps corresponding to each perspective image into the mesh generation model to generate a 3D mesh corresponding to the corresponding ERP image. However, depending on the embodiment, it is also possible to input each ERP depth map and ERP normal map into the mesh generation model to generate a 3D mesh.

[0093] Here, the mesh generation model may be one that generates a 3D mesh based on neural rendering, such as mono-SDF (Signed Distance Function). The mesh generation model may be one that is learned based on a loss function that includes depth loss and normal loss, which are the differences between the depth map and normal map of the generated 3D mesh and the depth map and normal map generated from the ERP image.

[0094] Here, when using fixed hyperparameters, it is necessary to set optimal hyperparameters based on the characteristics of each data set during training. However, finding and setting optimal hyperparameters each time is not easy, making it difficult to implement a mesh generation model capable of generating high-performance 3D meshes.

[0095] However, the computing device may learn by changing the respective weights applied to the depth loss and normal loss during learning of the mesh generation model according to a preset scheduling. That is, the depth weight applied to the depth loss may be scheduled to decrease from the initial depth weight at each epoch to reach the target depth weight, and the normal weight applied to the normal loss may be scheduled to increase from the initial normal weight at each epoch to reach the target normal weight. In this case, the initial depth weight may be set to be greater than the initial normal weight, and the target depth weight may be set to be less than the target normal weight.

[0096] In this case, since the depth weight is high and the normal weight is low at the beginning, structures such as walls within the 3D mesh can be generated to some extent according to the distance at the beginning, and as learning progresses, the depth weight becomes small and the normal weight becomes high, so that the noises on the plane of structures such as walls within the 3D mesh can be gradually removed and generated neatly.

[0097] That is, by learning a mesh generation model by changing the hyperparameters at each epoch rather than using fixed values ​​for the hyperparameters, it is possible to learn a mesh generation model that can generate high-performance 3D meshes without having to find optimal hyperparameters.

[0098] Additionally, in the case of a 3D mesh generated by a mesh generation model, it can generally provide information about the relative distance or ratio between internal points, but it may not be able to provide information about the actual distance between the points.

[0099] Accordingly, the computing device may learn by including an additional item in the depth loss when learning the mesh generation model so that the actual distance value within the 3D mesh can be reflected. That is, when the target environment is an indoor space, the learning may include an additional item corresponding to the difference between the depth value of the center point within the depth map of the 3D mesh and the camera height value of the camera that captured the ERP image by adding it to the depth loss. At this time, the computing device may learn by reflecting the additional item only for the perspective image corresponding to the floor surface of the target environment among the plurality of perspective images.

[0100] In the case of an ERP image generated by a 360-degree monocular camera, the depth value of the center point within the floor plane corresponds to the camera height of the 360-degree monocular camera. Therefore, if the loss function is set so that the depth value of the center point of the floor plane approaches the actual camera height, the depth value of the center point of the floor plane in the generated 3D mesh can match the actual camera height value. Accordingly, since the remaining depth values ​​within the 3D mesh also appear to correspond to actual distance values, the mesh generation model can generate a generated 3D mesh that reflects a metric scale that can measure distances, etc., of the actual target environment.

[0101] If the target environment is an indoor space, the computing device can detect planes included in a 3D mesh and perform post-processing to simplify the planes and remove noise (S160). Specifically, referring to FIG. 9, the computing device can first convert the 3D mesh into a point cloud and then segment each plane included in the point cloud using a segmentation technique such as RANSAC (S161).

[0102] Thereafter, based on the normal direction of each plane, the ceiling and floor planes can be distinguished (S162). Specifically, by comparing the normal direction of each plane with a predefined vertical direction, candidate planes that match the vertical direction within a margin of error can be found. Among the candidate planes, the candidate planes that contain the most points above and below the center of the global coordinate system can be found and set as the ceiling and floor planes, respectively.

[0103] Here, points located within a set range in the normal direction of the ceiling and floor surfaces can be projected onto the respective ceiling or floor surfaces and included in the ceiling or floor surfaces (S163). Through this, noises located around the ceiling or floor surfaces can be simplified by organizing them into the ceiling or floor surfaces.

[0104] Thereafter, the computing device can set the remaining points, excluding the ceiling and floor surfaces, as sides based on their respective normals. Specifically, points having normals perpendicular to the normal vectors of the ceiling or floor surfaces can be clustered, and at this time, points adjacent to these points can be further included and clustered based on their respective normal directions. Here, each of the clustered surfaces can be distinguished into its respective sides (S164). That is, points having normals perpendicular to the normal vectors of the ceiling or floor surfaces correspond to sides within the indoor space, and thus can be clustered.

[0105] Meanwhile, the Manhattan-world assumption states that each surface within an indoor space is created to be orthogonal to one another. When the Manhattan assumption is applied, the normal direction of the ceiling or floor surface can be set to the z-axis in the Cartesian coordinate system, and the remaining surfaces can be defined as having normal directions corresponding to the x-axis and y-axis, respectively. Accordingly, points having normal directions corresponding to the x-axis and y-axis can be clustered and distinguished into surfaces.

[0106] On the other hand, if the Manhattan assumption does not apply, the side faces may be perpendicular to the normals of the ceiling and floor, but may not form perpendicular angles between the side faces. Therefore, the side faces can be identified by clustering points whose normals are perpendicular to the normal vectors of the ceiling or floor.

[0107] Afterwards, for each side, points located within the set range in the normal direction of each side can be projected onto the corresponding side and included in each side (S165).

[0108] Once the ceiling, floor, and side surfaces are distinguished, the computing device can detect the tangent lines that connect the ceiling, floor, and side walls, and determine whether there are areas that cross the tangent lines. If there are areas that cross the tangent lines, the areas can be removed by projecting them onto the inside of the tangent lines (S166). In other words, since the ERP image is a photograph of an indoor space, if the Manhattan assumption is satisfied, the ceiling, floor, and side walls can be combined to create a rectangular parallelepiped shape. At this time, the areas that cross the tangent lines correspond to areas that protrude from the corresponding rectangular parallelepiped, and can therefore be considered areas corresponding to errors. Therefore, by projecting the outlier areas onto each plane and removing them, it is possible to simplify and generate the entire 3D mesh.

[0109] The present invention described above can be implemented as computer-readable code on a medium recording a program. The computer-readable medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. Furthermore, the medium may be a variety of recording or storage means, including a single or multiple hardware components, and is not limited to media directly connected to a computer system, but may also be distributed across a network. Examples of the medium include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and ROM, RAM, flash memory, and other media configured to store program instructions. Furthermore, other examples of media include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc. Therefore, the above detailed description should not be construed as limiting in all respects, but rather as illustrative. The scope of the present invention should be determined by a reasonable interpretation of the appended claims, and all changes within the equivalent scope of the present invention are included in the scope of the present invention.

[0110] The present invention is not limited to the above-described embodiments and the attached drawings. It will be apparent to those skilled in the art that components of the present invention can be substituted, modified, and altered without departing from the technical spirit of the present invention.

Claims

1. A method for generating a 3D mesh based on an ERP (Equirectangular Project) image using a computing device, A step of converting an ERP image for a target environment to generate multiple perspective images; A step of removing a perspective image including a specific material area using an inpainting technique if there is a perspective image among the plurality of perspective images; A step of generating a depth map and a normal map corresponding to the ERP image using the plurality of perspective images; and A method for generating a 3D mesh based on an ERP image, comprising the step of generating a 3D mesh corresponding to the depth map and the normal map using a mesh generation model.

2. In paragraph 1, the specific material area is A method for generating a three-dimensional mesh based on an ERP image, which includes a transparent or reflective surface.

3. In the first paragraph, the step of removing using the inpainting technique is A method for generating a 3D mesh based on an ERP image, wherein the specific material area is regenerated with the same material as an adjacent area of ​​the specific material area using the above-mentioned inpainting technique.

4. In the first paragraph, the step of removing using the infating technique is A step of receiving a masking image including a masking area in which the specific material area in the above perspective image is masked; and A method for generating a 3D mesh based on an ERP image, comprising a step of inpainting the inside of the masking area based on pixel values ​​of adjacent pixels located within an adjacent area that contacts the masking area.

5. In the fourth paragraph, the inpainting step A method for generating a 3D mesh based on an ERP image, wherein the pixel values ​​of adjacent pixels that are in horizontal contact with the masking area are interpolated to generate an interpolation value, and the pixel values ​​within the masking area are reset with the interpolation value.

6. In the first paragraph, the mesh generation model is It is generated by learning based on a loss function including depth loss and normal loss, which are the differences between the depth map and normal map of the above 3D mesh and the depth map and normal map generated from the ERP image. A method for generating a 3D mesh based on an ERP image, wherein each weight applied to the depth loss and normal loss during learning is changed according to a preset scheduling.

7. In paragraph 6, the scheduling is A method for generating a 3D mesh based on an ERP image, wherein the depth weight applied to the depth loss is set to decrease from the initial depth weight at each epoch to reach a target depth weight, and the normal weight applied to the normal loss is set to increase from the initial normal weight at each epoch to reach a target normal weight.

8. In paragraph 7, A method for generating a 3D mesh based on an ERP image, wherein the initial depth weight is greater than the initial normal weight, and the target depth weight is smaller than the target normal weight.

9. In paragraph 6, the mesh generation model is A method for generating a 3D mesh based on an ERP image, wherein if the target environment is an indoor space, an additional item corresponding to the difference between the depth value of the center point in the depth map of the 3D mesh and the camera height value of the camera that captured the ERP image is further included in the depth loss and learned.

10. In the 9th paragraph, the mesh generation model is A method for generating a 3D mesh based on an ERP image, wherein the additional items are reflected and learned for a perspective image corresponding to the floor surface of the target environment among the plurality of perspective images.

11. In paragraph 1, A method for generating a 3D mesh based on an ERP image, further comprising a post-processing step of detecting a plane included in the 3D mesh and simplifying the planes to remove noise if the target environment is an indoor space.

12. In the 11th paragraph, the post-processing step A step of converting the above 3D mesh into a point cloud and distinguishing each plane included in the point cloud using a segmentation technique; A step of distinguishing the ceiling plane and the floor plane based on the direction of the normal lines of the above planes; A step of projecting points located within a set range in the normal direction of the ceiling surface and floor surface onto the ceiling surface or floor surface and including them in the ceiling surface or floor surface; A step of clustering points having a normal line perpendicular to the normal vector of the ceiling surface or floor surface and dividing them into each side; A step of projecting points located within a set range in the normal direction of the above-mentioned sides onto the side and including them in each of the sides; and A method for generating a 3D mesh based on an ERP image, comprising the steps of detecting tangent lines where the ceiling surface and the floor surface and the side walls come into contact, and projecting areas extending beyond the tangent lines into the interior of the tangent lines, respectively.

13. In paragraph 12, the step of dividing into each aspect is A method for generating a 3D mesh based on an ERP image, wherein when the Manhattan-world assumption is applied, the normal direction of the ceiling or floor surface is set to the z-axis on the orthogonal coordinate system, and points having normal directions corresponding to the remaining x-axis and y-axis are each clustered and divided into the side surfaces.

14. A computing device including a processor and generating a three-dimensional mesh based on an ERP (Equirectangular Project) image, The above processor, Converting an ERP image to a target environment to generate multiple perspective images; If a perspective image including a specific material area exists among the above plurality of perspective images, the specific material area is removed using an inpainting technique; Using the plurality of perspective images, generating a depth map and a normal map corresponding to the ERP image; and A computing device, comprising: generating a three-dimensional mesh corresponding to the depth map and the normal map using a mesh generation model.

Citation Information

Patent Citations

  • KR20220167824A