Depth map image generation method and computing device therefor

The method addresses image deviation issues in generating panoramic depth map images by using a deep learning-based neural network with padding data, resulting in minimized errors and accurate data continuity.

WO2025110506A1PCT designated stage expired Publication Date: 2025-05-303I INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/016407
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-21
Filing Date
2024-10-25
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing technologies for generating depth map images from color images using deep learning suffer from image deviation between the left and right ends, particularly in panoramic images, leading to errors in areas where the two ends meet.

Method used

A method and computing device that utilize a deep learning-based artificial neural network to generate panoramic depth map images by inputting padding data to the left and right sides of panoramic images, minimizing errors between side data and preventing errors in areas where the two ends meet.

Benefits of technology

The method effectively generates panoramic depth map images with minimized errors between side data, ensuring accurate data continuity and smooth virtual space configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024016407_30052025_PF_FP_ABST
    Figure KR2024016407_30052025_PF_FP_ABST
Patent Text Reader

Abstract

One technical aspect of the present application presents a computing device. The computing device comprises: a memory for storing one or more instructions; and a processor for executing the one or more instructions stored in the memory, wherein the processor can generate a depth map image by executing the one or more instructions.
Need to check novelty before this filing date? Find Prior Art

Description

Method for generating a depth map image and computing device therefor

[0001] The present invention relates to a method for generating a depth map image and a computing device therefor.

[0002] Recently, virtual space implementation technology has been developed that provides an online virtual space corresponding to the actual space, allowing users to experience being in an actual space without having to visit the actual space in person.

[0003] 3D space reconstruction technology is used as a technology for implementing this type of virtual space, and in order to implement the space in three dimensions, depth information (depth map) images collected from each shooting point are required.

[0004] These depth information images can be extracted using positioning sensors such as LiDAR or parallax information from stereo cameras, but there are limitations such as the need for experts to take pictures with fairly expensive equipment to actually extract them.

[0005] Accordingly, technologies are being developed to generate depth information images from color images using deep learning artificial neural networks, etc.

[0006] However, in the case of these conventional technologies, there is a limitation that the image deviation between the left and right ends occurs, reducing the sense of reality.

[0007] In particular, in the case of panoramic images, they are used to provide 360 ​​VR or 3D space restoration, but in these panoramic images, when there is a step difference between the left and right ends, there is a limitation that the error is prominently displayed in a specific area of ​​the space where the two ends meet.

[0008] Prior art document: (Patent Document 0001) U.S. Patent Publication No. 11417096 ("Video format classification and metadata injection using machine learning", published on November 26, 2020)

[0009] One technical aspect of the present invention is to solve the problems of the above-mentioned prior art, and one technical aspect of the present invention is to provide a learning data generation method and a computing device therefor, which can generate a panoramic depth map image based on a panoramic color image using a deep learning-based artificial neural network.

[0010] In addition, one technical aspect of the present invention is to provide a learning data generation method and a computing device therefor, which can minimize errors between side data of a panoramic depth map image generated by a deep learning-based artificial neural network by inputting padding data to the left and right sides of a panoramic image.

[0011] The tasks of this application are not limited to the tasks mentioned above, and other tasks not mentioned will be clearly understood by those skilled in the art from the description below.

[0012] One technical aspect of the present application proposes a computing device. The computing device comprises a memory storing one or more instructions; and a processor executing the one or more instructions stored in the memory, wherein the processor is capable of generating a depth map image by executing the one or more instructions.

[0013] In one embodiment, the processor, by executing the one or more instructions, receives a base panorama color image and a base panorama depth map image corresponding to the base panorama color image, and generates a plurality of learning panorama color images and a plurality of learning panorama depth map images based on a spherical transformation of the base panorama color image and the base panorama depth map image. The plurality of learning panorama color images and the plurality of learning panorama depth map images may be images having padding data inserted at both ends.

[0014] In another embodiment, the processor sets a padding area for a query panoramic color image by executing the one or more instructions, inserts padding data into the padding area to generate a padded query panoramic color image, generates an estimated panoramic depth map image for the padded query panoramic color image using a pre-trained neural network, the estimated panoramic depth map image including the padding area, and deletes an area corresponding to the padding area from the estimated panoramic depth map image to generate an estimated panoramic depth map image corresponding to the query panoramic color image.

[0015] Another technical aspect of the present application proposes a method for generating a depth map image. The method for generating a depth map image is a method for generating a depth map image performed on a computing device, the computing device comprising: a memory storing one or more instructions; and a processor executing the one or more instructions stored in the memory, wherein the processor can perform the method for generating the depth map image by executing the one or more instructions.

[0016] In one embodiment, the method for generating a depth map image includes the steps of receiving a base panorama color image and a base panorama depth map image corresponding to the base panorama color image, and the steps of generating a plurality of learning panorama color images and a plurality of learning panorama depth map images based on a spherical transformation of the base panorama color image and the base panorama depth map image. The plurality of learning panorama color images and the plurality of learning panorama depth map images are images having padding data inserted at both ends.

[0017] In another embodiment, the depth map image generation method includes the steps of setting a padding area for a query panorama color image and inserting padding data into the padding area to generate a padded query panorama color image, generating an estimated panorama depth map image for the padded query panorama color image using a pre-trained neural network, the estimated panorama depth map image including the padding area, and generating an estimated panorama depth map image corresponding to the query panorama color image by deleting an area corresponding to the padding area from the estimated panorama depth map image.

[0018] Another technical aspect of the present application proposes a storage medium storing computer-readable instructions, wherein the instructions, when executed by a computing device, cause the computing device to perform an operation of generating a depth map image.

[0019] In one embodiment, the instructions, when executed by a computing device, cause the computing device to perform operations of receiving a base panorama color image and a base panorama depth map image corresponding to the base panorama color image, and generating a plurality of learning panorama color images and a plurality of learning panorama depth map images based on a spherical transformation of the base panorama color image and the base panorama depth map image. The plurality of learning panorama color images and the plurality of learning panorama depth map images are images having padding data inserted at both ends.

[0020] In another embodiment, the instructions, when executed by a computing device, cause the computing device to perform the following operations: setting a padding area for a query panoramic color image and inserting padding data into the padding area to generate a padded query panoramic color image; generating an estimated panoramic depth map image for the padded query panoramic color image using a pre-trained neural network, the estimated panoramic depth map image including the padding area; and deleting an area corresponding to the padding area from the estimated panoramic depth map image to generate an estimated panoramic depth map image corresponding to the query panoramic color image.

[0021] The solutions to the above problems do not exhaustively enumerate the features of the present invention. The various solutions to the problems of the present invention can be understood in more detail by referring to the specific embodiments described below.

[0022] According to one embodiment of the present invention, there is an effect of being able to generate a panoramic depth map image based on a panoramic color image using a deep learning-based artificial neural network.

[0023] In addition, according to one embodiment of the present invention, by inputting padding data on the left and right sides of a panoramic image, an error between the side data of a panoramic depth map image generated by a deep learning-based artificial neural network can be minimized, thereby having the effect of fundamentally preventing an error occurring in an area where both side ends of a panoramic image meet when constructing a virtual space.

[0024] The effects of the invention described above are not all enumerated in the various effects of the present application, and these various effects can be easily understood from the specific embodiments of the detailed description.

[0025] FIG. 1 is a diagram illustrating a system for providing a spherical image based on a depth map image according to one embodiment disclosed in the present application.

[0026] FIG. 2 is a diagram illustrating a learning data generation architecture according to one embodiment disclosed in the present application.

[0027] FIG. 3 is a flowchart illustrating a method for generating learning data according to one embodiment disclosed in the present application.

[0028] FIG. 4 is a diagram illustrating an example of generating a plurality of learning panoramic images based on a basic panoramic color image according to one embodiment disclosed in the present application.

[0029] FIG. 5 is a flowchart illustrating a data padding method for generating learning data according to one embodiment disclosed in the present application.

[0030] Figure 6 is a drawing illustrating an example of a method for padding data for the description of Figure 5.

[0031] FIG. 7 is a flowchart illustrating a method for setting padding data according to one embodiment disclosed in the present application.

[0032] Figure 8 is a drawing illustrating an example of a method for padding data for the description of Figure 7.

[0033] FIG. 9 is a diagram illustrating a training module according to one embodiment disclosed in the present application.

[0034] FIG. 10 is a flowchart illustrating a training method according to one embodiment disclosed in the present application.

[0035] FIG. 11 is a diagram illustrating an architecture for generating an estimated panoramic depth map image according to one embodiment disclosed in the present application.

[0036] FIG. 12 is a flowchart illustrating a method for generating an estimated panoramic depth map image according to one embodiment disclosed in the present application.

[0037] FIG. 13 is a diagram illustrating an example of an estimated panoramic depth map image generated according to one embodiment disclosed in the present application.

[0038] FIG. 14 is a block diagram illustrating a computing device according to one embodiment disclosed in the present application.

[0039] Hereinafter, preferred embodiments of the present invention will be described with reference to the attached drawings.

[0040] However, the embodiments of the present invention may be modified in various other forms, and the scope of the present invention is not limited to the embodiments described below. Furthermore, the embodiments of the present invention are provided to more fully explain the present invention to those of ordinary skill in the art.

[0041] That is, the above-mentioned purpose, features and advantages are described in detail below with reference to the attached drawings, so that those with ordinary skill in the art to which the present invention pertains can easily practice the technical idea of ​​the present invention. In describing the present invention, if it is determined that a detailed description of a known technology related to the present invention may unnecessarily obscure the gist of the present invention, a detailed description thereof will be omitted. Hereinafter, a preferred embodiment according to the present invention will be described in detail with reference to the attached drawings. In the drawings, the same reference numerals are used to indicate the same or similar components.

[0042] Additionally, the singular expressions used herein include plural expressions unless the context clearly dictates otherwise. In this application, terms such as "consist of" or "comprises" should not be construed to necessarily include all of the various components or various steps described in the specification, and should be construed to mean that some of the components or some of the steps may not be included, or that additional components or steps may be included.

[0043] Furthermore, various components and their subcomponents are described below to explain the system according to the present invention. These components and their subcomponents may be implemented in various forms, such as hardware, software, or a combination thereof. For example, each component may be implemented as an electronic component for performing its function, or may be implemented as software itself operable in an electronic system or as a functional element of such software. Alternatively, each component may be implemented as an electronic component and its corresponding operating software.

[0044] The various techniques described herein may be implemented with hardware, software, or, where appropriate, a combination of both. Terms such as "module," "unit," "server," and "system," as used herein, may be treated as equivalent to computer-related entities, i.e., hardware, a combination of hardware and software, software, or software in execution. Furthermore, each function executed in the system of the present invention may be configured as a module unit, recorded in a single physical memory, or recorded in a distributed manner between two or more memories and storage media.

[0045] Various embodiments of the present application may be implemented as software (e.g., a program) including one or more instructions stored in a storage medium that can be read by a machine (e.g., a user terminal (100) or a computing device (300)). For example, the processor (301) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the device to operate to perform at least one function according to the at least one instruction called. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' only means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium.

[0046] According to an embodiment, the method according to various embodiments disclosed in the present application may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read only memory (CDROM)), or may be distributed online (e.g., by download or upload) through an application store (e.g., Play Store) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0047] Although various flowcharts are disclosed to explain embodiments of the present invention, these are provided for the convenience of explaining each step, and each step is not necessarily performed in the order of the flowchart. That is, each step in the flowchart may be performed simultaneously, in the order shown in the flowchart, or in an order opposite to the order shown in the flowchart.

[0048] FIG. 1 is an exemplary drawing illustrating a system for providing a spherical image based on a depth map image according to one embodiment disclosed in the present application.

[0049] A system (10) for providing a spherical image based on a depth map image may include a user terminal (100), an image acquisition device (200), and a computing device (300).

[0050] The user terminal (100) is an electronic device that can be used by a user to access a computing device (300), and includes, for example, a mobile phone, a smart phone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation device, a personal computer (PC), a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, a head mounted display (HMD)), etc. However, in addition to these, the user terminal (100) may include an electronic device used for VR (Virtual Reality) and AR (Augmented Reality).

[0051] The image acquisition device (200) is a device that generates a color image and / or a depth map image used to generate a spherical image. In the drawings below, the color image is described as RGB, and the depth map image is described as Depth.

[0052] In the illustrated example, the image acquisition device (200) is illustrated as being divided into a distance measurement device (210) and an imaging device (220), but this is exemplary, and distance measurement and imaging may be performed using a single image acquisition device (200).

[0053] The imaging device (220) is a portable electronic device having a photographing function, which generates a color image expressed in color for a subject area - that is, an area captured in a color image.

[0054] In this application, the term "color image" encompasses all images expressed in color, and is not limited to a specific expression method. Therefore, in addition to color images, images such as CMYK (Cyan, Magenta, Yellow Key) images are collectively referred to as color images.

[0055] The imaging device (220) includes, for example, a mobile phone, a smart phone, a laptop computer, a personal digital assistant (PDA), a tablet PC, an ultrabook, a wearable device (e.g., a smart glass), etc.

[0056] The distance measuring device (210) is a device that can generate depth information for a subject area and create a depth map image.

[0057] In the present application specification, a depth map image encompasses an image that includes depth information with respect to a subject space. In other words, a depth map image means an image that expresses distance information from a capturing point to each point in a photographed subject space.

[0058] The distance measuring device (210) may include a predetermined sensor for distance measurement, such as a lidar sensor, an infrared sensor, an ultrasonic sensor, etc. Alternatively, the distance measuring imaging device (220) may include a stereo camera, a stereoscopic camera, a 3D depth camera (3D, depth camera), etc. that can measure distance information in place of a sensor.

[0059] The image generated by the imaging device (220) is called a basic panoramic color image, and the image generated by the distance measuring device (210) is called a 'basic panoramic depth map image'. In the present application specification, the color image and the depth map image are described by way of example as a panoramic image. The panoramic image is an image that can be expressed by covering at least 360 degrees in the left and right directions, and may be, for example, an equirectangular projection image.

[0060] The basic panoramic color image generated by the imaging device (220) and the basic panoramic depth map image generated by the distance measuring device (210) are generated for the same subject area under the same conditions (e.g., resolution, etc.), and therefore are matched 1:1 with each other.

[0061] The computing device (300) can receive a basic panoramic color image and a basic panoramic depth map image and perform learning. Here, the basic panoramic color image and the basic panoramic depth map image can be transmitted via a network.

[0062] The computing device (300) can generate a virtual spherical image implemented in 360 degrees based on the progressed learning (hereinafter referred to as a spherical image). In addition, the computing device (300) provides the generated spherical image to the user terminal (100). Here, the provision of the spherical image can be carried out in various forms, and as an example, it includes providing the spherical image to be operated on the user terminal (100), or as another example, providing a user interface for the spherical image implemented on the computing device (300).

[0063] The computing device (300) can generate multiple learning color images and learning panorama depth map images by converting the base panorama color image and the base panorama depth map image. This utilizes a characteristic environment that uses spherical images, and after spherizing the base panorama color image and the base panorama depth map image, multiple learning color images and learning panorama depth map images can be generated through slight adjustments.

[0064] A computing device (300) can generate a depth map image from a color image through an artificial neural network for deep learning. At this time, the computing device (300) can minimize errors between the side data of the panoramic depth map image generated by the deep learning-based artificial neural network by generating and setting padding data on the left and right sides of the panoramic image. As a result, errors occurring in the area where the two side edges of the panoramic image meet can be fundamentally prevented when constructing a virtual space.

[0065] Hereinafter, various embodiments will be described with reference to FIGS. 2 to 15.

[0066] FIG. 2 is a diagram illustrating a learning data generation architecture according to one embodiment disclosed in the present application.

[0067] The computing device (300) can generate a spherical image based on a basic panoramic color image transmitted from an imaging device (220) and a basic panoramic depth map image input from a distance measuring device (210).

[0068] The computing device (300) may include a learning data generation module (310), a neural network module (320), and a training module (330).

[0069] The learning data generation module (310) generates a plurality of learning data based on the basic panoramic color image and the basic panoramic depth map image.

[0070] That is, the learning data generation module (310) is provided with a base panorama color image and a base depth map image corresponding to the base panorama color image, and can generate a plurality of learning panorama color images and a plurality of learning panorama depth map images based on a spherical transformation for the base panorama color image and the base depth map image.

[0071] Here, multiple learning panorama color images and multiple learning panorama depth map images are images with padding data inserted on both sides.

[0072] The neural network module (320) learns based on learning data, and when a panoramic color image is input, it generates and outputs an estimated panoramic depth map image corresponding to it.

[0073] The neural network module (320) is a deep learning-based learning model that can generate an estimated panoramic depth map image corresponding to a color image using an artificial neural network for an input learning color image.

[0074] The neural network module (320) can be expressed as a mathematical model using nodes and edges. The neural network module (320) may have an architecture of a deep neural network (DNN) or an n-layer neural network. The DNN or n-layer neural network may correspond to a convolutional neural network (CNN), a convolutional neural network (CNN) based on a deep high-resolution network (HRNet), a recurrent neural network (RNN), deep belief networks, restricted Boltzman machines, etc.

[0075] For example, the neural network module (320) may receive a learning panorama color image as input and generate an estimated panorama depth map image for the image. In this example, in the initial operation, the neural network module (320) may generate an estimated panorama depth map image based on random values ​​at each node of the neural network because there is no learned content. The neural network module (320) may repeatedly perform feedback training on the generated estimated panorama depth map image to improve the accuracy of the estimated panorama depth map.

[0076] As another example, the neural network module (320) may receive a learning panorama color image and a set of learning panorama depth maps matching the same as input, learn the correlation between the panorama color image and the panorama depth map image based on the learning thereof, and generate an estimated panorama depth map image for the learning panorama color image based on the correlation. In this example as well, feedback training for the estimated panorama depth map image may be repeatedly performed to improve the accuracy of the estimated panorama depth map.

[0077] FIG. 3 is a flowchart illustrating a learning data generation method according to one embodiment disclosed in the present application, which will be described with reference to FIG. 3.

[0078] Referring to FIG. 3, the learning data generation module (310) can generate a basic spherical virtual image by performing spherical transformation on the basic panoramic color image and the basic depth map image (S310).

[0079] The learning data generation module (310) can generate multiple learning spherical images by changing the setting information of the basic spherical virtual image (S320).

[0080] The learning data generation module (310) can generate multiple intermediate panoramic images by plane-converting multiple learning spherical images (S330).

[0081] The learning data generation module (310) can generate multiple learning panoramic images by inserting padding data into the periphery of the left and right side data portions of multiple intermediate panoramic images (S340). In the drawings below, the data-padded panoramic images are displayed as squares, and the non-padded panoramic images are displayed as rectangles with one corner folded over.

[0082] In the description of Fig. 3, the term "multiple learning spherical images" refers to multiple learning spherical panoramic images and / or multiple learning spherical depth map images. Accordingly, a process for generating multiple learning images can be performed based on the generation of multiple learning spherical images for each image type.

[0083] For example, the learning data generation module (310) can perform the above-described process on a panoramic color image. That is, the learning data generation module (310) can spherically transform a basic panoramic color image to generate a basic spherical color image, and change the setting information of the basic spherical color image to generate a plurality of learning spherical color images. The learning data generation module (310) can plane-transform a plurality of learning spherical color images to generate a plurality of intermediate panoramic color images. Thereafter, data padding can be performed on both sides of the plurality of intermediate panoramic color images to generate a plurality of learning panoramic color images.

[0084] As another example, the learning data generation module (310) may perform the above-described process on a panoramic depth map image. That is, the learning data generation module (310) may spherically transform a basic panoramic depth map image to generate a basic spherical depth map image, change setting information of a basic spherical depth map virtual image to generate a plurality of learning spherical depth map images, and then plane-transform the plurality of learning spherical depth map images to generate a plurality of intermediate panoramic depth map images. Thereafter, data padding may be performed on both sides of the plurality of intermediate panoramic depth map images to generate a plurality of learning panoramic depth map images.

[0085] FIG. 4 is a drawing illustrating an example of generating a plurality of learning panoramic color images based on a basic panoramic color image according to one embodiment disclosed in the present application, which will be further described with reference to FIG. 4.

[0086] As shown in Figure 4 (a), the basic panoramic color image can be converted into a spherical shape. That is, it can be converted into a spherical omnidirectional virtual image (hereinafter referred to as a "spherical image") that provides a spherical omnidirectional viewpoint by the learning data generation module (310). The spherical image generated from the basic panoramic color image is referred to as a "basic spherical color image."

[0087] Figure (b) illustrates an example of generating multiple intermediate spherical color images by changing the configuration information based on the basic spherical color image of Figure (a). Here, the configuration information may be information on at least one of a rotation axis, a rotation direction, or a rotation angle with respect to the basic spherical image.

[0088] Multiple learning spherical depth map images can be generated with different configuration information. That is, at least one of the rotation axis, rotation direction, and rotation angle of the configuration information of the multiple learning spherical depth map images is set differently with respect to the base spherical image.

[0089] Thereafter, as shown in Figure (c), the learning data generation module (310) can generate multiple intermediate panoramic color images by plane-transforming each of the multiple learning spherical color images. Here, the plane-transformation is the inverse transformation of the spherical transformation, and one intermediate panoramic color image can be generated by plane-transforming one learning spherical image.

[0090] Thereafter, as shown in Figure (d), the learning data generation module (310) can generate a learning image by performing data padding on the intermediate image. That is, the learning data generation module (310) can generate a learning panorama color image by performing data padding on the intermediate panorama color image, and generate a learning panorama depth map image by performing data padding on the intermediate panorama depth map image.

[0091] In this way, by changing the setting information of the basic spherical color image to generate multiple training spherical images, it provides the effect of generating a very large amount of training data from a single basic panoramic color image. This is because in the case of a panoramic image, in order to express spherical data in a rectangular space, the expression is projected differently depending on the spatial position, such as an equirectangular projection. Therefore, if the axis, rotation direction, and rotation angle of the spherical image are changed differently and then reversely transformed into a flat panoramic image, completely different training images can be generated.

[0092] Although the above Fig. 4 illustrates the case of a panoramic color image, the process of Fig. 4 is equally applicable to a panoramic depth map image to create a data set whose elements are a 1:1 pair of a learning panoramic color image and a learning panoramic depth map image.

[0093] That is, if N learning panorama color images are generated from a base panorama color image according to a change setting in the setting information such as the example illustrated in Fig. 4, N learning panorama depth map images are generated from the base panorama depth map image by performing a change setting in the same setting information. In this case, the base panorama color image and the base panorama depth map image become corresponding pairs captured with the same camera pose and resolution. In addition, the learning panorama color image and the learning panorama depth map image having the same setting information are set as a pair, thereby generating N learning data pairs.

[0094] The learning data generation module (310) can pad data on the left and right sides of the panoramic image. That is, since the left and right sides of the panoramic image are in contact with each other and displayed in a specific area, it is natural for the data on both sides to have continuity. However, in the case of the estimated panoramic depth map image generated by the neural network module (320), the continuity characteristic of the two sides is weak, and when the panorama is spherically transformed based on this, there is a problem that a large deviation occurs in a specific area where the two sides meet. Accordingly, the learning data generation module (310) generates a type of additional padding data on both sides of the learning data, uses this to generate the learning and estimated depth map of the neural network module (320), and then deletes the data padding area. As a result, the estimated panoramic depth map generated by padding and deleting data can secure more accurate data continuity on both sides, thereby enabling smoother and more precise implementation when implementing virtualization.

[0095] Hereinafter, FIG. 5 is a flowchart illustrating a data padding method for generating learning data according to one embodiment disclosed in the present application, and FIG. 6 is a drawing illustrating an example of a method for padding data for the explanation of FIG. 5, and these will be further described with reference thereto.

[0096] Referring to FIG. 5, the learning data generation module (310) can set padding data areas on the left and right sides of the middle panoramic image (S510).

[0097] The learning data generation module (310) can set right-side padding data based on left-side data of the middle panoramic image (S520), and can set left-side padding data based on right-side data of the middle panoramic image (S530).

[0098] The learning data generation module (310) can set an intermediate panoramic image with padding data set as a learning panoramic image (S540).

[0099] In this way, the learning data generation module (310) can generate a plurality of learning panorama color images by setting padding data to the left and right ends of each of a plurality of intermediate panorama color images.

[0100] The learning data generation module (310) can similarly be applied to depth map images, and can generate multiple learning panorama depth map images by setting padding data to the left and right ends of multiple intermediate panorama depth map images, respectively.

[0101] Figure 6 (a) illustrates an intermediate panoramic image, and Figure 6 (b) illustrates a data-padded learning panoramic image.

[0102] In Figure 6 (b), it can be seen that padding data (621, 622) is set on both sides of the middle panoramic image (610). An example is shown in which one right-side padding data cell is set on the right side of the right-side data 19, where the right-side padding data cell can be set based on the left-side data 11.

[0103] Figure 6 illustrates an example in which the value of the right-side padding data cell is set as the left-side data, but is not limited thereto. That is, the size and value of the padding data cell may be set in various ways.

[0104] FIG. 7 and FIG. 8 are examples of setting padding data based on the difference between data on both sides. FIG. 7 is a flowchart explaining a method of setting padding data according to one embodiment disclosed in the present application, and FIG. 8 is a drawing explaining an example of a method of padding data for the explanation of FIG. 7.

[0105] Referring to FIG. 7, the learning data generation module (310) sets padding data areas at the left and right ends of the intermediate panoramic image, and then sets padding data values ​​to be inserted into the padding data area based on the difference between data on the opposite end of the one-side cell adjacent to the corresponding padding data cell.

[0106] In FIG. 7, the learning data generation module (310) can set a padding data area to include a plurality of padding data cell columns (S710).

[0107] The learning data generation module (310) can set the data value for the padding data cell of the first column to have a median value between the difference between the data on one side adjacent to the data on the opposite side (S720).

[0108] Alternatively, the learning data generation module (310) may set the data value for the padding data cells of the remaining columns to have a median value between the difference between the data of the adjacent previous padding data cell and the data on the opposite side (S720).

[0109] Figure 8 (a) illustrates an intermediate panoramic image, and Figure 8 (b) illustrates a data-padded learning panoramic image.

[0110] In Figure 8 (b), it can be seen that padding data (821, 822) is set on both sides of the middle panoramic image (810). A total of three columns of right-side padding data cells (822) are set on the right side of the first column right-side data 19.

[0111] The right-side padding data of the first row, first column is set to 18, and is set to have a median value of 18 based on the difference 2 between the adjacent one-side data 19 and the opposite-side data 17. Here, the median is set to be rounded up or down as an integer value, but either rounding up or down can be set so that the difference from the adjacent one-side data is small based on the first column.

[0112] For example, the right-hand side data of the second row is 29, and the right-hand side padding data of the first column is the median 1.5, which is the difference between the right-hand side data 29 of the second row and the left-hand side data 26, which is 3. To make it smaller than the right-hand side data, 1.5 is rounded down to 1, and 28, which is 29 minus 1, becomes the right-hand side padding data value of the first column.

[0113] By setting the padding data to be smaller than the adjacent side in this way, the continuity of the padding data can be set more smoothly, which has the effect of reducing errors.

[0114] Meanwhile, if it is not the first column padding data, the data value is set to have the median value between the adjacent padding data and the data on the opposite side. That is, the right padding data of the first row, second column is set to 17, which is rounded up based on the difference 0.5 between the right padding data 18 of the first column and the data 17 on the opposite side, and is set by subtracting 1 from the right padding data 18 of the first column. This is to round up to 1 when the difference value is 1 and the median is 0.5, and to prevent the data from remaining unchanged when rounding down to 0.5.

[0115] Hereinafter, training will be described with reference to FIGS. 9 and 10.

[0116] FIG. 9 is a diagram illustrating a training module according to one embodiment disclosed in the present application.

[0117] The neural network module (320) learns based on a learning panorama color image for learning and a learning panorama depth map image for the color image.

[0118] The neural network module (320) is a deep learning-based image conversion learning model, which can generate an estimated panoramic depth map image based on conversion through a learning neural network for an input learning color image.

[0119] The estimated panorama depth map image is a depth map generated by a trained neural network module (320). This estimated panorama depth map image is different from the trained panorama depth map image, which is a ground truth depth map generated by measuring a distance using a Lidar sensor, etc. Therefore, the neural network module (320) can be trained so that the difference between the estimated panorama depth map image and the trained panorama depth map (GT depth map) image is reduced, and training for this neural network module (320) is performed by a training module (330).

[0120] The training module (330) can compare the estimated depth map generated by the neural network module (320) with the learning depth map (S1010), and train the neural network module (320) based on the difference (S1020).

[0121] The training module (330) can perform training based on spherical transformation. For example, the training module (330) can spherically transform the estimated panoramic depth map and the learning panoramic depth map, respectively, and then train the neural network module (320) based on the difference between the spherically transformed estimated depth map and the spherically transformed learning depth map.

[0122] As illustrated in FIG. 9, the training module (330) may include a spherical transformation module (331), a loss calculation module (332), and an optimization module (333).

[0123] The spherical transformation module (331) performs spherical transformation to transform a panoramic image into a spherical shape. The spherical transformation module (331) can input an estimated panoramic depth map image and a learning panoramic depth map image and spherically transform them respectively.

[0124] In Fig. 9, the learning panorama depth map image spherically transformed by the spherical transformation module (331) is represented as 'learning depth map*', and the spherically transformed estimated panorama depth map image is represented as 'estimated depth map*'.

[0125] In one embodiment, the spherical transformation module (331) can perform spherical transformation using the mathematical expression 1 below.

[0126]

[0127] The loss calculation module (332) can calculate the loss between the spherically transformed estimated depth map (estimated depth map*) and the spherically transformed learning depth map (learning depth map*).

[0128] That is, the loss calculation module (332) can quantify the difference (loss value) between the spherically transformed estimated depth map and the spherically transformed learning depth map. For example, the loss value determined by the loss calculation module (332) can be determined in a range between 0 and 1.

[0129] The optimizing module (333) can receive the loss calculated from the loss calculation module (332) and perform optimization by changing the parameters of the neural network in response to the loss.

[0130] For example, the optimization module (333) may perform optimization by adjusting the weight parameter W of the neural network. As another example, the optimization module (333) may perform optimization by adjusting at least one of the weight parameter W and the bias b of the neural network.

[0131] Various optimization methods can be applied to the optimization module (333). For example, the optimization module (333) can perform optimization using Batch Gradient Descent, Stochastic Gradient Descent, Mini-Batch Gradient Descent, Momentum, Adagrad, RMSprop, etc.

[0132] The loss calculation module (332) can determine the loss by applying multiple loss calculation methods and calculating a representative value from the result values.

[0133] The loss calculation module (332) can calculate a first loss function result between a spherically transformed estimated depth map and a spherically transformed learning depth map as a first loss calculation method.

[0134] The following mathematical expression 2 is a mathematical expression that explains an example of the first loss calculation expression.

[0135]

[0136] Here, T represents the number of samples, y represents the learned depth map, and y* represents the estimated depth map.

[0137] The loss calculation module (332) can calculate a second loss function result between a spherically transformed estimated depth map and a spherically transformed learning depth map using a second loss calculation method.

[0138] The following mathematical expression 3 is a mathematical expression that explains an example of the second loss calculation expression.

[0139]

[0140] Here, T is the number of samples, and d is the difference between the learned depth map and the estimated depth map in log space.

[0141] The loss calculation module (332) can calculate a third loss function result between the spherically transformed estimated depth map and the spherically transformed learning depth map using the third loss calculation method.

[0142] The following mathematical expression 4 is a mathematical expression that explains an example of the third loss calculation expression.

[0143]

[0144] Here, ytrue represents the learned depth map, and ypredicted represents the estimated depth map.

[0145] The loss calculation module (332) can calculate representative values ​​for the first loss function result to the third loss function result and determine them as losses (S1204). Here, the average, median, mode, etc. can be applied as representative values.

[0146] FIG. 11 is a diagram illustrating an architecture for generating an estimated panoramic depth map image according to one embodiment disclosed in the present application.

[0147] While the architecture of FIG. 2 illustrates a learning architecture for a neural network module (320), the architecture of FIG. 11 illustrates a depth map generation architecture that actually generates an estimated panoramic depth map image from a query panoramic color image.

[0148] In Fig. 11, the description of the neural network module (320) and the training module (330) is omitted here as it can be understood based on the previously described content.

[0149] FIG. 12 is a flowchart illustrating a method for generating an estimated panoramic depth map image according to one embodiment disclosed in the present application.

[0150] The depth image generation module (340) receives a query panorama color image and inserts padding data into both sides of the query panorama color image to perform data padding on both sides (S1210). An image of a query panorama color image on which data padding on both sides has been performed is referred to as a "padded query panorama color image."

[0151] Here, data padding can be performed corresponding to data padding performed in the learning data generation module (320).

[0152] The depth image generation module (340) provides a ‘padded query panorama color image’ to the neural network module (320).

[0153] The neural network module (320) generates a ‘padded estimated panoramic depth map image’ for the padded query panoramic color image and provides it to the depth image generation module (340).

[0154] The depth image generation module (340) can generate a final estimated panoramic depth map image by deleting the data padding area, i.e., padding data on both sides, from the ‘padded estimated panoramic depth map image’.

[0155] The above description assumes data padding on both sides, but is not limited to this. Therefore, for panoramic images capable of top-bottom panoramas, where the data at the top and bottom are continuous, top-bottom data padding can also be performed.

[0156] By performing data padding like this, we can see that the continuity of both sides of the estimated panoramic depth map image is significantly improved.

[0157] FIG. 13 is a diagram illustrating an example of an estimated panoramic depth map image generated according to one embodiment disclosed in the present application.

[0158] Figure 13(a) illustrates an example of an estimated panoramic depth map image generated according to the embodiment illustrated in Figure 12.

[0159] Figure (b) shows the two ends of Figure (a) adjacent to each other, and as shown, it can be seen that the data (color values) at both ends have continuity.

[0160] FIG. 14 is a block diagram illustrating a computing device according to one embodiment disclosed in the present application. FIG. 14 is intended to provide a general and simplified description of a suitable computing environment in which embodiments of the computing device (300) may be implemented.

[0161] The computing device (300) may include a processing unit (1403) and a system memory (1401).

[0162] A computing device may include multiple processing units that cooperate to execute a program. Depending on the exact configuration and type of the computing device, the system memory (1401) may be volatile (e.g., random access memory (RAM), non-volatile (e.g., read-only memory (ROM), flash memory, etc.), or a combination thereof. The system memory (1401) includes a suitable operating system (1402) for controlling the operation of the platform, such as the WINDOWS operating system from Microsoft Corporation. The system memory (1401) may also include one or more software applications, such as program modules, applications, etc.

[0163] The computing device may include additional data storage (1404), such as a magnetic disk, an optical disk, or a tape. Such additional storage may be removable storage and / or non-removable storage. Computer-readable storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. System memory (1401), storage (1404) are both examples of computer-readable storage media. Computer-readable storage media may include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, CD-ROMs, DVDs or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that stores desired information and can be accessed by the computing device (1400).

[0164] The input devices (1405) of the computing device may include, for example, a keyboard, a mouse, a pen, a voice input device, a touch input device, and comparable input devices. The output devices (1406) may include, for example, a display, a speaker, a printer, and other types of output devices. Since these devices are well known in the art, a detailed description thereof will be omitted.

[0165] A computing device may also include a communication device (1407) that allows the device to communicate with other devices over a network, such as a wired or wireless network, a satellite link, a cellular link, a local area network, and comparable mechanisms, for example in a distributed computing environment. The communication device (1407) is one example of a communication medium, which may have computer-readable instructions, data structures, program modules, or other data therein. By way of example and not limitation, communication media includes wired media, such as a wired network or direct-connection, and wireless media, such as acoustic, RF, infrared, and other wireless media.

[0166] The present application described above is not limited by the above-described embodiments and attached drawings, but is limited by the patent claims described below, and a person having ordinary skill in the technical field to which the present application pertains can easily understand that the configuration of the present application can be variously changed and modified within a scope that does not deviate from the technical idea of ​​the present application.

[0167] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0168] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.

[0169]

[0170] <Task Information>

[0171] Assignment Number: 20242420

[0172] Ministry Name: Ministry of SMEs and Startups

[0173] Research Management Specialist Organization: Startup Promotion Agency

[0174] Research Project Name: 1000+ Ultra-Gap Startup Incubation Project

[0175] Research Project Name: AI Big Data Technology Field

[0176] Host organization: 3RI Co., Ltd.

[0177] Research period: April 19, 2024 - December 1, 2024

[0178]

[0179] <Explanation of symbols>

[0180] 100: User terminal

[0181] 200: Image acquisition device

[0182] 210: Distance measuring device

[0183] 220: Camera

[0184] 300: Computing Device

[0185] 310: Learning data generation module

[0186] 320: Neural Network Module

[0187] 330: Training Module

[0188] 331: Old conversion module

[0189] 332: Loss calculation module

[0190] 333: Optimizing Module

[0191]

[0192] According to one embodiment of the present invention, a method for generating a depth map image and a computing device therefor can generate a panoramic depth map image based on a panoramic color image using a deep learning-based artificial neural network, and thus has high industrial applicability.

[0193] In addition, according to one embodiment of the present invention, a method for generating a depth map image and a computing device therefor input padding data to the left and right sides of a panoramic image, thereby minimizing an error between the side data of a panoramic depth map image generated by a deep learning-based artificial neural network, and thereby fundamentally preventing an error occurring in an area where both sides of the panoramic image meet when constructing a virtual space, thus having high industrial applicability.

Claims

1. As a computing device, Memory that stores one or more instructions; and A processor that executes one or more instructions stored in said memory. Including, The above processor, by executing one or more of the instructions, Generating a depth map image Computing device.

2. In paragraph 1, The above processor, by executing one or more of the instructions, A base panorama color image and a base panorama depth map image corresponding to the base panorama color image are provided, Generating a plurality of learning panorama color images and a plurality of learning panorama depth map images based on a spherical transformation of the above base panorama color image and the above base panorama depth map image, The above multiple learning panorama color images and multiple learning panorama depth map images are images with padding data inserted on both sides. Computing device.

3. In paragraph 2, The above processor, by executing one or more of the instructions, The above basic panoramic color image is spherically converted to generate a basic spherical color image, the setting information of the above basic spherical color image is changed to generate a plurality of learning spherical color images, and the above multiple learning spherical color images are plane-converted to generate a plurality of intermediate panoramic color images. Computing device.

4. In paragraph 3, The above processor, by executing one or more of the instructions, The above-mentioned basic panoramic depth map image is spherically transformed to generate a basic spherical depth map image, the setting information of the above-mentioned basic spherical depth map virtual image is changed to generate a plurality of learning spherical depth map images, and the above-mentioned multiple learning spherical depth map images are plane-transformed to generate a plurality of intermediate panoramic depth map images. Computing device.

5. In paragraph 4, The above setting information is, Information about at least one of a rotation axis, a rotation direction or a rotation angle for the above-mentioned basic spherical color image or the above-mentioned basic spherical depth map image, Computing device.

6. In paragraph 5, The above multiple learning spherical depth map images are, At least one of the rotation axis, the rotation direction and the rotation angle of the setting information is set differently for the above basic spherical depth map image. Computing device.

7. In paragraph 4, The above processor, by executing one or more of the instructions, Generating the plurality of learning panorama color images by setting padding data to the left and right ends of the plurality of intermediate panorama color images, respectively. Computing device.

8. In paragraph 7, The above padding data is, It is set based on the difference between the data of the adjacent cell and the data of the opposite side. Computing device.

9. In paragraph 1, The above processor, by executing one or more of the instructions, A padded query panorama color image is generated by setting a padding area for a query panorama color image and inserting padding data into the padding area, and an estimated panorama depth map image for the padded query panorama color image is generated using a pre-trained neural network, wherein the estimated panorama depth map image includes the padding area, and an estimated panorama depth map image corresponding to the query panorama color image is generated by deleting an area corresponding to the padding area from the estimated panorama depth map image. Computing device.

10. In paragraph 9, The above processor, by executing one or more of the instructions, By setting padding data on the left and right sides of the above query panorama color image, the padded query panorama color image is generated. Computing device.

11. In paragraph 10, The above padding data is, It is set based on the difference between the data of the adjacent cell and the data of the opposite side. Computing device.

12. In paragraph 11, The above neural network, Learned based on multiple learning panorama color images and multiple learning panorama depth map images, with padding data inserted on both sides. Computing device.

13. In paragraph 12, Multiple learning panorama color images, A basic spherical color image is generated by spherically converting a basic panoramic color image, a plurality of learning spherical color images are generated by changing the setting information of the basic spherical color image, a plurality of intermediate panoramic color images are generated by plane converting the plurality of learning spherical color images, and then padding data is set at the left and right ends of each of the plurality of intermediate panoramic color images. Computing device.

14. In paragraph 12, Multiple learning panorama depth map images, A base spherical depth map image is generated by spherically converting a base panoramic depth map image, a plurality of learning spherical depth map images are generated by changing setting information of the base spherical depth map image, a plurality of intermediate panoramic depth map images are generated by plane converting the plurality of learning spherical depth map images, and then padding data is set at the left and right ends of each of the plurality of intermediate panoramic depth map images, thereby generating Computing device.

15. In paragraph 13, The above setting information is, Information about at least one of a rotation axis, a rotation direction or a rotation angle for the above-mentioned basic spherical color image or the above-mentioned basic spherical depth map image, Computing device.

16. A method for generating a depth map image performed on a computing device, The above computing device, Memory that stores one or more instructions; and A processor that executes one or more instructions stored in said memory. Including, The above processor performs the depth map image generation method by executing the one or more instructions. How to generate a depth map image.

17. In paragraph 16, The above depth map image generation method is, A step of providing a base panorama color image and a base panorama depth map image corresponding to the base panorama color image; and A step of generating a plurality of learning panorama color images and a plurality of learning panorama depth map images based on a spherical transformation for the base panorama color image and the base panorama depth map image; comprising; The above multiple learning panorama color images and multiple learning panorama depth map images are images with padding data inserted on both sides. How to generate a depth map image.

18. In paragraph 17, The step of generating the plurality of learning panorama color images and the plurality of learning panorama depth map images comprises: A step of generating a basic spherical color image by spherically converting the above basic panoramic color image; and A step of generating a plurality of learning spherical color images by changing the setting information of the above-mentioned basic spherical color image, and generating a plurality of intermediate panoramic color images by plane-transforming the above-mentioned plurality of learning spherical color images; comprising; How to generate a depth map image.

19. In Article 18, The step of generating the plurality of learning panorama color images and the plurality of learning panorama depth map images comprises: A step of generating a basic spherical depth map image by spherically converting the above basic panoramic depth map image; and A step of generating a plurality of learning spherical depth map images by changing the setting information of the above-mentioned basic spherical depth map virtual image, and generating a plurality of intermediate panoramic depth map images by plane-transforming the above-mentioned plurality of learning spherical depth map images; further comprising; How to generate a depth map image.

20. In paragraph 19, The above setting information is, Information about at least one of a rotation axis, a rotation direction or a rotation angle for the above-mentioned basic spherical color image or the above-mentioned basic spherical depth map image, How to generate a depth map image.

21. In paragraph 20, The above multiple learning spherical depth map images are, At least one of the rotation axis, the rotation direction and the rotation angle of the setting information is set differently for the above basic spherical depth map image. How to generate a depth map image.

22. In paragraph 19, The step of generating the plurality of learning panorama color images and the plurality of learning panorama depth map images comprises: A step of generating the plurality of learning panorama color images by setting padding data to the left and right ends of the plurality of intermediate panorama color images, respectively; further comprising; How to generate a depth map image.

23. In paragraph 19, The above padding data is, It is set based on the difference between the data of the adjacent cell and the data of the opposite side. How to generate a depth map image.

24. In paragraph 16, The above depth map image generation method is, A step of setting a padding area for a query panorama color image and inserting padding data into the padding area to generate a padded query panorama color image; A step of generating an estimated panoramic depth map image for the padded query panoramic color image using the pre-trained neural network, wherein the estimated panoramic depth map image includes a padding region; and A step of generating an estimated panoramic depth map image corresponding to the query panoramic color image by deleting an area corresponding to a padding area in the estimated panoramic depth map image; comprising; How to generate a depth map image.

25. In paragraph 24, The step of generating the above padded query panorama color image is: A step of generating the padded query panorama color image by setting padding data to the left and right ends of the query panorama color image, respectively; comprising; How to generate a depth map image.

26. In paragraph 25, The above padding data is, It is set based on the difference between the data of the adjacent cell and the data of the opposite side. How to generate a depth map image.

27. In paragraph 26, The above neural network, Learned based on multiple learning panorama color images and multiple learning panorama depth map images, with padding data inserted on both sides. How to generate a depth map image.

28. In paragraph 27, Multiple learning panorama color images, A basic spherical color image is generated by spherically converting a basic panoramic color image, a plurality of learning spherical color images are generated by changing the setting information of the basic spherical color image, a plurality of intermediate panoramic color images are generated by plane converting the plurality of learning spherical color images, and then padding data is set at the left and right ends of each of the plurality of intermediate panoramic color images. How to generate a depth map image.

29. In paragraph 27, Multiple learning panorama depth map images, A base spherical depth map image is generated by spherically converting a base panoramic depth map image, a plurality of learning spherical depth map images are generated by changing setting information of the base spherical depth map image, a plurality of intermediate panoramic depth map images are generated by plane converting the plurality of learning spherical depth map images, and then padding data is set at the left and right ends of each of the plurality of intermediate panoramic depth map images, thereby generating How to generate a depth map image.

30. In paragraph 28, The above setting information is, Information about at least one of a rotation axis, a rotation direction or a rotation angle for the above-mentioned basic spherical color image or the above-mentioned basic spherical depth map image, How to generate a depth map image.

31. In a storage medium storing computer-readable instructions, The above instructions, when executed by a computing device, cause the computing device to: Performing an operation to generate a depth map image, Storage medium.

32. In paragraph 31, The above instructions, when executed by a computing device, cause the computing device to: An operation of providing a base panorama color image and a base panorama depth map image corresponding to the base panorama color image; and An operation of generating a plurality of learning panorama color images and a plurality of learning panorama depth map images based on a spherical transformation of the above base panorama color image and the above base panorama depth map image; is performed; The above multiple learning panorama color images and multiple learning panorama depth map images are images with padding data inserted on both sides. Storage medium.

33. In paragraph 31, The above instructions, when executed by a computing device, cause the computing device to: An operation of setting a padding area for a query panorama color image and inserting padding data into the padding area to generate a padded query panorama color image; An operation of generating an estimated panoramic depth map image for the padded query panoramic color image using the pre-trained neural network, wherein the estimated panoramic depth map image includes a padding region; and An operation of generating an estimated panoramic depth map image corresponding to the query panoramic color image by deleting an area corresponding to a padding area in the estimated panoramic depth map image; Storage medium.

Citation Information

Patent Citations

  • Panorama-based self-supervised learning scene point cloud completion data set generation method

    CN113808261A

  • Method and apparatus for converting image from 2d to 3D using panorama image

    KR101370718B1

  • Device for Inserting Glenoid Baseplate and its Method

    KR1020240085616A

  • Hole filling method for virtual 3 dimensional model and computing device therefor

    KR102551097B1

  • Depth map image generation method and computing device therefor

    KR102551467B1