Computing device for generating learning data

KR103022907B1Active Publication Date: 2026-09-233I INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
KR1020230064433
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-05-18
Publication Date
2026-09-23
Estimated Expiration
2041-06-23

Smart Images

  • Figure 112023055358971-PAT00006_ABST
    Figure 112023055358971-PAT00006_ABST
Patent Text Reader

Abstract

A computing device for generating training data according to one technical aspect of the present application may include: a training data generation module that generates training data by changing the setting information of a basic spherical virtual image generated by spherical transformation based on a basic RGB image and a basic depth map image; a neural network module that generates an estimated depth map by performing training based on the training data; a training module that trains the neural network module by comparing the estimated depth map with the training data; and a virtual image providing module that generates a spherical virtual image and provides it to a user terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present application relates to a computing device for generating learning data. Background Technology

[0002] Recently, virtual space implementation technologies are being developed that provide an online virtual space corresponding to the actual space, enabling users to experience being in the actual space without having to visit it in person.

[0004] To implement such a virtual space, it is necessary to acquire planar images of the actual space to be implemented and generate three-dimensional virtual images based on them to provide the virtual space.

[0006] In the case of such conventional technology, virtual images are provided based on flat images, but there is a limitation in that distance information cannot be known in the conventional virtual space, resulting in a lack of realism and three-dimensional information. Prior art literature

[0007] U.S. Patent Application No. 16 / 880143 The problem to be solved

[0008] One technical aspect of the present application is intended to solve the problems of the prior art described above, and according to one embodiment disclosed in the present application, it is intended to provide distance information for a virtual space.

[0009] According to one embodiment disclosed in the present application, the purpose is to generate various training data sets using a single RGB image and a distance map image thereof.

[0010] According to one embodiment disclosed in the present application, the purpose is to generate a depth map image from an RGB image based on learning using a neural network model.

[0012] The problems of the present application are not limited to those mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the description below. means of solving the problem

[0013] One technical aspect of the present application proposes a computing device for generating training data. The computing device may include a training data generation module that generates training data by changing the setting information of a basic spherical virtual image generated by spherical transformation based on a basic RGB image and a basic depth map image; a neural network module that generates an estimated depth map by performing training based on the training data; a training module that trains the neural network module by comparing the estimated depth map with the training data; and a virtual image providing module that generates a spherical virtual image and provides it to a user terminal.

[0015] The means for solving the above-mentioned problem do not enumerate all the features of the present application. Various means for solving the problem of the present application may be understood in more detail by referring to the specific embodiments in the following detailed description. Effects of the invention

[0016] According to the present application, there is one or more of the following effects.

[0017] According to one embodiment disclosed in the present application, there is an effect of being able to provide distance information for a virtual space.

[0018] According to one embodiment disclosed in the present application, there is an effect of being able to generate a plurality of training data sets using one RGB image and a distance map image thereof.

[0019] According to one embodiment disclosed in the present application, a depth map image can be generated from an RGB image based on learning using a neural network model, and by using this, a virtual space containing depth information can be provided using only RGB images.

[0020] According to one embodiment disclosed in this application, in calculating the loss used for training a neural network, by using a combination of multiple functions, the loss range can be reduced to a minimum.

[0022] The effects of the present application are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description in the claims. Brief explanation of the drawing

[0023] FIG. 1 is an exemplary drawing for explaining a system that provides a spherical virtual image based on a depth map image according to one embodiment disclosed in the present application. FIG. 2 is a block diagram illustrating a computing device according to one embodiment disclosed in the present application. FIG. 3 is a diagram illustrating a learning data generation architecture according to one embodiment disclosed in the present application. FIG. 4 is a drawing illustrating an equirectangular projection image according to one embodiment disclosed in the present application and a spherical virtual image generated using the same. FIG. 5 is a diagram illustrating a method for generating learning data according to one embodiment disclosed in the present application. FIG. 6 is a diagram illustrating an example of generating a large amount of training data set based on basic RGB and basic depth maps according to one example. FIG. 7 is a drawing illustrating an example of a neural network architecture according to one embodiment disclosed in the present application. FIG. 8 is a drawing illustrating another example of a neural network architecture according to one embodiment disclosed in the present application. FIG. 9 is a diagram illustrating a training method using a neural network according to one embodiment disclosed in the present application. FIG. 10 is a drawing illustrating the difference between an equirectangular projection image according to one embodiment disclosed in the present application and a spherical virtual image generated using the same. FIG. 11 is a drawing illustrating a training method by a training module according to one embodiment disclosed in the present application. FIG. 12 is a drawing illustrating a loss calculation method according to one embodiment disclosed in the present application. Figure 13 is a diagram illustrating a learning RGB, a learning depth map, an estimated depth map, and a difference image between the estimated depth map and the learning depth map according to one example. FIG. 14 is a drawing illustrating a spherical virtual image generation architecture according to one embodiment disclosed in the present application. FIG. 15 is a drawing illustrating a method of providing a spherical virtual image to a user according to one embodiment disclosed in the present application. FIG. 16 is a drawing for explaining a spherical transformation according to one embodiment disclosed in the present application. Specific details for implementing the invention

[0024] Preferred embodiments of the present application will be described below with reference to the attached drawings.

[0025] However, the embodiments of the present application may be modified in various different forms, and the scope of the present application is not limited to the embodiments described below. Furthermore, the embodiments of the present application are provided to more fully explain the present application to those skilled in the art.

[0027] The various embodiments of this application and the terms used therein are not intended to limit the technical features described in this application to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this application, each of phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as “first,” “second,” or “first” or “second” may be used simply to distinguish a component from another component and do not limit the components in any other aspect (e.g., importance or order). Where any (e.g., first) component is referred to as “coupled,” “connected,” or “connected” to another (e.g., second) component, with or without the terms “functionally” or “communicationly,” it means that said component may be connected to said other component directly or through a third component.

[0028] As used in this application, the term "module" refers to a unit that processes at least one function or operation, which may be implemented as software or as a combination of hardware and software.

[0029] Various embodiments of the present application may be implemented as software (e.g., a program) comprising one or more instructions stored in a storage medium readable by a machine (e.g., a user terminal (100) or a computing device (300)). For example, a processor (301) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium.

[0030] According to the embodiments, the method according to the various embodiments disclosed in this application may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CDROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0031] According to various embodiments, each component (e.g., a module or a program) of the components described above may include a singular or multiple entities. According to various embodiments, one or more of the components or operations among the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., a module or a program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the components of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to the integration. According to various embodiments, operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0032] Various flowcharts are disclosed to explain the embodiments of the present application, but these are for the convenience of explaining each step and do not necessarily mean that each step must be performed in the order of the flowchart. That is, each step in the flowchart may be performed simultaneously, in the order according to the flowchart, or in the reverse order of the flowchart.

[0034] FIG. 1 is an exemplary drawing for explaining a system that provides a spherical virtual image based on a depth map image according to one embodiment disclosed in the present application.

[0035] A system (10) that provides a spherical virtual image based on a depth map image may include a user terminal (100), an image acquisition device (200), and a computing device (300).

[0036] A user terminal (100) is an electronic device that can be used by a user to access a computing device (300), and includes, for example, a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a personal computer (PC), a tablet PC, an ultrabook, a wearable device (for example, a smartwatch, a smart glass, a head-mounted display), etc. However, in addition to these, the user terminal (100) may include electronic devices used for VR (Virtual Reality) and AR (Augmented Reality).

[0037] The image acquisition device (200) is a device that generates RGB images and / or depth map images, which are used to generate spherical virtual images.

[0038] In the illustrated example, the image acquisition device (200) is illustrated as being divided into a distance measuring device (210) and an imaging device (220), but this is exemplary and distance measurement and imaging may be performed using a single image acquisition device (200).

[0039] The imaging device (220) is a portable electronic device having a shooting function and generates an RGB image that is expressed in color for the subject area—that is, the area captured in the RGB image.

[0040] In other words, in this application specification, RGB images encompass all images expressed in color and are not limited to a specific mode of expression. Accordingly, images expressed in CMYK (Cyan Magenta Yellow Key) are also collectively referred to as RGB images in this application specification.

[0041] The imaging device (220) includes, for example, a mobile phone, a smartphone, a laptop computer, a PDA (personal digital assistants), a tablet PC, an ultrabook, a wearable device (for example, a smart glass terminal), etc.

[0042] The distance measuring device (210) is a device capable of generating depth information for the subject area and generating a depth map image.

[0043] In the present application specification, a depth map image encompasses an image containing depth information with respect to the subject space. That is, a depth map image refers to an image expressed as distance information from the point of capture to each point for each point in the captured subject space.

[0044] The distance measuring device (210) may include a predetermined sensor for measuring distance, such as a lidar sensor, an infrared sensor, an ultrasonic sensor, etc. Alternatively, the distance measuring imaging device (220) may include a stereo camera, a stereoscopic camera, a 3D depth camera, etc., which can measure distance information in place of the sensor.

[0045] The image generated by the imaging device (220) is called a basic RGB image, and the image generated by the distance measuring device (210) is called a basic depth map image. Since the basic RGB image generated by the imaging device (220) and the basic depth map image generated by the distance measuring device (210) are generated for the same subject area under the same conditions (e.g., resolution, etc.), they are matched 1:1 with each other.

[0046] The computing device (300) can receive a basic RGB image and a basic depth map image and perform learning. Here, the basic RGB image and the basic depth map image can be transmitted via a network.

[0047] The computing device (300) can generate a spherical virtual image based on the learning performed. Additionally, the computing device (300) provides the generated spherical virtual image to the user terminal (100). Here, the provision of the spherical virtual image can be performed in various forms, for example, by providing the spherical virtual image to be operated on the user terminal (100), or for another example, by providing a user interface for the spherical virtual image implemented in the computing device (300).

[0048] The provision of a spherical virtual image from a computing device (300) to a user terminal (100) can also be provided via a network.

[0049] The computing device (300) can convert the base RGB image and the base depth map image to generate a number of training RGB images and training depth map images. This utilizes a characteristic environment that uses spherical virtual images, and can generate a number of training RGB images and training depth map images by sphericalizing the base RGB image and the base depth map image and then making slight adjustments.

[0051] Hereinafter, with reference to FIGS. 2 to 15, various embodiments of the components constituting a system (10) that provides a spherical virtual image will be described.

[0053] FIG. 2 is a block diagram illustrating a computing device according to one embodiment disclosed in the present application.

[0054] The computing device (300) may include a processor (301), memory (302), and a communication unit (303).

[0055] The processor (301) controls the overall operation of the computing device (300). For example, the processor (301) can perform the functions of the computing device (300) described in the present disclosure by executing one or more instructions stored in memory (302).

[0056] The processor (301) can generate a spherical virtual image based on a basic RGB image transmitted from an imaging device (220) and a basic depth map image input from a distance measuring device (210).

[0057] The processor (301) may include a training data generation module (310) that generates various training data based on a basic RGB image and a basic depth map image, a neural network module (320) that performs training based on the training data, a training module (330) that trains the neural network module (320) by comparing the estimated depth map and the training depth map, and a virtual image providing module (340) that generates a spherical virtual image and provides distance information of the subject area to a user terminal.

[0058] The training data generation module (310) can generate multiple training data, namely training RGB images and training depth map images, by converting the basic RGB images and basic depth map images into spheres and adjusting them.

[0059] For example, the learning data generation module (310) can convert the basic RGB image transmitted from the imaging device (220) and the basic depth map image transmitted from the distance measuring device (210) into a sphere. The converted image can acquire various learning data by changing the rotation angle based on the various axes of the sphere image. At this time, the learning RGB image refers to the RGB image provided to the neural network module (320) for learning, and the learning depth map image refers to the depth map image provided to the neural network module (320) for learning. Accordingly, the learning RGB image is an image generated from the basic RGB image, and the learning depth map image is an image generated from the basic depth map image.

[0060] The neural network module (320) learns based on a learning RGB image and a learning depth map image for the same. For example, the learning depth map image is associated with the learning RGB image in a 1:1 manner. The learning depth map image is a ground truth depth map because it is generated by measuring the distance using a Lidar sensor, etc., for the subject area where the learning RGB image was generated—including a distance estimation method using a stereo camera. After learning based on the learning RGB image and the learning depth map image, the neural network module (320) can generate an estimated depth map image for the input RGB image based on the learned content.

[0061] The training module (330) can train the neural network module (320) based on the accuracy of the estimated depth map generated by the neural network module (320).

[0062] For example, the training module (330) can compare the estimated depth map generated by the neural network module (320) for the training RGB image with the training depth map—which is the actual depth map—and continuously train the neural network module (320) so that the difference between the estimated depth map and the training depth map becomes smaller.

[0063] The neural network module (320) receives a query RGB image as input and generates an estimated depth map. The virtual image providing module (340) can generate a spherical virtual image based on the query RGB image and the estimated depth map. The spherical virtual image may be an image provided from the computing device (300) to the user terminal (100), for example, a virtual space that can be implemented on the user terminal (100).

[0064] The memory (302) can store a program for processing and controlling the processor (301) and can store data that is input to or output from the computing device (300). For example, the memory (302) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), RAM (Random Access Memory), SRAM (Static Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), PROM (Programmable Read-Only Memory), magnetic memory, a magnetic disk, and an optical disk.

[0065] The communication unit (303) may include one or more modules that enable communication between the computing device (300) and another electronic device, such as a user terminal (100) or an image acquisition device (200), or between the computing device (300) and a network where the other electronic device is located.

[0067] FIG. 3 is a diagram illustrating a learning data generation architecture according to one embodiment disclosed in the present application.

[0068] The basic RGB image shown in Fig. 3 is transmitted from the imaging device (220), and the basic depth map image is transmitted from the distance measuring device (210).

[0069] The base RGB image and base depth map image here may be equirectangular projection images used in omnidirectional virtual reality. The various types of RGB images and depth map images described below may be equirectangular projection images used to generate an omnidirectional virtual space.

[0070] FIG. 4 is a drawing illustrating an equirectangular projection image according to one embodiment disclosed in the present application and a spherical virtual image generated using the same.

[0071] As shown in the example illustrated in figure (a), the equirectangular projection image can be converted by the computing device (300) into a spherical omnidirectional virtual image (hereinafter referred to as the “spherical virtual image”) that provides a spherical view.

[0072] Figure (b) illustrates an example of converting the equirectangular projection image of Figure (a) into a spherical virtual image. The spherical virtual image shown in Figure (b) is a spherical virtual image generated based on an RGB equirectangular projection image, and relates to an example in which the distance values ​​of each pixel are set equally. Meanwhile, Figure (b) shows the spherical virtual image from outside the spherical virtual image, but this is for convenience of explanation. Accordingly, the spherical virtual image may provide a virtual image in all directions, 360 degrees left and right and 360 degrees up and down, from inside the illustrated spherical virtual image.

[0073] Below, a panoramic image (e.g., a 2:1 panoramic image) is described as an example of an equirectangular projection image. The panoramic image has the advantage of being able to produce a full-dimensional image of space with a single shot and to be easily converted in a spherical transformation.

[0074] However, since the image transmitted from the image acquisition device (200) being a panoramic image is one example for carrying out the present application, the image transmitted from the image acquisition device (200) may be a general image(s) taken according to the user's use and convenience. And the computing device (300) can convert these general images into equirectangular projection images.

[0075] The training data generation module (310) receives a basic RGB image and a basic depth map image as input and can generate a basic spherical virtual image by converting them into a sphere.

[0076] The base depth map image may be a panoramic depth map image containing distance information for the subject area. The base depth map image is matched 1:1 with a base RGB image having the same subject area.

[0077] The training data generation module (310) can generate multiple training data, namely training RGB images and training depth map images, by transforming the basic spherical virtual image in various ways. The training data generation module (310) can provide the generated training RGB images to the neural network module (320) and provide the training depth map images to the training module (330).

[0078] That is, the training data generation module (310) can generate a basic spherical virtual image using a basic RGB image and a basic depth map image, and generate a number of training spherical images by changing the setting information of the basic spherical virtual image, and then generate various training data sets based on this.

[0079] One embodiment of such a learning data generation module (310) will be described with further reference to FIGS. 5 and 6.

[0081] FIG. 5 is a diagram illustrating a method for generating training data according to one embodiment disclosed in the present application, and FIG. 6 is a reference diagram for illustrating the method for generating training data in FIG. 5.

[0082] Referring to FIG. 5, the training data generation module (310) can generate a basic spherical virtual image using a basic RGB image and a basic depth map image (S501).

[0083] For example, the training data generation module (310) can generate a basic spherical virtual image by converting the basic RGB image and the basic depth map image into a sphere.

[0084] This is exemplified in step (a) of Fig. 6. That is, the basic spherical virtual image may be a basic spherical virtual image obtained from a basic RGB image and a basic depth map image.

[0085] In one embodiment of generating a basic spherical virtual image, the training data generation module (310) can generate a basic spherical virtual image using a basic RGB image and generate a basic spherical virtual image by associating depth information corresponding to each pixel of the basic spherical virtual image with a basic depth map image.

[0086] For example, the training data generation module (310) can generate the basic spherical virtual image by spherically transforming the basic RGB image so that the distance of each pixel is represented as an equal distance. The training data generation module (310) can store distance information corresponding to each pixel of the basic RGB image in association with the basic spherical virtual image using the basic depth map image. For example, the distance information corresponding to each pixel of the RGB image can be stored in a table containing identification information for each pixel and distance information for that pixel.

[0087] In one embodiment, when a change in setting information for a basic spherical virtual image occurs, the training data generation module (310) can change the storage of distance information in response to such change. For example, if a rotation occurs at a specific angle in a specific direction relative to a specific rotation axis for the basic spherical virtual image, distance information can be obtained from a table by reflecting the change in the position of pixels that changes due to such rotation.

[0088] In another embodiment of generating a basic spherical virtual image, the training data generation module (310) can generate a spherical virtual image for the basic RGB image and the basic depth map image, respectively.

[0089] For example, the training data generation module (310) can generate a first basic spherical virtual image in which the distance of each pixel is expressed as an equal distance by spherical converting a basic RGB image, and generate a second basic spherical virtual image in which each pixel is expressed as distance information by spherical converting a basic depth map image.

[0090] In this embodiment, the learning data generation module (310) can generate a pair of learning RGB images and learning depth map images by changing the setting information equally for a pair of first basic spherical virtual images and a second basic spherical virtual image—the pair of first and second basic spherical virtual images with the changed setting information correspond to the learning spherical virtual image—and performing a plane transformation thereon.

[0091] In another embodiment of generating a basic spherical virtual image, the training data generation module (310) can generate a three-dimensional basic depth map image by reflecting both color information and distance information in a single pixel. That is, in the above-described embodiments, as shown in the example in FIG. 6, the distance of each pixel is set constant so that the shape of the basic spherical virtual image is displayed as a round sphere, but in this embodiment, since each pixel is displayed according to distance information, it is displayed as a three-dimensional shape in three-dimensional space rather than a round sphere.

[0092] For example, the training data generation module (310) can obtain color information at each pixel from the basic RGB image and obtain distance information at each pixel from the basic depth map image to set color information and distance information for each pixel. The training data generation module (310) can generate a basic spherical virtual image by expressing the set color information and distance information for each pixel in three-dimensional coordinates. This basic spherical virtual image is expressed as a three-dimensional shape displayed in three-dimensional space rather than a circular shape.

[0093] The training data generation module (310) can generate multiple training spherical images by changing the setting information of the basic spherical virtual image (S502). For example, the setting information may include the rotation axis, rotation direction, or rotation angle of the spherical image.

[0094] For example, the learning data generation module (310) can generate a plurality of learning spherical images from a basic spherical virtual image by changing at least one of the rotation axis, rotation direction, or rotation angle with respect to the basic spherical virtual image.

[0095] Step (b) of Fig. 6 illustrates an example of generating multiple training spherical images by changing the setting information of a basic spherical virtual image.

[0096] The training data generation module (310) can generate a plurality of training data sets—wherein a training data set means a pair of training RGB images and a training depth map image that matches one-to-one with them—by performing a plane transformation on a plurality of training spherical images (S503). Herein, the plane transformation is the inverse transformation of the spherical transformation, and a set of training RGB images and training depth map images can be generated by performing a plane transformation on a single training spherical image.

[0097] In this way, generating multiple training square images by changing the basic square virtual image setting information provides the effect of generating a large amount of training data from a single basic square image. That is, the accurate computational power of the neural network module (320) is based on a large amount of training data, but in practice, it is difficult to obtain a large amount of training data. However, in one embodiment of the present application, a large number of training square images can be generated by applying various transformations based on the basic square virtual image, and also, a large amount of training data set can be easily obtained by inverse transformation.

[0098] A number of training RGB images and training depth map images generated in this way can be provided to a neural network module (320) and used as training information.

[0100] FIG. 7 is a drawing illustrating one example of a neural network architecture according to one embodiment disclosed in the present application, and FIG. 8 is a drawing illustrating another example of a neural network architecture according to one embodiment disclosed in the present application.

[0101] For ease of explanation, the neural network module (320) illustrated in FIGS. 7 and 8 is described as being implemented using the computing device (300) described in FIG. 2. That is, it may be implemented by the execution of at least one instruction performed by the memory (302) and processor (301) of the computing device (300). However, the neural network module (320) may also be used in any other suitable device(s) and any other suitable system(s). Additionally, the neural network module (320) is described as being used to perform image processing-related tasks. However, the neural network module (320) may also be used to perform any other suitable tasks, including non-image processing tasks.

[0102] The neural network module (320) learns based on a learning RGB image and a depth map image for learning.

[0103] The neural network module (320) is a deep learning-based image conversion learning model and can generate an estimated depth map image based on conversion through a learning neural network for an input learning RGB image.

[0104] The neural network module (320) can be represented as a mathematical model using nodes and edges. The neural network module (320) may be an architecture of a Deep Neural Network (DNN) or an n-layer neural network. The DNN or n-layer neural network may correspond to Convolutional Neural Networks (CNN), a Convolutional Neural Network (CNN) based on HRNet (Deep High-Resolution Network), Recurrent Neural Networks (RNN), Deep Belief Networks, Restricted Boltzmann Machines, etc.

[0105] For example, the neural network module (320) can receive a training RGB image as input and generate an estimated depth map image therefrom, as shown in the example illustrated in FIG. 7. In this example, in the initial operation, the neural network module (320) has no learned content, so it can generate an estimated depth map image based on random values ​​at each node of the neural network. The neural network module (320) can improve the accuracy of the estimated depth map by repeatedly performing feedback training on the generated estimated depth map image.

[0106] As another example, the neural network module (320) receives a set of training RGB images and corresponding training depth maps as input, as shown in FIG. 8, learns the association between the RGB images and the depth map images based on the learning, and generates an estimated depth map image for the input training RGB based on such association. In this example as well, feedback training on the estimated depth map image can be repeatedly performed to improve the accuracy of the estimated depth map.

[0108] FIG. 9 is a drawing illustrating a training method using a neural network according to one embodiment disclosed in the present application, and will be described below with further reference to FIG. 9.

[0109] The neural network module (320) receives a training RGB image as input and generates a Predicted Depth map image based on the learned content (S901).

[0110] The learning RGB image refers to the RGB image provided to the neural network module (320) for learning. The learning depth map image refers to the depth map image provided to the neural network module (320) or the training module (330) for learning. The learning depth map image is associated with the learning RGB image in a 1:1 ratio. Since the learning depth map image is generated by measuring the distance using a Lidar sensor or the like for the subject area where the learning RGB image was generated, it is a ground truth depth map.

[0111] The neural network module (320) is a deep learning-based image conversion learning model and can generate an estimated depth map image based on conversion through a learning neural network for an input learning RGB image.

[0112] Afterwards, the neural network module (320) performs learning through the training process described below (S902). As described above, the neural network module (320) can perform learning on a number of learning RGB images and learning depth maps generated by the learning data generation module (310), so its accuracy can be easily increased.

[0113] The estimated depth map image is a depth map generated by the learned neural network module (320). This estimated depth map image differs from the learned depth map image, which is a ground truth depth map generated by measuring the distance using a Lidar sensor or the like. Therefore, the neural network module (320) can be trained so that the difference between this estimated depth map image and the learned depth map (ground truth depth map) image becomes smaller, and the training of this neural network module (320) is performed by the training module (330).

[0114] The training module (330) can compare the estimated depth map generated by the neural network module (320) with the training depth map and train the neural network module (320) based on the difference.

[0115] In one embodiment, the training module (330) can perform training based on a square transformation. For example, after the training module (330) performs a square transformation on the estimated depth map and the learning depth map, the neural network module (320) can train based on the difference between the square-transformed estimated depth map and the square-transformed learning depth map.

[0116] In this embodiment, the training RGB image, the estimated depth map image, and the training depth map image may all be equirectangular projection images. That is, since these equirectangular projection images are used in a spherically transformed state, in order to more accurately determine the difference in usage between the estimated depth map image and the training depth map image, the training module (330) performs training by comparing the estimated depth map and the training depth map after each of them has been spherically transformed. This will be explained further with reference to FIG. 10.

[0118] FIG. 10 is a drawing illustrating the difference between an equirectangular projection image according to one embodiment disclosed in the present application and a spherical virtual image generated using the same.

[0119] Areas A (1010) and B (1020) have the same area and shape in a spherical virtual image, but when converted to an equirectangular projection image, areas A’ (1011) and B’ (1021) have different areas and shapes. This is due to the conversion between the spherical virtual image and the planar equirectangular projection image (panoramic image).

[0120] Accordingly, the training module (330) can increase the accuracy of the training and, accordingly, increase the accuracy of the estimated depth map image by performing training after spherical transformation of the estimated depth map and the training depth map, respectively.

[0122] FIG. 11 is a drawing illustrating a training method using a training module according to one embodiment disclosed in the present application, and will be described below with reference to FIG. 8 and FIG. 11.

[0123] As illustrated in FIG. 8, the training module (330) may include a spherical conversion module (331), a loss calculation module (332), and an optimizing module (333).

[0124] Referring further to FIG. 11, the spherical transformation module (331) performs a spherical transformation so that the equirectangular projection image corresponds to the spherical transformed image. The spherical transformation module (331) receives an estimated depth map image and a training depth map image as inputs and can perform a spherical transformation on each of them (S1101).

[0125] In FIG. 8, the learning depth map image converted into a sphere by the sphere conversion module (331) is labeled as the 'learning depth map*', and the estimated depth map image converted into a sphere is labeled as the 'estimated depth map*'.

[0126] In one embodiment, the spherical conversion module (331) can perform spherical conversion using the following mathematical formula 1.

[0127] [Mathematical Formula 1]

[0128]

[0129] A detailed explanation of mathematical formula 1 can be easily understood by referring to the explanation shown in Fig. 16.

[0130] The loss calculation module (332) can calculate the loss between the square-transformed estimated depth map (estimated depth map*) and the square-transformed training depth map (training depth map*) (S1102).

[0131] That is, the loss calculation module (332) can quantify the difference (loss value) between the spherically transformed estimated depth map and the spherically transformed learning depth map. For example, the loss value determined by the loss calculation module (332) can be determined in the range between 0 and 1.

[0132] The optimizing module (333) receives the loss calculated from the loss calculation module (332) and can perform optimization by changing the parameters of the neural network in response to the loss (S1103).

[0133] For example, the optimizing module (333) can perform optimization by adjusting the weight parameter W of the neural network. For another example, the optimizing module (333) can perform optimization by adjusting at least one of the weight parameter W and the bias b of the neural network.

[0134] Various types of optimizing methods can be applied to the optimizing module (333). For example, the optimizing module (333) can perform optimization using Batch Gradient Descent, Stochastic Gradient Descent, Mini-Batch Gradient Descent, Momentum, Adagrad, RMSprop, etc.

[0136] FIG. 12 is a drawing illustrating a loss calculation method according to one embodiment disclosed in the present application.

[0137] In one embodiment illustrated in FIG. 12, the loss calculation module (332) may apply a plurality of loss calculation methods and determine the loss by calculating a representative value from the resulting values.

[0138] Referring to FIG. 12, the loss calculation module (332) can calculate the result of a first loss function between a square-transformed estimated depth map and a square-transformed learning depth map using a first loss calculation method (S1201).

[0139] The following mathematical formula 2 is a mathematical formula that explains an example of the first loss calculation formula.

[0140] [Mathematical Formula 2]

[0141]

[0142] Here, T represents the number of samples, y represents the training depth map, and y* represents the estimated depth map.

[0143] The loss calculation module (332) can calculate the result of the second loss function between the square-transformed estimated depth map and the square-transformed learning depth map using the second loss calculation method (S1202).

[0144] The following mathematical formula 3 is a mathematical formula that explains an example of the second loss calculation formula.

[0145] [Mathematical Formula 3]

[0146]

[0147] Here, T is the number of samples, and d is the difference between the training depth map and the estimated depth map in log space.

[0148] The loss calculation module (332) can calculate the result of the third loss function between the square-transformed estimated depth map and the square-transformed learning depth map using the third loss calculation method (S1203).

[0149] The following mathematical formula 4 is a mathematical formula that explains an example of the third loss calculation formula.

[0150] [Mathematical Formula 4]

[0151]

[0152] Here, y true is the learning depth map, y predicted represents the estimated depth map.

[0153] The loss calculation module (332) can calculate a representative value for the first loss function result to the third loss function result and determine it as a loss (S1204). Here, the mean, median, mode, etc., can be applied as the representative value.

[0155] FIG. 13 is a diagram illustrating a learning RGB image, a learning depth map image, an estimated depth map image, and a difference image between the estimated depth map image and the learning depth map image according to one example.

[0156] Figure (a) illustrates an example of a training RGB image input to a neural network module (320).

[0157] Figure (b) illustrates an example of an estimated depth map image generated by a neural network module (320) that receives a training RGB image as input.

[0158] Figure (c) illustrates an example of a training depth map image that provides actual depth values ​​corresponding to a training RGB image.

[0159] Figure (d) illustrates a difference image between an estimated depth map image and a training depth map image. However, Figure (d) is for the purpose of providing an intuitive explanation, and as previously mentioned, in one embodiment disclosed in this application, when the training module (330) calculates the difference after performing a spherical transformation on the estimated depth map image and the training depth map image, it may be displayed differently from Figure (d).

[0161] FIG. 14 is a drawing illustrating a spherical virtual image generation architecture according to one embodiment disclosed in the present application, and FIG. 15 is a drawing illustrating a method of providing a spherical virtual image to a user according to one embodiment disclosed in the present application.

[0162] With reference to FIGS. 14 and 15, a method for providing a spherical virtual image to a user according to one embodiment disclosed in the present application will be described.

[0163] When the neural network module (320) receives a query RGB image, it generates an estimated depth map corresponding to the query RGB image using a neural network trained as described above (S1501).

[0164] Here, the query RGB image is an RGB image used to generate a spherical virtual image, and is an image that does not have a corresponding ground truth map. Therefore, an estimated depth map is generated using a neural network module (320) and used to generate a spherical virtual image.

[0165] The neural network module (320) provides the generated estimated depth map to the virtual image providing module (340).

[0166] The virtual image providing module (340) can generate a spherical virtual image based on the query RGB image and the estimated depth map image provided by the neural network module (320) (S1502).

[0167] For example, the virtual image providing module (340) can check the estimated depth map generated by the neural network module (320) and generate a spherical virtual image using the query RGB image and the estimated depth map.

[0168] Here, spherical virtual images refer collectively to images intended to provide a virtual space that users can experience.

[0169] For example, a spherical virtual image is generated based on a query RGB image, and for each pixel of such a spherical virtual image, distance information from each pixel included in an estimated depth map image may be included. Figure 4(b) illustrates an example of such a spherical virtual image, and Figure 4(b) illustrates an example in which the virtual image is displayed in the form of a perfect sphere, with each pixel indicated as being at the same distance. Even though it is displayed in this way, since it includes distance information for each pixel, distance information can be obtained for each pixel within the spherical virtual image.

[0170] As another example, a spherical virtual image can display the position and color of each pixel in a 3D coordinate system using color information for each pixel obtained from a query RGB image and distance information for each pixel obtained from an estimated depth map image. Another example of such a spherical virtual image can be displayed as a three-dimensional space displayed in a 3D coordinate system.

[0171] That is, the spherical virtual image may include distance information for at least one point (e.g., a pixel) included in the virtual image. Here, the distance information is determined based on an estimated depth map.

[0172] The virtual image providing module (340) can provide a spherical virtual image to the user. For example, the virtual image providing module (340) can provide a user interface to the user terminal (100) that includes an access function to the spherical virtual image.

[0173] The virtual image providing module (340) can receive a user request from a user through a user interface (S1503). For example, the virtual image providing module (340) can receive a request to check the distance to at least one point within a spherical virtual image, i.e., a user query. In response to the user query, the virtual image providing module (340) can check the distance information for the point within the spherical virtual image and provide it to the user. For example, in the spherical virtual image provided to the user terminal (100), the user can set a desired object or location in space, and the virtual image providing module (340) can check the distance information and provide it to the user terminal (100).

[0175] The present application described above is not limited by the aforementioned embodiments and attached drawings, but is limited by the claims set forth below, and it is readily apparent to those skilled in the art to which the present application belongs that the configuration of the present application can be varied and modified within the scope of the technical concept of the present application. Explanation of the symbols

[0176] 100 : User terminal 200 : Image acquisition device 300: Computing device

Claims

Claim 1 A computing device for generating training data, comprising: a training data generation module that generates training data by changing the setting information of a basic spherical virtual image generated by spherical transformation based on a basic RGB image and a basic depth map image; a neural network module that generates an estimated depth map by performing training based on the training data; a training module that trains the neural network module by comparing the estimated depth map with the training data; and a virtual image providing module that generates a spherical virtual image and provides it to a user terminal; wherein the training module is characterized by spherical transforming the estimated depth map and the training data, and training the neural network module based on the difference between the spherical transformed estimated depth map and the spherical transformed training data. Claim 2 A computing device for generating learning data, wherein, in paragraph 1, the setting information is information regarding at least one of a rotation axis, a rotation direction, or a rotation angle for the base spherical virtual image. Claim 3 A computing device for generating training data, wherein the training data is a training RGB image and a training depth map image generated by changing the setting information of the basic spherical virtual image based on the basic RGB image and the basic depth map image. Claim 4 A computing device for generating training data according to paragraph 2, wherein the training data generation module generates a plurality of training spherical images from a base spherical virtual image by changing at least one of the rotation axis, rotation direction, or rotation angle with respect to the base spherical virtual image, and generates a plurality of training RGB images and a plurality of training depth map images that are each matched 1:1 by planar transformation of each of the plurality of training spherical images. Claim 5 A computing device for generating training data, wherein, in paragraph 4, the training RGB image and the training depth map image are equirectangular projection images. Claim 6 A computing device for generating training data according to claim 1, wherein the training data generation module generates a basic spherical virtual image in which the distance of each pixel is expressed as an equal distance by spherical transforming the basic RGB image, and stores distance information corresponding to each pixel of the basic RGB image in association with the basic spherical virtual image using the basic depth map image. Claim 7 A computing device for generating training data, wherein the training data generation module generates a first basic spherical virtual image in which the distance of each pixel is expressed as an equal distance by spherical transforming the basic RGB image, and generates a second basic spherical virtual image in which each pixel is expressed as distance information by spherical transforming the basic depth map image. Claim 8 A computing device for generating training data according to claim 1, wherein the training data generation module obtains color information at each pixel from the basic RGB image and obtains distance information at each pixel from the basic depth map image, sets color information and distance information for each pixel, and generates the basic spherical virtual image by expressing the set color information and distance information for each pixel in three-dimensional coordinates. Claim 9 A computing device for generating training data, wherein, in claim 1, the neural network module repeatedly performs feedback training using the estimated depth map image. Claim 10 delete

Citation Information

Patent Citations

  • Employing three-dimensional (3D) data predicted from two-dimensional (2D) images using neural networks for 3D modeling applications and other applications

    US20190026957A1