Method and system for inpainting true ortho image

The use of a machine learning model to inpaint occluded areas in true ortho images addresses the duplication issue in 3D models, enhancing the visual quality and accuracy of 3D representations.

WO2025165174A1PCT designated stage Publication Date: 2025-08-07NAVER CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/099049
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-02
Filing Date
2025-01-16
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

True ortho images, created from aerial photographs, suffer from occlusion areas that cause duplication of objects or shapes in 3D models, degrading their quality.

Method used

A method and system using a machine learning model to inpaint occluded areas in true ortho images, generating high-quality images by inpainting target regions based on mask information, thereby eliminating duplication issues.

Benefits of technology

The inpainted images improve the visual quality of 3D models by removing duplication, resulting in cleaner and more accurate 3D representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025099049_07082025_PF_FP_ABST
    Figure KR2025099049_07082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for inpainting a true ortho image, performed by at least one processor. This method comprises the steps of: receiving a first true ortho image of a target region and first mask information associated with a target area from among the target region; and on the basis of the first true ortho image and the first mask information, generating a second true ortho image corresponding to an image in which the target area is inpainted, by using a machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Inpainting method and system for real-time video

[0001] The present disclosure relates to a method and system for inpainting a sedation image, and more particularly, to a method and system for generating a sedation image in which a target area is inpainted based on a sedation image of a target area and mask information associated with the target area among the target area, using a machine learning model.

[0002] A true ortho image is a vertical image created using a digital surface model, minimizing distortion caused by changes in the height of terrain features. This image provides the user with a view similar to what they would see from a vertical perspective. These true ortho images can be created using aerial photographs taken from above the target area using drones, satellites, or other means.

[0003] However, because true aerial photography images are viewed from above, they can create occlusion areas not visible in aerial photographs. When using these true aerial photography images to create 3D models, such as digital twins, the occlusion areas can cause duplicate objects or shapes in the 3D model. This can degrade the quality of the 3D model.

[0004] The present disclosure provides a method and system (device) for inpainting a true-life image to solve the above-mentioned problems.

[0005] The present disclosure can be implemented in various ways, including as a method, a device (system), or a computer program stored on a readable storage medium.

[0006] According to one embodiment of the present disclosure, a method for inpainting a true ortho image may include the steps of receiving a first true ortho image of a target region and first mask information associated with a target region among the target region, and using a machine learning model, generating a second true ortho image corresponding to an image in which the target region is inpainted based on the first true ortho image and the first mask information.

[0007] A computer program stored in a computer-readable recording medium may be provided to execute a method according to one embodiment of the present disclosure on a computer.

[0008] A system according to one embodiment of the present disclosure comprises a communication module, a memory, and at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory, wherein the at least one program may include instructions for receiving a first true ortho image of a target region and first mask information associated with a target region among the target region, and generating a second true ortho image corresponding to an image in which the target region is inpainted based on the first true ortho image and the first mask information using a machine learning model.

[0009] According to some embodiments of the present disclosure, by inpainting occluded areas or areas associated therewith in a true-image image, the doubling issue, which occurs when identical objects or shapes in the occluded areas of the true-image image are duplicated in a 3D model, can be eliminated. Accordingly, by combining a 3D model with a true-image image in which occluded areas, etc., have been inpainted, a high-quality true-image image with the doubling issue eliminated can be provided.

[0010] According to some embodiments of the present disclosure, a machine learning model can be used to inpaint specific objects within a real-time video or other type of image. This can remove unnecessary portions or correct distortions in the video or image, resulting in a visually clean and high-quality video or image.

[0011] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs (referred to as “one skilled in the art”) from the description of the claims.

[0012] Embodiments of the present disclosure will be described below with reference to the accompanying drawings, wherein like reference numerals represent similar elements, but are not limited thereto.

[0013] FIG. 1 illustrates an example of a true image with a target area inpainted according to one embodiment of the present disclosure.

[0014] FIG. 2 is a block diagram showing the internal configuration of an information processing system according to one embodiment of the present disclosure.

[0015] FIG. 3 is a diagram showing the internal configuration of a processor of an information processing system according to one embodiment of the present disclosure.

[0016] FIG. 4 is a drawing showing an example of a method for inpainting a true image according to one embodiment of the present disclosure.

[0017] FIG. 5 is a diagram showing an example of a first true image and mask information according to one embodiment of the present disclosure.

[0018] FIG. 6 is a diagram illustrating an example of an inpainted second true image according to one embodiment of the present disclosure.

[0019] FIG. 7 is a diagram showing an example of a three-dimensional city model generated using a true image according to one embodiment of the present disclosure.

[0020] FIG. 8 is a diagram illustrating an example of a three-dimensional city model in which a texture image of a building model is inpainted according to one embodiment of the present disclosure.

[0021] FIG. 9 is a diagram illustrating an example of a three-dimensional city model in which a texture image of a building model is inpainted according to one embodiment of the present disclosure.

[0022] FIG. 10 is a flowchart illustrating an example of a method for inpainting a true image according to one embodiment of the present disclosure.

[0023] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the attached drawings. However, in the following description, specific descriptions of widely known functions or configurations will be omitted if they may unnecessarily obscure the gist of the present disclosure.

[0024] In the attached drawings, identical or corresponding components are assigned the same reference numerals. Furthermore, in the description of the embodiments below, duplicate descriptions of identical or corresponding components may be omitted. However, even if a description of a component is omitted, it is not intended that such component is not included in any embodiment.

[0025] The advantages and features of the disclosed embodiments, and methods for achieving them, will become clearer with reference to the embodiments described below, along with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure the completeness of the disclosure and to fully inform those skilled in the art of the scope of the invention.

[0026] The terms used in this specification will be briefly explained, followed by a detailed description of the disclosed embodiments. The terms used in this specification have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of engineers working in the relevant field, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.

[0027] In this specification, singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, plural expressions include singular expressions unless the context clearly indicates otherwise. When a part of the specification is said to include a component, this does not exclude other components, but rather implies that other components may be included, unless otherwise specifically stated.

[0028] Also, the term 'module' or 'part' used in the specification means a software or hardware component, and the 'module' or 'part' performs certain roles. However, the 'module' or 'part' is not limited to software or hardware. The 'module' or 'part' may be configured to reside on an addressable storage medium and may be configured to execute one or more processors. Thus, as an example, the 'module' or 'part' may include at least one of components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, or variables. The functionality provided within the components and 'modules' or 'parts' may be combined into a smaller number of components and 'modules' or 'parts', or further separated into additional components and 'modules' or 'parts'.

[0029] According to one embodiment of the present disclosure, a 'module' or 'unit' may be implemented as a processor and a memory. 'Processor' should be broadly construed to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, and the like. In some circumstances, a 'processor' may also refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), and the like. A 'processor' may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors in conjunction with a DSP core, or any other such combination of configurations. In addition, 'memory' should be broadly construed to include any electronic component capable of storing electronic information. 'Memory' may refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage, registers, etc. Memory is said to be in electronic communication with the processor if the processor can read information from, and / or write information to, the memory. Memory integrated in a processor is in electronic communication with the processor.

[0030] In the present disclosure, the "system" may include, but is not limited to, at least one of a server device and a cloud device. For example, the system may be comprised of one or more server devices. As another example, the system may be comprised of one or more cloud devices. As yet another example, the system may be configured and operated by a combination of a server device and a cloud device.

[0031] In the present disclosure, 'display' may refer to any display device associated with a computing device, for example, any display device capable of displaying any information / data controlled by or provided from the computing device.

[0032] In the present disclosure, 'each of the plurality of As' or 'each of the plurality of As' may refer to each of all components included in the plurality of As, or may refer to each of some components included in the plurality of As.

[0033] In this disclosure, a "machine learning model" may include any model used to infer an answer to a given input. In one embodiment, the machine learning model may include an artificial neural network model including an input layer, multiple hidden layers, and an output layer. Here, each layer may include multiple nodes. In this disclosure, each of the multiple machine learning models is described as a separate machine learning model, but this is not limited thereto, and some or all of the multiple machine learning models may be implemented as a single machine learning model. Furthermore, a single machine learning model may include multiple machine learning models. In this disclosure, the terms "machine learning model" and "artificial neural network model" may be used interchangeably to refer to the same or similar models.

[0034] In the present disclosure, a "true ortho image" may refer to a vertical image created by minimizing distortion due to height changes of terrain features using a digital surface model among images that provide a view similar to that seen by a user from above in a vertical direction. In addition, an "ortho image" may refer to an image that is converted into an image that looks like all objects when viewed from a vertical direction by correcting geometric distortion information due to terrain ups and downs, such as height differences or inclination, from aerial photographs taken by aircraft such as drones, aircraft, or satellites.

[0035] In the present disclosure, "inpainting" refers to an image processing technique that removes or restores damaged portions from an image. Here, the damaged portions may include missing areas, damaged pixels, noise, etc. in the image. In the present disclosure, inpainting a target area may refer to visually improving the quality of an image by removing or restoring at least a portion of the target area based on the surrounding area.

[0036] FIG. 1 illustrates an example of a true ortho image in which a target area (122) is inpainted according to one embodiment of the present disclosure. In one embodiment, the true ortho image can be used not only in general map services, precision map production, etc., but also in the creation of 3D spatial maps or 3D models for digital twin environments. For example, the true ortho image can be combined with a digital elevation model and / or a 3D building model to create a 3D city model.

[0037] The first image (110) is an example of an image that combines a first tectonic image of a target area and a digital elevation model representing the terrain. Referring to the first image (110), since the surface corresponding to the river area in the digital elevation model is located below the surrounding terrain, when the bridge area of ​​the tectonic image is combined with the river area, the bridge area may be displayed as being bent downward. Accordingly, a bridge located over a river in the first image (110) may also be displayed as being bent downward.

[0038] The second image (120) is an example of an image that combines a first tactic image of the target area, a digital elevation model, and a 3D building model including information such as a bridge. Referring to the second image (120), a 3D model of a bridge, etc. may be combined with the first image (110) that combines the first tactic image and the digital elevation model. In this case, since the 3D bridge model is added again to the bridge area (i.e., the target area (122)) on the first image (110), a doubling issue, such as the bridge being displayed twice over a river in the target area (122), may occur. This doubling issue may mainly occur due to an occlusion area by one or more 3D objects existing in the target area in the tactic image. This doubling issue may cause a perception distortion of the corresponding area on the tactic image and may deteriorate the visual quality of the 3D model combined with the corresponding area.

[0039] According to some embodiments of the present disclosure, the third image (130) is an example that represents an image that combines a second anatomical image in which a target area (122) of a target area is inpainted, a digital elevation model, and a three-dimensional building model. In one embodiment, a processor (e.g., a processor of an information processing system described below in FIG. 2) may receive a first anatomical image of the target area and first mask information associated with the target area (122) of the target area. In addition, the processor may generate a second anatomical image corresponding to an image in which the target area (122) is inpainted based on the first anatomical image and the first mask information using a machine learning model. Here, the machine learning model may be pre-trained to input an image and mask information, and to inpaint an area corresponding to the mask information within the image to output a visually enhanced image. An example of generating an inpainted image is described in detail below with reference to FIGS. 5 and 6. Accordingly, doubling issues, such as those of the target area (122), may be removed from the second anatomical image.

[0040] Although Figure 1 illustrates an example of eliminating the doubling issue by inpainting a true orthophoto image, the present disclosure is not limited thereto. For example, by applying the inpainting method of the present disclosure to a general orthophoto image, the visual quality of the orthophoto image can be improved.

[0041] Inpainting a true-image using this method eliminates the doubling issue that occurs when two-dimensional regions of the same object or shape are displayed overlapping with their corresponding three-dimensional models in occluded areas of the true-image. Consequently, the inpainted true-image can be used to provide an image that reflects the shape of a high-quality three-dimensional model.

[0042] FIG. 2 is a block diagram illustrating the internal configuration of an information processing system (200) according to one embodiment of the present disclosure. The information processing system (200) may include a memory (210), a processor (220), a communication module (230), and an input / output interface (240). The information processing system (200) may be configured to communicate information and / or data with an external system via a network using the communication module (230).

[0043] The memory (210) may include any non-transitory computer-readable recording medium. According to one embodiment, the memory (210) may include a non-volatile permanent mass storage device such as a read-only memory (ROM), a disk drive, a solid state drive (SSD), a flash memory, etc. As another example, a non-volatile permanent mass storage device such as a ROM, an SSD, a flash memory, a disk drive, etc. may be included in the information processing system (200) as a separate permanent storage device distinct from the memory. In addition, the memory (210) may store an operating system and at least one program code (e.g., code for executing inpainted real-time image generation, etc., which is installed and operated in the information processing system (200).

[0044] These software components may be loaded from a computer-readable recording medium separate from the memory (210). This separate computer-readable recording medium may include a recording medium directly connectable to the information processing system (200), for example, a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, a memory card, etc. As another example, the software components may be loaded into the memory (210) through a communication module (230) other than a computer-readable recording medium. For example, at least one program may be loaded into the memory (210) based on a computer program (e.g., a program for executing inpainted real-time image generation, etc.) that is installed by files provided by developers or a file distribution system that distributes installation files of applications through the communication module (230).

[0045] The processor (220) may be configured to process commands of a computer program by performing basic arithmetic, logic, and input / output operations. The commands may be provided to a user terminal (not shown) or another external system via the memory (210) or the communication module (230). For example, the processor (220) may generate a second true image corresponding to an image in which a target area is inpainted based on the first true image and the first mask information using a machine learning model. In addition, the processor (220) of the information processing system (200) may be configured to manage, process, and / or store information and / or data received from a plurality of user terminals and / or a plurality of external systems.

[0046] The communication module (230) may provide a configuration or function for a user terminal (not shown) and the information processing system (200) to communicate with each other via a network, and may provide a configuration or function for the information processing system (200) to communicate with an external system (e.g., a separate cloud system, etc.). For example, control signals, commands, data, etc. provided under the control of the processor (220) of the information processing system (200) may be transmitted to the user terminal and / or the external system via the communication module (230) and the network through the communication module of the user terminal and / or the external system. For example, the information processing system (200) may transmit a request for inpainting of a real-time image to the user terminal via the communication module (230).

[0047] In addition, the input / output interface (240) of the information processing system (200) may be a means for interfacing with a device (not shown) for input or output that is connected to the information processing system (200) or that the information processing system (200) may include. In FIG. 2, the input / output interface (240) is illustrated as an element configured separately from the processor (220), but is not limited thereto, and the input / output interface (240) may be configured to be included in the processor (220). The information processing system (200) may include more components than those illustrated in FIG. 2. However, there is no need to explicitly illustrate most of the conventional technology components.

[0048] FIG. 3 is a diagram illustrating the internal configuration of a processor (220) of an information processing system according to one embodiment of the present disclosure. According to one embodiment, the processor (220) may include a learning unit (310), an inpainting unit (320), and a 3D city model generation unit (330). The internal configuration of the processor (220) of the information processing system illustrated in FIG. 3 is merely an example and may be implemented differently. For example, at least a portion of the configuration of the processor (220) may be omitted, another configuration may be added, and at least a portion of the operations or processes performed by the processor (220) may be performed by a processor of a user terminal that is communicatively connected to the information processing system.

[0049] The learning unit (310) can train a machine learning model to inpaint an image, such as a true-image, or a portion of an image. Specifically, the learning unit (310) can pre-train a machine learning model to input an image and mask information, inpaint an area corresponding to the mask information in the image, and output a visually improved image. In addition, the learning unit (310) can use at least one of an aerial image, a ground image, or a synthetic image as image learning data for the machine learning model. Here, the aerial image, ground image, or synthetic image used as learning data can include a high-resolution image (e.g., 512 x 512, etc.) to improve the detailed quality of inpainting, and can include images that include occluded areas as well as images that do not include occluded areas. Additionally, the learning unit (310) can use mask information associated with any area in the aerial image, ground image, or synthetic image as mask information learning data for the machine learning model.

[0050] In one embodiment, the machine learning model trained to inpaint a portion of an image or video, such as a true-image video, may be, but is not limited to, a convolutional neural network (CNN)-based model including multiple convolution layers, a stable diffusion-based model, a generative adversarial network (GAN)-based model, and the like.

[0051] The inpainting unit (320) can generate a second true image in which the target area is inpainted based on a first true image of the target area and mask information associated with the target area among the target area. Specifically, the inpainting unit (320) can receive the first true image and mask information associated with the target area. In addition, the inpainting unit (320) can generate a second true image corresponding to an image in which the target area is inpainted based on the first true image and the mask information using a machine learning model learned by the learning unit (310). The inpainting unit (320) can store the generated second true image in a database or transmit it to the 3D city model generation unit (330).

[0052] In one embodiment, the inpainting unit (320) can inpaint not only the real-time image but also the texture image of the 3D city model. Specifically, the inpainting unit (320) can generate a texture image in which a specific object is inpainted in the texture image of the 3D city model based on the texture image of the 3D city model and mask information associated with the specific object using the machine learning model learned by the learning unit (310). Here, the texture image may include a 2D image expressing the surface color, texture, etc. of a building included in the 3D city model. In addition, the inpainting unit (320) can generate a texture image in which a specific texture is inpainted in the texture image of the 3D city model based on the texture image of the 3D city model and mask information associated with the specific texture using the machine learning model learned by the learning unit (310). An example of inpainting a texture image is described in detail below with reference to FIGS. 8 and 9.

[0053] The 3D city model generation unit (330) can generate a 3D city model using the second true-image image generated by the inpainting unit (320). Specifically, the 3D city model generation unit (330) can receive a digital elevation model and a 3D building model of the target area. In addition, the 3D city model generation unit (330) can generate a 3D city model based on the second true-image image, the digital elevation model, and the 3D building model. In this way, by generating a 3D city model using the improved second true-image image, the doubling issue in the 3D city model can be eliminated.

[0054] FIG. 4 is a diagram illustrating an example of a method for inpainting a 3D image according to one embodiment of the present disclosure. In one embodiment, a processor (e.g., 220 of FIG. 2 ) may receive a first 3D image (410) representing a target region and mask information (420) associated with a target region among the target region. Here, the target region may include an occluded region caused by one or more three-dimensional objects present in the target region in the first 3D image (410).

[0055] In one embodiment, the processor may input the first true image (410) and mask information (420) into a pre-trained machine learning model (430). Here, the machine learning model (430) may be pre-trained to input the image and mask information and inpaint an area corresponding to the mask information within the image to output a visually improved image. In addition, the machine learning model (430) may be pre-trained using training data including at least one of an aerial image, a ground image, or a synthetic image (e.g., an image synthesized from an aerial image, a ground image, etc.), and mask information associated with an arbitrary area within the aerial image, the ground image, or the synthetic image. Accordingly, the machine learning model (430) may generate a second true image (440) based on the first true image (410) and the mask information (420).

[0056] In one embodiment, the mask information (420) may include first mask information that masks a target region and a region including line characteristics associated with the target region, and second mask information that masks only the target region on the source image. In this case, the machine learning model (430) may generate an intermediate image in which the target region and a region including line characteristics associated with the target region are inpainted, based on the first true image (410) and the first mask information. Additionally, the processor may generate a second true image (440) in which the target region is inpainted, based on the intermediate image generated by the machine learning model (430), the source image of the first true image (410), and the second mask information.

[0057] For example, the first mask information may include mask information of a bridge region corresponding to the target region and a road region connected to the bridge corresponding to an region including line characteristics associated with the target region. In this case, the machine learning model (430) may generate an intermediate image in which the bridge region and the road region connected to the bridge are inpainted in the first true image (410). In addition, the processor may generate a second true image (440) in which only the bridge region is inpainted based on the intermediate image, the source image, and the second mask information in which only the bridge region is masked. In other words, the road region may be supplemented again with the road region of the source image. Accordingly, a true image in which the target region is inpainted more naturally may be generated.

[0058] In one embodiment, the machine learning model can generate a second immersive image with specific objects inpainted on it based on mask information associated with specific objects (or non-building objects, such as cars or street trees) in the first immersive image. For example, the machine learning model can generate the second immersive image by removing and supplementing street trees, etc., present in the first immersive image. Accordingly, a high-quality 3D city model can be generated based on the second immersive image, the digital elevation model, and the 3D building model.

[0059] This configuration allows for inpainting specific objects within an image using a machine learning model. This allows for the processing of 2D images to produce cleaner, higher-quality 3D building models.

[0060] FIG. 5 is a diagram showing examples of a first sedation image and mask information according to an embodiment of the present disclosure, and FIG. 6 is a diagram showing examples of an inpainted second sedation image according to an embodiment of the present disclosure. The first image (510) is an example showing a first sedation image of a target area. In addition, the second image (520) is an example showing mask information (522) associated with a target area (512) among the target areas. Here, the mask information (522) may correspond to the target area (512) in terms of location, size, shape, etc. Additionally, the third image (610) is an example showing a second sedation image visually improved through inpainting.

[0061] In one embodiment, mask information (522) associated with a target area (512) among target regions may be received. Here, the mask information (522) may be received in a form in which the target area (512) is masked. Alternatively, the mask information (522) may be generated by masking the target area (512) in the first image (510) using a separate tool, etc. In addition, the masked area included in the mask information (522) may be corrected to more accurately correspond to the position, size, shape, etc. of the target area (512).

[0062] In one embodiment, the machine learning model can generate a second true image corresponding to an image in which the target area (512) of the first true image is inpainted, based on the first true image and mask information (522). Specifically, the machine learning model can remove an object included in the target area (512) of the first true image. Furthermore, the machine learning model can restore the inpainted area (612) based on information of a surrounding area. For example, a bridge included in the target area (512) of the first true image can be removed based on the mask information (522), and a occluded river under the bridge can be restored, thereby generating a second true image. The inpainted area (612) of the second true image generated in this way can be combined with a three-dimensional model (e.g., a model of an overpass or a bridge) that can be connected to a surrounding area (e.g., a road).

[0063] FIG. 7 is a diagram illustrating an example of a 3D city model generated using a true image according to one embodiment of the present disclosure. The first image (710) illustrates an example of an image of a 3D city model generated using a first true image. In addition, the second image (720) illustrates an example of an image of a 3D city model generated using a second true image in which a target area (712) of the first true image is inpainted.

[0064] In one embodiment, the target area to be inpainted may be determined not only by the bridge but also by three-dimensional objects present in the target area. For example, the target area (712) may include an occluded area created by a building's overpass. By inpainting this target area (712) using a machine learning model, the occluded road area can be restored and displayed in the inpainted area (722).

[0065] FIG. 8 is a diagram illustrating an example of a 3D city model in which a texture image of a building model is inpainted according to one embodiment of the present disclosure. The first image (810) illustrates an example of a building image of a 3D city model generated using a true-image image. In addition, the second image (820) illustrates an example of an image in which a building of the 3D city model is inpainted.

[0066] In one embodiment, a machine learning model may be pre-trained to inpaint a texture image of a building model. Specifically, the machine learning model may be pre-trained to input an image and mask information, and to inpaint an area corresponding to the mask information within the image to output a visually enhanced image. Here, the image may include not only a true-image image but also a texture image of a 3D city model. For example, the machine learning model may inpaint a target area (812) including an image of an open louver of the building based on mask information associated with a louver included in the target area (812) of the 3D city model. Accordingly, the machine learning model may generate a texture image of the 3D city model including an image of a closed louver in the inpainting area (822).

[0067] FIG. 9 is a diagram illustrating an example of a 3D city model in which a texture image of a building model is inpainted according to one embodiment of the present disclosure. The first image (910) illustrates an example of a building image of a 3D city model generated using a true-image image. In addition, the second image (920) illustrates an example of an image in which a building of the 3D city model is inpainted.

[0068] In one embodiment, a machine learning model can be pre-trained to inpaint a texture image of a building model. Specifically, the machine learning model can be pre-trained to input an image and mask information, inpaint an area corresponding to the mask information within the image, and output a visually improved image. Here, the image can include not only a true-image image but also a texture image of a 3D city model. For example, the machine learning model can inpaint a target area (912) including an image of a car reflected on building glass based on the texture image of the 3D city model and mask information associated with the building glass texture included in the target area (912). Accordingly, the machine learning model can generate a texture image of the 3D city model inpainted with the same texture by removing the car reflected on the building glass from the inpainting area (922). Through such a configuration, the 3D city model can be improved.

[0069] FIG. 10 is a flowchart illustrating an example of a method (1000) for inpainting a sedated image according to one embodiment of the present disclosure. In one embodiment, the method (1000) for inpainting a sedated image may be performed by at least one processor. The method (1000) may begin with the processor receiving a first sedated image of a target region and first mask information associated with a target region among the target region (S1010). Here, the target region may include an occluded region caused by one or more three-dimensional objects present in the target region.

[0070] Thereafter, the processor may generate a second true image corresponding to the image in which the target area is inpainted based on the first true image and the first mask information using a machine learning model (S1020). Here, the machine learning model may be pre-trained to input an image and mask information, inpaint an area corresponding to the mask information in the image, and output a visually improved image. In addition, the machine learning model may be pre-trained using learning data including at least one of an aerial image, a ground image, or a synthetic image, and mask information associated with any area in the aerial image, the ground image, or the synthetic image. Additionally, the processor may add a model of one or more three-dimensional objects by overlapping them in an occluded area in the second true image.

[0071] In one embodiment, the processor may receive second mask information associated with a specific object within the true image. In this case, the processor may generate a second true image in which the specific object is inpainted based on the second mask information.

[0072] In one embodiment, the processor may generate a three-dimensional city model based on a second tactile image, a digital elevation model of the target area, and a building model. Specifically, the processor may receive a texture image of the three-dimensional city model and third mask information associated with the texture. Furthermore, the processor may generate an improved texture image based on the texture image and the third mask information using a machine learning model. Based on the improved texture image, the processor may generate a three-dimensional city model.

[0073] In one embodiment, the first mask information may include mask information that masks a target region and an area including line characteristics associated with the target region. In this case, the processor may use a machine learning model to generate an intermediate image in which the target region and an area including line characteristics associated with the target region are inpainted. Furthermore, the processor may generate a second true image based on the intermediate image, the source image of the first true image, and fourth mask information that masks only the target region on the source image.

[0074] The above-described method may be provided as a computer program stored on a computer-readable recording medium for execution on a computer. The medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording means or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program instructions, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.

[0075] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, these techniques may be implemented in hardware, firmware, software, or a combination thereof. Those skilled in the art will appreciate that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software will depend on the particular application and the design requirements imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementations should not be construed as departing from the scope of the present disclosure.

[0076] In a hardware implementation, the processing units used to perform the techniques may be implemented within one or more ASICs, DSPs, GPUs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, a computer, or a combination thereof.

[0077] Accordingly, the various exemplary logical blocks, modules, and circuits described in connection with the present disclosure may be implemented or performed by any combination of a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or those designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0078] In a firmware and / or software implementation, the techniques may be implemented as instructions stored on a computer-readable medium, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, a compact disc (CD), a magnetic or optical data storage device, etc. The instructions may be executable by one or more processors and may cause the processor(s) to perform certain aspects of the functionality described herein.

[0079] When implemented in software, the techniques may be stored on or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media includes both computer storage media and communication media, including any medium that facilitates transfer of a computer program from one place to another. Storage media may be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. In addition, any connection is suitably made to a computer-readable medium.

[0080] For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, digital subscriber line, or wireless technologies such as infrared, radio, and microwave are included within the definition of media. Disk and disc, as used herein, includes compact discs, laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks usually reproduce data magnetically, whereas discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0081] A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. Alternatively, the processor and the storage medium may reside as discrete components in the user terminal.

[0082] While the embodiments described above have been described as utilizing aspects of the presently disclosed subject matter in one or more standalone computer systems, the present disclosure is not limited thereto and may be implemented in conjunction with any computing environment, such as a network or distributed computing environment. Furthermore, aspects of the present disclosure may be implemented in multiple processing chips or devices, and storage may be similarly affected across multiple devices. Such devices may include personal computers, network servers, and portable devices.

[0083] While the present disclosure has been described in connection with certain embodiments herein, various modifications and variations may be made without departing from the scope of the present disclosure, which would be apparent to those skilled in the art. Furthermore, such modifications and variations are intended to fall within the scope of the claims appended to this specification.

Claims

1. A method for inpainting a real-time image, performed by at least one processor, A step of receiving a first true ortho image of a target area and first mask information associated with a target area among the target area; and A step of generating a second true image corresponding to an image in which the target area is inpainted based on the first true image and the first mask information using a machine learning model. An inpainting method comprising:

2. In paragraph 1, An inpainting method, wherein the target area includes an occlusion area caused by one or more three-dimensional objects present in the target area.

3. In paragraph 2, A step of adding a model of one or more three-dimensional objects by overlapping them in the occluded area of the second true image. An inpainting method comprising:

4. In paragraph 1, The above machine learning model is an inpainting method that is pre-trained to input image and mask information, inpaint an area corresponding to the mask information within the image, and output a visually improved image.

5. In paragraph 4, An inpainting method wherein the machine learning model is pre-trained using training data including at least one of an aerial image, a ground image, or a synthetic image, and mask information associated with any area within the aerial image, the ground image, or the synthetic image.

6. In paragraph 1, A step of receiving second mask information associated with a specific object in the above-mentioned true image. Including more, The step of generating the second true image is as follows: A step of generating a second true image in which the specific object is inpainted based on the second mask information. An inpainting method comprising:

7. In paragraph 1, A step of creating a 3D city model based on the second true image, the digital elevation model of the target area, and the building model. An inpainting method comprising:

8. In paragraph 7, The steps of creating the above 3D city model are: A step of receiving a texture image of the above 3D city model and third mask information associated with the texture; A step of generating an improved texture image based on the texture image and the third mask information using the machine learning model; and A step of generating the 3D city model based on the improved texture image. An inpainting method comprising:

9. In paragraph 1, An inpainting method, wherein the first mask information includes mask information that masks an area including the target area and line characteristics associated with the target area.

10. In paragraph 9, The step of generating the second true image is as follows: A step of generating an intermediate image in which the target area and an area including line characteristics associated with the target area are inpainted using the machine learning model. An inpainting method comprising:

11. In paragraph 10, The step of generating the second true image is as follows: A step of generating the second true image based on the intermediate image, the source image of the first true image, and fourth mask information that masks only the target area on the source image. An inpainting method comprising:

12. As a system, Communication module; memory; and At least one processor connected to said memory and configured to execute at least one computer-readable program contained in said memory, At least one program above, Receive a first true ortho image of the target area and first mask information associated with the target area among the target areas, A system comprising instructions for generating a second true image corresponding to an image in which the target area is inpainted, based on the first true image and the first mask information, using a machine learning model.

Citation Information

Patent Citations

  • True orthophoto generation method

    CN106875364A

  • Method for fusing digital surface model data based on true orthophoto

    CN113362439A

  • Smart city operation management platform based on one-network unified management of CIM model

    CN114331020A

  • Supporting system for monitoring the automated external defibrillator

    KR1020230037091A

  • Square duct

    KR1020240044640A