Method and system for generating synthetic image
By employing different image generation models for virtual and real domain styles, the method addresses the gap in autonomous driving simulator training data, producing high-quality, realistic images that improve learning effectiveness and reduce costs.
Patent Information
- Application Number
- PCT/KR2025/005689
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-26
- Filing Date
- 2025-04-28
- Publication Date
- 2025-10-30
Smart Images

Figure KR2025005689_30102025_PF_FP_ABST
Abstract
Description
Method and system for creating a virtual image
[0001] The present disclosure relates to a method and system for generating a virtual image used in autonomous driving simulation.
[0002] Recently, with the advancement of automotive technologies such as IT, electricity, and electronics, autonomous driving technology, which utilizes all of these technologies, is attracting attention. Autonomous driving technology controls a vehicle without driver intervention. It monitors the driving environment through various sensors installed in the vehicle and makes driving decisions.
[0003] Meanwhile, autonomous driving simulators are trained using virtual images that resemble the actual driving environment of a vehicle as training data. However, there is a gap between the virtual (synthetic) images used as training data for autonomous driving simulators and the real (real) images. Failure to properly address this gap can lead to reduced learning effectiveness. This problem degrades the performance of autonomous driving simulators and limits the usability of autonomous driving technology in application fields.
[0004] The present disclosure provides a virtual image generation method and device (system) to solve the above problems.
[0005] The present disclosure can be implemented in various ways, including as a method, a device (system), or a computer program stored on a readable storage medium.
[0006] According to one embodiment of the present disclosure, a virtual image generation method, performed by at least one processor, comprises the steps of: obtaining a target image of a first domain style; generating a first image of the first domain style representing a first type of object in the target image; generating a second image of the first domain style representing a second type of object in the target image; generating a first partial virtual image of the second domain style based on the first image using a first image generation model; generating a second partial virtual image of the second domain style based on the second image using a second image generation model; and generating a virtual image of the second domain style based on the first partial virtual image and the second partial virtual image, wherein the first domain style and the second domain style are different.
[0007] According to one embodiment of the present disclosure, the first domain style is a virtual domain style and the second domain style is a real domain style.
[0008] According to one embodiment of the present disclosure, the first type of object is an object that is defined by being distinguished by instance object, and the second type of object is an object that is defined by being distinguished by class according to the properties of the object.
[0009] According to one embodiment of the present disclosure, the first image may include RGB information for an object of a first type, and the second image may include segmentation information for an object of a second type.
[0010] According to one embodiment of the present disclosure, the first image generation model may be a model trained to generate an output image of a second domain style based on an input image of a first domain style.
[0011] According to one embodiment of the present disclosure, the second image generation model may be a model trained to generate an output image of a second domain style based on segmentation information.
[0012] According to one embodiment of the present disclosure, the step of generating a virtual image of a second domain style may include the step of generating a combined image by combining a first partial virtual image and a second partial virtual image, the step of extracting at least a portion of an area where the first partial virtual image and the second partial virtual image are adjacent within the combined image, and the step of converting first color information for at least a portion of the area.
[0013] According to one embodiment of the present disclosure, the step of converting the first color information may include the step of extracting second color information corresponding to at least a portion of the region from the target image, and the step of converting the first color information for at least a portion of the region within the combined image into the second color information.
[0014] According to one embodiment of the present disclosure, the first type of object includes a dynamic object and a first static object, the second type of object includes a second static object, and the first static object may be an object associated with traffic information.
[0015] A computer-readable recording medium having recorded thereon instructions for executing a method according to one embodiment of the present disclosure on a computer may be provided.
[0016] An information processing system according to one embodiment of the present disclosure includes a communication module, a memory, and at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory, wherein the at least one program includes instructions for obtaining a target image of a first domain style, generating a first image of the first domain style representing a first type of object in the target image, generating a second image of the first domain style representing a second type of object in the target image, generating a first partial virtual image of the second domain style based on the first image using a first image generation model, generating a second partial virtual image of the second domain style based on the second image using the second image generation model, and generating a virtual image of the second domain style based on the first partial virtual image and the second partial virtual image, wherein the first domain style is different from the second domain style.
[0017] According to some embodiments of the present disclosure, a high-quality image with higher realism can be generated by generating a virtual image using different image generation models depending on the type of object for one image.
[0018] According to some embodiments of the present disclosure, it is possible to reduce the time and cost required to implement a first domain style target image similar to actual reality using computer graphics or the like.
[0019] According to some embodiments of the present disclosure, a second partial virtual image that directly reflects the style of objects existing in real reality can be generated using a second image generation model.
[0020] According to some embodiments of the present disclosure, the color of the boundary between the first partial virtual image and the second partial virtual image within the combined image is corrected to naturally connect, thereby generating a more natural, high-quality second domain style virtual image.
[0021] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs (referred to as “one skilled in the art”) from the description of the claims.
[0022] Embodiments of the present disclosure will be described below with reference to the accompanying drawings, wherein like reference numerals represent similar elements, but are not limited thereto.
[0023] FIG. 1 is a diagram illustrating an example of generating a virtual image of a second domain style from a target image of a first domain style according to one embodiment of the present disclosure.
[0024] FIG. 2 is a schematic diagram showing a configuration in which an information processing system is connected to enable communication with a plurality of user terminals to generate a virtual image according to one embodiment of the present disclosure.
[0025] FIG. 3 is a block diagram showing the internal configuration of a user terminal and an information processing system according to one embodiment of the present disclosure.
[0026] FIG. 4 is a diagram illustrating an example of generating a first partial virtual image based on a first image using a first image generation model according to one embodiment of the present disclosure.
[0027] FIG. 5 is a diagram illustrating an example of generating a second partial virtual image based on a second image using a second image generation model according to one embodiment of the present disclosure.
[0028] FIG. 6 is a diagram illustrating an example of generating a second domain style virtual image based on a first partial virtual image and a second partial virtual image according to one embodiment of the present disclosure.
[0029] FIG. 7 is a diagram illustrating an example of generating a virtual image of a second domain style based on a combined image according to one embodiment of the present disclosure.
[0030] FIG. 8 is a drawing showing an example of an image generated according to a virtual image generation method according to one embodiment of the present disclosure.
[0031] FIG. 9 is a flowchart illustrating a virtual image generation method according to one embodiment of the present disclosure.
[0032] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the attached drawings. However, in the following description, specific descriptions of widely known functions or configurations will be omitted if they may unnecessarily obscure the gist of the present disclosure.
[0033] In the attached drawings, identical or corresponding components are assigned the same reference numerals. Furthermore, in the description of the embodiments below, duplicate descriptions of identical or corresponding components may be omitted. However, even if a description of a component is omitted, it is not intended that such component is not included in any embodiment.
[0034] The advantages and features of the disclosed embodiments, and methods for achieving them, will become clearer with reference to the embodiments described below, along with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure the completeness of the disclosure and to fully inform those skilled in the art of the scope of the invention.
[0035] The terms used in this specification will be briefly explained, followed by a detailed description of the disclosed embodiments. The terms used in this specification have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of engineers working in the relevant field, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.
[0036] In this specification, singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, plural expressions include singular expressions unless the context clearly indicates otherwise. When a part of the specification is said to include a component, this does not exclude other components, but rather implies that other components may be included, unless otherwise specifically stated.
[0037] Also, the term 'module' or 'part' used in the specification means a software or hardware component, and the 'module' or 'part' performs certain roles. However, the 'module' or 'part' is not limited to software or hardware. The 'module' or 'part' may be configured to reside on an addressable storage medium and may be configured to execute one or more processors. Thus, as an example, the 'module' or 'part' may include at least one of components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, or variables. The functionality provided within the components and 'modules' or 'parts' may be combined into a smaller number of components and 'modules' or 'parts', or further separated into additional components and 'modules' or 'parts'.
[0038] According to one embodiment of the present disclosure, a 'module' or 'unit' may be implemented as a processor and a memory. 'Processor' should be broadly construed to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, and the like. In some circumstances, a 'processor' may also refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), and the like. A 'processor' may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors in conjunction with a DSP core, or any other such combination of configurations. In addition, 'memory' should be broadly construed to include any electronic component capable of storing electronic information. 'Memory' may refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage, registers, etc. Memory is said to be in electronic communication with the processor if the processor can read information from, and / or write information to, the memory. Memory integrated in a processor is in electronic communication with the processor.
[0039] In the present disclosure, the "system" may include, but is not limited to, at least one of a server device and a cloud device. For example, the system may be comprised of one or more server devices. As another example, the system may be comprised of one or more cloud devices. As yet another example, the system may be configured and operated by a combination of a server device and a cloud device.
[0040] In the present disclosure, 'each of the plurality of As' or 'each of the plurality of As' may refer to each of all components included in the plurality of As, or may refer to each of some components included in the plurality of As.
[0041] In the present disclosure, the term "domain style" refers to the visual characteristics and / or artistic style of an image, and may represent a unique combination of the field of view (FOV) of the camera that captured the image, camera parameters, color, texture, pattern, shape, and other visual elements that define the overall appearance and aesthetic quality of the image. For example, the domain style of an image may include a virtual domain style such as computer graphics (e.g., computer game graphics), or a real-world domain style such as a real-world image captured by a specific camera. Furthermore, when different cameras capture the real world, images captured by each camera may have different domain styles depending on the various characteristics of the cameras.
[0042] FIG. 1 is a diagram illustrating an example of generating a virtual image (180) of a second domain style from a target image (110) of a first domain style according to one embodiment of the present disclosure. As illustrated, a processor (e.g., at least one processor of an information processing system that generates a virtual image) may obtain a target image (110) of the first domain style. Here, the first domain style may be a virtual domain style generated through a computer simulation or a computer game, but is not limited thereto. For example, the first domain style may include various types of domain styles (e.g., a cartoon image style, a pointillism image style, etc.).
[0043] In one embodiment, the processor may identify / extract a first type of object within a target image (110) of a first domain style. Here, the first type of object may refer to an object that is defined by being distinguished by instance object. As a specific example, the first type of object may include an object that requires a clear distinction and definition of the boundary for each object even among objects of the same class. For example, the first type of object may include dynamic objects such as vehicles, pedestrians, and bicycles. Additionally, the first type of object may include an object related to vehicle driving, which includes fine-grained information that does not allow even minor content information damage associated with the object. For example, the first type of object may include static objects related to traffic information such as traffic signs, traffic lights, and lanes.
[0044] In one embodiment, the processor may generate a first image (120) of a first domain style representing an object of a first type. Here, the first image (120) may include RGB information for an object of the first type identified / extracted from a target image (110) of the first domain style.
[0045] In one embodiment, the processor may identify / extract a second type of object within a target image (110) of a first domain style. Here, the second type of object is an object defined by being classified into a class according to the object's properties, and the second type of object may refer to an object that does not require distinction by instance object. For example, the second type of object may include static objects with little relevance to vehicle driving, such as buildings and trees. In addition, the second type of object may include objects with ambiguous boundaries or objects with little need to clearly define their shapes, such as the sky and clouds. Here, the type of class is variable and may vary depending on the application to which it is applied. In addition, all objects except the first type of object may be classified and defined as the second type of object.
[0046] In one embodiment, the processor may generate a second image (130) of a first domain style representing an object of a second type. Here, the second image (130) may include segmentation information for the object of the second type identified / extracted / generated from a target image (110) of the first domain style.
[0047] In one embodiment, the processor may use the first image generation model (140) to generate a first partial virtual image (160) of a second domain style based on the first image (120). Here, the first domain style and the second domain style may be different from each other. For example, the first domain style may be a virtual domain style and the second domain style may be a real-world domain style, such as a real-world image captured by a specific camera, but is not limited thereto. For example, the first domain style and the second domain style may be two different domain styles among various types of domain styles (e.g., a cartoon image style, a pointillism image style, a hand-drawn image style, etc.).
[0048] In one embodiment, the first image generation model (140) may be a model (e.g., a neural network model) trained to receive an image of a first domain style as input and generate an image of a second domain style as output. For example, the first image generation model (140) may be a model trained based on a pair of a first training image of the first domain style and a second training image of the second domain style. Accordingly, the first image generation model (140) may generate a first partial virtual image (160) of the second domain style based on the first image (120) of the first domain style. An example of generating a first partial virtual image (160) based on the first image (120) by the first image generation model (140) is described in detail below with reference to FIG. 4.
[0049] In one embodiment, the processor may generate a second partial virtual image (170) of a second domain style based on the second image (130) using a second image generation model (150). Here, the second image generation model (150) may be a model trained to generate an image of the second domain style based on segmentation information. For example, the second image generation model (150) may be trained based on a pair of a third learning image of the second domain style and segmentation information generated from the third learning image of the second domain style. Accordingly, the second image generation model (150) may generate a second partial virtual image (170) in which objects of the second type of the second domain style are generated within a corresponding segmentation area based on segmentation information for objects of the second type within the second image (130). An example of generating a second partial virtual image (170) based on a second image (130) by a second image generation model (150) is described in detail below based on FIG. 5.
[0050] In one embodiment, the processor may generate a second domain style virtual image (180) based on the first partial virtual image (160) and the second partial virtual image (170). For example, the processor may generate a combined image by combining the first partial virtual image (160) for a first type of object and the second partial virtual image (170) for a second type of object. In addition, the processor may perform a post-processing operation on an area where the first partial virtual image (160) and the second partial virtual image (180) are adjacent in the combined image, thereby generating the second domain style virtual image (180). An example of generating the second domain style virtual image (180) based on the first partial virtual image (160) and the second partial virtual image (170) is described in detail below with reference to FIGS. 6 and 7.
[0051] By this configuration, the processor can generate a high-quality image with higher realism by generating a virtual image using different image generation models depending on the type of object for one image.
[0052] FIG. 2 is a schematic diagram illustrating a configuration in which an information processing system (230) is connected to a plurality of user terminals (210_1, 210_2, 210_3) so as to be able to communicate with each other to generate a virtual image according to one embodiment of the present disclosure. As illustrated, the plurality of user terminals (210_1, 210_2, 210_3) may be connected to the information processing system (230) capable of generating a virtual image via a network (220). Here, the plurality of user terminals (210_1, 210_2, 210_3) may include terminals of users who receive the generated virtual images.
[0053] In one embodiment, the information processing system (230) may include one or more server devices and / or databases capable of storing, providing, and executing computer executable programs (e.g., downloadable applications) and data associated with virtual image creation, or one or more distributed computing devices and / or distributed databases based on cloud computing services.
[0054] The virtual image provided by the information processing system (230) may be provided to the user through an image generation application web browser or web browser extension program installed on each of a plurality of user terminals (210_1, 210_2, 210_3). For example, the information processing system (230) may provide information corresponding to a virtual image generation request received from a user terminal (210_1, 210_2, 210_3) or perform corresponding processing through an image generation application, etc.
[0055] A plurality of user terminals (210_1, 210_2, 210_3) can communicate with an information processing system (230) via a network (220). The network (220) can be configured to enable communication between the plurality of user terminals (210_1, 210_2, 210_3) and the information processing system (230). Depending on the installation environment, the network (220) can be configured as a wired network such as Ethernet, a wired home network (Power Line Communication), a telephone line communication device, and RS-serial communication, a wireless network such as a mobile communication network, WLAN (Wireless LAN), Wi-Fi, Bluetooth, and ZigBee, or a combination thereof. The communication method is not limited, and may include not only a communication method utilizing a communication network (e.g., a mobile communication network, wired Internet, wireless Internet, broadcasting network, satellite network, etc.) that the network (220) may include, but also short-range wireless communication between user terminals (210_1, 210_2, 210_3).
[0056] In FIG. 2, a mobile phone terminal (210_1), a tablet terminal (210_2), and a PC terminal (210_3) are illustrated as examples of user terminals, but are not limited thereto, and the user terminals (210_1, 210_2, 210_3) may be any computing device capable of wired and / or wireless communication and capable of installing and executing a virtual image creation service application or web browser, or a virtual image creation service application or web browser, etc. For example, the user terminal may include an AI speaker, a smartphone, a mobile phone, a navigation device, a computer, a laptop, a digital broadcasting terminal, a PDA (Personal Digital Assistants), a PMP (Portable Multimedia Player), a tablet PC, a game console, a wearable device, an IoT (internet of things) device, a VR (virtual reality) device, an AR (augmented reality) device, a set-top box, etc. In addition, although FIG. 2 illustrates three user terminals (210_1, 210_2, 210_3) communicating with the information processing system (230) via the network (220), this is not limited thereto, and a different number of user terminals may be configured to communicate with the information processing system (230) via the network (220).
[0057] Although FIG. 2 illustrates an exemplary configuration in which user terminals (210_1, 210_2, 210_3) communicate with an information processing system (230) to receive generated virtual images, the present invention is not limited thereto. For example, user terminals (210_1, 210_2, 210_3) may directly generate virtual images without communicating with the information processing system (230).
[0058] FIG. 3 is a block diagram showing the internal configuration of a user terminal (210) and an information processing system (230) according to one embodiment of the present disclosure. The user terminal (210) may refer to any computing device capable of executing applications, web browsers, etc., and capable of wired / wireless communication, and may include, for example, a mobile phone terminal (210_1), a tablet terminal (210_2), a PC terminal (210_3), etc. of FIG. 2. As illustrated, the user terminal (210) may include a memory (312), a processor (314), a communication module (316), and an input / output interface (318). Similarly, the information processing system (230) may include a memory (332), a processor (334), a communication module (336), and an input / output interface (338). As illustrated in FIG. 3, the user terminal (210) and the information processing system (230) may be configured to communicate information and / or data via a network (220) using respective communication modules (316, 336). In addition, the input / output device (320) may be configured to input information and / or data to the user terminal (210) or output information and / or data generated from the user terminal (210) via the input / output interface (318).
[0059] The memory (312, 332) may include any non-transitory computer-readable recording medium. According to one embodiment, the memory (312, 332) may include a permanent mass storage device such as a read-only memory (ROM), a disk drive, a solid-state drive (SSD), or flash memory. As another example, a permanent mass storage device such as a ROM, an SSD, a flash memory, or a disk drive may be included in the user terminal (210) or the information processing system (230) as a separate permanent storage device distinct from the memory. In addition, an operating system and at least one program code may be stored in the memory (312, 332).
[0060] These software components may be loaded from a computer-readable recording medium separate from the memory (312, 332). This separate computer-readable recording medium may include a recording medium directly connectable to the user terminal (210) and the information processing system (230), and may include, for example, a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, a memory card, etc. As another example, the software components may be loaded into the memory (312, 332) through a communication module (316, 336) other than a computer-readable recording medium. For example, at least one program may be loaded into the memory (312, 332) based on a computer program that is installed by files provided by developers or a file distribution system that distributes installation files of applications through a network (220).
[0061] The processor (314, 334) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (314, 334) by a memory (312, 332) or a communication module (316, 336). For example, the processor (314, 334) may be configured to execute instructions received according to program code stored in a storage device such as the memory (312, 332).
[0062] The communication module (316, 336) may provide a configuration or function for the user terminal (210) and the information processing system (230) to communicate with each other via the network (220), and may provide a configuration or function for the user terminal (210) and / or the information processing system (230) to communicate with another user terminal or another system (e.g., a separate cloud system, etc.). For example, a request or data (e.g., an image generation model learning request, a virtual image generation request, etc.) generated by the processor (314) of the user terminal (210) according to a program code stored in a recording device such as a memory (312) may be transmitted to the information processing system (230) via the network (220) under the control of the communication module (316). Conversely, a control signal or command provided under the control of the processor (334) of the information processing system (230) can be received by the user terminal (210) through the communication module (316) of the user terminal (210) via the communication module (336) and the network (220).
[0063] The input / output interface (318) may be a means for interfacing with an input / output device (320). As an example, the input device may include a device such as a camera, keyboard, microphone, mouse, etc., including an audio sensor and / or an image sensor, and the output device may include a device such as a display, a speaker, a haptic feedback device, etc. As another example, the input / output interface (318) may be a means for interfacing with a device that has a configuration or function integrated into one for performing input and output, such as a touch screen. For example, when the processor (314) of the user terminal (210) processes a command of a computer program loaded into the memory (312), a service screen configured using information and / or data provided by the information processing system (230) or another user terminal may be displayed on the display through the input / output interface (318). In FIG. 3, the input / output device (320) is illustrated as not being included in the user terminal (210), but is not limited thereto, and may be configured as a single device with the user terminal (210). In addition, the input / output interface (338) of the information processing system (230) may be a means for interfacing with a device (not shown) for input or output that is connected to the information processing system (230) or that the information processing system (230) may include. In FIG. 3, the input / output interfaces (318, 338) are illustrated as elements configured separately from the processors (314, 334), but are not limited thereto, and the input / output interfaces (318, 338) may be configured to be included in the processors (314, 334).
[0064] The user terminal (210) and information processing system (230) may include more components than those illustrated in FIG. 3. However, it is not necessary to explicitly illustrate most conventional components. In one embodiment, the user terminal (210) may be implemented to include at least some of the input / output devices (320) described above. Furthermore, the user terminal (210) may further include other components, such as a transceiver, a Global Positioning System (GPS) module, a camera, various sensors, a database, and the like.
[0065] While a program for learning an artificial neural network model, image generation application, etc. is running, the processor (314) can receive text, images, videos, voices, and / or actions, etc. input or selected through an input device such as a camera, microphone, including a touch screen, keyboard, audio sensor, and / or image sensor connected to an input / output interface (318), and can store the received text, images, videos, voices, and / or actions, etc. in a memory (312) or provide them to an information processing system (230) through a communication module (316) and a network (220).
[0066] The processor (314) of the user terminal (210) may be configured to manage, process, and / or store information and / or data received from an input / output device (320), another user terminal, an information processing system (230), and / or multiple external systems. The information and / or data processed by the processor (314) may be provided to the information processing system (230) via a communication module (316) and a network (220). The processor (314) of the user terminal (210) may transmit the information and / or data to the input / output device (320) via an input / output interface (318) and output the information and / or data. For example, the processor (314) may output or display the received information and / or data on the screen of the user terminal (210).
[0067] The processor (334) of the information processing system (230) may be configured to manage, process, and / or store information and / or data received from multiple user terminals (210) and / or multiple external systems. Information and / or data processed by the processor (334) may be provided to the user terminal (210) via a communication module (336) and a network (220).
[0068] FIG. 4 is a diagram illustrating an example of generating a first partial virtual image (430) based on a first image (410) using a first image generation model (420) according to one embodiment of the present disclosure. In one embodiment, the first image generation model (420) may receive a first image (410) of a first domain style. Here, the first domain style may be, but is not limited to, a virtual domain style. In addition, the first image (410) may be an image representing a first type of object within a target image of the first domain style. The first type of object is an object defined separately for each instance object, and may include dynamic objects (e.g., vehicles, pedestrians, bicycles, etc.) and static objects related to traffic information (e.g., traffic signs, traffic lights, lanes, etc.).
[0069] In one embodiment, the first image (410) may include RGB information for a first type of object. For example, the processor may identify / extract a first type of object within a target image of a first domain style, and then generate the first image (410) based on RGB information for pixels in an area constituting the identified first type of object.
[0070] In one embodiment, the first image generation model (420) may be a model trained to generate an output image of a second domain style based on an input image of a first domain style. For example, the first image generation model (420) may be a model trained based on a pair of a first training image of the first domain style and a second training image of the second domain style. In addition, the first training image and the second training image may include RGB information about objects in the image. Here, the domain styles of the first training image of the first domain style and the second training image of the second domain style are different from each other, but the appearances of the objects in the image may be the same or similar on a pixel-wise basis.
[0071] Accordingly, the first image generation model (420) can generate an output image in which the mood, color, light intensity, etc. of the input image are changed while maintaining the appearance of the objects in the input image. That is, the first image generation model (420) can generate a first partial virtual image (430) in which the mood, color, light intensity, etc. of the image are changed while maintaining the appearance of the first type of object in the first image (410) the same, based on the first image (410). In other words, although the domain styles of the first image (410) of the first domain style and the first partial virtual image (430) of the second domain style are different from each other, the appearances of the objects in the images may be the same or similar on a pixel-wise basis.
[0072] The image generated by the first image generation model (420) learned as described above has the advantage that the shapes of objects in the image are not distorted or the contents within the objects are not changed. However, if the appearance, texture, etc. of the objects are completely maintained, the effect of changing the domain style may be reduced even if the domain styles of the input image and the output image are different. For example, when generating an output image in a real domain style from an input image in a virtual domain style, the appearance, texture, etc. of objects generated through computer simulation or computer game may be implemented as is in the output image, which may reduce the realism of the output image.
[0073] Accordingly, in one embodiment, the processor may identify only the first type of object within the target image of the first domain style, and generate a first image (410) representing the first type of object. Thereafter, the first image generation model (420) is used to generate a first partial virtual image (430) of the second domain style based on the first image (410). For example, among the first type of objects included within the first image (410), an object related to vehicle driving (e.g., a traffic sign, a lane) may not be distorted in the appearance and content (e.g., the content of a traffic sign or the direction of a lane, etc.) of the object even after the first partial virtual image is generated by the first image generation model (420).
[0074] On the other hand, the second image representing a second type of object (e.g., a building, the sky, a tree, etc.) that is relatively less important for vehicle driving is used to generate a second partial virtual image using a second image generation model rather than the first image generation model (420). An example of generating a second partial virtual image based on the second image using the second image generation model is described in detail in FIG. 5.
[0075] In Fig. 4, for the sake of convenience of explanation, the first image (410) and the first partial virtual image (430) are illustrated as including not only the first type of object but also the second type of object. However, it can be understood that the first image generation model (420) selectively identifies / extracts RGB information for only the first type of object among the objects in the first image (410) to generate the first partial virtual image (430).
[0076] FIG. 5 is a diagram illustrating an example of generating a second partial virtual image (530) based on a second image (510) using a second image generation model (520) according to one embodiment of the present disclosure. In one embodiment, the second image generation model (520) may receive the second image (510). For example, the second image (510) may include segmentation information for a second type of object.
[0077] In one embodiment, the processor may identify / extract a second type of object within a target image of a first domain style, and perform semantic segmentation on the identified second type of object to generate a second image (510). Here, the first domain style may be, but is not limited to, a virtual domain style. In addition, the second image (510) may be an image representing a second type of object within the target image of the first domain style. The second type of object is an object defined by being classified into a class according to the object's properties, and may include static objects (e.g., buildings, the sky, trees, etc.) that have little relevance to traffic information. All objects except for the first type of object may be classified and defined as the second type of object.
[0078] In one embodiment, the second image generation model (520) may be a model trained to generate an output image (e.g., an RGB image) of a second domain style based on segmentation information. For example, the second image generation model (520) may be trained based on a pair of third training images of the second domain style and segmentation information for objects within the third training images. Here, the second domain style may be a real-world domain style, such as a real-world image captured by a specific camera.
[0079] In one embodiment, the second image generation model (520) can generate a second partial virtual image (530) of a second domain style based on segmentation information associated with a second type of object in the second image (510). For example, the second image generation model (520) can generate a second partial virtual image (530) of a second domain style in which an object of the same class as the second type of object is generated in a segmentation area corresponding to the second type of object.
[0080] The second image generation model (520) learned as described above does not generate an output image in which the appearance and / or content of objects in the input image are identically implemented at the pixel level, but can be learned so that the style of objects likely to exist in reality is directly reflected in the objects in the output image. For example, the second image generation model (520) can generate a second partial virtual image (530) in which an image in which the style of objects likely to exist in reality is directly reflected as an object of the same class as the second type of object is generated in the segmentation area of the second type of objects included in the second image (510). As a specific example, the image of an object existing in reality may include images of objects (e.g., terrain of each country, buildings, etc.) that are difficult to implement through computer simulations or computer games.
[0081] By this configuration, the time and cost required to implement a first domain style target image similar to actual reality using computer graphics, etc. can be reduced, and a more realistic second partial virtual image that directly reflects the style of objects existing in actual reality can be generated using a second image generation model (520).
[0082] In FIG. 5, for convenience of explanation, it is illustrated that not only objects of the second type but also objects of the first type are included in the second image (510) and the second partial virtual image (530). However, it can be understood that the second image generation model (520) selectively identifies / extracts segmentation information for only objects of the second type among the objects in the second image (510) to generate the second partial virtual image (530).
[0083] FIG. 6 is a diagram illustrating an example of generating a virtual image (660) of a second domain style based on a first partial virtual image (610, 620) and a second partial virtual image (630, 640) according to one embodiment of the present disclosure. In one embodiment, the processor may receive a first partial virtual image (610, 620) of a second domain style generated by a first image generation model. The first partial virtual image (610, 620) may be an image generated based on a first image representing a first type of object within a target image of the first domain style. Accordingly, the first partial virtual image (610, 620) may be an image of a second domain style in which an object of the first type is generated. Referring to FIG. 6 , it can be confirmed that an image for an object of the first type (e.g., a vehicle, a lane, etc.) is generated within the first partial virtual image (620).
[0084] In one embodiment, the processor may receive a second partial virtual image (630, 640) of a second domain style generated by a second image generation model. The second partial virtual image (630, 640) may be an image generated based on a second image representing a second type of object within a target image of the first domain style. Accordingly, the second partial virtual image (630, 640) may be an image of the second domain style in which an object of the second type is generated. Referring to FIG. 6, it can be confirmed that an image of an object of the second type (e.g., a building, the sky, a tree, etc.) is generated within the second partial virtual image (640).
[0085] In one embodiment, the processor can generate a combined image by combining a first partial virtual image (620) and a second partial virtual image (640). Since the first partial virtual image (620) generates a first type of object, and the second partial virtual image (640) generates a second type of object (all objects except for objects of the first type), the first partial virtual image (620) and the second partial virtual image (640) can be combined so that they are perfectly adjacent to each other without any empty or overlapping areas in the combined image. However, in this case, some areas where the first partial virtual image (620) and the second partial virtual image (640) are adjacent may be unnatural. Accordingly, the processor can perform a post-processing operation (650) on the combined image to generate a second domain style virtual image (660). A detailed description thereof will be described in detail later with reference to FIG. 7.
[0086] FIG. 7 is a diagram illustrating an example of generating a second domain style virtual image (750) based on a combined image (710) according to one embodiment of the present disclosure. In one embodiment, the processor may generate the combined image (710) by combining a first partial virtual image and a second partial virtual image. The processor may extract at least a portion of the area where the first partial virtual image and the second partial virtual image are adjacent within the combined image (710), but is not limited thereto.
[0087] In one embodiment, the processor may produce first color information (720) representing color information for the combined image (710). Additionally, the processor may produce second color information (740) representing color information for the target image (730) of the first domain style.
[0088] In one embodiment, the first color information (720) and the second color information (740) may be produced through Fourier transform. For example, the processor may perform a Fourier transform on the combined image (710) to obtain an amplitude map and a phase map. Here, the amplitude map is information related to light, color, etc. of the combined image (710) and may correspond to the first color information (720). For example, the amplitude map may be expressed as a two-dimensional coordinate system in which the horizontal and vertical axes each represent a horizontal frequency and a vertical frequency of the combined image (710), and each coordinate value on the coordinate system may represent the amplitude of a frequency component corresponding to the corresponding coordinate (e.g., the brightness of a pixel in the combined image (710). In addition, the phase map may be edge information for objects in the combined image (710). For example, a phase map can be expressed as a two-dimensional coordinate system in which the horizontal and vertical axes each represent the horizontal frequency and vertical frequency of the combined image (710), and each coordinate value on the coordinate system can represent the phase of the frequency component corresponding to the coordinate (e.g., edge information for objects, spatial arrangement information, etc.).
[0089] Similarly, the processor may perform a Fourier transform on the target image (730) of the first domain style to obtain an amplitude map and a phase map. Here, the amplitude map is information related to light, color, etc. of the target image (730) of the first domain style, and may correspond to the second color information (740). In addition, the phase map may be edge information for objects within the target image (730) of the first domain style.
[0090] In one embodiment, the processor may convert the first color information (720) for the combined image (710) into the second color information (740) for the target image (730) of the first domain style. For example, the processor may convert the first color information (720) for the first area of the amplitude map (hereinafter referred to as the 'first amplitude map') for the combined image (710) into the second color information (740) for the second area of the amplitude map (hereinafter referred to as the 'second amplitude map') for the target image (730) of the first domain style. Here, the first area is an area close to the origin on the coordinate system of the first amplitude map, and may be determined as an area of a low-frequency band. That is, the first area may represent an area in which fluctuations in light, color, etc. according to the location of pixels on the combined image (710) are relatively small. In addition, the second region may be determined as a region close to the origin on the coordinate system of the second amplitude map, and may be determined as a region of a low-frequency band. That is, the second region may represent a region in which there is relatively little variation in light, color, etc. according to the location of pixels on the target image (730) of the first domain style. The second region of the second amplitude map may be an region corresponding to the first region of the first amplitude map. The shape, size, location, etc. of the first region and / or the second region may be determined differently depending on the resolution of the image, the target color conversion intensity, etc.
[0091] In one embodiment, the processor may inject second color information (740) for a second region of the second amplitude map into a first region of the first amplitude map. Thereafter, the processor may generate a virtual image (750) of a second domain style by performing an inverse Fourier transform on the amplitude map and the phase map of the combined image (710) into which the second color information (740) has been injected. For example, the processor may generate a virtual image (750) of a second domain style by transforming only the first color information (720) associated with the first region of the first amplitude map while maintaining color information associated with a region excluding the first region of the first amplitude map (e.g., a region of a high-frequency band) and phase information associated with the phase map of the combined image (710). Accordingly, the shapes of objects in the virtual image (750) of the second domain style are maintained to be identical / similar to the shapes of objects in the combined image (710), and the overall color of the virtual image (750) of the second domain style can be corrected to be similar to the overall color of the target image (710).
[0092] In another embodiment, the processor may produce first color information (720) for a first region, which is at least a portion of an area extracted within the combined image (710). Furthermore, the processor may produce second color information (740) for a second region, which is at least a portion of an area extracted within the target image (730) of the first domain style. The second region extracted within the target image (730) of the first domain style may be an area corresponding to the first region extracted within the combined image (710).
[0093] In one embodiment, the processor can convert first color information (720) for a first area within the combined image (710) into second color information (740) for a second area within the target image (730) of the first domain style. For example, the processor can inject second color information (740) for a second area within the target image (730) of the first domain style into the first area within the combined image (710) based on an amplitude map and a phase map obtained from each of the combined image (710) and the target image (730) of the first domain style. By this configuration, the color of the boundary between the first partial virtual image and the second partial virtual image within the combined image (710) is corrected to be naturally connected, so that a more natural and high-quality virtual image (750) of the second domain style can be generated.
[0094] FIG. 8 is a diagram illustrating examples of images generated according to a virtual image generation method according to one embodiment of the present disclosure. The first image is an example of a target image (810) of a first domain style. The first domain style may be a virtual domain style generated through a computer simulation or a computer game, etc. In other words, the first image may be a target image (810) virtually generated to generate a virtual image according to the method of the present disclosure.
[0095] The second image is an example of a first partial virtual image (820) of a second domain style generated using the first image generation model. The second domain style may be a real-world domain style, such as one captured by a specific camera. Referring to the second image, although the domain style of the first partial virtual image (820) is different from that of the target image (810), it can be confirmed that the appearances of objects in the image are identical or similar on a pixel-wise basis. Specifically, it can be confirmed that the appearances of objects in the target image (810) in the first partial virtual image (820) are maintained completely identically (or similarly), while the mood, color, and intensity of light of the image are changed.
[0096] The third image is an example of a second partial virtual image (830) of a second domain style generated using a second image generation model. Referring to the third image, it can be seen that the second partial virtual image (830) is an image of an object having the same class as the objects in the target image (810) but a different style, in a segmentation area for objects in the target image (810) of the first domain style. At this time, it can be seen that the shape of the second type of object is partially distorted in the second partial virtual image (830). For example, it can be seen that the lane, which is an object associated with vehicle driving, is distorted. Additionally, it can be seen that the second partial virtual image (830) has a different domain style from the target image (810).
[0097] The fourth image is an example of a second domain style virtual image (840) generated based on a first partial virtual image (820) associated with a first type of object and a second partial virtual image (830) associated with a second type of object. Specifically, the processor may generate a first image associated with a first type of object within a first domain style target image (810), and generate the first partial virtual image (820) based on the first image using a first image generation model. In addition, the processor may generate a second image associated with a second type of object within the first domain style target image (810), and generate the second partial virtual image (830) based on the second image using a second image generation model. Additionally, the first partial virtual image (820) and the second partial virtual image (830) may be combined and then a post-processing operation may be performed to generate the second domain style virtual image (840). Referring to the virtual image (840) of the second domain style and the target image (810) of the first domain style, it can be confirmed that the first type of object (e.g., a vehicle, a lane, etc.) in the virtual image (840) of the second domain style has the same appearance (or is similar) as the first type of object included in the target image (810) of the first domain style, and only the color, light intensity, etc. have been changed. In addition, it can be confirmed that the second type of object (e.g., the sky, a building, etc.) in the virtual image (840) of the second domain style has the same class information as the second type of object included in the target image (810) of the first domain style, but the style of the object has been changed.
[0098] FIG. 9 is a flowchart illustrating a virtual image generation method (900) according to one embodiment of the present disclosure. In one embodiment, the method (900) may be performed by at least one processor of an information processing system. The method (900) may begin with the processor acquiring a target image of a first domain style (S910). Here, the first domain style may be a virtual domain style.
[0099] Then, the processor can generate a first image of the first domain style representing a first type of object within the target image (S920). Here, the first type of object may be an object defined by being distinguished by instance object. Additionally or alternatively, the first type of object may include a dynamic object and a first static object. The first static object may be an object associated with traffic information. In addition, the first image may include RGB information for the first type of object.
[0100] Then, the processor can generate a second image of the first domain style representing a second type of object within the target image (S930). Here, the second type of object may be an object defined by a class according to the object's properties. Additionally or alternatively, the second type of object may include a second static object. Furthermore, the second image may include segmentation information for the second type of object.
[0101] Then, the processor can generate a first partial virtual image of a second domain style based on the first image using the first image generation model (S940). Here, the second domain style may be a real-world domain style. The first image generation model may be a model trained to generate an output image of the second domain style based on an input image of the first domain style.
[0102] Then, the processor can generate a second partial virtual image of the second domain style based on the second image using the second image generation model (S950). The second image generation model may be a model trained to generate an output image of the second domain style based on segmentation information.
[0103] Then, the processor can generate a second domain style virtual image based on the first partial virtual image and the second partial virtual image (S960). The step of generating the second domain style virtual image may include the step of generating a combined image by combining the first partial virtual image and the second partial virtual image, the step of extracting at least a portion of the area where the first partial virtual image and the second partial virtual image are adjacent within the combined image, and the step of converting first color information for at least a portion of the area.
[0104] In one embodiment, the processor may extract second color information corresponding to at least a portion of the target image to convert the first color information. Thereafter, the processor may convert the first color information for at least a portion of the combined image into second color information.
[0105] The above-described method may be provided as a computer program stored on a computer-readable recording medium for execution on a computer. The medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording means or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program instructions, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.
[0106] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, these techniques may be implemented in hardware, firmware, software, or a combination thereof. Those skilled in the art will appreciate that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software will depend on the particular application and the design requirements imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementations should not be construed as departing from the scope of the present disclosure.
[0107] In a hardware implementation, the processing units used to perform the techniques may be implemented within one or more ASICs, DSPs, GPUs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, a computer, or a combination thereof.
[0108] Accordingly, the various exemplary logical blocks, modules, and circuits described in connection with the present disclosure may be implemented or performed by any combination of a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or those designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0109] In a firmware and / or software implementation, the techniques may be implemented as instructions stored on a computer-readable medium, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, a compact disc (CD), a magnetic or optical data storage device, etc. The instructions may be executable by one or more processors and may cause the processor(s) to perform certain aspects of the functionality described herein.
[0110] When implemented in software, the techniques may be stored on or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media includes both computer storage media and communication media, including any medium that facilitates transfer of a computer program from one place to another. Storage media may be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. In addition, any connection is suitably made to a computer-readable medium.
[0111] For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, digital subscriber line, or wireless technologies such as infrared, radio, and microwave are included within the definition of media. Disk and disc, as used herein, includes compact discs, laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks usually reproduce data magnetically, whereas discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0112] A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. Alternatively, the processor and the storage medium may reside as discrete components in the user terminal.
[0113] While the embodiments described above have been described as utilizing aspects of the presently disclosed subject matter in one or more standalone computer systems, the present disclosure is not limited thereto and may be implemented in conjunction with any computing environment, such as a network or distributed computing environment. Furthermore, aspects of the present disclosure may be implemented in multiple processing chips or devices, and storage may be similarly affected across multiple devices. Such devices may include personal computers, network servers, and portable devices.
[0114] While the present disclosure has been described in connection with certain embodiments herein, various modifications and variations may be made without departing from the scope of the present disclosure, which would be apparent to those skilled in the art. Furthermore, such modifications and variations are intended to fall within the scope of the claims appended to this specification.
Claims
1. A method for creating a virtual image performed by at least one processor, Step of obtaining a target image of the first domain style; generating a first image of the first domain style representing a first type of object within the target image; generating a second image of the first domain style representing a second type of object within the target image; A step of generating a first partial virtual image of a second domain style based on the first image using a first image generation model; A step of generating a second partial virtual image of the second domain style based on the second image using a second image generation model; and A step of generating a virtual image of the second domain style based on the first partial virtual image and the second partial virtual image. Including, A method for generating a virtual image, wherein the first domain style and the second domain style are different.
2. In paragraph 1, The above first domain style is a virtual domain style, A method for creating a virtual image, wherein the second domain style is a real domain style.
3. In paragraph 1, The above first type of object is an object defined by classifying it by instance object, A method for creating a virtual image, wherein the above second type of object is an object defined by class according to the object's properties.
4. In paragraph 1, The first image includes RGB information for the first type of object, A method for generating a virtual image, wherein the second image includes segmentation information for an object of the second type.
5. In paragraph 1, The above first image generation model is, A virtual image generation method, wherein the model is trained to generate an output image of the second domain style based on an input image of the first domain style.
6. In paragraph 1, The second image generation model is, A virtual image generation method, wherein the model is trained to generate an output image of the second domain style based on segmentation information.
7. In paragraph 1, The step of creating a virtual image of the above second domain style is: A step of combining the first partial virtual image and the second partial virtual image to generate a combined image; A step of extracting at least a portion of an area where the first partial virtual image and the second partial virtual image are adjacent within the combined image; and A step of converting first color information for at least some of the above regions A method for creating a virtual image, comprising:
8. In paragraph 7, The step of converting the above first color information is: A step of extracting second color information corresponding to at least a portion of the target image; and A step of converting the first color information for at least some area within the combined image into the second color information. A method for creating a virtual image, comprising:
9. In paragraph 1, The first type of object includes a dynamic object and a first static object, The second type of object includes a second static object, A method for creating a virtual image, wherein the first static object is an object associated with traffic information.
10. A computer-readable non-transitory recording medium recording commands for executing the method according to paragraph 1 on a computer.
11. As an information processing system, Communication module; memory; and At least one processor connected to said memory and configured to execute at least one computer-readable program contained in said memory Including, At least one program above, Obtain the target image of the first domain style, Generating a first image of the first domain style representing a first type of object within the target image, Generating a second image of the first domain style representing a second type of object within the target image; Using the first image generation model, a first partial virtual image of a second domain style based on the first image is generated, Using the second image generation model, a second partial virtual image of the second domain style is generated based on the second image, Includes commands for generating a virtual image of the second domain style based on the first partial virtual image and the second partial virtual image, An information processing system wherein the first domain style is different from the second domain style.
Citation Information
Patent Citations
Hair style composition system and method the same
KR102284305B1
Interior partition
KR102468691B1
Method and system for evaluating autonomous vehicle based on mixed reality image
KR102548579B1
Method and system for training image generation model using content information
KR102624083B1