Hybrid Reality Video Generation Method and System Related to Traveling Environment

By generating mixed reality videos that integrate actual driving environment footage with virtual objects, the system addresses the challenges of testing autonomous vehicles, offering a cost-effective and location-independent solution for simulating diverse driving scenarios.

JP7694991B2Active Publication Date: 2025-06-18モライ インコーポレーティッド
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024527732
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-11-30
Filing Date
2023-11-01
Publication Date
2025-06-18
Estimated Expiration
2043-11-01

Smart Images

  • Figure 0007694991000001
    Figure 0007694991000001
  • Figure 0007694991000002
    Figure 0007694991000002
  • Figure 0007694991000003
    Figure 0007694991000003
Patent Text Reader

Abstract

The present disclosure relates to a method for generating mixed reality images related to a driving environment executed by at least one processor. [Solution] The mixed reality image generating method includes the steps of acquiring an actual image related to a driving environment, generating a first virtual image including objects of the driving environment, and generating a mixed reality image including objects of the driving environment based on the actual image and the first virtual image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method and a system for generating a mixed reality video related to a driving environment. Specifically, the present disclosure relates to a method and a system for generating a mixed reality video including an object in a driving environment based on an actual video related to the driving environment and a virtual video including an object in the driving environment.

Background Art

[0002] In recent years, with the advent of the Fourth Industrial Revolution era, autonomous vehicle-related technologies have attracted attention as future vehicles. Autonomous driving technology to which various advanced technologies including IT technology are applied is increasing in its phase as a new growth driver of the automotive industry.

[0003] On the other hand, an autonomous vehicle generally travels to a given destination by recognizing the surrounding environment and judging the driving situation to control the vehicle without the intervention of a driver, and thus is related to various social problems such as safety regulations and operation regulations. Therefore, in order to solve such social problems and commercialize autonomous driving technology, continuous performance tests are required.

[0004] However, as vehicle parts and software become more diverse and complex, there are problems in that there are location constraints and it takes a lot of time and cost to actually realize and test the driving environment in order to evaluate an autonomous vehicle.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The present disclosure provides a method for generating a mixed reality video related to a driving environment, a computer-readable non-transitory recording medium recording instruction words, and a system (device) for solving the above problems.

Means for Solving the Problems

[0006] The present disclosure can be implemented by various methods including a method, a system (apparatus), or a computer program stored in a readable storage medium.

[0007] According to an embodiment of the present disclosure, a method for generating a mixed reality video related to a driving environment, which is executed by at least one processor, includes: acquiring an actual video related to the driving environment; generating a first virtual video including objects in the driving environment; and generating a mixed reality video including objects in the driving environment based on the actual image and the first virtual video.

[0008] According to an embodiment of the present disclosure, the first virtual video includes an image generated from a simulation model that realizes the driving environment.

[0009] According to an embodiment of the present disclosure, the method further includes generating a second virtual video related to the driving environment, and the step of generating a mixed reality video based on the actual video and the first virtual video includes generating a mixed reality video based on the actual video, the first virtual video, and the second virtual video.

[0010] According to an embodiment of the present disclosure, the second virtual video includes a semantic image generated from a simulation model that realizes the driving environment.

[0011] According to an embodiment of the present disclosure, the step of acquiring an actual video related to the driving environment includes acquiring an actual video using a first camera mounted on a vehicle.

[0012] According to an embodiment of the present disclosure, the first virtual video is generated from data acquired using a second camera mounted on a vehicle realized in a simulation model as the same vehicle model as the vehicle, and the second camera is a virtual camera realized in the simulation model as a camera having the same external and internal parameters as the first camera.

[0013] According to an embodiment of the present disclosure, the second virtual video is generated from data obtained using a third camera mounted on a vehicle realized in a simulation model as the same vehicle model as the vehicle, and the third camera is a virtual camera realized in the simulation model as a camera having the same external and internal parameters as the first camera.

[0014] According to an embodiment of the present disclosure, the step of generating a mixed reality video based on the actual video and the first virtual video includes extracting a portion corresponding to an object from the first virtual video, and generating a mixed reality video based on the portion corresponding to the object and the actual video.

[0015] According to an embodiment of the present disclosure, the step of extracting a portion corresponding to an object includes extracting a portion corresponding to the object using the second virtual video.

[0016] According to an embodiment of the present disclosure, the step of extracting a portion corresponding to an object based on the second virtual video includes filtering the second virtual video to generate a mask image for the object, and extracting a portion corresponding to the object from the first virtual video using the mask image.

[0017] According to an embodiment of the present disclosure, the step of generating a mixed reality video based on the portion corresponding to the object and the actual video includes removing a portion of the actual video corresponding to the portion corresponding to the object.

[0018] According to an embodiment of the present disclosure, the step of removing a portion of the actual video corresponding to the portion corresponding to the object includes removing a portion of the actual video corresponding to the portion corresponding to the object using the second virtual video.

[0019] According to an embodiment of the present disclosure, the step of removing a part of the actual video corresponding to the part corresponding to the object using the second virtual video includes generating a mask inverse image from the second virtual video, and removing a part of the actual video corresponding to the part corresponding to the object using the mask inverse image.

[0020] According to an embodiment of the present disclosure, the step of generating a mask inverse image includes generating a mask image for the object from the second virtual video, and inverting the mask image to generate a mask inverse image.

[0021] A computer-readable non-transitory recording medium recording instructions for executing the above-described method according to an embodiment of the present disclosure on a computer is provided.

[0022] An information processing system according to an embodiment of the present disclosure includes a communication module, a memory, and at least one processor coupled to the memory and configured to execute at least one computer-readable program included in the memory. The at least one program includes instructions for acquiring an actual video related to a driving environment, generating a first virtual video including an object in the driving environment, and generating a mixed reality video including the object in the driving environment based on the actual video and the first virtual video.

Advantages of the Invention

[0023] According to an embodiment of the present disclosure, by fusing virtual objects into an actual driving environment, various test situations can be realized without being restricted by location constraints, and an autonomous driving vehicle can be tested.

[0024] According to an embodiment of the present disclosure, by visualizing and providing virtual objects in an actual driving environment, a tester riding in an autonomous driving vehicle can visually recognize a virtual test situation for the actual driving environment, and tests for various driving environments can be smoothly performed.

[0025] According to an embodiment of the present disclosure, instead of actually arranging physical objects for testing an autonomous vehicle, by providing an image in which virtual objects are projected onto a driving environment during the test process of the autonomous vehicle, the time and cost required for the test can be saved.

[0026] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by those having ordinary knowledge in the technical field to which the present disclosure belongs (referred to as "ordinary technicians (persons skilled in the art)") from the description of the claims.

[0027] Embodiments of the present disclosure are described with reference to the accompanying drawings described below, where like reference numerals indicate like elements, but are not limited thereto.

Brief Description of the Drawings

[0028]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0029] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the accompanying drawings. In the following description, specific descriptions of well-known functions and configurations may be omitted if they may unnecessarily obscure the gist of the present disclosure.

[0030] In the accompanying drawings, the same or corresponding components are given the same reference numerals. Also, in the description of the following embodiments, the description of the same or corresponding components may be omitted. However, even if the description of a component is omitted, it is not intended that such a component is not included in the embodiments having such a component.

[0031] The advantages and features of the disclosed embodiments, and the methods for achieving them, will become clear by referring to the embodiments described below together with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and can be realized in various different forms. These embodiments are provided only to make the present disclosure complete and to properly inform those skilled in the art of the scope of the invention.

[0032] Briefly explain the terms used in this specification and specifically describe the disclosed embodiments. The terms used in this specification are selected to be as general as possible and currently widely used while considering the functions in this disclosure. However, this may change due to the intentions or precedents of those skilled in the relevant fields, the emergence of new technologies, etc. In addition, in certain cases, there are terms arbitrarily selected by the applicant, and in such cases, the meaning thereof will be described in detail in the description of the corresponding invention. Therefore, the terms used in this disclosure should not be merely the names of the terms, but should be defined based on the meaning of the terms and the content throughout this disclosure.

[0033] In this specification, singular expressions include plural expressions unless clearly specified as singular in the context. Also, plural expressions include singular expressions unless clearly specified as plural in the context. When a certain part of the specification states that it includes a certain component, it means that, unless otherwise stated, it may further include other components rather than excluding other components.

[0034] Furthermore, as used herein, the terms "module" or "portion" mean software or hardware components, and the "module" or "portion" serves some role. However, the "module" or "portion" is not meant in the sense of being limited to software or hardware. The "module" or "portion" may be configured to be on an addressable storage medium and may also be configured to constitute one or more processors. Thus, by way of example, the "module" or "portion" may include components such as software components, object-oriented software components, class components, and task components, and at least one of a process, a function, an attribute, a procedure, a subroutine, a segment of program code, a driver, firmware, microcode, a circuit, data, a database, a data structure, a table, an array, or a variable. The components and the "module" or "portion" may be combined with a smaller number of components and "modules" or "portions" in which the provided function is, or may be further separated into additional components and "modules" or "portions".

[0035] According to an embodiment of the present disclosure, a "module" or a "unit" can be implemented by a processor and a memory. The "processor" should be broadly interpreted to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, and the like. In some environments, the "processor" may also refer to a custom semiconductor (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), and the like. The "processor" may refer to a combination of processing devices such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors coupled with a DSP core, or any other such configuration combination. Also, the "memory" should be broadly interpreted to include any electronic component capable of storing electronic information. The "memory" may refer to various types of processor-readable media such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage devices, registers, and the like. If the processor can read information from and / or write information to the memory, the memory is said to be in electronic communication with the processor. The memory integrated with the processor is in electronic communication with the processor.

[0036] In the present disclosure, "mixed reality (MR)" may refer to combining the virtual world and the real world to create new information such as a new environment or visualization. That is, it may refer to a technology that joins virtual reality to the real world so that real physical objects and virtual objects can interact with each other.

[0037] In the present disclosure, "actual video" may refer to video of a physical object or environment that has been materialized. For example, it may refer to video of an actually existing vehicle, roadway, and surrounding objects.

[0038] In the present disclosure, "virtual video" may refer to video of an object realized virtually, rather than video of a physical object that has been materialized. For example, it may refer to video of an object in a virtual reality environment.

[0039] In the present disclosure, "semantic image" may refer to an image to which semantic segmentation technology has been applied. In this case, semantic segmentation technology may refer to technology for predicting the class to which each pixel in an image belongs, as technology for dividing objects in the image into semantic units (e.g., objects).

[0040] In the present disclosure, "mask image" may refer to an image used to mask a specific part of an image. In this case, masking may refer to an operation for distinguishing between the area to which an effect is applied and the remaining area when applying a specific effect to a video.

[0041] FIG. 1 is an exemplary diagram in which a method for generating mixed reality video related to a driving environment according to an embodiment of the present disclosure is used. As shown in FIG. 1, the mixed reality video generation system includes a virtual video generation unit 120, an actual video processing unit 150, a virtual video processing unit 160, and a mixed reality video generation unit 170, and mixed reality video 190 can be generated from actual video 110 and virtual videos 130 and 140.

[0042] According to one embodiment, the actual video 110 can be generated from the first camera 100. In this case, the first camera 100 can be a camera mounted on a vehicle. Also, the actual video 110 can be a video related to the driving environment. For example, the first camera 100 mounted on the vehicle can generate the actual video 110 in which the driving road and surrounding objects are photographed. Then, the generated actual video 110 can be transmitted to the actual video processing unit 150.

[0043] According to one embodiment, the virtual video generation unit 120 can generate the virtual videos 130 and 140 and transmit them to the actual video processing unit 150 and the virtual video processing unit 160. In this case, the first virtual video 130 can include an image generated from a simulation model that realizes the driving environment. Also, the first virtual video 130 can be a video including objects in the driving environment (for example, traffic cones located on the road, other driving vehicles, etc.). The second virtual video 140 can include a semantic image generated from a simulation model that realizes the driving environment.

[0044] According to one embodiment, the first virtual video 130 can be generated from the second camera 122. Also, the second virtual video 140 can be generated from the third camera 124. The second camera 122 can be a virtual camera mounted on a vehicle realized in a simulation model as the same vehicle model as the vehicle on which the first camera 100 is mounted. Similarly, the third camera 124 can also be a virtual camera mounted on a vehicle realized in a simulation model as the same vehicle model as the vehicle on which the first camera 100 is mounted. Also, the second camera 122 and the third camera 124 can be virtual cameras realized in a simulation model as cameras with the same external parameters (for example, the orientation of the camera, etc.) and internal parameters (for example, focal length, optical center, etc.) as the first camera.

[0045] According to an embodiment, based on the actual video 110 and the first virtual video 130, a mixed reality video 190 including objects in the driving environment can be generated. For this purpose, an actual video 152 processed from the actual video 110 can be generated by the actual video processing unit 150. Also, a virtual video 162 processed from the first virtual video 130 can be generated by the virtual video processing unit 160. The processed actual video 152 and the processed virtual video 162 can be fused by the mixed reality video generation unit 170 to generate a mixed reality video 190 (180). Details regarding this will be described later with reference to FIGS. 4 to 6.

[0046] According to an embodiment, the second virtual video 140 can be used to generate the processed actual video 152 and the processed virtual video 162. For example, the second virtual video 140 can be used to extract a portion corresponding to an object from the first virtual video 130. Also, the second virtual video 140 can be used to remove the video of the portion where the object is located in the actual video 110. Details regarding this will be described later with reference to FIGS. 5 and 6.

[0047] With the configuration described above, instead of actually arranging physical objects for testing an autonomous vehicle, by providing a video in which virtual objects are projected onto the driving environment during the test process of the autonomous vehicle, the time and cost required for the test can be saved.

[0048] FIG. 2 is a block diagram showing the internal configuration of an information processing system 200 according to an embodiment of the present disclosure. The information processing system 200 corresponds to a mixed reality video generation system including the virtual video generation unit 120, the actual video processing unit 150, the virtual video processing unit 160, and the mixed reality video generation unit 170 of FIG. 1, and may include a memory 210, a processor 220, a communication module 230, and an input / output interface 240. The information processing system 200 can be configured to communicate information and / or data with an external system via a network using the communication module 230.

[0049] Memory 210 may include any non-transitory computer-readable recording medium. According to one embodiment, memory 210 may include a random access memory (RAM), a read-only memory (ROM), a disk drive, a solid state drive (SSD), a permanent mass storage device such as a flash memory, etc. As another example, a permanent mass storage device such as a ROM, an SSD, a flash memory, a disk drive, etc. may be included in the information processing system 200 as a separate permanent storage device distinct from the memory. Also, an operating system and at least one program code (for example, code for generating a mixed reality video provided and driven in the information processing system 200, etc.) may be stored in the memory 210.

[0050] Such software components may be loaded from a computer-readable recording medium separate from the memory 210. Such a separate computer-readable recording medium may include a recording medium directly connectable to such an information processing system 200, but may include, for example, a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, a memory card, etc. As another example, the software components may be loaded into the memory 210 via the communication module 230, which is not a computer-readable recording medium. For example, at least one program may be loaded into the memory 210 based on a computer program (for example, a program for generating a mixed reality video, etc.) installed by a file provided by a file distribution system that distributes developer or application installation files via the communication module 230.

[0051] The processor 220 may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. The instructions may be provided by the memory 210 or the communication module 230 to a user terminal (not shown) or other external systems. For example, the processor 220 may receive a plurality of information (e.g., actual video) necessary for generating mixed reality video, and generate mixed reality video related to the driving environment based on the received plurality of information.

[0052] The communication module 230 can provide a configuration or function for the information processing system 200 and a user terminal (not shown) to communicate with each other via a network, and may provide a configuration or function for the information processing system 200 to communicate with an external system (as an example, a separate cloud system, etc.). As an example, control signals, instructions, data, etc. provided under the control of the processor 220 of the information processing system 200 can be transmitted to the user terminal and / or the external system via the communication module 230 and the network, through the communication module of the user terminal and / or the external system. For example, the user terminal may receive the generated mixed reality video related to the driving environment.

[0053] Also, the input / output interface 240 of the information processing system 200 may be means for interfacing with the information processing system 200 or a device (not shown) for input or output that the information processing system 200 may include. In FIG. 2, the input / output interface 240 is shown as an element configured separately from the processor 220, but is not limited thereto, and the input / output interface 240 may be configured to be included in the processor 220. The information processing system 200 may include more components than those shown in FIG. 2. However, it is not necessary to clearly illustrate most of the conventional components.

[0054] The processor 220 of the information processing system 200 can be configured to manage, process, and / or store information and / or data received from a plurality of user terminals and / or a plurality of external systems. According to one embodiment, the processor 220 can receive actual images related to the driving environment. Thereafter, the processor 220 can generate a mixed reality image including objects in the driving environment based on the received actual images and the generated virtual images.

[0055] FIG. 3 is a diagram showing the internal configuration of the processor 220 of the information processing system 200 according to an embodiment of the present disclosure. According to one embodiment, the processor 220 can be configured to include a virtual image generation unit 120, an actual image processing unit 150, a virtual image processing unit 160, and a mixed reality image generation unit 170. The internal configuration of the processor 220 of the information processing system shown in FIG. 3 is merely an example, and at least a part of the configuration of the processor 220 may be omitted or other configurations may be added, and at least a part of the operations or processes executed by the processor 220 may be realized differently, such as being executed by the processor of a user terminal communicably connected to the information processing system. Note that in FIG. 3, the configuration of the processor 220 is described by being divided according to each function, but this does not necessarily mean that it is physically divided. For example, the actual image processing unit 150 and the virtual image processing unit 160 are described separately, but this is for helping the understanding of the invention and is not limited thereto.

[0056] According to one embodiment, the virtual image generation unit 120 can generate a virtual image for generating a mixed reality image related to the driving environment. For example, the virtual image generation unit 120 can generate a first virtual image including objects in the driving environment (e.g., traffic cones located on the road, other driving vehicles, etc.) and / or a second virtual image related to the driving environment. The first virtual image may include an image generated from a simulation model that realizes the driving environment. The second virtual image may include a semantic image generated from a simulation model that realizes the driving environment. In one embodiment, the first virtual image and the second virtual image can be generated from data acquired using a second camera and a third camera mounted on a vehicle realized in the simulation model. In this case, the vehicle realized in the simulation model can be the same model as the vehicle equipped with the first camera for acquiring actual images. Also, the second camera and the third camera can be virtual cameras realized in the simulation model as cameras having the same external parameters (e.g., camera orientation, etc.) and internal parameters (e.g., focal length, optical center, etc.) as the first camera 100.

[0057] According to one embodiment, the actual image processing unit 150 can generate a processed actual image for generating a mixed reality image from the actual image. For example, the actual image processing unit 150 can generate a processed actual image by removing a part of the actual image. Also, the actual image processing unit 150 can utilize the second virtual image to generate a processed actual image. For example, the actual image processing unit 150 can generate a mask inverse image from the second virtual image and generate a processed actual image by removing a part of the actual image using the mask inverse image. Details regarding this will be described later with reference to FIG. 6.

[0058] According to an embodiment, the virtual video processing unit 160 can generate a virtual video processed from the first virtual video. For example, the virtual video processing unit 160 can generate a processed virtual video by extracting a portion corresponding to an object from the first virtual video. Also, the virtual video processing unit 160 can use the second virtual video to generate the processed virtual video. For example, the virtual video processing unit 160 can generate a mask image from the second virtual video and use the mask image to extract a portion corresponding to an object from the first virtual video, thereby generating a processed virtual video. Details regarding this will be described later with reference to FIG. 5.

[0059] According to an embodiment, the mixed reality video generation unit 170 can generate a mixed reality video based on the processed virtual video and the processed actual video. For example, the mixed reality video generation unit 170 can generate a mixed reality video by fusing the processed virtual video and the processed actual video. Details regarding this will be described later with reference to FIG. 4.

[0060] FIG. 4 is an exemplary diagram in which a mixed reality video 190 is generated from an actual video 110 and virtual videos 130, 140 according to an embodiment of the present disclosure. As shown in the figure, the mixed reality video 190 can be generated based on the first virtual video 130, the second virtual video 140, and the actual video 110.

[0061] According to an embodiment, the actual video 110 may be a video related to the driving environment. For example, the actual video 110 may be a video that captures a driving road for testing an autonomous vehicle. The first virtual video 130 may include objects in the driving environment as an image generated from a simulation model that realizes the same driving environment as the actual video. For example, the first virtual video 130 may be a video captured by a camera in the simulation model of peripheral vehicles or traffic cones existing in the simulation model that realizes the driving road of the actual video. The second virtual video 140 may be a semantic image generated from a simulation model that realizes the same driving environment as the actual video. For example, the second virtual video 140 may be a video in which semantic segmentation technology is applied to a video captured by a camera in the simulation model of peripheral vehicles or traffic cones existing in the simulation model that realizes the driving road of the actual video, and is a video segmented by object.

[0062] According to an embodiment, the first virtual video 130 may be processed to generate a processed virtual video 162 (410). More specifically, the processed virtual video 162 may be generated by being processed by the virtual video processing unit 160 and extracting portions corresponding to the objects from the first virtual video 130. For example, by extracting portions corresponding to the vehicles and traffic cones existing in the first virtual video, a processed virtual video 162 in which only the extracted portions exist may be generated. Also, in the process of processing the first virtual video 130, the second virtual video 140 may be used. For example, a mask image for the vehicles and traffic cones existing in the first virtual video may be generated using the second virtual video 140, and a video in which portions corresponding to the vehicles and traffic cones are extracted may be generated using the generated mask image. Details regarding this will be described later with reference to FIG. 5.

[0063] According to an embodiment, the actual video 110 can be processed to generate a processed actual video 152 (420). More specifically, the processed actual video 152 can be generated by processing in the actual video processing unit 150 and removing a part of the actual video 110 corresponding to the part corresponding to the object of the first virtual video in the actual video 110. For example, a part of the actual video 110 corresponding to the part where the vehicle and traffic cone of the first virtual video are located can be removed to generate the processed actual video 152. Also, in the process of processing the actual video 110, the second virtual video 140 can be used. For example, using the second virtual video 140, a mask inverse image for the vehicle and traffic cone existing in the first virtual video is generated, and using the generated mask inverse image, a video with a part of the actual video 110 removed can be generated. Details regarding this will be described later with reference to FIG. 6.

[0064] According to an embodiment, the processed virtual video 162 and the processed actual video 152 can be fused to generate a mixed reality video 190 (180). More specifically, the mixed reality video 190 can be generated by the mixed reality video generation unit 170 and projecting the objects existing in the processed virtual video 162 onto the processed actual video 152. For example, the mixed reality video 190 can be generated by projecting and fusing the vehicle and traffic cone existing in the processed virtual video 162 onto the removed part of the processed actual video 152.

[0065] FIG. 5 is an exemplary diagram showing that a processed virtual video 162 is generated from the first virtual video 130 and the second virtual video 140 according to an embodiment of the present disclosure. According to an embodiment, the processed virtual video 162 can be generated by extracting the part corresponding to the object from the first virtual video 130 (530). For example, by extracting the part corresponding to the vehicle and traffic cone from the first virtual video 130, a processed virtual video 162 with the remaining part excluding the part corresponding to the vehicle and traffic cone removed can be generated.

[0066] According to an embodiment, a mask image 520 may be used to extract a portion corresponding to an object from the first virtual video 130. For example, the mask image 520 is an image that can mask a portion corresponding to an object in the first virtual video 130. By applying the mask image 520 to the first virtual video 130, portions other than the masked object portion can be removed. In this way, by masking the portion corresponding to the object in the first virtual video 130 using the mask image 520, the portion corresponding to the object can be extracted. In this case, the mask image 520 may be a binary type image.

[0067] According to an embodiment, the mask image 520 may be generated from the second virtual video 140. For example, the second virtual video 140 is an image separated by object using semantic segmentation technology. The mask image 520 may be generated (510) by filtering a portion corresponding to an object in the second virtual video.

[0068] FIG. 6 is an exemplary diagram in which a processed real video 152 is generated from the real video 110 and the second virtual video 140 according to an embodiment of the present disclosure. According to an embodiment, the processed real video 152 may be generated (630) by removing a part of the real video 110 corresponding to the portion corresponding to the object of the first virtual video 130. For example, the processed real video 152 may be generated by removing a part of the real video 110 corresponding to the portion where the vehicle and traffic cone of the first virtual video are arranged in the real video 110.

[0069] According to an embodiment, a mask inverse image 620 may be used to remove a part of the real video 110. The mask inverse image is an image that can mask the remaining portion excluding the portion corresponding to the object in the first virtual video. By applying the mask inverse image 620 to the real video 110, only the portion corresponding to the portion corresponding to the object of the first virtual video can be removed.

[0070] According to one embodiment, the mask inverse image 620 can be generated from the second virtual image 140. For example, a mask image 520 is generated (510) by filtering the second virtual image 140, and the mask inverse image 620 can be generated (610) by inverting the mask image 520. The mask inverse image 620 can be a binary type image similar to the mask image 520.

[0071] FIG. 7 is a diagram showing an example of the first virtual image 130 according to one embodiment of the present disclosure. According to one embodiment, the first virtual image 130 can be an image generated from a simulation model that realizes the same driving environment as the actual image. More specifically, the first virtual image 130 can be generated from data obtained using a camera mounted on a vehicle realized in the simulation model as the same vehicle model as the vehicle on which the camera that captured the actual image was mounted. Also, the camera used to generate the first virtual image 130 can be a virtual camera realized in the simulation model as a camera having the same external parameters (e.g., camera orientation, etc.) and internal parameters (e.g., focal length, optical center, etc.) as the camera that captured the actual image.

[0072] According to one embodiment, the first virtual image 130 can include objects in the driving environment. For example, the first virtual image 130 can be an image including vehicles 720 traveling around and traffic cones 710_1, 710_2, 710_3, 710_4 placed around. Such objects in the first virtual image 130 can be extracted and processed by the mask image and fused with the actual image to generate a mixed reality image.

[0073] FIG. 8 is a diagram showing an example of an actual video 110 according to an embodiment of the present disclosure. According to one embodiment, the actual video 110 may be a video related to a driving environment. For example, the actual video 110 may be a video including a driving road and objects located in the vicinity (e.g., surrounding vehicles, etc.). Also, the actual video 110 may be acquired by a camera mounted on the vehicle.

[0074] According to one embodiment, the actual video 110 may be processed for generating a mixed reality video. More specifically, a part corresponding to a part of the object of the first virtual video in the actual video 110 may be removed. For example, a part 820 corresponding to a part where the vehicle of the first virtual video is located and parts 810_1, 810_2, 810_3, 810_4 corresponding to parts where traffic cones are located may be removed by a mask inverse image.

[0075] FIG. 9 is a diagram showing an example of a mixed reality video 190 according to an embodiment of the present disclosure. According to one embodiment, the mixed reality video 190 may be generated by fusing a processed virtual video and a processed actual video. The mixed reality video 190 may be generated by projecting a part corresponding to an object of the first virtual video extracted from the actual video. For example, by combining the surrounding vehicles 920 of the first virtual video and the traffic cones 910_1, 910_2, 910_3, 910_4 in the actual video including the driving road, a mixed reality video 190 in which virtual objects and the actual video are fused may be generated.

[0076] In this way, by visualizing and providing virtual objects in the actual driving environment, a tester riding in an autonomous vehicle can visually recognize the test situation, and the test can be smoothly performed. Also, instead of actually arranging physical objects for testing the autonomous vehicle, by providing a video in which virtual objects are projected in the driving environment during the test process of the autonomous vehicle, the time and cost required for the test can be saved.

[0077] FIG. 10 is a flowchart showing a mixed reality video generation method 1000 related to a driving environment according to an embodiment of the present disclosure. The method 1000 can be executed by at least one processor (e.g., processor 220) of an information processing system. As shown in the figure, the method 1000 can be started by acquiring an actual video related to the driving environment (S1010). In one embodiment, the actual video related to the driving environment can be acquired using a first camera mounted on a vehicle.

[0078] Thereafter, the processor can generate a first virtual video including objects in the driving environment (S1020). In one embodiment, the first virtual video can include an image generated from a simulation model that realizes the driving environment. Also, the first virtual video can be generated from data acquired using a second camera mounted on a vehicle that is realized in the simulation model as the same vehicle model as the vehicle on which the camera used to acquire the actual video is mounted. The second camera can be a virtual camera realized in the simulation model as a camera having the same external and internal parameters as the first camera.

[0079] Finally, the processor can generate a mixed reality video including objects in the driving environment based on the actual video and the first virtual video (S1030). In one embodiment, the processor can extract a portion corresponding to an object from the first virtual video and generate a mixed reality video based on the portion corresponding to the object and the actual video. Also, the processor can remove a part of the actual video corresponding to the portion corresponding to the object.

[0080] In one embodiment, the processor may generate a second virtual image related to the driving environment and generate a mixed reality image based on the actual image, the first virtual image, and the second virtual image. The second virtual image may include a semantic image generated from a simulation model that realizes the driving environment. Also, the second virtual image is generated from data obtained using a third camera mounted on a vehicle realized in the simulation model as the same vehicle model as the vehicle on which the camera used to acquire the actual image is mounted, and the third camera may be a virtual camera realized in the simulation model as a camera having the same external and internal parameters as the first camera.

[0081] In one embodiment, the processor may use the second virtual image to extract a portion corresponding to an object from the first virtual image. The processor may filter the second virtual image to generate a mask image for the object and use the mask image to extract a portion corresponding to the object from the first virtual image.

[0082] In one embodiment, the processor may use the second virtual image to remove a part of the actual image corresponding to the portion of the object in the first virtual image. The processor may generate an inverse mask image from the second virtual image and use the mask image to remove a part of the actual image corresponding to the portion of the object. The inverse mask image may be generated by generating a mask image for the object from the second virtual image and inverting the mask image to generate the inverse mask image.

[0083] The flowchart shown in FIG. 10 and the foregoing description are merely examples, and in some embodiments, they may be implemented differently. For example, in some embodiments, the order of each step may be changed, some steps may be repeated, some steps may be omitted, or some steps may be added.

[0084] The prior method can be provided as a computer program stored in a computer-readable recording medium for execution by a computer. The medium can be one that continuously stores a computer-executable program or temporarily stores it for execution or download. Further, the medium can be various recording or storage means in a form combined with single or multiple hardware, but is not limited to a medium directly connected to any computer system and can also exist distributed on a network. Examples of the medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto optical media such as floptical disks, and those configured to store program instruction words including ROM, RAM, flash memory, etc. Also, as examples of other media, there can be mentioned recording media or storage media managed by app stores that distribute applications, sites that supply or distribute various other software, servers, etc.

[0085] The methods, operations, or techniques of the present disclosure can also be implemented by various means. For example, such techniques can be implemented by hardware, firmware, software, or combinations thereof. Those of ordinary skill in the art will understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure of this application can also be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate such interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been generally described above in terms of their functional aspects. Whether such functions are implemented as hardware or software depends on the specific application and the design requirements imposed on the overall system. Those of ordinary skill in the art can implement the functions described in various ways for each specific application, but such implementation should not be construed as departing from the scope of the present disclosure.

[0086] In a hardware implementation, the processing units used to execute the techniques can be implemented within one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described in the present disclosure, computers, or combinations thereof.

[0087] Accordingly, the various illustrative logical blocks, modules, and circuits described in connection with the present disclosure may be implemented or executed by any combination of a general purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of devices designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other configuration.

[0088] In a firmware and / or software implementation, the techniques may be realized in instructions stored on a computer-readable medium such as a random access memory (RAM), read-only memory (ROM), nonvolatile random access memory (NVRAM), PROM (programmable read-only memory), EPROM (erasable programmable read-only memory), EEPROM (electrically erasable PROM), flash memory, compact disc (CD), magnetic or optical data storage device, etc. The instructions may be executable by one or more processors and may also cause the one or more processors to perform particular aspects of the functions described in the present disclosure.

[0089] The embodiments described above are described as utilizing aspects of the presently disclosed subject matter in one or more stand-alone computer systems, but the disclosure is not limited thereto and may be implemented in conjunction with any computing environment such as a network or a distributed computing environment. Further, aspects of the subject matter in the disclosure may be implemented on multiple process chips or devices, and storage may be similarly affected across multiple devices. Such devices may include PCs, network servers, and portable devices as well.

[0090] Although the present disclosure has been described in connection with some embodiments, various modifications and changes can be made without departing from the scope of the present disclosure that can be understood by those of ordinary skill in the art to which the invention of the present disclosure pertains. Also, such modifications and changes should be considered to fall within the scope of the appended claims of this specification.

Claims

1. In a method for generating mixed reality video related to a driving environment, which is executed by at least one processor, obtaining an actual video related to the driving environment; generating a first virtual video including objects in the driving environment; generating a second virtual video related to the driving environment to process the actual video and the first virtual video; generating a mixed reality video including objects in the driving environment based on the actual video, the first virtual video, and the second virtual video; including the second virtual video includes a semantic image generated from a simulation model that realizes the driving environment, The step of generating a mixed reality video based on the actual video, the first virtual video, and the second virtual video includes: generating a virtual video processed by extracting a part corresponding to an object from the first virtual video using the second virtual video; generating an actual video processed by removing a part of the actual video corresponding to the part corresponding to the object using the second virtual video; generating a mixed reality video based on the processed virtual video and the processed actual video. A method for generating a mixed reality video.

2. The method for generating a mixed reality video according to claim 1, wherein the first virtual video includes an image generated from a simulation model that realizes the driving environment.

3. The step of obtaining an actual video related to the driving environment includes: The method for generating a mixed reality video according to claim 1, including obtaining the actual video using a first camera mounted on a vehicle.

4. The first virtual video is Generated from data acquired using a second camera mounted on a vehicle realized as a simulation model as the same vehicle model as the said vehicle, The second camera, The method for generating a mixed reality video according to claim 3, wherein the second camera is a virtual camera realized as a simulation model as a camera having the same external parameters and internal parameters as the first camera.

5. The second virtual video, Generated from data acquired using a third camera mounted on a vehicle realized as a simulation model as the same vehicle model as the said vehicle, The third camera, The method for generating a mixed reality video according to claim 3, wherein the third camera is a virtual camera realized as a simulation model as a camera having the same external parameters and internal parameters as the first camera.

6. The step of generating a virtual video processed by extracting a portion corresponding to the object from the first virtual video using the second virtual video includes: Filtering the second virtual video to generate a mask image for the object; Extracting a portion corresponding to the object from the first virtual video using the mask image. The method for generating a mixed reality video according to claim 1.

7. The step of generating an actual video processed by removing a part of the actual video corresponding to the portion corresponding to the object using the second virtual video includes: Generating a mask inverse image from the second virtual video; Removing a part of the actual video corresponding to the portion corresponding to the object using the mask inverse image. The method for generating a mixed reality video according to claim 1.

8. The step of generating the mask inverse image includes: generating a mask image for the object from the second virtual image; inverting the mask image to generate the mask inverse image; The mixed reality video generation method according to claim 7.

9. A computer-readable non-transitory recording medium recording instruction codes for executing the method according to claim 1 by a computer.

10. As an information processing system, a communication module; a memory; at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory; The at least one program includes: acquiring actual video related to the driving environment; generating a first virtual image including an object in the driving environment; generating a second virtual image related to the driving environment to process the actual video and the first virtual image; including instruction codes for generating a mixed reality video including an object in the driving environment based on the actual video, the first virtual image, and the second virtual image; The second virtual image includes a semantic image generated from a simulation model realizing the driving environment; Generating a mixed reality video based on the actual video, the first virtual image, and the second virtual image includes: generating a virtual video processed by extracting a portion corresponding to an object from the first virtual image using the second virtual image; generating an actual video processed by removing a part of the actual video corresponding to the portion corresponding to the object using the second virtual image; An information processing system including generating a mixed reality video based on the processed virtual video and the processed actual video.

Citation Information

Patent Citations

  • Generation of a virtual world to assess real-world video analysis performance

    JP2017151973A

  • Image synthesizer

    JP2017219969A

  • Method and apparatus for identifying driving lane

    JP2019067364A

  • Image processing device and program

    JP2020135525A

  • Use of expansion of image having object obtained by simulation for machine learning model training in autonomous driving application

    JP2021176077A