Device for generating pseudo image of dangerous scene during vehicle driving

The device uses a machine learning model to superimpose realistic images of pedestrians or obstacles onto unsafe driving scenarios, addressing the unnatural appearance of such images in existing systems and improving driver awareness.

JP2025161171APending Publication Date: 2025-10-24TOYOTA JIDOSHA KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024064135
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-11
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing technologies struggle to generate realistic images of dangerous driving situations, as superimposed images of pedestrians or obstacles often appear unnaturally, failing to effectively alert drivers to potential dangers.

Method used

A device uses a machine learning algorithm to create a model that superimposes images of pedestrians or obstacles onto unsafe driving scenarios, utilizing corrected vehicle surroundings images and semantic segmentation, video inpainting, and GANs to ensure realistic and natural integration.

Benefits of technology

The device generates near-miss images that realistically depict pedestrians or obstacles, enhancing the driver's perception of danger and reducing the effort required for image creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025161171000001_ABST
    Figure 2025161171000001_ABST
Patent Text Reader

Abstract

To generate a pseudo image on which an image of a pedestrian or the like is superimposed so as to more realistically appear in images around a vehicle photographed during various unsafe driving as an image of a dangerous scene to be shown to a person to be diagnosed or a person to be educated when diagnosing or educating a vehicle driving technique.SOLUTION: A device 1 for generating a pseudo image of a dangerous scene during driving of a vehicle 10 generates a near-miss image in which an image of a pedestrian or the like is superimposed on an unsafe driving image photographed when an on-vehicle camera 12 executes unsafe driving. The near-miss image is generated by inputting the unsafe driving image to an image generation model that has learned to output one vehicle surrounding image when one corrected vehicle surrounding image obtained from one vehicle surrounding image in which the image of the pedestrian or the like is captured is input according to a machine learning algorithm using a plurality of vehicle surrounding images in the image of the pedestrian or the like is captured and a corrected vehicle surrounding image obtained by correcting the vehicle surrounding images to a state in which the image of the pedestrian or the like is not captured as learning data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an apparatus for generating simulated images of dangerous situations that may be encountered while driving a vehicle such as an automobile. [Background technology]

[0002] When diagnosing or training the driving skills of vehicles such as automobiles, images of dangerous situations, such as pedestrians suddenly appearing, are sometimes displayed on a television, PC, or driving simulator monitor screen and shown to a person being diagnosed or trained who will be the driver of the vehicle, in order to analyze the behavior of the person being diagnosed or trained and to alert them to dangerous situations while driving the vehicle. To this end, a technology has been proposed for generating pseudo-images of dangerous situations while driving a vehicle to be shown to the person being diagnosed or trained. For example, Patent Document 1 discloses a technology that displays superimposed images in nearly real time on a display unit positioned so as to block the driver's direct view of the vehicle, thereby recreating dangerous situations with a high degree of realism for the driver while the vehicle is in motion. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2012-22215 Summary of the Invention [Problem to be solved by the invention]

[0004] As described above, when diagnosing or training vehicle driving skills, images of dangerous situations that a person being diagnosed or trained may encounter while driving a vehicle may be shown to the person being diagnosed or trained. For example, images (moving or still images) of the area around the vehicle captured by an in-vehicle camera such as a drive recorder capturing images of the vehicle's surroundings when the person being diagnosed or trained is actually engaged in driving that could have resulted in a dangerous situation (driving that would have resulted in contact with a pedestrian or other obstacle (e.g., pedestrian) had it been present - hereinafter referred to as "unsafe driving") may be used. These images are superimposed on the road surface to make the image of a pedestrian or other obstacle appear at an appropriate position (hereinafter referred to as "near-miss image"). In this case, it is desirable that the image of a pedestrian or other obstacle in the near-miss image appear as realistically as possible on the road surface (if the image of a pedestrian or other obstacle appears unrealistically and unnaturally, the person being diagnosed or trained may not realize the danger even when viewing such an image). In this regard, since there are various situations in which unsafe driving is carried out, it is advantageous if images of pedestrians, etc. can be superimposed as realistically as possible on the image of the road surface in any of the images of unsafe driving in various situations.

[0005] In view of the above circumstances, the main objective of the present invention is to generate an image in which images of pedestrians and the like are superimposed so as to appear more realistically in an image of the surroundings of a vehicle taken during various unsafe driving situations. [Means for solving the problem]

[0006] According to the present invention, the above problem is solved by a device for generating a pseudo image of a dangerous scene while driving a vehicle, the device comprising: an unsafe driving image acquisition means for acquiring an unsafe driving image, which is an image taken while the vehicle is traveling by an on-board camera that captures images of the surroundings of the vehicle, when it is determined that unsafe driving has been performed; a near-miss image generating means for generating a near-miss image, which is an image in which an image of a pedestrian or other obstacle is superimposed on the unsafe driving image; Including, The near-miss image generating means uses as learning data a plurality of vehicle surroundings images, which are images taken while the vehicle is traveling by an on-board camera that captures images of the surroundings of the vehicle and which show images of pedestrians or other obstacles, and corrected vehicle surroundings images obtained by correcting each of the plurality of vehicle surroundings images so that images of the pedestrians or other obstacles are absent, to create an image generation model constructed according to a machine learning algorithm, the image generation model being trained to output a vehicle surroundings image showing an image of the pedestrian or other obstacle when a corrected vehicle surroundings image obtained from a vehicle surroundings image showing an image of the pedestrian or other obstacle is input, and when the unsafe driving image is input to the image generation model, a near-miss image in which an image of the pedestrian or other obstacle is superimposed on the unsafe driving image is output. This is achieved by:

[0007] In the above-described configuration of the present invention, a "dangerous situation while driving a vehicle" refers to a situation in which a pedestrian or other obstacle (e.g., a pedestrian) would have been present while the vehicle was being driven, and the "pseudo image" refers to an image in which an image of a pedestrian or other obstacle is superimposed on an image in which no pedestrian or other obstacle is actually present. The "on-board camera capturing images of the surroundings of the vehicle" refers to an on-board camera capturing images of the surroundings in any direction of the front, rear, left, or right of the vehicle, and may be a drive recorder camera. As described above, "unsafe driving" refers to a situation in which a pedestrian or other obstacle (e.g., a pedestrian) would have been present, and may be detected by any model that detects unsafe driving in any manner using images from the on-board camera. A "near miss image" refers to an image in which an image of a pedestrian or other obstacle is superimposed as realistically as possible on an "unsafe driving image" when unsafe driving is determined to have been performed, i.e., a "pseudo image of a dangerous situation while driving a vehicle." The "plurality of vehicle surroundings images, which are images taken by on-board cameras that capture images of the vehicle's surroundings while the vehicle is traveling and which contain images of pedestrians or other obstacles," refer to images taken by on-board cameras of various vehicles while the vehicle is traveling, which contain images of pedestrians or other obstacles, and may be images stored in any open data source of images, for example, an image database on a cloud network. The "corrected vehicle surroundings image obtained by correcting each of the plurality of vehicle surroundings images so that there are no images of pedestrians or other obstacles" refers to an image of a location where there are no images of pedestrians or other obstacles, which is obtained by removing images of pedestrians or other obstacles from the vehicle surroundings images that contain images of pedestrians or other obstacles, and may be prepared by any method. Such a corrected vehicle surroundings image can be generated, for example, by extracting and removing images of pedestrians or other obstacles from a vehicle surroundings image that contains such images using semantic segmentation technology, and then masking the removed areas of the image using video inpainting (image inpainting) technology or the like, thereby correcting the image so that it appears natural and continuous with the surrounding images.The "image generation model" in the near-accident image generating means is configured by learning to use multiple vehicle surroundings images and corrected vehicle surroundings images generated from them as learning data, and when a corrected vehicle surroundings image (not showing an image of a pedestrian or other obstacle) obtained from a vehicle surroundings image showing an image of a pedestrian or other obstacle is input, it outputs the vehicle surroundings image showing an image of the pedestrian or other obstacle. Thus, when an unsafe driving image is input to the image generation model in the near-accident image generating means, a near-accident image in which an image of a pedestrian or other obstacle is superimposed on the unsafe driving image is output. Note that an algorithm such as GAN (generative adversarial network) may be used to generate the near-accident image.

[0008] According to the configuration of the present invention, in generating near miss images, which are pseudo images of dangerous situations while driving a vehicle, a model is used that is trained to output a vehicle surroundings image on the corrected vehicle surroundings image using multiple vehicle surroundings images that show images of pedestrians, etc., and corrected vehicle surroundings images obtained from each of the vehicle surroundings images that do not show images of pedestrians, etc. In this model, images of pedestrians, etc. that were originally shown on the corrected vehicle surroundings image that does not show images of pedestrians, etc. are superimposed, so that images of pedestrians, etc. are expected to appear in a realistic and natural manner in the output vehicle surroundings image. Therefore, when an unsafe driving image is input, images of pedestrians, etc. are expected to appear in a realistic and natural manner in the unsafe driving image. As a result, it is expected that a diagnosed or trained person who views the near miss images generated by the device of the present invention will feel a stronger sense of danger from unsafe driving. [Effects of the Invention]

[0009] Thus, in the device for generating pseudo images of dangerous situations while driving a vehicle according to the present invention, an image generation model trained by a machine learning algorithm is used to superimpose an image of a pedestrian or the like onto any unsafe driving image, so that an image containing the original image of the pedestrian or the like is output on an image obtained by removing the image of the pedestrian or the like from an image containing the pedestrian or the like. This generates a "near miss image" by superimposing an image of a pedestrian or the like onto an arbitrary unsafe driving image. According to this method, the image generation model adjusts the position and size of the superimposed image of the pedestrian or the like within the image so that the image of the road surface around the vehicle in the unsafe driving image appears realistic and natural. This allows the image of the pedestrian or the like to appear more realistic and natural than if the image of the pedestrian or the like were simply pasted onto the image of the road surface around the vehicle. Alternatively, the image generation model eliminates the need for the image creator to adjust the position, orientation, and size of the superimposed image of the pedestrian or the like onto the image of the road surface around the vehicle so that the image appears as realistic as possible, significantly reducing the effort and time required by the image creator.

[0010] Other objects and advantages of the present invention will become apparent from the following description of preferred embodiments of the invention. [Brief explanation of the drawings]

[0011] [Figure 1] Fig. 1(A) is a schematic diagram of a vehicle driving skill diagnosis or training system that uses images generated by the device of this embodiment. Fig. 1(B) is a block diagram showing the configuration of the device of this embodiment. [Figure 2] Fig. 2(A) is a diagram that schematically shows an image of the area around a vehicle that includes images of pedestrians, etc. Fig. 2(B) is a diagram that schematically shows images of pedestrians, etc. extracted from the image of Fig. 2(A). Fig. 2(C) is a diagram that schematically shows the state in which the images of pedestrians, etc. have been removed from the image of Fig. 2(A). Fig. 2(D) is a diagram that schematically shows an image in which the areas in the image of Fig. 2(C) that contained images of pedestrians, etc. have been masked so that they naturally blend into the background, resulting in an image without images of pedestrians, etc. [Figure 3]FIG. 3(A) is a schematic diagram of an input image and an output image during learning (training) of an image generation model in the device of this embodiment, and FIG. 3(B) is a schematic diagram of an input image and an output image in the image generation model. [Explanation of symbols]

[0012] 1... Near miss image generating device, 2... display monitor, 10... vehicle, 12... in-vehicle camera, 13... image memory, 30... network, P... person to be diagnosed or person to be educated BEST MODE FOR CARRYING OUT THE INVENTION

[0013] Configuration for diagnosis or training of vehicle driving skills As already mentioned, the device according to this embodiment is a device that generates pseudo images of dangerous situations while driving a vehicle to be shown to a person being diagnosed or trained in vehicle driving skills diagnosis or training. In vehicle driving skill diagnosis or training using images generated by the device according to this embodiment, as shown in Fig. 1(A), images (moving or still images) of dangerous situations such as a pedestrian suddenly appearing are displayed on a monitor screen 2 of a television, PC, or driving simulator and shown to the person being diagnosed or trained P, thereby analyzing the behavior of the person being diagnosed or trained and alerting them to dangerous situations while driving a vehicle. Since it is not possible to intentionally place pedestrians or the like in actual dangerous situations when taking images of such dangerous scenes, the device (near miss image generating device) 1 according to this embodiment generates "near miss images" in which images of pedestrians or the like are superimposed on images taken by an in-vehicle camera 12 such as a drive recorder when the diagnosed or trainee P engages in unsafe driving, i.e., driving that could have led to a dangerous situation, while driving their own vehicle 10, thereby attempting to make the diagnosed or trainee P who sees the images realize the danger of their own unsafe driving. In this regard, it is preferable to be able to generate near miss images in which images of pedestrians or the like are superimposed so that they appear as realistically as possible in images of unsafe driving in various situations (unsafe driving images). Therefore, as will be described in detail below, the device 1 according to this embodiment collects images of the surroundings of the vehicle taken by on-board cameras of various or multiple vehicles under various circumstances while the vehicle is traveling, and images showing images of pedestrians or other obstacles (vehicle surrounding images) from various available sources such as an open data source 30 on a network, and uses the collected vehicle surrounding images to prepare an image generation model that generates pseudo images in which images of pedestrians, etc. are superimposed as realistically as possible on images of the surroundings of the vehicle in which no images of pedestrians, etc. are shown, according to a machine learning algorithm, in an attempt to generate more realistic near-miss images. The near-miss image generation device 1 is configured as a computer device that operates according to a program.

[0014] Configuration and operation of the near miss image generating device (a) Overview As shown in FIG. 1B, the near-miss image generating device according to this embodiment includes a pseudo image generating unit and a model learning unit. The pseudo image generating unit, in brief, generates a pseudo image (near-miss image) by superimposing images of pedestrians and other similar objects onto unsafe driving images captured by an onboard camera 12 during unsafe driving behavior in a vehicle 10 driven by a person being diagnosed or trained for driving skill diagnosis or training. The model learning unit, in brief, collects images of the vehicle's surroundings, including images of pedestrians and other similar objects, captured by an onboard camera such as a drive recorder (DR) under various circumstances while the vehicle is traveling, and uses these images to construct an image generation model that superimposes images of pedestrians and other similar objects onto unsafe driving images so that they appear as realistically as possible. The pseudo image generating unit and the model learning unit will be described in detail below.

[0015] (b) Model learning section In the model learning unit, first, vehicle surrounding images, which show images of pedestrians and the like, captured by an onboard camera such as a drive recorder (DR) while various vehicles are traveling under various conditions, such as various times and weather conditions, are stored in an image database (DR image DB). As already mentioned, the vehicle surrounding images may be collected from any location, for example, from an open data source provided via a cloud network, the Internet, etc. (They do not have to be captured in a vehicle driven by the person being diagnosed or trained). The images may be either moving images or still images.

[0016] Next, from the images around the vehicle stored in the image database, the learning data generation unit generates images to be used for learning or training the image generation model. Specifically, first, from the image around the vehicle that shows images of pedestrians P as shown in Figure 2(A), the image of pedestrians P as shown in Figure 2(B) is extracted pixel by pixel using semantic segmentation technology or the like, and the image of pedestrians P is removed from the image around the vehicle. As a result, the area S where the image of pedestrians P was located becomes shadowed as shown in Figure 2(C). Therefore, this area S is masked using video inpainting technology or the like to make it naturally continuous with its surroundings, and a pseudo image is generated as if the image of pedestrians P had not been included in the original image around the vehicle, as shown in Figure 2(D). Then, the vehicle surroundings image without the image of the pedestrian etc. P (corrected vehicle surroundings image) and the vehicle surroundings image with the original image of the pedestrian etc. P (original vehicle surroundings image) are stored in a database (learning data DB) as data for learning the image generation model.

[0017] Thus, the model generation unit uses the above-mentioned learning data to learn or train the image generation model according to an appropriate machine learning algorithm. In this training, an original vehicle surroundings image containing images of pedestrians, etc. P and a corrected vehicle surroundings image generated from the original vehicle surroundings image without images of pedestrians, etc. P are used as one set of learning data, and the image generation model is trained so that, for one corrected vehicle surroundings image, the original vehicle surroundings image that served as the basis for the corrected vehicle surroundings image is generated, as shown in FIG. 3(A). Through this learning, the image generation model generates images of pedestrians, etc. P in the corrected vehicle surroundings image. Since the images of pedestrians, etc. P match the original vehicle surroundings image, they are expected to be superimposed on images of the road surface, etc., in a realistic and natural manner. Furthermore, by using images captured under various conditions, such as at different times and in different weather conditions, as learning data, it is expected that images of pedestrians, etc. corresponding to the conditions can be superimposed on images under various conditions (for example, an image of a pedestrian holding an umbrella can be superimposed on a rainy day). As a machine learning algorithm, any form of GAN (generative adversarial network) may be used. The generated image generation model is used in the near-miss image generation section of the pseudo image generation section.

[0018] (c) Pseudo image generation unit The pseudo-image generation unit first collects images (which may be moving or still images) taken by an on-board camera such as a drive recorder while the person being diagnosed or trained is driving the vehicle, and index values ​​representing the vehicle's driving state (vehicle speed, wheel speed, acceleration, yaw rate, etc.), and stores them in a database. Next, using the stored camera images and index values ​​representing the vehicle's driving state, it is determined whether the person being diagnosed or trained has engaged in unsafe driving in any manner (unsafe driving determination unit). Here, existing unsafe driving detection technology using machine learning or the like may be used. Then, the camera images taken when unsafe driving is determined to have been engaged in may be stored in a database (unsafe driving image DB). Thereafter, in the near-miss image generation unit, the unsafe driving images stored in the unsafe driving image DB are input into the image generation model generated by the model generation unit. As a result, as already mentioned, the image generation model is configured to generate an image in which an image of a pedestrian or the like is superimposed on the image of the road surface or the like in a realistic and natural manner on an image of the vehicle surroundings in which no image of a pedestrian or the like is reflected on the road surface under various circumstances. Therefore, an image (near miss image) is generated in which an image of a pedestrian or the like that was not actually present is superimposed on an unsafe driving image, as shown in FIG. 3(B). The near miss image may be a moving image or a still image. The generated near miss image is displayed on a television monitor or a terminal display and shown to the person being diagnosed or educated.

[0019] As described above, according to the present embodiment, images in which images of pedestrians, etc. are pseudo-superimposed so as to appear realistically on unsafe driving images can be generated as images shown to a person being diagnosed or trained when diagnosing or training vehicle driving skills, using as learning data images of on-board cameras that capture images of pedestrians, etc. in various situations and images from which the images of pedestrians, etc. have been removed from those images, using an image generation model constructed according to a machine learning algorithm. In the present embodiment, when an image of a pedestrian, etc. is embedded in each unsafe driving image, the image generation model trained as described above automatically determines the position, orientation, and size of the image of the pedestrian, etc. in the vehicle surroundings image, taking into account the positional relationship and dimensional relationship between the image of the pedestrian, etc. and the image surrounding it. Therefore, it is expected that near miss images, which are images of unsafe driving in which images of pedestrians, etc. appear in a more realistic and natural manner than before, can be generated.

[0020] The above description has been made in relation to the embodiments of the present invention, but it will be apparent that many modifications and changes will be readily apparent to those skilled in the art, and the present invention is not limited to the above-described exemplary embodiments, but can be applied to various devices without departing from the concept of the present invention.

Claims

[Claim 1] A device for generating a pseudo image of a dangerous situation while driving a vehicle, comprising: an unsafe driving image acquisition means for acquiring an unsafe driving image, which is an image taken while the vehicle is traveling by an on-board camera that captures images of the surroundings of the vehicle, when it is determined that unsafe driving has been performed; a near-miss image generating means for generating a near-miss image, which is an image in which an image of a pedestrian or other obstacle is superimposed on the unsafe driving image; Including, The near-miss image generation means uses as learning data a plurality of vehicle surroundings images taken while the vehicle is traveling by an onboard camera that captures images of the vehicle's surroundings, which include images of pedestrians or other obstacles, and corrected vehicle surroundings images obtained by correcting each of the plurality of vehicle surroundings images so that images of pedestrians or other obstacles are absent, to create an image generation model constructed according to a machine learning algorithm, the image generation model being trained so that when a corrected vehicle surroundings image obtained from a vehicle surroundings image that includes an image of a pedestrian or other obstacle is input, the image generation model outputs the vehicle surroundings image that includes an image of the pedestrian or other obstacle, and when the unsafe driving image is input to the image generation model, a near-miss image is output in which the image of the pedestrian or other obstacle is superimposed on the unsafe driving image.

Citation Information

Patent Citations

  • Vehicle dangerous scene replay apparatus

    JP2012022215A