Image generation device and image generation program

The image generating device automates the placement of CG images within recognized road areas, addressing the time-consuming manual process of adding objects to captured images, thereby enhancing the efficiency of training data generation.

JP2026013181APending Publication Date: 2026-01-28TOYOTA JIDOSHA KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024113447
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2026-01-28

AI Technical Summary

Technical Problem

Existing methods require manual identification of CG image positions, making the process time-consuming when adding CG images to captured images for training data generation.

Method used

An image generating device that automatically designates initial and subsequent object drawing areas within recognized road areas in successive images, using a three-dimensional recognition model to reduce user effort by generating images with added objects.

Benefits of technology

Reduces the effort required for generating training data by automating the placement of CG images, allowing for efficient creation of images with added objects without manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026013181000001_ABST
    Figure 2026013181000001_ABST
Patent Text Reader

Abstract

To provide an image generation device and a program for reducing the labor of a user when generating image data for teacher data.SOLUTION: An image generation device 1 that generates image data representing continuous images for learning includes an acquisition unit 31 that acquires image data representing continuous captured images obtained by imaging a road from different imaging points, a road recognition unit 32 that three dimensionally recognizes a region of a road in each captured image represented by the image data, and an image generation unit 32 that generates, in an initial image that is one image of the continuous captured images: An initial region designation unit 33 that designates an initial object drawing region at an arbitrary position within a predetermined distance range that can be visually recognized by a driver in a recognized road region, a subsequent region designation unit 34 that designates a subsequent object drawing region at a position corresponding to the initial object drawing region in a subsequent image subsequent to an initial image among consecutive captured images, and an image addition unit 35 that generates data of an image in which an additional image of an arbitrary object is added to the object drawing region designated in each image of the consecutive captured images are provided.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an image generating device and an image generating program. [Background technology]

[0002] A learning device that learns a model for discriminating objects present on a road has been known (Patent Document 1). The learning device described in Patent Document 1 adds computer graphic images (CG images) of objects present on the road to captured images of the road to create images for training data. In particular, the learning device described in Patent Document 1 changes the position and size at which the CG image is added based on the amount of movement of the vehicle when adding the CG image of the object to the captured image.

[0003] It has also been known that a learning model is generated that uses a feature point map with depth information as training data to estimate the three-dimensional coordinates of an object from image data, and that image data of an object is input into the learning model to estimate the three-dimensional coordinates of the object (Patent Document 2). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2022-154193 [Patent Document 2] Patent Publication No. 2021-117130 Summary of the Invention [Problem to be solved by the invention]

[0005] In Patent Document 1, for images following the first image to which a CG image is added, the position to which the CG image should be added is automatically identified based on the amount of vehicle movement. However, when adding a CG image to an arbitrary captured image for the first time, the position to add the CG image must be manually identified. Therefore, generating an image with an object added is time-consuming.

[0006] In view of the above problems, an object of the present disclosure is to reduce the effort required by users when generating image data for training data. [Means for solving the problem]

[0007] The gist of the present disclosure is as follows.

[0008] (1) An image generating device that generates image data representing successive images for learning, an acquisition unit that acquires image data representing successive images of a road captured from different imaging points; a road recognition unit that recognizes a road area in each captured image represented by the image data in three dimensions; an initial area designation unit that designates an initial object drawing area at an arbitrary position within a predetermined distance range that is visible to the driver within a recognized road area in an initial image that is one of the consecutive captured images; a subsequent area designation unit that designates a subsequent object drawing area at a position corresponding to the initial object drawing area in a subsequent image following the initial image among the consecutive captured images; an image adding unit that generates image data in which an additional image of an arbitrary object is added to the object drawing area specified in each of the consecutive captured images. (2) The image generating device according to (1) above, wherein the initial area designation unit designates the position farthest from the imaging point within a predetermined distance range visible to the driver as the initial object drawing area. (3) The image generation device described in (1) or (2) above, wherein the image addition unit generates image data representing an additional image of an arbitrary object from random noise based on text input, and generates image data in which the additional image representing the object has been added to the object drawing area, and in the subsequent image, generates image data representing the additional image to be added in the subsequent image from the same random noise as the random noise used to generate image data representing the additional image to be added in the initial image. (4) The image generating device described in (3) above, wherein the image addition unit generates image data in the subsequent image based on text input having the same content as the text input used to generate image data representing the additional image to be added in the initial image, except for text regarding the appearance of objects in the additional image to be added. (5) An image generation program for generating image data representing successive images for learning, acquiring image data representing successive images of a road surface taken from different imaging locations; three-dimensionally recognizing a road area in each captured image represented by the image data; In an initial image, which is one of the consecutive captured images, designating an initial object drawing area at an arbitrary position within a predetermined distance range that is visible to the driver within the recognized road area; specifying a subsequent object drawing area at a position corresponding to the initial object drawing area in a subsequent image following the initial image among the consecutive captured images; generating image data in which an additional image of an arbitrary object is added to the object drawing area specified in each of the consecutive captured images; An image generation program that causes a computer to execute the above. [Effects of the Invention]

[0009] According to the present disclosure, the user's effort in generating image data for training data can be reduced. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating a configuration of an image generating device according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating a vehicle that transmits image data to an image generation device for use in the image generation device. [Figure 3] FIG. 3 is a diagram showing a state in which roads are recognized by the road recognition unit in one captured image. [Figure 4] FIG. 4 is a diagram similar to FIG. 3, showing the visible region specified by the initial region specifying unit in one captured image. [Figure 5] FIG. 5 is a diagram showing a schematic representation of a mask that can be applied to an arbitrary image. [Figure 6] FIG. 6 shows an image to which additional images generated using an image generation AI model have been added. [Figure 7] FIG. 7 is a flowchart showing the flow of the image generation process. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, the embodiments will be described in detail with reference to the drawings. In the following description, like components are designated by like reference numerals.

[0012] The configuration of an image generating device 1 will be described with reference to Figures 1 and 2. Figure 1 is a diagram schematically illustrating the configuration of an image generating device 1 according to an embodiment. The image generating device 1 generates image data representing a series of images (i.e., a video) to be used as training data in training a machine learning model.

[0013] In this embodiment, the machine learning model is a model for identifying objects (obstacles) on the road in an image of the road ahead captured by an exterior camera attached to the vehicle. Therefore, when image data captured by the exterior camera is input, if the image represented by the data includes an object (e.g., cargo, tires, cardboard boxes, rocks, branches, etc.) that exists on the road as an obstacle, the machine learning model outputs the position of the object, etc.

[0014] To train such a machine learning model, image data of consecutive images of the road ahead of the vehicle captured by an exterior camera is required as training data. However, it is difficult to actually capture and prepare a large number of consecutive images that actually capture such obstacles. Therefore, image generation device 1 generates images by adding additional images of any obstacles to consecutive images captured by the exterior camera of a moving vehicle. This makes it possible to generate an image in which an obstacle is present on the road even if the obstacle is not captured in the images captured by the exterior camera, thereby reducing the burden of generating training data used in the machine learning model described above.

[0015] 2 is a diagram that schematically shows a vehicle 100 that transmits image data used in image generation device 1 to image generation device 1. As shown in FIG. 2, vehicle 100 is equipped with an exterior camera 101 that captures an image ahead of vehicle 100. Image generation device 1 may be mounted on vehicle 100, or may be formed in a server that can communicate with vehicle 100.

[0016] In this embodiment, the exterior camera 101 is disposed inside the windshield of the vehicle 100 and captures images of the area in front of the vehicle 100. Therefore, the exterior camera 101 captures images of the road in front of the vehicle 100 while the vehicle 100 is traveling. The exterior camera 101 captures images of the area in front of the vehicle 100 at predetermined imaging intervals and generates image data of continuous images showing the area in front of the vehicle 100.

[0017] 1, the image generating device 1 includes a communication interface 10, a storage unit 20, and a processor 30. The communication interface 10, the storage unit 20, and the processor 30 may be separate circuits, or may be configured as a single integrated circuit.

[0018] The communication interface 10 is an interface circuit for connecting the image generation device 1 to a device external to the image generation device 1. The image generation device 1 transmits and receives data to and from the external device via the communication interface 10. The external device may include, for example, an exterior camera 101 of any vehicle 100 or a vehicle storage device (not shown) that stores data of images captured by the exterior camera 101. The external device may also include a learning device that trains a machine learning model. In addition, the external device may include a user input device (e.g., a keyboard, a mouse, etc.) and a user output device (e.g., a display, a speaker, etc.). In this embodiment, the communication interface 10 receives data of images of the front of the vehicle 100 while the vehicle 100 is traveling, captured by the exterior camera 101, from the exterior camera 101 of any vehicle 100 or the vehicle storage device, and stores the data in the storage unit 20. The communication interface 10 also transmits data of images for learning generated by the image generation device 1 to the learning device.

[0019] The storage unit 20 is a non-transitory storage medium for storing data. The storage unit 20 includes, for example, at least one of a volatile semiconductor memory, a non-volatile semiconductor memory, a hard disk drive (HDD), and a solid-state drive (SSD). The storage unit 20 stores computer programs executed by the processor 30, particularly an image generation program for executing an image generation process. The storage unit 20 also stores data used in the computer programs executed by the processor 30, such as data of images of the area ahead of the vehicle 100 received from the outside via the communication interface 10. In addition, the storage unit 20 stores data of images generated by the processor 30.

[0020] The processor 30 has one or more central processing units (CPUs) and their peripheral circuits. The processor 30 may further have other arithmetic circuits such as a logic unit or a numerical operation unit. The processor 30 executes a computer program stored in the storage unit 20. In particular, in this embodiment, the processor 30 executes an image generation program stored in the storage unit 20.

[0021] 1, the processor 30 includes an acquisition unit 31, a road recognition unit 32, an initial area designation unit 33, a subsequent area designation unit 34, and an image addition unit 35. Each of these units included in the processor 30 is, for example, a functional module realized by a computer program running on the processor 30. Alternatively, each of the units included in the processor 30 may be implemented in the image generation device 1 as an independent integrated circuit, microprocessor, or firmware.

[0022] The acquisition unit 31 acquires image data representing successive captured images of a road taken from different imaging points. In this embodiment, image data of successive captured images taken by the exterior camera 101 of the vehicle 100 while it is traveling is stored in the storage unit 20 via the communication interface 10. Therefore, the acquisition unit 31 acquires the image data of such successive captured images stored in the storage unit 20 from the storage unit 20. Since such successive captured image data was captured while the vehicle 100 was traveling, it represents successive captured images of a road taken from imaging points that are spaced a short distance apart.

[0023] The road recognition unit 32 three-dimensionally recognizes the road area in each captured image represented by the image data acquired by the acquisition unit 31. In this embodiment, the consecutive captured images represented by the image data are captured from imaging locations that are slightly different from each other. As a result, in this embodiment, the road recognition unit 32 three-dimensionally recognizes the road area based on the image data acquired by the acquisition unit 31 without using distance data (e.g., detection data from a distance measurement sensor such as LiDAR or millimeter-wave radar) between the vehicle 100 and an object ahead of the vehicle 100 at the time the image data was captured. That is, in this embodiment, the road recognition unit 32 three-dimensionally recognizes the road area in each captured image based only on the image data representing the consecutive captured images acquired by the acquisition unit 31. In particular, in this embodiment, the road recognition unit 32 three-dimensionally recognizes the road area in each captured image from the image data representing the consecutive captured images using a three-dimensional recognition model such as SfM (Structure from Motion).

[0024] Fig. 3 is a diagram showing a state in which a road is recognized by the road recognition unit 32 in one captured image. In Fig. 3, the road recognized by the road recognition unit 32 is represented by a point cloud PC including a plurality of points P that represent a three-dimensional position relative to the vehicle 100 (particularly, the exterior camera 101 of the vehicle 100). That is, in this embodiment, the road recognition unit 32 calculates the point cloud PC that represents the three-dimensional position of the road. Therefore, the road recognition unit 32 three-dimensionally recognizes the area of ​​the road in each captured image represented by the image data acquired by the acquisition unit 31, and outputs data of the point cloud PC, which is a collection of a plurality of points P that represent the three-dimensional position of the road.

[0025] In this embodiment, a three-dimensional recognition model such as SfM is used to recognize roads in a captured image three-dimensionally based on image data representing successive captured images. As a result, detection data from a distance measurement sensor such as a LiDAR or millimeter-wave radar is not required to recognize roads in a captured image. Therefore, roads can be recognized based on a small amount of data.

[0026] In this embodiment, the road recognition unit 32 uses SfM to generate point cloud PC data that three-dimensionally represents a road from image data that represents consecutive captured images, and recognizes the road three-dimensionally using this point cloud PC. However, the road recognition unit 32 may use any three-dimensional recognition model other than SfM as long as it can three-dimensionally recognize the road area in the captured images based on image data that represents consecutive captured images.

[0027] The initial area designation unit 33 designates an initial object drawing area at any position within a predetermined distance range that can be seen by the driver in the recognized road area in an initial image, which is one of the consecutive captured images.

[0028] The initial region designation unit 33 first sets one image from the consecutive captured images included in the image data acquired by the acquisition unit 31 as the initial image (hereinafter, the time when this initial image appears among the consecutive captured images is defined as time t=0). The initial image setting by the initial region designation unit 33 may be performed based on input from the user via an input device. In this case, the user specifies an image to be used as the initial image via the input device, and the initial region designation unit 33 sets the image specified by the user as the initial image. Alternatively, the initial region designation unit 33 may automatically set the initial image from the consecutive captured images based on parameters such as the frequency of appearance of obstacles set by the user via the input device.

[0029] In addition, when the initial image is set, the initial area designation unit 33 identifies, as the visible area, the area within the initial image that is recognized by the road recognition unit 32 as having a road and that is within the distance range that can be seen by the driver.

[0030] FIG. 4 is a diagram similar to FIG. 3, showing the visible area specified by the initial area designation unit 33 in one captured image. In particular, in the example shown in FIG. 4, the area represented by the point cloud PC' is specified as the visible area. Here, the point cloud PC' shown in FIG. 4 does not include measurement points P located in an area away from the image capture point (i.e., the exterior camera 101 of the vehicle 100) among the point cloud PC shown in FIG. 3. Therefore, the point cloud PC' shown in FIG. 4 is a point cloud from which points P representing positions farther than a predetermined visibility limit distance have been removed from the point cloud PC representing the road in three dimensions. Therefore, in this embodiment, the visible area specified by the initial area designation unit 33 is represented by the point cloud PC' located at a distance from the vehicle 100 equal to or less than the visibility limit distance among the point cloud PC including multiple points P representing the three-dimensional positions of the road.

[0031] In this embodiment, the initial area designation unit 33 designates an initial object drawing area at an arbitrary position within the visible area identified in this manner. That is, the initial area designation unit 33 designates the object drawing area within the area represented by the point cloud PC' in FIG. 4.

[0032] In particular, in this embodiment, the initial area designation unit 33 designates the area of ​​the visible area that is the farthest from the imaging point in the traveling direction of the vehicle 100 as the initial object drawing area. Therefore, the area in the zone that is the farthest from the vehicle 100 within the visibility limit distance that the driver can see is designated as the initial object drawing area. In the example shown in FIG. 4, the initial object drawing area is designated at an arbitrary position in zone I that is the farthest from the vehicle 100 in the traveling direction of the vehicle 100, out of the area represented by the point cloud PC'.

[0033] Here, the object drawing area is an area in which an additional image of an object to be added is drawn. In this embodiment, the object drawing area in the initial image is automatically designated by the initial area designation unit 33. This eliminates the need for the user to designate an area in which to add an object, thereby reducing the user's effort when generating image data for training data. Also, in this embodiment, the initial area designation unit 33 designates the farthest area of ​​the visible area as the initial object drawing area. As a result, the object to be added is drawn at the farthest position visible to the driver, and a natural image is generated without the added object suddenly appearing.

[0034] The subsequent area designation unit 34 designates a subsequent object drawing area in a subsequent image following the initial image among the consecutive captured images at a position corresponding to the object drawing area in the initial image. As described above, the road recognition unit 32 recognizes the road area three-dimensionally, and therefore also recognizes the positional relationship of the road between different images. Therefore, when the road recognition unit 32 recognizes the road area, it recognizes a point on the road in another captured image that corresponds to an arbitrary point on the road in an arbitrary captured image. The subsequent area designation unit 34 designates the object drawing area in the subsequent screen based on the positional relationship of corresponding points in different captured images recognized in this way.

[0035] Specifically, the subsequent area designation unit 34 designates an object drawing area in a captured image subsequent to the initial image at a position corresponding to the object drawing area in the initial image. Then, when an object drawing area is designated in a captured image subsequent to the initial image, the subsequent area designation unit 34 designates an object drawing area in a position corresponding to this object drawing area in the subsequent captured image, and repeats this operation. As a result, object drawing areas are designated at corresponding positions in a plurality of consecutive captured images.

[0036] Further, the subsequent area designation unit 34 designates the object drawing area so that the size of the object drawing area changes depending on the three-dimensional distance from the imaging point to the object drawing area. Since successive captured images are basically images captured by the forward-moving vehicle 100, the distance from the imaging point to the object drawing area becomes shorter in later images. Therefore, the subsequent area designation unit 34 designates the object drawing area so that the object drawing area becomes larger in later images.

[0037] In this embodiment, an object drawing area in a subsequent image is specified based on the positional relationship of the road between different images recognized by the road recognition unit 32. As a result, travel data of the vehicle 100 (e.g., data such as the speed and acceleration of the vehicle 100, the steering angle of the vehicle 100, etc.) is not required to specify the object drawing area in the subsequent image. Therefore, the object drawing area in the subsequent image can be appropriately specified based on less data.

[0038] The image adding unit 35 generates image data by adding an additional image of an arbitrary object to a designated object drawing area in each of the consecutive captured images. In this embodiment, the image adding unit 35 generates image data by adding an additional image of the same object (obstacle) to each of the consecutive captured images.

[0039] In this embodiment, when a user inputs text via an input device, the image adding unit 35 generates image data representing an additional image of an arbitrary object according to the text input from random noise based on the text input, and also generates image data in which the additional image represented by this image data is added to the object drawing area. For example, when the user inputs the text "cardboard box," the image adding unit 35 generates image data of an additional image representing "cardboard box" according to the text input. Then, the image adding unit 35 generates image data of the additional image in which this "cardboard box" is added to the object drawing area of ​​each captured image.

[0040] Furthermore, the added image generated by the image adding unit 35 changes depending on the random noise initially applied. Therefore, even if the same text is input by the user, if the initial random noise is different, a different image will be generated according to the text input. For example, if the text "cardboard box" is input, and the initial random noise is different, image data representing an image of a cardboard box with a different shape, color, printing on the surface, etc. will be generated. On the other hand, if the text "cardboard box" is input, and the initial random noise is the same, image data representing an image of a cardboard box with the same shape, color, and printing will be generated.

[0041] In this embodiment, the image adding unit 35 generates image data using an image generation AI model such as Stable Diffusion. In particular, with Stable Diffusion, when generating an image, text about the image to be generated is input by the user, and data representing the image is generated according to the text. Also, with Stable Diffusion, when generating an image, random noise is first input, and data representing the image according to the input text is generated based on the random noise.

[0042] Specifically, the image adding unit 35 first generates image data by adding an additional image of an arbitrary object to the initial image at time t=0. The image adding unit 35 generates a mask by painting the object drawing area specified by the initial area specifying unit 33 with a single color (for example, white). In addition, the image adding unit 35 overlays the mask thus generated on the initial image at time t=0. As a result, an image is generated in which part of the initial image is painted white with the mask.

[0043] 5A and 5B are diagrams schematically illustrating masks that are added to an arbitrary image. Fig. 5A shows the mask M that is added to the initial image at time t = 0. As shown in Fig. 5A, the mask M that is added to the initial image at time t = 0 is formed in the area specified by the initial area specification unit 33, i.e., the area of ​​the visible area that is farthest from the imaging point.

[0044] In addition, the image addition unit 35 adds an additional image generated using an image generation AI model based on text input by the user and arbitrary random noise to the area where the mask M is provided for the initial image at time t=0.

[0045] FIG. 6 is a diagram showing an image to which an additional image generated using an image generation AI model has been added. FIG. 6(A) shows an image in which a generated additional image (image surrounded by a square in the figure) has been added to the area of ​​mask M shown in FIG. 5(A) of the initial image at time t=0. In particular, in FIG. 6(A), an additional image of cardboard has been added. As a result, in this embodiment, a user can add an additional image to the initial image according to the text input simply by inputting text related to the image to be added.

[0046] Next, the image adding unit 35 generates image data by adding an additional image of an arbitrary object to the subsequent image at each time t=n (n is a value greater than 0) after time t=0. For each subsequent image, the image adding unit 35 generates a mask by painting the object drawing area specified by the subsequent area specifying unit 34 with a single color (for example, white). In addition, the image adding unit 35 overlays the mask thus generated on the subsequent image at time t=n. As a result, an image is generated in which a part of the subsequent image at time t=n is painted white with the mask.

[0047] Fig. 5(B) shows the mask M added to the subsequent image at time t = N. As shown in Fig. 5(B), the mask M added to the subsequent image at time t = n is formed in the area specified by the subsequent area specification unit 34, i.e., the area corresponding to the object drawing area in the initial image.

[0048] In addition, the image addition unit 35 adds an additional image generated using an image generation AI model based on text input by the user and arbitrary random noise to the area where the mask M is provided for the subsequent image at time t=n.

[0049] FIG. 6(B) shows an image in which a generated image (an image surrounded by a square in the figure) has been added to the region of the mask M shown in FIG. 5(B) of the subsequent image at time t=n. In this embodiment, when generating image data representing the additional image to be added to the subsequent image, the same random noise as the random noise used to generate the image data representing the additional image to be added to the initial image is used. In addition, in this embodiment, when generating image data representing the image to be added to the subsequent image, text input having the same content as the text input used to generate the image data representing the additional image to be added to the initial image is used. As a result, in FIG. 6(B), an additional image of a cardboard box similar to the additional image of a cardboard box in FIG. 6(A) has been added. As a result, in this embodiment, a user can add generated images according to this text input to subsequent images by simply inputting text related to an image to be added to a series of captured images once.

[0050] In the above embodiment, when generating image data representing additional images to be added to subsequent images, the same text input is used as the text input used to generate image data representing additional images to be added to the initial image. However, the text related to the appearance of the object in the image data representing the additional images to be added may be different for each subsequent image. For example, for a subsequent image following an initial image, text indicating that the orientation of the object has changed by an arbitrary angle relative to the initial image may be added to the text used to generate the image. This allows image data including a more appropriate image of the object to be generated for the subsequent image. In any case, it can be said that the image adding unit 35 generates image data for the subsequent images based on text input with the same content as the text input used to generate image data representing the additional images to be added to the initial image, except for the text related to the appearance of the object in the image data representing the additional images.

[0051] Next, the flow of image generation processing for creating successive images for learning will be described with reference to Fig. 7. Fig. 7 is a flowchart showing the flow of image generation processing. The image generation processing shown in Fig. 7 is executed by processor 30.

[0052] When the image generation process is started, first, the acquisition unit 31 acquires, from the storage unit 20, image data representing successive captured images of a road taken while the vehicle 100 is traveling (step S11). Therefore, the acquisition unit 31 acquires, for example, a video that has been captured while the vehicle 100 is traveling, and then transmitted from the vehicle 100 to the image generation device 1 and stored in the storage unit 20.

[0053] Next, the road recognition unit 32 uses a three-dimensional recognition model (e.g., SfM) to three-dimensionally recognize the road area in each captured image represented by the image data acquired by the acquisition unit 31 (step S12). In this embodiment, the road recognition unit 32 three-dimensionally recognizes the road area for all captured images included in the acquired image data. However, the road recognition unit 32 may three-dimensionally recognize the road area only for some consecutive captured images included in the acquired image data.

[0054] Next, the initial area designation unit 33 designates an initial object drawing area at a predetermined position recognized by the road recognition unit 32 in one image (initial image) of the consecutive captured images acquired by the acquisition unit 31 (step S13). The initial area designation unit 33 may automatically designate the initial image, or may automatically designate the initial object drawing area.

[0055] Next, the subsequent area designation unit 34 designates a subsequent object drawing area at a position corresponding to the initial object drawing area in a subsequent image following the initial image among the consecutive captured images acquired by the acquisition unit 31 (step S14). The subsequent area designation unit 34 automatically designates the subsequent object drawing area.

[0056] Next, the image adding unit 35 uses the image generation AI model to generate image data of an image in which an additional image of an arbitrary object has been added to the object drawing area specified by the initial area specifying unit 33 and the subsequent area specifying unit 34 (step S15). The image adding unit 35 generates image data of an image in which an additional image of the same object has been added to all captured images in which an object drawing area has been specified. As a result, consecutive images in which an arbitrary object (obstacle) exists on the road are generated from consecutive captured images of a road in which no obstacle exists on the road.

[0057] Although preferred embodiments according to the present disclosure have been described above, the present disclosure is not limited to these embodiments, and various modifications and changes can be made within the scope of the claims. [Explanation of symbols]

[0058] 1. Image generation device 20 Memory section 30 processors 100 vehicles 101 Exterior camera

Claims

1. An image generating device that generates image data representing successive images for learning, an acquisition unit that acquires image data representing successive images of a road captured from different imaging points; a road recognition unit that recognizes a road area in each captured image represented by the image data in three dimensions; an initial area designation unit that designates an initial object drawing area at an arbitrary position within a predetermined distance range that is visible to the driver within a recognized road area in an initial image that is one of the consecutive captured images; a subsequent area designation unit that designates a subsequent object drawing area at a position corresponding to the initial object drawing area in a subsequent image following the initial image among the consecutive captured images; an image adding unit that generates image data in which an additional image of an arbitrary object is added to the object drawing area specified in each of the consecutive captured images.

2. The image generating device according to claim 1 , wherein the initial area designation unit designates, as the initial object rendering area, an area that is farthest from an imaging point within a predetermined distance range that can be seen by a driver.

3. 3. The image generating device according to claim 1, wherein the image addition unit generates image data representing an additional image of an arbitrary object from random noise based on text input, and generates image data in which the additional image representing the object has been added to the object drawing area, and in the subsequent image, generates image data representing the additional image to be added in the subsequent image from the same random noise as the random noise used to generate image data representing the additional image to be added in the initial image.

4. 4. The image generating device of claim 3, wherein the image adding unit generates image data for the subsequent image based on text input having the same content as the text input used to generate image data representing the additional image to be added for the initial image, except for text related to how an object in the additional image to be added appears.

5. An image generation program for generating image data representing successive images for learning, acquiring image data representing successive images of a road surface taken from different imaging locations; three-dimensionally recognizing a road area in each captured image represented by the image data; In an initial image, which is one of the consecutive captured images, designating an initial object drawing area at an arbitrary position within a predetermined distance range that is visible to the driver within the recognized road area; specifying a subsequent object drawing area at a position corresponding to the initial object drawing area in a subsequent image following the initial image among the consecutive captured images; generating image data in which an additional image of an arbitrary object is added to the object drawing area specified in each of the consecutive captured images; An image generation program that causes a computer to execute the above.

Citation Information

Patent Citations

  • Device and method for estimating three-dimensional position

    JP2021117130A

  • Learning device, learning method, program, and object detection device

    JP2022154193A