Learning image generation device, learning image generation method, and program

The training image generation device generates training fisheye images with virtual road markings by converting and adding virtual paint, addressing the need for actual image capture and reducing data preparation load in vehicle connection state detection.

JP2025185434APending Publication Date: 2025-12-22TOYOTA JIDOSHA KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024093676
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-10
Publication Date
2025-12-22

AI Technical Summary

Technical Problem

Existing technologies for detecting the connection state between towing and towed vehicles using image processing do not utilize learning models and require capturing fisheye images with road paint, which increases data preparation load.

Method used

A training image generation device and method that generates training fisheye images by converting fisheye original images to planar images, adding virtual road paint, and then reversing the conversion to create images with virtual road markings, allowing training without actual capture.

Benefits of technology

Enables training of models using fisheye images with road paint without physically capturing such images, reducing data preparation load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025185434000001_ABST
    Figure 2025185434000001_ABST
Patent Text Reader

Abstract

To make it possible, without requiring actually photographing a fisheye image containing road surface paint, to learn a model by using the fisheye image as learning data.SOLUTION: A learning image generation device 1 which generates a learning fisheye image used for learning of a model used for estimation of a condition of a trailer on the basis of a fisheye image photographed by a camera mounted on a vehicle towing the trailer through a towing bar includes: an acquisition unit 3A which acquires a learning fisheye original image photographed by a learning camera mounted on a learning vehicle towing the learning trailer through a learning towing bar; an image conversion unit 3B which generates a plane normal conversion image by applying plane normal image conversion to the learning fish eye original image; an adding unit 3C which adds virtual road surface paint to the plane normal conversion image; and a learning image generation unit 3D which generates a learning fisheye image by applying a conversion opposite to the plane normal conversion to the plane normal conversion (plane image with the virtual road surface paint) after addition of the virtual paint.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a learning image generation device, a learning image generation method, and a program. [Background technology]

[0002] Patent Document 1 describes a towing vehicle in which a camera unit with a wide-angle or fisheye lens is mounted on the wall below the rear hatch. Patent Document 1 also describes that the image data captured by the camera unit can be used to detect the connection state between the towing vehicle and the towed vehicle (for example, the connection angle, whether they are connected, etc.). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2018-176788 Summary of the Invention [Problem to be solved by the invention]

[0004] Incidentally, although Patent Document 1 describes that the connection state between the towing vehicle and the towed vehicle is detected by image processing, it does not state whether a model that requires learning is used in that image processing. If a model is used to detect the connection state between the towing vehicle and the towed vehicle, it is necessary to suppress the increase in the load of preparing learning data used to train that model.

[0005] In view of the above, the present disclosure aims to provide a training image generation device, a training image generation method, and a program that enable training of a model using fisheye images including road paint as training data, without the need to actually capture fisheye images including road paint using a training camera. is. [Means for solving the problem]

[0006] (1) One aspect of the present disclosure is a training image generation device that generates training fisheye images as training data used to train a model used to estimate the state of a trailer based on fisheye images taken by a camera mounted on a vehicle towing a trailer via a tow bar. The training image generation device includes: an acquisition unit that acquires training fisheye original images taken by a training camera mounted on a training vehicle towing a training trailer via a training tow bar; an image conversion unit that generates a planar orthogonalized converted image by performing a planar orthogonalization transformation, which is a transformation from a fisheye image to a planar image, on the training fisheye original image acquired by the acquisition unit; an addition unit that adds virtual road paint to the planar orthogonalized converted image generated by the image conversion unit; and a training image generation unit that generates the training fisheye image by performing a transformation inverse to the planar orthogonalization transformation performed by the image conversion unit on a virtual road paint-added planar image, which is the planar orthogonalized converted image after the virtual road paint has been added by the addition unit.

[0007] (2) One aspect of the present disclosure includes a training image generation step in which a training image generation device generates training fisheye images as training data used to train a model used to estimate the state of a trailer based on fisheye images taken by a camera mounted on a vehicle towing a trailer via a tow bar; an acquisition step in which the training image generation device acquires training fisheye original images taken by a training camera mounted on a training vehicle towing a training trailer via a training tow bar; and a fisheye image-to-planar image conversion step in which the training image generation device converts the training fisheye original images acquired in the acquisition step. This learning image generation method includes an image transformation step of generating a planar-normalized transformed image by executing a planar normalization transformation, which is a transformation, and an additional step of the learning image generation device adding virtual road paint to the planar normalized transformed image generated in the image transformation step, wherein in the learning image generation step, the learning fisheye image is generated by executing a transformation that is the reverse of the planar normalization transformation executed in the image transformation step on a planar image with virtual road paint, which is the planar normalized transformed image after the virtual road paint has been added in the additional step.

[0008] (3) One aspect of the present disclosure is a program for causing a processor to execute the following steps: a training image generation step for generating training fisheye images as training data used in training a model used to estimate the state of a trailer based on fisheye images taken by a camera mounted on a vehicle towing a trailer via a tow bar; an acquisition step for acquiring training fisheye original images taken by a training camera mounted on a training vehicle towing a training trailer via a training tow bar; an image transformation step for generating planar orthogonal transformed images by performing a planar orthogonal transformation, which is a transformation from a fisheye image to a planar image, on the training fisheye original images acquired in the acquisition step; and an addition step for adding virtual road paint to the planar orthogonal transformed images generated in the image transformation step, wherein in the training image generation step, the training fisheye images are generated by performing an inverse transformation of the planar orthogonal transformation performed in the image transformation step on a virtual road paint-added planar image, which is the planar orthogonal transformed image after the virtual road paint has been added in the addition step. [Effects of the Invention]

[0009] According to the present disclosure, it is possible to train a model using fisheye images containing road paint as training data, without the need to actually capture fisheye images containing road paint using a training camera. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a diagram illustrating an example of a learning image generating device 1 according to a first embodiment. [Figure 2] FIG. 1 is a diagram showing an example of a learning vehicle L1 equipped with a learning camera L11. [Figure 3] 10 is a diagram showing an example of a planar normalized converted image IM2 etc. generated by an image conversion unit 3B. FIG. [Figure 4] FIG. 10 is a diagram showing an example of a training fish-eye image IM4 generated by a training image generating unit 3D. [Figure 5] This figure explains the state of trailer R2 (hitch angle Φ of trailer R2) estimated using a model trained using training fisheye image IM4 including virtual road surface paint VP (demarcation line) generated by the training image generation unit 3D of the training image generation device 1 as training data. [Figure 6] 1 is a flowchart illustrating an example of processing executed in the learning image generating device 1 of the first embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of a training image generating device, a training image generating method, and a program according to the present disclosure will be described with reference to the drawings.

[0012] First Embodiment FIG. 1 is a diagram showing an example of a learning image generating device 1 according to the first embodiment. 1, the training image generating device 1 is configured by a microcomputer having a communication interface (I / F) 11, a memory 12, and a processor 13. The communication interface 11 has an interface circuit for connecting the training image generating device 1 to a device external to the training image generating device 1 (for example, a storage device (not shown) that stores training fisheye original images IM1 (see FIG. 2(B)) captured by a training camera L11 (see FIG. 2(A))).

[0013] Fig. 2 shows an example of a learning vehicle L1 equipped with a learning camera L11, etc. In detail, Fig. 2(A) shows an example of a learning vehicle L1 equipped with a learning camera L11, etc., and Fig. 2(B) shows an example of a learning fisheye original image IM1 captured by the learning camera L11. In the example shown in FIG. 2, a learning camera L11 is disposed at the rear end L1R of the learning vehicle L1. The learning camera L11 captures images of the rear of the learning vehicle L1 (the right side in FIG. 2(A)). As shown in FIG. 2(A), the learning vehicle L1 tows a learning trailer L2 via a learning tow bar L3. The learning trailer L2 is connected to the learning vehicle L1 so as to be rotatable around a hitch ball (not shown). As shown in FIG. 2(B), the learning fisheye original image IM1 includes a portion of the learning vehicle L1, the learning trailer L2, and the learning tow bar L3. On the other hand, in the example shown in Figure 2, as shown in Figure 2(A), no marking lines are painted as road surface markings on the road surface on which the training vehicle L1 and training trailer L2 are traveling, and as shown in Figure 2(B), no marking lines as road surface markings are included in the training fisheye original image IM1.

[0014] 1, the memory 12 stores programs and various data used in the processing executed by the processor 13. The processor 13 has a function as an acquisition unit 3A, a function as an image conversion unit 3B, a function as an addition unit 3C, and a function as a learning image generation unit 3D. The acquisition unit 3A acquires a training fisheye original image IM1 captured by the training camera L11. In detail, the acquisition unit 3A acquires a training fisheye original image IM1 that does not include road markings such as lane markings, as shown in the example of FIG. 2(B). The image conversion unit 3B generates a planar normal image conversion image IM2 (see Figure 3(A)) by performing a planar normal image conversion, which is a conversion from a fisheye image to a planar image, on the learning fisheye original image IM1 acquired by the acquisition unit 3A. The addition unit 3C generates a planar image IM3 with virtual road surface paint (see Figure 3(C)) by adding virtual road surface paint VP (demarcation line) (see Figure 3(B)) to the planar normal image converted image IM2 generated by the image conversion unit 3B.

[0015] Fig. 3 shows an example of the planar normal image converted image IM2 etc. generated by the image conversion unit 3B. In detail, Fig. 3(A) shows an example of the planar normal image converted image IM2 generated by the image conversion unit 3B, Fig. 3(B) shows an example of virtual road paint VP (marking lines) added by the addition unit 3C to the planar normal image converted image IM2 shown in Fig. 3(A), and Fig. 3(C) shows an example of a planar image IM3 with virtual road paint, which is the planar normal image converted image after the virtual road paint VP (marking lines) has been added by the addition unit 3C. In the example shown in Figure 3, the addition unit 3C combines the planar normalized converted image IM2 shown in Figure 3(A) with the virtual road paint VP (demarcation line) shown in Figure 3(B) to generate a planar image IM3 with virtual road paint shown in Figure 3(C).

[0016] In the example shown in Figure 1, the training image generation unit 3D performs a transformation on the planar image IM3 with virtual road paint that is the inverse of the planar image normalization transformation performed by the image conversion unit 3B, thereby generating a training fisheye image IM4 (see Figure 4) that includes virtual road paint VP (demarcation lines).

[0017] FIG. 4 is a diagram showing an example of a training fish-eye image IM4 generated by the training image generating unit 3D. In the example shown in Figure 4, the training image generation unit 3D performs a transformation that is the inverse of the planar normalization transformation on the planar image IM3 with virtual road paint shown in Figure 3(C), generating a training fisheye image IM4 that includes the virtual road paint VP (demarcation line).

[0018] 1 to 4, the training vehicle L1 and the training trailer L2 do not need to actually travel on a road surface on which road surface paint (marking lines) is painted, and a training fisheye image IM4 including marking lines (virtual road surface paint VP) can be obtained in the same way as if the training vehicle L1 and the training trailer L2 actually traveled on a road surface on which road surface paint (marking lines) is painted. In other words, the example shown in FIGS. 1 to 4 makes it possible to train a model using a fisheye image including road surface paint (marking lines) as training data, without actually having to photograph a fisheye image including road surface paint (marking lines) with the training camera L11.

[0019] In one application example of the training image generation device 1 of the first embodiment, a training fisheye image IM4 including virtual road surface paint VP (demarcation line) generated by the training image generation unit 3D of the training image generation device 1 is used as training data for training a model used to estimate the state of trailer R2 (see Figure 5) (e.g., the hitch angle Φ of trailer R2) based on a fisheye image taken by a camera R11 (see Figure 5) mounted on a vehicle R1 (see Figure 5) towing trailer R2 (see Figure 5) via a tow bar R3 (see Figure 5).

[0020] Figure 5 is a diagram for explaining the state of trailer R2 (hitch angle Φ of trailer R2) estimated using a model trained using a training fisheye image IM4 including virtual road surface paint VP (demarcation line) generated by the training image generation unit 3D of the training image generation device 1 as training data. In the example shown in FIG. 5, a camera R11 is disposed at the rear end R1R of a vehicle R1. The camera R11 photographs the rear of the vehicle R1 (the right side in FIG. 5). The vehicle R1 tows a trailer R2 via a tow bar R3. The trailer R2 is connected to the vehicle R1 so as to be rotatable around a hitch ball (not shown). The fisheye image captured by the camera R11 includes a portion of the vehicle R1, the trailer R2, and the tow bar R3. In one application example of the training image generation device 1 of the first embodiment (the example shown in Figures 1 to 5), the hitch angle Φ of trailer R2 is estimated based on a fisheye image (an image including trailer R2, etc.) taken by camera R11 (see Figure 5) by using a model obtained by performing training using training data, which is a data set of a training fisheye image IM4 including virtual road surface paint VP (demarcation line) generated by the training image generation unit 3D of the training image generation device 1, and a label indicating the hitch angle θ (see Figure 2(A)) of the training trailer L2 at the time of capturing the training fisheye original image IM1 (see Figure 2(B)) corresponding to the training fisheye image IM4. In detail, for training the model, a dataset of a training fisheye image IM4 including virtual road surface paint VP (demarcation line) and a label indicating the hitch angle θ (see Figure 2(A)) of the training trailer L2 at the time when the training fisheye original image IM1 corresponding to the training fisheye image IM4 was captured is used as training data, and a dataset of a training fisheye original image IM1 not including the virtual road surface paint VP (demarcation line) and a label indicating the hitch angle θ (see Figure 2(A)) of the training trailer L2 at the time when the training fisheye original image IM1 was captured is also used as training data.

[0021] In one application example of the training image generation device 1 of the first embodiment (the example shown in Figures 1 to 5), a training fisheye image IM4 including virtual road paint VP (marking line) generated by the training image generation unit 3D of the training image generation device 1 is used to train a model used to estimate the hitch angle Φ of the trailer R2, but in other application examples, the training fisheye image IM4 including virtual road paint VP (marking line) may be used to train a model used to estimate the state of the trailer R2 other than the hitch angle Φ of the trailer R2, such as estimating the position (coordinates) of the ends (left end and right end) of the trailer R2.

[0022] FIG. 6 is a flowchart illustrating an example of processing executed in the learning image generating device 1 of the first embodiment. In the example shown in FIG. 6, in step S10, the acquisition unit 3A acquires a learning fish-eye original image IM1 that does not include road paint (demarcation lines) and is captured by the learning camera L11. In step S11, the image conversion unit 3B performs planar image-normalization conversion on the learning fish-eye original image IM1 acquired in step S10, thereby generating a planar image-normalized converted image IM2. In step S12, the adding unit 3C generates a virtual road paint-added planar image IM3 by adding virtual road paint VP (demarcation lines) to the planar normal image IM2 generated in step S11. In step S13, the training image generation unit 3D performs a transformation on the planar image IM3 with virtual road paint that is the inverse of the planar normalization transformation performed in step S11, thereby generating a training fisheye image IM4 including the virtual road paint VP (demarcation line). It is possible to learn the following.

[0023] Second Embodiment As described above, in the first embodiment (the example shown in FIGS. 1 to 6), the acquisition unit 3A acquires a learning fish-eye original image IM1 that does not include road surface paint such as marking lines. On the other hand, in the second embodiment, the acquisition unit 3A acquires a learning fisheye original image IM1 that does not include road markings as road surface paint (more specifically, road markings such as maximum speed, no turn, left turn arrow, straight arrow, right turn arrow, etc.).

[0024] As described above, in the first embodiment (the example shown in Figures 1 to 6), the addition unit 3C generates a planar image IM3 with virtual road paint by adding virtual road paint VP (demarcation lines) to the planar normal image conversion image IM2 generated by the image conversion unit 3B. On the other hand, in the second embodiment, the adding unit 3C generates a virtual road paint-added planar image IM3 by adding virtual road paint VP (road markings) to the planar normal image converted image IM2 generated by the image converting unit 3B.

[0025] As described above, in the first embodiment (the example shown in Figures 1 to 6), the training image generation unit 3D generates a training fisheye image IM4 including virtual road paint VP (demarcation lines) by performing a transformation on the planar image IM3 with virtual road paint that is the inverse of the planar normalization transformation performed by the image conversion unit 3B. On the other hand, in the second embodiment, the training image generation unit 3D performs a transformation on the planar image IM3 with virtual road paint that is the inverse of the planar normalization transformation performed by the image conversion unit 3B, thereby generating a training fisheye image IM4 including virtual road paint VP (road marking). In the second embodiment, the training vehicle L1 and the training trailer L2 do not need to actually travel on a road surface on which road paint (road markings) are painted, and a training fisheye image IM4 including road markings (virtual road paint VP) can be obtained in the same way as if the training vehicle L1 and the training trailer L2 actually traveled on a road surface on which road paint (road markings) are painted. In other words, the second embodiment makes it possible to train a model using a fisheye image including road paint (road markings) as training data, without the need to actually capture a fisheye image including road paint (road markings) with the training camera L11.

[0026] As described above, embodiments of the training image generation device, training image generation method, and program of the present disclosure have been described with reference to the drawings. However, the training image generation device, training image generation method, and program of the present disclosure are not limited to the above-described embodiments and may be modified as appropriate without departing from the spirit of the present disclosure. The configurations of the above-described embodiments may be combined as appropriate. In the above-described embodiments, the processing performed in the training image generation device 1 has been described as software processing performed by executing a program. However, the processing performed in the training image generation device 1 may be processing performed by hardware. Alternatively, the processing performed in the training image generation device 1 may be processing that combines both software and hardware. Furthermore, the program stored in the memory 12 of the training image generation device 1 (a program that realizes the functions of the processor 13 of the training image generation device 1) may be recorded on a computer-readable storage medium such as a semiconductor memory, a magnetic recording medium, an optical recording medium, etc., and provided, distributed, etc. [Explanation of symbols]

[0027] 1...Learning image generation device, 11...Communication interface, 12...Memory, 13...Processor, 3A...Acquisition unit, 3B...Image conversion unit, 3C...Addition unit, 3D...Learning image generation unit, IM1...Learning fisheye original image, IM2...Planar normalized converted image, IM3...Planar image with virtual road surface paint, IM4...Learning fisheye image, VP...Virtual road surface paint, L1...Learning vehicle, L1R...Rear end, L11...Learning camera, L2...Learning trailer, L3...Learning tow bar

Claims

1. 1. A training image generation device that generates training fisheye images as training data used in training a model used to estimate a state of a trailer based on fisheye images taken by a camera mounted on a vehicle towing a trailer via a tow bar, comprising: an acquisition unit that acquires a training fisheye original image taken by a training camera mounted on a training vehicle that tows a training trailer via a training tow bar; an image conversion unit that generates a planar-normalized converted image by performing planar normalization conversion, which is a conversion from a fisheye image to a planar image, on the learning fisheye original image acquired by the acquisition unit; an adding unit that adds virtual road surface paint to the planar normalized transformed image generated by the image transforming unit; and a training image generation unit that generates the training fisheye image by performing a transformation that is the inverse of the planar normalization transformation performed by the image conversion unit on the planar image with virtual road paint, which is the planar normalization transformed image after the virtual road paint has been added by the addition unit.

2. a training image generation step in which a training image generation device generates training fisheye images as training data used to train a model used to estimate the state of the trailer based on fisheye images taken by a camera mounted on a vehicle towing a trailer via a tow bar; an acquisition step in which the learning image generation device acquires a learning fisheye original image taken by a learning camera mounted on a learning vehicle towing a learning trailer via a learning tow bar; an image conversion step in which the learning image generation device performs planar image normalization conversion, which is a conversion from a fisheye image to a planar image, on the learning fisheye original image acquired in the acquisition step, to generate a planar image normalized converted image; an additional step of the learning image generating device adding virtual road paint to the planar normalized transformed image generated in the image transforming step; In the learning image generating step, the learning fisheye image is generated by performing a transformation inverse to the planar normalization transformation performed in the image transforming step on the planar image with virtual road paint, which is the planar normalization transformed image after the virtual road paint is added in the adding step. A method for generating images for training.

3. The processor a training image generation step of generating training fisheye images as training data used to train a model used to estimate the state of the trailer based on fisheye images taken by a camera mounted on a vehicle towing a trailer via a tow bar; an acquisition step of acquiring a training fisheye original image taken by a training camera mounted on a training vehicle towing a training trailer via a training tow bar; an image conversion step of generating a planar-normalized converted image by performing planar normalization conversion, which is a conversion from a fisheye image to a planar image, on the learning fisheye original image acquired in the acquisition step; and adding a virtual road surface paint to the planar normalized transformed image generated in the image transformation step, In the learning image generation step, the learning fisheye image is generated by performing a transformation that is the inverse of the planar normalization transformation performed in the image transformation step on the planar image with virtual road paint, which is the planar normalization transformed image after the virtual road paint is added in the addition step.

Citation Information

Patent Citations

  • Traction support device

    JP2018176788A