Learning device, learning method, storage medium and object detection device
By correcting the pitch amount of the image captured by the moving object and adding computer graphics images, an appropriate learning model is generated, which solves the problem of object discrimination in the image captured by the moving object and improves the accuracy of object recognition.
Patent Information
- Application Number
- CN202210123471.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-30
- Filing Date
- 2022-02-09
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-02-09
AI Technical Summary
In the prior art, when a mobile body is equipped with a camera to capture an image, it is difficult to appropriately generate a learning completion model for discriminating objects on the road.
The learning device performs pitch correction, sunlight direction estimation and movement calculation of the captured image, adds computer graphic images, generates learning images and learns model parameters, and forms an appropriate learning completion model.
It realizes the appropriate generation of object discrimination models on the road, and improves the accuracy and reliability of object recognition.
Smart Images

Figure CN115147528B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a learning device, a learning method, a storage medium and an object detection device. Background Art
[0002] Conventionally, a technique has been disclosed in which teaching data for learning is created with respect to three-dimensional graphics drawn by a graphics rendering unit and a deep learning recognition unit is caused to learn the data (Patent Document 1).
[0003]
Prior technical literature
[0004] [Patent Literature]
[0005] [Patent Document 1] International Publication No. 2017 / 171005 Summary of the Invention
[0006] Technical problem to be solved by the invention
[0007] Conventional technologies may not be able to appropriately generate a learned model for identifying objects present on a road in an image captured by a camera mounted on a moving object.
[0008] The present invention has been made in consideration of such circumstances, and one of its objects is to provide a learning device, a learning method, a storage medium, and an object detection device that can appropriately generate a learned model for distinguishing objects existing on the road.
[0009] Technical solutions to technical problems
[0010] The learning device, learning method, storage medium, and object detection device of the present invention adopt the following structures.
[0011] (1): A learning device according to one embodiment of the present invention, wherein the learning device comprises: a captured image acquisition unit that acquires a captured image of a road; a CG image appending unit that appends a computer graphic image of an object existing on the road to a live image obtained based on the captured image; and a learning processing unit that uses the category of the added computer graphic image as teaching data to learn the parameters of the learned model in a manner such that the category of the object is output when an image is input.
[0012] (2): In the above-mentioned scheme (1), the learning device further includes a sunlight direction estimating unit, which estimates the sunlight direction based on the live image obtained based on the captured image, and the CG image adding unit gives the computer graphic image of the object a shadow obtained based on the sunlight direction.
[0013] (3): In the above-mentioned scheme (1) or (2), the captured image is captured by a camera mounted on a moving body, and the learning device further comprises: a pitch amount estimating unit, which estimates the pitch amount of the moving body at each shooting time point based on the captured image; and a first correction unit, which performs a first correction on the captured image to eliminate the pitch amount and generate the live image.
[0014] (4): In the above-mentioned scheme (3), the learning device further includes a second correction unit, which performs a second correction on the live image obtained by adding the computer graphic image to restore the first correction to generate a learning image, and the learning processing unit uses the learning image as learning data to learn the parameters of the learned model.
[0015] (5): In the above-mentioned scheme (1) or (2), the captured image is a captured image taken by a camera mounted on a moving body, the live image is the captured image, the learning device further includes a pitch amount estimation unit, which estimates the pitch amount of the moving body at each shooting time point based on the captured image, and the CG image appending unit appends the computer graphics image at a position corresponding to the pitch amount in the live image.
[0016] (6): In any one of the above schemes (1) to (5), the captured image is captured by a camera mounted on a moving body, the learning device further includes a movement amount acquisition unit, which acquires the movement amount of the moving body, and the CG image appending unit determines the position and size of the computer graphics image based on the movement amount of the moving body.
[0017] (7): Another embodiment of the present invention is a learning method that is executed using a computer, wherein the learning method comprises the following steps: obtaining a photographic image obtained by photographing a road; appending a computer graphic image of an object existing on the road to a real-life image obtained based on the photographic image; and using the category of the appended computer graphic image as teaching data to learn the parameters of the learned model in such a manner that the category of the object is output when an image is input.
[0018] (8): A storage medium according to another embodiment of the present invention stores a program, wherein the program causes a computer to perform the following processing: obtaining a photographic image obtained by photographing a road; appending a computer graphic image of an object existing on the road to a real-life image obtained based on the photographic image; and learning the parameters of the learned model by using the category of the appended computer graphic image as teaching data to output the category of the object when an image is input.
[0019] (9): Another embodiment of the present invention relates to an object detection device, which is mounted on a mobile body, wherein the object detection device inputs a captured image obtained by a camera mounted on the mobile body on at least a road in the direction of travel of the mobile body into the learned model obtained by learning by the learning device of any one of the embodiments (1) to (6), thereby determining whether an object on the road reflected in the captured image is an object that the mobile body should avoid contacting.
[0020] (10): A learning device according to another embodiment of the present invention, wherein the learning device comprises: a captured image acquisition unit that acquires a captured image of a road; a CG image appending unit that appends a computer graphic image of an object existing on the road to a live image obtained based on the captured image; and a learning processing unit that uses the position of the appended computer graphic image as teaching data to learn the parameters of the learned model in a manner such that the position of the object is output when an image is input.
[0021] (11): Another embodiment of the present invention is a learning method that is performed using a computer, wherein the learning method includes the following processing: obtaining a photographic image obtained by photographing a road; appending a computer graphic image of an object existing on the road to a real-life image obtained based on the photographic image; and using the position of the appended computer graphic image as teaching data to learn the parameters of the learned model in a manner that outputs the position of the object when an image is input.
[0022] (12): A storage medium according to another embodiment of the present invention stores a program, wherein the program causes a computer to perform the following processing: obtaining a photographic image obtained by photographing a road; appending a computer graphic image of an object existing on the road to a real-life image obtained based on the photographic image; and learning the parameters of the learned model by using the position of the appended computer graphic image as teaching data to output the position of the object when an image is input.
[0023] Effects of the Invention
[0024] According to the above-mentioned solutions (1) to (8) and (10) to (12), a learned model for distinguishing objects existing on the road can be appropriately generated.
[0025] According to the above-mentioned solution (9), objects on the road can be appropriately discriminated using the learned model obtained by appropriate learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 1 is a diagram showing an example of the configuration of the learning device 100 .
[0027] Figure 21 is a diagram illustrating the flow of processing performed by each component of the learning device 100 .
[0028] Figure 3 IM2 is a diagram showing how the position of the road surface in the corrected image IM2 changes according to the slope.
[0029] Figure 4 It is a diagram for explaining the processing of the pitch amount estimating unit 120 .
[0030] Figure 5 It is a diagram for explaining the content of area correction.
[0031] Figure 6 It is a diagram for explaining the content of the up-down shift processing.
[0032] Figure 7 This is a diagram for explaining the details of the process of obtaining the offset d.
[0033] Figure 8 It is a diagram for explaining the process of estimating the pitch amount.
[0034] Figure 9 It is a diagram for explaining the process of estimating the pitch amount.
[0035] Figure 10 It is a diagram for explaining the processing of the lateral movement amount estimating unit 131 .
[0036] Figure 11 This is a diagram for explaining the contents of left-right shifting and image clipping processing.
[0037] Figure 12 It is a diagram for explaining the processing of the speed estimating unit 132 .
[0038] Figure 13 This is a diagram for explaining the contents of the enlargement / reduction and image cropping processing.
[0039] Figure 14 1 is a diagram showing an example of the configuration and usage environment of the object detection device 200 .
[0040] Description of reference numerals:
[0041] 10. Camera
[0042] 100 Learning Devices
[0043] 110 captured image acquisition unit
[0044] 120 Pitch estimation unit
[0045] 122 First Amendment
[0046] 124 Second Amendment
[0047] 126 Sunlight Direction Estimation
[0048] 130 Movement amount estimation unit
[0049] 140 CG image additions
[0050] 150 Learning Processing Department
[0051] 170 Storage Department
[0052] 200 Object Detection Device
[0053] 252 Learning completed model. DETAILED DESCRIPTION
[0054] The following describes embodiments of the learning device, learning method, storage medium, and object detection device of the present invention with reference to the accompanying drawings. The learning device is a device that generates a learned model for object discrimination used by an object detection device mounted on a mobile body. Examples of mobile bodies include self-propelled devices such as four-wheeled vehicles, two-wheeled vehicles, micromobilities, and robots, or mobile devices such as smartphones that are carried by a self-moving mobile body or carried by a person. In the following description, the mobile body is assumed to be a four-wheeled vehicle, and the mobile body is referred to as a "vehicle" for explanation.
[0055] [Learning device]
[0056] Figure 1This figure shows an example of the structure of the learning device 100. The object detection device 200 includes, for example, a captured image acquisition unit 110, a pitch amount estimation unit 120, a first correction unit 122, a second correction unit 124, a sunlight direction estimation unit 126, a movement amount acquisition unit 130, a CG (computer graphics) image addition unit 140, a learning processing unit 150, and a storage unit 170. The movement amount acquisition unit 130 includes, for example, a lateral movement amount estimation unit 131 and a speed estimation unit 132. The CG image addition unit 140 includes, for example, a CG image generation unit 141, a CG image movement amount / magnification ratio calculation unit 142, a road surface position estimation unit 143, a CG image position determination unit 144, and an image synthesis unit 145. Each component other than the storage unit 170 is implemented by, for example, a hardware processor such as a CPU (Central Processing Unit) executing a program (software). Some or all of these components may be implemented by hardware (including circuitry) such as LSI (Large Scale Integration), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), or GPU (Graphics Processing Unit), or by a combination of software and hardware. The program may be pre-stored in a storage device (a storage device with a non-transitory storage medium) such as an HDD (Hard Disk Drive) or flash memory, or may be stored in a removable storage medium (a non-transitory storage medium) such as a DVD or CD-ROM and installed by attaching the storage medium to a drive device. The storage unit 170 stores data and information such as actual driving images 171 and machine learning model definition information 172.
[0057] The captured image acquisition unit 110 acquires, for example, a captured image that is obtained by capturing at least a road in the direction of travel of the vehicle and captured by a camera mounted on the vehicle while the vehicle is moving (of course, this also includes the case of a temporary stop). A captured image is a moving image in which images are connected in a time series. The captured image can also be captured by a fixed-point observation camera or by a camera of a smartphone. For example, the actual driving image 171 stored in the storage unit 170 is an example of the captured image. The captured image acquisition unit 110 reads the actual driving image 171 from the storage unit 170 and provides it to other functional units (for example, expands it to a shared area of RAM (Random Access Memory)). The actual driving image 171 is an image captured along with the movement of the vehicle in a vehicle equipped with a camera that captures the side in the direction of travel. The actual driving image 171 may be provided to the learning device 100 by an onboard communication device via a network such as a WAN (Wide Area Network) or a LAN (Local Area Network), or may be transferred to the learning device 100 from various portable storage devices and stored in the storage unit 170 .
[0058] For the subsequent processing of the functional unit, refer to Figure 2 To explain. Figure 2 1 is a diagram illustrating the flow of processing performed by each unit of the learning device 100. Here, the actual travel image 171 is referred to as an actual travel image IM1.
[0059] The pitch amount estimation unit 120 estimates the pitch amount of the vehicle at each shooting time point in the actual driving image IM1. For example, the pitch amount estimation unit 120 compares the actual driving image IM1 at the shooting time (hereinafter simply referred to as time) k with the actual driving image IM1 at time k-1, and estimates the pitch amount of the vehicle between time k and time k-1 as the pitch amount at time k. The pitch amount estimation unit 120 performs the relevant processing for each time. The pitch amount refers to the amount of rotation around the axis with the left and right directions of the vehicle as the axis. When there is a pitch amount, longitudinal fluctuations are generated in the actual driving image IM1 between time k and time k-1, so it is necessary to perform a process to correct the pitch amount in advance for the processing of the CG image appending unit 140 described later. Details of the processing of the pitch amount estimation unit 120 will be described later.
[0060] The first correction unit 122 performs a first correction on the actual driving image IM1 to eliminate vertical fluctuations in the image corresponding to the pitch amount estimated by the pitch amount estimating unit 120. For example, the first correction unit 122 determines the amount of the first correction using a table or map that specifies the vertical fluctuations of each pixel in the image with respect to the pitch amount. The image obtained by the first correction unit 122 is referred to as the corrected image IM2. The corrected image IM2 is an example of a "live image obtained based on a captured image."
[0061] The second correction unit 124 performs a second correction to restore the first correction on the CG-added image IM3 obtained by adding a CG image to the corrected image IM2 by the CG image adding unit 140 , thereby generating a learning image IM4 .
[0062] The sunlight direction estimating unit 126 estimates the sunlight direction of the scenery reflected in the actual driving image IM1 based on the corrected image IM2. For example, the sunlight direction estimating unit 126 identifies the shadow portion of the corrected image IM2 and estimates the three-dimensional object that creates the shadow. The sunlight direction obtained based on the identified positional relationship between the two is converted from the image plane into real space, thereby deriving the sunlight direction in real space.
[0063] The movement amount acquisition unit 130 estimates the movement amount of the vehicle. Specifically, the movement amount acquisition unit 130 estimates the following two movement amounts.
[0064] The lateral movement amount estimating unit 131 of the movement amount acquiring unit 130 compares the actual driving image IM1 at time k with the actual driving image IM1 at time k-1, and estimates the lateral movement amount of the vehicle between time k and time k-1 (referred to as the movement amount in a direction perpendicular to the vehicle's central axis or the movement amount in the road width direction). Details of the processing performed by the lateral movement amount estimating unit 131 will be described later.
[0065] The speed estimating unit 132 of the movement amount acquiring unit 130 compares the actual driving image IM1 at time k with the actual driving image IM1 at time k-1, for example, to estimate the vehicle speed at time k. Details of the processing of the speed estimating unit 132 will be described later.
[0066] The CG image adding unit 140 adds a CG image of an object existing on the road to the corrected image IM2 .
[0067] The CG image generation unit 141 of the CG image appending unit 140 reads, for example, from the storage unit 170 a template image of a number of objects that may be present on the road as fallen or placed objects (e.g., tires, corrugated cardboard boxes, bicycles, steel frames, etc.), which are objects that vehicles should avoid contact with, and transforms or re-renders the template image to a size corresponding to the magnification factor described later. Furthermore, the template image is given a shadow based on the sunlight direction estimated by the sunlight direction estimation unit 126 to generate a CG image. Furthermore, the CG image generation unit 141 generates CG images of objects that vehicles do not need to avoid, such as road markings, manholes, and road surface materials. The CG image generation unit 141 generates CG images of road markings that are not collected as captured images, road markings with white spots, and frequently occurring fallen objects at various angles and in various states (such as crushed, damaged, or stained). This makes it possible to expand the range covered by the learning process compared to the case where a learned model is generated using only captured images, reduce the probability of recognition as an unlearned object, and suppress the occurrence of erroneous avoidance in the vehicle.
[0068] The CG image movement amount / magnification ratio calculation unit 142 of the CG image adding unit 140 determines the position and size of the CG image based on the vehicle's movement amount estimated by the movement amount acquisition unit 130. Specifically, the CG image movement amount / magnification ratio calculation unit 142 adds the image displacement amount (calculated by projecting the displacement amount on an imaginary plane when viewed from above onto the image plane) based on the distance traveled by the vehicle from time k-1 to time k to the position of the CG image at time k-1, thereby calculating the position of the CG image at time k. The CG image movement amount / magnification ratio calculation unit 142 outputs the calculated position of the CG image at time k to the CG image position determination unit 144. Furthermore, the CG image movement amount / magnification ratio calculation unit 142 calculates a magnification ratio for the CG image at time k-1 based on the distance traveled by the vehicle from time k-1 to time k. The CG image movement amount / magnification ratio calculation unit 142 outputs the magnification ratio to the CG image generation unit 141. When the vehicle is moving forward, the magnification ratio becomes a value greater than 1, and when the vehicle is moving backward, the magnification ratio becomes a value less than 1 (ie, is reduced).
[0069] The road surface position estimating unit 143 of the CG image adding unit 140 estimates the position of the road surface in the actual space in the traveling direction of the mobile object 1 based on the corrected image IM2 , and establishes a correspondence relationship between each position on the road surface and each position in the corrected image IM2 . Figure 3This diagram shows how the position of the road surface in the corrected image IM2 changes depending on the slope. As shown in the figure, if the road surface in real space in the direction of travel of the mobile body 1 is uphill or downhill (especially when the vehicle is flat and has a slope in the direction of travel), or if the vehicle is turning left or right, the position of the road surface in the corrected image IM2 changes, so the position change on the CG image should be added. The road surface position estimation unit 143 performs the above-mentioned processing to identify the position on this image. The processing of the road surface position estimation unit 143 can be processing to establish a correspondence between each position on the image in an ideal state where the road surface is level and there are no turns, and each position on the corrected image IM2 representing the actual state.
[0070] The CG image position determination unit 144 of the CG image adding unit 140 determines the position of the CG image on the image based on the calculated position (corresponding to the position in the aforementioned ideal state) obtained from the CG image movement amount / magnification ratio calculation unit 142 and the position of the road surface estimated by the road surface position estimation unit 143.
[0071] The image synthesis unit 145 of the CG image adding unit 140 superimposes (adds) the CG image generated by the CG image generating unit 141 onto the corrected image IM2 at the position of the CG image determined by the CG image position determining unit 144. This generates a CG-added image IM3. The CG image adding unit 140 may erase and overwrite the road surface image at the position occupied by the CG image, or may retain the road surface image and add the pixel values of the CG image.
[0072] As described above, the second correction unit 124 performs the second correction on the CG-added image IM3 to restore the first correction, thereby generating the learning image IM4 .
[0073] The learning processing unit 150 learns the parameters of the machine learning model defined by the machine learning model definition information 172. The machine learning model definition information 172 is information that defines the number of nodes, connection relationships, etc. of the machine learning model. Figure 2 The example shown is a DNN (Deep Neural Network). The learning processing unit 150 uses the learning image IM4 as learning data and the category of the additional CG image as teaching data. Using methods such as backpropagation, the learning unit learns the parameters of the machine learning model so that, when the image is input, the category of the object (the object the vehicle should avoid contact with) is output. The machine learning model, after parameter learning, is used as a learned model in the vehicle-mounted device. Therefore, it can be said that the learning processing unit 150 learns the parameters of the learned model.
[0074] The learning processing unit 150 may also use the learning image IM4 as learning data and the position of the additional CG image as teaching data instead of (or in addition to) this, and learn the parameters of the machine learning model by outputting the position of the object (the object that the vehicle should avoid contacting) when the image is input through methods such as back propagation.
[0075] In the above description, the CG image adding unit 140 adds a CG image of an object on the road to the corrected image IM2 obtained by performing a first correction on the actual driving image IM1 to eliminate vertical fluctuations in the image corresponding to the pitch amount estimated by the pitch amount estimating unit 120. The CG image adding unit 140 also performs a second correction to the CG-added image IM3 to which the CG image is added, thereby reversing the first correction, thereby generating a learning image IM4. Alternatively, the CG image adding unit 140 may add a CG image to the actual driving image IM1 at a position where the position of the CG image determined by the CG image position determining unit 144 is corrected based on the pitch amount estimated by the pitch amount estimating unit 120. In this case, the second correction is unnecessary.
[0076] Through the above-described processing, a learned model for distinguishing objects existing on the road can be appropriately generated.
[0077] Hereinafter, the processing of the pitch amount estimating unit 120 , the lateral movement amount estimating unit 131 , and the speed estimating unit 132 will be described in more detail. Figure 4 This diagram illustrates the processing performed by the pitch amount estimation unit 120. The pitch amount estimation unit 120 grayscales the actual driving image IM1 at time k and time k-1. Hereinafter, these grayscaled images are referred to as grayscale images. The pitch amount estimation unit 120 performs image cropping processing on the grayscale image at time k, cropping the image by deleting unnecessary portions.
[0078] The pitch amount estimation unit 120 performs regional correction based on the vehicle's speed on the grayscale image at time k-1. The vehicle's speed may be the result of processing by the speed estimation unit 132, described later, or may be the speed obtained from a speed sensor in conjunction with camera imaging in the vehicle (the speed provided with the captured image). Figure 5This is a diagram for explaining the content of the area correction. The pitch amount estimation unit 120 calculates the distance the vehicle has moved from time k-1 to time k from the speed of the vehicle, and enlarges the grayscale image at time k-1 with a magnification based on the distance moved (i.e., a magnification based on the speed). This is because: for the portion obtained by photographing the same part, if the vehicle is moving forward, the time k will be closer than the time k-1, and the time k will be more enlarged. Through area correction, the pitch amount estimation unit 120 makes the same object appear in the grayscale image at time k and time k-1 with the same size, so that the comparison between the images can be performed accurately. The pitch amount estimation unit 120 trims the upper, lower, left, and right ends of the enlarged image to return it to the original size.
[0079] return Figure 4 The pitch amount estimating unit 120 performs a vertical shift process on the grayscale image at time k-1 after the region correction, thereby generating N images. Figure 6 This is a diagram for explaining the content of the up-down shift processing. The pitch estimation unit 120 generates a plurality of (here, N) comparison target images with different offsets in the upward and downward directions, respectively, with respect to the grayscale image at the time k-1 after the region correction. In the figure, the dotted line represents the grayscale image before the offset is generated. Furthermore, between the region (cropped region) obtained by cropping the region of the determined position from the grayscale image at the time k and the cropped region of the comparison target image, N difference images ( Figure 3 Return Figure 4 The pitch amount estimation unit 120 performs binarization processing by comparing the pixel value of each pixel in the difference image with a threshold value, assigning a pixel value of 1 if the pixel value is above the threshold value, and assigning a pixel value of zero if the pixel value is below the threshold value. The pitch amount estimation unit 120 then selects the target difference image with the smallest total pixel value (the largest number of zero pixels) and calculates the offset required to generate the target difference image. Figure 7 This figure explains the process of determining the offset d. In the example shown, the difference target image from the comparison target image obtained by shifting the offset d by -2, that is, shifting it downward by 2 pixels, is selected as the difference target image with the smallest total pixel value.
[0080] Next, the pitch amount estimating unit 120 estimates the pitch amount based on the offset amount d. Figure 8 and Figure 9This figure illustrates the process of estimating the pitch amount. The relationship between the offset d and the pitch amount θ is expressed by the following equation. In the equation, l is the distance in real space to the vanishing point VP in the image (e.g., approximately 100 to 200 meters), and h is twice the height from the bottom edge of the image to the vanishing point VP. The pitch amount estimation unit 120 calculates the pitch amount θ based on this equation. It should be noted that while h varies depending on the image size, the pitch amount θ is essentially restored to the offset in the image and used, so this does not pose a problem.
[0081] tan{(α / 2)+θ}={d+(h / 2)} / l
[0082] Figure 10 This diagram illustrates the processing performed by the lateral movement estimation unit 131. The lateral movement estimation unit 131 grayscales the actual driving image IM1 at time k and time k-1. (The grayscale processing may also be shared with the pitch amount estimation unit 120. The same applies to the speed estimation unit 132.) The lateral movement estimation unit 131 performs image cropping processing on the grayscale image at time k-1, deleting unnecessary portions.
[0083] The lateral movement amount estimation unit 131 performs regional correction based on the vehicle speed on the grayscale image at time k. This is similar to the pitch amount estimation unit 120 , and the enlargement / reduction direction is the same as that of the pitch amount estimation unit 120 .
[0084] The lateral movement amount estimating unit 131 performs left-right shifting and image clipping processing on the grayscale image at time k that has undergone area correction, thereby generating M images. Figure 11 This is a diagram for explaining the contents of the left-right shift and image clipping processing. The lateral movement amount estimation unit 131 generates a plurality of (here, M) comparison target images with different shift amounts in the left and right directions, respectively, for the grayscale image at the time k where the area correction is performed. In the figure, the dotted line represents the grayscale image before the shift occurs. Furthermore, M difference images (pixels with the difference between pixels) are generated between the area (clipped area) obtained by clipping the area of the determined position from the grayscale image at the time k-1 and the clipped area of the comparison target image. Figure 10 Return Figure 10 The lateral movement amount estimation unit 131 performs binarization processing by comparing the pixel value of each pixel in the difference image with a threshold value, assigning a pixel value of 1 if the pixel value is above the threshold value, and assigning a pixel value of zero if the pixel value is below the threshold value. The lateral movement amount estimation unit 131 then selects the target difference image with the smallest total pixel value (the largest number of zero pixels) and calculates the lateral movement amount ΔY based on the shift amount required to generate the target difference image.
[0085] Figure 12 This diagram is used to explain the processing of the speed estimation unit 132. The speed estimation unit 132 grayscales the actual driving image IM1 at time k and time k-1. The speed estimation unit 132 performs image clipping processing on the grayscale image at time k-1 to delete unnecessary parts.
[0086] The speed estimation unit 132 performs enlargement / reduction and image clipping processing on the grayscale image at time k to generate R images. Figure 13 This is a diagram for explaining the contents of the zooming / reduction and image clipping processing. The speed estimation unit 132 generates a plurality of (here, R) comparison target images with different zoom ratios (reduction ratios) in the zooming direction and the reduction direction, respectively, with respect to the grayscale image at time k. In the figure, the dotted line represents the grayscale image before zooming in or out. Furthermore, R difference images ( R ) are generated between the area (clipped area) obtained by clipping the area at the position determined from the grayscale image at time k-1 and the clipped area of the comparison target image, with the difference between the pixels as the pixel. Figure 10 Return Figure 12 The velocity estimation unit 132 performs binarization processing by comparing the pixel value of each pixel in the difference image with a threshold value, assigning a pixel value of 1 if the pixel value is above the threshold value, and assigning a pixel value of zero if the pixel value is below the threshold value. The velocity estimation unit 132 then selects the target difference image with the smallest total pixel value (the largest number of zero pixels), and calculates the velocity V based on the magnification (reduction) ratio used to generate the target difference image.
[0087] According to the embodiment of the learning device described above, it comprises: a captured image acquisition unit 110, which acquires a captured image of at least a road in the direction of travel of the moving body and is captured by a camera mounted on the moving body while the moving body is moving; a CG image appending unit 140, which appends a computer graphic image of an object existing on the road to the live image IM2 obtained based on the captured image; and a learning processing unit 150, which uses the position of the added computer graphic image as teaching data to learn the parameters of the learned model in a manner that outputs the position of the object when the image is input, thereby being able to appropriately generate a learned model for distinguishing objects existing on the road.
[0088] [Object detection device]
[0089] The following describes an embodiment of an object detection device that utilizes a learned model generated by learning device 100. The object detection device is, for example, mounted on a mobile object. Examples of the mobile object include a four-wheeled vehicle, a two-wheeled vehicle, a micromobile object, a robot, and the like. In the following description, the mobile object is assumed to be a four-wheeled vehicle and is referred to as a "vehicle."
[0090] Figure 14 1 is a diagram showing an example of the configuration and usage environment of the object detection device 200. The object detection device 200 communicates with the camera 10, the travel control device 300, the reporting device 310, and the like.
[0091] The camera 10 is mounted on the back of the vehicle's windshield, for example, to capture images of at least the road in the vehicle's direction of travel and output the captured images to the object detection device 200. It should be noted that a sensor fusion device or the like may be located between the camera 10 and the object detection device 200, but this description is omitted.
[0092] The driving control device 300 may be, for example, an automatic driving control device that causes the vehicle to drive autonomously, or a driving support device that performs inter-vehicle distance control, automatic braking control, automatic lane change control, etc. The reporting device 310 may be a speaker, vibrator, light-emitting device, display device, etc., for outputting information to the vehicle's occupants.
[0093] The object detection device 200 includes, for example, an acquisition unit 210, a low-resolution processing unit 220, a high-resolution processing unit 230, and a storage unit 250. The storage unit 250 stores a learned model 252 obtained by learning by the learning device 100. The acquisition unit 210, the low-resolution processing unit 220, and the high-resolution processing unit 230 are each implemented by executing a program (software) through a hardware processor such as a CPU. Some or all of these components can also be implemented by hardware (including circuitry) such as an LSI, ASIC, FPGA, GPU, or the like, or by the collaboration of software and hardware. The program can be pre-stored in a storage device such as an HDD or flash memory (a storage device having a non-temporary storage medium), or can be stored in a removable storage medium such as a DVD or CD-ROM (a non-temporary storage medium), and installed by assembling the storage medium in a drive device.
[0094] The acquisition unit 210 acquires a captured image from the camera 10. The acquisition unit 210 stores (data of) the acquired captured image in a working memory such as a RAM.
[0095] The low-resolution processing unit 220 performs thinning processing on the captured image, for example, to generate a low-resolution image with lower quality than the captured image. For example, a low-resolution image has fewer pixels than the captured image. The low-resolution processing unit 220 extracts a region containing a characteristic feature from the low-resolution image and outputs this region as the target area to the high-resolution processing unit 230. The specific example of the process for extracting this region is not particularly limited, and any method may be employed.
[0096] The high-resolution processing unit 230 cuts out a portion of the captured image corresponding to the target area and inputs the image of this portion into the learned model 252. The learned model 252 determines whether the image projected onto the target area is a road sign, a fallen object (an object learned using CG images in the learning device 100), or an unidentified object (an object that has not been learned).
[0097] The determination results of the high-resolution processing unit 230 are output to the driving control device 300 and / or the reporting device 310. The driving control device 300 performs automatic braking control, automatic steering control, and other functions to prevent the vehicle from colliding with objects identified as "fallen objects" (actually, areas on the image) and unknown objects (unlearned objects). The reporting device 310 outputs an alarm using various methods when the TTC (Time To Collision) between an object identified as "fallen objects" (same as above) and the vehicle falls below a threshold.
[0098] According to the embodiment of the object detection device described above, objects on the road can be appropriately discriminated using the appropriately learned model 252 .
[0099] The above-described embodiment can be expressed as follows.
[0100] A learning device comprising:
[0101] a storage device storing a program; and
[0102] Hardware processor,
[0103] The hardware processor executes the program stored in the storage device to perform the following processing:
[0104] Obtaining images taken on the road;
[0105] adding a computer graphic image of an object existing on the road to a live image obtained based on the captured image;
[0106] The added category of the computer graphics image is used as teaching data, and the parameters of the learned model are learned in such a manner that the category of the object is output when an image is input.
[0107] The above-described embodiment can also be expressed as follows.
[0108] A learning device comprising:
[0109] a storage device storing a program; and
[0110] Hardware processor,
[0111] The hardware processor executes the program stored in the storage device to perform the following processing:
[0112] Obtaining images taken on the road;
[0113] adding a computer graphic image of an object existing on the road to a live image obtained based on the captured image;
[0114] The position of the added computer graphics image is used as teaching data, and the parameters of the learned model are learned so that the position of the object is output when the image is input.
[0115] While specific embodiments of the present invention have been described above, the present invention is not limited to these embodiments at all, and various modifications and substitutions can be made without departing from the spirit of the present invention.
Claims
1. A learning device, wherein: The learning device comprises: a captured image acquisition unit that acquires a captured image of a road captured by a camera mounted on the mobile object; a CG image adding unit for adding a computer graphic image of an object existing on the road to a live image obtained based on the captured image; a learning processing unit that uses the added category of the computer graphics image as teaching data to learn parameters of the learned model so as to output the category of the object when the image is input; a pitch amount estimating unit for estimating the pitch amount of the mobile object at each shooting time point based on the captured image; a first correction unit configured to perform a first correction on the captured image to eliminate the pitch amount and generate the live image; as well as a second correction unit for performing a second correction on the live image obtained by adding the computer graphic image, thereby restoring the first correction, to generate a learning image; The learning processing unit learns parameters of the learned model using the learning image as learning data.
2. A learning device, wherein: The learning device comprises: a captured image acquisition unit that acquires a captured image of a road captured by a camera mounted on the mobile object; a CG image adding unit for adding a computer graphic image of an object existing on the road to the live image as the captured image; a learning processing unit that uses the added category of the computer graphics image as teaching data to learn parameters of the learned model so as to output the category of the object when the image is input; as well as a pitch amount estimating unit for estimating the pitch amount of the mobile object at each shooting time point based on the shot image; The CG image adding unit adds the computer graphics image at a position corresponding to the pitch amount in the live image.
3. The learning device according to claim 1 or 2, wherein: The learning device further includes a sunlight direction estimating unit that estimates the sunlight direction based on a live image obtained based on the captured image. The CG image adding unit adds a shadow based on the sunlight direction to the computer graphic image of the object.
4. The learning device according to claim 1 or 2, wherein: The learning device further includes a movement amount acquisition unit that acquires the movement amount of the moving object. The CG image adding unit determines the position and size of the computer graphics image based on the movement amount of the moving object.
5. A learning method, which is a learning method performed using a computer, wherein: The learning method has the following processing: Obtaining an image captured by a camera mounted on a mobile object. estimating a pitch amount of the mobile object at each shooting time point based on the captured images; performing a first correction on the captured image to eliminate the pitch amount to generate a live image based on the captured image; adding computer graphics images of objects on the road to the live image; Using the category of the additional computer graphics image as teaching data, the parameters of the learned model are learned so that the category of the object is output when the image is input; performing a second correction on the live image obtained by adding the computer graphic image, which is a restoration of the first correction, to generate a learning image; as well as The parameters of the learned model are learned using the learning images as learning data.
6. A storage medium storing a program, wherein: The program causes the computer to execute the following processing: Obtaining an image captured by a camera mounted on a mobile object. estimating a pitch amount of the mobile object at each shooting time point based on the captured images; performing a first correction on the captured image to eliminate the pitch amount to generate a live image based on the captured image; adding computer graphics images of objects on the road to the live image; Using the category of the additional computer graphics image as teaching data, the parameters of the learned model are learned so that the category of the object is output when the image is input; performing a second correction on the live image obtained by adding the computer graphic image, which is a restoration of the first correction, to generate a learning image; as well as The parameters of the learned model are learned using the learning images as learning data.
7. A learning method, which is a learning method performed using a computer, wherein: The learning method has the following processing: Obtaining an image captured by a camera mounted on a mobile object. estimating a pitch amount of the mobile object at each shooting time point based on the captured images; adding a computer graphic image of an object existing on the road at a position corresponding to the pitch amount in the live image, to the live image as the captured image; as well as The added category of the computer graphics image is used as teaching data, and the parameters of the learned model are learned so that the category of the object is output when an image is input.
8. A storage medium storing a program, wherein: The program causes the computer to execute the following processing: Obtaining an image captured by a camera mounted on a mobile object. estimating a pitch amount of the mobile object at each shooting time point based on the captured images; adding a computer graphic image of an object existing on the road at a position corresponding to the pitch amount in the live image, to the live image as the captured image; as well as The added category of the computer graphics image is used as teaching data, and the parameters of the learned model are learned so that the category of the object is output when an image is input.
9. An object detection device, mounted on a mobile object, wherein: The object detection device inputs a captured image obtained by a camera mounted on the mobile body on at least the road in the direction of travel of the mobile body into the learned model obtained by learning with the learning device described in any one of claims 1 to 4, thereby determining whether the object on the road reflected in the captured image is an object that the mobile body should avoid contact with.
10. A learning device, wherein: The learning device comprises: a captured image acquisition unit that acquires a captured image of a road captured by a camera mounted on the mobile object; a CG image adding unit for adding a computer graphic image of an object existing on the road to a live image obtained based on the captured image; a learning processing unit that uses the position of the added computer graphics image as teaching data to learn parameters of the learned model so as to output the position of the object when the image is input; a pitch amount estimating unit for estimating the pitch amount of the mobile object at each shooting time point based on the captured image; a first correction unit configured to perform a first correction on the captured image to eliminate the pitch amount and generate the live image; as well as a second correction unit for performing a second correction on the live image obtained by adding the computer graphic image, thereby restoring the first correction, to generate a learning image; The learning processing unit learns parameters of the learned model using the learning image as learning data.
11. A learning method performed using a computer, wherein: The learning method includes the following processing: Obtaining an image captured by a camera mounted on a mobile object. estimating a pitch amount of the mobile object at each shooting time point based on the captured images; performing a first correction on the captured image to eliminate the pitch amount to generate a live image based on the captured image; adding computer graphics images of objects on the road to the live image; Using the position of the added computer graphics image as teaching data, the parameters of the learned model are learned so that the position of the object is output when the image is input; performing a second correction on the live image obtained by adding the computer graphic image, which is a restoration of the first correction, to generate a learning image; as well as The parameters of the learned model are learned using the learning images as learning data.
12. A storage medium storing a program, wherein: The program causes the computer to execute the following processing: Obtaining an image captured by a camera mounted on a mobile object. estimating a pitch amount of the mobile object at each shooting time point based on the captured images; performing a first correction on the captured image to eliminate the pitch amount to generate a live image based on the captured image; adding computer graphics images of objects on the road to the live image; Using the position of the added computer graphics image as teaching data, the parameters of the learned model are learned so that the position of the object is output when the image is input; performing a second correction on the live image obtained by adding the computer graphic image, which is a restoration of the first correction, to generate a learning image; as well as The parameters of the learned model are learned using the learning images as learning data.
13. A learning device, wherein: The learning device comprises: a captured image acquisition unit that acquires a captured image of a road captured by a camera mounted on the mobile object; a CG image adding unit for adding a computer graphic image of an object existing on the road to the live image as the captured image; a learning processing unit that uses the position of the added computer graphics image as teaching data to learn parameters of the learned model so as to output the position of the object when the image is input; as well as a pitch amount estimating unit for estimating the pitch amount of the mobile object at each shooting time point based on the shot image; The CG image adding unit adds the computer graphics image at a position corresponding to the pitch amount in the live image.
14. A learning method, which is a learning method performed using a computer, wherein: The learning method includes the following processing: Obtaining an image captured by a camera mounted on a mobile object. estimating a pitch amount of the mobile object at each shooting time point based on the captured images; adding a computer graphic image of an object existing on the road at a position corresponding to the pitch amount in the live image, to the live image as the captured image; as well as The position of the added computer graphics image is used as teaching data, and the parameters of the learned model are learned so that the position of the object is output when the image is input.
15. A storage medium storing a program, wherein: The program causes the computer to execute the following processing: Obtaining an image captured by a camera mounted on a mobile object. estimating a pitch amount of the mobile object at each shooting time point based on the captured images; adding a computer graphic image of an object existing on the road at a position corresponding to the pitch amount in the live image, to the live image as the captured image; as well as The position of the added computer graphics image is used as teaching data, and the parameters of the learned model are learned so that the position of the object is output when the image is input.
Citation Information
Patent Citations
3-d graphic generation, artificial intelligence verification and learning system, program, and method
WO2017171005A1
Vehicle surrounding monitoring device
JP2011182260A
3-d graphic generation, artificial intelligence verification and learning system, program, and method
US20180308281A1