Vehicle estimation device

The vehicle estimation device addresses the challenge of accurately differentiating between large and small vehicles by employing advanced image processing techniques, enabling reliable traffic volume surveys even under adverse conditions.

JP7696561B2Active Publication Date: 2025-06-23田中 成典 +3
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
JP2021068265
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-04-14
Publication Date
2025-06-23
Estimated Expiration
2041-04-14

AI Technical Summary

Technical Problem

Existing systems struggle to accurately differentiate between large and small vehicles in traffic volume surveys, especially under conditions like night-time imaging, poor weather, or when license plate characters cannot be read.

Method used

A vehicle estimation device that includes image processing modules for extracting and converting vehicle images, estimating vehicle parts, and determining vehicle size based on learned models, even under challenging imaging conditions.

Benefits of technology

The device achieves accurate classification of vehicles as large or small, enhancing the reliability of traffic volume surveys and overcoming limitations posed by environmental factors and image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696561000001
    Figure 0007696561000001
  • Figure 0007696561000002
    Figure 0007696561000002
  • Figure 0007696561000003
    Figure 0007696561000003
Patent Text Reader

Abstract

To provide a vehicle estimation device capable of highly flexible vehicle determination.SOLUTION: Vehicle image extracting means 4 recognizes a vehicle included in a captured image, surrounds the vehicle with a boundary box, and outputs the surrounded vehicle. Converted image generation means 6 converts a vehicle image into a converted vehicle image suitable for portion estimation by portion estimation vehicle image generation means 8. The portion estimation vehicle image generation means 8 outputs an estimated portion vehicle image in which each portion of the vehicle is colored separately with respect to the converted vehicle image. Large / small size estimation means 10 estimates whether the vehicle is a large-sized vehicle or a small-sized vehicle based on the portion estimation vehicle image. Since it is determined whether the vehicle is the large-sized vehicle or small-sized vehicle based on the portion estimation vehicle image, more accurate estimation can be performed. Also, the converted vehicle image suitable for the portion estimation is generated by the converted image generating means 6. Therefore, an appropriate estimation can be made even if a quality of the captured image is not favorable.SELECTED DRAWING: Figure 1a
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an apparatus for estimating whether a vehicle is a large vehicle or a small vehicle based on a captured image.

Background Art

[0002] For traffic volume surveys and management in facilities, it is necessary to determine whether a vehicle is a large vehicle or a small vehicle. For example, in traffic volume surveys, at intersections, etc., investigators visually classify the types of passing vehicles, such as large or small, and count the traffic volume for each classification.

[0003] In recent years, it has been possible to discriminate the types of objects captured using deep learning (for example, the YOLO model). An apparatus that learns this object detection model with various vehicles and discriminates the types of vehicles using the learned model can be realized.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, when applying an apparatus using the above-mentioned learned model to traffic volume surveys, etc., even if the type of vehicle could be estimated, it was not possible to estimate whether it was a large vehicle or a small vehicle. In particular, it was difficult to determine the size (for example, large truck, small truck) of the same type of vehicle (for example, truck).

[0006] In addition, a system has been proposed that reads the characters on a license plate of a vehicle and estimates the type and other information from the characters (Patent Document 1). According to this system, not only the type of the vehicle but also the distinction between large vehicles and small vehicles appearing in the characters on the license plate can be obtained. By performing estimation based on the characters on such a license plate, not only the type but also the distinction between large and small can be made possible.

[0007] However, due to factors such as the imaging distance, position, and weather, there is a possibility that the characters on the license plate cannot be read. For this reason, there were limitations in using it for traffic volume surveys and the like.

[0008] In addition, in Patent Documents 2 and 3, based on the wheelbase of the tire, it is determined whether it is a large vehicle or a small vehicle.

[0009] However, since the determination is based only on the wheelbase of the tire, its accuracy was not always sufficient.

[0010] Also, in any of the above prior arts, when the captured image is insufficient, such as not being clear due to being at night or having partial defects, it was not possible to determine whether it is a large vehicle or a small vehicle.

[0011] An object of the present invention is to solve the above problems and provide a vehicle estimation device capable of performing highly flexible vehicle determination.

Means for Solving the Problems

[0012] The independently applicable features of this invention are listed below.

[0013] (1)(2) The vehicle estimation device according to the present invention includes a vehicle image extraction means for recognizing a vehicle in a captured image obtained by capturing a traveling vehicle and extracting a vehicle image by a bounding box surrounding the vehicle, a conversion image generation means for outputting a converted vehicle image suitable for extracting parts of the vehicle based on the vehicle image, a part estimation vehicle image generation means for outputting a part estimation vehicle image in which part regions including the front and rear tires and the side surface of the vehicle are painted based on the converted vehicle image, and a size estimation means for estimating whether the vehicle is a large vehicle or a small vehicle based on the part estimation vehicle image.

[0014] The vehicle image is converted into a converted vehicle image and then a part estimation vehicle image is generated, and size estimation is performed based on this. Therefore, it is possible to accurately estimate large vehicles and small vehicles.

[0015] (3) In the vehicle estimation device according to the present invention, the vehicle image extraction means performs extraction processing using a learned extraction model learned to recognize a vehicle in the captured image based on the captured image and extract a vehicle image by a bounding box surrounding the vehicle, or the conversion image generation means performs generation processing using a learned generation model learned to generate a converted vehicle image suitable for extracting parts of the vehicle based on the vehicle image, or the part estimation vehicle image generation means performs generation processing using a learned part estimation model learned to generate a part estimation vehicle image in which part regions including the front and rear tires and the side surface of the vehicle are painted based on the converted vehicle image, or the size estimation means performs estimation processing using a learned size estimation model learned to estimate whether the vehicle is a large vehicle or a small vehicle based on the part estimation vehicle image.

[0016] Therefore, these means can be constructed by learning processing.

[0017] (4) The vehicle estimation device according to the present invention is characterized in that the conversion image generation means uses a learned conversion model that has been learned to output a converted vehicle image as if it were captured during the day, upon receiving at least the vehicle image captured at night.

[0018] Therefore, even for a dark image captured at night, it is possible to estimate large vehicles and small vehicles.

[0019] (5) The vehicle estimation device according to the present invention is characterized in that the conversion image generation means determines whether the vehicle image was captured at night based on the imaging time of the vehicle image or the data content of the vehicle image, converts the vehicle image determined to be captured at night into a converted vehicle image using the learned conversion model, and outputs the vehicle image not determined to be captured at night as the converted vehicle image as it is.

[0020] Therefore, conversion can be performed only for dark images captured at night.

[0021] (6) The vehicle estimation device according to the present invention is characterized in that the conversion image generation means uses a learned conversion model that has been learned to output a converted vehicle image with the color of the vehicle body changed, upon receiving the vehicle image.

[0022] Therefore, in relation to the background and the like, it is possible to convert to a vehicle body color that facilitates part estimation and target estimation.

[0023] (7) The vehicle estimation device according to the present invention is characterized in that the conversion image generation means determines whether the color of the vehicle body in the vehicle image is different from the color of the vehicle body after conversion based on the vehicle image data, converts the vehicle image determined to be different into a converted vehicle image using the learned conversion model, and outputs the vehicle image determined to be the same color as the converted vehicle image as it is.

[0024] Therefore, conversion can be performed only for vehicle images that require conversion.

[0025] (8) The vehicle estimation device according to the present invention is characterized in that the conversion image generation means uses a learned conversion model that has been trained to receive the vehicle image and output a converted vehicle image with reduced reflection on the vehicle body.

[0026] Therefore, it is possible to convert the vehicle image with reduced reflection into a vehicle image that is easy to perform part estimation and the like.

[0027] (9) The vehicle estimation device according to the present invention is characterized in that the conversion image generation means determines whether there is reflection on the vehicle body in the vehicle image based on the imaging time of the vehicle image or the vehicle image data, converts the vehicle image determined to have reflection into a converted vehicle image using the learned conversion model, and outputs the vehicle image determined not to have reflection as the converted vehicle image as it is.

[0028] Therefore, conversion can be performed only on the vehicle image with reflection.

[0029] (10) The vehicle estimation device according to the present invention is characterized in that the conversion image generation means uses a learned conversion model that has been trained to receive the vehicle image having at least a missing part and output a converted vehicle image with the missing part complemented.

[0030] Therefore, it is possible to convert the vehicle image without missing parts into a vehicle image that is easy to perform part estimation and the like.

[0031] (11) The vehicle estimation device according to the present invention is characterized in that the conversion image generation means determines whether there is a missing part in the vehicle image, outputs a converted vehicle image using the learned conversion model if it is determined that there is a missing part, and outputs the vehicle image as the converted vehicle image as it is if it is determined that there is no missing part.

[0032] Therefore, conversion can be performed only on the vehicle image with a missing part.

[0033] (12)(13) The vehicle estimation device according to the present invention further includes grayscale conversion means for converting the vehicle image or the converted vehicle image into grayscale to output a grayscale vehicle image or a grayscale converted vehicle image, and the size estimation means estimates whether the vehicle is a large vehicle or a small vehicle based on the part-estimated vehicle image and the grayscale vehicle image or the grayscale converted vehicle image.

[0034] Therefore, it is possible to perform size estimation considering the information included in the original image while suppressing misjudgment of size estimation due to color differences.

[0035] (14) The vehicle estimation device according to the present invention is characterized in that the size estimation means outputs the vehicle type including the distinction between the estimated large vehicle and small vehicle.

[0036] Therefore, it is possible to perform vehicle determination considering size estimation.

[0037] (15)(16) The vehicle estimation device according to the present invention includes vehicle passage determination means for detecting that the bounding box straddles a passing line in the continuous captured images based on the passing line set in the captured image of a predetermined area and determining the passage of the vehicle.

[0038] Therefore, it is possible to measure the passage of the vehicle.

[0039] (17) The vehicle estimation device according to the present invention is characterized in that two passing lines are set, and the vehicle passage determination means determines that the vehicle has passed when the bounding box intersects both of the two passing lines.

[0040] Therefore, it is possible to more accurately determine the passage of the vehicle.

[0041] (18) The vehicle estimation device according to the present invention is characterized in that two passing lines, a first passing line and a second passing line, are set in front of the moving direction of the vehicle, and the vehicle passing determination means determines that the vehicle has passed when the bounding box intersects the second passing line after intersecting the first passing line.

[0042] Therefore, the passing direction can also be determined.

[0043] (19) The vehicle estimation device according to the present invention is characterized in that the vehicle passing determination means determines whether the center position of the bounding box has crossed a passing line on the far side with respect to the camera in the captured image among the first passing line or the second passing line, and determines whether the point at the bottom of the bounding box has crossed a passing line on the nearer side than the camera in the captured image among the first passing line or the second passing line.

[0044] Therefore, the screen can be effectively utilized to determine the passage of the passing line in a large vehicle image.

[0045] (20) The vehicle estimation device according to the present invention is characterized in that the vehicle estimation device is constructed as a server device.

[0046] Therefore, the captured image can be transmitted from the terminal device and utilized.

[0047] (21)(22) The vehicle passing detection device according to the present invention includes an imaging unit that captures an image of a traveling vehicle and outputs a captured image, a conversion image generation means that outputs a converted vehicle image suitable for recognizing the vehicle in the captured image based on the captured image, a vehicle image extraction means that recognizes the vehicle in the converted captured image and extracts a vehicle image by a bounding box surrounding the vehicle, and a vehicle passing determination means that determines the passage of the vehicle by detecting that the bounding box has straddled the passing line in consecutive captured images based on a passing line set in the captured image of the predetermined area.

[0048] Therefore, even for an imaging image not suitable for vehicle extraction such as night-time imaging, it is possible to determine the passage of a vehicle.

[0049] In the embodiment, "vehicle image extraction means" corresponds to step S4.

[0050] In the embodiment, "converted image generation means" corresponds to step S10.

[0051] In the embodiment, "part-estimated vehicle image generation means" corresponds to step S12.

[0052] In the embodiment, "size estimation means" corresponds to step S14.

[0053] The "program" is a concept including not only a program directly executable by a CPU, but also a program in source form, a compressed program, an encrypted program, etc.

Brief Description of the Drawings

[0054]

Figure 1a

Figure 1b

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24a

Figure 24b

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Mode for Carrying Out the Invention

[0055] 1. First Embodiment 1.1 Functional Configuration FIG. 1 shows the functional configuration of a vehicle estimation device according to an embodiment of the present invention. The imaging unit 2 is installed to image a location where automobiles pass, such as a road, and outputs an imaging image. The vehicle image extraction means 4 recognizes the vehicles included in the imaging image and outputs the vehicles surrounded by a bounding box. The vehicle image extraction means 4 can use, for example, an estimation model learned based on the imaging image and an image of the vehicle portion of the imaging image surrounded by a bounding box.

[0056] The conversion image generation means 6 converts the vehicle image into a conversion vehicle image that is preferable for performing part estimation by the part estimation vehicle image generation means 8. The conversion image generation means 6 can use, for example, a generator that converts the vehicle image into a conversion vehicle image. In this case, the generator may be learned by a discriminator that discriminates between the generated conversion vehicle image and the corresponding real vehicle image.

[0057] The vehicle part estimation image generation means 8 outputs a vehicle part estimation image in which each part of the vehicle is colored separately for the converted vehicle image. This vehicle part estimation image generation means 8 can use, for example, an estimation model learned based on the converted vehicle image and an image in which the converted vehicle image is colored separately for each part. As the estimation model, semantic segmentation is preferable.

[0058] The size estimation means 10 estimates whether the vehicle is a large vehicle or a small vehicle based on the vehicle part estimation image. This size estimation means 10 can use, for example, an estimation model learned based on the vehicle part estimation image and data indicating whether the vehicle is a large vehicle or a small vehicle.

[0059] As described above, in this embodiment, since it is determined whether the vehicle is a large vehicle or a small vehicle based on the vehicle part estimation image, a more accurate estimation can be performed. Further, the conversion image generation means 6 generates a converted vehicle image preferable for part estimation. Therefore, an appropriate estimation can be performed even if the quality of the captured image is not preferable.

[0060] 1.2 Hardware Configuration Fig. 2 shows the hardware configuration of the vehicle estimation device. A memory 32, a display 34, a camera 2, an SSD 36, a DVD-ROM drive 38, a keyboard / mouse 40, and a communication circuit 42 are connected to the CPU 30.

[0061] The communication circuit 42 is for connecting to the Internet. An operating system 44 and a vehicle estimation program 46 are recorded on the SSD 36. The vehicle estimation program 46 exhibits its functions in cooperation with the operating system 44.

[0062] These programs were recorded on a DVD-ROM 48 and installed on the SSD 36 via the DVD-ROM drive 38.

[0063] 1.3 Vehicle Estimation Processing Fig. 3 shows a flowchart of the vehicle estimation program 46. The CPU 30 acquires a color imaging image captured by the camera 2 and records it in the SSD 36. In this embodiment, imaging is performed as a video. The captured video is recorded in the SSD 36. Fig. 4A shows an example of the imaging image. In this embodiment, the camera 2 is provided on a utility pole or in the street, and imaging is performed at a fixed angle from above.

[0064] The CPU 30 acquires one frame of the color imaging image recorded in the SSD 36 (step S2). In this embodiment, when the imaging image is recorded, it is read out and processed in real time. However, the imaging image may be recorded once and then the vehicle estimation process may be performed later.

[0065] The CPU 30 recognizes the vehicles included in the acquired one-frame imaging image, and generates a bounding box (rectangular region) in which the vehicle is inscribed, as shown in Fig. 4B (step S4). In this embodiment, a learned inference model (for example, a machine learning model by deep learning) is used for the process of recognizing the vehicle and generating the bounding box.

[0066] In this embodiment, a model for object detection by YOLO (You Only Look Onece) is used to recognize the position and region of the vehicle from the imaging image, and the vehicle is surrounded by a bounding box and output. Note that an end-to-end method other than YOLO may be used. Also, a model for object detection such as a region proposal method (Region Proposal Method) such as a sliding window approach or R-CNN may be used.

[0067] The learning of the object detection model is performed as follows. In this embodiment, the learning process is performed using the computer in Fig. 2, but other computers may be used (the same applies to other machine learning models below).

[0068] In the learning process, first, a large number of image data including vehicles as shown in FIG. 5A are prepared and recorded in the SSD 36. Preferably, there are a large number of images of various vehicle types.

[0069] The location and angle for imaging the vehicle to obtain the image data for learning are preferably as close as possible to the location and angle in the actual operation. For example, if imaging and operation are performed at an angle as shown in FIG. 4A, it is preferable that the angle is close to this.

[0070] On the other hand, if a general-purpose object detection model regardless of the camera installation position is to be constructed, it is preferable to prepare learning data with various angles, locations, and backgrounds.

[0071] Each piece of image data (original image data) is displayed on the display 34. While looking at the image, the user operates the keyboard / mouse 40 to surround the vehicle part with a rectangle (bounding box) as shown in FIG. 5B.

[0072] As a result, as shown in FIG. 5C, the analysis data in which the upper left coordinates and lower right coordinates (coordinates in the two-dimensional image) of the bounding box are recorded are associated with the original image data and recorded in the SSD 36.

[0073] The above process is performed for all the original images. As a result, a large number of learning data will be recorded in the SSD 36.

[0074] Figure 6 shows a flowchart of a process for training an object detection model (YOLO) for vehicle recognition. The CPU 30 acquires an original image from the SSD 36 (step S52). For example, it reads out an original image as shown in FIG. 5A. Next, the CPU 30 performs vehicle recognition on the original image using the object detection model (YOLO), and outputs the coordinates of the bounding box (surrounding the recognized vehicle) and the estimated vehicle type (step S54). That is, it outputs the region of the classified object.

[0075] Next, the CPU 30 reads out the analysis data (FIG. 5C), which is the coordinates of the bounding box and the vehicle type information recorded in association with the original image (step S58). Subsequently, the CPU 30 uses the analysis data of FIGS. 5A and 5C as teacher data, and learns the parameters for the recognition in step S54 based on the result of vehicle recognition (step S60).

[0076] In the above manner, an object detection model, which is a learned inference model, is generated. Initially, it is an unlearned or under-learned inference model, but by repeating the above process, a sufficiently learned inference model can be obtained.

[0077] Note that in the above, not only is the vehicle recognized, but its vehicle type (passenger car, truck, bus, etc.) is also estimated. With this learned object detection model (YOLO), it is not theoretically impossible to estimate not only the vehicle type but also large vehicles and small vehicles. However, unlike the estimation of the vehicle type, accurately estimating large and small vehicles is difficult with estimation based on the captured image.

[0078] Note that in this embodiment, not only photos from the front of the vehicle but also photos from the rear and side are used as training data. Therefore, an inference model can be generated that can perform inference on vehicles imaged from any direction. If it is only necessary to judge vehicles from a specific direction such as the front during operation, the training data can be created based on photos from only a specific direction such as the front.

[0079] Returning to the flowchart of FIG. 3, in step S4, the CPU 30 uses the object detection model learned as described above to recognize the vehicle depicted in the captured image (FIG. 4A), and calculates the bounding box and vehicle type shown in FIG. 4B.

[0080] Note that in this embodiment, vehicle recognition is performed based on a color captured image, but vehicle recognition may be performed based on a grayscale image (black and white image).

[0081] FIG. 4B shows an image in which a vehicle is recognized and surrounded by a bounding box. Also, although not shown, the vehicle type estimated in association with each bounding box is recorded (recorded in a format such as that in FIG. 3C).

[0082] Subsequently, the CPU 30 cuts out the image of each vehicle based on the bounding box (step S8). An example of the cut-out vehicle image is shown in FIG. 7.

[0083] The CPU 30 performs day-night conversion processing based on the cut-out vehicle image (step S10). Note that the size of the cut-out image varies depending on the distance to the vehicle. Therefore, in this embodiment, vehicle images of different sizes according to the distance are converted (normalized) to a unified size before estimation. For example, the entire vehicle image is evenly enlarged or reduced so that the horizontal width of the vehicle image (the horizontal width of the bounding box) becomes a determined width.

[0084] Here, the day-night conversion processing is processing for converting a vehicle image captured at night to be like a vehicle image captured during the day. This enables discrimination between large vehicles and small vehicles, which was difficult in vehicle images captured at night. Note that vehicle images captured during the day are output without performing style conversion (suppressing style conversion as much as possible) as daytime images.

[0085] Note that although the shape of the vehicle may be slightly deformed by the day-night conversion process, since the purpose of this device is to distinguish between large vehicles and small vehicles, deformations to such an extent that they do not affect this estimation result do not need to be considered a problem.

[0086] In this embodiment, a pre-trained style conversion model is used for the day-night conversion process. In this embodiment, CycleGAN is used as the style conversion model, but other models such as GAN may also be used.

[0087] The configuration of the style conversion model by CycleGAN is shown in FIG. 8. The day-night conversion model 70 is trained to convert a given daytime vehicle image (bright vehicle image) into a nighttime vehicle image (dark vehicle image). On the other hand, the night-day conversion model 72 is trained to convert a given nighttime vehicle image into a daytime vehicle image.

[0088] In this embodiment, the pre-trained night-day conversion model 72 is used for the day-night conversion process. The first discrimination model 74, the second discrimination model 76, and the night-day conversion model 70 are those used for training.

[0089] Hereinafter, the training of the style conversion model in FIG. 8 will be described. In the training process, first, a large number of nighttime vehicle images as shown in FIG. 9A and daytime vehicle images as shown in FIG. 9B are prepared and recorded in the SSD 36. Here, the daytime vehicle images and nighttime vehicle images to be prepared do not necessarily have to be images of the same vehicle.

[0090] When realizing the day-night conversion process by a CNN or the like without using an adversarial generation network such as CycleGAN, it is advisable to prepare pairs of images of the same vehicle taken at the same angle during the day and at night, and train so that when a nighttime vehicle image is input, a daytime vehicle image is output.

[0091] In this embodiment, since CycleGAN is used as the style conversion model, vehicle images with no corresponding relationship as shown in FIG. 9 can be used.

[0092] The learning process using the image in FIG. 9 is as follows. In FIG. 8, a predetermined number of daytime vehicle images (FIG. 9B) are given to the day-night conversion model 70, and image conversion is performed to generate a predetermined number of converted night vehicle images. At the time of this image generation, the day-night conversion model 70 is not learned.

[0093] Next, a predetermined number is selected from the prepared night vehicle images in FIG. 9A, and the original flag "1" is attached to them. On the other hand, the flag "0" indicating not original is attached to the above-mentioned generated predetermined number of converted night vehicle images. The night vehicle images and the converted night vehicle images obtained in this way are given to the second discrimination model 76, and learning is performed so that it can correctly discriminate whether the given image is original or not. Also, the day-night conversion model 72 is learned according to the correct / incorrect answers of the discrimination by the second discrimination model.

[0094] In the above manner, learning is performed so that the discrimination ability of the second discrimination model (the ability to distinguish between original and non-original images) and the conversion ability of the day-night conversion model (the ability to convert daytime images into nighttime images) are improved.

[0095] In the same manner as above, the day-night conversion model 72 and the first discrimination model 74 are also learned.

[0096] In this embodiment, the day-night conversion process in FIG. 3, step S10 is performed using the day-night conversion model 72 learned as described above. For example, when the cut-out vehicle image is as shown in FIG. 10A, it is converted into a daytime vehicle image (converted daytime vehicle image) as shown in FIG. 10B (day-night conversion).

[0097] Note that in this embodiment, a color vehicle image is input to obtain a color converted daytime vehicle image, but a grayscale image (black and white image) may be input to obtain a grayscale converted daytime vehicle image.

[0098] Note that when a daytime vehicle image is input to the night-day conversion model 72 that has been trained to generate a daytime vehicle image from a nighttime vehicle image, a daytime vehicle image with almost no change is output as a converted image (converted daytime vehicle image) (day-day conversion). This is presumably because no features of the nighttime image in the nighttime vehicle image are found in the daytime vehicle image, so no part to be converted is found and it is output with almost no change. Therefore, all vehicle images can be input to the night-day conversion model 72 for processing without distinguishing whether they are nighttime images or daytime images.

[0099] Next, based on the converted daytime vehicle images (including not only those night-day converted but also those day-day converted) obtained as described above, the CPU 30 estimates each part of the vehicle and generates a part-estimated vehicle image in which each part is colored separately (step S12).

[0100] In this embodiment, the SegNet, which is a semantic segmentation model, is used as the part-estimation model for generating a part-estimated vehicle image in which each part is colored separately. Note that other semantic segmentation models such as U-Net may be used instead of SegNet.

[0101] The learning of the part-estimation model is performed as follows. In the learning process, first, a large number of vehicle images as shown in FIGS. 11A and 11B are prepared and recorded in the SSD 36. Preferably, there are a large number of images with various vehicle types and backgrounds. In particular, it is preferable that the images well balance both small vehicles (such as passenger cars and small trucks) and large vehicles (such as buses and ordinary trucks).

[0102] The location and angle for imaging the vehicle to obtain the image data for learning are preferably as close as possible to those in the actual operation. For example, if imaging and operation are performed at an angle as shown in FIG. 4A, it is preferable that the angle is close to this.

[0103] On the one hand, if a general-purpose system that is not restricted by the installation position of the camera is to be constructed, it is preferable to use vehicle images at various angles as learning data.

[0104] In creating the learning data, the vehicle images in FIGS. 11A and 12A are displayed on the display 34. While viewing the images, the user operates the keyboard / mouse 40 to label the front license plate, rear license plate, front tire, rear tire, left tire, right tire, front, windshield, left side (including the side glass), right side (including the side glass), rear, rear glass, top, and background (parts other than the automobile) with different colors. Examples of the labeled partial vehicle images are shown in FIGS. 11B and 12B. In this example, the front is painted red, the right side is painted orange, the left side is painted pink, and the front license plate is painted gray so that they can be distinguished from each other.

[0105] The generated partial vehicle images are recorded on the hard disk 36 in association with the original vehicle images. In this embodiment, when recording the label images color-coded by the user, they are recorded as images color-coded in grayscale and distinguishable individually.

[0106] In the above manner, a large number of vehicle images and the corresponding partial vehicle images will be recorded on the hard disk 36.

[0107] FIG. 13 shows a conceptual configuration diagram of the SegNet model used in this embodiment. The SegNet model is a model suitable for learning and use to output an image O indicating to which class (here, labels such as the front and left tire) each pixel of the input image I belongs.

[0108] The convolutional layer and the pooling layer extract the features of the image, and the image compressed by pooling is restored to the resolution of the original image by the unpooling layer. At this time, the pooling index used in pooling is transmitted to the corresponding unpooling layer to restore the information contained in the original image.

[0109] Based on the prepared vehicle images and partial vehicle images in FIGS. 11 and 12, a flowchart for learning the partial estimation model in FIG. 13 is shown in FIG. 15.

[0110] First, the CPU 30 acquires a vehicle image from the SSD 36 (step S62). For example, it reads vehicle images such as those in FIGS. 11A and 12A. Next, the CPU 30 gives the vehicle image to the partial estimation model (SegNet) to obtain the partial estimation vehicle image (step S64). Note that the vehicle image is given as RGB three-channel data as shown in FIG. 14. In contrast, the obtained partial estimation vehicle image is single-channel grayscale data.

[0111] Subsequently, the CPU 30 acquires the partial vehicle image prepared corresponding to the vehicle image (step S66). The CPU 30 learns the parameters such as the convolutional layer so that the partial estimation vehicle image matches the partial vehicle image (so that the error becomes small) (step S68).

[0112] The above processing is repeated for a predetermined number of vehicle images, and learning is performed so that a correct partial estimation image is output (steps S60 and S70). In this embodiment, learning is performed using a color vehicle image and a grayscale partial vehicle image.

[0113] In the above embodiment, learning is performed based on the image of each vehicle. However, learning may also be performed based on an image of a plurality of vehicles captured and the corresponding partial vehicle images.

[0114] Returning to FIG. 3, in step S12, the CPU 30 uses the part estimation model learned as described above to estimate the part areas of the converted vehicle image and generate a part-estimated vehicle image (step S12). For example, based on the converted vehicle image (color) shown in FIG. 16A, a part-estimated vehicle image (grayscale) as shown in FIG. 16B can be obtained. In FIG. 16B, the front, front license plate, windshield, left side, top, left front tire, and left rear tire are recognized and colored.

[0115] Note that in this embodiment, for trucks, the cargo bed is regarded as the vehicle body, and containers and equipment provided thereon are not treated as parts of the vehicle (background) (refer to the first image from the right in FIG. 11B and the second image from the right in FIG. 12B). This is because the number of such vehicles is not large enough to distinguish them by coloring, and there are many variations in the equipment installed on the cargo bed (there are many variations such as special vehicles like garbage collection trucks, concrete mixer trucks, and crane trucks), and there is a possibility that appropriate learning cannot be performed.

[0116] Note that in this embodiment, in step S10, after converting to a daytime image, the part estimation process is performed. Therefore, even for an image captured at night, a part-estimated vehicle image can be obtained with high accuracy.

[0117] Next, the CPU 30 estimates whether the vehicle is a large vehicle or a small vehicle based on the above part-estimated vehicle image using the large / small estimation model (step S14).

[0118] As described above, in this embodiment, in step S4, vehicle types such as passenger cars, freight cars, and buses are obtained. However, among freight cars, there are small freight cars (small trucks) and ordinary freight cars (ordinary trucks). In order to distinguish between them, large / small estimation is performed. The former is small, and the latter is large.

[0119] In addition, among passenger cars, there are cases where they may be misrecognized as large buses like wagons or minibuses due to their shape. Therefore, in this embodiment, in order to distinguish between them, large / small estimation is performed. The former is small, and the latter is large.

[0120] As described above, by distinguishing between large and small, together with the determination result of step S4, it is possible to obtain passenger cars (small cars), small trucks (small cars), buses (large cars), and ordinary trucks (large cars), which are classifications used in road traffic sensors and the like.

[0121] In this embodiment, a convolutional neural network model (CNN) is used as the large / small estimation model for estimating large and small vehicles. The learning of the large / small estimation model is performed as follows. In the learning process, first, a large number of partial vehicle images (which may also be partial estimated vehicle images) with flags indicating large and small vehicles are prepared and recorded in SSD36. Note that the partial vehicle images may be painted based on the vehicle image, may be painted based on the converted vehicle image, or may include both.

[0122] FIG. 17 shows a flowchart of the learning process. First, a large number of partial vehicle images with flags indicating large and small vehicles are prepared and recorded in SSD36. Such partial vehicle images can use those shown in FIGS. 11B and 12B.

[0123] When preparing the learning data, each partial vehicle image is displayed on the display 34, and while the operator views the image, the operator operates the keyboard / mouse 40 to input the distinction between large vehicles and small vehicles. Note that since it may be difficult to make a determination based only on the partial vehicle image, the original vehicle image is also displayed. As described above, passenger cars (including minibuses and station wagons) and small trucks are regarded as small vehicles, and buses and ordinary trucks are regarded as large vehicles. This distinction between large vehicles and small vehicles is recorded in the SSD 36 in association with the partial vehicle image. Note that the one shown in FIG. 11B is a small vehicle, and the one shown in FIG. 12B is a large vehicle. Note that the distinction between large vehicles and small vehicles may be input simultaneously when generating the partial estimated vehicle image.

[0124] As described above, a large number of partial vehicle images with flags indicating the distinction between large vehicles and small vehicles are recorded in the SSD 36.

[0125] In the learning process, the CPU 30 acquires a partial vehicle image from the SSD 36 (step S82). For example, it reads out partial vehicle images such as those in FIGS. 11B and 12B. Next, the CPU 30 performs estimation of large or small on the partial vehicle image by the size estimation engine (step S84).

[0126] Next, the CPU 30 reads out the distinction between large and small attached to the partial vehicle image (step S86). Subsequently, the CPU 30 uses the read distinction between large and small as teacher data and learns the parameters of the CNN based on the estimation result in step S84 (step S88).

[0127] When learning is performed based on a predetermined number of partial vehicle images, the CPU 30 ends the learning process (steps S80, S90).

[0128] Note that the reason why this estimation functions is presumably that even if the external shapes of large vehicles and small vehicles are similar, the ratio occupied by the windshield and the tire intervals before and after are different.

[0129] In the above-described embodiment, learning is performed based on an image of each vehicle. However, learning may be performed based on images of a plurality of vehicles.

[0130] In step S14 of FIG. 3, the CPU 30 uses the size estimation model learned as described above to estimate large or small based on the part estimation image. Next, the CPU 30 corrects the vehicle type estimation in step S4 based on the large / small estimation result (step S14).

[0131] In step S4, vehicle type estimation of a passenger car, a freight car, and a bus was performed. In this embodiment, in step S14, as shown in FIG. 18, based on the large / small estimation, the vehicle type estimation is corrected to be accurate.

[0132] The CPU 30 determines whether the vehicle in the vehicle image is the first one to be imaged (step S16). If it is the first one to be imaged, the result of the vehicle estimation regarding the vehicle is recorded (step S20).

[0133] On the other hand, if it has been imaged in the frames so far in the video, since the result of the vehicle estimation has already been recorded, the current estimation result is added to it. Note that whether it is the same vehicle can be determined by the similarity of the vehicle images.

[0134] As described above, for the vehicle extracted in step S8, vehicle type estimation including whether it is a large vehicle or a small vehicle can be performed.

[0135] The CPU 30 repeats the above estimation for the number of vehicle images (steps S6, S22). Thereby, vehicle type estimation can be performed for each vehicle imaged in the image captured by the camera.

[0136] Next, the CPU 30 captures the next captured image (step S202), and performs vehicle type estimation for each vehicle in the same manner as above, and records the result (steps S4 to S22).

[0137] By repeating the above, for each vehicle, a plurality of vehicle type estimation results can be obtained. The CPU 30 integrates these vehicle type estimation results and determines them as the final vehicle type estimation result for the vehicle (step S24). For example, among the multiple vehicle type estimation results, the most frequent vehicle type estimation result is set as the final vehicle type estimation result.

[0138] In the above manner, it is possible to accurately determine the vehicle type including at least the determination of large vehicles and small vehicles while performing real-time imaging.

[0139] 1.4 Others (1) In the above embodiment, in step S14, the estimation of large and small is performed based on the part estimation image. However, it may also be used for estimation including the converted vehicle image generated in step S10 (or the vehicle image extracted in step S8). Since the converted vehicle image (vehicle image) surrounded by the bounding box also includes the background at the time of imaging, the information from the background image is also used for the estimation of large and small, and the estimation accuracy is improved.

[0140] In this case, the learning of the large / small estimation model will be performed using both the part vehicle image and the converted vehicle image (vehicle image). When using grayscale for the part vehicle image and color for the converted vehicle image (vehicle image), as the image data given to the large / small estimation model, the converted vehicle image (vehicle image) may be given in 1 to 3 channels of RGB, and the part vehicle image (part estimation vehicle image) may be given in 4 channels.

[0141] That is, as shown in FIG. 19, the grayscale data of the part vehicle image (part estimation vehicle image) may be added as 4 channels at the corresponding positions of the RGB data of the color.

[0142] In addition, in the above-described modification, the vehicle image used for learning and estimation is a color image. However, a converted vehicle image (vehicle image) that has been grayscale-converted may be used for learning and estimation. By performing grayscale conversion, it is possible to reduce false estimation caused by color.

[0143] (2) In the above-described embodiment, when converting the night vehicle image into the day vehicle image in step S10, the day vehicle image is also provided to the conversion model to obtain the night-day conversion image. However, as shown in the flowchart of FIG. 20, it may be determined whether the vehicle image is a night image, and conversion may be performed only when it is a night image, and it may be used without conversion when it is a day image.

[0144] Note that the determination as to whether the vehicle image in step S101 is a night vehicle image can be made based on the magnitude of the dynamic range of the vehicle image (the dynamic range of a night image is small) or the imaging time.

[0145] (3) In the above-described embodiment, night-day conversion is performed in step S10. However, instead of this, or in addition to this, a conversion model for converting the color of the vehicle body may be used.

[0146] For example, a vehicle image of a vehicle body in blue as shown in FIG. 21A is converted into a vehicle image of a vehicle body in white as shown in FIG. 21B to obtain a converted vehicle image.

[0147] In this case, a CycleGAN model as shown in FIG. 8 can be learned and used. For learning, a vehicle image with a colored vehicle body and a vehicle image with a white vehicle body can be used to perform learning in the same manner as in the case of night-day conversion.

[0148] By making the vehicle body white, the generation accuracy of the part-estimated vehicle image is improved. In addition, by making the vehicle body white, when this is used for size estimation, false estimation caused by color (or density when grayscale-converted) can be eliminated. Further, by converting to a white vehicle body, there is no reflection as a result, and false determination is less likely to occur.

[0149] When a vehicle image with a white vehicle body is input into the learned vehicle body color conversion model, a vehicle image that is also white is output. Therefore, regardless of the color of the vehicle body, conversion can be performed by the vehicle body color conversion model and processing can be carried out.

[0150] Also, in the same way as in the case of night-day conversion, it can be determined whether the color of the vehicle body is white. If it is a color other than white, conversion is performed by the vehicle body color conversion model, and in the case of other colors, conversion by the vehicle body color conversion model is not performed and it is output as it is.

[0151] Note that whether the vehicle body is white can be tentatively determined by performing part estimation and determining whether the parts determined to be the side or the front are white.

[0152] As for which color is preferably converted, it may be determined in consideration of factors such as being easy to perform part estimation and size estimation in relation to the background and the like.

[0153] Also, a vehicle body with a pattern may be converted into a single-color vehicle body without a pattern.

[0154] (4) In the above embodiment, in step S10, night-day conversion is performed. However, instead of this, or in addition to this, a conversion model that converts to eliminate reflections on the vehicle body or glass may be used.

[0155] For example, an image with reflections on the front glass or the front as shown in Fig. 22A can be converted into an image without reflections as shown in Fig. 22B and output as a converted vehicle image.

[0156] In this case, a CycleGAN model as shown in Fig. 8 can be learned and used. For learning, a vehicle image with reflections and a vehicle image without reflections can be used to perform learning in the same way as in the case of night-day conversion.

[0157] By obtaining a vehicle image without reflections, the generation accuracy of the part-estimated vehicle image can be improved. Further, by using a vehicle image without reflections, when this is used for size estimation, erroneous estimation caused by reflections can be eliminated.

[0158] When a vehicle image without reflections is input to a learned reflection removal conversion model, a vehicle image without reflections is also output. Therefore, processing can be performed by performing conversion using the reflection removal conversion model regardless of the presence or absence of reflections.

[0159] Also, similar to the case of day-night conversion, it is possible to determine whether there are reflections, and if there are reflections, perform conversion using the reflection removal conversion model, and if there are no reflections, not perform conversion using the reflection removal conversion model and output the image as it is. Note that it is assumed that there are reflections during a predetermined time period in the middle of the day on a sunny day, and for vehicle images captured during this time period, conversion using the reflection removal conversion model is performed, and in other cases, conversion using the reflection removal conversion model is not performed and the image is output as it is.

[0160] Note that whether the vehicle body is white can be determined by tentatively performing part estimation and finding that the density of the parts determined to be the side or the front is not uniform.

[0161] (5) In the above embodiment, in step S10, day-night conversion is performed. However, instead of this, or in addition to this, when a part of the vehicle body is missing, a conversion model that performs conversion to fill in the missing part may be used. Such a defect occurs when a vehicle comes to the edge of the camera's imaging angle or when vehicles overlap.

[0162] For example, an image with a missing side or top as shown in Fig. 23A can be converted into an image without such a defect as shown in Fig. 23A and output as a converted vehicle image.

[0163] In this case, a CycleGAN model as shown in FIG. 8 can be trained and used. For training, a defective vehicle image and a non-defective vehicle image can be used to perform training in the same way as in the case of night-day conversion.

[0164] By using a non-defective vehicle image, the generation accuracy of the partial estimation vehicle image can be improved. Also, by using a non-defective vehicle image, when this is used for size estimation, misestimation due to defects can be eliminated.

[0165] When a non-defective vehicle image is input to the trained defect completion conversion model, a non-defective vehicle image is also output. Therefore, regardless of the presence or absence of defects, processing can be performed by performing conversion using the defect completion conversion model.

[0166] Also, similar to the case of night-day conversion, it is possible to determine whether there is a defect, and if there is a defect, perform conversion using the defect completion conversion model, and if there is no defect, not perform conversion using the defect completion conversion model and output it as it is.

[0167] Note that whether there is a defect can be tentatively determined by performing partial estimation, and if any part is in contact with the image edge over a predetermined length or more, it can be determined that there is a defect. Also, in step S4, if the bounding boxes overlap, it can be determined that there is a defect.

[0168] Also, two or more of the above-described examples of image conversion can be combined and implemented.

[0169] (6) In the above embodiment, estimation is performed after normalizing the size of the vehicle image. However, estimation may be performed without normalization.

[0170] (7) In the above embodiment, first, the vehicle type is estimated, then the large / small size is estimated, corrected, and the final vehicle type is obtained. However, in step S4, the final vehicle type may be estimated at once using the partial estimation image of the vehicle.

[0171] (8) In the above embodiment, the vehicle type is estimated in real time. However, the above processing may be performed based on the recorded captured image to estimate the vehicle type. In this case, the image captured by the camera 2 may be recorded on a portable recording medium and read into the SSD 36 for processing.

[0172] (9) In the above embodiment, as the part estimation image, one in which part regions such as the front, front license plate, side surface, tire, and windshield are clarified is used. However, in the case of large vehicles and small vehicles, a part estimation image including a part where the overall size can be known and a part where the size of the windshield, tire interval, or license plate can be known can be used. For example, a part estimation image including at least the windshield and one of the side surfaces may be used. Also, a part estimation image including at least one of the front tires and one of the rear tires and one of the side surfaces may be used.

[0173] (10) In the above embodiment, a common learned inference model for all vehicles is constructed to estimate large vehicles and small vehicles. However, a learned inference model may be constructed for each vehicle type estimated in step S4, and large vehicles and small vehicles may be estimated for each vehicle type. Also, the vehicle type may be determined.

[0174] (11) In the above embodiment, large vehicles and small vehicles are inferred based on the part estimation image. However, instead of this, or in addition to this, the size of the bounding box frame when the vehicle image is normalized, the area occupied by the vehicle, etc. may be used as the basis for inference.

[0175] (12) In the above embodiment, vehicles moving in various directions are imaged. However, at the time of installing the camera, if it is set so that only vehicles moving in one direction (such as only vehicles coming towards) are imaged, the estimation process becomes easy and the accuracy is also improved.

[0176] (13) In the above embodiment, the determination of large or small is made based on a plurality of vehicle images of the same vehicle. However, this may also be done based on a single image.

[0177] (14) In the above embodiment, as shown in Fig. 1a, the device is configured. However, as shown in Fig. 1b, a converted image may be generated for the entire captured image, and a vehicle image may be extracted based on the converted image.

[0178] Also, the device may be configured by the imaging unit 2, the converted image generation means 6, and the vehicle image extraction means 4 (including the vehicle type estimation process). Thereby, even for a captured image at night or the like, it becomes possible to estimate the vehicle type. In this case, the converted image generation means 6 generates a converted image suitable for the process by the vehicle image extraction means 4.

[0179] (15) In the above embodiment, it is configured as a stand-alone device, but it may also be configured as a server device provided with the vehicle image extraction means 4, the converted image generation means 6, the part estimation vehicle image generation means 8, and the size estimation means 10. In this case, the captured image is transmitted from the terminal device to the server device.

[0180] (16) The above embodiment and its modification examples can be implemented in combination with other embodiments and modification examples as long as they do not conflict with their essence.

[0181] 2. Second Embodiment 2.1 Functional Configuration Fig. 24a shows the functional configuration of the vehicle estimation device according to the second embodiment of the present invention. In this embodiment, in addition to size estimation, it is configured as a device for measuring traffic volume.

[0182] The imaging unit 2, the vehicle image extraction means 4, the converted image generation means 6, the part estimation vehicle image generation means 8, and the size estimation means 10 have the same configuration as in the first embodiment.

[0183] The vehicle passage determination means 54 acquires an imaging image surrounded by the boundary box 60 of the vehicle from the vehicle image extraction means 4. The vehicle passage determination means 54 determines whether the boundary box 60 straddles a passing line 62 set so as to intersect with the movement locus of the automobile in the continuously captured images, and determines whether the vehicle has passed.

[0184] By determining the vehicles that have passed by the vehicle passage determination means 54 and counting the number thereof, the traffic volume can be measured. Since the presence or absence of passage is determined using the passing line 62 and the boundary box 60, the passage of the automobile can be accurately determined.

[0185] Also, since information on whether the vehicle for which passage has been determined is a large vehicle or a small vehicle is obtained from the large / small estimation means 10, the traffic volume can be measured including that information.

[0186] 2.2 Hardware Configuration The hardware configuration is the same as that shown in FIG. 2 in the first embodiment.

[0187] 3.3 Traffic Volume Measurement Processing FIGS. 25 to 27 show a flowchart of a vehicle estimation program 46 having a traffic volume measurement processing function. The CPU 30 acquires an imaging image by the camera 2 and records it in the SSD 36. In this embodiment, imaging is performed as a video. FIG. 4A shows an example of the imaging image. In this embodiment, the camera 2 is provided on a utility pole or in the street, and imaging is performed at a fixed angle from above.

[0188] The CPU 30 acquires one frame of the imaging image recorded in the SSD 36 (step S202). In this embodiment, the imaging image is read out and processed in real time when it is recorded. However, the imaging image may be once recorded and then the traffic volume measurement processing may be performed later.

[0189] The CPU 30 recognizes the vehicles included in the captured image of one frame, and as shown in FIG. 4B, generates a bounding box (rectangular area) in which the vehicle is inscribed (step S204). This process can use a learned object detection model in the same manner as in the first embodiment.

[0190] Next, the CPU 30 assigns a vehicle ID to each bounding box recognized as shown in FIG. 4B (steps S206 to S216). Here, the vehicle ID is an ID for specifying the vehicle captured in the captured image. In this embodiment, the same vehicle ID is assigned to the same vehicle even in images of different frames. Hereinafter, the assignment of the vehicle ID will be described in detail.

[0191] The CPU 30 selects one of the bounding boxes of the vehicles in the captured image as the target bounding box (step S208). The CPU 30 determines whether a vehicle similar to the vehicle in this bounding box (for example, determining the similarity based on the image feature amount) was in a position in the reverse (rear) direction with respect to the vehicle traveling direction on the screen in the previous frame (step S210). If so, the vehicle ID given to the vehicle in the previous frame is inherited and given to the vehicle in this frame (step S212). If not, it is determined that the vehicle has been imaged for the first time, and a new vehicle ID is assigned to the vehicle (step S214).

[0192] Note that due to noise images or the like, it may be determined that there are no similar vehicles even though they appeared in the previous frame. In this case, it is not appropriate to assign a new vehicle ID. Therefore, in this embodiment, for the bounding box that completely exceeds the line 200 in FIG. 4A, the process is always performed as if the same vehicle appeared in the previous frame. Specifically, considering the average distance that the vehicle moves in one frame, the vehicle existing at that position in the previous frame is regarded as the same vehicle.

[0193] When the CPU 30 performs the above processing on one boundary box, it performs the same processing on the next boundary box. By doing this for all boundary boxes, vehicle IDs can be assigned to all boundary boxes in the captured image of one frame.

[0194] For example, as shown in Fig. 28A, assume that vehicle IDs (C1 to C4) are assigned to all boundary boxes (vehicle images are omitted) in the captured image of one frame.

[0195] The CPU 30 focuses on one of these boundary boxes (step S220). For example, assume that it focuses on the boundary box with vehicle ID = C4. The CPU 30 calculates the positions of the center of gravity and the center point of the bottom side of boundary box C4 (step S222). In the figure, this is indicated by a point.

[0196] The CPU 30 determines whether this boundary box C4 has passed the first passing line 62a (step S224). Here, the first passing line 62a is preset by the keyboard / mouse 40 while the operator is displaying the captured image on the display 34. The same applies to the second passing line 62b.

[0197] The first passing line 62a is provided upstream of the second passing line 62b in the passing direction of the vehicle.

[0198] In this embodiment, the CPU 30 determines that the first passing line 62a has been passed when two conditions are met: 1) the center of gravity of the boundary box is on the downstream side of the first passing line 62a, and 2) in the previous frame, the center of gravity of the boundary box with the same vehicle ID was on the upstream side of the first passing line 62a.

[0199] In Fig. 28A, since the boundary box C4 is on the upstream side of the first passing line 62a, it is not determined that the first passing line 62a has been passed.

[0200] Therefore, since the first passing flag remains down (the first passing flag is down in the initial state), after step S228, the processing for the boundary box C4 ends.

[0201] In the same manner as described above, the CPU 30 also performs the processing of steps S218 to S236 for the other boundary boxes C1, C2, and C3. When the processing for all the target vehicle IDs is completed, the CPU 30 acquires the captured image of the next frame from the SSD 36 (step S202).

[0202] The CPU 30 recognizes the vehicles in the acquired captured image and generates boundary boxes (step S204). Further, vehicle IDs are assigned to these boundary boxes (steps S206 to S216). This state is shown in FIG. 28B.

[0203] Subsequently, for each boundary box in this frame, it is determined whether the first passing line 62a and the second passing line 62b have been passed. For example, the center of gravity of the boundary box C4 is on the downstream side of the first passing line 62a, and in the image of the previous frame (FIG. 28A), it was on the upstream side of the first passing line 62a. Therefore, the CPU 30 determines that the boundary box C4 has passed the first passing line 62a and sets the first passing flag 62a (step S226).

[0204] Subsequently, the CPU 30 determines whether the center point of the bottom side of the boundary box C4 has passed the second passing line 62b (step S230). The CPU 30 determines that the second passing line 62b has been passed when two conditions are satisfied: 1) the center point of the bottom side of the boundary box is on the downstream side of the second passing line 62b, and 2) in the previous frame, the center point of the bottom side of the boundary box with the same vehicle ID was on the upstream side of the second passing line 62b.

[0205] In the state of Fig. 28B, since the center point of the bottom side of the boundary box C4 is upstream of the second passing line 62b, it is determined that it has not passed through the second passing line 62b.

[0206] Therefore, since the second passing flag remains down (the second passing flag is down in the initial state), after step S230, the processing for the boundary box C4 ends.

[0207] The CPU 30 performs the processing of steps S218 to S236 for the other boundary boxes C1, C2, and C3 in the same manner as above. After finishing the processing for all the target vehicle IDs, the CPU 30 acquires the captured image of the next frame from the SSD 36 (step S202).

[0208] In this way, the frame images of Fig. 28C, Fig. 28D, Fig. 28E, and Fig. 28F are sequentially processed. The boundary box C4 passes through the second passing line 62b in the frame image of Fig. 28F. Therefore, the CPU 30 proceeds from step S230 to S232 and sets the second passing flag.

[0209] As a result, since both the first passing flag and the second passing flag are set, the CPU 30 determines that the vehicle corresponding to the boundary box C4 is a passing vehicle (step S238). Then, the passing count is incremented and recorded.

[0210] Note that the vehicle in the boundary box C4 whose passing has been confirmed and counted in this way does not need to be the target of subsequent processing, so it is removed from the processing target vehicles targeted by the processing of steps S218 to S236.

[0211] Note that in this embodiment, vehicles in the oncoming lane that pass through the second passing line 62b first and then pass through the first passing line 62a are not counted as passing vehicles.

[0212] As described above, for the vehicle that has passed through the second passing line 62b after passing through the first passing line 62a, counting can be performed as a passing vehicle.

[0213] Subsequently, the CPU 30 estimates the vehicle type including large and small based on the vehicle image. The flowchart thereof is shown in FIG. 27. This process is the same as that of the first embodiment. However, in this embodiment, since a vehicle ID is attached, the estimation result is recorded using this.

[0214] Therefore, the CPU 30 can count the traffic volume including the classification of large and small using the estimation results of large vehicles and small vehicles. For example, the passing volume can be counted for each vehicle type as shown in FIG. 18.

[0215] 2.4 Others (1) In the above embodiment, for the vehicle that has passed through the second passing line 62b after passing through the first passing line 62a, counting is performed as a passing vehicle. However, for the vehicle that has passed through the first passing line 62a after passing through the second passing line 62b, counting may be performed as a passing vehicle in the opposite direction.

[0216] (2) In the above embodiment, for the first passing line 62a, the presence or absence of passing is determined by the center of gravity, and for the second passing line 62b, it is determined by the center point of the bottom side. However, for both, the determination may be made using the center of gravity or the center point of the bottom side.

[0217] In the passing line on the side farther from the camera in the captured image (line 62a in FIG. 28), it is preferable to determine passing by the center position of the bounding box (such as the center of gravity, the center, a point near the center, etc.). This is because passing can be reliably determined.

[0218] Also, in the passing line on the side closer to the camera in the captured image (line 62b in FIG. 28), it is preferable to determine passing by a point on the bottom side of the bounding box (not limited to the center point). This is because the screen can be effectively utilized to prevent passing detection omission.

[0219] (3) In the above embodiment, the passage is determined using two passing lines. However, three or more passing lines may be used, and when these are passed through in order, it may be determined that the vehicle has passed. Also, only one passing line may be provided, and when the vehicle has passed through this passing line, it may be determined that the vehicle has passed. In this case, it may be determined from the position of the vehicle in the time-series frames in which direction the vehicle has passed.

[0220] (4) In the above embodiment, the traffic volume of a lane in which the vehicle moves in one direction is measured. However, as shown in FIG. 29, an intersection is used as a captured image, and passing lines 62a to 62h (in the figure, a plurality of passing lines are provided for each road, but one may also be sufficient) are provided at the entrance (exit) of the intersection, and by counting including the passing direction, the traffic volume of the intersection can be grasped in detail.

[0221] As a result, it is possible to count from which direction and in which direction the vehicle has come, such as from I to IV, I to VI, I to VIII, I to II (U-turn), etc. (the same applies to II and below).

[0222] Based on this count data, a signal control system as shown in FIG. 30 can be constructed. Based on the traffic volume of the intersection measured by the traffic volume measurement device (for example, the traffic volume 30 minutes ago), the signal control device controls the green time, red time, green arrow time, etc. of the signal. The traffic volume measurement device may also serve as the signal control device.

[0223] For example, the signal control device compares the total value of the vehicle passing volume from I to VI and the vehicle passing volume from V to II with the total value of the vehicle passing volume from III to VIII and the vehicle passing volume from VII to IV, and controls the ratio of the green time of signal machine 250 (256) to the green time of signal machine 252 (254) according to the ratio. Also, according to the ratio of the vehicle passing volume from I to VI to the vehicle passing volume from I to IV, the ratio of the green time to the green arrow time of signal machine 250 is controlled.

[0224] Also, the traffic volume may be counted on the road before entering the intersection, the passing volume at the future intersection may be predicted, and the traffic signal may be controlled based on this prediction.

[0225] (6) The vehicle passing volume counted according to the above embodiment can be used for the design of the length of the right-turn lane at the intersection. For example, at the intersection in FIG. 30, the length of the right-turn lane provided at I can be calculated based on the vehicle passing volume from I to IV.

[0226] For example, traffic volume information at such an intersection may be recorded in a server device, this information may be acquired by road design software installed in a terminal device, and based on this, the length of the right-turn lane may be proposed and presented to a designer.

[0227] (7) Based on the vehicle passing volume counted according to the above embodiment, when the cumulative road passing volume exceeds a threshold value, a warning indicating that it is time for road repair may be output.

[0228] (8) In the above embodiment, the apparatus is configured as shown in FIG. 24a. However, in the same manner as the modification in the first embodiment (see FIG. 1b), a converted image may be generated for the entire captured image, and then the vehicle image may be extracted.

[0229] Also, as shown in FIG. 24b, without performing size estimation, a vehicle image may be extracted based on the converted image, and the passing of the vehicle may be determined.

[0230] (9) In the above embodiment, the passing line 62 is used to determine the passing of the vehicle. However, passing may be determined by other methods (such as whether the bounding box 60 has disappeared from the screen).

[0231] (10) In the above-described embodiment, it is configured as a stand-alone device, but it may also be configured as a server device including a vehicle image extraction means 4, a converted image generation means 6, a part-estimated vehicle image generation means 8, a size estimation means 10, and a vehicle passage determination means 54. In this case, the captured image is transmitted from the terminal device to the server device.

[0232] (11) The above-described embodiment and its modifications can be implemented in combination with other embodiments and modifications as long as they do not conflict with the essence thereof.

[0233] 3. Third Embodiment 3.1 Functional Configuration FIG. 31 shows a functional block diagram of an admission management system according to an embodiment of the present invention. The imaging unit 2 is provided, for example, at the entrance of a parking lot such as a facility, and captures an image of a vehicle approaching the entrance gate. The vehicle type estimation and counting means 100 estimates the vehicle type based on the captured image by, for example, the method described in the second embodiment, and counts the number of passing vehicles for each vehicle type. The gate control means 150 controls the opening and closing of the gate based on the estimated vehicle type. For example, depending on whether it is a large vehicle or a small vehicle, it controls whether to open the gate to the parking lot for large vehicles and the gate to the parking lot for small vehicles.

[0234] 3.2 System Configuration and Operation FIG. 32 shows the appearance of the entrance gate of the admission management system. A gate 120 for small vehicles and a gate 140 for large vehicles are provided. The camera 2 captures an image of a vehicle attempting to enter these gates.

[0235] FIG. 33 shows the hardware configuration. The computer 160 is the same as the configuration of FIG. 2 in the first embodiment. However, the CPU 30 is capable of giving commands to the gate control unit 180 that controls the opening and closing of the gates 120 and 140.

[0236] Figure 34 shows a flowchart of the control program. In step S110, the CPU 30 acquires an image captured by the camera 2. If the vehicle is not recognized in the captured image, step S110 is repeated.

[0237] When the vehicle is recognized, the CPU 30 estimates the vehicle type (large or small) by the process described in the first embodiment (step S114). If it is estimated to be large, the CPU 30 instructs the gate control unit 180 to open the large vehicle gate 140 (step S116). Also, if it is estimated to be small, the CPU 30 instructs the gate control unit 180 to open the small vehicle gate 120 (step S116).

[0238] After opening the gate, when receiving the output from the passing detection sensor (not shown), the CPU 30 closes the gate (step S118).

[0239] As described above, it is possible to automatically determine whether it is a large or small vehicle, select the destination gate, and open and close it. Also, since the number of large and small vehicles entering is counted, it is possible to display on a display (not shown) that the large vehicle area or the small vehicle area is full.

[0240] 3.3 Others (1) In the above embodiment, the parking lot has been described, but it can be similarly applied to other facilities such as drive-through safaris. Also, in the above, the gate is selectively opened and closed based on the determination of large or small, but the gate may be selectively opened and closed according to a more detailed vehicle type classification.

[0241] Also, it may be determined whether to open the gate according to a predetermined vehicle type (or according to large or small). For example, for a bridge dedicated to small vehicles, it is possible to control not to open the gate for large vehicles at its entrance.

[0242] (2) In the above embodiment, the case of controlling the opening and closing of the gate has been described. However, depending on vehicle type estimation (including estimation of large and small vehicles), parking fees and facility usage fees may be calculated. This can be used at the entrance of a highway or the like.

[0243] (3) The above embodiment and its modifications can be implemented in combination with other embodiments and modifications as long as they do not conflict with their essence.

Claims

1. Vehicle image extraction means for recognizing a vehicle in a captured image of a traveling vehicle at an angle capable of imaging the front and rear tires and the sides, and extracting a vehicle image by a bounding box surrounding the vehicle; Conversion image generation means for outputting a converted vehicle image obtained by converting the vehicle itself in the vehicle image into an image suitable for extracting parts of the vehicle based on the vehicle image; Part-estimated vehicle image generation means for generating a part-estimated vehicle image in which part regions including the front and rear tires and the sides of the vehicle are painted separately based on the converted vehicle image; Size estimation means for estimating whether the vehicle is a large vehicle or a small vehicle based on the part-estimated vehicle image; In a vehicle estimation device comprising: The conversion image generation means generates a converted vehicle image using a learned model trained to convert into a converted vehicle image suitable for extracting parts of the vehicle. The part-estimated vehicle image generation means generates a part-estimated vehicle image using a learned model trained based on an image of the vehicle and a part vehicle image in which part regions including at least the front and rear tires and the sides of the vehicle are painted separately in different colors or densities so as to be distinguishable. The size estimation means is configured to estimate whether the vehicle is a large vehicle or a small vehicle using a learned model trained to estimate whether the vehicle is a large vehicle or a small vehicle based on the part-estimated vehicle image and a determination of whether it is a large vehicle or a small vehicle. A vehicle estimation device characterized by that.

2. A vehicle estimation program for realizing a vehicle estimation device by a computer, the computer being Vehicle image extraction means for recognizing a vehicle in a captured image of a traveling vehicle at an angle capable of imaging the front and rear tires and the sides, and extracting a vehicle image by a bounding box surrounding the vehicle; Conversion image generation means for outputting a converted vehicle image obtained by converting the vehicle itself in the vehicle image into an image suitable for extracting parts of the vehicle based on the vehicle image; Based on the converted vehicle image, there is a site-estimated vehicle image generation means for generating a site-estimated vehicle image in which site areas including the front and rear tires and the side surfaces of the vehicle are painted separately, In a vehicle estimation program for functioning as a size estimation means for estimating whether the vehicle is a large vehicle or a small vehicle based on the site-estimated vehicle image, The converted image generation means generates a converted vehicle image using a learned model that has been learned to convert the vehicle image into a converted vehicle image suitable for extracting the site of the vehicle, The site-estimated vehicle image generation means generates a site-estimated vehicle image using a learned model that has been learned based on the vehicle image and a site vehicle image in which each site area including at least the front and rear tires and the side surfaces of the vehicle is painted separately with different colors or densities so as to be distinguishable, The size estimation means is configured to estimate whether the vehicle is a large vehicle or a small vehicle using a learned model that has been learned to estimate whether the vehicle is a large vehicle or a small vehicle based on the site-estimated vehicle image and the determination of whether it is a large vehicle or a small vehicle. A vehicle estimation program characterized by that.

3. In the apparatus of claim 1 or the program of claim 2, The converted image generation means uses a learned conversion model that has been learned to receive at least the vehicle image captured at night and output a converted vehicle image as if it were captured during the day. An apparatus or program characterized by that.

4. In the apparatus of claim 3 or the program of claim 3, The converted image generation means determines whether the vehicle image was captured at night based on the imaging time of the vehicle image or the data content of the vehicle image, and uses the learned conversion model for the vehicle image determined to be captured at night as a converted vehicle image, and outputs the vehicle image not determined to be captured at night as a converted vehicle image as it is. An apparatus or program characterized by that.

5. In any of the apparatuses or programs of claims 1 to 4, The conversion image generation means uses a learned conversion model that is trained to receive the vehicle image and output a converted vehicle image with the color of the vehicle body changed to white. The device or program is characterized by this. **Claim 6** In the device or program according to claim 5, the conversion image generation means determines whether the color of the vehicle body in the vehicle image is different from white based on the vehicle image, uses the learned conversion model for the vehicle image determined to be different, and outputs the vehicle image determined to be the same color as the converted vehicle image as it is. The device or program is characterized by this. **Claim 7** In the device or program according to any one of claims 1 to 6, the conversion image generation means uses a learned conversion model that is trained to receive the vehicle image and output a converted vehicle image with reduced reflection on the vehicle body. The device or program is characterized by this. **Claim 8** In the device or program according to claim 7, the conversion image generation means determines whether there is reflection on the vehicle body in the vehicle image based on the imaging time of the vehicle image or the vehicle image data, uses the learned conversion model for the vehicle image determined to have reflection, and outputs the vehicle image determined not to have reflection as the converted vehicle image as it is. The device or program is characterized by this. **Claim 9** In the device or program according to any one of claims 1 to 8, the conversion image generation means uses a learned conversion model that is trained to receive at least the vehicle image with a missing part and output a converted vehicle image with the missing part complemented. The device or program is characterized by this. **Claim 10** In the device or program according to claim 9, The conversion image generation means determines whether there is a missing part in the vehicle image. If it is determined that there is a missing part, a converted vehicle image is output using the learned conversion model. If it is determined that there is no missing part, the vehicle image is output as the converted vehicle image as it is. The apparatus or program is characterized by this.

11. In the apparatus according to any one of Claims 1, 3 to 10, further comprising grayscale conversion means for converting the vehicle image or the converted vehicle image into grayscale to output a grayscale vehicle image or a grayscale converted vehicle image, The size estimation means estimates whether the vehicle is a large vehicle or a small vehicle based on the part-estimated vehicle image and the grayscale vehicle image or the grayscale converted vehicle image. The apparatus is characterized by this.

12. In the program according to any one of Claims 2 to 10, the program further causes a computer to function as grayscale conversion means for converting the vehicle image or the converted vehicle image into grayscale to output a grayscale vehicle image or a grayscale converted vehicle image, The size estimation means estimates whether the vehicle is a large vehicle or a small vehicle based on the part-estimated vehicle image and the grayscale vehicle image or the grayscale converted vehicle image. The program is characterized by this.

13. In the apparatus or program according to any one of Claims 1 to 12, The size estimation means outputs a vehicle type including the distinction between the estimated large vehicle and small vehicle. The apparatus or program is characterized by this.

14. In the apparatus according to any one of Claims 1, 3 to 13, further comprising vehicle passage determination means for detecting that the bounding box has straddled the passing line in the consecutive captured images based on the passing line set in the captured image and determining the passage of the vehicle.

15. In any of the programs according to claims 2 to 13, the program further causes a computer to function as vehicle passage determination means for determining passage of a vehicle by detecting that the bounding box straddles a passing line set in the captured image in consecutive captured images based on the passing line.

16. In the apparatus or program according to claim 14 or 15, two passing lines are set, and the vehicle passage determination means determines that the vehicle has passed when the bounding box intersects both of the two passing lines. An apparatus or program characterized by that.

17. In any of the apparatuses or programs according to claims 14 to 16, two passing lines, a first passing line and a second passing line, are set in front of the moving direction of the vehicle, and the vehicle passage determination means determines that the vehicle has passed when the bounding box intersects the second passing line after intersecting the first passing line. An apparatus or program characterized by that.

18. In the apparatus or program according to claim 17, the vehicle passage determination means determines whether the intersection has occurred based on whether the center position of the bounding box has passed through the passing line on the side farther from the camera in the captured image among the first passing line or the second passing line, and determines whether the intersection has occurred based on whether the point at the bottom of the bounding box has passed through the passing line on the side closer to the camera than the first passing line or the second passing line in the captured image. An apparatus or program characterized by that.

19. In any of the apparatuses or programs according to claims 1 to 18, the vehicle estimation device is constructed as a server device. An apparatus or program characterized by that.

20. An imaging unit that images a traveling vehicle and outputs an imaging image, conversion image generation means for outputting a converted vehicle image suitable for recognizing the vehicle in the imaging image based on the imaging image, vehicle image extraction means for recognizing the vehicle in the converted imaging image and extracting a vehicle image by a bounding box surrounding the vehicle, vehicle passage determination means for determining passage of a vehicle by detecting that the bounding box has straddled the passing line in the continuous imaging images based on the passing line set in the imaging image, In a vehicle passage detection device comprising: Two passing lines, a first passing line and a second passing line, are set in front of the moving direction of the vehicle, The vehicle passage determination means determines that the vehicle has passed when the bounding box intersects the second passing line after intersecting the first passing line. The vehicle passage determination means determines whether the intersection has occurred based on whether the center position of the bounding box has passed through the passing line on the side farther from the camera in the imaging image among the first passing line or the second passing line, and determines whether the intersection has occurred based on whether the point at the bottom of the bounding box has passed through the passing line on the side closer to the camera than the first passing line or the second passing line in the imaging image. A vehicle passage detection device characterized by this.

21. A vehicle passage detection program for realizing a vehicle passage detection device by a computer, the computer being conversion image generation means for outputting a converted vehicle image suitable for recognizing the vehicle in the imaging image based on the imaging image of the traveling vehicle captured, vehicle image extraction means for recognizing the vehicle in the converted imaging image and extracting a vehicle image by a bounding box surrounding the vehicle, In a vehicle passing detection program for functioning as vehicle passing determination means for determining passing of a vehicle by detecting that the bounding box has straddled a passing line in consecutive captured images based on the passing line set in the captured image, Two passing lines, i.e., a first passing line and a second passing line, are set in front of the moving direction of the vehicle, The vehicle passing determination means determines that the vehicle has passed when the bounding box intersects the second passing line after intersecting the first passing line. The vehicle passing determination means determines whether or not the intersection has occurred based on whether or not the center position of the bounding box has passed through the passing line on the side farther from the camera in the captured image among the first passing line and the second passing line, and determines whether or not the intersection has occurred based on whether or not the point at the bottom of the bounding box has passed through the passing line on the side closer to the camera than the first passing line or the second passing line in the captured image. A vehicle passing detection program characterized by the above.

Citation Information

Patent Citations

  • Vehicle entrance detector

    JP1998208058A

  • Method and device for discriminating vehicle kind in the daytime

    JP1999353581A

  • Vehicle traffic quantity measuring system

    JP2002042113A

  • Vehicle detection device

    JP2013257720A

  • Vehicle type discrimination device and vehicle type discrimination method

    JP2017045137A