Information processing apparatus, region division method, and program
The information processing apparatus uses a neural network model with a context-enhanced traffic segmentation model to accurately segment road and congestion areas from aerial photographs, addressing scale variation issues and enhancing boundary clarity for real-time traffic analysis.
Patent Information
- Application Number
- JP2022102136
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-06-24
AI Technical Summary
Existing methods struggle to accurately segment traffic congestion areas from aerial photographs, as images taken by traffic cameras provide limited road condition information for small areas and fail to distinguish between road and congestion zones effectively.
An information processing apparatus using a neural network model with a context-enhanced traffic segmentation model, comprising a Feature Pyramid Network, Original Traffic Module, and Context Attention Module, to simultaneously identify road and traffic congestion areas, addressing scale variation issues through multilevel prediction and context enhancement.
Enables accurate and automatic segmentation of road and traffic congestion areas from aerial images, improving boundary clarity and enabling real-time traffic condition analysis.
Smart Images

Figure 0007713197000001 
Figure 0007713197000002 
Figure 0007713197000003
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for extracting road information from images of aerial photographs.
Background Art
[0002] By detecting traffic congestion or estimating traffic density from images of aerial photographs, real-time traffic condition information can be provided to urban monitoring systems and drivers. With such traffic condition information, for example, an appropriate driving route can be determined.
[0003] In the technique disclosed in Non-Patent Document 2, which is a conventional technique for detecting congestion, a technique for treating congestion detection as a classification problem by estimating traffic density from an image taken by a camera installed at an intersection is disclosed.
[0004] Also, in the technique disclosed in Non-Patent Document 3, data is collected using an open-source application programming interface (API) provided by the Land Transport Authority (LTA), and a convolutional neural network (CNN) for estimating traffic density is proposed.
[0005] However, since the images used in Non-Patent Documents 2 and 3 were taken by traffic cameras (cameras installed at intersections etc.), they can provide road condition information for only a very small area.
[0006] Also, many methods for extracting roads from images of aerial photographs have been proposed based on semantic segmentation. For example, Non-Patent Document 1 discloses a technique for performing road extraction by decomposing it into three mutually related subtasks, namely, road surface segmentation, road edge detection, and road centerline extraction. However, although the road network can be extracted, it has not been possible to segment (divide) the traffic congestion areas from the images of aerial photographs.
Prior Art Documents
Non-Patent Literature
[0007]
Non-Patent Literature 1
Non-Patent Literature 2
Non-Patent Literature 3
Summary of the Invention
Problems to be Solved by the Invention
[0008] The present invention has been made in view of the above points, and an object thereof is to provide a technique that enables the identification of traffic congestion areas on roads from aerial photo images.
Means for Solving the Problems
[0009] According to the disclosed technique, there is provided an acquisition unit that acquires an image taken from above, a calculation unit that simultaneously identifies road areas and traffic congestion areas from the image using a neural network model, and an output processing unit that outputs the identification result obtained by the calculation unit. An information processing apparatus including these components is provided.
Advantages of the Invention
[0010] According to the disclosed technique, a technique is provided that enables the identification of traffic congestion areas on roads from aerial photo images.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Embodiments for Carrying Out the Invention
[0012] Hereinafter, embodiments of the present invention (hereinafter referred to as "the present embodiments") will be described with reference to the drawings. The embodiments described below are merely examples, and the embodiments to which the present invention is applied are not limited to the following embodiments.
[0013] In addition, in this specification and the claims, "division", "segmentation", "extraction", "classification", "categorization", and "segmentation" may be used synonymously with each other. That is, "division", "segmentation", "extraction", "classification", "categorization", and "segmentation" described in the specification or the claims may be replaced with any one of the others.
[0014] Also, "aerial photograph" may be replaced with "aerial photo". The image of an aerial photograph may be replaced with "aerial image" or "aerial photo". An aerial photograph / aerial photo is a photograph taken of the ground from a flying object in the sky, and the flying object is not limited to a specific one. For example, the flying object may be an airplane, a satellite, or a drone.
[0015] Also, "enhance" means, for example, improving the accuracy of the classification of traffic jam areas or clarifying the boundary between traffic jam areas and areas other than traffic jams.
[0016] (Example of device configuration, example of operation) FIG. 1 shows a configuration example of an information processing apparatus 100 according to the present embodiment. As shown in FIG. 1, the information processing apparatus 100 includes an aerial photograph collection unit 110, an algorithm calculation unit 120, an output processing unit 130, and a learning unit 140. Note that the aerial photograph collection unit 110 may be referred to as an acquisition unit. Also, the algorithm calculation unit 120 may be referred to as a calculation unit.
[0017] Referring to the flowchart of FIG. 2, the processing flow during inference (testing) by the information processing apparatus 100 will be described. In S101, the aerial photo collection unit 110 acquires photos or videos (moving images) taken by a drone, satellite, aircraft, etc. These photos and videos will be collectively referred to as "aerial photos". The images of the aerial photos acquired by the aerial photo collection unit 110 are input to the algorithm calculation unit 120.
[0018] The algorithm calculation unit 120 has a neural network model (end-to-end model) described later. Here, it is assumed that the model has been trained. In S102, the algorithm calculation unit 120 inputs the image of the aerial photo into the model, and as an output from the model, acquires an image in which the road area and the traffic jam area (the traffic jam area in the road area and the area other than the traffic jam area are distinguished) are segmented (classified).
[0019] In S103, the output processing unit 130 may output the image obtained by the algorithm calculation unit 120 (the image in which the road area and the traffic jam area are segmented) as it is, or may perform processing on the image and output the processed image. For example, the output processing unit 130 can also extract and output only the traffic jam area on the road from the image obtained by the algorithm calculation unit 120.
[0020] During the learning of the model, a large number of aerial photo images and their label data (for example, data in which roads, traffic jams, and areas other than traffic jams are labeled on the images) are used. The learning unit 140 inputs the aerial photo image into the model and adjusts the parameters (weights) of the model so that the error between the output from the model and the correct answer is minimized.
[0021] Note that the device for learning (the device equipped with the learning unit 140) and the device for inference may be separate devices. In this case, the device for learning may be called a learning device. Hereinafter, the configuration and operation of the algorithm calculation unit 120 will be described in detail. Also, the device for inference does not necessarily have to be equipped with the learning unit 140.
[0022] The algorithm calculation unit 120 has a neural network model. With this model, it is possible to simultaneously segment (segment) the road surface and the traffic jam areas on the road surface in the aerial photo image.
[0023] In the image of the aerial photo, since the size (scale) of the vehicle seen from above is smaller than the size of the road surface seen from above, segmenting the vehicle is generally very difficult. This is called the scale variation problem. Therefore, in the prior art, it is very difficult to accurately segment the boundary between the traffic jam area and the non-traffic jam area.
[0024] The model constituting the algorithm calculation unit 120 according to the present embodiment solves the above problems and can accurately segment the traffic jam area from the aerial photo image.
[0025] Hereinafter, the configuration and operation of the context-enhanced traffic segmentation model in the present embodiment will be described in detail. Hereinafter, for convenience of description, the context-enhanced traffic segmentation model may be referred to as the "model".
[0026] (Overall configuration of the model) Fig. 3 shows an example of the overall configuration of the context-enhanced traffic segmentation model. As shown in Fig. 3, this model has a Feature Pyramid Network (FPN) 210, an Original Traffic Module 220, and a Context Attention Module 230.
[0027] The Context Attention Module 230 has a Global Context Generator 240 and an Attention Computation Block 250.
[0028] Since multilevel prediction is effective for scale variation problems, first, an aerial photo image is input into a feature pyramid network 210 that performs multilevel prediction. The feature pyramid network 210 generates a feature pyramid consisting of five feature maps (P2 to P6) at different scales from the input image. The feature map of P6 contains the semantic information of the highest (topmost) level.
[0029] The five feature maps (P2 to P6) are input into the original traffic module 220, and the original traffic module 220 generates segmentation results for the road surface and the original traffic jam.
[0030] Also, the feature map of P6 is input into the global context generator 240, and the global context generator 240 generates a global context feature from the feature map of P6.
[0031] The global context feature and the original traffic jam segmentation result are input into the attention calculation block 250, and the attention calculation block 250 outputs an enhanced (higher-quality) traffic jam segmentation (the congested area).
[0032] The final output can be obtained by combining the enhanced traffic jam segmentation and the segmentation of the road surface obtained by the original traffic module 220.
[0033] (Original Traffic Module 220) Next, the original traffic module 220 will be described. To segment the congested area on the road, it is necessary to segment the vehicle group on the road. However, when viewed from above, the scale of the vehicle relative to the road is small, causing a scale variation problem. The original traffic module 200 has a configuration of a multiscale feature fusion network to address this problem.
[0034] Figure 4 shows a configuration example of the original traffic module 220. As shown in Figure 4, the original traffic module 220 includes Convolution layers 221 and a Fusion Layer 222. The Convolution layers 221 include three consecutive 3×3 convolution layers for each feature map.
[0035] As shown in Figure 4, each feature map P ∈ R 256×H×W (size: 256×H×W) is input into the convolution layer 221. For each feature map, the same convolution process is performed. Through the convolution process, a new feature map ~ P ∈ R 1×H×W is generated. In the text of this specification, for the convenience of description, the symbols described at the beginning of the characters are described before the characters. " ~ P" is an example of this.
[0036] Next, each feature map is resized to the scale of the original input R h×w by bilinear interpolation. Then, five feature maps are fused by concatenation, and the fused feature map is input into one 3×3 convolution layer (fusion layer 222) to obtain the original segmentation result having a traffic jam map A ∈ R 1×h×w and a road surface map B ∈ R 1×h×w .
[0037] (Context Attention Module 230) Next, the context attention module 230 will be described.
[0038] Generally, in the image of an aerial photograph, since the boundary between the traffic jam area on the road and other areas is ambiguous, it is difficult to explicitly separate the traffic jam area and the passable road area. The context attention module 230 solves this problem and clarifies the boundary between the traffic jam area and other areas. The context attention module 230 uses the global context feature map to improve the original traffic jam map obtained by the original traffic module 220 and clarify the above boundary.
[0039] The global context module 240 includes a pyramid pooling module (PPM). As described above, the feature map P6 has the most powerful semantic information, and the global context module 240 takes the feature map P6 as input.
[0040] That is, first, the global context module 240 applies the pyramid pooling module (PPM) to the feature map of P6 to further utilize the region representation and context dependence. The global context module 240 obtains a global context feature map C ∈ R 1×h×w to obtain.
[0041] In the pyramid pooling module, by using a plurality of grids having a pyramidal size hierarchy and performing pooling on the input, global (global) rough context information indicating how much the features of each class are included in each grid can be obtained.
[0042] The global context feature map and the original traffic jam map are input to the attention calculation block 250.
[0043] Fig. 5 shows the processing configuration of the attention calculation block 250. A neural network is configured to enable this processing.
[0044] As shown in FIG. 5, first, the original traffic jam map A and the global context feature map C are each downsampled to obtain { - A, - C} ∈ R 1×h / 4×w / 4 . Next, these are reshaped (transformed) to obtain two new feature maps { ~ A, ~ C} ∈ R 1×n . Here, n = h / 4 × w / 4, indicating the number of pixels in the feature map.
[0045] As shown in FIG. 5 and the following formula (1), matrix multiplication is performed between A and the transposed ~ A and the transposed ~ C, and the context attention map S ∈ R n×n is calculated by the softmax layer.
[0046] S = Softmax( ~ C T × ~ A) (1) To enhance (clarify) the traffic jam area using the context information, as shown in FIG. 5 and the following formula (2), the context attention map S is multiplied by ~ A, and the product is reshaped to R 1×h / 4×w / 4 . Then, - A is added to obtain the traffic jam map ~ A S ∈ R 1×h / 4×w / 4 enhanced by the context. Here, α is a learnable weight parameter initialized as 0.
[0047] ~ A S = α(Reshape( ~ A × S)) + - A (2) And then, ~ A S is resized to obtain the final result A S ∈ R 1×h×w of the enhanced traffic jam segmentation.
[0048] The above is the processing of the context attention module 230. Finally, the enhanced traffic congestion segmentation and the original road surface segmentation are combined (merged), input into a 3×3 convolutional layer, and the final traffic congestion segmentation result is generated. In the final traffic congestion segmentation result, for example, on the image of the aerial photo, the road area is segmented and shown, and the traffic congestion area and the non-traffic congestion area in the road area are separately shown.
[0049] (Summary of the context-enhanced traffic segmentation model) As described above, in this embodiment, the context-enhanced traffic segmentation model divides (segments) traffic congestion and road surface from the image of the aerial photo in an end-to-end manner. The "context-enhanced traffic segmentation model" is a traffic segmentation model whose performance is enhanced by context.
[0050] The model in this embodiment is composed of two modules (the original traffic module 220 and the context attention module 230) that enable explicit division of traffic congestion and road surface.
[0051] The original traffic module 220 is a module for solving the scale variation problem in the image of the aerial photo. That is, in this module, based on the feature pyramid, a multi-scale feature map is used, and further features are extracted by the convolutional layer 221. Then, a plurality of features of different scales are fused by the fusion layer 222 to obtain the original (initial) segmentation of traffic congestion and road surface.
[0052] The context attention module 230 enhances the boundaries of traffic congestion. The context attention module 230 consists of an attention calculation block 250 and a corresponding global context generator 240. The feature map at the top level of the feature pyramid contains the strongest semantic information. Therefore, it is input into the global context generator 240, and a global context map is obtained through pyramid pooling operation. Then, in the attention calculation block 250, an attention map between the global context map and the original segmentation of traffic congestion is calculated. Finally, the attention map is used to strengthen (clarify) the boundaries of traffic congestion to obtain the final traffic congestion segmentation result. Thereby, the road surface and the congestion area can be simultaneously and accurately distinguished. Also, in the image of the aerial photograph, the scale problem that the scale of the vehicle is smaller than the scale of the road surface seen from the air and the segmentation of the vehicle is very difficult is solved.
[0053] (Hardware configuration example) The information processing apparatus 100 can be realized, for example, by causing a computer to execute a program. This computer may be a physical computer or a virtual machine on the cloud.
[0054] That is, the information processing apparatus 100 can be realized by executing a program corresponding to the processing performed by the information processing apparatus 100 using hardware resources such as a CPU and a memory built in the computer. The above program can be recorded on a computer-readable recording medium (such as a portable memory), saved, distributed, or provided through a network such as the Internet or e-mail.
[0055] FIG. 6 is a diagram showing an example of the hardware configuration of the above computer. The computer in FIG. 6 includes a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, etc., which are mutually connected by a bus BS.
[0056] A program for realizing the processing on the computer is provided by a recording medium 1001 such as a CD-ROM or a memory card, for example. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 via the drive device 1000 into the auxiliary storage device 1002. However, the installation of the program does not necessarily have to be performed from the recording medium 1001, and it may be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program and also stores necessary files, data, etc.
[0057] When an instruction to start the program is given, the memory device 1003 reads out and stores the program from the auxiliary storage device 1002. The CPU 1004 realizes the functions related to the information processing device 100 according to the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network, various measurement devices, a motion intervention device, etc. The display device 1006 displays a GUI (Graphical User Interface) etc. by the program. The input device 1007 is composed of a keyboard, a mouse, buttons, or a touch panel, etc., and is used to input various operation instructions. The output device 1008 outputs the calculation result.
[0058] (Effects of the Embodiment) With the technology according to this embodiment, instead of extracting road and traffic jam information from an image with human eyes as in the prior art, the road area and the traffic jam area can be automatically output. Further, considering the scale variation problem on a conventional image (such as a small car and a large road), the area can be accurately segmented from the image. Also, the boundary line part dividing between areas can be smoothly displayed.
[0059] In addition, with the segmentation map obtained by the technology according to this embodiment, visually, the location of the traffic jam and the ratio of the traffic jammed area to the road area can be grasped. Also, since the passable locations can be grasped even in case of a traffic jam, for example, it can be determined whether an emergency vehicle can pass through. Such a point is superior to traffic jam detection and density estimation in the prior art.
[0060] (Supplementary Note) Regarding the above embodiments, the following supplementary claims are further disclosed. (Supplementary Claim 1) A memory, a processor, and comprising wherein the processor acquires an image taken from above, simultaneously classifies the road area and the traffic jam area from the image using a neural network model, and outputs the obtained classification result an information processing apparatus. (Supplementary Claim 2) The model comprises a first module that classifies the road area and the traffic jam area in the image using a plurality of feature maps obtained from the image, and a second module that enhances the traffic jam area obtained by the first module using a specific feature map among the plurality of feature maps The information processing apparatus according to Supplementary Claim 1. (Supplementary Claim 3) The plurality of feature maps are generated by a feature pyramid network included in the model, and the specific feature map is the feature map at the highest level among the plurality of feature maps The information processing apparatus according to Supplementary Note 2. (Supplementary Note 4) The second module A context generator that generates global context from the specific feature map, An attention calculation block that uses the global context and the traffic jam area obtained by the first module to generate a traffic jam area with higher accuracy than the traffic jam area The information processing apparatus according to Supplementary Note 2 or 3, comprising (Supplementary Note 5) A region classification method executed by an information processing apparatus, comprising An acquisition step of acquiring an image taken from above, A calculation step of simultaneously classifying a road region and a traffic jam region from the image using a neural network model, An output step of outputting the classification result obtained by the calculation step A region classification method comprising (Supplementary Note 6) A non-transitory storage medium storing a program for causing a computer to function as each part in the information processing apparatus according to any one of Supplementary Notes 1 to 4.
[0061] As described above, the present embodiment has been described, but the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.
Explanation of Signs
[0062] 100 Information processing apparatus 110 Aerial photo collection unit 120 Algorithm calculation unit 130 Output processing unit 140 Learning unit 210 Feature pyramid network 220 Original traffic module 221 Convolutional layer 222 Fusion layer 230 Context attention module 240 Global Context Generator 250 Attention Calculation Block 1000 Drive Device 1001 Recording Medium 1002 Auxiliary Storage Device 1003 Memory Device 1004 CPU 1005 Interface Device 1006 Display Device 1007 Input Device 1008 Output Device
Claims
1. An acquisition unit that acquires an image taken from above; A calculation unit that simultaneously classifies a road area and a traffic jam area from the image using a neural network model; An output processing unit that outputs the classification result obtained by the calculation unit An information processing apparatus comprising:
2. The model is A first module that classifies a road area and a traffic jam area in the image using a plurality of feature maps obtained from the image; A second module that enhances the traffic jam area obtained by the first module using a specific feature map among the plurality of feature maps The information processing apparatus according to claim 1, comprising:
3. The plurality of feature maps are generated by a feature pyramid network included in the model, and the specific feature map is the highest-level feature map among the plurality of feature maps The information processing apparatus according to claim 2.
4. The second module is A context generator that generates a global context from the specific feature map; An attention calculation block that generates a traffic jam area with higher accuracy than the traffic jam area using the global context and the traffic jam area obtained by the first module The information processing apparatus according to claim 2, comprising:
5. A region classification method executed by an information processing apparatus, comprising: An acquisition step of acquiring an image taken from above; A calculation step of simultaneously classifying a road area and a traffic jam area from the image using a neural network model; An output step of outputting the classification result obtained by the calculation step A region classification method comprising:
6. A program for causing a computer to function as each unit in the information processing apparatus according to any one of claims 1 to 4.
Citation Information
Patent Citations
A system for planetary-scale analysis.
JP2019513315A
Learning method and learning device for improving segmentation performance in road obstacle detection required to satisfy level 4 and level 5 of autonomous vehicles by using laplacian pyramid network, and testing method and testing device using the same
JP2020119500A
Traffic abnormality detection device, method, and electronic apparatus
JP2022017185A
Deep Learning Methods For Estimating Density and / or Flow of Objects, and Related Methods and Software
US20200118423A1