Program, method, system, and road map creation method
A computer program efficiently detects and evaluates road surface damage by generating a bird's-eye view with assigned damage information, addressing the inefficiencies of manual recording methods.
Patent Information
- Application Number
- JP2022116149
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-07-21
- Publication Date
- 2026-02-09
- Estimated Expiration
- 2042-07-21
AI Technical Summary
The traditional method of manually recording the location of road surface damage requires a significant amount of work.
A system that uses a computer program to detect road surface damage by acquiring a video from a moving vehicle, performing image processing to generate point cloud data, and generating a bird's-eye view with damage information assigned to each unit area.
Efficient evaluation of the state of cracks on a road surface is achieved, reducing manual labor and improving data accuracy.
Smart Images

Figure 0007812303000001 
Figure 0007812303000002 
Figure 0007812303000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a program, a method, and a system ,oh and how to create road maps. [Background technology]
[0002] In recent years, as disclosed in Patent Document 1, for example, a system has become known that uses a camera to acquire an image of the road surface and detects damage to the road surface.
[0003] When using such a system to compile the crack rate for the entire road, the location of the damage on the road surface had to be recorded manually. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-15417 Summary of the Invention [Problem to be solved by the invention]
[0005] The traditional method of manually recording the location of road surface damage requires a huge amount of work.
[0006] The present disclosure aims to provide a system that can efficiently evaluate the state of cracks on a road surface. [Means for solving the problem]
[0007] The program disclosed herein is a program used in a computer that detects road surface damage, and causes a processor of the computer to execute the following steps: acquiring a video of the road surface captured from a moving vehicle; detecting damage to the road surface captured in the frames by performing identification using a trained model on multiple frames included in the captured video; performing image processing to acquire, from the multiple frames, point cloud data indicating the relative positional relationship of the subjects captured in the frames and the estimated shooting positions of the camera that captured each frame; generating a bird's-eye view of the road surface using the acquired point cloud data and images captured in the multiple frames; and assigning detected damage information to each divided unit area of the generated bird's-eye view using the estimated shooting positions to generate road surface damage data having damage information for each unit area. [Effects of the Invention]
[0008] According to the present disclosure, the state of cracks in a road surface can be efficiently evaluated. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a block diagram illustrating an example of the configuration of a damage detection system according to an embodiment of the present invention. [Figure 2] 2 is a block diagram showing an example of a functional configuration of the mobile terminal shown in FIG. 1. FIG. [Figure 3] FIG. 2 is a diagram illustrating a state in which a terminal device is mounted on a vehicle. [Figure 4] 2 is a block diagram showing an example of a functional configuration of a server shown in FIG. 1. FIG. [Figure 5] 10A and 10B are diagrams illustrating detection of damage from a frame according to the present embodiment. [Figure 6] 10A and 10B are diagrams illustrating selection of key frames according to the present embodiment. [Figure 7] 1A to 1C are diagrams illustrating the principle of image processing according to the present embodiment. [Figure 8] 5A to 5C are diagrams illustrating data acquired by image processing according to the present embodiment. [Figure 9] 10A and 10B are diagrams illustrating characteristics of a road surface image in a frame according to the present embodiment. [Figure 10] FIG. 10 is a diagram illustrating scale estimation according to the embodiment. [Figure 11] FIG. 10 is a diagram illustrating rendering of point cloud data onto a plane. [Figure 12] FIG. 10 is a diagram illustrating allocation of damage data to bird's-eye view data. [Figure 13] FIG. 2 is a diagram illustrating a first process executed by the damage detection system. [Figure 14] FIG. 10 is a diagram illustrating an outline of a process executed by a server. [Figure 15] FIG. 10 is a diagram illustrating a second process executed by the damage detection system. [Figure 16] FIG. 10 is a diagram illustrating a third process executed by the damage detection system. [Figure 17] FIG. 10 is a diagram illustrating a fourth process executed by the damage detection system. [Figure 18] FIG. 4 is a diagram showing a part of bird's-eye view data according to the embodiment. [Figure 19] FIG. 4 is a diagram showing road surface damage data according to the present embodiment. [Figure 20] FIG. 2 is a diagram showing a road damage map according to the embodiment. [Figure 21] FIG. 21 is an enlarged view of a section of the road damage map shown in FIG. 20. [Figure 22] FIG. 21 is a diagram visualizing the damage state for each unit area of the part of the road surface shown in FIG. 20. [Figure 23] 10A and 10B are diagrams illustrating a process of allocating damage information in a server according to a modified example. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments will be described with reference to the drawings. Note that the embodiments described below do not unduly limit the content of the present disclosure described in the claims. Furthermore, not all of the configurations described in the embodiments are necessarily essential components of the present disclosure.
[0011] <1. Overview of the embodiment> An overview of an embodiment of a damage detection system 1 (hereinafter simply referred to as system 1) according to the present disclosure will be described. The system 1 according to this embodiment uses images or videos taken from the traveling vehicle 100 to identify damage to the road surface on the road on which the vehicle 100 is traveling, and detects the presence or absence of damage.
[0012] Furthermore, the system 1 uses images or videos taken from the traveling vehicle 100 to generate various data related to the road surface in the evaluation section. Representative examples of the various types of data generated by the system 1 are defined below. Bird's-eye view data (bird's-eye view): Information showing the plan view of the road surface in the section where damage is being evaluated Road surface damage data: Information on damage information assigned to each unit area in bird's-eye view data Road damage map: A road map with road surface damage data for each evaluation section superimposed on the map information. The process of creating each of these data will be described later.
[0013] <2. Overall system configuration> First, the configuration of a system 1 according to this embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing an example of the configuration of the system 1 according to this embodiment. The system 1 shown in Fig. 1 includes a terminal device 10 and an information processing server 20 (hereinafter simply referred to as the server 20). The terminal device 10 and the server 20 are connected to each other so as to be able to communicate with each other via a network 30. The network 30 is configured as a wired or wireless network.
[0014] The terminal device 10 and the server 20 are connected to the network 30 using a wireless base station 31 or a wireless LAN standard 32. The terminal device 10 and the server 20 may also be connected to the network 30 by wired communication.
[0015] Here, a collection of devices (for example, the terminal device 10 and the server 20) that make up the system 1 can be understood as a single "information processing device." In other words, the system 1 can be realized as a collection of multiple devices, and the allocation of multiple functions for realizing the system 1 can be appropriately determined based on the processing capacity of the hardware of each device.
[0016] The terminal device 10 is a terminal mounted on the vehicle 100. The terminal device 10 is, for example, a general-purpose mobile terminal such as a smartphone, a tablet terminal, or a laptop computer. A plurality of smartphones, tablet terminals, laptop computers, etc. may be used as the terminal device 10. Such a mobile terminal may be a terminal mounted on the vehicle 100 for the purpose of checking the driving route of the vehicle 100 while driving.
[0017] The terminal device 10 may be a general-purpose drive recorder mounted on the vehicle 100 for the purpose of recording the situation when an accident occurs, and equipped with the functions of the terminal device 10 described below. Furthermore, the terminal device 10 may complement this function with the image capturing function of a general-purpose drive recorder.
[0018] Furthermore, in the present invention, the terminal device 10 that photographs the condition of the road surface is not limited to a dedicated device used only for photographing the road surface, but may also be a dual-purpose device that photographs the road surface incidentally by utilizing the photographing function possessed by the terminal device 10 that is installed for other purposes.
[0019] The terminal device 10 includes a processor 11, a memory 12, a storage 13, a communication IF 14, and an input / output IF 15. The processor 11 is hardware for executing an instruction set written in a program, and is composed of an arithmetic unit, a register, a peripheral circuit, and the like.
[0020] The memory 12 is for temporarily storing programs and data to be processed by the programs, and is a volatile memory such as a DRAM (Dynamic Random Access Memory).
[0021] The storage 13 is a storage device for saving data, such as a flash memory, a hard disc drive (HDD), or a solid state drive (SSD).
[0022] The communication IF 14 is an interface for inputting and outputting signals so that the system 1 can communicate with external devices.
[0023] The input / output IF 15 functions as an interface with an input device (for example, a pointing device such as a mouse, a keyboard) for receiving input operations from the user, and an output device (for example, a display, a speaker, etc.) for presenting information to the user.
[0024] The server 20 includes a processor 21, a memory 22, a storage 23, a communication IF 24, and an input / output IF 25. The configurations of the processor 21, the memory 22, the storage 23, the communication IF 24, and the input / output IF 25 are similar to the configurations of the processor 11, the memory 12, the storage 13, the communication IF 14, and the input / output IF 15 described above, respectively, and therefore will not be described again.
[0025] The server 20 is realized by a general-purpose computer, a mainframe, etc. The server 20 may be realized by one computer or a combination of multiple computers.
[0026] 3. Configuration of Terminal Device 10 Fig. 2 is a block diagram showing the functional configuration of the terminal device 10. As shown in Fig. 2, the terminal device 10 includes a plurality of antennas (antenna 111, antenna 112), wireless communication units (first wireless communication unit 121, second wireless communication unit 122) corresponding to the respective antennas, an operation reception unit 130, an imaging unit 150, a storage unit 160, a control unit 170, and a GPS antenna 180. Note that the terminal device 10 may include a keyboard as the operation reception unit 130.
[0027] The terminal device 10 also has functions and configurations (for example, a battery for storing power, a power supply circuit for controlling the supply of power from the battery to each circuit, etc.) that are not specifically shown in Fig. 2. As shown in Fig. 2, each block included in the terminal device 10 is electrically connected by a bus or the like.
[0028] The antenna 111 emits a signal emitted by the terminal device 10 as a radio wave. The antenna 111 also receives a radio wave from space and provides the received signal to the first radio communication unit 121.
[0029] The antenna 112 emits a signal emitted by the terminal device 10 as a radio wave. The antenna 112 also receives a radio wave from space and provides the received signal to the second radio communication unit 122.
[0030] The first wireless communication unit 121 performs modulation / demodulation processing and the like for transmitting and receiving signals via the antenna 111 so that the terminal device 10 can communicate with other wireless devices. The second wireless communication unit 122 performs modulation / demodulation processing and the like for transmitting and receiving signals via the antenna 112 so that the terminal device 10 can communicate with other wireless devices. The first wireless communication unit 121 and the second wireless communication unit 122 are communication modules including a tuner, an RSSI (Received Signal Strength Indicator) calculation circuit, a CRC (Cyclic Redundancy Check) calculation circuit, a high-frequency circuit, etc. The first wireless communication unit 121 and the second wireless communication unit 122 perform modulation / demodulation and frequency conversion of wireless signals transmitted and received by the terminal device 10, and provide the received signals to the control unit 170.
[0031] The operation reception unit 130 has a mechanism for receiving input operations from the user. Specifically, the operation reception unit 130 includes a display 131. The operation reception unit 130 is configured as a touch screen that uses a capacitive touch panel to detect the position of the user's touch on the touch panel. The operation terminal does not necessarily have to include a touch panel.
[0032] Display 131 displays data such as images, videos, and text under the control of control unit 170. Display 131 is realized by, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display.
[0033] The image capturing unit 150 is a device that is mounted on the vehicle 100 and captures an image of the road surface. Specifically, the image capturing unit 150 is a camera that receives light with a light receiving element and outputs image capturing data 161. The image capturing unit 150 can capture not only still images but also moving images. When a moving image is captured by the image capturing unit 150, each frame of the image constituting the moving image is treated as captured data 161. Each image (hereinafter referred to as a frame) constituting the captured data 161 has information relating to the time of capture. FIG. 3 is a diagram illustrating a state in which the terminal device 10 is mounted on a vehicle 100.
[0034] 3, the terminal device 10 is mounted inside the vehicle 100 with the image capturing unit 150 facing forward (in the traveling direction) of the vehicle 100. In this embodiment, the image capturing unit 150 continuously captures the state of the road surface ahead of the vehicle 100 to obtain video data of the road surface. When the photographing unit 150 acquires image data, the photographing unit 150 photographs the state of the road surface ahead intermittently at predetermined time intervals.
[0035] 2 detects the position of the terminal device 10, thereby detecting the position of the vehicle 100 on which the terminal device 10 is mounted, and acquires information related to the route traveled by the vehicle 100. The GPS antenna 180 acquires the acquired information related to the position of the vehicle 100 and the travel route together with time information, and transmits it to the transmitting / receiving unit 172 of the control unit 170. The transmitting / receiving unit 172 transmits the received information related to the position of the vehicle 100 and the travel route together with information related to the time when the position information was acquired to the server 20.
[0036] The storage unit 160 is configured with, for example, a flash memory or the like, and stores data and programs used by the terminal device 10. In one aspect, the storage unit 160 stores shooting data 161 and frame data 162. The frame data 162 is image data in which information relating to the shooting time is associated with a plurality of pieces of shooting data 161 through processing by the data processing unit 173, which will be described later.
[0037] The control unit 170 reads a program stored in the storage unit 160 and executes instructions contained in the program to control the operation of the terminal device 10. The control unit 170 is, for example, an application program pre-installed in the terminal device 10. By operating in accordance with the program, the control unit 170 fulfills the functions of an input receiving unit 171, a transmitting / receiving unit 172, a data processing unit 173, and a display processing unit 174.
[0038] The input receiving unit 171 performs processing to receive input operations made by the user to the terminal device 10 .
[0039] The transmitting / receiving unit 172 performs processing for the terminal device 10 to transmit and receive data to and from an external device such as the server 20 in accordance with a communication protocol. The transmitting / receiving unit 172 executes processing to transmit the frame data 162 stored in the storage unit 160 to the server 20.
[0040] The data processing unit 173 associates the position information of the vehicle 100 acquired by the GPS antenna 180 with each frame constituting the photographed data 161 based on the photographed time. The data processing unit 173 stores the image data associated with the position information in the storage unit 160 as frame data 162.
[0041] The display processing unit 174 performs processing to present information to the user. The display processing unit 174 performs processing to display a display image on the display 131, etc.
[0042] <4. Server 20 Configuration> 4 is a diagram showing the functional configuration of the server 20. As shown in FIG. 4, the server 20 functions as a communication unit 201, a storage unit 202, and a control unit 203. The communication unit 201 performs processing for the server 20 to communicate with external devices.
[0043] <4-1. Configuration of storage unit 202> The memory unit 202 stores data and programs used by the server 20. The memory unit 202 stores frame data 2021, point cloud data 2022, attitude data (attitude information) 2023, damage data 2024, bird's-eye view data 2025, road surface damage data 2026, a road damage map 2027, a first trained model 2028, and a second trained model 2029.
[0044] The frame data 2021 is image information transmitted from the terminal device 10. That is, the frame data 2021 is image data in which the data processing unit 173 of the terminal device 10 associates position information with the photographed data 161 photographed by the photographing unit 150 of the terminal device 10.
[0045] The point cloud data 2022 is data that indicates the relative positional relationship of a subject captured in multiple frames, and is acquired by processing in the image processing unit 2034, which will be described later.
[0046] The posture data (posture information) 2023 is information relating to the posture of the image capturing unit 150 that captured the frame. The posture data 2023 includes the following information for each frame. Information about the estimated shooting position of the shooting unit 150 at the time the frame was shot Information about the estimated shooting direction of the image capturing unit 150 at the time the frame was captured The posture data 2023 is acquired by processing in the image processing unit 2034, which will be described later.
[0047] Damage data (damage information) 2024 is data related to damage detected on the frame. The damage information includes information related to the position and type (including the degree) of the damage based on the plane coordinate system of the frame, which is two-dimensional information. The damage data 2024 is acquired by processing by a damage detection unit 2032, which will be described later.
[0048] The bird's-eye view data 2025 is data showing a plan view of the road surface in the damage evaluation section. The bird's-eye view data 2025 is data that is rendered as a plan view of the road surface occupying the evaluation section by image processing by the image processing unit 2034, which will be described later, for a plurality of frame data 2021 captured by the image capturing unit 150 facing in the traveling direction of the vehicle 100.
[0049] The road surface damage data 2026 is data in which damage information for each unit area is assigned to the bird's-eye view data 2025. The road surface damage data 2026 is generated by processing in a data generating unit 2037, which will be described later.
[0050] The road damage map 2027 is data in which road surface damage data 2026 for each evaluation section is superimposed on map information. The road damage map 2027 is generated by processing in a data generation unit 2037, which will be described later.
[0051] The first trained model 2028 is a model that identifies whether or not there is a damaged portion of the road surface in an image. The first trained model 2028 is obtained by having a machine learning model perform machine learning based on training data in accordance with a model training program. For example, in this embodiment, the first trained model 2028 is trained to output the location and type of damage for input image information.
[0052] In this case, the learning data is, for example, image information of road surface conditions photographed in the past as input data, and the correct output data for the input image information is information regarding the position of the damaged area in the image's planar coordinate system and information regarding the type of damage. The first trained model 2028 is stored in the storage unit 202 in a trained state, but may be retrained as needed based on the image data 161 captured by the image capturing unit 150.
[0053] Specifically, the first trained model 2028 compares the feature amounts (reference feature amounts) of images of road surfaces having various types of damage with the feature amounts of the photographed data 161 to be evaluated, and compares the similarity. In this way, the first trained model 2028 determines whether any type of damage is present on the road surface captured in the frame. When evaluating the similarity, if the feature amounts of the photographed data 161 to be evaluated are within a predetermined threshold range with respect to the reference feature amounts, it is determined that the damage is present. Note that reference feature amounts are set for each of the various types of damage.
[0054] The first trained model 2028 according to this embodiment is, for example, a parameterized composite function in which multiple functions are combined. The parameterized composite function is defined by a combination of multiple adjustable functions and parameters. The first trained model 2028 according to this embodiment may be any parameterized composite function that meets the above requirements, but is a multi-layered neural network model (hereinafter referred to as a "multi-layered network"). The first trained model 2028 using a multi-layered network has an input layer, an output layer, and at least one intermediate layer or hidden layer provided between the input layer and the output layer.
[0055] The multi-layered network according to the present embodiment may be, for example, a deep neural network (DNN), which is a multi-layered neural network that is the subject of deep learning. As the DNN, for example, a convolution neural network (CNN) that targets images may be used.
[0056] The second trained model 2029 is a model that performs a process (road surface semantic segmentation) to extract road areas from the subject captured in the frame. The second trained model 2029 is obtained by having a machine learning model perform machine learning based on training data in accordance with a model training program. For example, in this embodiment, the second trained model 2029 is trained to extract and output portions corresponding to road areas from input image information.
[0057] In this case, the learning data may be, for example, image information of a photographed landscape including a road surface as input data, and information obtained by extracting a road surface area from the input image information as correct output data. The second trained model 2029 is stored in the memory unit 202 in a trained state, but may be retrained at any time based on the image data 161 captured by the image capturing unit 150.
[0058] Specifically, the second trained model 2029 compares the feature amounts (reference feature amounts) of an image of a road surface area in image information of a photographed scene including a road surface with the feature amounts for each area of the frame to be evaluated to compare the similarities. As a result, the second trained model 2029 determines and extracts areas of the frame that have a high similarity to the reference feature amounts as road surface areas. When evaluating the similarity, areas that are within a preset threshold range of the reference feature amounts among the feature amounts for each area of the frame to be evaluated are determined to be road surface areas.
[0059] The second trained model 2029 according to this embodiment is, for example, a parameterized composite function in which multiple functions are combined. The parameterized composite function is defined by a combination of multiple adjustable functions and parameters. The second trained model 2029 according to this embodiment may be any parameterized composite function that meets the above requirements, but is a multi-layer network. The second trained model 2029 using a multi-layer network has an input layer, an output layer, and at least one intermediate layer or hidden layer provided between the input layer and the output layer.
[0060] <4-2. Configuration of the control unit 203> The control unit 203 performs functions as a transmission / reception unit 2031, a damage detection unit 2032, a frame selection unit 2033, an image processing unit 2034, a road surface extraction unit 2035, a scale estimation unit 2036, and a data generation unit 2037 by the processor 21 of the server 20 performing processing according to a program.
[0061] The transmitting / receiving unit 2031 controls the process by which the server 20 transmits signals to external devices in accordance with a communication protocol, and the process by which the server 20 receives signals from external devices in accordance with a communication protocol.
[0062] The damage detection unit 2032 detects damage to the road surface captured in each frame by performing classification using a trained model on multiple frames included in the captured video. Specifically, the damage detection unit 2032 classifies the presence or absence of road surface damage for each frame data 2021, and outputs information about the detected damage as damage information in association with position information in a plane coordinate system of the frame. The damage detection unit 2032 inputs the frame data 2021 to the first trained model 2028, thereby acquiring road surface damage information output from the first trained model 2028. Figure 5 is a diagram illustrating detection of damage from a frame.
[0063] As shown in FIG. 5, the damage detection unit 2032 outputs a bounding box B that encloses damage on the road surface for each frame. By outputting a bounding box B that encloses the damage in this way, not only is it possible to determine whether or not there is damage on the road surface, but also to identify the location of the damage, thereby obtaining more detailed information about the state of the damage. The location of the damage is identified by its position coordinates in the plane coordinate system of the frame. The location of the damage is emphasized by the bounding box B, making it easier to confirm the detected damage information. Note that the information about the location of the damage includes, in addition to the position coordinates in the plane coordinate system of the frame, location information at which the frame was captured, which is information associated with the frame as location information acquired by the GPS antenna 180.
[0064] The damage detection unit 2032 is not limited to a detection method using a bounding box, and may instead use a method of creating a damage map by segmentation. In this case, the damage detection unit 2032 uses segmentation to output a damage map that indicates the location of each type of damage for the subject captured in the frame. In this case, the damage map indicates the location of each pixel in the image and the area it occupies, and therefore provides more information than a bounding box.
[0065] The frame selection unit 2033 shown in FIG. 4 selects key frames from the frame data 2021. A key frame refers to a frame suitable for image processing, which will be described later. Specifically, a key frame refers to a frame among multiple frames in which the subjects overlap sufficiently with each other and in which the subjects change significantly over time. In other words, to be selected as a key frame, it is necessary for the degree of overlap of the subjects and the amount of change in the subjects to be well balanced between before and after the time series. FIG. 6 is a diagram explaining the selection of key frames.
[0066] 6, the frame selection unit 2033 selects frames suitable for image processing (described later) as key frames KF from multiple frames (frame group FG) captured by the image capture unit 150 through motion estimation processing. In the motion estimation processing, the frame selection unit 2033 evaluates the relationship between the amount of camera motion between two chronologically adjacent frames and the three-dimensional structure of the captured scene. As a result, the frame selection unit 2033 extracts frames suitable for the three-dimensional structure as key frames KF.
[0067] The image processing unit 2034 shown in Fig. 4 performs image processing involving three-dimensional construction, called SfM (Structure from Motion) processing, on multiple frames. The SfM processing by the image processing unit 2034 will be described in detail with reference to Fig. 7. Fig. 7 is a diagram illustrating the principle of image processing by the image processing unit 2034.
[0068] As shown in FIG. 7, in the SfM process, the shooting positions (P1 to P3) and shooting directions of multiple images (F1 to F3) taken by a camera are estimated. Then, in the SfM process, the same point (X j ) for each image position (X 1j ~X 3j ) to generate a 3D model of the entire subject. The created 3D model is composed of a point cloud, which is made up of many points. The point cloud obtained through SfM processing is merely information on relative positional relationships, and is data without defined geospatial coordinates.
[0069] That is, the image processing unit 2034 generates point cloud data from the parallax with respect to the traveling direction of the vehicle 100 for each frame for feature points extracted from the road surface and surrounding scenery, which are the subject of each frame.
[0070] That is, the image processing unit 2034 estimates, from multiple frames arranged in time series, information about the posture of the imaging unit 150 for each frame. The information about the posture of the imaging unit 150 includes an estimated imaging position and an estimated imaging direction of the imaging unit 150.
[0071] The image processing unit 2034 then extracts a part of the subject that is commonly captured in several frames as a feature point, and generates point data corresponding to the feature point from the parallax to the feature point for several frames in which the feature point is captured. The feature point may be a road surface area or a non-road surface area of the subject captured in the frames.
[0072] The image processing unit 2034 shown in FIG. 4 performs this processing on a plurality of feature points and on all of a plurality of frames arranged in time series, thereby acquiring point cloud data 2022 that indicates the entire subject captured in all of the frames. The image processing unit 2034 can also perform image processing on the selected key frames.
[0073] 8 is a diagram illustrating data acquired by image processing. As shown in FIG. 8, the data acquired by image processing includes the following information: Point cloud data showing the relative positional relationship of the photographed subject (Point Cloud Data 2022) Posture information of the imaging unit 150 that captured each frame (estimated imaging position and estimated imaging direction of the imaging unit 150) 8 indicates the locus L of the estimated imaging position of the imaging unit 150. Furthermore, the cone C indicates the imaging direction of the imaging unit 150. The position of the cone C indicates the estimated imaging position. Each point data is a 3D reconstruction of the extracted feature points. Of these, points displayed in black indicate points on the road surface, and points displayed in white indicate points outside the road surface.
[0074] Here, the characteristics of the road surface image captured in the frame captured by the image capturing unit 150 will be described. FIG. 9 is a diagram illustrating the characteristics of road surface images captured in frames. FIGS. 9A and 9B are frames adjacent to each other in a chronological order. In FIGS. 9A and 9B, comparison area A1 is hatched to represent the same area. Comparing these figures, images of non-road surface areas such as houses and roadside trees contain many areas where the angle at which the subject is viewed over time and the scale change little. On the other hand, images of road surface areas are characterized by large changes in the angle at which the subject is viewed over time and the scale. Furthermore, by examining comparison area A1 in the two figures, it is confirmed that road surface images are characterized by extreme changes in scale from frame to frame.
[0075] That is, road surface images captured from the traveling vehicle 100 have the characteristics of having few feature points and large changes in scale. Of these, the changes in scale become more noticeable by selecting key frames. 8, the point cloud data 2022 obtained by image processing by the image processing unit 2034 is sparse (coarse) compared to the point cloud obtained by general SfM processing. Here, general SfM processing refers to SfM processing performed using multiple images acquired in a shooting procedure in which the change in distance between the subject and the camera for each image is small.
[0076] 4 extracts a road surface area from each of a plurality of key frames selected from the frame data 2021. The road surface extraction unit 2035 inputs each key frame to the second trained model 2029, thereby acquiring the road surface area for each key frame output from the second trained model 2029.
[0077] The scale estimation unit 2036 estimates the scale of the point cloud data 2022. FIG. 10 is a plan view of the road seen from above. As shown in Fig. 10, the scale estimation unit 2036 estimates the scale of the estimated shooting trajectory L by matching the travel route R of the vehicle 100 with the estimated shooting trajectory L of the shooting unit 150.
[0078] In other words, since the estimated shooting trajectory L obtained by image processing does not have defined geographical coordinates, the scale of the estimated shooting trajectory L can be estimated using the position information of the shooting unit 150, and the estimated scale can be used as the scale of the point cloud data 2022. Here, the travel route R of the vehicle 100 is position information of the image capturing unit 150 acquired by the GPS antenna 180, and can be obtained from data associated with each frame. The travel route R includes position information of the vehicle 100 (image capturing unit 150) throughout the entire evaluation section from the start to the end of the evaluation.
[0079] On the other hand, the estimated shooting trajectory L of the shooting unit 150 is information in chronological order of the estimated shooting positions of the shooting unit 150 acquired in image processing by the image processing unit 2034. The estimated shooting trajectory L includes information on the estimated shooting positions of the shooting unit 150 throughout the entire evaluation section from the start to the end of the evaluation. By overlaying this information, the scale of the point cloud data 2022, which is the relative positional relationship, can be estimated.
[0080] The scale estimation process may be performed as part of the image processing by the image processing unit 2034. In this case, the point cloud data 2022 with the estimated scale is acquired in the image processing by the image processing unit 2034, and the individual scale estimation process can be omitted.
[0081] The data generation unit 2037 shown in FIG. 4 fits a plane to the acquired point cloud data 2022 and renders a frame image corresponding to the plane, thereby generating bird's-eye view data 2025 showing the road surface. Fig. 11 is a diagram illustrating rendering of point cloud data 2022 onto a plane. Fig. 11A is a diagram illustrating fitting of a virtual plane to point cloud data, and Fig. 11B is a diagram illustrating the process of projecting information onto the plane.
[0082] As shown in Fig. 11A, the data generation unit 2037 uses information about the road surface area extracted by the road surface extraction unit 2035 to determine the point cloud belonging to the road surface from the point cloud data 2022. Then, the data generation unit 2037 fits a virtual plane to the point cloud belonging to the road surface. This allows a plane of a certain area in the local camera coordinate system to be estimated. As shown in Fig. 11B, by repeating this plane estimation along the traveling direction, multiple pieces of plane data SD are created.
[0083] Next, the data generation unit 2037 renders a road surface image onto each piece of created plane data SD. Specifically, rendering is performed by projecting the images captured by each camera onto the plane data SD using the estimated shooting positions and estimated shooting directions of the multiple cameras observing the plane data SD. At this time, the position of the plane data SD is fine-tuned and the images are synthesized.
[0084] Next, the data generation unit 2037 sequentially joins together the created multiple pieces of plane data SD along the traveling direction of the vehicle 100. At this time, the data generation unit 2037 fine-tunes the position of the plane data SD and performs image synthesis processing so as to join them smoothly.
[0085] The data generating unit 2037 connects together the plane data SD for the entire evaluation section to generate bird's-eye view data 2025. Details of the generated bird's-eye view data 2025 will be described later.
[0086] In this way, in the system 1, the road surface extraction unit 2035 extracts the road surface area in the frame (key frame), and uses information about the extracted road surface area to convert the point cloud data 2022 into bird's-eye view data 2025. By performing such processing, the system 1 makes it possible to generate bird's-eye view data 2025 with a certain amount of information from the sparse point cloud data 2022.
[0087] The data generating unit 2037 also generates road surface damage data 2026 by allocating the detected damage information to the generated bird's-eye view data 2025. Fig. 12 is a diagram illustrating allocation of damage data 2024 to bird's-eye view data 2025.
[0088] 12, the data generating unit 2037 divides (segments) the bird's-eye view data 2025 into predetermined unit areas (grids G). For example, a unit area may be an area measuring 50 cm square. Then, the data generating unit 2037 uses the posture information of the imaging unit 150 to allocate damage data 2024 to each grid.
[0089] At this time, the data generation unit 2037 refers to the estimated photographing position and estimated photographing direction of the photographing unit 150, and for the damage detected from each frame, determines and projects the corresponding position in the bird's-eye view data 2025, which is a different coordinate system. That is, in a frame in which damage is detected, the data generation unit 2037 converts the position of the damage specified in the plane coordinate system of that frame into a different coordinate system, that of the bird's-eye view, using the estimated shooting position and estimated shooting direction of that frame. In this way, the data generation unit 2037 projects the damage information onto the bird's-eye view (assignment for each grid). The data generation unit 2037 performs this process for all frames in which damage is detected, thereby generating road surface damage data 2026. Note that the data generation unit 2037 may assign damage data 2024 for each grid using only the estimated shooting position out of the attitude information of the imaging unit 150. The road surface damage data 2026 generated by the data generating unit 2037 will be described in detail later.
[0090] Here, since image processing is performed only on key frames, the posture of the camera (estimated shooting position and estimated shooting direction) is not obtained for all frames in which damage is detected. Therefore, the data generation unit 2037 identifies a key frame KF that is closest in time series to the frame in which damage is detected. Then, the data generation unit 2037 assigns damage information detected in the frame to each unit area in the bird's-eye view data 2025 using the estimated shooting position and estimated shooting direction obtained by image processing in the key frame that is closest in time to the frame in which damage is detected.
[0091] Furthermore, the data generating unit 2037 superimposes the generated road surface damage data 2026 on map information displaying the road surface for each evaluation section to generate a road damage map 2027 having damage information for each evaluation section. The generated road damage map 2027 will be described in detail later.
[0092] <5. System 1 Processing> Next, the processing of the system 1 will be described.
[0093] <5-1. First process> 13 is a diagram illustrating a first process executed by the system 1. In the first process, the terminal device 10 mainly takes an image of the road surface.
[0094] As shown in FIG. 13, first, the image capturing unit 150 of the terminal device 10 mounted on the traveling vehicle 100 captures a video of the road surface (step S101). Specifically, the photographing unit 150 continuously photographs a video of the road surface ahead of the traveling vehicle 100. At this time, the photographing direction of the photographing unit 150 is set so that not only the road surface but also the surrounding scenery is included as a subject of the frame. The photographing unit 150 sends photographed data 161 acquired by photographing to the input receiving unit 171 of the control unit 170. The control unit 170 stores the photographed data 161 accepted by the input receiving unit 171 in the storage unit 160. At this time, the GPS antenna 180 continuously acquires location information of the terminal device 10.
[0095] After step S101, the control unit 170 associates position information with the captured frame (step S102). Specifically, the data processing unit 173 of the control unit 170 associates positional information at which each frame constituting the shooting data 161 was shot based on time information. The data processing unit 173 stores each frame associated with the positional information in the storage unit 160 as frame data 162.
[0096] After step S102, the terminal device 10 transmits the frame data 162 to the server 20 (step S103). Specifically, the control unit 170 of the terminal device 10 transmits the frame data 162 to the server 20 via the transmitting / receiving unit 172 .
[0097] After step S103, the server 20 acquires the frame data 162 (step S201). Specifically, the control unit 203 of the server 20 receives the frame data 162 transmitted via the transmitting / receiving unit 2031 and stores it in the storage unit 202 . This completes the first process.
[0098] <5-2. Overview of Server 20 Processing> FIG. 14 is a diagram illustrating an outline of the processing executed by the server 20. As shown in FIG. The server 20 mainly performs two routes of processing. As a second process (S21), the server 20 detects damage from each frame. Furthermore, the server 20 executes a process of generating bird's-eye view data 2025 by road surface reconstruction as a third process (S22). Then, the server 20 performs processing to generate road surface damage data 2026 and a road damage map 2027 as a fourth processing step (S23). This completes the processing of the server 20. Each of these processes will now be described in detail.
[0099] <5-3. Second Processing> 15 is a diagram illustrating the second processing executed by the system 1. In the second processing, the server 20 mainly performs damage detection on the frame data 2021.
[0100] As shown in FIG. 15, the server 20 inputs each frame of the frame data 2021 into the trained model to detect damage (step S211). Specifically, the damage detection unit 2032 in the control unit 203 of the server 20 inputs each frame of the frame data 2021 stored in the memory unit 202 into the first trained model 2028, thereby acquiring damage information output from the first trained model 2028. The acquired damage information is stored in the memory unit 202. The damage detection unit 2032 executes this process for all frames to be evaluated. This completes the second process.
[0101] <5-4. Third Treatment> 16 is a diagram illustrating the third process executed by the system 1. In the third process, the server 20 mainly generates bird's-eye view data 2025 by reconstructing the road surface.
[0102] As shown in FIG. 16, first, the server 20 selects a key frame (step S221). Specifically, the frame selection unit 2033 of the server 20 performs the above-described motion estimation on the frame data 2021 stored in the storage unit 202, thereby selecting frames suitable for a three-dimensional configuration as key frames.
[0103] After step 221, the server 20 acquires point cloud data 2022 and the camera posture by image processing (step 222). Specifically, the image processing unit 2034 in the control unit 203 of the server 20 performs the above-described image processing on the selected key frames to obtain the following information. Point cloud data showing the relative positions of subjects captured in multiple keyframes2022 Information about the estimated shooting position of the shooting unit 150 at the time when each key frame was shot Information about the estimated shooting direction of the shooting unit 150 at the timing when each key frame was shot The acquired data is stored in the storage unit 202 as point cloud data 2022 and orientation data 2023.
[0104] After step 221, in parallel with step 222, the server 20 extracts a road surface area for each key frame (step 223). Specifically, the road surface extraction unit 2035 in the control unit 203 of the server 20 inputs the key frames into the second trained model 2029 to obtain data on the road surface area extracted for each key frame.
[0105] After step 222, the server 20 generates bird's-eye view data 2025 using the point cloud data 2022 (step 224). Specifically, the data generation unit 2037 in the control unit 203 of the server 20 estimates planar data from the acquired point cloud data 2022 in accordance with the procedure described above. The data generation unit 2037 generates bird's-eye view data 2025, which is a plan view of the road surface of the entire evaluation section, by rendering frame images on the estimated planar data. The generated bird's-eye view data 2025 is stored in the storage unit 202.
[0106] After step 224, the server 20 estimates the scale (step 225). Specifically, the scale estimation unit 2036 in the control unit 203 of the server 20 estimates the scale of the point cloud data 2022 by matching the travel route of the vehicle 100 with the estimated photographing trajectory of the photographing unit 150. Information about the estimated scale is stored in the storage unit 202 in association with bird's-eye view data 2025. FIG. 18 is a diagram showing a portion of the bird's-eye view data 2025.
[0107] 18, the bird's-eye view data 2025 is expressed in a coordinate system in which the evaluation section of the road surface is viewed from above. In this figure, the photographed data 161 is acquired by a vehicle traveling in the left lane. The bird's-eye view data 2025 has a data structure in which a plurality of planar data are connected along the traveling direction of the vehicle 100 from the start point to the end point of the evaluation section. In this diagram, the symbol L indicates the locus of the estimated photographing position by the photographing unit 150. Furthermore, the symbol C indicates the estimated photographing direction of the photographing unit 150 for each frame.
[0108] After step 225 shown in FIG. 16, the server 20 divides the bird's-eye view data 2025 into grids (step 226). Specifically, the data generation unit 2037 in the control unit 203 of the server 20 divides the road surface area of the bird's-eye view data 2025 into predetermined unit areas. At this time, the scale estimated by the scale estimation unit 2036 is used. As a result, divisions for each unit area are formed in the bird's-eye view data 2025. This completes the third process. <5-5. Fourth Treatment> 17 is a diagram illustrating the fourth process executed by the system 1. In the fourth process, the server 20 mainly generates road surface damage data 2026 and a road damage map 2027.
[0109] As shown in FIG. 17, first, the server 20 assigns damage information to bird's-eye view data 2025 to generate road surface damage data 2026 (step S231). Specifically, the data generation unit 2037 in the control unit 203 of the server 20 references the estimated shooting position and estimated shooting direction of the shooting unit 150 and projects the damage information onto the bird's-eye view data 2025. The generated road surface damage data 2026 is stored in the storage unit 202. Fig. 19 is a diagram showing the generated road surface damage data 2026.
[0110] As shown in FIG. 19, the road surface damage data 2026 displays grids G that divide unit areas. Damage information is assigned to the grids G of the road surface damage data 2026. In the example shown, symbol D1 indicates that there is minor damage, and symbol D2 indicates that there is severe damage. As shown in the figure, by inputting the road surface damage data 2026, in which damage information is assigned for each unit area, into a crack rate calculation model that calculates the crack rate, the crack rate in the evaluation section can be quantitatively evaluated.
[0111] After step 231, the server 20 generates a road map by superimposing the road surface damage data 2026 on the map information (step S232). Specifically, the data generation unit 2037 in the control unit 203 of the server 20 superimposes the generated road surface damage data 2026 on map information displaying the road surface for each evaluation section to create a road damage map 2027 having damage information for each evaluation section. The generated road damage map 2027 is stored in the storage unit 202. This completes the fourth process. FIG. 20 is a diagram showing the generated road damage map 2027.
[0112] 20, a road damage map 2027 has multiple pieces of road surface damage data 2026 for each evaluation section superimposed on map information. In addition, the display mode is color-coded according to the degree of damage, so that the degree of road surface damage for a road region occupying a certain area can be intuitively grasped.
[0113] FIG. 21 is an enlarged view of a portion of the road damage map 2027 shown in FIG. 20. In the example shown, the display format changes depending on the degree of damage. Specifically, in FIG. 20, the distribution of crack rates every 20 meters is displayed as a color map. In this way, by changing the display format depending on the degree of damage, the user can visually grasp the crack rate in the evaluation section.
[0114] FIG. 22 is a diagram visualizing the damage state for each unit area for a portion of the road surface shown in FIG. 20. In the road damage map 2027, the damage state for each unit area is visualized, allowing the user to check the detailed damage state for each unit area. In this display state, by specifying each grid to which damage information is assigned, the frame in which damage was actually detected at that location may be displayed, and the specific damage state may be presented to the user.
[0115] <6.Summary> As described above, in the system 1 according to this embodiment, point cloud data 2022 is acquired by image processing of frames of a video captured by the imaging unit 150 mounted on the vehicle 100, and bird's-eye view data 2025 showing the road surface is generated from the point cloud data 2022. Then, in the system 1, damage data 2024 detected from the frames is assigned to the bird's-eye view data 2025, thereby generating road surface damage data 2026 having damage information for each unit area. By using the road surface damage data 2026 generated in this manner, a user of the system 1 can efficiently evaluate the state of cracks in the road surface in the evaluation section and the proportion of cracks in the road surface (crack rate).
[0116] Furthermore, in system 1, the user does not make individual preparations such as placing multiple signs at predetermined intervals around the road surface to be evaluated, which serve as distance standards, before capturing the video. That is, in system 1, the user performs image processing using captured data 161 acquired through a simple capturing procedure in which the user simply captures the scenery seen from the window of traveling vehicle 100 using capturing unit 150.
[0117] Furthermore, the system 1 creates a road damage map 2027 having damage information for each evaluation section by overlaying the generated road surface damage data 2026 on map information that displays the road surface for each evaluation section. Therefore, by checking the road damage map 2027, it is possible to visualize from a bird's-eye view the state of road surface damage in a certain area, not just the evaluation section of a single road surface.
[0118] Furthermore, before generating bird's-eye view data 2025, system 1 extracts a road surface area from the subject captured in each frame for each of a plurality of frames included in the captured video. Then, when generating bird's-eye view data 2025, system 1 determines the point cloud that constitutes the extracted road surface area from point cloud data 2022 and fits a virtual plane to estimate a plurality of plane data from point cloud data 2022. Then, frame images are rendered on each estimated plane data. Therefore, even if the point cloud data 2022 obtained from the photographic data 161, which has the characteristics of having few features on the road surface and large changes in scale between adjacent frames in a time series, is sparse, it is possible to generate bird's-eye view data 2025 with an amount of information sufficient for practical use in road surface evaluation.
[0119] Furthermore, prior to the step of performing image processing, the system 1 selects multiple key frames from multiple frames included in the captured video. Then, when performing image processing, the system 1 performs image processing on the selected key frames. Therefore, by performing image processing only on key frames that are particularly easy to use for image processing, the load required for image processing can be reduced. Even if the point cloud data 2022 becomes sparser due to the selection of key frames, it is possible to provide the bird's-eye view data 2025 with a certain amount of information by performing rendering based on the extraction of road surface areas.
[0120] Furthermore, in the system 1, in the step of performing image processing, the estimated photographing position and estimated photographing direction of the photographing section 150 that photographed each of the multiple frames are acquired as posture information of the photographing section 150. Then, in the step of generating road surface damage data 2026, the system 1 estimates the positional relationship between a frame in which damage is detected and a plurality of key frames using the acquired estimated shooting position and estimated shooting direction. Then, the system 1 assigns damage information detected in the frame to each unit area using the identified positional relationship. In this way, by determining the position in the bird's-eye view data 2025 of the damage detected from the frame using the estimated shooting position and estimated shooting direction of the shooting unit 150 estimated by image processing, the position of the damage can be assigned to the bird's-eye view data 2025 with high accuracy.
[0121] Furthermore, in system 1, damage information is allocated to each unit area by dividing bird's-eye view data 2025 into unit areas and projecting the damage information onto bird's-eye view data 2025 with reference to the estimated photographing position and estimated photographing direction of photographing unit 150. This makes it possible to accurately convert information about the photographing position of a frame and the position of damage detected in a frame expressed in the plane coordinate system of the frame into a position in bird's-eye view data 2025, which is a different coordinate system.
[0122] Furthermore, by using the estimated shooting position and estimated shooting direction for each frame obtained by image processing rather than the position information obtained by the GPS antenna 180, the position on the bird's-eye view data 2025 can be determined with high accuracy.
[0123] Furthermore, the system 1 estimates the scale of the point cloud data 2022 by matching the travel path of the vehicle 100 obtained from the position information of the image capturing unit 150 with the estimated image capturing trajectory obtained from the estimated image capturing position of the image capturing unit 150. Then, when allocating damage information to the bird's-eye view data 2025, grid division is performed into predetermined unit areas using the estimated scale. Therefore, it is possible to accurately divide the unit areas for the bird's-eye view data 2025 generated from the relative point cloud data 2022.
[0124] Furthermore, position information acquired from the GPS antenna 180 is generally prone to contain a certain amount of error. In contrast, in System 1, the position information acquired from the GPS antenna 180 is only used for scale estimation and for overlaying road surface damage information on the road map. That is, as described above, in System 1, the estimated shooting position and estimated shooting direction for each frame acquired by image processing are used to assign the damage position to the bird's-eye view, so the accuracy of the damage position in the road surface damage data can be ensured.
[0125] <7. Variations> FIG. 23 is a diagram for explaining the processing of the data generating unit 2037 according to the modified example. As shown in Figure 23, when generating road surface damage data 2026, the data generation unit 2037 of the modified example uses the posture of the camera (estimated shooting position and estimated shooting direction) obtained from the key frames KF before and after the frame DF in which the damage was detected.
[0126] Specifically, the data generation unit 2037 identifies two key frames KF located before and after the frame DF in which damage was detected in time series. Then, the data generation unit 2037 estimates the change in the posture of the camera between the two key frames KF from the estimated shooting positions and estimated shooting directions in the two identified key frames KF, and estimates the shooting position and shooting direction of the camera at the shooting time of the damaged frame. Alternatively, the data generation unit 2037 may perform image matching between the key frame KF and the frame in which damage was confirmed. In this case, the data generation unit 2037 may associate the frame in which damage was confirmed with the point cloud and estimate the attitude (estimated shooting position and estimated shooting direction) of the imaging unit that captured the frame in which damage was detected.
[0127] <8. Other Modifications> In the above-described embodiment, an example has been shown in which the photographing unit 150 of the terminal device 10 photographs the scenery ahead of the traveling vehicle 100, but the present invention is not limited to this. For example, the above-described processes may be performed using photographed data 161 obtained by photographing the road surface behind or to the side of the traveling vehicle 100 with the photographing unit 150.
[0128] Furthermore, in the above-described embodiment, the data generating unit 2037 determines and assigns the position on the bird's-eye view data 2025 of the damage information detected from the frame by using both the estimated shooting position and the estimated shooting direction of the photographing unit 150 among the posture information of the photographing unit 150, but this is not limited to this. That is, the data generating unit 2037 may determine the position on the bird's-eye view data 2025 of the damage information detected from the frame by using at least one of the estimated shooting position and the estimated shooting direction of the photographing unit 150 among the posture information of the photographing unit 150.
[0129] In the above embodiment, the scale estimation unit 2036 estimates the scale of the point cloud data 2022 by matching the travel path of the vehicle 100 with the estimated photographing trajectory obtained from the estimated photographing position of the photographing unit 150, but this process may be omitted. That is, the data generation unit 2037 may determine the position on the bird's-eye view data 2025 of the damage information detected from the frame using only the attitude information of the photographing unit 150 acquired by image processing.
[0130] Furthermore, in the above-described embodiment, the image processing by the image processing unit 2034 and the road surface extraction processing by the road surface extraction unit 2035 are performed on key frames selected by the frame selection unit 2033, but this is not limited to this. That is, the selection of key frames may be omitted, and the image processing by the image processing unit 2034 and the road surface extraction processing by the road surface extraction unit 2035 may be performed on all frame data 2021.
[0131] In the above embodiment, the road surface area is extracted by the road surface extraction unit 2035, but this process may be omitted. In this case, when generating the bird's-eye view data 2025, the data generation unit 2037 may fit a plane to the entire point cloud data 2022, and then render a frame image for the fitted plane data.
[0132] In the above embodiment, the terminal device 10 transmits captured video data to the server 20, but this is not limiting. For example, video data captured by a drive recorder or the like may be input to the server 20 using a storage medium such as an SD card. In other words, the means by which the server 20 acquires video data can be selected arbitrarily. The vehicle used to capture video of the road surface is not limited to a vehicle dedicated to inspecting the road surface, but can be any vehicle selected, such as a delivery truck, a public transportation bus, or an ordinary passenger car.
[0133] Although several embodiments of the present disclosure have been described above, these embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and modifications are intended to be included in the scope of the inventions and their equivalents as defined in the claims, as well as in the scope and spirit of the inventions.
[0134] <Additional Notes> The matters explained in the embodiment are additionally noted below.
[0135] (Appendix 1) A program used in a computer to detect road surface damage, The computer processor A step of acquiring a video of a road surface captured from a traveling vehicle (step S201); A step of detecting damage to a road surface captured in a frame by performing classification using a trained model on a plurality of frames included in the video (step S211); performing image processing to acquire, from the plurality of frames, point cloud data indicating relative positional relationships of the subjects photographed in the frames and estimated photographing positions of the photographing units that photographed each frame (step S222); generating a bird's-eye view of the road surface using the acquired point cloud data and the images captured in the plurality of frames; A program that executes a step (step S231) of assigning detected damage information to each divided unit area using the estimated shooting position for the generated bird's-eye view, and generating road surface damage data having damage information for each unit area.
[0136] (Appendix 2) The processor further A program described in Appendix 1 that executes a step (step S232) of overlaying the generated road surface damage data on map information displaying the road surface for each evaluation section, thereby creating a road damage map having damage information for each evaluation section.
[0137] (Appendix 3) Prior to the step of generating a bird's-eye view (step S224), For each of a plurality of frames included in the captured video, a step (step S223) is executed in which a road surface area is extracted from the subject captured in each frame; A program described in Appendix 1 or 2, in which, in the step of generating a bird's-eye view (step S224), a point cloud that constitutes the extracted road surface area is identified from the point cloud data, and a virtual plane is fitted to the identified point cloud, and the road surface image captured in the frame is rendered onto planar data.
[0138] (Appendix 4) Prior to the step of performing image processing (step S222), A step (step S221) is executed in which a plurality of key frames, which are frames in which the subject changes significantly over time, are selected from a plurality of frames included in the captured moving image; 4. The program according to any one of Supplementary Notes 1 to 3, wherein in the step of performing image processing (step S222), image processing is performed on the selected key frames.
[0139] (Appendix 5) In the step of performing image processing (step S222), Acquire an estimated photographing direction together with an estimated photographing position of the photographing unit; In the step of generating road surface damage data (step S231), 5. A program according to any one of appendices 1 to 4, which estimates the positional relationship between a frame in which damage is detected and multiple key frames using the estimated shooting position and estimated shooting direction obtained, and assigns the damage information detected in the frame to each unit area using the identified positional relationship.
[0140] (Appendix 6) In the step of performing image processing (step S222), Acquire an estimated photographing direction together with an estimated photographing position of the photographing unit; In the step of generating road surface damage data (step S231), 6. A program according to any one of appendices 1 to 5, which assigns damage information to each unit area by projecting it onto a bird's-eye view divided into unit areas, with reference to the estimated shooting position and estimated shooting direction of the shooting unit.
[0141] (Appendix 7) The processor further acquiring position information of the image capturing unit that changes as the vehicle travels, together with time information; a step of associating position information with frames captured in a step of capturing a video of a road surface (step S101) based on the time of capturing the video (step S102); Prior to the step of generating road surface damage data (step S231), a step of estimating the scale of the point cloud data by combining the vehicle travel route obtained from the position information and the estimated photographing trajectory obtained from the estimated photographing position of the photographing unit (step S225); and a step of dividing the bird's-eye view into unit areas using the estimated scale (step S226).
[0142] (Appendix 8) 1. A computer-implemented method for detecting road surface damage, comprising: The computer's processor A step of capturing a video of a road surface by a camera unit mounted on a vehicle (step S201); A step of detecting damage to a road surface captured in a frame by performing classification using a trained model on a plurality of frames included in the captured video (step S211); performing image processing to acquire, from the plurality of frames, point cloud data indicating relative positional relationships of the subjects photographed in the frames and attitude information of the photographing units that photographed each of the frames (step S222); A step of generating a bird's-eye view showing the road surface by rendering the road surface image captured in the frame onto planar data fitted to the acquired point cloud data (step S224); A method that executes a step (step S231) of assigning detected damage information to each divided unit area using posture information for the generated bird's-eye view, and generating road surface damage data having damage information for each unit area.
[0143] (Appendix 9) A system including a computer for detecting road surface damage, The computer's processor A means for acquiring a video of a road surface taken from a moving vehicle; A means for detecting road surface damage captured in a frame by performing classification using a trained model on multiple frames included in the captured video; an image processing means for acquiring, from a plurality of frames, point cloud data indicating relative positional relationships of subjects photographed in the frames and estimated photographing positions of the photographing units that photographed each frame; A means for generating a bird's-eye view showing the road surface by rendering a road surface image captured in a frame onto planar data fitted to the acquired point cloud data; A system comprising: a means for assigning detected damage information to each divided unit area using the estimated shooting position for the generated bird's-eye view, and generating road surface damage data having damage information for each unit area.
[0144] (Appendix 10) A method for creating a road map in which road surface damage data having damage information for each unit area is superimposed on map information for each evaluation section, comprising: The computer's processor A step of acquiring a video of a road surface captured from a traveling vehicle (step S201); A step of detecting damage to a road surface captured in a frame by performing classification using a trained model on a plurality of frames included in the captured video (step S211); performing image processing to acquire, from the plurality of frames, point cloud data indicating relative positional relationships of the subjects photographed in the frames and estimated photographing positions of the photographing units that photographed each frame (step S222); A step of generating a bird's-eye view showing the road surface by rendering the road surface image captured in the frame onto planar data fitted to the acquired point cloud data (step S224); A step (step S231) of allocating the detected damage information to each partitioned unit area of the generated bird's-eye view using the estimated photographing position, and generating road surface damage data having damage information for each unit area; A method for executing a step (step S232) of superimposing the generated road surface damage data on map information displaying the road surface for each evaluation section to create a road map having damage information for each evaluation section.
[0145] (Appendix 11) A road map in which road surface damage data having damage information for each unit area is superimposed on map information for each evaluation section. [Explanation of symbols]
[0146] 1...Damage detection system 10...Terminal device 150...Photography department 160...Storage section 170...Control unit 171...input reception section 172...Transmitter / receiver 173...Data Processing Department 174...Display processing unit 20...Information processing server 201…Communications Department 202...Storage section 2021…Frame data 2022…Point cloud data 2023…Attitude data 2024…Damage data 2025…Bird's-eye view data 2026…Road surface damage data 2027…Road Damage Map 2028…First trained model 2029…Second trained model 203...Control unit 2031...Transmitter / receiver 2032…Damage detection unit 2033...Frame selection section 2034...Image processing unit 2035...Road surface extraction part 2036…Scale estimation section 2037…Data Generation Department
Claims
1. A program used in a computer to detect road surface damage, a processor of the computer; acquiring a video of a road surface captured from a moving vehicle; A step of detecting damage to the road surface captured in a plurality of frames included in the video by performing classification using a trained model; performing image processing to acquire, from the plurality of frames, point cloud data indicating relative positional relationships of subjects photographed in the frames and estimated photographing positions of the photographing units that photographed each of the frames; generating a bird's-eye view of the road surface using the acquired point cloud data and the images captured in the plurality of frames; A program that executes a step of assigning detected damage information to each divided unit area using the estimated shooting position for the generated bird's-eye view, and generating road surface damage data having damage information for each unit area.
2. The processor further comprises: The program of claim 1 executes a step of superimposing the generated road surface damage data on map information displaying the road surface for each evaluation section to create a road damage map having damage information for each evaluation section.
3. Prior to the step of generating the bird's-eye view, the processor extracting a road surface area from a subject captured in each of a plurality of frames included in the captured video; The program according to claim 1, wherein in the step of generating the bird's-eye view, a point cloud that constitutes the extracted road surface area is identified from the point cloud data, and a road surface image captured in the frame is rendered onto planar data that has been fitted to a virtual plane based on the identified point cloud.
4. Prior to the step of performing the image processing, the processor selecting, from a plurality of frames included in the captured moving image, a plurality of key frames which are frames in which a subject changes significantly over time; The program according to claim 1 , wherein the image processing step performs the image processing on the selected key frames.
5. In the step of performing image processing, acquiring an estimated photographing direction together with the estimated photographing position of the photographing unit; In the step of generating road surface damage data, 5. The program according to claim 4, wherein for a frame in which damage is detected, the estimated shooting position and the estimated shooting direction in the frame in which the damage is detected are estimated from the estimated shooting positions and the estimated shooting directions in the key frames located before and after the key frames in a time series, and the damage information detected in the frame is assigned to each unit area.
6. In the step of performing image processing, Acquire an estimated photographing direction together with an estimated photographing position of the photographing unit; In the step of generating road surface damage data, The program according to claim 1 , wherein the damage information is assigned to each unit area by projecting the damage information onto the bird's-eye view divided into unit areas with reference to the estimated shooting position and the estimated shooting direction of the shooting unit.
7. The processor further comprises: acquiring position information of the photographing unit that changes as the traveling vehicle travels, together with time information; a step of associating the position information with the frames captured in the step of capturing the video of the road surface based on the capturing time, Prior to the step of generating the road surface damage data, the processor a step of estimating a scale of the point cloud data by combining a travel path of the traveling vehicle obtained from the position information and an estimated photographing trajectory obtained from the estimated photographing position of the photographing unit; The program according to claim 6 , further comprising: a step of dividing the bird's-eye view into the unit areas using the estimated scale.
8. 1. A computer-implemented method for detecting road surface damage, comprising: capturing a video of a road surface by a camera unit mounted on the vehicle; A step of detecting damage to the road surface captured in a plurality of frames included in the captured video by performing classification using a trained model; performing image processing to acquire, from the plurality of frames, point cloud data indicating relative positional relationships of the subjects photographed in the frames and attitude information of the photographing unit that photographed each of the frames; generating a bird's-eye view of the road surface using the acquired point cloud data and the images captured in the plurality of frames; A method comprising: assigning detected damage information to each divided unit area using the posture information for the generated bird's-eye view, and generating road surface damage data having damage information for each unit area.
9. A system including a computer for detecting road surface damage, a processor of the computer, A means for acquiring a video of a road surface captured from a traveling vehicle, and a means for detecting damage to the road surface captured in a plurality of frames included in the captured video by performing classification using a trained model; an image processing means for acquiring, from the plurality of frames, point cloud data indicating relative positional relationships of subjects photographed in the frames and estimated photographing positions of the photographing units that photographed each of the frames; a means for generating a bird's-eye view showing the road surface using the acquired point cloud data and the images captured in the plurality of frames; A system comprising: a means for assigning detected damage information to each divided unit area of the generated bird's-eye view using the estimated shooting position, and generating road surface damage data having damage information for each unit area.
10. A method for creating a road map in which road surface damage data having damage information for each unit area is superimposed on map information for each evaluation section, comprising: The computer's processor acquiring a video of a road surface captured from a moving vehicle; A step of detecting damage to the road surface captured in a plurality of frames included in the captured video by performing classification using a trained model; performing image processing to acquire, from the plurality of frames, point cloud data indicating relative positional relationships of subjects photographed in the frames and estimated photographing positions of the photographing units that photographed each of the frames; generating a bird's-eye view of the road surface using the acquired point cloud data and the images captured in the plurality of frames; a step of allocating detected damage information for each of the partitioned unit areas of the generated bird's-eye view using the estimated photographing position, and generating road surface damage data having damage information for each of the unit areas; A method for executing a step of superimposing the generated road surface damage data on map information displaying the road surface for each evaluation section, thereby creating a road map having damage information for each evaluation section.
Citation Information
Patent Citations
Road detection method and system based on infrared binocular structured light
CN114519732A
Pavement crack analyzer, pavement crack analysis method, and pavement crack analysis program
JP2018021375A
Hot water storage type water heater
JP2019015417A
Road surface property measurement device, road surface property measurement system, road surface property measurement method and road surface property measurement program
JP2020144079A
Ortho-image creation system, ortho-image creation method, Anti-aircraft mark used therefor, and road investigation method
JP2020204602A