Image processing program, image processor, and method for processing image

The image processing program enhances deformation detection in infrastructure by using neural networks for region-specific detection and filtering, addressing the limitations of existing technologies in detecting multiple types of deformations and reducing overdetection.

JP2025137013APending Publication Date: 2025-09-19KK TOSHIBA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024035979
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-08
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing image processing technologies are limited in their ability to accurately detect multiple types of deformations in infrastructure, such as road surfaces, and often result in overdetection outside the targeted regions.

Method used

An image processing program that utilizes a neural network to limit deformation detection to specific regions based on deformation type, employing weakly supervised learning and semantic segmentation to enhance accuracy and reduce overdetection.

Benefits of technology

Accurately detects deformations in specific regions, reducing overdetection and improving detection precision by filtering results to match the targeted areas, enabling high-accuracy deformation identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025137013000001_ABST
    Figure 2025137013000001_ABST
Patent Text Reader

Abstract

To provide an image processing program for accurately detecting abnormalities of each region from a plurality of regions.SOLUTION: The image processing program according to the present embodiment causes a computer to execute: a region limiting procedure for limiting the position of an abnormality according to an abnormality type on the basis of information representing the position of the abnormality detected from the image by using a neural network and the abnormality type; and an information output procedure for outputting at least one of an image, the information representing the position of the abnormality, the abnormality type, and information representing the limited position of the abnormality.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing program, an image processing device, and an image processing method. [Background technology]

[0002] Progress is being made in automating the inspection of infrastructure such as roads and power equipment. To automate inspections, technology is being developed that detects abnormalities (damage or abnormal conditions) in the inspection target from images.

[0003] For example, there is a road surface anomaly detection technology that detects anomalies in road surfaces from camera images. Specifically, to reduce the effort required for teaching, a machine learning model is generated using weakly supervised learning, which teaches only the presence or absence of anomalies in the image, and the trained model is then used to detect the location of anomalies in the image.

[0004] There is also a known technology that uses a camera to capture images of the road surface and uses a trained model to detect the presence or absence of road surface damage. However, this technology is limited to detecting abnormalities within the road area and is unable to detect multiple types of abnormalities. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Patent No. 6955295 [Non-patent literature]

[0006] [Non-Patent Document 1] Satoshi Ito, "Detecting Multiple Types of Road Deformation Using Multiple Instance Learning," ViEW Vision Technology Practical Use Workshop, OS1-O3, December 2022.<URL:https: / / view.tc-iaip.org / view / 2022 / > [Non-patent document 2] Toshiba Corporation and Toshiba Digital Solutions Corporation, "Development of road surface anomaly detection AI that contributes to the maintenance and long-term operation of expressways, demonstrating real-time detection of potholes that could lead to serious accidents - Promoting faster and more advanced daily expressway inspections with AI that can easily detect a variety of anomalies and be implemented on different roads", public press release,<URL:https: / / www.global.toshiba / jp / technology / corporate / rdc / rd / topics / 23 / 2309-01.html> Summary of the Invention [Problem to be solved by the invention]

[0007] The problem to be solved by the present invention is to provide an image processing program that can accurately detect deformations in each of a plurality of regions. [Means for solving the problem]

[0008] The image processing program of the embodiment causes a computer to execute an area limitation procedure that limits the location of the deformation based on information representing the location of the deformation detected from an image using a neural network and the type of deformation, according to the type of deformation, and an information output procedure that outputs at least one of an image, information representing the location of the deformation, and information representing the location of the limited deformation. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a diagram showing an example of a road surface deformation detection system 1 according to an embodiment. [Figure 2] FIG. 10 is a diagram showing an example of teaching data in weakly supervised learning according to the embodiment. [Figure 3] A diagram explaining Multiple Instance Learning for images. [Figure 4] A diagram showing a method for learning a first inference model in an embodiment. [Figure 5] FIG. 4 is a diagram showing an example of a region limiting unit in the embodiment. [Figure 6] 5A to 5C are diagrams showing an example of processing by an area detection unit in the embodiment. [Figure 7] 5A and 5B are diagrams showing examples of outputs from an information output unit in the embodiment. [Figure 8] 5A and 5B are diagrams showing examples of outputs from an information output unit in the embodiment. [Figure 9] 5A and 5B are diagrams showing examples of outputs from an information output unit in the embodiment. [Figure 10] FIG. 6 is a diagram showing an example of processing by an information output unit in the embodiment. [Figure 11] 5A and 5B are diagrams showing examples of outputs from an information output unit in the embodiment. [Figure 12] 5A and 5B are diagrams showing examples of outputs from an information output unit in the embodiment. [Figure 13] 10 is a flowchart illustrating an example of image processing executed by the image processing apparatus according to the embodiment. [Figure 14] FIG. 4 is a diagram showing an example of a display screen according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of the image processing program, apparatus, and method according to the present invention will be described in detail below with reference to the accompanying drawings.

[0011] In this embodiment, an example will be described in which an expressway is used as an inspection target and a road surface deformation detection system 1 is used to detect road surface deformation as a detection target. However, the inspection target and the detection target are not limited to this.

[0012] 1 is a diagram showing an example of a road surface deformation detection system 1 according to an embodiment. The road surface deformation detection system 1 includes an imaging device 22 and an image processing device 10, for example.

[0013] The imaging device 22 may be provided on a mobile object, or may be a camera that can be carried by a user. In this embodiment, an example in which the imaging device 22 is provided on a mobile object will be described.

[0014] The image processing device 10 acquires captured images from the imaging device 22. The image processing device 10 is an image processing device that detects abnormalities in the road surface, etc. from the captured images. The image processing device 10 has an image acquisition unit 40A, a abnormality detection unit 40B, an area limitation unit 40C, a storage unit 40D, and an information output unit 40E.

[0015] The image acquisition unit 40A acquires the captured image taken by the imaging device 22. Thereafter, the acquired captured image is output to the deformation detection unit 40B and the area limitation unit 40C. The image may also be output to the information output unit 40E. The image acquisition unit 40A may also be configured to acquire the captured image by accessing a database of images acquired by the imaging device 22. The captured image may also be acquired from a portable storage medium such as a memory card.

[0016] The deformation detection unit 40B processes the image input from the image acquisition unit 40A. The image processing detects the deformation area and the deformation type. That is, the deformation detection unit 40B executes a detection process for the detection target on the captured image acquired by the image acquisition unit 40A. For example, the deformation detection unit 40B detects a deformation on the image acquired by the image acquisition unit 40A. When a deformation is detected on the image, the deformation area and the deformation type, which are deformation information related to the deformation, are output to the area limitation unit 40C.

[0017] A defect is an abnormality that appears on an object being inspected. For example, a defect on a road surface is a crack, rut, pothole, sinkhole, depression, step, etc. that has appeared on the road surface. For example, a defect on a sound insulation wall is damage, crack, rust, bolt abnormality, etc. The severity of the defect may be defined according to its size, depth, etc. For example, the severity of a pothole may be defined as AA or A depending on its size and depth.

[0018] The deformation information is information related to the deformation, and indicates the deformation type and deformation area, for example. For example, if a pothole is detected on the road surface in the captured image, "pothole" is output as the deformation type. If damage is detected on a sound insulation wall, "damage to sound insulation wall" is output as the deformation type.

[0019] The deformation detection unit 40B may also output a likelihood map as deformation information. The likelihood map may be output for each type of deformation. The information indicating the position of the deformation includes information on the deformation area and the likelihood map.

[0020] A likelihood map is a score map that indicates the likelihood of an anomaly at each position in an image. For example, the vertical and horizontal lengths of the likelihood map are the same as those of the input image. For example, the score takes a continuous value from 0 to 1, and the larger the value, the more likely a particular anomaly is. The likelihood map makes it possible to visualize which positions in the input image are likely to be anomalies.

[0021] In addition to continuous values, the likelihood map values ​​may be binarized and output, with values ​​equal to or greater than a predetermined threshold set to 1 and values ​​other than that set to 0. The predetermined threshold may be input by the user, or a fixed value may be set in advance. A different predetermined threshold may be set for each type of deformation.

[0022] For example, when a pothole on the road surface and damage to a sound barrier are detected in the captured image, the abnormality area is output as indicating that there are two abnormalities.

[0023] The deformation detection unit 40B performs a detection process on the captured image acquired by the image acquisition unit 40A using the first inference model.

[0024] The first inference model may be a rule-based model or an inference model constructed by machine learning. For example, the first inference model is a neural network. For example, the first inference model is a neural network trained by supervised learning. The type of deformation may be output from the neural network or input from an external source.

[0025] The first inference model may be, for example, a method of calculating a likelihood map of an input image based on the difference in feature values ​​when an input image and a normal image are input into a neural network that has been pre-trained with a large number of images. Here, the normal image is an image that depicts the normal state of the inspection target. The pre-trained neural network may be, for example, a neural network trained using the ImageNet dataset, which is a technology described below.

[0026] ImageNet: A large-scale hierarchical image database, CVPR 2009, pp.248-255(2009), J. Deng et al.

[0027] In addition, for example, a normal image similar to the input image is searched for from a group of normal images collected in advance and used as a reference image, and the difference in feature values ​​obtained from a trained neural network between the input image and the reference image is calculated to calculate a likelihood map. For example, the following technology can be used.

[0028] Sub-Image Anomaly Detection with Deep Pyramid Correspondences, arXiv:2005.02357(2020), N. Cohen and Y. Hoshen.

[0029] Furthermore, as in the above, the difference in feature amounts between the input image and the reference image is calculated to calculate the likelihood map. However, for example, in the technology described below, the difference in feature amounts between the reference images is further calculated to calculate a normal image variation map. The normal image variation map is a score map that indicates the variation of the normal pattern at each position of the image. The likelihood map is corrected using the normal image variation map. For example, the technology described below can be used.

[0030] Unsupervised Anomaly Localization Using Locally Adaptive Query-Dependent Scores, ICIAP 2022(LNCS, vol.13232), pp.300-311(2022), N. Kawamura.

[0031] Furthermore, for example, a learning data set may be prepared that combines an image of a road with a circumscribing rectangle indicating the position of a deformation in the input image as a teaching label, and the first inference model may be trained using this data set.

[0032] However, preparing a circumscribing rectangle of the deformed part as a teaching label poses problems such as the teaching work taking time and the fact that there are individual differences in the judgment of the deformed part depending on the teaching worker.

[0033] Therefore, the first inference model may be trained using weakly supervised learning that combines only the image and the presence or absence of anomalies shown in the image as teaching labels. As a specific example, the first inference model performs weakly supervised learning using data in which images are assigned multi-labels of 0 or 1 to indicate the presence or absence of multiple types of anomalies.

[0034] Figure 2 shows an example of training data for images and training labels in weakly supervised learning. For example, the road surface image 31 in Figure 2 indicates that "it does not contain vertical or horizontal cracks, but it does contain tortoiseshell cracks and potholes." The Multiple Instance Learning framework can be applied to a setting where only the above-mentioned images and the presence or absence of abnormalities in the images are used as training labels.

[0035] Multiple-Instance learning is a type of supervised learning. In general supervised learning, a label y is assigned to each sample x. In contrast, in Multiple-Instance learning, a correct label is assigned to a set of multiple instances. In Multiple-Instance learning, this set is called a bag. For example, let's explain the correct label for a two-class classification task. If all instances in the bag are negative, it is considered a negative example and is assigned the label y=0, and if there is at least one positive instance in the bag, it is considered a positive example and is assigned the label y=1.

[0036] Figure 3 is a diagram explaining Multiple Instance Learning for an image. Consider the case where there is only one type of anomaly to be detected. If anomaly detection for a road image is interpreted as Multiple Instance Learning, each region (i, j) in the image corresponds to an instance. Furthermore, the image corresponds to the set bag that groups together the instances.

[0037] Figure 3 shows the relationship between the label t_ij for each region (i, j) in an image and the label y for the image. Here, 1 indicates that there is an abnormality, and 0 indicates that there is no abnormality. An image 401 that includes even one region with an abnormality is assigned the label y=1 indicating that there is an abnormality, and an image 402 in which all regions are normal is assigned the label y=0 indicating that there is no abnormality. In other words, as shown in equation (1), the maximum value of the label t_ij for the region and the label y for the image are equal.

[0038] y=max ij t ij ···(1)

[0039] Next, Figure 4 shows an example of a training method for the first inference model. An example of a method for training the first inference model using Multiple Instance learning is described below. For example, the first inference model is configured using a Fully Convolutional Network, which is composed of convolution operations with spatial locality. The first inference model is used to calculate a score s_ij that represents the likelihood of an abnormality for each region (i, j) in the image. Here, s_ij is a continuous value between 0 and 1, with larger values ​​indicating a greater likelihood of an abnormality. The first inference model is used to calculate scores s_ij for all regions in the image, and the maximum value is taken as the score p for the image. The score p is expressed as in equation (2). In Figure 4, K indicates the number of types of abnormalities. In other words, it corresponds to the number of score maps to be calculated. If there is only one type of abnormality to be detected, K=1.

[0040] p=max ij s ij ···(2)

[0041] Using the score p for this image and the label y for the image, BinaryCrossEntropyLoss is calculated, and the first inference model is optimized so that it decreases. BinaryCrossEntropyLoss(L) is expressed as in Equation (3).

[0042] L=-ylogp-(1-y)log(1-p)···(3)

[0043] This learning method allows the first inference model to be trained so that the score s_ij for an area is low for areas without deformation and high for areas with deformation. The above describes the case where there is one type of deformation, but if there are multiple types, it is sufficient to calculate a score map for each type of deformation. In this embodiment, GlobalMaxPooling is used to convert the score map into a score for the image. Note that the likelihood map may be binarized when output.

[0044] 5 is a diagram showing an example of details of the region limiting section in the embodiment. The region limiting section 40C includes an image region estimating section 41 and a filtering section .

[0045] The image region estimation unit 41 acquires an input image from, for example, the image acquisition unit 40A. The image region estimation unit 41 segments various regions in the input image, such as the road surface, sound barrier, and bridge, and outputs a region estimation image containing region labels of the input image as region estimation results to the filtering unit 42. The type of deformation may also be input from an input unit selectable by the user.

[0046] The filtering unit 42, for example, acquires the deformation area and deformation type from the deformation detection unit 40B. It also acquires an area estimation image from the image area estimation unit 41. The filtering unit 42 uses the deformation area, deformation type, and area estimation image to limit the area to a specific deformation area. After that, the filtering unit 42 outputs deformation information, such as filtered images filtered for each area, to the memory unit 40D. Note that the filtering unit 42 may also output the information to the information output unit 40E.

[0047] 6 is a diagram showing a specific example of the processing by the area limiting unit 40C. For example, when taking an image on a highway and wanting to identify a deformation of a sound insulating wall, the deformation detecting unit 40B may overdetect anomalies in the road surface and other areas other than the sound insulating wall as the deformation area.

[0048] The image area estimation unit 41 segments the sound insulating wall from the input image and outputs an area estimation image to the filtering unit 42. Furthermore, the deformation detection unit 40B outputs the detected deformation type and the deformation area including overdetection to the filtering unit 42. The filtering unit 42 then identifies deformations within the sound insulating wall area using the information input from the image area estimation unit 41 and the deformation detection unit 40B.

[0049] The image region estimation unit 41 performs segmentation on the input image and divides the region labels of the input image into regions that can be divided. Then, it outputs the region estimation image as the region estimation result to the filtering unit 42. A second inference model is used for this region division.

[0050] The second inference model can be, for example, a semantic segmentation model. Semantic segmentation is a type of image recognition technology that estimates the label or category of each pixel in an input image, such as what is depicted in the image. Examples of labels or categories include a sound barrier, road surface, or guardrail.

[0051] Furthermore, the region limiting unit 40C performs segmentation on the image using a second inference model that divides the region labels of the image that can be divided into regions, and extracts regions from the image according to the type of deformation.

[0052] The semantic segmentation model can use the following technologies. Note that any technology that can estimate labels and categories on a pixel-by-pixel basis is sufficient.

[0053] U-Net:Convolutional Networks for Biomedical Image Segmentation,arxiv.org,1505.04597. <URL:https: / / arxiv.org / abs / 1505.04597>

[0054] SAM:Segment Anything ,Meta AI. <URL:https: / / segment-anything.com / >

[0055] Context Net:Exploring Context and Detail for Semantic Segmentation in Real-time,arxiv.org,1805.04554. <URL:https: / / arxiv.org / abs / 1805.04554>

[0056] As shown in FIG. 6, the region limiting unit 40C performs a region estimation process on the input image in the image region estimation unit 41, and obtains region estimation results of labels and categories of each region in the image.

[0057] The region estimation image in the image region estimation unit 41 is an example of a region estimation result expressed as an image. For example, the input image is divided into regions so that each region can be identified, with sound barriers colored orange, road surfaces colored blue, guardrails colored green, cars colored white, and everything else colored black. Note that the region estimation image is an example of a representation of the region estimation result, and is not limited to this, as long as it can identify the label of each region of the input image.

[0058] Next, the filtering unit 42 filters the area estimation image and the deformation type and deformation area output from the deformation detection unit 40B. The deformation area and deformation type are limited to match the area of ​​the division information from the image area estimation unit 41. The deformation area is also filtered from the image according to information indicating the position of the limited deformation.

[0059] First, the region estimation image is divided by the type of deformation. For example, if the type of deformation is a crack in a sound insulation wall, only the sound insulation wall region is extracted from the region estimation image. If the type of deformation is a damage to the road surface, only the road surface region is extracted from the region estimation image. Here, the sound insulation wall region is divided into M wall , road surface area is M road Let's say.

[0060] Next, the score of the likelihood map shown in the deformation area is set to S. If the likelihood map is, for example, a binary map that assigns a value of 1 to positions where a deformation is detected and a value of 0 to other positions where no deformation is detected, the filtering unit 42 performs an AND operation on each of the areas divided by deformation type and the deformation area.

[0061] The filtered map is called S', and the crack filtering map of the sound insulation wall is called S' crack , the road surface damage filtering map is S′ pothole In this case, each region is expressed as follows:

[0062] S' crack =M wall ∩S

[0063] S' pothole =M road ∩S

[0064] Filtering map S' of cracks in a noise barrier crack For example, the deformation area is limited to the deformation within the sound barrier area, and the deformation detected outside the sound barrier area can be removed. pothole In this case, the abnormality area is limited to the abnormalities within the road surface area, and abnormalities detected outside the road surface area can be removed. The same is true for the case of a guardrail. Therefore, overdetection detected in other areas can be removed, thereby suppressing overdetection.

[0065] Furthermore, even if the likelihood map is not a binary map but is expressed as a continuous score between 0 and 1, filtering can be performed using the following formula. Note that A○B represents multiplication for each element. Note that "○" represents the Hadamard product.

[0066] S' crack =M wall ○S crack

[0067] S' pothole =M road ○S pothole

[0068] The storage unit 40D acquires data such as images relating to the deformation from the region limiting unit 40C. The storage unit 40D sends the acquired data to the information output unit 40E. Note that the information output unit 40E may acquire the data from the region limiting unit 40C.

[0069] The information output unit 40E outputs the deformation information, text, images related to the deformation, etc. based on the acquired results to a display unit such as a display screen. The output pattern here is not important as long as the results are clear. For example, the output information may include at least one of an image, information indicating the position of the deformation, the type of deformation, detection information, and area limitation information. Note that outputting includes temporarily storing and displaying.

[0070] 7 to 9 show examples of display screens of the detection results acquired by the information output unit 40E in this embodiment. The output results shown below may be displayed on a display unit provided in the information output unit 40E, or may be output to an external display unit and then displayed.

[0071] Figure 7 shows the case where only text is output. The area name, type of deformation, and coordinates of the deformation are output from the results of the filtering unit 42. In addition, information about the deformation and the parameters used during detection may also be displayed.

[0072] Next, Fig. 8 is an example of a binarized display of the likelihood map of the deformation area. As shown in Fig. 8, the information output unit 40E outputs the input image, the area estimation image output for the input image using semantic segmentation, and an image of the deformation area showing the deformation in the image. Furthermore, images divided into each area according to the type of deformation are output. Specifically, an image divided into sound barriers and road surfaces is output. Then, an output deformation image for each limited area is output. The text shown in Fig. 7 is also output.

[0073] Figure 9 is an example in which the likelihood map of the deformation area is displayed as a continuous value between 0 and 1 without being binarized. In addition to the output results shown in Figure 8, Figure 9 also allows you to set a deformation score threshold for the image of the deformation area that indicates the deformation within the image. An example is shown in which the deformation score threshold is set to 0.300. It is possible to adjust the output of deformation areas with a certain score or higher based on the likelihood map. The deformation score is also output for the output deformation image.

[0074] Next, Fig. 10 is a diagram showing an example of the synthesis process of the filtering map of the information output unit 40E. As shown in Fig. 10, the output deformation image is output by synthesizing the input image and the output deformation area by the filtering unit 42 as shown in Figs. 8 and 9.

[0075] For example, the input image I input ,Output deformation image I output Then, the crack filtering map S′ of the sound insulation wall crack and the road surface damage filtering map S′ pothole Since and are combined, they are expressed as follows:

[0076] I output =I input +S′ crack

[0077] I output =I input +S′ pothole

[0078] At this time, the information output unit 40E outputs the input image I input You can also make edits such as darkening the image overall. In this case, you can display the area name, deformation type, deformation coordinates, and maximum likelihood map score of the output data displayed as text on the screen. Alternatively, you can output a file containing the set deformation score threshold and the text output data separately from the image. The file here can be a txt, json, binary file, or any other file that can read the results.

[0079] Fig. 11 shows an example in which the information output unit 40E outputs the output results of each area together. Specifically, the results of the filtering unit 42, including the cracks in the sound insulation wall and the damage to the road surface, combined with the results of Fig. 8, are output together and superimposed on the input image. The output deformation image I in this example is output is expressed as follows: S' crack and S' potholeand may be overlaid with different colors.

[0080] I output =I input +S′ crack +S′ pothole

[0081] 12 is an example in which the likelihood map of the deformation area in FIG. 11 is displayed as a continuous value between 0 and 1 without being binarized. Specifically, the results in FIG. 9 are also output.

[0082] Note that a file containing the textually displayed deformation score threshold and output data may be output separately. In other words, any file that can read the output results will suffice.

[0083] 13 is a flowchart showing an example of image processing executed by the image processing device 10 according to the embodiment. First, the image acquisition unit 40A acquires a captured image from the imaging device 22 or the like (step S100).

[0084] Next, the deformation detection unit 40B executes a deformation detection process using the first inference model for the captured image acquired by the image acquisition unit 40A. Specifically, the deformation detection process detects a deformation area and a deformation type (step S101).

[0085] Next, the deformation detection unit 40B of the image processing device 10 determines whether a deformation of the detection target has been detected in the captured image. If a deformation of the detection target has been detected, the deformation detection unit 40B transmits the detected type of deformation and deformation information to the area limiting unit 40C. In addition, the image acquisition unit 40A transmits the captured image to the area limiting unit 40C (step S102).

[0086] Next, the image region estimation unit 41 of the region limiting unit 40C divides the acquired input image into multiple regions using semantic segmentation, and then transmits the region estimation image divided into multiple regions to the filtering unit 42 of the region limiting unit 40C (step S103).

[0087] Next, the filtering unit 42 performs filtering to limit the deformation area according to the deformation type, using the deformation type and deformation area output from the deformation detection unit 40B and the area estimation image output from the image area estimation unit 41. This makes it possible to limit the deformation and deformation area of ​​the road surface and the sound insulation wall (step S104).

[0088] The filtering unit 42 then stores the text information created from the deformation information or the output deformation image, which is a combination of the input image and the output result with the deformation area limited, in the storage unit 40D, or transmits it to the information output unit 40E. The deformation information is then displayed on the display unit from the information output unit 40E. The display content displays at least one piece of information, such as text or an output deformation image. There is no limit to the number of items displayed (step S105).

[0089] In this embodiment, the deformation detection unit 40B detects the deformation type and deformation area using a first inference model. The image area estimation unit 41 of the area limitation unit 40C generates an area estimation image by dividing the input image into areas. The filtering unit 42 of the area limitation unit 40C performs filtering using the area estimation image, deformation type, and deformation area of ​​the acquired division information. This allows multiple deformation areas in the image to be accurately detected by limiting them to each area.

[0090] That is, the image processing device includes an area limiting unit 40C that limits the position of a deformation from an image according to the type of deformation based on information indicating the position of the deformation and the type of deformation detected from the image using a neural network that detects information about the deformation from the image, and an information output procedure that outputs at least one of the image, the information indicating the position of the deformation, the type of deformation, the detection information, and the area limiting information. Note that outputting may include storing or displaying.

[0091] Figure 14 is an example showing the points on the travel route where the images in which abnormalities were detected were taken. For example, in Figure 14, the inspection data item in the upper right shows four locations on the travel route where abnormalities were detected as the number of detected abnormalities. ID: 001 is selected. When selected, an image of the abnormality and detailed information are displayed in the selected abnormality item in the lower right.

[0092] By selecting "Go to detection details" in the item of the selected abnormality, a separate window will open, allowing you to transition to a screen displaying details of the detection, such as an enlarged view or a score map. You can also transition to a detection result display screen, as shown in the output example of the information output unit 40E. While FIG. 14 shows that the inspection target is limited to the road surface, the inspection target may also be selected from road surfaces, guardrails, sound barriers, etc.

[0093] In this embodiment, the road surface deformation detection system 1 has been described as an example in which an expressway is used as an inspection target, and deformations in the road surface and sound barriers are detected as detection targets. However, the inspection targets and detection targets are not limited to these. For example, inspection targets include general roads, bridges, wind turbine blades, steel towers, solar panels, power plants, and substations. Deformations that are detected include potholes, holes, peeling, cracks, depressions, rust, fallen objects, and landslides. Note that the inspection targets are not limited to roads, and may also be industrial products. This allows for quality inspections before shipment.

[0094] In this embodiment, an example in which image processing is performed on an image captured by an imaging device mounted on a moving object has been described, but the present invention is not limited to this. For example, the image processing device may perform image processing on an image captured manually or an image captured by a fixed camera.

[0095] As explained above, the image processing program according to this embodiment causes a computer to execute an area limitation procedure that limits the position of a deformation according to the type of deformation based on information indicating the position of the deformation detected from an image using a neural network that detects information related to the deformation from the image and the type of deformation, and an information output procedure that outputs at least one of the image, information indicating the position of the deformation, the type of deformation, and information indicating the position of the limited deformation. Note that outputting includes storing and displaying.

[0096] This allows the location of the deformation to be limited according to the type of deformation based on the information indicating the position of the deformation and the type of deformation, thereby limiting the deformation to each region in the image, preventing overdetection outside the region and improving detection accuracy.

[0097] Furthermore, the information representing the position of the abnormality in the image processing program according to this embodiment is characterized by including a likelihood map representing the likelihood of an abnormality at each position in the image.

[0098] This allows the deformation to be detected with high accuracy by representing the deformation in a likelihood map.

[0099] Furthermore, the information representing the position of the deformation in the image processing program according to this embodiment is characterized by including information on the deformation area.

[0100] This makes it easier for users to check the deformation area by representing the deformation area.

[0101] In addition, the neural network of the image processing program according to this embodiment is characterized by detecting information about deformations from images by learning from data that instructs the presence or absence of various types of deformations in predetermined units for the image, and outputting likelihoods in units smaller than the predetermined units.

[0102] This allows the system to learn from the data that indicates whether or not there is an abnormality, making it possible to accurately detect an abnormality from an image.

[0103] In addition, the area limitation procedure of the image processing program of this embodiment is characterized by using a second inference model that divides the area labels of a divisible image into areas, performing segmentation on the image, and extracting areas from the image according to the type of deformation.

[0104] This allows the second inference model to be used to divide the image and extract areas according to the type of anomaly, preventing overdetection outside the area and improving detection accuracy.

[0105] Furthermore, the area limiting procedure of the image processing program according to this embodiment is characterized in that the area of ​​the abnormality is filtered from the image according to information representing the position of the abnormality to be limited.

[0106] This allows filtering for each area, and allows the deformation to be limited to each area.

[0107] In addition, the image processing program according to this embodiment executes an information output procedure that outputs at least one of an image, information representing the position of the deformation, the type of deformation, detection information, and area limitation information.

[0108] This allows the user to easily grasp various pieces of information.

[0109] The image processing program according to this embodiment executes an input procedure for the type of deformation that can be selected by the user.

[0110] This allows you to select the appropriate type of deformation depending on the object you want to detect.

[0111] In addition, the image processing device of this embodiment is equipped with an area limitation unit that limits the position of the deformation according to the type of deformation based on information representing the position of the deformation detected from the image using a neural network that detects information regarding the deformation from the image, and the type of deformation, and an information output unit that outputs at least one of the image, information representing the position of the deformation, the type of deformation, and information representing the position of the limited deformation.

[0112] In addition, the image processing method of this embodiment includes an area limitation process that limits the position of the deformation according to the type of deformation based on information representing the position of the deformation detected from the image and the type of deformation using a neural network that detects information about the deformation from the image, and an information output process that outputs at least one of the image, information representing the position of the deformation, the type of deformation, and information representing the position of the limited deformation.

[0113] The present invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be created by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined. [Explanation of symbols]

[0114] 10: Image processing device 40A: Image acquisition unit 40B: Deformation detection unit 40C: Area limited part 40D: Storage section 40E: Information output section 41: Image area estimation unit 42: Filtering section

Claims

1. On the computer, A region limiting procedure for limiting the position of the deformation from the image according to the type of deformation based on information representing the position of the deformation detected from the image using a neural network and the type of deformation; An information output procedure for outputting at least one of the image, information representing the position of the deformation, the deformation type, and information representing the limited position of the deformation; An image processing program for executing the above.

2. The information representing the position of the deformation includes a likelihood map representing the likelihood of deformation at each position of the image, The image processing program according to claim 1 .

3. The information representing the position of the deformation includes information on the deformation area, The image processing program according to claim 1 .

4. The neural network is characterized in that it learns with data that teaches the presence or absence of each type of deformation in a predetermined unit for the image, and detects information about the deformation from the image by outputting a likelihood in a unit smaller than the predetermined unit. The image processing program according to claim 1 .

5. The region limitation procedure uses a second inference model that divides the region label of the divisible image, performs segmentation on the image, and extracts a region from the image according to the deformation type. The image processing program according to claim 1 .

6. The region limitation procedure is characterized in that the deformation region is filtered from the image according to information representing the position of the deformation to be limited. The image processing program according to claim 1 .

7. The deformation type is an input procedure selectable by a user; 2. The image processing program according to claim 1, for executing the above.

8. An area limiting unit that limits the position of the deformation according to the type of deformation based on information representing the position of the deformation detected from the image using a neural network and the type of deformation; An information output unit that outputs at least one of the image, information representing the position of the deformation, the deformation type, and information representing the limited position of the deformation; An image processing device comprising:

9. A region limiting step of limiting the position of the deformation according to the type of deformation based on information representing the position of the deformation detected from the image using a neural network and the type of deformation; An information output process for outputting at least one of the image, information representing the position of the deformation, the deformation type, and information representing the limited position of the deformation; An image processing method comprising:

Citation Information

Patent Citations

  • Inspection support system, server device, inspection support method, and inspection support program

    JP2021022257A

  • Deterioration state diagnosis method for structures

    JP2021032042A

  • Device, method, and program for specifying sewer damage

    JP2021156655A

  • Determination program determination method and information processing device

    JP2023045787A

  • DAMAGE ASSESSMENT DEVICE, METHOD, AND PROGRAM

    JP7429774B2