Image processing program, image processor, and method for processing image

The image processing program employs a first inference model for deformation detection and a second inference model for area detection, leveraging weakly supervised learning and semantic segmentation to improve the accuracy of road surface deformation analysis.

JP2025137151APending Publication Date: 2025-09-19KK TOSHIBA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024036197
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-08
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Conventional image processing technologies for detecting road surface deformations struggle to accurately determine the extent of deformations and improve the accuracy of damaged area detection, despite advancements in anomaly detection using machine learning.

Method used

An image processing program that utilizes a first inference model for deformation detection and a second inference model for area detection, combined with gaze information generation, to accurately identify and outline deformation areas on road surfaces using weakly supervised learning and semantic segmentation.

Benefits of technology

Enhances the precision of detecting and outlining deformation areas on road surfaces by integrating gaze information and semantic segmentation, enabling accurate detection and characterization of deformation extent and location.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025137151000001_ABST
    Figure 2025137151000001_ABST
Patent Text Reader

Abstract

To provide an image processing program capable of accurately detecting a region to detect in an image.SOLUTION: An image processing program according to the present embodiment causes a computer to execute: a gaze information generation procedure for generating gaze information for detecting the region of a detection target on the basis of detection information regarding the detection target detected from an image by using a first inference model; a region detection procedure for performing region detection by using a second inference model on the basis of the image in which the detection target is detected and the gaze information; and an information output procedure for outputting at least one of the image, the detection information, the gaze information, and the result of the region detection.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing program, an image processing device, and an image processing method. [Background technology]

[0002] In recent years, progress has been made in automating the inspection of infrastructure such as roads and power equipment. To automate inspections, technology has been developed that detects abnormalities (damage or abnormal conditions) in the inspection target from images.

[0003] For example, there is a road surface anomaly detection technology that detects anomalies in road surfaces from camera images. Specifically, to reduce the effort required for teaching, a machine learning model is generated using weakly supervised learning, which teaches only the presence or absence of anomalies in the image, and the trained model is then used to detect the location of anomalies in the image.

[0004] There is also a known technology that uses a camera to photograph road surfaces and uses a trained model to detect the presence or absence of road surface damage. Specifically, to address the issue of reduced detection accuracy due to differences in the appearance of images depending on the time of day they are taken, detection accuracy is improved by re-evaluating the presence or absence of damage using multiple images taken at different times and in the same location.

[0005] However, while conventional techniques can detect the presence or absence of deformation in detected deformation images, they have the problem that the extent of the deformation is unclear. Also, although the presence or absence of damage is reevaluated, the accuracy of detecting the damaged area cannot be improved. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Patent No. 6955295 [Non-patent literature]

[0007] [Non-Patent Document 1] Ito, “Detecting Multiple Types of Road Deformation Using Multiple Instance Learning,” ViEW Vision Technology Practical Use Workshop, OS1-O3, Dec. 2022.<URL:https: / / view.tc-iaip.org / view / 2022 / > [Non-patent document 2] Toshiba Corporation and Toshiba Digital Solutions Corporation, "Development of road surface defect detection AI contributes to the maintenance and long-term operation of expressways, demonstrating real-time detection of potholes that could lead to serious accidents - AI that can detect a variety of defects and is easy to implement on different roads promotes faster and more sophisticated daily expressway inspections," public press release, Research and Development Center | Toshiba.<URL:https: / / www.global.toshiba / jp / technology / corporate / rdc / rd / topics / 23 / 2309-01.html> Summary of the Invention [Problem to be solved by the invention]

[0008] The problem to be solved by the present invention is to provide an image processing program that can accurately detect a detection target area in an image. [Means for solving the problem]

[0009] The image processing program of the embodiment causes a computer to execute a gaze information generation procedure that generates gaze information for area detection of a detection target based on detection information regarding the detection target detected from an image using a first inference model, an area detection procedure that performs area detection using a second inference model based on the image in which the detection target is detected and the gaze information, and an information output procedure that outputs at least one of the image, the detection information, the gaze information, and the result of the area detection. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a diagram showing an example of a road surface deformation detection system 1 according to an embodiment. [Figure 2]1 is a functional block diagram of an example of a road surface deformation detection system 1 according to an embodiment. [Figure 3] FIG. 10 is a diagram showing an example of teaching data in weakly supervised learning according to the embodiment. [Figure 4] A diagram explaining Multiple Instance Learning for images. [Figure 5] A diagram showing an example of a learning method for a first inference model in an embodiment. [Figure 6] FIG. 3 is a diagram showing an example of a gaze information generating unit in the embodiment. [Figure 7] 10A and 10B are diagrams showing an example of expressing a plurality of types of normal regions in an embodiment. [Figure 8] FIG. 10 is a diagram showing an example of a screen on which a user inputs generated text. [Figure 9] FIG. 10 is a diagram showing an example of a screen on which a user selects a language model to be used for text generation. [Figure 10] 5A to 5C are diagrams showing an example of processing by an area detection unit in the embodiment. [Figure 11] FIG. 10 is a diagram showing an example of a detection result display screen in the embodiment. [Figure 12] 10 is a flowchart illustrating an example of image processing executed by the image processing apparatus according to the embodiment. [Figure 13] FIG. 10 is a diagram showing an example of target points generated by a gaze information generating unit in the second embodiment. [Figure 14] FIG. 10 is a diagram showing an example of a selection screen for gaze information according to the embodiment. [Figure 15] FIG. 10 is a diagram showing an example of a detection result display screen in the second embodiment. [Figure 16] FIG. 4 is a diagram showing an example of a display screen according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of an image processing apparatus according to the present invention will be described in detail below with reference to the accompanying drawings.

[0012] (First embodiment) In the first embodiment, an example will be described in which an expressway is used as an inspection target and a road surface deformation detection system 1 detects road surface deformation as a detection target. However, the inspection target and the detection target are not limited to this.

[0013] 1 is a diagram showing an example of a road surface deformation detection system 1 according to the first embodiment. The road surface deformation detection system 1 includes, as an example, an image processing device 10 and a mobile object 20. The image processing device 10 and the mobile object 20 are connected to each other wirelessly or via a wire so as to be able to communicate with each other.

[0014] The image processing device 10 is an example of an image processing device of this embodiment. The image processing device 10 is an image processing device that detects road surface abnormalities from captured images. The image processing device 10 may be configured to be provided inside a moving object 20, which will be described later.

[0015] The mobile object 20 is, for example, a mobile object used to collect captured images in the road surface abnormality detection system 1. The mobile object 20 is equipped with an imaging device 22 and transmits captured road surface images to the image processing device 10. The mobile object 20 only needs to be capable of moving at least to capture images of the detection target. For example, the mobile object 20 may be a robot, automobile, ship, flying object, or the like having a position movement mechanism. The mobile object 20 may move autonomously or may be moved by operation by a user.

[0016] The imaging device 22 does not have to be provided in the moving object 20. For example, it may be a camera that the user can carry in his / her hand. In the first embodiment, an example in which the imaging device 22 is provided in the moving object 20 will be described.

[0017] In the first embodiment, the mobile object 20 is described as an example of an automobile that moves by operation by a user. For example, the mobile object 20 is an inspection vehicle driven by a user to inspect a highway. The inspection vehicle routinely inspects the road surface of the highway and incidental features of the highway for abnormalities. The mobile object 20 is also equipped with a camera as an imaging device 22.

[0018] 2 is a functional block diagram of an example of a road surface deformation detection system 1. The road surface deformation detection system 1 includes an image processing device 10 and a mobile object 20. The image processing device 10 and the mobile object 20 are connected to each other so as to be able to communicate with each other via a network N or the like.

[0019] The mobile object 20 includes an imaging device 22, a communication unit 24, a control unit 26, and a storage unit 28. The imaging device 22, the communication unit 24, the control unit 26, and the storage unit 28 are communicably connected via a bus 29 or the like.

[0020] The imaging device 22 captures an image of an inspection target in the moving body 20. For example, the imaging device 22 is a camera, a smartphone, or the like. The imaging device 22 may be any device capable of capturing an image of an inspection target, and the location and number of devices to be installed are not limited. For example, the imaging device 22 may be installed in front, on the side, or behind the inspection vehicle. The imaging device 22 may also be installed inside and outside the inspection vehicle. The imaging device 22 does not have to be installed on the vehicle.

[0021] The control unit 26 transmits the imaging results of the imaging device 22 from the communication unit 24 to the image processing device 10. In the first embodiment, the control unit 26 transmits the imaging results of the imaging device 22 to the image processing device 10 at predetermined time intervals. That is, the control unit 26 of the moving object 20 sequentially transmits chronologically consecutive imaging results to the image processing device 10. The control unit 26 may also temporarily transmit the imaging results to the storage unit 28 for storage. In this case, the control unit 26 may transmit the imaging results stored in the storage unit 28 to the image processing device 10 all at once, regardless of the chronological order.

[0022] The storage unit 28 stores various types of data and is, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory (registered trademark), a hard disk, an optical disk, or the like.

[0023] Next, the image processing device 10 includes an image processing unit 40, a storage unit 42, an input unit 44, a display unit 46, and a communication unit 48. The image processing unit 40, the storage unit 42, the input unit 44, the display unit 46, and the communication unit 48 are communicatively connected via a bus 49 or the like.

[0024] First, the storage unit 42 stores various types of data. The storage unit 42 is, for example, a semiconductor memory element such as a RAM or a flash memory, a hard disk, an optical disk, or the like. The storage unit 42 may be a storage device provided outside the image processing device 10. The storage unit 42 may also be a storage medium that stores or temporarily stores programs and various types of information downloaded via a LAN (Local Area Network), the Internet, or the like.

[0025] The input unit 44 receives various operations from the user and is, for example, a keyboard, a pointing device such as a mouse, a microphone, and the like.

[0026] The display unit 46 displays various types of information. For example, the display unit 46 displays the captured image transmitted from the imaging device 22, the image processing results by the image processing unit 40, etc. The display unit 46 may be configured integrally with the input unit 44 and configured as a touch panel.

[0027] Furthermore, the input unit 44 and the display unit 46 may be configured as separate units from the image processing device 10. In this case, the input unit 44 and the display unit 46 may be configured to be communicably connected to the image processing device 10.

[0028] The communication unit 48 is a communication interface that can communicate with the mobile object 20 and the imaging device 22 via the network N, for example.

[0029] At least one of the storage unit 42 and the image processing unit 40 may be mounted on an external information processing device such as a server device connected via the network N and the communication unit 48.

[0030] Furthermore, at least one of the functional units included in the image processing unit 40, which will be described later, may be mounted on an external information processing device such as a server device connected to the image processing unit 40 via the network N and the communication unit 48.

[0031] The image processing unit 40 executes various types of image processing in the image processing device 10. The image processing unit 40 includes an image acquisition unit 40A, a detection unit 40B, a gaze information generation unit 40C, an area detection unit 40D, and an information output unit 40E.

[0032] The image acquisition unit 40A, the detection unit 40B, the gaze information generation unit 40C, the area detection unit 40D, and the information output unit 40E are realized by one or more processors. For example, each of the above units may be realized by having a processor such as a CPU (Central Processing Unit) execute a program, i.e., by software. Each of the above units may be realized by a processor such as a dedicated IC, i.e., by hardware. Each of the above units may be realized by a combination of software and hardware. When multiple processors are used, each processor may realize one of the units, or two or more of the units.

[0033] The image acquisition unit 40A acquires captured images of the inspection target. For example, while the mobile object 20 is traveling on the expressway to be inspected, the image acquisition unit 40A acquires images captured by the imaging device 22. The image acquisition unit 40A may acquire at least one frame image from the road surface video generated by the imaging device 22 as the image to be processed. The image acquisition unit 40A may also be configured to access a database of videos captured by the imaging device 22 to acquire the captured images. The captured images may also be acquired from a portable storage medium such as a memory card.

[0034] The detection unit 40B executes a detection process for the detection target on the captured image acquired by the image acquisition unit 40A. For example, the detection unit 40B detects a deformation of the road surface on the image of the road surface acquired by the image acquisition unit 40A. When a deformation of the road surface is detected on the image, the detection unit 40B outputs deformation information on the deformation. The deformation information may also be temporarily stored in the memory unit 42.

[0035] A deformation is an abnormality that appears on an object being inspected. For example, a road surface deformation is a crack, rut, pothole, sinkhole, depression, step, etc. that has appeared on the road surface. The severity of the deformation may be defined according to its size, depth, etc. For example, the severity of a pothole may be defined as AA or A depending on its size and depth.

[0036] The detection unit 40B also outputs deformation information. Deformation information is information related to deformation. For example, the detection unit 40B outputs the deformation type as deformation information. For example, if the detection unit 40B detects a pothole on the road surface captured in the captured image, it outputs "pothole" as the deformation type. The detection unit 40B may also output "AA pothole" as the deformation type, with the severity level added.

[0037] The detection unit 40B may also output a likelihood map as the abnormality information. A likelihood map is a score map that indicates the likelihood of an abnormality at each position in the image. For example, the vertical and horizontal lengths of the likelihood map are the same as those of the input image. For example, the score takes on a continuous value from 0 to 1, with larger values ​​indicating a more likely specific abnormality. The likelihood map can visualize which positions in the input image are likely to be abnormal.

[0038] Furthermore, the likelihood map values ​​may be binarized and output as deformation information, with values ​​equal to or greater than a predetermined threshold being set to 1 and values ​​other than that being set to 0. Alternatively, the predetermined threshold may be input by the user, or a fixed value may be set in advance. Different predetermined thresholds may be set for different types of deformation.

[0039] In the first embodiment, an example will be described in which the type of deformation is output as the deformation information. The detection unit 40B executes a detection process on the captured image acquired by the image acquisition unit 40A using a first inference model.

[0040] The first inference model may be a rule-based inference model or an inference model constructed by machine learning. For example, the first inference model is a neural network. For example, the first inference model is a neural network trained by supervised learning.

[0041] The first inference model may be, for example, a method of calculating a likelihood map of an input image based on the difference in features when the input image and a normal image are input into a neural network that has been pre-trained with a large number of images. Here, the normal image is an image that depicts the normal state of the inspection target. The pre-trained neural network may be, for example, a neural network trained using the ImageNet dataset, which is a technology described below.

[0042] ImageNet: A large-scale hierarchical image database, CVPR 2009, pp.248-255(2009), J. Deng et al.

[0043] In addition, for example, a normal image similar to the input image is searched for from a group of normal images collected in advance and used as a reference image, and the difference in feature values ​​obtained from a trained neural network between the input image and the reference image is calculated to calculate a likelihood map. For example, the following technology can be used.

[0044] Sub-Image Anomaly Detection with Deep Pyramid Correspondences, arXiv:2005.02357(2020), N. Cohen and Y. Hoshen.

[0045] Furthermore, as in the above, the difference in feature amounts between the input image and the reference image is calculated to calculate the likelihood map. However, for example, in the technology described below, the difference in feature amounts between the reference images is further calculated to calculate a normal image variation map. The normal image variation map is a score map that indicates the variation of the normal pattern at each position of the image. The likelihood map is corrected using the normal image variation map. For example, the technology described below can be used.

[0046] Unsupervised Anomaly Localization Using Locally Adaptive Query-Dependent Scores, ICIAP 2022(LNCS, vol.13232), pp.300-311(2022), N. Kawamura.

[0047] Furthermore, for example, a learning data set may be prepared that combines an image of a road with a circumscribing rectangle indicating the position of a deformation in the input image as a teaching label, and the first inference model may be trained using this data set.

[0048] However, preparing circumscribing rectangles that indicate the location of the abnormality as teaching labels takes time for the teaching work, and there are also issues such as individual differences in the way teaching workers judge the abnormality.

[0049] Therefore, the first inference model may be trained using weakly supervised learning that combines only the image and the presence or absence of anomalies shown in the image as teaching labels. As a specific example, the first inference model is weakly supervised learning that uses data in which images are assigned multi-labels of 0 or 1 to indicate the presence or absence of multiple types of anomalies.

[0050] Figure 3 shows an example of training data for images and training labels in weakly supervised learning. For example, the road surface image 31 in Figure 3 indicates that "it does not contain vertical or horizontal cracks, but it does contain tortoiseshell cracks and potholes." In this case, it does not indicate which parts of the road surface image 31 are deformed. The Multiple Instance Learning framework can be applied to a setting where only the above-mentioned images and the presence or absence of deformities shown in the images are used as training labels.

[0051] Multiple-Instance learning is a type of supervised learning. In general supervised learning, a label y is assigned to each sample x. In contrast, in Multiple-Instance learning, a correct label is assigned to a set of multiple instances. In Multiple-Instance learning, this set is called a bag. For example, let's explain the correct label for a two-class classification task. If all instances in the bag are negative, it is considered a negative example and is assigned the label y=0, and if there is at least one positive instance in the bag, it is considered a positive example and is assigned the label y=1.

[0052] Figure 4 is a diagram explaining Multiple Instance Learning for an image. Consider the case where there is only one type of anomaly to be detected. If anomaly detection for a road image is interpreted as Multiple Instance Learning, each region (i, j) in the image corresponds to an instance. Here, i is an index indicating the horizontal position in the image, and j is an index indicating the horizontal position in the image. For example, each region (i, j) may be in pixel units, or in patch units obtained by dividing the image. Furthermore, the image corresponds to a set bag that collects instances.

[0053] Figure 4 shows the relationship between the label t_ij for each region (i, j) in an image and the label y for the image. Here, 1 indicates that there is an abnormality, and 0 indicates that there is no abnormality. An image 401 that includes even one region with an abnormality is assigned the label y=1 indicating that there is an abnormality, and an image 402 in which all regions are normal is assigned the label y=0 indicating that there is no abnormality. In other words, as shown in equation (1), the maximum value of the label t_ij for the region and the label y for the image are equal.

[0054] y=max ij t ij ···(1)

[0055] Figure 5 shows an example of a training method for the first inference model. An example of a method for training the first inference model using Multiple Instance learning is described below. For example, the first inference model is configured using a Fully Convolutional Network, which is made up of convolution operations with spatial locality. The first inference model is used to calculate a score s_ij that represents the likelihood of an abnormality for each region (i, j) in the image. Here, s_ij is a continuous value between 0 and 1, with larger values ​​indicating a greater likelihood of an abnormality. The first inference model is used to calculate scores s_ij for all regions in the image, and the maximum value is taken as the score p for the image. The score p is expressed as in equation (2). In Figure 5, K indicates the number of types of abnormalities. In other words, it corresponds to the number of score maps to be calculated. If there is only one type of abnormality to be detected, K=1.

[0056] p=max ij s ij ···(2)

[0057] Using the score p for this image and the label y for the image, BinaryCrossEntropyLoss is calculated, and the first inference model is optimized so that it decreases. BinaryCrossEntropyLoss(L) is expressed as in Equation (3).

[0058] L=-ylogp-(1-y)log(1-p)···(3)

[0059] This learning method allows the first inference model to be trained so that the score s_ij for an area is low for areas without deformation and high for areas with deformation. The above describes the case where there is one type of deformation, but if there are multiple types, it is sufficient to calculate a score map for each type of deformation. Furthermore, in the first embodiment, GlobalMaxPooling is used to convert the score map into a score for the image.

[0060] In other words, the first inference model learns from data instructed to detect whether or not there is a deformation in a specific detection target in a predetermined unit for the image, outputs likelihood in units smaller than the predetermined unit, and is trained so that the maximum value of the output likelihood matches the presence or absence of a deformation in the instructed detection target.

[0061] Returning to FIG. 2, the gaze information generation unit 40C calculates gaze information for detecting a deformed area based on the deformation information of the image in which the deformation is detected. The gaze information is transmitted to the area detection unit 40D and used to detect the deformed area. For example, the gaze information is text expressing a deformation based on the deformation type or text expressing a normal area according to the deformation type. An example of a normal area is a background area. The gaze information may also be position information based on a likelihood map. The gaze information generation unit 40C may also generate at least one of text expressing a deformation based on the deformation type and text expressing a normal area according to the deformation type.

[0062] Next, the gaze information generation unit 40C in the first embodiment will be described using Figure 6. In the first embodiment, the gaze information generation unit 40C receives the deformation type as deformation information from the detection unit 40B and generates text expressing the deformation based on the deformation type. Furthermore, in the first embodiment, the gaze information generation unit 40C generates text expressing the surrounding normal area related to the deformation type in addition to the text expressing the deformation.

[0063] The gaze information generating unit 40C generates at least one text representing a deformation and one text representing a normal area based on the type of deformation. Note that a plurality of texts may be generated. For example, a set of texts representing deformations generated by the gaze information generating unit 40C is called G anormal Let's say. G anormal has text representing N variations as elements.

[0064] G anormal =[g 1 a ,g 2 a ,g 3 a,···,g N a ]

[0065] where g 1 a is one piece of text that represents a deformation. For example, when a pothole is received as the deformation type, the gaze information generating unit 40C generates a text g that represents a pothole as follows: i a Generate g 3 a The text is generated in the same manner thereafter.

[0066] g 1 a =“a hole in a road surface”

[0067] g 2 a =“a depression in a road surface”

[0068] In addition, we define a set of texts representing the surrounding normal areas related to the anomaly type as G normal Let's say. G normal has text elements that represent the surrounding normal areas associated with the M anomaly types.

[0069] G normal =[g 1 n ,g 2 n ,g 3 n ,···,g M n ]

[0070] where g 1 n is a piece of text that expresses a surrounding normal area related to the deformation type. For example, when a pothole is received as the deformation type, the gaze information generating unit 40C generates text g that expresses a normal area related to the pothole as follows: i n Generate g 3 nThe text is generated in the same manner thereafter.

[0071] g 1 n =“smooth roadway”

[0072] g 2 n =“well-maintained road surface”

[0073] Then, the gaze information generating unit 40C performs the above G anormal and G normal is transmitted as gaze information to the region detection unit 40D. The gaze information generation unit 40C may also generate text that expresses a plurality of types of normal regions.

[0074] The detection information also includes the type of deformation of the detection target, and the gaze information generation unit 40C generates, as gaze information, text representing a deformation that characterizes at least one specific detection target and text representing normality that characterizes a background area related to at least one specific detection target based on the type of deformation of the detection target.

[0075] An example of text expressing multiple types of normal areas will be explained using FIG. 7. For example, consider a sound barrier, which is an ancillary feature of a highway, as an inspection target. For example, consider a case where a flying object attached to the sound barrier is detected as an abnormality. The detection unit 40B transmits "flying object" as the type of abnormality. For example, if a flying object is attached to the top of the sound barrier, the related normal areas include "sound barrier" and "sky." In this case, the gaze information generation unit 40C may generate text expressing a normal sound barrier and text expressing the sky. Specifically, it may be done as follows. g 3 n The text is generated in the same manner thereafter.

[0076] g 1 n =“fine soundproof wall”

[0077] g 2 n= “blue sky”

[0078] 8 is an example of a display screen on which the user inputs the generated text. The generated text may be set in advance by the user based on the inspection target. For example, when the user sets the text in advance, the user inputs the corresponding text from the input unit. As shown in FIG. 8, for example, the information output unit 40E displays a screen on the display unit 46 on which the user inputs the generated text and the corresponding background area.

[0079] 9 is an example of a display screen on which a user selects a language model to be used for text generation. The text to be generated may be generated using a large-scale language model (LLM). When using a large-scale language model (LLM), the gaze information generation unit 40C is provided with a trained large-scale language model, and generates text using the large-scale language model based on the type of deformation output by the detection unit 40B.

[0080] For example, the gaze information generation unit 40C generates multiple texts by inputting "List multiple sentences that describe potholes on a road surface" into the large-scale language model. Alternatively, the gaze information generation unit 40C may be provided with multiple large-scale language models, and the user may select which large-scale language model to use to generate text. In this case, for example, the information output unit 40E displays on the display unit 46 a screen for inputting the large-scale language model to be selected by the user.

[0081] That is, there may be a UI in which the user specifies the LLM that generates the text. Also, there may be a UI in which the user inputs the text. Note that the text input screen in Figure 8 and the text generation method selection screen in Figure 9 may be integrated.

[0082] 10 shows an example of processing by the area detection unit 40D in the first embodiment. For an image in which an abnormality has been detected by the detection unit 40B, the area detection unit 40D detects an abnormality area using the gaze information generated by the gaze information generation unit 40C. That is, when an abnormality has been detected by the detection unit 40B, the area detection unit 40D receives the image in which the abnormality has been detected from the image acquisition unit 40A, and detects an abnormality area for that image based on the gaze information. The area detection unit 40D also performs area detection using a second inference model.

[0083] Furthermore, gaze information, which is abnormality information related to the abnormality of the detection target and normality information of the background area related to the abnormality of the detection target, is sent to the area detection unit 40D. The area detection unit 40D then detects the abnormal area of ​​the detection target and the normal area of ​​the background area, and integrates the detection results to detect the abnormal area.

[0084] The second inference model may be rule-based or may be constructed using machine learning. For example, the second inference model may be a trained semantic segmentation model. Semantic segmentation is a method for predicting, for each pixel in an image, which class the pixel belongs to.

[0085] In the first embodiment, an example will be described in which a trained semantic segmentation model is used as the second inference model. For example, a case will be described in which the following segment anything model (SAM) is used as the second inference model.

[0086] SAM:Segment Anything | Meta AI (segment-anything.com) <URL:https: / / segment-anything.com / >

[0087] SAM is a basic model for area detection. SAM receives a prompt and an image and performs area detection on the image. The prompts that can be input to SAM are text, target points, and rectangles. The area detection unit 40D uses the gaze information generated by the gaze information generation unit 40C as a prompt to input to SAM.

[0088] In the first embodiment, the gaze information generating unit 40C generates gaze information by using the G of the text. anormal and G normal As described above, the set of G generated by the gaze information generating unit 40C is used as a prompt to input to the SAM. anormal has a text element that describes the deformation. normal has text elements that represent normal areas related to the abnormality.

[0089] The region detection unit 40D detects the G anormal element g of i a ,G normal element g i n Using the prompt, we detect normal regions and abnormal regions, respectively. For example, let us consider the input image as x. The text representing the abnormality is g. i a Using the prompt, segmentation is performed on the input image x using SAM. The resulting mask of the abnormal region is then denoted as M i anormal Then, we can write it as follows:

[0090] M i anormal =SAM(x,g i a )

[0091] Here, the mask indicates which pixels in the input image correspond to the abnormal area. For example, the mask is an image of the same size as the input image, and pixels detected as abnormal areas are assigned a value of 1, and other pixels are assigned a value of 0.

[0092] Also, for example, text g representing normal regions related to the abnormality i n The segmentation is performed on the input image x by SAM, and the resulting mask of the normal region is denoted as M. i normal Then, we can write it as follows:

[0093] M i normal =SAM(x,g i n )

[0094] Also, N texts G anormal The mask group of N abnormality regions obtained from the above is expressed as follows:

[0095] [M 1 anormal ,M 2 anormal ,···,M N anormal ]

[0096] Also, M texts G normal The mask set of M normal regions obtained from is expressed as follows:

[0097] [M 1 normal ,M 2 normal ,···,M M normal ]

[0098] The region detection unit 40D detects the mask group [M 1 anormal ,M 2 anormal ,···,M N anormal ] and the normal region mask group [M 1 normal ,M 2 normal ,···,M M normal ] is integrated to detect areas of deformation.

[0099] An example of a method for integrating the detection results of the abnormal region and the detection results of the normal region will be described. For example, first, the region detection unit 40D integrates the detection results of the abnormal region with the mask group [M 1 anormal ,M 2 anormal ,···,M N anormal ] to select one mask M anormal For example, the area detection unit 40D generates [M 1 anormal ,M 2 anormal ,···,M N anormal ] according to a predefined rule, and select one mask from M anormal It is also possible to do so.

[0100] For example, the mask group of the deformation region [M 1 anormal ,M 2 anormal ,···,M N anormal ], select the mask with the largest area, M anormal Alternatively, the user may select any number of masks from among N and integrate them. When the user specifies the mask to be used for integration, the user specifies the mask to be used for integration from the input unit. For example, the information output unit 40E displays the mask on the display unit 46 by superimposing it on the input image, and the user selects the mask to be used for integration while looking at the display screen.

[0101] Also, for example, [M 1 anormal ,M 2 anormal ,···,M N anormal ] can be used to take the union or intersection of all elements of M anormal can be shown as follows:

[0102] M anormal =U N i M i anormal

[0103] Next, the region detection unit 40D detects the mask M anormal Similarly, the mask group of normal regions [M 1 normal ,M 2 normal ,···,M M normal ] to the normal region mask M normal Generate.

[0104] Next, the region detection unit 40D detects the mask M anormal and the normal region mask M normal The mask M of the affected area is then integrated to detect the affected area. anormal and the normal region mask M normal As an example of how to integrate the mask M anormal Then, the mask M of the deformed area anormal and the normal region mask M normal The final mask of the deformation area is generated by removing the overlapping area. out anormal Then, we can write it as follows:

[0105] M out anormal =M anormal ∧¬M normal

[0106] Finally, the region detection unit 40D generates a final mask M out anormal is transmitted to the information output unit 40E. Alternatively, the information may be output from the area detection unit 40D to the storage unit 42 and stored therein.

[0107] While there is little training data for abnormalities, there is a lot of training data for normal road surfaces. Therefore, normal areas can be detected more accurately than abnormalities. Deformed areas can be detected more accurately by combining the detection results for abnormal areas and normal areas. However, it is also possible to detect only either abnormal areas or normal areas.

[0108] 11 shows an example of a detection result display screen in the first embodiment. The information output unit 40E outputs various information. For example, when a change in the road surface is detected by the detection unit 40B, the information output unit 40E outputs information relating to the change.

[0109] For example, the image in which a deformation is detected and the type of deformation detected may be displayed. In addition, the area detection results by the area detection unit 40D may be displayed. The detected area may be superimposed on the input image. The area detection unit 40D may also display all area detection results performed for each prompt. The gaze information generation unit 40C may also display the generated prompt. The damage type, damage level, size, area, and position may also be output. The shooting time and shooting position of the image in which a deformation is detected may also be output. At least one of the image, detection information, gaze information, and area detection results may also be output.

[0110] 12 is a flowchart showing an example of image processing executed by the image processing device 10 of the first embodiment. First, the image acquisition unit 40A acquires a captured image of the road surface or the like from the imaging device 22. Then, the image acquisition unit 40A sends the acquired captured image to the detection unit 40B (step S100).

[0111] Next, the detection unit 40B executes a deformation detection process using the first inference model on the captured image acquired from the image acquisition unit 40A, and acquires deformation information such as the type of deformation (step S101).

[0112] Next, the detection unit 40B determines whether a deformation of the detection target is detected in the captured image. If a deformation of the detection target is detected, the detection unit 40B transmits deformation information to the gaze information generation unit 40C. In addition, the image acquisition unit 40A transmits the captured image to the area detection unit 40D (step S102).

[0113] Next, the gaze information generation unit 40C generates gaze information such as text expressing the deformation and text expressing the normal area around the deformation based on the acquired deformation information such as the deformation type. At this time, either the text expressing the deformation or the text expressing the normal area around the deformation may be generated. The generated gaze information is sent to the area detection unit 40D (step S103).

[0114] Next, the area detection unit 40D executes area detection processing using the captured image acquired from the image acquisition unit 40A and the gaze information such as text generated by the gaze information generation unit 40C. Then, the area detection result is transmitted to the information output unit 40E. Note that the area detection result may be output to the storage unit 42 and stored therein (step S104).

[0115] As a result, in the first embodiment, gaze information such as text about the deformed area and the related normal area is generated from the deformation information such as the deformation type obtained using the model learned by weakly supervised learning in the detection unit 40B. Then, by detecting the deformed and normal areas and integrating the results of the area detection, the shape, coordinates, etc. of the deformed area in the image can be accurately detected.

[0116] That is, the system includes a gaze information generation unit 40C that generates gaze information for detecting the area of ​​the detection target based on detection information about the detection target detected from the image using a first inference model that detects the detection target from the image, a region detection unit 40D that performs region detection using a second inference model based on the image in which the detection target is detected and the gaze information, and an information output unit 40E that outputs at least one of the image, the detection information, the gaze information, and the result of the region detection. Note that although it is said to output, it may also be stored in a memory unit or displayed.

[0117] (Second embodiment) In the second embodiment, a detection unit 40B outputs a likelihood map as deformation information, a gaze information generation unit 40C generates position information based on the likelihood map as gaze information, and a region detection unit 40D performs region detection based on the position information.

[0118] The second embodiment differs from the first embodiment in the detection unit 40B, gaze information generation unit 40C, area detection unit 40D, and information output unit 40E of the image processing unit 40. The other functional configuration blocks are the same as those in the first embodiment, so their description will be omitted. The detection unit 40B, gaze information generation unit 40C, area detection unit 40D, and information output unit 40E in the second embodiment will be described below.

[0119] The detection unit 40B executes a detection process for the detection target on the captured image acquired by the image acquisition unit 40A. For example, the detection unit 40B detects a deformation of the road surface on the image of the road surface acquired by the image acquisition unit 40A. When a deformation of the road surface is detected on the image, the detection unit 40B outputs deformation information related to the deformation. The detection unit 40B outputs a likelihood map as the deformation information.

[0120] A likelihood map is a score map that indicates the likelihood of an abnormality at each position in an image. For example, the vertical and horizontal lengths of the likelihood map are the same as those of the input image. For example, the score takes a continuous value from 0 to 1, and the larger the value, the more likely a specific abnormality is. The likelihood map can visualize which positions in the input image are likely to be abnormal. In the second embodiment, an example of outputting a likelihood map as abnormality information will be described.

[0121] The likelihood map is obtained by calculating the score s_ij for all regions in the image using the first inference model described in the first embodiment.

[0122] The gaze information generating unit 40C calculates gaze information for detecting a modified area based on the deformation information of the image in which the deformation is detected.

[0123] In the first embodiment, the detection unit 40B outputs the deformation type as the deformation information. Furthermore, the gaze information generation unit 40C generates text representing a deformation area and text representing a normal area based on the deformation type as the gaze information.

[0124] 13 shows an example of a likelihood map and symmetric points generated by the gaze information generator 40C in the second embodiment. In the second embodiment, the gaze information generator 40C receives a likelihood map as abnormality information from the detector 40B, and generates position information of the abnormality region and position information of the normal region based on the likelihood map.

[0125] For example, the gaze information generating unit 40C generates target points and a rectangle as position information. In the second embodiment, an example of outputting target points as position information will be described. The target points indicate an area (i, j) on the likelihood map. The gaze information generating unit 40C generates a set S of target points indicating the position information of the deformation area. anormal Generate.

[0126] S anormal has target points that indicate the location information of the deformed area as elements. For example, S anormal If S has N target points as elements, then anormal =[s 1 a ,s 2 a ,s 3 a ,···,s N a ], where s i a is one of the target points indicating the position information of the abnormal region. normal Generate.

[0127] S normal has target points that indicate the position information of the normal region as elements. For example, S normal If S has M target points as elements, then normal =[s 1 n ,s 2 n ,s 3 n ,···,s M n ], where s i n is one of the target points that indicates the position information of the normal region.

[0128] 13, the gaze information generating unit 40C determines whether S anormal For example, the gaze information generating unit 40C generates S as a target point by using the coordinates with the highest likelihood in the likelihood map. anormal In addition, if there are multiple coordinates with the highest likelihood, all the coordinates are treated as target points and S anormal Alternatively, the user can specify the number of coordinates to be selected as target points, and the target points are selected in descending order of likelihood until the specified number is reached. anormal Alternatively, a user-specified number of coordinates may be randomly selected from among the coordinates having a likelihood equal to or greater than a user-specified threshold, and these coordinates may be used as target points. anormal Alternatively, the center of gravity of the coordinates with a likelihood equal to or greater than a threshold specified by the user can be used as the target point. anormal It can also be an element of.

[0129] Next, the gaze information generating unit 40C normal For example, the gaze information generating unit 40C outputs information on areas with low likelihood as target points indicating the position information of normal areas related to the abnormal area. Since likelihood is a score that indicates the likelihood of an abnormality, areas with low likelihood are considered to be normal areas. For example, as target points of normal areas related to the abnormal area, coordinates with low likelihood around areas with high likelihood are used as target points, and S normal For example, the gaze information generating unit 40C may be an element of S anormal One element of s i a The area with the nearest likelihood of 0 is taken as the target point of the normal area and is called S normal In addition, coordinates less than the threshold value that exist in the vicinity of coordinates with likelihoods equal to or greater than the threshold value specified by the user are set as target points in the normal region and are included in S normal It can also be an element of.

[0130] FIG. 13 shows, for example, the target point s 1 a ,s 2 a ,s 3 aThe central ellipse surrounded by indicates the deformation. Around the deformation, target points s indicating the position information of the deformation area are 1 a ,s 2 a ,s 3 a The target point s indicating the position information of the normal area is displayed outside of it. 1 n ,s 2 n ,s 3 n is displayed.

[0131] The detection information also includes a likelihood map that represents the likelihood of each position in the image being a detection target, and the gaze information generation unit 40C generates, as gaze information, position information of at least one abnormal area that is considered to be a detection target and position information of at least one normal area that is considered to be background based on the likelihood map.

[0132] 14 shows an example of a selection screen for selecting a method to be used for generating gaze information. The gaze information used by the gaze information generating unit 40C may be selected from a user-selectable input unit. In addition, the following (1) to (4) are possible displays for selecting gaze information.

[0133] (1) Allows the user to select either a target point or a rectangle as gaze information. (2) The number of target points can be selected. (3) There may be a UI that allows the user to set the sampling method for the target points. (4) There may be a UI that allows the user to manually set the target point, i.e., select the position while viewing the likelihood map.

[0134] Then, the gaze information generating unit 40C performs the above S anormal and S normal is transmitted as gaze information to the area detection unit 40D.

[0135] Next, the area detection unit 40D detects an abnormal area based on the image in which the abnormality has been detected by the detection unit 40B and the gaze information generated by the gaze information generation unit 40C.

[0136] When a deformation is detected by the detection unit 40B, the area detection unit 40D receives the image in which the deformation is detected from the image acquisition unit 40A. Then, the area detection unit 40D detects the deformation area in the image based on the gaze information. The area detection unit 40D performs area detection using a second inference model.

[0137] In the second embodiment, as in the first embodiment, a segment anything model (SAM) is used as the second inference model. anormal and S normal is used as a prompt to be input to the SAM. anormal has target points that indicate the location information of the deformed area as elements. normal has target points indicating the position information of normal areas as elements.

[0138] The area detection unit 40D detects s i a ,s i n As a prompt, the normal area and the abnormal area are detected. For example, the input image is x. The position information of the abnormal area is s i a Using the prompt, segmentation is performed on the input image x by SAM. The resulting mask of the abnormal region is then denoted as M i anormal Then, we can write it as follows:

[0139] M i anormal =SAM(x,s i a )

[0140] Also, for example, the position information s of the normal region i n The mask of normal regions obtained by segmenting the input image x using SAM is called M. inormal Then, we can write it as follows:

[0141] M i normal =SAM(x,s i n )

[0142] The area detection unit 40D is S anormal Alternatively, the region detection unit 40D may perform region detection for each of the N elements of S to generate a mask. normal You can also generate a mask by performing region detection on each of the M elements of S. SAM can also generate a mask by receiving multiple target points as a single prompt. In this case, S anormal You can generate a mask using the following prompt:

[0143] M anormal =SAM(x,s anormal )

[0144] The region detection unit 40D detects the abnormal region by integrating the mask of the abnormal region and the mask of the normal region. The method of integrating the masks is the same as in the first embodiment, so a description thereof will be omitted.

[0145] While there is little learning data for abnormalities, there is a lot of learning data for normal road surfaces. Therefore, normal areas can be detected more accurately than abnormalities. By integrating the detection results of abnormal areas and normal areas, abnormal areas can be detected more accurately. Note that either the abnormal area or the normal area may be detected. Furthermore, if multiple pieces of position information are generated by the gaze information generation unit 40C, the area detection unit 40D may perform area detection multiple times depending on the number of pieces of position information generated.

[0146] 15 shows an example of a display screen for the detection results in the second embodiment. The information output unit 40E outputs various information in the same manner as in the first embodiment. The information output unit 40E displays the input image of the image acquisition unit 40A, the region detection results of the detection unit 40B, the likelihood map of the detection unit 40B, and the abnormal region detection results of the region detection unit 40D.

[0147] In the second embodiment, the gaze information generator 40C generates gaze information on abnormalities and related normal areas from a likelihood map obtained by the detector 40B using a model trained by weakly supervised learning. The area detector 40D then performs area detection for each area and integrates the results of the area detection. This allows accurate detection of the shape, position, and type of abnormal areas in the image.

[0148] Figure 16 is an example of a display screen that displays the location on the travel route where the image in which the abnormality was detected was captured. For example, in Figure 16, the inspection data item in the upper right shows four locations on the travel route where abnormalities were detected as the number of detected abnormalities. ID: 001 is selected. When selected, an image of the abnormality and detailed information are displayed in the selected abnormality item in the lower right.

[0149] By selecting "Go to detection details" in the item of the selected abnormality, a separate window will open, allowing you to move to a screen displaying details of the detection, such as an enlarged view or a score map. You can also move to the detection result display screens shown in Figure 11 and Figure 15. Note that in Figure 16, the inspection target is limited to the road surface, but the inspection target may also be selected from road surfaces, guardrails, sound barriers, etc.

[0150] In this embodiment, the road surface deformation detection system 1 has been described as an example in which an expressway is the inspection target and deformations of the road surface are detected as the detection target. However, the inspection target and the detection target are not limited to these. For example, inspection targets include general roads, bridges, wind turbine blades, steel towers, solar panels, power plants, and substations. Deformations that are the detection target include potholes, holes, peeling, cracks, depressions, rust, fallen objects, and landslides. Note that the inspection target is not limited to roads, but may also be industrial products. This allows for quality inspections before shipping.

[0151] In this embodiment, an example in which image processing is performed on an image captured by an imaging device mounted on a moving object has been described, but the present invention is not limited to this. For example, the image processing device may perform image processing on an image captured manually or an image captured by a fixed camera.

[0152] In this embodiment, the case where the detection target is an abnormality region and the background region related to the abnormality is a normal region has been described, but the present invention is not limited to this form.

[0153] As described above, the image processing program of this embodiment causes a computer to execute a gaze information generation procedure that uses a first inference model to detect a detection target from an image and generates gaze information for performing area detection of the detection target based on detection information regarding the detection target detected from the image; an area detection procedure that uses a second inference model to perform area detection based on the image in which the detection target is detected and the gaze information; and an information output procedure that outputs at least one of an image, the detection information, the gaze information, and the result of the area detection.

[0154] As a result, by detecting the area of ​​the detection target using the detection information and gaze information, the shape of the detection target in the image and information about the detection target can be accurately detected from the gaze information.

[0155] In addition, the gaze information of the image processing program according to this embodiment includes at least one of information related to the detection target and information on the background area related to the detection target, and the area detection procedure includes at least one of area detection of the detection target and detection of the background area.

[0156] This makes it possible to detect information about the detection target by performing area detection from information about the detection target or information about a background area related to the detection target.

[0157] Furthermore, the area detection means of the image processing program according to this embodiment detects the area of ​​the detection target and the background area, and integrates the detection results to perform area detection.

[0158] This allows area detection to be performed from information about the detection target and its related background area, and by integrating the respective detection results, it is possible to more accurately detect the shapes of the detection target and background area in the image, as well as information about the detection target.

[0159] In addition, the detection information of the image processing program of this embodiment includes the type of detection target, and the gaze information generation procedure generates, as gaze information, a text prompt that characterizes at least one specific detection target and a text prompt that characterizes a background area related to the at least one specific detection target based on the type of detection target.

[0160] This allows text prompts to be generated as information about the detection target and its related background area, and by detecting each area based on the text prompts and integrating the detection results, it is possible to accurately detect the shape, type, etc. of the detection target and background area in the image, as well as information about the detection target.

[0161] In addition, the detection information of the image processing program according to this embodiment includes a likelihood map that indicates the likelihood of each position in the image being a detection target, and the gaze information generation procedure generates, as gaze information, position information of at least one area considered to be a detection target and position information of at least one area considered to be background based on the likelihood map.

[0162] This allows position information to be generated as information on the detection target and its related background area, and by detecting each area based on the position information and integrating the detection results, it is possible to accurately detect the shapes and position information of the detection target and background area in the image, as well as likelihood map information related to the detection target.

[0163] In addition, the first inference model of the image processing program of this embodiment learns using data that instructs the presence or absence of a specific detection target in an image in predetermined units, outputs likelihood in units smaller than the predetermined units, and learns so that the presence or absence of the instructed detection target matches the maximum value of the output likelihood.

[0164] This reduces the workload of teaching the first inference model learning data by using a machine learning model trained with data that indicates the presence or absence of a detection target.

[0165] The image processing program according to the present embodiment also executes an information output procedure for outputting at least one of an image, detection information, gaze information, and a region detection result.

[0166] This allows the user to easily grasp various pieces of information.

[0167] Furthermore, the gaze information of the image processing program according to this embodiment executes an input procedure that can be selected by the user.

[0168] This allows the user to select appropriate gaze information depending on the detection target.

[0169] In addition, the image processing device of this embodiment includes a gaze information generation unit that generates gaze information for performing area detection of the detection target based on detection information regarding the detection target detected from the image using a first inference model that detects the detection target from the image, an area detection unit that performs area detection using a second inference model based on the image in which the detection target is detected and the gaze information, and an information output unit that outputs at least one of the image, the detection information, the gaze information, and the result of the area detection.

[0170] In addition, the image processing method of this embodiment includes a gaze information generation process that uses a first inference model to detect the detection target from the image to generate gaze information for performing area detection of the detection target based on detection information regarding the detection target detected from the image, an area detection process that uses a second inference model to perform area detection based on the image in which the detection target is detected and the gaze information, and an information output process that outputs at least one of the image, the detection information, the gaze information, and the result of area detection.

[0171] The present invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be created by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined. [Explanation of symbols]

[0172] 40: Image processing section 40A: Image acquisition unit 40B: Detection unit 40C: Gaze information generation unit 40D: Area detection unit 40E: Information output section

Claims

1. On the computer, a gaze information generation step of generating gaze information for detecting an area of ​​a detection target based on detection information regarding the detection target detected from an image using a first inference model; an area detection step of performing the area detection using a second inference model based on the image in which the detection target is detected and the gaze information; an information output step of outputting at least one of the image, the detection information, the gaze information, and the result of the area detection; An image processing program for executing the above.

2. the gaze information includes at least one of information related to the detection target and information on a background area related to the detection target, the region detection step includes at least one of detecting a region of the detection target and detecting a background region; The image processing program according to claim 1 .

3. the region detection step performs region detection of the detection target and region detection of the background region, and integrates the detection results to perform region detection. The image processing program according to claim 2 .

4. the detection information includes a type of detection target, the gaze information generating step generates, as the gaze information, a text prompt characterizing at least one specific detection target and a text prompt characterizing a background region related to the at least one specific detection target, based on a type of the detection target. The image processing program according to claim 3 .

5. the detection information includes a likelihood map representing the likelihood of the detection target at each position in the image; the gaze information generating step generates, as the gaze information, position information of at least one region regarded as the detection target and position information of at least one region regarded as a background based on the likelihood map. The image processing program according to claim 3 .

6. The first inference model learns from data that indicates the presence or absence of a specific detection target in a predetermined unit for the image, and outputs a likelihood in a unit smaller than the predetermined unit; The system is trained so that the presence or absence of the target to be detected matches the maximum likelihood value to be output. The image processing program according to claim 1 .

7. The gaze information includes a user-selectable input procedure; 2. The image processing program according to claim 1, for executing the above steps.

8. a gaze information generating unit that generates gaze information for detecting an area of ​​a detection target based on detection information regarding the detection target detected from an image using a first inference model; an area detection unit that performs the area detection using a second inference model based on the image in which the detection target is detected and the gaze information; an information output unit that outputs at least one of the image, the detection information, the gaze information, and the result of the area detection; An image processing device comprising:

9. a gaze information generating step of generating gaze information for detecting an area of ​​a detection target based on detection information relating to the detection target detected from an image using a first inference model; an area detection step of detecting the area using a second inference model based on the image in which the detection target is detected and the gaze information; an information output step of outputting at least one of the image, the detection information, the gaze information, and the result of the area detection; An image processing method comprising:

Citation Information

Patent Citations

  • Information processor, information processing method, and program

    JP2018198053A

  • Inspection support system, server device, inspection support method, and inspection support program

    JP2021022257A

  • Gazed object recognition system and method

    JP2021149446A

  • Method and device for image processing, computer readable storage medium, and electronic device

    US20190377944A1

  • Identification device, identification program, and identification method

    JP6955295B1