Information processing apparatus, information processing method, and program

The information processing apparatus addresses the challenge of complex deep learning engine tuning by allowing users to easily operate and combine multiple machine learning models for image processing and detection, enhancing usability and flexibility.

JP7685871B2Active Publication Date: 2025-05-30NTT COMWARE CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021078610
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-06
Publication Date
2025-05-30
Estimated Expiration
2041-05-06

AI Technical Summary

Technical Problem

Existing systems require specialized knowledge to tune deep learning engines for image processing and detection, making it difficult for users to change settings or use multiple deep learning engines effectively.

Method used

An information processing apparatus that uses a combination of multiple machine learning models, including classification, detection, and segmentation models, which can be easily operated and configured by users through a simple interface, allowing for flexible processing and output combination based on user-defined operations.

Benefits of technology

Enables users to perform complex detection tasks using multiple machine learning models with a simple operation, improving usability and flexibility without requiring specialized knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007685871000001
    Figure 0007685871000001
  • Figure 0007685871000002
    Figure 0007685871000002
  • Figure 0007685871000003
    Figure 0007685871000003
Patent Text Reader

Abstract

To use detection using a plurality of machine learning models, with a simple operation.SOLUTION: An information processing device includes: a reception unit that receives an operation of a user; a setting unit that sets a combination of a plurality of machine learning models having different output formats, based on the operation received by the reception unit; and an execution unit that causes the plurality of machine learning models to execute detection processing, based on the combination set by the setting unit.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] Image and other detection technologies are rapidly becoming more widespread due to the advancement of AI (Artificial Intelligence) and DL (Deep Learning) technologies, the shift to Open Source Software (OSS), and the development of cloud computing using GPUs (Graphics Processing Units).In addition, devices that handle point cloud data are becoming more widespread, and point cloud data is beginning to be used in fields such as digital twin technology and autonomous driving technology.

[0003] Furthermore, systems that use, for example, DL technology to manage defects in facilities, etc. have been known for some time. For example, Patent Document 1 discloses an information processing device that includes: a learning unit that performs deep learning on an image of a surface based on teaching data that indicates the state of the surface, in order to comprehensively determine the state of the surface of a building, etc.; an image acquisition unit that acquires an image of an area that includes the surface, which is associated with position information that indicates the position where the image was captured; and a state determination unit that determines the state of the surface in the image acquired by the image acquisition unit based on the learning result of the learning unit. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2020-140334 Summary of the Invention [Problem to be solved by the invention]

[0005] Although the above-described information processing devices use cameras to detect buildings and the like, they incorporate image processing and deep learning engines specialized for specific fields and applications, and therefore require engineers with specialized knowledge to tune the deep learning engine, etc. This makes it difficult for users of the information processing devices to change settings such as tuning the deep learning engine.

[0006] Furthermore, with the spread of DL technology, systems with deep learning engines are often used in ensembles when DL technology is used in many situations. Even in this case, the output format and output content differ depending on the deep learning engine, so it is necessary to tune each system individually to build the system. In particular, when the output content differs between multiple deep learning engines, it is impossible to design a single system for general use. Therefore, if you want to tune each deep learning engine, change the deep learning engine, or rearrange the output results, you need to change the entire system each time.

[0007] The present invention has been made in consideration of the above-mentioned problems, and aims to provide an information processing device, an information processing method, and a program that enable detection using multiple machine learning models with simple operations. [Means for solving the problem]

[0008] (1) One aspect of the present invention is a system including: a reception unit that receives a user's operation; and, based on the operation received by the reception unit, The machine learning system includes a first machine learning model that outputs a classification result for classifying an object, a second machine learning model that outputs a detection result for detecting whether or not the object has a defect, and a third machine learning model that outputs a segmentation result for determining a defective area in the object. Combining multiple machine learning models and expression information representing a processing order of the first machine learning model, the second machine learning model, and the third machine learning model. a setting unit that sets the combination set by the setting unit; inputting an image of a detection target to each of the first machine learning model, the second machine learning model, and the third machine learning model included in Let the , combining the processing results of the machine learning models based on the formula information. and an execution unit.

[0009] (2) One aspect of the present invention is the information processing device described above, each of the first machine learning model, the second machine learning model, and the third machine learning model; corresponds to a class name, and the setting unit sets the each of the first machine learning model, the second machine learning model, and the third machine learning model; may be specified.

[0010] (3) In one aspect of the present invention, in the information processing device described above, the setting unit performs the following operation based on the operation received by the receiving unit: a first machine learning model, the second machine learning model, and the third machine learning model Set the processing order for the combination of The execution unit executes the processing based on the processing order set by the setting unit. each of the first machine learning model, the second machine learning model, and the third machine learning model; may perform the detection process.

[0011] (4) In one aspect of the present invention, in the information processing device described above, the setting unit performs the following operation based on the operation received by the receiving unit: each of the first machine learning model, the second machine learning model, and the third machine learning model; a result generation process based on the output result of the setting unit, and the execution unit may execute the result generation process set by the setting unit.

[0012] (5) In one aspect of the present invention, in the information processing device described above, the setting unit performs the following operation based on the operation received by the receiving unit: each of the first machine learning model, the second machine learning model, and the third machine learning model; and a post-processing unit for processing the output result of the image processing unit, and the execution unit may execute the post-processing set by the setting unit.

[0013] (6) In one aspect of the present invention, in the information processing device, the setting unit performs the setting process based on the operation received by the receiving unit. each of the first machine learning model, the second machine learning model, and the third machine learning model; and the execution unit executes the preprocessing set by the setting unit and outputs the result of the preprocessing to the each of the first machine learning model, the second machine learning model, and the third machine learning model; may be used as input for

[0015] ( 7 ) One aspect of the present invention is The computer A step of accepting a user's operation, and based on the accepted operation, The machine learning system includes a first machine learning model that outputs a classification result for classifying an object, a second machine learning model that outputs a detection result for detecting whether or not the object has a defect, and a third machine learning model that outputs a segmentation result for determining a defective area in the object. Combining multiple machine learning models and expression information representing a processing order of the first machine learning model, the second machine learning model, and the third machine learning model. and setting the plurality of machine learning models based on the set combination. inputting an image of a detection target to each of the first machine learning model, the second machine learning model, and the third machine learning model included inLet the combining the processing result of the first machine learning model, the processing result of the second machine learning model, and the processing result of the third machine learning model based on the formula information; Step and Execute , an information processing method.

[0016] ( 8 ) One aspect of the present invention is a method for controlling a computer to receive a user's operation; and based on the received operation, The machine learning system includes a first machine learning model that outputs a classification result for classifying an object, a second machine learning model that outputs a detection result for detecting whether or not the object has a defect, and a third machine learning model that outputs a segmentation result for determining a defective area in the object. Combining multiple machine learning models and expression information representing a processing order of the first machine learning model, the second machine learning model, and the third machine learning model. and setting the plurality of machine learning models based on the set combination. inputting an image of a detection target to each of the first machine learning model, the second machine learning model, and the third machine learning model included in Let the combining the processing result of the first machine learning model, the processing result of the second machine learning model, and the processing result of the third machine learning model based on the formula information; It is a program that executes a process including steps. [Effects of the Invention]

[0017] According to one aspect of the present invention, detection using multiple machine learning models can be performed with simple operations. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a block diagram showing an example of the configuration of an object detection system 1 according to a first embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a configuration of assets and data. [Figure 3] FIG. 10 is a diagram illustrating an example of a screen displayed by a visualization unit. [Figure 4] FIG. 10 is a diagram illustrating another example of a screen displayed by the visualization unit. [Figure 5] FIG. 10 is a diagram illustrating an example of an annotation setting screen by a learning setting unit. [Figure 6] FIG. 10 is a diagram showing an example of a learning setting screen by a learning setting unit. [Figure 7] FIG. 10 is a diagram showing an example of a parameter setting screen by a learning setting unit. [Figure 8] FIG. 10 is a diagram illustrating an example of a setting screen for merging teacher data and classes by the annotation management unit. [Figure 9] FIG. 10 is a diagram showing an example of a setting screen for images and parameters by a detection setting unit. [Figure 10] FIG. 10 is a diagram showing an example of a travel route setting screen. [Figure 11] FIG. 10 is a diagram illustrating an example of a task set in task management. [Figure 12] FIG. 10 is a diagram illustrating an example of a task for a robot set in task management. [Figure 13] FIG. 10 is a block diagram illustrating an example of an information processing device according to a second embodiment. [Figure 14] FIG. 2 is a block diagram illustrating a processing unit in an information processing device. [Figure 15] FIG. 10 is a diagram for explaining an example of the operation of a preprocessing unit. [Figure 16] FIG. 10 is a diagram for explaining another example of the operation of the preprocessing unit. [Figure 17] FIG. 10 is a diagram illustrating an example of a sequence of an information processing device according to the second embodiment. [Figure 18] FIG. 10 is a diagram showing an example of a screen displayed on the user terminal device 700 when formula information is generated. [Figure 19] FIG. 1 is a diagram illustrating an example of input data and output data of a classification-type machine learning model. [Figure 20] FIG. 10 is a diagram illustrating an example of input data and output data of a detection-type machine learning model. [Figure 21] FIG. 10 is a diagram illustrating an example of input data and output data of a segmentation-type machine learning model. [Figure 22] FIG. 6 illustrates an example of processing and output classes of the combiner 650 for symbols for combining and combining machine learning models. [Figure 23] FIG. 10 is a diagram illustrating an example of use of each combination of machine learning models. [Figure 24] FIG. 10 is a diagram illustrating an example of AND processing. [Figure 25] FIG. 10 is a diagram illustrating an example of OR processing. [Figure 26] FIG. 10 is a diagram illustrating an example of DIV processing. [Figure 27] FIG. 10 is a diagram illustrating an example of NOT processing. DETAILED DESCRIPTION OF THE INVENTION

[0019] An information processing device, an information processing method, and a program to which the present invention is applied will be described below with reference to the drawings.

[0020] First Embodiment [Outline of the first embodiment] The object detection system 1 of the first embodiment uses AI technology to detect objects, detect areas, or provide results of object classification. A user, for example, purchases the object detection system 1 incorporating a machine learning model for utilizing AI technology. The object detection system 1 configures learning and detection settings for the machine learning model based on the user's operations, and then performs learning and detection using the machine learning model constructed based on the configured learning and detection settings. To allow the user to configure and manage the machine learning model, the object detection system 1 accepts the user's operations and visualizes various information. Thus, by introducing the object detection system 1, the user can proactively and consistently perform the entire process, from configuring learning and detection settings for the machine learning model to learning and detection using the machine learning model, even without specialized knowledge of computers or data science for managing detection targets such as equipment. Furthermore, the object detection system 1 not only performs the learning and detection settings for the machine learning model and learning and detection using the machine learning model, but also consistently manages the operation of the machine learning model, inputs data into the machine learning model, creates a machine learning model for the detection target image, and performs detection and visualization.

[0021] [Configuration of Object Detection System 1] FIG. 1 is a block diagram showing an example configuration of an object detection system 1 according to a first embodiment. The object detection system 1 includes, for example, a web server unit 100, a user terminal device 200, a robot operation unit 202, an AI execution unit 300, a data management unit 400, a database 402, and a system setting unit 500. The web server unit 100, the user terminal device 200, the robot operation unit 202, the AI ​​execution unit 300, the data management unit 400, the database 402, and the system setting unit 500 are connected to a communication network. Each device connected to the communication network includes a communication interface such as a network interface card (NIC) or a wireless communication module (not shown in FIG. 1). The communication network includes, for example, the Internet, a wide area network (WAN), a local area network (LAN), a cellular network, etc.

[0022] The user terminal device 200 is a device having a display device and an operation device, such as a personal computer, a smartphone, a tablet terminal, etc. The user terminal device 200 transmits information based on the operation of the user of the object detection system 1 to the web server unit 100, and acquires display information to be presented to the user of the object detection system 1 from the web server unit 100.

[0023] The robot operation unit 202 is an operation device that generates operation information for operating the robot 210. The robot operation unit 202 may be integrated with, for example, the user terminal device 200. The robot 210 is equipped with a camera device and a moving mechanism and is an autonomous device that captures and moves images. The robot operation unit 202 may set the robot's movement path, movement period, movement speed, etc. The operation information generated by the robot operation unit 202 is supplied to the robot 210 via the web server unit 100, but may also be supplied directly from the robot operation unit 202 to the robot 210. Note that while the robot 210 is provided in this embodiment, this is not limited thereto. Furthermore, the robot 210 may not be provided if there is another means for acquiring images for machine learning model training and detection. The other means for acquiring images for machine learning model training and detection may be, for example, an automatic piloting system for a drone. Even if the object detection system 1 does not set the drone's path, etc., the object detection system 1 may cooperate with the automatic piloting system to acquire images captured by the drone.

[0024] The data management unit 400 is, for example, a general-purpose database management device, and manages data registered in the database 402. The database 402 may be realized, for example, by a storage area network (SAN), a network attached storage (NAS), or a cloud-based storage service. The data management unit 400 processes two types of data: assets and files. Assets are data for managing the object detection system 1. Files are actual files used for learning and detection in the object detection system 1. Actual files store data such as images, point cloud data, and tuning parameters for machine learning models, while files store data such as images, point cloud data, tuning parameters, model files created by learning, and setting parameters. An asset has a data structure including, for example, its own asset ID, parent asset ID (multiple IDs can be set; the root asset can be omitted, but the others must be set), management metadata (key-value), and asset type. A file has a data structure that includes, for example, actual file data (including data such as configuration parameters), its own file ID, the asset ID to which the file is linked (also referred to as the parent asset for convenience), management metadata (key-value), and information such as file type, location information, camera shooting orientation information, and defect class.

[0025] The data management unit 400 performs operations such as asset registration, deletion, and update, and searches for child assets and parent assets using asset IDs as keys, based on information obtained via the communication interface. For example, based on information obtained via the communication interface, the data management unit 400 performs operations such as file registration, deletion, and update, search for parent assets using file IDs as keys, search for files using asset IDs as keys, and obtain a list of assets or files that match the key-value of management metadata. An access management ID (dataset ID) is set for each file and asset, and the data management unit 400 may use the access management ID (dataset ID) to enable access restrictions on an asset-by-asset and file-by-file basis.

[0026] FIG. 2 is a diagram showing an example of the configuration of assets and data. In the database 402, a hierarchy is set for a route in the order of location ID, building, building part, and date, and multiple image data are registered corresponding to the date. In addition, in the database 402, a hierarchy is set for a route in the order of location ID, tree, and date, and multiple image data are registered corresponding to the date. Note that restrictions and rules do not need to be set for asset hierarchies, and files (image data) do not need to be linked to child assets in the lowest hierarchical level. Furthermore, one file may be allowed to have multiple parent assets. For example, selecting a building asset may display point cloud data corresponding to the building part, and selecting the date "y2 / m2 / d2" may switch to displaying point cloud data corresponding to the date "y2 / m2 / d2."

[0027] The AI ​​execution unit 300 is a device that performs processing using AI technology and may be configured to include a GPU server. In the embodiment, the processing using AI technology is a learning process for a machine learning model and a detection process using the machine learning model. Note that the detection process may also be referred to as a recognition process, a classification process, or a determination process. The AI ​​execution unit 300 includes, for example, a learning unit 310 and a detection unit 320. Functional units such as the learning unit 310 and the detection unit 320 are realized by a processor such as a central processing unit (CPU) executing a program stored in a program memory. The program is stored in a storage device such as a hard disk drive (HDD) or flash memory of the AI ​​execution unit 300. Some or all of these functional units may be realized by hardware such as a large-scale integration (LSI), an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA), or may be realized by a combination of software and hardware.

[0028] The learning unit 310 constructs a machine learning model to detect defects and the like in the object. The machine learning model is, for example, a convolutional neural network (CNN), but is not limited to this and may be any known machine learning model. Examples of machine learning models include Mask R-CNN, YOLO, SSD, and DeepLab. The detection unit 320 performs calculations using the machine learning model to detect defects and the like in the object. The machine learning model in the AI ​​execution unit 300 may be any of object detection, classification, and area detection. An object detection machine learning model, for example, detects the location of a defect, such as damage to a building as the object. A classification machine learning model, for example, sets multiple categories in advance and determines which category the object belongs to. An area detection machine learning model, for example, visualizes an area where a defect is suspected to occur or outputs the percentage of defects (numerical value). The learning unit 310 acquires training data, inputs the training data into the machine learning model, updates the parameters of the machine learning model to obtain an appropriate output, and stores the updated parameters as the learning result. The detection unit 320 inputs image data of the object into a machine learning model, and detects defects or the like in the object based on the output of the machine learning model.

[0029] The AI ​​execution unit 300 checks, for example, the tasks registered in the database 402 and the parent assets (= task assets) linked to the tasks. If a task is found, the AI ​​execution unit 300 executes learning and / or detection of a machine learning model and registers the learning results and / or detection results in the database 402. In learning and / or detection of a machine learning model, the AI ​​execution unit 300 may associate one task asset with one GPU server, or may associate multiple task assets with one GPU server. The AI ​​execution unit 300 also executes learning and / or detection for a machine learning model registered in a task among multiple machine learning models. Furthermore, in detection using a machine learning model, if expression information (described later) is registered in association with the task, the AI ​​execution unit 300 executes detection processing using multiple machine learning models in accordance with the expression information. Furthermore, the AI ​​execution unit 300 may execute detection processing using one machine learning model multiple times. For example, the AI ​​execution unit 300 may execute detection processing for class a and detection processing for class b using a certain machine learning model A, and output the same for class a and class b, or may divide the output according to the difference in the confidence level of class a as a result of the detection processing for class a using machine learning model A.

[0030] The system setting unit 500 is a device that performs settings such as user settings for users of the object detection system 1. The system setting unit 500 performs tasks such as creating and changing user information, setting access restrictions for each function in the object detection system 1, and setting the servers used by the learning unit 310 and the detection unit 320, and registers the user information and setting information in a database (not shown). The system setting unit 500 may temporarily stop the server during periods when no tasks are scheduled for the AI ​​execution unit 300. By temporarily stopping the server, the system setting unit 500 can reduce unnecessary charges, for example, when the AI ​​execution unit 300 uses a server that charges by the hour.

[0031] The web server unit 100 is a server device that controls each unit in the object detection system 1. The web server unit 100 includes, for example, a user interface unit 110, a data input unit 120, a learning setting unit 130, a detection setting unit 140, a path setting unit 150, a task management unit 160, a management unit 170, and an ensemble setting unit 180. The user interface unit 110 includes, for example, an operation reception unit 112 and a visualization unit 114. Functional units such as the user interface unit 110, the data input unit 120, the learning setting unit 130, the detection setting unit 140, the path setting unit 150, the task management unit 160, the management unit 170, and the ensemble setting unit 180 are realized by a processor such as a CPU executing a program stored in a program memory. The program is stored in advance in a storage device such as a hard disk drive (HDD) or flash memory of the web server unit 100.

[0032] The operation accepting unit 112 accepts an operation from the user based on information from the user terminal device 200. The visualization unit 114 performs processing to visualize information to be presented to the user.

[0033] The visualization unit 114 displays assets and files based on, for example, user operations, information for setting up learning and detection, and information for presenting learning results and detection results. Specifically, the visualization unit 114 displays captured images and videos supplied from the robot operation unit 202, detection results by the AI ​​execution unit 300 (e.g., images in which defects are detected), point cloud data and CAD models registered in the database 402, and the like. When location information is associated with an asset or a file, the visualization unit 114 displays the asset or file by linking it to the location information. The location information is, for example, latitude and longitude information, altitude information, or relative coordinate information indicating a distance from a reference point in the X, Y, and Z directions. When the object detection system 1 holds relative location information as location information, it holds information linking the reference position with GPS information and can convert the relative coordinates and GPS information based on the linking information.

[0034] When an asset or file is selected, the visualization unit 114 may, for example, display details of the asset or file, switch the display position of the point cloud or map, or switch between displaying and hiding the shooting location. Furthermore, when a shooting location on a point cloud or map is specified, the visualization unit 114 may display images (such as a captured image or a defect detection image) related to the specified point cloud or shooting location. Furthermore, the visualization unit 114 may display the original image and the image determined by the AI ​​execution unit 300 side by side, and may enlarge and display each of the original image and the determined image based on a user operation. Furthermore, the visualization unit 114 may narrow down the assets or files by setting a filter such as a class in which a defect was detected by the AI ​​execution unit 300, and may plot on a map or point cloud the shooting locations where only the specified defect was detected based on a user operation.

[0035] 3 is a diagram showing an example of a screen displayed by the visualization unit. For example, the visualization unit 114 visualizes date information, layer information, and asset information related to point cloud data or images while displaying the point cloud data, CAD images, or captured images using a 3D viewer. The layer information indicates, for example, that the images superimposed in the 3D viewer are point cloud data and a CAD model. This allows the user to determine whether to set the visualized point cloud data as training data or as a detection target.

[0036] 4 is a diagram showing another example of a screen displayed by the visualization unit. The visualization unit 114 can display a hierarchy of assets and data as point cloud data to be displayed in a 3D viewer, for example. The visualization unit 114 can present to the user, for example, that assets such as the Suzaku Gate roof and Suzaku Gate building components, and multiple roadside trees are registered in the lower layers of the object "Suzaku Gate," and allow the user to decide whether to add or delete point cloud data or images to the registered assets.

[0037] For example, the data input unit 120 acquires image data already registered in the database 402 during learning and inputs the data into the learning unit 310. Also, for example, the data input unit 120 acquires image data captured by the robot 210 during detection and inputs the image data into the detection unit 320.

[0038] The learning setting unit 130 sets a learning process for a machine learning model for detecting an object based on an operation received by the operation receiving unit 112. The learning setting unit 130 sets information necessary for the learning process, such as the machine learning model to be learned, parameters to be tuned in the machine learning model, and images to be input to the machine learning model. The learning setting unit 130, for example, causes the visualization unit 114 to display images captured by the robot 210 or images registered in the database 402, and designates images specified by user operation as training data. The learning setting unit 130 can also add or delete images to or from the training data based on user operation.

[0039] Furthermore, the learning setting unit 130 may prepare a template of a machine learning model whose parameters have been tuned, and use the machine learning model for learning and detection. For tuning, for example, an image that can be automatically calculated from the image set of training data, the annotation size, the image size, etc. may be used. Furthermore, the learning setting unit 130 may delete the machine learning model used for learning.

[0040] Furthermore, the learning setting unit 130 includes an annotation management unit 132 that adds and changes attributes to training data used in the learning process of the machine learning model based on user operations and integrates attributes based on user operations. FIG. 5 is a diagram illustrating an example of an annotation setting screen by the learning setting unit. When setting training data, the annotation management unit 132 displays an image and allows the user to select a dead or damaged area on a tree included in the image, thereby setting training data with the attribute of the dead or damaged area added. FIG. 6 is a diagram illustrating an example of a learning setting screen by the learning setting unit. For example, the annotation management unit 132 can display an image with an attribute of a pruning scar added, and set a learning setting pattern 1, a machine learning model, and parameters for the machine learning model for the image. FIG. 7 is a diagram illustrating an example of a parameter setting screen by the learning setting unit. Based on user operations, the annotation management unit 132 can set a region detection type machine learning model to detect decay fungi in wood, and can also set machine learning parameters such as the number of epochs, number of iterations, batch size, anchor size, anchor scale, image size, resize mode, data expansion (increase in annotation data), etc. It is desirable that for these machine learning parameters, numerical values ​​or the like are prepared as templates for each machine learning parameter, and the user can select a desired numerical value from the template, thereby tuning the learning process of the machine learning model without having to set the machine learning parameters individually.

[0041] FIG. 8 illustrates an example of a screen for setting up merging of training data and classes by the annotation management unit. The annotation management unit 132 has functions for copying and backing up annotations and merging multiple annotations. The annotation merging function can, for example, merge separate annotation data into a single annotation data, or merge classes with different names defined in the same annotation data into the same class name. This can streamline tasks such as combining annotation data after multiple people share the tasks of setting, adding, and deleting annotation data, or merging annotation data classes and re-learning if the estimation accuracy does not improve after initial finely divided learning. The annotation management unit can, for example, merge annotation data after multiple users have performed separate annotation tasks, or merge specified classes from finely divided classes into the same class, allowing the learning unit to perform learning processing on the merged class. For example, the annotation management unit may have annotated Mushroom A and Mushroom B as separate classes based on the annotation work, but during the learning process, it can train both as a unified class called Mushroom. Furthermore, if the annotation management unit detects certain annotation data with high accuracy after setting it, it can back up that annotation data, and the backed-up annotation data can be used even if additional annotation data is reset.

[0042] The detection setting unit 140 sets a process for detecting an object based on an operation received by the operation receiving unit 112. The detection setting unit 140 causes the detection unit 320 to detect images such as defects using a machine learning model created by the learning unit 310 for captured images, images displayed by the visualization unit 114, etc. The detection setting unit 140 can set a series of processes, such as storing images detected by the detection unit 320 in association with original detection images (teacher data) in the database 402 and displaying them by the visualization unit 114. If location information is added to the original detection images, the detection setting unit 140 may set the detected images to also be assigned location information. Furthermore, the detection setting unit 140 may set the deletion of some images from the processing results of the detection unit 320.

[0043] FIG. 9 is a diagram showing an example of a setting screen for images and parameters set by the detection setting unit. When setting the detection process, the detection setting unit 140 causes the visualization unit 114 to display the setting screen shown in FIG. 9. For example, the detection setting unit 140 can select an image to which the attribute of the Suzaku Gate roof has been added and set a machine learning model. The detection setting unit 140 may set the detection type, reliability of the detection result, and class name of the machine learning model. When the detection start button is selected with the parameters specified, the detection setting unit 140 causes the detection unit 320 to execute the detection process.

[0044] The detection setting unit 140 may cause the detection unit 320 to perform detection using multiple machine learning models for one detection target image. In this case, the detection setting unit 140 can set a combination of multiple machine learning models, the processing order for the combination, and the like, by having the user select an add model button based on a user operation, for example.

[0045] The route setting unit 150 sets a travel route for the ground-traveling robot 210 or a drone (not shown). FIG. 10 is a diagram showing an example of a travel route setting screen. The route setting unit 150 displays a map screen as shown in FIG. 10 and sets a route based on a user's operation on the map. The route setting unit 150 sets camera settings such as an area in which the robot 210 will move, a robot ID, an operation mode such as start / stop of shooting, a moving speed, and a shooting point based on a user's operation. The route setting unit 150 may set a shooting interval when capturing still images as a camera setting. The route setting unit 150 registers information related to the route setting in the database 402. As a result, the information related to the route setting registered in the database 402 can be acquired by the robot operation unit 202. The route setting unit 150 can operate the robot 210 in response to selection of the start running button.

[0046] The task management unit 160 manages tasks such as learning processing and detection processing of machine learning models. FIG. 11 is a diagram illustrating an example of a task set in task management. For example, the task management unit 160 registers, for each task, the priority, start time, end time, type of machine learning model, machine learning model ID, class ID, detection result, and task status. FIG. 12 is a diagram illustrating an example of a task for a robot set in task management. For example, the task management unit 160 registers, for each task, the priority, start time, end time, route parameters, route ID, and task status. The route parameters include, for example, the area, robot ID, operation mode of the robot 210, movement speed, camera control settings attached to the robot 210 or connected to the robot 210 or a network, such as the shooting interval and zoom magnification.

[0047] For example, when the start time of a registered task arrives, the task management unit 160 executes processing using a machine learning model registered in association with the task, thereby updating the processing result and status. The task management unit 160 may also rearrange the execution order or delete tasks in response to a user operation. Furthermore, the task management unit 160 may also suspend or resume tasks. For example, the task management unit 160 may cooperate with the system setting unit 500 to automatically stop a device for running a machine learning model when there are no suspended tasks or unprocessed tasks, and may automatically start a device for running the machine learning model when a task is added.

[0048] The management unit 170 has a function of, for example, setting the users who use the object detection system 1 based on user operations. In addition, the management unit 170 also performs, for example, setting and changing access restrictions for assets, obtaining and displaying the access status for assets, and correcting data in visualization processing.

[0049] The ensemble setting unit 180 is an example of a setting unit that sets a combination of multiple machine learning models with different output formats based on a user operation. The output formats include, for example, detection-type output, classification-type output, and area detection-type output. The ensemble setting unit 180 will be described in a second embodiment.

[0050] <Advantages of the First Embodiment> As described above, according to the embodiment of the object detection system 1, an object detection device can be realized that includes an operation reception unit 112 that receives user operations, a visualization unit 114 that presents information to the user, a data management unit 400 that manages image data and point cloud data including position information of the object, a learning setting unit 130 that sets a learning process for a machine learning model for detecting the object based on the operation received by the operation reception unit 112, a learning unit 310 that executes the learning process of the machine learning model set by the learning setting unit 130 using the image data and point cloud data and visualizes the results of the learning process using the visualization unit 114, a detection setting unit 140 that sets a detection process using the machine learning model learned by the learning unit 310 based on the operation received by the operation reception unit 112, and a detection unit 320 that executes the detection process set by the detection setting unit 140 using the image data and point cloud data and visualizes the results of the detection process using the visualization unit 114.

[0051] According to the object detection system 1 of the embodiment, an object detection device that can be consistently managed by a user can be realized by using at least one of the following machine learning models: a detection type machine learning model that detects whether or not an object has a defect; a classification type machine learning model that classifies objects; and a segmentation type machine learning model that determines defective areas in an object.

[0052] According to the object detection system 1 of the embodiment, the learning setting unit 130 sets a learning process for each of the detection-type machine learning model, the classification-type machine learning model, and the segmentation-type machine learning model, the learning unit 310 executes the learning process for each of the detection-type machine learning model, the classification-type machine learning model, and the segmentation-type machine learning model based on the learning process set for each of the machine learning models, the detection setting unit 140 sets a detection process for one of the detection-type machine learning model, the classification-type machine learning model, and the segmentation-type machine learning model, and the detection unit 320 executes the detection process using the one machine learning model set by the detection setting unit 140. As a result, according to the object detection system 1, the user can consistently manage the three types of machine learning models to detect defects and the like in the target object.

[0053] The object detection system 1 of the embodiment further includes a path setting unit 150 that sets a path for the robot 210 based on an operation received by the operation receiving unit 112, and the data management unit 400 can acquire image data and point cloud data including position information of the target object from the robot 210. As a result, the object detection system 1 allows the user to consistently manage the path of the robot 210 as well.

[0054] According to the embodiment of the object detection system 1, the system is provided with a task management unit 160 that manages tasks for the robot 210, the learning setting unit 130, the learning unit 310, the detection setting unit 140, and the detection unit 320 based on operations received by the operation reception unit 112, thereby enabling the user to consistently manage these tasks.

[0055] <Second embodiment> Next, a second embodiment will be described. The information processing device according to the second embodiment includes, for example, a reception unit that receives user operations, a setting unit that sets a combination of multiple machine learning models with different output formats based on the operation received by the reception unit, and an execution unit that causes the multiple machine learning models to execute detection processing based on the combination set by the setting unit. The information processing device may be applied to the ensemble setting unit 180 in the first embodiment, but is not limited thereto and can be applied to any system that performs information processing using multiple machine learning models.

[0056] In a second embodiment, the combination and processing order of multiple machine learning engines are expressed by formula information. The formula information is created based on user operations. In the embodiment, machine learning models in the formula information are written as A, B, C, and classes detected or classified by the machine learning models are written as a, b, c. For example, when machine learning model A detects classes a, b, and c, the machine learning model is written as A(a, b, c). When machine learning model A classifies classes a, b, and c, the machine learning model is also written as A(a, b, c).

[0057] The class names detected by different machine learning models may be the same, such as machine learning model A detecting the class "tree" and machine learning model B detecting the class "tree." In this case, it is desirable to distinguish and register the class "tree" of machine learning model A as A-tree and the class "tree" of machine learning model B as B-tree in the information processing device 600. For the sake of simplicity, the embodiments will be described in such a way that machine learning models A and B do not overlap.

[0058] The process of combining output results of machine learning models is performed based on the combining symbol included in the formula information. For example, a combining process that outputs an overlapping region between an area where class a is detected and an area where class b is detected is represented by the combining symbol "and" or "*." A combining process that outputs a combined region where class a is detected and an area where class b is detected is represented by the combining symbol "or" or "+." A combining process that excludes a combined region where class a is detected and an area where class b is detected is represented by the combining symbol "not" or "¬." In the embodiment, "and" is replaced with "*" and "or" is replaced with "+," but the priority of operations in the embodiment differs from the priority of actual linear algebraic operations. That is, in mathematics, when a + b * c is written, b * c is calculated first, but in the embodiment, a + b is calculated first. However, when a + (b * c) is written in the formula information, b * c is calculated first. The combining symbol "div" in the embodiment will be described later.

[0059] FIG. 13 is a block diagram showing an example of an information processing device according to the second embodiment. The information processing device 600 is connected to, for example, a user terminal device 700. The user terminal device 700 corresponds to, for example, the user terminal device 200 described above. The user terminal device 700 provides information based on a user's operation to the information processing device 600. The user terminal device 700 provides information for setting a combination of multiple machine learning models with different output formats to the information processing device 600. The information for setting a combination of multiple machine learning models with different output formats is information expressed by an equation (hereinafter referred to as equation information), as will be described later. The user terminal device 700 displays the processing results of the information processing device 600.

[0060] The information processing device of the second embodiment includes, for example, a data processing unit 610, a preprocessing unit 620, a DL unit 630, a postprocessing unit 640, a combining unit 650, and a result generation unit 660. The information processing device 600 divides the preprocessing unit 620, the DL unit 630, and the result generation unit 660 into parts, and combines the functions of the preprocessing unit 620, the DL unit 630, and the result generation unit 660 based on expression information.

[0061] 14 is a block diagram illustrating processing units in an information processing device. The preprocessing unit 620 performs preprocessing on the entire image, the DL unit 630 performs detection processing on the entire image for each machine learning model specified by the formula information, and each machine learning model outputs the detection results to the combination unit 650 on a class-by-class basis. The combination unit 650 combines the detection results on a class-by-class basis, the DL unit 630 supplies the processing results on a class-by-class basis to the combination unit 650, which combines the detection results on a class-by-class basis, the postprocessing unit 640 performs postprocessing on a class-by-class basis, and the result generation unit 660 outputs the processing results for the entire image. The result generation unit 660 may output the processing results for the entire image on a class-by-class basis, or may output the processing results in which the post-processing results are sequentially overwritten on the input image as needed.

[0062] The data processing unit 610 is an example of a receiving unit that receives formula information supplied from the user terminal device 700. The data processing unit 610 converts the formula information into a processing order of pre-processing, DL processing, and post-processing. The data processing unit 610 controls the pre-processing unit 620, DL unit 630, and post-processing unit 640 so that processing is performed in the converted processing order. The combining unit 650 combines the processing results by providing the processing results of the pre-processing unit 620 to the DL unit 630 and the processing results of the DL unit 630 to the post-processing unit 640. The result generation unit 660 outputs the processing results provided by the combining unit 650 to the user terminal device 700.

[0063] The preprocessing in the preprocessing unit 620 is a process executed before DL processing. The processes executed by the preprocessing unit 620 include, for example, a blur detection process 622 and an orthogonal conversion process 624, but may also include brightness or luminance adjustment processes and image division processes. The blur detection process 622 is a process for filtering blurred images with a small amount of edge detection in the image data. The orthogonal conversion process 624 is a process for correcting image distortion using a predetermined conversion formula.

[0064] The pre-processing unit 620 may perform processing to adjust the amount of processing (resources) of at least one of the downstream DL unit 630, post-processing unit 640, and result generation unit 660. For example, when the DL unit 630 combines multiple DL processes, the upper limit of the processable image size, etc., may differ depending on the type of DL process and the specifications of the GPU in the information processing device. In this case, for example, the pre-processing unit 620 may perform processing to reduce the image to the upper limit that can be processed by the DL process with the strictest resource restrictions among the DL processes, and the image reduction processing is processing to reduce the amount of image data by compressing the image or cropping the image.

[0065] FIG. 15 is a diagram illustrating an example of the operation of the preprocessing unit. Preprocessing is defined as processing performed on the entire image A before processing using a machine learning model in the DL unit 630. The preprocessing unit 620 performs, for example, processing to exclude images that cannot be detected with high accuracy due to blurring from among the images to be supplied to the DL unit 630, and processing to exclude images that cannot be detected with high accuracy due to distortion from among the images to be supplied to the DL unit 630. The content of the preprocessing is described in the header of the formula information, as will be described later. In other words, the preprocessing is defined in the formula information as an element different from the element that defines the combination of machine learning models. Note that the processing results of the preprocessing (e.g., the detection result of a blurred image) may be output from the result generating unit 660, and the processing results using the machine learning model may not be output.

[0066] 16 is a diagram illustrating another example of the operation of the pre-processing unit. The pre-processing unit 620A may divide an image A into multiple images A-1, A-2, and A-3 by BBox decomposition (tree detection). Image A-1 is processed by the DL unit 630-1, the combining unit 650-1, the post-processing unit 640-1, and the result generation unit 660-1. Image A-2 is processed by the DL unit 630-2, the combining unit 650-2, the post-processing unit 640-2, and the result generation unit 660-2. Image A-3 is processed by the DL unit 630-3, the combining unit 650-3, the post-processing unit 640-3, and the result generation unit 660-3. The pre-processing unit 620A may divide image A according to a predetermined rule, but is not limited to this. The pre-processing unit 620A may output an image obtained by extracting a predetermined object from the image using a machine learning model as the divided image. In this case, the output format of the machine learning model may be matched to the output format of the downstream DL unit 630. In addition to or instead of outputting the processing results for the divided multiple images A-1, A-2, and A-3, the result generation unit 660 may collectively output the processing results for the divided multiple images A-1, A-2, and A-3.

[0067] The post-processing in the post-processing unit 640 is a process executed after the DL process. The processes executed by the post-processing unit 640 include, for example, a filter process 642. The filter process 642 is a process based on the size of the detection area, and for example, is a process that eliminates data about a detection area whose proportion to the entire image is smaller than a predetermined value. The filter process 642 is also a process based on the detection position, and for example, is a process that eliminates data about a detection position that is shifted a predetermined amount from the center of the image.

[0068] The result generation unit 660 performs processing to generate an output result of the information processing device 600. The processing executed by the result generation unit 660 includes a mask generation processing 662, a bbox processing 664, a mosaic processing 666, and a data output processing 668. The mask generation processing 662 is processing to generate a mask that paints a detection area with a predetermined color or a mask that surrounds the detection area. The bbox processing 664 is processing to acquire rectangular coordinates that surround the detection area. The mosaic processing 666 is processing to apply mosaic processing to the detection position and detection area in the image, or to areas other than the detection position and detection area. The data output processing 668 is processing to calculate and output information such as the proportion of the image occupied by the detection area or the proportion of a detection area occupied by another detection area. For example, when detecting tree areas and mushroom areas by DL processing, the data output processing 668 may calculate and output the proportion of the tree area to the entire image or the proportion of the mushroom area to the tree area.

[0069] Although the information processing device 600 of this embodiment includes the pre-processing unit 620 and the post-processing unit 640, these units may not be included. Furthermore, the pre-processing unit 620 and the post-processing unit 640 may perform processing other than the processing described in the embodiment.

[0070] The processes executed in the DL unit 630 correspond to three types of machine learning models: classification, detection, and segmentation, and the combination of machine learning models and the processing order are defined by the formula information. The processes executed in the DL unit 630 include, for example, DeepLab process 631, Mask RCNN process 633, SSD process 635, Yolo process 637, and CNN process 639. DeepLab process 631 is DL process using a segmentation-type machine learning model. Mask RCNN process 633 is DL process using detection and segmentation-type machine learning models. SSD process 635 is DL process using a detection-type machine learning model. Yolo process 637 is DL process using a detection-type machine learning model. CNN (Convolutional Neural Network) process 639 is DL process using detection and classification-type machine learning models.

[0071] For example, when AND-combining the detection results of A(a) and B(b), as in the case of classification-type machine learning model A(a) * segmentation-type machine learning model B(b), if class b has been detected by machine learning model B and class a is True, the combination unit 650 sets the output result of the AND-combination to "True." If A(a) * B(b) is class a * class b, the combination unit 650 retains the region detection result of class b. Conversely, if class b * class a, the combination unit 650 outputs the detection result of class a that does not have region data for class b. When OR-combining the detection results, as in the case of class a + class b, if class a is "True" and class b is "False," the combination unit 650 outputs "True" with no region data. Furthermore, when class a*→class b is written, if class a is "True" and class b is "False", the combining unit 650 outputs a detection result of "True" that does not have area data if an area of ​​class b is not detected, and outputs a detection result of "False" if an area of ​​class b is detected.

[0072] 17 is a diagram showing an example of a sequence of the information processing device in the second embodiment. First, the data processing unit 610 starts up, for example, in response to the registration of a task, and downloads formula information, image data, setting parameters, etc. The data processing unit 610 may periodically monitor tasks in the database and start up when a task to be executed by the data processing unit 610 is registered.

[0073] First, the data processing unit 610 controls the preprocessing unit 620 to sequentially start the preprocessing described in the formula information and acquires the processing results. At this time, if the formula information includes multiple preprocessing processes, the data processing unit 610 executes the preprocessing processes sequentially according to the order of the preprocessing processes described in the formula information. Next, the data processing unit 610 supplies the image data as the processing results of the preprocessing to all machine learning models included in the formula information and acquires detection results from all machine learning models.

[0074] Next, the data processing unit 610 causes the combining unit 650, the post-processing unit 640, and the result generating unit 660 to execute processing for each expression. First, the data processing unit 610 causes the combining unit 650 to combine classes in accordance with the expression information, causes the post-processing unit 640 to execute post-processing described in the expression information, and causes the result generating unit 660 to overwrite (merge) the results of the post-processing. Next, the data processing unit 610 performs processing to store the results generated by the result generating unit 660 in a database or the like, and processing to notify the user terminal device 700 of the results generated by the result generating unit 660.

[0075] Next, details of formula input by the user terminal device 700 and processing based on the formula will be described. FIG. 18 is a diagram showing an example of a screen displayed on the user terminal device 700 when generating formula information. When generating formula information, the user terminal device 700 displays an ensemble design screen (GUI) as shown in FIG. 18. Note that the ensemble design screen is not limited to the ensemble design screen shown in FIG. 18, and any other type of screen may be used as long as it allows input of a mathematical formula. The ensemble design screen includes an ensemble class name (Ensembble Class Name), an input field for post-processing, an input field for multiple parameters (Param 1, Param 2), an input field for a link symbol, an input field for a model name, and a save button. The ensemble class is set, for example, for each output unit from the result generation unit 660. The input field for parameter 1 includes, for each machine learning model, an input field for a machine learning model, an input field for a class, and an input field for a not flag, and further includes an input field for a link symbol between the machine learning models entered in the input field for parameter 1. The input field for parameter 2 includes, for each machine learning model, an input field for a machine learning model, an input field for a class, and an input field for a not flag. Furthermore, between the input field for parameter 1 and the input field for parameter 2, there is provided an input field for a combination symbol for combining the detection result of the machine learning model input in the input field for parameter 1 with the detection result of the machine learning model input in the input field for parameter 2. Furthermore, the ensemble design screen includes an Add to Parameter button following parameter 2 and an Add button for an ensemble class with an output different from parameters 1 and 2. Note that post-processing and result generation processing can be written so that they are simply connected after DL processing. The user terminal device 700 can set pre-processing that is set for the entire image, and combining processing and result output processing that are set on a class-by-class basis.

[0076] For example, if machine learning model A outputs classes a, b, and c, machine learning model B outputs classes d, e, and f, and machine learning model C outputs class g, and there are post-processing X and Y and result generation processes α and β, the user terminal device 700 can write the formula (a+b)*c*X*Y*α, but cannot write the post-processing and result generation process before DL processing. In other words, the user terminal device 700 restricts the input of the formula (a+b*α)*c.

[0077] The output type of the result generation unit 660 is the same as the type of the latter part of the formula. For example, if a formula includes a classification-type machine learning model A (a) and a detection-type machine learning model B (b), the output type of class a + class b is the output type of class b (detection type), and the output type of class b + class a is the output type of class a (classification type). Similarly, the output type of class a * class b is the output type of class b (detection type), and the output type of class b * class a is the output type of class a (classification type). As an exception, the output type of an OR operation between a detection-type machine learning model and a segmentation-type machine learning model has both the detection type and the segmentation type. Note that post-processing or result generation processing may be added to convert the detection-type output type to the segmentation-type output type and unify it into the segmentation type.

[0078] The user terminal device 700 may add a function to convert the output type of DL processing when creating formula information. For example, when formula information is created using an existing machine learning model, if the parameters required for the machine learning model's output cannot be output as is, the detection process of the machine learning model may include a parameter conversion process. For example, the output of a bbox-type machine learning model may be converted to a detection-type output, the output of a segmentation-type machine learning model may be converted to only a segmentation-type output, or the output of a classification-type machine learning engine may be converted to output in all types (detection, classification, and segmentation). Note that bbox-type machine learning models and segmentation-type machine learning models must be in a format that allows output results to be obtained for each class. Furthermore, the output of a classification-type machine learning model retains all classes that the machine learning model may output, with areas corresponding to classes recognized by the classification-type machine learning model being "True" and areas other than "True" being "False." Furthermore, in the case of a machine learning model such as maskrcnn that has both segmentation-type and bbox-type output types, both types are output. If an output that has both segmentation type output and bbox type output is not permitted, only segmentation type output may be output.

[0079] Figure 19 is a diagram showing an example of input data and output data in a classification-type machine learning model, Figure 20 is a diagram showing an example of input data and output data in a detection-type machine learning model, and Figure 21 is a diagram showing an example of input data and output data in a segmentation-type machine learning model.

[0080] In the information processing device 600, the input data to and output data from the DL unit 630 are as shown in FIGS. As shown in FIG. 19(a), input data for a classification-type machine learning model is, for example, image data (image), a class list [classname1, classname2, classname3], a parameter {"confidence": 0.8}, and parameters specific to the machine learning model. The class list indicates a list of classes that are acceptable for output from the classification-type machine learning model, and the parameter indicates that the confidence level for acceptable output from the classification-type machine learning model is 0.8. As shown in FIG. 19(b), output data for the classification-type machine learning model is, for example, output data for each class, and is data indicating that the detection result for class 1 is "True," the detection result for class 2 is "False," and the output result for class 3 is "False."

[0081] As shown in FIG. 20(a), input data for a detection-type machine learning model is, for example, image data (image), a class list [classname1, classname2, classname3], a parameter {"confidence": 0.8}, and parameters specific to the machine learning model. The class list indicates a list of classes for which output from the detection-type machine learning model is permitted, and the parameter indicates that the confidence level for which output from the detection-type machine learning model is permitted is 0.8. As shown in FIGS. 20(b) and 20(c), output data for a detection-type machine learning model is, for example, output data for each class, and includes data indicating that the detection result for class 1 is "True" and the coordinate values ​​of two bbox-type rectangles, the detection result for class 2 is "True" and the coordinate values ​​of one bbox-type rectangle, and the output result for class 3 is "False."

[0082] As shown in FIG. 21(a), input data for a segmentation-type machine learning model is, for example, image data (image), a class list [classname1, classname2, classname3], a parameter {"confidence": 0.8}, and parameters specific to the machine learning model. The class list indicates a list of classes for which output from the segmentation-type machine learning model is permitted, and the parameter indicates that the confidence level for which output from the segmentation-type machine learning model is permitted is 0.8. As shown in FIGS. 21(b) and 21(c), output data for the segmentation-type machine learning model is, for example, output data for each class, including a detection result for class 1 of "True" and a matrix corresponding to the area of ​​class 1, a detection result for class 2 of "True" and a matrix corresponding to the area of ​​class 2, and data indicating that the output result for class 3 is "False."

[0083] The expression information includes, for example, a header that defines the pre-processing unit 620, a body that defines the combining unit 650, the post-processing unit 640, and the result generating unit 660, and parameters that supplement the body and header. There is only one header for multiple machine learning models included in one piece of expression information, and a body is defined for each class to be detected. The expression information has, for example, the following format: {header:[X,Y], body:["x=a+b+c"],["y=c*e"], parameter:{ name:{"a":["project1","tree"],"b":["project2",mashroom_B"],,,,,} option:{a:{confidence=0.6}} } }

[0084] In the above formula, X and Y correspond to pre-processing, x and y correspond to post-processing or result generation processing, and a to e correspond to classes. Note that the parameter confidence is set as an option, but the confidence may be set for each class (for example, the confidence for class a is 0.6 or more), and DL processing may be performed for each confidence setting.

[0085] The header describes symbols corresponding to processes in order to define the processing order. For example, if X is orthogonal conversion and Y is edge detection, the processing order will be different between header:[X,Y] and header:[Y,X]. The body describes symbols corresponding to processes according to the processing order, and expressions using the symbols or "+" and AND "*" are used to connect the symbols corresponding to processes. The body can be defined with multiple expressions for one image. Parameters include, for example, information about the machine learning model corresponding to the class name included in the expression, and information indicating that a confidence filter should be applied to the output value of the machine learning model.

[0086] As with the header, the processing order of the body can be changed by the order in which the classes are written. For example, there is the task of adjusting the confidence parameter in DL processing to reduce false positives. In this task, for the same detection result, for example, by defining that areas with a confidence of 0.7 are colored red, areas with a confidence of 0.7 to 0.8 are colored yellow, and areas with a confidence of 0.8 to 0.9 are colored green, it is possible to overwrite the areas initially colored red for a single image with a confidence of 0.7 to 0.8 with yellow, and areas with a confidence of 0.8 to 0.9 with green. This makes it easy to present confidence to the user, even when it is difficult to display confidence for each pixel on an image, as in the output results of a segmentation-based machine learning model. This helps to set the optimal confidence for DL ​​processing.

[0087] The formula information may have the following format, for example: In this formula information, the header includes the process "resize ("process":["resize"])" and the process method ""param":{"resize":"compression_gpumemory"}". The body includes x1 and x2 corresponding to the post-processing and result generation process, classes c1 to c4 included in x1 and x2, and color1 and color2 corresponding to the coloring process. For each class, c1 to c4 include, for example, "rule":"DL" indicating DL processing, "engine":"segmentation2" indicating the DL processing type, "modelId":3947008441290434" indicating the ID of the machine learning model, "class":"02sonsyou_miki" indicating the class, and "confidence":0.1 indicating the confidence level. Furthermore, color1 and color2 include, for example, "rule":"Draw" indicating the drawing process, "method":"paint" indicating the drawing process method, and "color":[255,0,0] indicating color information. Note that the class may be paint indicating painting or mosaic indicating mosaic processing. "header":{"process":["resize"],"param":{"resize":"compression_gpumemory"}}, "body":["x1=c1+c2*color1","x2=c3+c4*color2"], "param":{"name":{"x1":"sonsyo","x2":"hukyu"},"variable":{ "c1":{"rule":"DL","engine":"segmentation2","modelId":3947008441290434,"class":"02sonsyou_miki","confidence":0.1}, "c2":{"rule":"DL","engine":"segmentation2","modelId":5042429856406092,"class":"02sonsyou_eda","confidence":0.7}, "c3":{"rule":"DL","engine":"segmentation2","modelId":1669202033922186,"class":"01fukyu_miki","confidence":0.7}, "c4":{"rule":"DL","engine":"segmentation2","modelId":1669202033922186,"class":"01fukyu_eda","confidence":0.7}, "color1":{"rule":"Draw","method":"paint","color":[255,0,0]}, "color2":{"rule":"Draw","method":"paint","color":[0,255,0]}}

[0088] 22 is a diagram showing an example of the processing and output class of the combining unit 650 for a combination of machine learning models and a symbol for combining. For example, when the formula information is cls (classification-type machine learning model) * dtc (detection-type machine learning model), the combining unit 650 leaves as the processing result an image that is classified into a class specified by the classification-type machine learning model and in which a class specified by the detection-type machine learning model is detected, and the result generating unit 660 outputs the processing result in a classification output format (dtc). Furthermore, when the formula information is dtc-dtc, the combining unit 650 performs processing corresponding to the symbol "-(dev)" by deleting the detection results of classes specified by the detection-type machine learning models that overlap in both detection-type machine learning models.

[0089] 23 is a diagram showing examples of use of each combination of machine learning models. For example, when the formula information is cls (classification-type machine learning model) * dtc (detection-type machine learning model), it is possible to detect images that are classified into the specified class of wooden buildings and in which mushrooms are detected.

[0090] Fig. 24 is a diagram showing an example of AND processing. As shown in Fig. 24, the combining unit 650 combines the detection results of two machine learning models connected by AND(*) so that both detection results overlap, and outputs them in an output format corresponding to the class specified by the machine learning model defined after AND(*).

[0091] Fig. 25 is a diagram showing an example of OR processing. As shown in Fig. 25, the combining unit 650 combines the detection results so that at least one of the detection results of two machine learning models connected by OR (+) is included, and outputs the combined results in an output format corresponding to the class specified by the machine learning model defined after AND (*).

[0092] Fig. 26 is a diagram showing an example of DIV processing. As shown in Fig. 26, the combining unit 650 performs processing to subtract the processing result of the rear machine learning model connected by DIV(-) from the processing result of the front machine learning model.

[0093] 27 is a diagram showing an example of NOT processing. As shown in FIG. 27, the combining unit 650 outputs the result of inverting the processing result of the machine learning model written after NOT(-).

[0094] When the combining unit 650 holds both the box type and the segmentation type (seg) (dtc·seg) (where "·" is an optional combining symbol), for example, if the formula information is (dtc·seg) + seg, the output is ((dtc·seg).bbox or seg.seg) or ((dtc·seg).seg or seg.seg), and the combining can be realized by combining the operations dtc+seg and seg+seg.

[0095] For example, if the expression information is (dtc·seg) * seg, the combining unit 650 can realize the combination by combining the operations dtc*seg and seg*seg, as follows: ((dtc·seg).bbox and seg.seg) or ((dtc·seg).seg and seg.seg), output: if (dtc·seg).bbox == seg.seg : seg = seg.seg. Furthermore, if the order is reversed, the output type will be (dtc·seg), but no special operation will appear.

[0096] Furthermore, when the output format of the formula information has both bbox and segmentation types, such as (dtc·seg), the combining unit 650 converts it to seg+seg, and converts the coordinate values ​​[x1, y1, x2, y2] in the bbox output format into matrices [x][y] in the segmentation output format. This allows the combining unit 650 to perform operations between the converted segmentation type matrices.

[0097] According to the information processing device 600 of the second embodiment, an information processing device can be realized that includes a user terminal device 700 as a reception unit that receives user operations; a data processing unit 610 as a setting unit that sets a combination of multiple machine learning models with different output formats based on the operation received by the user terminal device 700; and a DL unit 630 as an execution unit that causes the multiple machine learning models to perform a detection process based on the combination set by the data processing unit 610. Furthermore, the information processing device 600 can set a processing order for the combination of multiple machine learning models based on user operations, thereby causing the multiple machine learning models to perform a detection process based on the set processing order. Furthermore, the information processing device 600 can set formula information representing the combination of machine learning models, operations between the machine learning models, and the processing order based on user operations, and combine the processing results of the machine learning models based on the formula information. The information processing device 600 can perform, for example, object detection by combining machine learning models with different output methods, such as segmentation (region), detection, and classification, and similar types of machine learning models with different output formats.

[0098] Furthermore, according to the information processing device 600, multiple machine learning models are associated with class names, and multiple machine learning models are identified based on the class names specified by user operation. Therefore, by simply specifying the class names, a combination of machine learning models can be easily used without being aware of the machine learning models.

[0099] Furthermore, the information processing device 600 can set pre-processing, post-processing, and result generation processing for processing information to be input to multiple machine learning models based on user operations, making it possible to break down and combine image processing techniques, including machine learning models, into parts. This allows the information processing device 600 to have high scalability and prevent the need to change the entire system according to demand.

[0100] Although each embodiment and variant has been described, these are merely examples and are not intended to be limiting. For example, one aspect of the present invention may be realized by combining any of the embodiments or variants, or a part of each embodiment or a part of each variant, with one or more other embodiments or one or more other variants.

[0101] In addition, the various processes described above related to the object detection device 100 may be performed by recording a program for executing each process of the object detection system 1 and the information processing device 600 in this embodiment on a computer-readable recording medium, and reading and executing the program recorded on the recording medium into a computer system.

[0102] Note that the term "computer system" here may include hardware such as the OS and peripheral devices. Furthermore, if a WWW system is used, the term "computer system" also includes the homepage provision environment (or display environment). Furthermore, "computer-readable recording media" refers to storage devices such as flexible disks, magneto-optical disks, ROMs, and writable non-volatile memory such as flash memory, portable media such as CD-ROMs, and hard disks built into computer systems.

[0103] Furthermore, the term "computer-readable recording medium" also includes a storage medium that stores a program for a certain period of time, such as a volatile memory (e.g., DRAM (Dynamic Random Access Memory)) within a computer system that serves as a server or client when the program is transmitted via a network such as the Internet or a communication line such as a telephone line. The program may also be transmitted from a computer system that stores the program in a storage device or the like to another computer system via a transmission medium or by transmission waves in the transmission medium.

[0104] Here, the "transmission medium" for transmitting the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The program may also be one that realizes part of the above-mentioned functions. Furthermore, it may be a so-called differential file (differential program) that can realize the above-mentioned functions in combination with a program already recorded in a computer system. [Explanation of symbols]

[0105] 1. Object detection system 100 Web Server 110 User Interface Section 112 Operation reception section 114 Visualization section 120 Data Input Department 130 Learning Settings 132 Annotation Management Unit 140 Detection setting section 150 Route setting section 160 Task Management Department 170 Management Department 180 Ensemble Settings 200 User terminal device 202 Robot Operation Department 210 Robot 300 AI Execution Department 310 Learning Department 320 Detector 400 Data Management Department 402 Database 500 System Settings 600 Information Processing Devices 610 Data Processing Unit 620 Pretreatment section 630 DL section 640 Post-processing section 650 Joint 660 Result generation section 700 User terminal device

Claims

1. A reception unit that receives a user's operation, Based on the operation received by the reception unit, a combination of a plurality of machine learning models including a first machine learning model that outputs a classification type result for classifying an object, a second machine learning model that outputs a detection type result for detecting the presence or absence of a defect in the object, and a third machine learning model that outputs a segmentation type result for determining a defective area in the object, and a setting unit that sets formula information representing the processing order of the first machine learning model, the second machine learning model, and the third machine learning model, An execution unit that inputs one image to be detected to each of the first machine learning model, the second machine learning model, and the third machine learning model included in the combination set by the setting unit, executes a detection process on each of the first machine learning model, the second machine learning model, and the third machine learning model, and combines the processing results of the machine learning models based on the formula information, An information processing apparatus comprising:

2. Each of the first machine learning model, the second machine learning model, and the third machine learning model corresponds to a class name, The setting unit specifies each of the first machine learning model, the second machine learning model, and the third machine learning model based on the class name specified by the reception unit, The information processing apparatus according to claim 1.

3. The setting unit sets the processing order in the combination of the first machine learning model, the second machine learning model, and the third machine learning model based on the operation received by the reception unit, The execution unit executes a detection process on each of the first machine learning model, the second machine learning model, and the third machine learning model based on the processing order set by the setting unit, The information processing apparatus according to claim 1 or 2.

4. The setting unit sets a result generation process based on the output results of each of the first machine learning model, the second machine learning model, and the third machine learning model based on the operation received by the reception unit, The execution unit executes the result generation process set by the setting unit, The information processing apparatus according to claim 1 or 2.

5. The setting unit sets post-processing for processing the output results of the first machine learning model, the second machine learning model, and the third machine learning model respectively, based on the operation received by the reception unit. The execution unit executes the post-processing set by the setting unit. The information processing apparatus according to claim 1 or 2.

6. The setting unit sets pre-processing for processing information to be input to each of the first machine learning model, the second machine learning model, and the third machine learning model, based on the operation received by the reception unit. The execution unit executes the pre-processing set by the setting unit, and uses the result of the pre-processing as the input to each of the first machine learning model, the second machine learning model, and the third machine learning model. The information processing apparatus according to claim 1 or 2.

7. A computer receives a user operation; sets a combination of a plurality of machine learning models including a first machine learning model that outputs a classification type result for classifying an object, a second machine learning model that outputs a detection type result for detecting the presence or absence of a defect in the object, and a third machine learning model that outputs a segmentation type result for determining a defective area in the object, and formula information representing the processing order of the first machine learning model, the second machine learning model, and the third machine learning model, based on the received operation; inputs an image of one detection target to each of the first machine learning model, the second machine learning model, and the third machine learning model included in the plurality of machine learning models based on the set combination, causes each of the first machine learning model, the second machine learning model, and the third machine learning model to execute a detection process, and combines the processing result of the first machine learning model, the processing result of the second machine learning model, and the processing result of the third machine learning model based on the formula information; An information processing method for executing the above.

8. In the computer receives a user operation; A combination of a plurality of machine learning models including a first machine learning model that outputs a classification type result for classifying an object based on a received operation, a second machine learning model that outputs a detection type result for detecting the presence or absence of a defect in the object, and a third machine learning model that outputs a segmentation type result for determining a defective area in the object, and a step of setting expression information representing the processing order of the first machine learning model, the second machine learning model, and the third machine learning model. Based on the set combination, input an image of one detection target into each of the first machine learning model, the second machine learning model, and the third machine learning model included in the plurality of machine learning models, execute a detection process on each of the first machine learning model, the second machine learning model, and the third machine learning model, and combine the processing result of the first machine learning model, the processing result of the second machine learning model, and the processing result of the third machine learning model based on the expression information. A program for executing a process including the above.

Citation Information

Patent Citations

  • Medical image processing apparatus, system and program

    JP2020101860A

  • Defect inspection device, defect inspection method, and program therefor

    JP2020106467A

  • Same structure detection device, same structure detection method, and same structure detection program

    JP2020140334A