Program, method, information processing device and system

By executing a program that evaluates the inference quality of a learning model for each data group within a dataset, the program addresses the inability of existing techniques to assess model performance comprehensively, leading to improved model evaluation and performance.

JP2025083210AActive Publication Date: 2025-05-30ADANSONS INC
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2023196973
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-30
Estimated Expiration
2043-11-20

AI Technical Summary

Technical Problem

The existing techniques are unable to evaluate the inference quality of a learning model for each part, leading to incomplete assessment of model performance.

Method used

A program is executed on a computer to store a learning model, acquire a dataset with metadata, and evaluate the inference quality of the learning model for each data group within the dataset.

Benefits of technology

This approach allows for the comprehensive evaluation of the inference quality of a learning model for each part, enabling targeted improvements and better model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025083210000001_ABST
    Figure 2025083210000001_ABST
Patent Text Reader

Abstract

To enable inference quality of a learning model to be evaluated part by part.SOLUTION: There is provided a program to be executed by a computer including a processor and a storage unit. The program causes the processor to execute: a model storage step of storing a learning model; a data acquisition step of acquiring a dataset consisting of a plurality of pieces of data respectively associated with pieces of metadata; and a model evaluation step of evaluating, for each among a plurality of groups including one or more pieces of data in the data set acquired in the data acquisition step, inference quality of the learning model for the one or more pieces of data included in said group.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a program, a method, an information processing apparatus, and a system.

Background Art

[0002] Techniques for verifying the reliability of artificial intelligence models and the like are known. Patent Document 1 discloses a technique for reducing a decrease in recognition accuracy even when the number of objects is large. Patent Document 2 discloses a technique for providing an image analysis apparatus or the like that recognizes an object based on a reference image from an analysis target image even when there are a plurality of feature points of the same or similar local feature amounts in the reference image.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] There is a problem that the inference quality of a learning model cannot be evaluated for each part. Therefore, the present disclosure has been made to solve the above problems, and an object thereof is to provide a technique for evaluating the inference quality of a learning model for each part.

Means for Solving the Problems

[0005] A program for causing a computer including a processor and a storage unit to execute, the program causing the processor to execute a model storage step of storing a learning model, a data acquisition step of acquiring a data set including a plurality of data each associated with metadata, and a model evaluation step of evaluating the inference quality of the learning model for one or more data included in each of a plurality of groups including one or more data in the data set acquired in the data acquisition step.

Advantages of the Invention

[0006] According to the present disclosure, the inference quality of a learning model can be evaluated for each part.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Mode for Carrying Out the Invention

[0008] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In all the drawings for describing the embodiments, the same reference numerals are given to common components, and repeated descriptions are omitted. Note that the following embodiments do not unduly limit the content of the present disclosure described in the claims. Also, not all of the components shown in the embodiments are essential components of the present disclosure. Also, each figure is a schematic diagram and is not necessarily drawn precisely.

[0009] <Configuration of System 1> The system 1 in the present disclosure is an information processing system for evaluating the quality of a learning model. The learning model includes any artificial intelligence model such as a machine learning, artificial intelligence, and deep learning model. The quality of the learning model includes indicators indicating any quality related to artificial intelligence models such as generality, accuracy, robustness, speed and efficiency, and reliability. The system 1 includes information processing devices of the server 10 and the user terminal 20 connected via the network N. FIG. 1 is a block diagram showing the functional configuration of the system 1. FIG. 2 is a block diagram showing the functional configuration of the server 10. FIG. 3 is a block diagram showing the functional configuration of the user terminal 20. FIG. 12 is a block diagram showing the functional configuration of the combination model.

[0010] Each information processing device is composed of a computer including an arithmetic unit and a storage unit. The basic hardware configuration of the computer and the basic functional configuration of the computer realized by the hardware configuration will be described later. For each of the server 10 and the user terminal 20, descriptions overlapping with the basic hardware configuration of the computer and the basic functional configuration of the computer described later will be omitted. In the present disclosure, as an example, the user terminal 20 and the server 10 have different device configurations, but it is not limited thereto. Specifically, the configuration of the user terminal 20 may include all the configurations of the server 10. In this case, the user terminal 20 can execute the information processing according to the present disclosure with a single stand-alone configuration. In addition, the hardware configuration for realizing the information processing system may be realized in any system configuration as long as it can execute the information processing according to the present disclosure.

[0011] <Configuration of Server 10> The server 10 is an information processing device that provides an information processing service for evaluating the quality of a learning model. The server 10 includes a storage unit 101 and a control unit 104.

[0012] <Configuration of the Storage Unit 101 of Server 10> The storage unit 101 of the server 10 includes an application program 1011, a user table 1012, a data table 1013, an evaluation table 1014, a model table 1021, a group table 1022, and a sub-model table 1023.

[0013] The application program 1011 is a program for causing the control unit 104 of the server 10 to function as each functional unit. The application program 1011 includes applications such as a web browser application.

[0014] The user table 1012 is a table that stores and manages information of member users (hereinafter referred to as users) who use the service. By registering for the service, the information of the user is stored in a new record of the user table 1012. Thereby, the user can use the service according to the present disclosure. The user table 1012 is a table having columns of user ID and user name, with the user ID as the primary key. Figure 4 is a diagram showing the data structure of the user table 1012.

[0015] The user ID is an item for storing user identification information for identifying a user. The user identification information is an item for which a unique value is set for each user. The user name is an item for storing the user's name. The user name may be set to any string such as a nickname instead of the real name.

[0016] The data table 1013 is a table for storing and managing a data set (test data) used to evaluate the quality (generality) of a learning model (main model). Note that the data table 1013 may be configured to store any data set used as input to the learning model such as training data and validation data. The data table 1013 is a table having columns of data ID, user ID, data, and metadata, with the data ID as the primary key. Figure 5 is a diagram showing the data structure of the data table 1013.

[0017] The data ID is an item for storing data identification information for identifying data. The data identification information is an item for which a unique value is set for each data information. The user ID is an item for storing user identification information for identifying a user. The data includes any structured and unstructured data such as image data, video data, text data, and audio data. In the evaluation of the quality of artificial intelligence models such as object detection models used for autonomous driving and the like, the data includes the following information. · Image data The image data includes still image data from in-vehicle cameras. The still image data includes road signs, other vehicles, pedestrians, obstacles, etc. In the present disclosure, the image data may also include video data. · Video data The video data includes time-series video data output from in-vehicle cameras and other sensors. The video data includes videos of the movement of other vehicles, pedestrians, bicycles, etc. related to traffic conditions, intersections, crosswalks, changes in traffic lights, the appearance of the road being traveled and the road with less traffic, etc. The video data includes videos related to various weather conditions such as sunny, rainy, snowy, foggy, etc. related to transcription conditions, videos during the day, at night, and at dawn and dusk. The video data includes videos related to road types such as highways, urban areas, rural roads, mountain roads, etc. The video data includes videos of signs, signals, road markings, etc. related to traffic rules. · Text data The text data includes log data from vehicle sensors and data obtained from external information sources (e.g., traffic information and weather forecasts). · Audio data The audio data includes audio input from in-vehicle microphones and information on ambient sounds input from external microphones. The metadata includes annotation information stored in association with the data, correct labels in object detection models, classification models, etc. In the image data and the like of the data used in artificial intelligence models such as object detection models used for autonomous driving and the like, when creating the data, it includes any incidental information obtained from log data from vehicle sensors and external information sources. It includes the following information. · Data measured by various sensors such as cameras, sunlight sensors, acceleration sensors, etc., separately from the data input to the learning model such as the main model Log data such as event occurrence records and surrounding traffic conditions collected separately from the data input to the learning model such as the main model Information that supplements the data input to the learning model such as the main model or the environment in which the learning model of a vehicle or the like is applied, such as the measurement date and time, location, situation, and information on events that occurred before and after Information representing the electronic attributes of data, such as the file size Data obtained by subjecting one or more of the above data to arbitrary processing Any other information that is related to or is presumed to be related to one or more of the above data In addition, the metadata in the present disclosure may include internal feature amounts output from an intermediate layer of a learning model such as a main model, output data output from a learning model such as a main model, input data input to a learning model such as a main model, data obtained by subjecting the input data to arbitrary processing, and combinations of a plurality of these metadata. The metadata in the present disclosure includes first metadata including internal feature amounts output from an intermediate layer of a learning model such as a main model, second metadata not including internal feature amounts output from an intermediate layer of a learning model such as a main model, third metadata that is output data output from a learning model such as a main model, fourth metadata that is input data input to a learning model such as a main model, fifth metadata that is data obtained by subjecting the input data to arbitrary processing, and any combination of one or more of the first metadata to the fifth metadata. In addition, the metadata in the present disclosure may include data obtained by quantitatively measuring the difference between the prediction of the model and the correct data calculated at the time of learning evaluation, the values of the evaluation function and the loss function, values indicating patterns of mistakes (Example 1: mispredicting a dog as a cat, Example 2: mispredicting a cat as a dog), and values obtained by quantifying the reliability of the prediction.

[0018] The internal feature amounts include, for example, the following information. In a convolutional neural network (CNN) when classifying an input image, from after the input layer that receives the image, the output data of each layer up to the fully connected layer that outputs the input data to be input to the activation function used in the output layer (classification layer) such as the Softmax function that outputs class classification using the data after convolution is included in the internal feature amount. In a CNN, based on the internal feature amount output from the fully connected layer, the final class classification is executed. The internal feature amount includes, in addition to the output data output from each layer of the above network, information used or generated in the calculation process such as weight parameters and gradients used in the calculation of each layer. In tasks other than classification, the output data that generates the final prediction for each task and the information used or generated in the calculation process thereof are included in the internal feature amount. For example, in object detection, it includes the coordinate position of the object (bounding box) and the class classification of the object. In the present disclosure, when performing performance evaluation and quality evaluation (including quality evaluation processing) of the main model, one or more appropriate internal feature amounts may be selected from the above for each task. Specifically, the quality evaluation processing may include a step of selecting one or more internal feature amounts effective for quality evaluation from among a plurality of candidates for internal feature amounts (feature amount candidates) based on the evaluation result of the main model. For example, a first quality evaluation process is executed based on a combination of a plurality of internal feature amounts (referred to as a first internal feature amount set and a second internal feature amount set), and the evaluation result for the first internal feature amount set and the evaluation result for the second internal feature amount set are compared. When the evaluation result in the first internal feature amount set is superior to the evaluation result in the second internal feature amount set, the process may include setting the internal feature amount used for the second quality evaluation process to the first internal feature amount set.

[0019] The evaluation table 1014 is a table for storing and managing the evaluation results (evaluation information) of the main model for each data. The evaluation table 1014 is a table having columns of main model ID, data ID, inference data, and error data. Figure 6 is a diagram showing the data structure of the evaluation table 1014.

[0020] The main model ID is an item for storing main model identification information for identifying the main model. The data ID is an item for storing data identification information for identifying the data. The inference data stores the output data (inference) output from the main model when the data (test data) specified by the data ID is applied as input data to the main model (learning model) specified by the main model ID. The error data is an item for storing information indicating the content, type, etc. of the error determined as a result of comparing the inference data with the metadata of the data specified by the data ID. Specifically, the error data includes information indicating the type of error (error code) and the content of the error (string information indicating the content of the error) as follows. ·Label Error: Refers to an error in labeling within the dataset. A label error may have an adverse effect on model learning because the training data is inaccurate. For example, if an image of a cat is labeled "dog", the model will learn this incorrect information. Label errors often occur due to human errors during data collection or annotation, or errors in automated labeling processes. ·EdgeCase: Indicates the case where the model makes a mistake between categories that are similar or have only minor differences, such as "passenger car" and "pickup truck" in a classification task. ·False Positive (FP): Indicates the case where it is incorrectly predicted as a certain class but is actually not that class. ·False Negative (FN): Indicates the case where it should be predicted as a certain class but is not predicted as such. ·Overfitting: This refers to a situation where the model is overly adapted to the training data and cannot generalize well to new, unknown data. ·Bias: This indicates a case where the prediction is biased towards a specific class or characteristic. Variance: This indicates a case where the prediction varies greatly across different datasets or environments.

[0021] The model table 1021 is a table for storing and managing the main model to be evaluated for quality. In this disclosure, as an example, the learning model is permanently stored in the model table 1021, but it is not limited thereto. For example, it may be configured to be stored in a volatile storage medium such as a memory (RAM, etc.) of the server 10 that can be read and written at high speed. Specifically, the learning model received from outside the server 10 may be temporarily deployed to the volatile storage medium and processed when executing the quality evaluation process and the combined model inference process according to this disclosure. In this case, information processing can be executed at high speed even for large-scale learning models. The model table 1021 is a table having columns for the main model ID, user ID, main model, and model quality, with the main model ID as the primary key. Figure 7 is a diagram showing the data structure of the model table 1021.

[0022] The main model ID is an item for storing the main model identification information for identifying the main model. The main model identification information is an item for which a unique value is set for each main model. The user ID is an item for storing the user identification information for identifying the user. The main model is an item for storing the data of the learning model related to the main model to be evaluated for quality. The learning model is an inference model that outputs (infers) output data in response to the input of input data. The input data may include information related to image data, video data, text data, and audio data. The output data may include information related to metadata. The learning process of the learning model will be described later. The learning model is a type of, for example, machine learning, artificial intelligence, deep learning model, etc. The learning model does not necessarily have to be a single learning model, and it may also be realized by switching between multiple independent learning models. As an example of the learning model, a deep learning model by a deep neural network in deep learning will be described. The learning model does not necessarily have to be a deep learning model, and any machine learning or artificial intelligence model may also be used. The model quality is an item that stores information indicating the evaluation result of the quality of the main model. Specifically, the model quality is an item that stores an index indicating how appropriately the learning model for the main model can make inferences about the data set related to the data. In the present disclosure, the model quality includes the coverage, which is the number of data sets that can be appropriately inferred (no errors occurred) out of the total number of data sets of the data. The model quality includes evaluation indexes of inference quality such as accuracy rate, precision rate, recall rate, F1 score, mean absolute error, mean squared error, etc.

[0023] The group table 1022 is a table for storing and managing information related to groups (group information). The group table 1022 is a table having columns of group ID, main model ID, data IDs, number of data, metadata, error data, and group quality, with the group ID as the primary key. FIG. 8 is a diagram showing the data structure of the group table 1022.

[0024] The group ID is an item for storing group identification information for identifying a group. The group identification information is an item for which a unique value is set for each group information. The main model ID is an item for storing main model identification information for identifying the main model. Data IDs are items that store data identification information of data belonging to a group. One or more pieces of data identification information are stored in the data IDs. The number of data is an item that stores the number of data belonging to a group. Metadata is an item that stores metadata representing (characterizing) a group. The metadata does not necessarily have to be stored in association with a group. In the present disclosure, it is sufficient that at least one of the metadata and the error data is stored in association with each group. Error data is an item that stores information indicating the content and type of an error representing (characterizing) a group. Group quality is an item that stores information indicating the evaluation result of the inference quality of the main model in the data belonging to a group. The group quality includes information comprehensively understanding the execution result of the model, its importance, and its effect. Specifically, the group quality includes the following information. Evaluation indicators of inference quality such as accuracy rate, precision rate, recall rate, F1 score, mean absolute error, and mean squared error. counts_total: Indicates the total number of data belonging to a group. This enables grasping the scale of the data to be evaluated. model_effect_score: Indicates the influence of specific data on the output result of the model. Specifically, it is represented as the sum of the loss functions for each data belonging to a group, and is an indicator showing how much negative impact it has on the main model. Instead of the loss function, any indicator that quantifies the quality of data prediction or the degree of negative impact in operation may be used. priority_score: Indicates an indicator weighted based on other metrics such as ease of improvement, necessity, and urgency of improvement with respect to model_effect_score. For example, when a specific data is very easy to improve and is weighted based on a predetermined priority with respect to model_effect_score, the resulting score is the priority_score.

[0025] The sub-model table 1023 is a table for storing and managing sub-models. The sub-model table 1023 is a table having columns of sub-model ID, main-model ID, sub-model, and application conditions, with the sub-model ID as the primary key. Figure 9 is a diagram showing the data structure of the sub-model table 1023.

[0026] The sub-model ID is an item for storing sub-model identification information for identifying a sub-model. The sub-model identification information is an item for which a unique value is set for each sub-model information. The main-model ID is an item for storing main-model identification information for identifying a main model. The sub-model is an item for storing data of a learning model related to the sub-model. The application conditions are items for defining the range of application conditions of input data to which the sub-model is applied instead of the main model. Specifically, information indicating the range of metadata of the input data is stored in the application conditions. Also, the application conditions may include one or more group IDs (group IDs) for specifying a group to which the sub-model is applied. For example, based on the group IDs, data IDs are specified by referring to the group ID item of the group table 1022. Based on the specified data IDs, metadata is specified by referring to the data ID item of the data table 1013. It may be configured such that the range of application conditions of input data to which the sub-model is applied instead of the main model is defined based on the metadata.

[0027] <Configuration of the control unit 104 of the server 10> The control unit 104 of the server 10 includes a user registration control unit 1041. The control unit 104 realizes each functional unit by executing the application program 1011 stored in the storage unit 101.

[0028] The user registration control unit 1041 performs a process of storing information of a user who wishes to use the service according to the present disclosure in the user table 1012. The information stored in the user table 1012 is such that the user opens a web page or the like operated by the service provider from an arbitrary information processing terminal, inputs information into a predetermined input form, and transmits it to the server 10. The user registration control unit 1041 stores the received information in a new record of the user table 1012, and the user registration is completed. As a result, the user stored in the user table 1012 can use the service. Prior to the registration of user information by the user registration control unit 1041 in the user table 1012, the service provider may perform a predetermined review to restrict whether the user can use the service. The user ID may be any character string or number that can identify the user, any character string or number desired by the user, or the user registration control unit 1041 may automatically set any character string or number.

[0029] <Configuration of the user terminal 20> The user terminal 20 is an information processing device operated by a user who uses the service. The user terminal 20 may be, for example, a mobile terminal such as a smartphone or a tablet, or a stationary PC (Personal Computer) or a laptop PC. It may also be a wearable terminal such as an HMD (Head Mount Display) or a wristwatch-type terminal. The user terminal 20 includes a storage unit 201, a control unit 204, an input device 206, and an output device 208.

[0030] <Configuration of the storage unit 201 of the user terminal 20> The storage unit 201 of the user terminal 20 includes a user ID 2011 and an application program 2012.

[0031] User ID 2011 is the user's account ID. The user sends User ID 2011 from user terminal 20 to server 10. Server 10 identifies the user based on User ID 2011 and provides the services according to the present disclosure to the user. Note that User ID 2011 includes information such as a session ID temporarily assigned by server 10 for identifying the user using user terminal 20.

[0032] Application program 2012 may be pre-stored in storage unit 201 or may be configured to be downloaded from a web server or the like operated by a service provider via a communication IF. Application program 2012 includes applications such as a web browser application. Application program 2012 includes an interpreter-type programming language such as JavaScript (registered trademark) that is executed on a web browser application stored in user terminal 20.

[0033] <Configuration of control unit 204 of user terminal 20> The control unit 204 of user terminal 20 includes an input control unit 2041 and an output control unit 2042. The control unit 204 realizes each functional unit by executing the application program 2012 stored in the storage unit 201.

[0034] <Configuration of input device 206 of user terminal 20> The input device 206 of user terminal 20 includes a camera 2061, a microphone 2062, a position information sensor 2063, a motion sensor 2064, and a touch device 2065.

[0035] <Configuration of output device 208 of user terminal 20> The output device 208 of user terminal 20 includes a display 2081 and a speaker 2082.

[0036] <Operation of system 1> The following describes each process of system 1. FIG. 10 is a flowchart showing the operation of the quality evaluation process. FIG. 11 is a flowchart showing the operation of the combined model inference process. FIG. 13 is a first screen example showing the operation of the quality evaluation process. FIG. 14 is a second screen example showing the operation of the quality evaluation process.

[0037] <Quality Evaluation Process> The quality evaluation process is a process of evaluating the quality of a learning model (main model) and presenting the evaluation result. The quality evaluation process may include a process of creating a sub-model with better quality for a part of the main model based on the evaluation result.

[0038] <Overview of Quality Evaluation Process> The quality evaluation process evaluates by applying a data set consisting of a plurality of test data to the main model, groups (clusters) the data set based on the evaluation result or metadata, evaluates the quality of the main model, visualizes the evaluation result, accepts the selection of a predetermined group, and creates a sub-model that improves the main model within the application range determined based on the selected predetermined group. It is a series of processes.

[0039] <Details of Quality Evaluation Process> The details of the quality evaluation process will be described below.

[0040] In step S101, the control unit 104 of the server 10 executes a model storage step of storing the learning model. The user operates the input device 206 of the user terminal 20 to input the URL of a page (quality evaluation processing page) for executing quality evaluation processing in a web browser or the like, and opens the quality evaluation processing page. The control unit 204 of the user terminal 20 sends a request for opening the quality evaluation processing page to the server 10. The control unit 104 of the server 10 generates a quality evaluation processing page based on the received request and sends it to the user terminal 20. The control unit 204 of the user terminal 20 displays the received quality evaluation processing page on the display 2081 of the user terminal 20. The user operates the input device 206 of the user terminal 20 to select a file upload button or the like provided on the quality evaluation processing page, thereby selecting a learning model to be evaluated for quality in the quality evaluation processing, which is stored in the storage unit 201 of the user terminal 20 or any location such as a predetermined cloud service. The control unit 204 of the user terminal 20 sends the user ID 2011 and the selected learning model to the server 10. The control unit 104 of the server 10 stores the received user ID and learning model in the user ID and main model items of a new record in the model table 1021, respectively. New serial numbers are assigned to the main model identification information for the main model ID. In the present disclosure, the learning model stored in this step is referred to as the main model. Note that the user may also be configured to store the learning model to be evaluated for quality in the quality evaluation processing in the model table 1021 by operating the input device 206 of the user terminal 20 or by executing an arbitrary program at the command line.

[0041] In step S102, the control unit 104 of the server 10 executes a data acquisition step of acquiring a data set composed of a plurality of data each associated with metadata. The user selects a dataset consisting of a plurality of test data used to evaluate the quality of the learning model in the quality evaluation process by operating the input device 206 of the user terminal 20 to select a file upload button or the like provided on the quality evaluation process page. Each of the plurality of test data is stored in association with metadata. The control unit 204 of the user terminal 20 transmits the user ID 2011 and the selected dataset to the server 10. The control unit 104 of the server 10 stores the received user ID, the plurality of test data included in the dataset, and the metadata in the user ID, data, and metadata items of a new record in the data table 1013, respectively. In the data and metadata items of the data table 1013, the plurality of test data included in the dataset are stored in association with the metadata. A new serial number is assigned to the data ID for data identification information. Note that the user may be configured to store the plurality of test data and metadata used to evaluate the quality of the learning model in the quality evaluation process in the data table 1013 by operating the input device 206 of the user terminal 20 or by executing an arbitrary program on the command line.

[0042] The control unit 104 of the server 10 applies each of the plurality of data (test data) stored in the data table 1013 as input data to the main model and obtains internal feature amounts and a plurality of output data as inference results. The control unit 104 of the server 10 evaluates the inference quality of the main model regarding whether the output data is correct or incorrect by comparing each of the plurality of output data with the metadata (correct answer data) stored in association with each of the internal feature amounts and the input data (test data). Specifically, the control unit 104 of the server 10 determines that it is correct when the output data matches the correct answer data and incorrect when it does not match. For example, when the main model is a classification model, the main model outputs an output label (classification label) in response to the input of input data. The control unit 104 of the server 10 compares the output output label with the correct label included in the metadata stored in association with each of the input data. The control unit 104 of the server 10 determines that it is correct when the output label matches the correct label for a plurality of input data, and incorrect when they do not match. In addition, when it is impossible to determine correct / incorrect for a combination of output data and correct data, or when there is no correct data, the inference quality of the main model may be evaluated using any index.

[0043] When the output data of the server 10 is incorrect (when the inference result is wrong), the control unit 104 of the server 10 identifies information (error information) indicating the content of the error related to the reason for the incorrect answer. For example, the information indicating the content of the error includes error codes such as Label Error and EdgeCase (error type information regarding the type of error in the inference result). The information indicating the content of the error includes, for example, information (string information, etc.) regarding the content of the error such as "a label of a cat is attached to a dog image" and "the position and size of the bounding box are inaccurate" when the error code is Label Error. For example, when the error code is Edge Case, it includes information indicating the content of the error such as "detection failure under specific light conditions" and "false detection of an object with an unusual pose or viewpoint". Note that the information indicating the content of the error may be identified by the control unit 104 of the server 10 using any machine learning model, deep learning model, artificial intelligence model, etc. based on the output data and metadata, or may be identified by manual work by the user. As a result, an evaluation result of the inference quality including an index for measuring the accuracy and error of the prediction of the main model with respect to the dataset can be obtained.

[0044] The control unit 104 of the server 10 stores the main model ID, the data ID of the input data, the output data, and the information indicating the content of the error in the main model ID, data ID, inference data, and error data items of a new record in the evaluation table 1014, respectively. Note that when the inference result is "correct" or an inference result that can be regarded as having no error defined by any arbitrary index, the error data may be configured to store null or blank (store nothing). As a result, the evaluation result of the inference quality for the test data is stored in the evaluation table 1014 as evaluation information.

[0045] In the present disclosure, an example in which the quality evaluation process is executed based on the error information is disclosed as an example. However, when the output data is correct (when the inference result is correct), it may be configured to specify information (correct answer information) indicating the content of the correct answer regarding the reason for the correct answer. Specifically, instead of the error information, correct answer information may be used to execute the clustering process (first embodiment) and the clustering process (second embodiment) described later. Also, instead of the error information, correct answer information may be used to execute the quality evaluation process and the combined model inference process. Also, both the error information and the correct answer information (correct / error information) may be used to execute the clustering process (first embodiment) and the clustering process (second embodiment) described later. Also, the correct / error information may be used to execute the quality evaluation process and the combined model inference process.

[0046] <Clustering Process (First Embodiment)> In step S103, the control unit 104 of the server 10 executes clustering based on the similarity of the metadata for the data set acquired in the data acquisition step, and executes a clustering step of identifying the clusters formed by the clustering as groups. Specifically, the control unit 104 of the server 10 searches for the user ID item in the data table 1013 based on the user ID 2011, and acquires a data set composed of a plurality of data IDs, data, and metadata. The control unit 104 of the server 10 executes clustering processing for classifying a plurality of data into different groups based on the similarity of metadata for the plurality of data (test data) included in the data set.

[0047] Specifically, the similarity of metadata is calculated as the distance between a plurality of data in each vector space such as internal feature amounts, shooting date and time, time zone (morning, noon, night, etc.), vehicle running speed, vehicle position information, weather information, etc. As the distance between data, any distance measure such as cosine similarity, Manhattan distance, Euclidean distance can be selected. The control unit 104 of the server 10 classifies each of the plurality of data included in the data set into a plurality of groups (clusters) by using an arbitrary clustering method such as hierarchical clustering, K-means, DBSCAN, etc. based on the distance between the plurality of data.

[0048] The control unit 104 of the server 10 stores, for each of the plurality of groups, the main model ID, the data ID of the data classified into the group, and the number of data classified into the group in the main model ID, data IDs, and number of data items of a new record in the group table 1022. A group ID is newly assigned to the group identification information. Note that group metadata (for example, representative vector, centroid vector, etc. in the cluster) characterizing each group may be stored in the metadata item of the group table 1022. The group metadata may be specified by the method described in the clustering process (second embodiment).

[0049] The control unit 104 of the server 10 executes a group quality evaluation step of obtaining an inference result for each of one or more pieces of data included in a group specified in the clustering step by applying each of the one or more pieces of data as input data of the learning model stored in the model storage step. The group quality evaluation step obtains error information regarding an error in the inference result for each of the one or more pieces of data. Specifically, the control unit 104 of the server 10 refers to the evaluation table 1014 for a predetermined group and obtains evaluation information calculated for the data included in the group. The control unit 104 of the server 10 searches the data ID item of the evaluation table 1014 based on a plurality of data IDs stored in the data IDs of the group information for the predetermined group information, and obtains items of a plurality of inference data and error data for each of the plurality of pieces of data. Note that evaluation information may be calculated in this step.

[0050] The control unit 104 of the server 10 executes a group quality specification step of specifying a group inference quality that characterizes a group for at least a part of a plurality of groups based on the inference result obtained in the group quality evaluation step.

[0051] The control unit 104 of the server 10 calculates evaluation indexes (group inference quality) of the inference quality of the main model in a predetermined group, such as accuracy rate, precision rate, recall rate, F1 score, mean absolute error, and mean squared error, according to the content (number of correct and incorrect answers) of a plurality of error data included in the predetermined group for the predetermined group. In addition, the control unit 104 of the server 10 may calculate evaluation indexes such as counts_total, model_effect_score, and priority_score, which are described in the group quality item of the group table 1022. In addition, for a predetermined group, the control unit 104 of the server 10 may output the group inference quality by applying the acquired plurality of inference data and error data as input data to an arbitrary machine learning model, deep learning model, artificial intelligence model, or the like.

[0052] The group quality identification step identifies group error information and group error type information that characterize the group for at least a part of the plurality of groups. Specifically, for a predetermined group, the control unit 104 of the server 10 identifies, as error data (group error data) that characterizes the predetermined group, the error data with the largest number among the acquired plurality of error data. For example, when, for a predetermined group, the error code related to Label Error is the error code with the largest number, Label Error is identified as the error code (group error information, group error type information) that characterizes the predetermined group. Note that when the inference quality is sufficiently good, such as when the evaluation index of the inference quality for a predetermined group is greater than a predetermined value, a configuration may be adopted in which the error data that characterizes the predetermined group is not identified. That is, the error data that characterizes the group is not necessarily identified for all groups, and a configuration may be adopted in which the error data that characterizes the group is identified for some groups. In addition, for a predetermined group, the control unit 104 of the server 10 may output group error data, group error information, and group error type information by applying the acquired plurality of inference data and error data as input data to an arbitrary machine learning model, deep learning model, artificial intelligence model, or the like.

[0053] The group quality evaluation step acquires group error type information regarding the type of error in the inference result for each of one or more data.

[0054] The control unit 104 of the server 10 executes a group quality storage step of storing the group inference quality specified in the group quality specification step in association with at least a part of a plurality of groups. The group quality storage step stores the group error information and the group error type information specified in the group quality specification step in association with at least a part of a plurality of groups. Specifically, the control unit 104 of the server 10 stores, in the group quality and error data items of the record specified by the group ID of the group in the group table 1022, each of the evaluation index of the inference quality of the main model calculated for the group and the error code characterizing the group. Thereby, in the group table 1022, error data and group quality are stored in association with each group classified by clustering.

[0055] <Clustering process (second embodiment)> In step S103, the control unit 104 of the server 10 executes a data inference step of obtaining an inference result for each of a plurality of data by applying each of the plurality of data included in the data set obtained in the data acquisition step as input data of the learning model stored in the model storage step. The data inference step obtains error information regarding an error in the inference result and error type information regarding the type of error in the inference result for at least a part of the plurality of data. Specifically, the control unit 104 of the server 10 searches the item of the main model ID in the evaluation table 1014 based on the main model ID, and obtains evaluation information including the data ID, the inference data, and the error data. Thereby, the control unit 104 of the server 10 can obtain the inference result based on the main model of the data set stored in the data table 1013.

[0056] In step S103, the control unit 104 of the server 10 executes clustering according to the inference result obtained in the data inference step, and executes a clustering step of identifying the clusters formed by the clustering as groups. The clustering step executes clustering based on the similarity of the error information and error type information obtained in the data inference step for the data set obtained in the data acquisition step, and identifies the clusters formed by the clustering as groups. Specifically, the control unit 104 of the server 10 executes clustering processing for classifying a plurality of data (test data) included in the data set into a plurality of different groups based on the similarity of the error information included in the error data.

[0057] Specifically, the error information includes information regarding the type of error such as an error code. The similarity of the error information is calculated as the distance between a plurality of data in a vector space defined based on the type of error. As the distance between data, any distance metric such as cosine similarity, Manhattan distance, Euclidean distance can be selected. The control unit 104 of the server 10 classifies each of the plurality of data included in the data set into a plurality of groups (clusters) by using any clustering method such as hierarchical clustering, K - means, DBSCAN based on the distance between the plurality of data.

[0058] Note that the error information and the error type itself may be used as classification labels and classified into groups. For example, the data set may be classified according to each error code, such as a group of LabelError error data, a group of EdgeCase error data, etc. Note that data with nothing stored in the error data (data whose inference result is "correct") may be classified into a group indicating no error.

[0059] The control unit 104 of the server 10 stores, for each of a plurality of groups, the main model ID, the data ID of the data classified into the group, and the number of data classified into the group, in the main model ID, data IDs, and number of data items of a new record in the group table 1022. A group ID is newly numbered with group identification information. Note that group error data, group error information, and group error type information (for example, representative vectors, centroid vectors, etc. of error data in a cluster) that characterize each group may be stored in the error data item of the group table 1022. The group error data, group error information, and group error type information may be specified by the method described in the clustering process (first embodiment).

[0060] The control unit 104 of the server 10 executes a group metadata specifying step of specifying group metadata that characterizes at least a part of the plurality of groups based on the metadata associated with one or more data included in the groups specified in the clustering step. Specifically, the control unit 104 of the server 10 refers to the data table 1013 for a predetermined group and acquires the metadata stored in association with the data included in the group. The control unit 104 of the server 10 searches the data ID item of the data table 1013 based on the plurality of data IDs stored in the data IDs of the predetermined group information, and acquires the items of the plurality of metadata stored in association with each of the plurality of data. For a given group, the control unit 104 of the server 10 identifies, as metadata (group metadata) characterizing the given group, the metadata that appears most frequently among the acquired multiple metadata. For example, items of metadata that are common among a plurality of data included in the group (among data at a predetermined ratio or more) (the time is at night, the weather is cloudy, the vehicle speed is 50 km to 60 km) may be used as group metadata. Alternatively, metadata characterizing each group (for example, a representative vector, a centroid vector, etc. in a cluster) may be used as group metadata. Alternatively, for a given group, the control unit 104 of the server 10 may output group metadata by applying the acquired multiple metadata as input data to an arbitrary machine learning model, deep learning model, artificial intelligence model, etc.

[0061] The control unit 104 of the server 10 executes a group metadata storage step of associating and storing the group metadata identified in the group metadata identification step with at least a part of the multiple groups. Specifically, for a given group, the control unit 104 of the server 10 stores the identified group metadata in the metadata item of the record specified by the group ID of the group in the group table 1022.

[0062] <Main Model Quality Evaluation Process> In step S104, the control unit 104 of the server 10 executes a model evaluation step of evaluating the inference quality of the learning model for one or more data included in the group for each of the multiple groups including one or more data in the data set acquired in the data acquisition step. The model evaluation step evaluates the inference quality of the learning model for one or more data included in the group identified in the clustering step. The model evaluation step calculates the inference accuracy of the learning model. Specifically, the control unit 104 of the server 10 retrieves the main model ID in the group table 1022 based on the main model ID, thereby obtaining the group ID, the number of data, error data, and group quality. The control unit 104 of the server 10 can evaluate the inference quality of the learning model for each group by referring to the error data (group error data) and group quality for each group. For example, the control unit 104 of the server 10 calculates evaluation results related to the inference quality, such as the number of groups with sufficiently good inference quality, the number of groups with insufficiently good inference quality, the degree of inference quality for each group (determined based on the evaluation index), and the number of data included in each group, by comparing the evaluation index of the inference quality for each group with a predetermined value. In addition, the control unit 104 of the server 10 retrieves the inference data and error data by searching for the main model ID in the evaluation table 1014 based on the main model ID. The control unit 104 of the server 10 calculates the coverage rate, which is the number of data sets that could be appropriately inferred (without errors) in the number of data sets, as an evaluation result related to the inference quality based on the obtained inference data and error data. In addition, the control unit 104 of the server 10 may output an evaluation result related to the inference quality by applying it to an arbitrary machine learning model, deep learning model, artificial intelligence model, or the like. The control unit 104 of the server 10 stores the evaluation result related to the inference quality in the model quality item of the record specified based on the main model ID in the model table 1021. Thereby, the model quality is stored in association with the main model.

[0063] <Evaluation result presentation process> In step S105, the control unit 104 of the server 10 executes a quality presentation step of presenting a plurality of groups in association with the inference quality evaluated in the model evaluation step. The control unit 104 of the server 10 searches the group table 1022 based on the main model ID and acquires group information. The control unit 104 of the server 10 transmits the acquired group information to the user terminal 20. The control unit 204 of the user terminal 20 generates an evaluation result presentation screen based on the received group information, displays it on the display 2081 of the user terminal 20, and presents it to the user. As a result, the quality of the entire learning model can be interpreted globally for each group range of the dataset. For example, it is possible to visually and intuitively confirm what percentage of the dataset has good inference quality and what percentage has poor inference quality.

[0064] FIG. 13 is a first screen example of the evaluation result presentation screen. The evaluation result presentation screen D1 shows a list of 204 records of group information in a table format having columns of indicators D11, D12, D13 regarding group quality included in the group information and the number of data D14. In the present disclosure, only groups with a group quality below a predetermined value (defective) (referred to as hot spots) are listed, and groups with a group quality above the predetermined value (good) are excluded. Note that both groups with a group quality below a predetermined value (defective) (referred to as hot spots) and groups with a group quality above the predetermined value (good) may be listed. The user can sort the presented multiple pieces of group information according to the order of indicators D11, D12, D13 regarding group quality, the number of data D14, and the like. In addition to the table format, not only the similarity between group metadata but also a hierarchical structure can be defined for each group based on the group quality, and it may be visualized in any format such as a treemap.

[0065] The quality presentation step presents a plurality of groups in association with an influence degree regarding the degree of influence on the learning model calculated based on the inference quality evaluated in the model evaluation step. The evaluation result presentation screen D1 includes the model_effect_score. Thus, the overall quality of the learning model can be interpreted globally according to the degree of influence on the learning model for each group range of the dataset.

[0066] FIG. 14 is a second screen example of the evaluation result presentation screen. The evaluation result presentation screen D3 includes a heatmap D30. The heatmap is a mapping of a multi-dimensional space pasted based on the metadata of the dataset into a two-dimensional space by an arbitrary subspace, multi-dimensional scaling method, etc. The heatmap D30 includes points D31, D32,.... Note that the points D31, D32,... are drawn with color-coding according to the evaluation index of the inference quality included in the group quality of the corresponding group. Specifically, it may be drawn step by step in green as the inference quality is better and in red as the inference quality is worse. The heatmap D30 shows the spread of the metadata space, and the positioning of each group in the space is drawn as points D31, D32,.... By overlooking the heatmap D30, the user can visually and intuitively confirm where in the metadata space the groups with good or bad inference quality are positioned (whether the main model is weak in which metadata area).

[0067] <Group Selection> In step S106, the control unit 204 of the user terminal 20 executes a group selection step of receiving a selection of a predetermined group from the user among a plurality of groups. The user can select the groups (each row) included in the evaluation result presentation screen D1 by operating the input device 206 of the user terminal 20. The user can select the points D31, D32,... included in the evaluation result presentation screen D3 by operating the input device 206 of the user terminal 20. The control unit 204 of the user terminal 20 acquires and receives the group ID associated with the selected group. Note that step S106 may be omitted. For example, the control unit 104 of the server 10 may be configured to automatically select a group that satisfies a predetermined condition such as the group quality being equal to or less than a predetermined value (for example, the inference quality is poor, etc.) from the group table 1022 without receiving a selection operation from the user. Also, the user may be configured to execute a group selection step of receiving a selection of a predetermined group among a plurality of groups by operating the input device 206 of the user terminal 20 or by executing an arbitrary program on the command line.

[0068] <Sub-model creation> In step S107, the control unit 104 of the server 10 executes a model modification step of creating a modification function that modifies the learning model stored in the model storage step based on one or more data included in the group. The model modification step creates a modification function that modifies the learning model stored in the model storage step based on one or more data included in a predetermined group selected in the group selection step. The model modification step creates a modification function that modifies the learning model based on the inference quality of the group identified by applying one or more data included in the group as input data to the learning model.

[0069] Specifically, the control unit 104 of the server 10 searches for the group ID in the group table 1022 based on the group ID of the group selected in step S106, and acquires group information including data IDs, metadata, error data, and group quality. The control unit 104 of the server 10 transmits the group information to the user terminal 20. The control unit 204 of the user terminal 20 displays and presents the received group information on the display 2081 of the user terminal 20. The user can check the group with poor inference quality together with the contents of the metadata, error data, and group quality. The user creates a learned model (modified model, sub-model) that modifies the main model while referring to group information such as metadata, error data, and group quality by operating the input device 206 of the user terminal 20. Specifically, for a group with poor inference quality, the user creates a sub-model with better inference quality than the main model for the dataset included in the group. Although the following methods can be considered as methods for creating the sub-model, any method can be applied. · Use, as the sub-model, the main model re-learned based on the test data included in the group. · Use, as the sub-model, the one modified by feature engineering such as deleting unimportant features of the main model, creating new features, and converting existing features. The feature engineering may be executed automatically or by the user. · Use, as the sub-model, the one with hyperparameters such as the learning rate, regularization parameter, and model depth of the main model adjusted. The adjustment may be executed automatically or by the user. · The main model fine-tuned based on predetermined data may be used as the sub-model. In the present disclosure, as an example, an example of creating a modified model is taken as an example, but it is not limited thereto. For example, it may be configured to create an arbitrary function (one that maps an input value to a predetermined output value), a constant function that outputs a predetermined constant regardless of the input value, and other modification functions (pre-processing and post-processing functions) including the modified model already described. In the embodiments of the present disclosure, the case of using the modified model as the modification function is described as an example, but the scope of application of the present invention is not limited thereto.

[0070] In step S107, the control unit 104 of the server 10 executes a condition storage step of storing, in association with a correction function, an application condition for applying the correction function created in the model correction step. The condition storage step stores, in association with the correction function, the group metadata stored in association with the group of correction functions created in the model correction step as an application condition. Specifically, the control unit 104 of the server 10 stores the main model ID, the created sub-model, and the metadata (group metadata) included in the group information of the selected group in the main model ID, sub-model, and application condition items of a new record in the sub-model table 1023, respectively. Note that, in the present disclosure, the range of metadata is described as an example of the application condition, but it is not limited thereto. For example, any one or a combination of two or more of the metadata included in the group information, error data (group error data), group quality, etc. may be used as the application condition. Further, the application condition may include a condition for information regarding an internal feature amount output from an intermediate layer of the main model. That is, the input data may be once input to the main model, and the condition regarding the output data output from the main model may be included in the application condition. In addition, it may be assumed that the user can define an arbitrary condition as the application condition of the sub-model.

[0071] Note that, in the present disclosure, the created correction function is used to create a combined model in step S108 described later, but it is not limited thereto. For example, the correction function may be used for verification of a data set or correction of a data set, and the configuration may be such that output data therefor can be output. For example, when the quality of the data set created by the user is low (for example, when an image of a "dog" is annotated as "cat", etc.), although the correction function outputs an apparently incorrect inference result (for example, "dog" for the annotation "cat"), there are cases where the inference result is actually correct (the "dog" output by the correction function is actually correct). In such a case, the correction function may output information for notifying the user or the like with an alert or the like. The control unit 104 of the server 10 notifies the user of the output content according to the information output by the correction function. Further, a correction function may be configured to correct the data set.

[0072] <Combined model creation> In step S108, the control unit 104 of the server 10 executes a combined model creation step of creating a combined model based on the learning model and the correction function. Specifically, the control unit 104 of the server 10 creates a combined model by combining the main model and one or more sub-models. FIG. 12 is a block diagram showing the functional configuration of the combined model. The operation of the combined model will be described later in the combined model inference process. Note that the combined model may be configured as a single model different from the main model and one or more sub-models, or may be configured as a model combining the main model and one or more sub-models. In the present disclosure, the combined model includes a model in which one or more sub-models are combined to input input data to a specific area of the input data of the main model. The combined model includes a model in which one or more sub-models output output data with respect to a specific area of the output data of the main model.

[0073] When the input data is included in the application conditions stored in the condition storage step, the combined model created in the combined model creation step outputs the output data by the correction model associated with the application conditions, and when the input data is not included in the application conditions stored in the condition storage step, outputs the output data by the learning model. When the combined model created in the combined model creation step outputs the output data by the correction model associated with the group metadata if the metadata associated with the input data is included in the group metadata, and outputs the output data by the learning model if the metadata associated with the input data is not included in the group metadata.

[0074] The inference process by the combined model will be described in detail in the combined model inference process.

[0075] <Combined Model Inference Process> The combined model inference process is a process of inferring output data for input data based on a combined model that combines two types of learning models, a main model and a submodel.

[0076] <Overview of Combined Model Inference Process> The combined model inference process is a series of processes that receive input data, select a learning model to be applied to the received input data from among the main model and one or more submodels based on the received input data, and output the output data as an inference result by applying the input data to the learning model.

[0077] <Details of Combined Model Inference Process> The details of the combined model inference process will be described below.

[0078] In step S301, the control unit 104 of the server 10 receives the input data. The data to be received may be test data included in the data set received in step S101 of the quality evaluation process. Note that the input data may be received from the user, or actual environment data captured by an in-vehicle camera or the like in a production environment such as an arbitrary information processing service may be received as the input data. The method of receiving the input data is not limited.

[0079] In step S302, the control unit 104 of the server 10 determines whether (or not it matches) the input data is included in the application conditions of the sub-model table 1023 based on the metadata included in the input data.

[0080] When the control unit 104 of the server 10 can identify the sub-model information related to the matching application conditions (when one or more records can be extracted from the sub-model table 1023), it identifies the sub-model included in the extracted records of the sub-model table 1023. When multiple records are extracted, it may be configured to select only one sub-model based on an arbitrary algorithm. For example, a priority order for extraction may be assigned to each sub-model, or it may be randomly extracted. When the control unit 104 of the server 10 cannot identify the sub-model information related to the matching application conditions (when one or more records cannot be extracted from the sub-model table 1023), it identifies the main model. Thereby, for each range of the metadata of the input data, a combined model can be created in which the model to be applied is switched to either the correction model or the learning model. The combined model can output output data with better inference quality than the main model for the input data.

[0081] In the present disclosure, the range of metadata is described as an example of the application conditions, but it is not limited thereto. The control unit 104 of the server 10 may be configured to identify the sub-model when the input data received in step S301 is included in the application conditions, and to identify the main model when it is not included in the application conditions.

[0082] Also, the selection of the applicable model does not necessarily have to be executed before inputting the input data into either the main model or the submodel. For example, when the application conditions of the submodel include conditions regarding the internal feature quantities output from the intermediate layer of the main model after inputting the input data into the main model, it may be configured such that the application of the submodel is selected according to the information of the output data output from the input data. That is, the output data output from the main model may be included in the application conditions. In this case, the input data input to the submodel does not necessarily have to be the input data received in step S301, and the output data of the main model may be input to the submodel. Thus, in the combination model in the present disclosure, it is not limited to the case where the input data received in step S301 is selectively input to either the main model or the submodel, and the submodel may be selected according to the content of the output data output from the main model. In addition, when the quality of the output data of the main model according to the input data is not sufficient (such as when the reliability, accuracy, etc. are lower than a predetermined value), the output data from the submodel may be used as the output data of the combination model.

[0083] In step S303, the control unit 104 of the server 10 inputs the input data received in step S301 as input data to the submodel or the main model specified in step S302. The control unit 104 of the server 10 acquires the output data output from the submodel or the main model. When the submodel is specified in step S302, the control unit 104 of the server 10 acquires the inference result of the submodel with respect to the input data as the output data. When the main model is specified in step S302, the control unit 104 of the server 10 acquires the inference result of the main model with respect to the input data as the output data. Accordingly, for each range of input data (range of groups), a combined model can be created by switching the model to be applied between a modified model and a learning model. The combined model can output output data with excellent inference quality for the input data input to the combined model.

[0084] In step S302, when the application condition of the sub-model includes a condition regarding the internal feature amount output from the intermediate layer of the main model for the input data input to the main model, it is preferable to use the output data from the sub-model as the output data of the combined data. Accordingly, output data with excellent inference quality can be output for the input data input to the combined model.

[0085] FIG. 12 is a block diagram showing the functional configuration of the combined model. The combined model inference process will be described based on the functional block diagram of the combined model. The combined model M1 outputs output data M18 in response to the input of input data M11. The input data M11 input to the combined model M1 is subjected to validation M12 regarding the input data. The input data M11 is input to either the main model M13 or the sub-model M14 according to the validation result for the input data M11. The main model M13 outputs output data in response to the input of the input data M11, and validation M15 regarding the output data is executed. Similarly, the sub-model M14 outputs output data in response to the input of the input data M11, and validation M16 regarding the output data is executed. When the application condition includes a condition regarding the internal feature amount output from the intermediate layer of the main model, according to the result of validation M15, when it meets the application condition, the input data M11 or the output data of the main model is input to the sub-model M14. In this case, the output data of the sub-model M14 is output as the output data M18 of the combined model M1, and the output data of the main model M13 is not output as the output data M18 of the combined model M1. The output data from the main model M13 or the sub-model M14 is output as the output data M18 of the combined model M1. Note that the output data from the sub-model M14 is manually verified M17 according to the result of the validation M16 regarding the output data, and the output data after such manual verification is also reflected in the output data M18. Note that it is not always necessary to execute the verification M17 for the output data from the sub-model M14. For example, when the quality of the dataset created by the user is low (for example, when an image of a "dog" is annotated as "cat", etc.), although the correction function outputs an apparently incorrect inference result (for example, "dog" for the annotation "cat"), in fact, the inference result may be correct (the "dog" output by the correction function is actually correct). In such a case, for the validation M16 regarding the output data of the sub-model, processing for executing a notification such as an alert to the user is performed, and the manual verification M17 is executed. Note that the main model M13 and the sub-model M14 mentioned here may be configured by combining a plurality of models, correction functions, etc., or by passing through a further sub-model after the validation M16.

[0086] <Basic Hardware Configuration of Computer> FIG. 15 is a block diagram showing the basic hardware configuration of a computer 90. The computer 90 includes at least a processor 901, a main memory device 902, an auxiliary storage device 903, and a communication IF 991 (Interface). These are electrically connected to each other by a communication bus 921.

[0087] The processor 901 is hardware for executing an instruction set described in a program. The processor 901 is composed of an arithmetic unit, registers, peripheral circuits, and the like.

[0088] The main memory device 902 is for temporarily storing programs and data processed by programs, etc. For example, it is a volatile memory such as DRAM (Dynamic Random Access Memory).

[0089] The auxiliary storage device 903 is a storage device for storing data and programs. For example, it includes flash memory, HDD (Hard Disc Drive), magneto-optical disk, CD-ROM, DVD-ROM, semiconductor memory, etc.

[0090] The communication IF 991 is an interface for inputting and outputting signals for communicating with other computers via a network using wired or wireless communication standards. The network is composed of various mobile communication systems constructed by the Internet, LAN, wireless base stations, etc. For example, the network includes 3G, 4G, 5G mobile communication systems, LTE (Long Term Evolution), wireless networks (e.g., Wi-Fi (registered trademark)) that can be connected to the Internet by a predetermined access point, etc. When connecting wirelessly, communication protocols such as Z-Wave (registered trademark), ZigBee (registered trademark), Bluetooth (registered trademark), etc. are included. When connecting wired, the network also includes those directly connected by a USB (Universal Serial Bus) cable, etc.

[0091] Note that all or part of each hardware configuration can be distributed and provided in a plurality of computers 90, and the computer 90 can be virtually realized by connecting them to each other via a network. Thus, the computer 90 is a concept that includes not only a single housing or a computer 90 housed in a case, but also a virtualized computer system.

[0092] <Basic Functional Configuration of Computer 90> The functional configuration of a computer realized by the basic hardware configuration of computer 90 (Fig. 15) will be described. The computer includes at least functional units of a control unit, a storage unit, and a communication unit.

[0093] Note that the functional units included in computer 90 can also be realized by dispersing all or part of each functional unit among a plurality of computers 90 interconnected by a network. Computer 90 is a concept that includes not only a single computer 90 but also a virtualized computer system.

[0094] The control unit is realized by the processor 901 reading out various programs stored in the auxiliary storage device 903 and expanding them in the main storage device 902, and executing processing according to the programs. The control unit can realize a functional unit that performs various information processes according to the type of program. Thereby, the computer is realized as an information processing device that performs information processing.

[0095] The storage unit is realized by the main storage device 902 and the auxiliary storage device 903. The storage unit stores data, various programs, and various databases. Also, the processor 901 can secure a storage area corresponding to the storage unit in the main storage device 902 or the auxiliary storage device 903 according to the program. Further, the control unit can cause the processor 901 to execute processes of adding, updating, and deleting data stored in the storage unit according to various programs.

[0096] The database refers to a relational database and is for managing a data set called a table, which is structurally defined by rows and columns, and a master, in association with each other. In a database, a table is called a table, a master, a column of a table is called a column, and a row of a table is called a record. In a relational database, the relationships between tables and masters can be set and associated. Normally, for each table and each master, a column serving as a primary key for uniquely identifying records is set, but setting a primary key for a column is not essential. The control unit can cause the processor 901 to add, delete, or update records in specific tables and masters stored in the storage unit according to various programs. Also, by storing data, various programs, and various databases in the storage unit, the information processing apparatus and information processing system according to the present disclosure can be regarded as being manufactured.

[0097] Note that the databases and masters in the present disclosure may include any data structure (such as a list, dictionary, associative array, object, etc.) in which information is structurally defined. The data structure shall also include data that can be regarded as a data structure by combining data with functions, classes, methods, etc. described in any programming language.

[0098] The communication unit is realized by the communication IF 991. The communication unit realizes the function of communicating with other computers 90 via a network. The communication unit can receive information transmitted from other computers 90 and input it to the control unit. The control unit can cause the processor 901 to execute information processing on the received information according to various programs. Also, the communication unit can transmit information output from the control unit to other computers 90.

[0099] <Supplementary Note> The matters described in each of the above embodiments are appended below.

[0100] (Supplementary Note 1) A program for causing a computer including a processor and a storage unit to execute: a model storage step (S101) in which the processor stores a learning model; a data acquisition step (S102) in which the processor acquires a data set including a plurality of data each associated with metadata; and a model evaluation step (S104) in which, for each of a plurality of groups each including one or more data from the data set acquired in the data acquisition step, the processor evaluates the inference quality of the learning model for the one or more data included in the group. Accordingly, the inference quality of the learning model can be evaluated for each group. For example, a region in which the learning model is weak can be specified as a group.

[0101] (Appendix 2) The program according to Appendix 1, wherein the model evaluation step (S104) is a step of calculating the inference accuracy of the learning model. Accordingly, the inference accuracy of the learning model can be evaluated for each group.

[0102] (Appendix 3) A data inference step (S103) in which the processor applies each of the plurality of data included in the data set acquired in the data acquisition step as input data of the learning model stored in the model storage step, thereby obtaining an inference result for each of the plurality of data; and a clustering step (S103) in which the processor performs clustering according to the inference results obtained in the data inference step and specifies, as a group, a cluster formed by the clustering. The program according to Appendix 1 or 2, wherein the model evaluation step (S104) is a step of evaluating the inference quality of the learning model for one or more data included in the group specified in the clustering step. Accordingly, the inference quality of the learning model can be evaluated for each range of clusters specified according to the inference results.

[0103] (Appendix 4) The data inference step (S103) is a step of obtaining error information regarding an error in an inference result for at least a part of a plurality of data, and the clustering step (S103) is a step of performing clustering based on the similarity of the error information obtained in the data inference step on the data set obtained in the data acquisition step, and specifying, as a group, the cluster constituted by the clustering. The program according to Supplementary Note 3. Thereby, the inference quality of the learning model can be evaluated for each range of clusters specified according to the similarity of the errors in the inference results.

[0104] (Supplementary Note 5) The data inference step (S103) is a step of obtaining error type information regarding the type of an error in an inference result for at least a part of a plurality of data, and the clustering step (S103) is a step of performing clustering based on the similarity of the error type information obtained in the data inference step on the data set obtained in the data acquisition step, and specifying, as a group, the cluster constituted by the clustering. The program according to Supplementary Note 4. Thereby, the inference quality of the learning model can be evaluated for each range of clusters specified according to the similarity of the types of errors in the inference results.

[0105] (Supplementary Note 6) A group metadata specifying step (S103) in which a processor specifies group metadata characterizing the group for at least a part of a plurality of groups based on metadata associated with one or more data included in the group specified in the clustering step, and a group metadata storage step (S103) in which the group metadata specified in the group metadata specifying step is stored in association with at least a part of a plurality of groups. The program according to any one of Supplementary Notes 3 to 5. Accordingly, for each range of clusters specified according to the inference result, the content of the cluster can be interpreted based on the group metadata. For example, for a cluster with poor inference quality, the cause can be interpreted based on the metadata.

[0106] (Appendix 7) The processor executes a clustering step (S103) of performing clustering based on the similarity of metadata on the data set acquired in the data acquisition step and specifying the clusters constituted by the clustering as groups, and the model evaluation step (S104) is a step of evaluating the inference quality of the learning model with respect to one or more data included in the groups specified in the clustering step. The program according to Appendix 1 or 2. Accordingly, the inference quality of the learning model can be evaluated for each range of clusters specified according to the similarity of the metadata.

[0107] (Appendix 8) The processor applies each of one or more data included in the groups specified in the clustering step as input data of the learning model stored in the model storage step, thereby obtaining an inference result for each of the one or more data, a group quality evaluation step (S103), a group quality specification step (S103) of specifying the group inference quality characterizing the group for at least a part of the plurality of groups based on the inference results obtained in the group quality evaluation step, and a group quality storage step (S103) of storing the group inference quality specified in the group quality specification step in association with at least a part of the plurality of groups. The program according to Appendix 7. Accordingly, the inference quality of the learning model can be evaluated for each range of clusters specified according to the similarity of the metadata.

[0108] (Appendix 9) The group quality evaluation step (S103) is a step of obtaining error information regarding an error in an inference result for each of one or more pieces of data, the group quality specification step (S103) is a step of specifying group error information that characterizes the group for at least a part of a plurality of groups, and the group quality storage step (S103) is a step of storing the group error information specified in the group quality specification step in association with at least a part of a plurality of groups. The program described in Supplementary Note 8. Thereby, it is possible to evaluate the inference quality regarding the error in the inference result of the learning model for each range of clusters specified according to the similarity of the metadata.

[0109] (Supplementary Note 10) The group quality evaluation step (S103) is a step of obtaining group error type information regarding the type of error in the inference result for each of one or more pieces of data, the group quality specification step (S103) is a step of specifying group error type information that characterizes the group for at least a part of a plurality of groups, and the group quality storage step (S103) is a step of storing the group error type information specified in the group quality specification step in association with at least a part of a plurality of groups. The program described in Supplementary Note 9. Thereby, it is possible to evaluate the inference quality regarding the type of error in the inference result of the learning model for each range of clusters specified according to the similarity of the metadata.

[0110] (Supplementary Note 11) A quality presentation step (S105) in which a processor presents a plurality of groups in association with the inference quality evaluated in the model evaluation step. The program according to any one of Supplementary Notes 1 to 10. Thereby, the quality of the entire learning model can be comprehensively interpreted for each group range of the dataset. For example, it is possible to visually and intuitively confirm what percentage of the dataset has good inference quality and what percentage has poor inference quality.

[0111] (Appendix 12) The quality presentation step (S105) is a step of presenting a plurality of groups in association with an influence degree regarding the degree of influence on a learning model calculated based on the inference quality evaluated in the model evaluation step, and executes the program described in Appendix 11. Thereby, the quality of the entire learning model can be globally interpreted according to the degree of influence on the learning model for each group range of the dataset.

[0112] (Appendix 13) The processor executes a model modification step (S107) of creating a modification function for modifying the learning model stored in the model storage step based on one or more data included in the group, and executes the program described in any one of Appendices 1 to 12. Thereby, a modification function for modifying the learning model for each group can be created.

[0113] (Appendix 14) The model modification step (S107) is a step of creating a modification function for modifying the learning model based on the inference quality of the group specified by applying one or more data included in the group as input data of the learning model, and executes the program described in Appendix 13. Thereby, for example, a modification function for modifying the learning model can be created for a specific group with low inference quality of the learning model.

[0114] (Appendix 15) The processor executes a group selection step (S106) of receiving a selection of a predetermined group from among a plurality of groups from the user, and the model modification step (S107) is a step of creating a modification function for modifying the learning model stored in the model storage step based on one or more data included in the predetermined group selected in the group selection step, and executes the program described in Appendix 13. Accordingly, according to the selection of the group by the user, a correction function for correcting the learning model for the selected group can be created.

[0115] (Appendix 16) A condition storage step (S107) in which the processor stores, in association with the correction function, an application condition for applying the correction function created in the model correction step, and a combined model creation step (S108) in which a combined model is created based on the learning model and the correction function. The combined model created in the combined model creation step outputs output data by the correction function associated with the application condition when the input data is included in the application condition stored in the condition storage step, and outputs output data by the learning model when the input data is not included in the application condition stored in the condition storage step. A program according to any one of Appendices 13 to 15. Accordingly, for each range of input data (range of groups), a combined model can be created by switching the model to be applied between the correction function and the learning model. The combined model can output output data with excellent inference quality for the input data.

[0116] (Appendix 17) A group metadata specifying step (S103) in which a processor specifies group metadata characterizing at least a part of a plurality of groups based on metadata associated with one or more data included in the group, and a group metadata storage step (S103) in which the group metadata specified in the group metadata specifying step is stored in association with at least a part of the plurality of groups, and a condition storage step (S107) is a step of storing in association with a correction function, the group metadata stored in association with the group of correction functions created in the model correction step as an application condition, and the combined model created in the combined model creation step outputs output data by a correction function associated with the group metadata when the metadata associated with the input data is included in the group metadata, and outputs output data by the learning model when the metadata associated with the input data is not included in the group metadata, the program according to Appendix 16. Thereby, for each range of metadata of the input data, a combined model can be created by switching the model to be applied to either a correction function or a learning model. The combined model can output output data with excellent inference quality for the input data.

[0117] (Appendix 18) A method executed by a computer including a processor and a memory, the method including the processor executing all steps executed in the invention according to any one of Appendices 1 to 17. Thereby, the inference quality of the learning model can be evaluated for each group. For example, an area where the learning model is weak can be specified as a group.

[0118] (Appendix 19) An information processing apparatus including a control unit and a storage unit, the control unit executing all steps executed in the invention according to any one of Appendices 1 to 17. As a result, the inference quality of the learning model can be evaluated for each group. For example, areas that the learning model is weak in can be identified as groups.

[0119] (Appendix 20) A system comprising means for performing all steps executed in the invention according to any one of Appendices 1 to 17. As a result, the inference quality of the learning model can be evaluated for each group. For example, areas that the learning model is weak in can be identified as groups.

Explanation of Signs

[0120] 1 System, 10 Server, 101 Storage Unit, 104 Control Unit, 106 Input Device, 108 Output Device, 20 User Terminal, 201 Storage Unit, 204 Control Unit, 206 Input Device, 208 Output Device

Claims

1. A program for causing a computer including a processor and a memory unit to execute, wherein the processor performs a model storage step of storing a learning model, performs a data acquisition step of acquiring a data set including a plurality of data each associated with metadata, and performs a model evaluation step of evaluating the inference quality of the learning model for one or more data included in each of a plurality of groups including the data set acquired in the data acquisition step, and the program.

2. The model evaluation step is a step of calculating the inference accuracy of the learning model, The program according to claim 1.

3. wherein the processor performs a data inference step of applying each of the plurality of data included in the data set acquired in the data acquisition step as input data of the learning model stored in the model storage step to obtain an inference result for each of the plurality of data, performs a clustering step of performing clustering according to the inference result obtained in the data inference step and specifying a cluster formed by the clustering as a group, and the model evaluation step is a step of evaluating the inference quality of the learning model for the one or more data included in the group specified in the clustering step, The program according to claim 1.

4. The data inference step is a step of obtaining error information regarding an error in the inference result for at least a part of the plurality of data, and the clustering step is a step of performing clustering based on the similarity of the error information obtained in the data inference step on the data set acquired in the data acquisition step and specifying a cluster formed by the clustering as a group, The program according to claim 3.

5. The data inference step is a step of obtaining error type information regarding the type of error in the inference result for at least a part of the plurality of data, The clustering step is to perform clustering on the data set obtained in the data acquisition step based on the similarity of the error type information obtained in the data inference step, and identify the clusters formed by the clustering as groups. The program according to claim 4.

6. The processor A group metadata identification step of identifying group metadata characterizing at least a part of the plurality of groups based on metadata associated with one or more data included in the group identified in the clustering step; A group metadata storage step of storing the group metadata identified in the group metadata identification step in association with at least a part of the plurality of groups; executes The program according to claim 3.

7. The processor A clustering step of performing clustering on the data set obtained in the data acquisition step based on the similarity of the metadata, and identifying the clusters formed by the clustering as groups; executes The model evaluation step is to evaluate the inference quality of the learning model for the one or more data included in the group identified in the clustering step. The program according to claim 1.

8. The processor A group quality evaluation step of obtaining an inference result for each of the one or more data by applying each of the one or more data included in the group identified in the clustering step as input data of the learning model stored in the model storage step; A group quality identification step of identifying a group inference quality characterizing at least a part of the plurality of groups based on the inference results obtained in the group quality evaluation step; A group quality storage step of storing the group inference quality identified in the group quality identification step in association with at least a part of the plurality of groups; executes The program according to claim 7.

9. The group quality evaluation step is a step of obtaining error information regarding an error in an inference result for each of the one or more pieces of data, The group quality specification step is a step of specifying group error information that characterizes the group for at least a part of the plurality of groups, The group quality storage step is a step of storing the group error information specified in the group quality specification step in association with at least a part of the plurality of groups, The program according to claim 8.

10. The group quality evaluation step is a step of obtaining group error type information regarding the type of error in the inference result for each of the one or more pieces of data, The group quality specification step is a step of specifying the group error type information that characterizes the group for at least a part of the plurality of groups, The group quality storage step is a step of storing the group error type information specified in the group quality specification step in association with at least a part of the plurality of groups, The program according to claim 9.

11. The processor A quality presentation step of presenting the plurality of groups in association with the inference quality evaluated in the model evaluation step, executes, The program according to any one of claims 1 to 10.

12. The quality presentation step is a step of presenting the plurality of groups in association with the degree of influence on the learning model calculated based on the inference quality evaluated in the model evaluation step, executes, The program according to claim 11.

13. The processor A model modification step of creating a modification function for modifying the learning model stored in the model storage step based on one or more pieces of data included in the group, executes, The program according to any one of claims 1 to 10.

14. The model modification step is a step of creating the modification function for modifying the learning model based on the inference quality of the group specified by applying one or more pieces of data included in the group as input data of the learning model, The program according to claim 13.

15. The processor A group selection step of receiving a selection of a predetermined group from among the plurality of groups from a user; Execute, The model modification step is a step of creating a modification function that modifies the learning model stored in the model storage step based on one or more pieces of data included in the predetermined group selected in the group selection step. The program according to claim 13.

16. The processor A condition storage step of storing, in association with the modification function, an application condition for applying the modification function created in the model modification step; A combined model creation step of creating a combined model based on the learning model and the modification function; Execute, The combined model created in the combined model creation step When the input data is included in the application conditions stored in the condition storage step, outputs output data by the modification function associated with the application conditions; When the input data is not included in the application conditions stored in the condition storage step, outputs output data by the learning model. The program according to claim 13.

17. The processor A group metadata identification step of identifying group metadata characterizing at least a part of the plurality of groups based on metadata associated with one or more pieces of data included in the group; A group metadata storage step of storing the group metadata identified in the group metadata identification step in association with at least a part of the plurality of groups; Execute, The condition storage step is a step of storing, in association with the modification function, the group metadata stored in association with the group of the modification function created in the model modification step as the application condition; The combined model created in the combined model creation step When the metadata associated with the input data is included in the group metadata, outputs output data by the modification function associated with the group metadata; When the metadata associated with the input data is not included in the group metadata, outputs output data by the learning model. The program according to claim 16.

18. A method executed on a computer comprising a processor and a memory, the method wherein the processor executes all steps executed in the invention according to any one of claims 1 to 10.

19. An information processing apparatus comprising a control unit and a storage unit, the information processing apparatus wherein the control unit executes all steps executed in the invention according to any one of claims 1 to 10.

20. A system comprising means for executing all steps executed in the invention according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method of creating learned model, method of classifying data, computer and program

    JP2020052935A

  • Information processing device, information processing method and program

    JP2022175062A

  • Data processing method and apparatus for training depth information estimation model

    JP2023147276A

  • Active learning to reduce noise in labels

    US20190354810A1

  • Providing performance views associated with performance of a machine learning system

    US20200349466A1