Program, method, information processing device, and system

The system evaluates the inference quality of a learning model for each part by storing the model, acquiring a dataset with metadata, and assessing performance within data groups, addressing the limitations of existing technologies in model evaluation.

WO2025109867A1PCT designated stage expired Publication Date: 2025-05-30ADANSONS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/034613
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-09-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing technologies lack the capability to evaluate the inference quality of a learning model for each part, leading to incomplete assessment of model performance.

Method used

A program and system that store a learning model, acquire a dataset with metadata, and evaluate the inference quality of the learning model for each data group within the dataset, allowing for detailed performance assessment.

Benefits of technology

Enables comprehensive evaluation of the inference quality of a learning model for each part, identifying areas of weakness and improving model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024034613_30052025_PF_FP_ABST
    Figure JP2024034613_30052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a program for causing a computer comprising a processor and a storage unit to execute, via the processor: a model storage step for storing a learning model; a data acquisition step for acquiring a data set comprising a plurality of pieces of data respectively associated with pieces of metadata; and a model evaluation step for evaluating, for each among a plurality of groups including one or a plurality of pieces of data in the data set acquired in the data acquisition step, the quality of inference by the learning model pertaining to the one or more pieces of data included in said group.
Need to check novelty before this filing date? Find Prior Art

Description

Program, method, information processing device, and system

[0001] The present disclosure relates to a program, a method, an information processing device, and a system.

[0002] There are known techniques for verifying the reliability of artificial intelligence models, etc. Patent Document 1 discloses a technique for reducing a decrease in recognition accuracy even when there are a large number of objects. Patent Document 2 discloses a technique for providing an image analysis device or the like that recognizes an object based on a reference image from an image to be analyzed, even when the reference image contains multiple feature points with the same or similar local features.

[0003] JP 2022-188727 A JP 2012-190089 A

[0004] There is a problem in that the inference quality of a learning model cannot be evaluated for each part. Therefore, the present disclosure has been made to solve the above problem, and its purpose is to provide a technology for evaluating the inference quality of a learning model for each part.

[0005] A program to be executed by a computer having a processor and a memory unit, wherein the processor executes a model storage step of storing a learning model, a data acquisition step of acquiring a dataset consisting of multiple data each associated with metadata, and a model evaluation step of evaluating, for each of multiple groups containing one or more data, the inference quality of a learning model for one or more data included in the group for the dataset acquired in the data acquisition step.

[0006] According to the present disclosure, the inference quality of a learning model can be evaluated on a part-by-part basis.

[0007] FIG. 1 is a block diagram showing the functional configuration of the system 1. FIG. 1 is a block diagram showing the functional configuration of the server 10. FIG. 1 is a block diagram showing the functional configuration of the user terminal 20. FIG. 2 is a diagram showing the data structure of a user table 1012. FIG. 3 is a diagram showing the data structure of a data table 1013. FIG. 4 is a diagram showing the data structure of an evaluation table 1014. FIG. 5 is a diagram showing the data structure of a model table 1021. FIG. 6 is a diagram showing the data structure of a group table 1022. FIG. 7 is a diagram showing the data structure of a sub-model table 1023. A flowchart showing the operation of quality evaluation processing. A flowchart showing the operation of bond model inference processing. A block diagram showing the functional configuration of a bond model. A first example screen showing the operation of quality evaluation processing. A second example screen showing the operation of quality evaluation processing. A block diagram showing the basic hardware configuration of a computer 90.

[0008] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In all drawings describing the embodiments, common components are designated by the same reference numerals, and repeated description will be omitted. Note that the following embodiments do not unduly limit the content of the present disclosure described in the claims. Furthermore, not all components shown in the embodiments are necessarily essential components of the present disclosure. Furthermore, each drawing is a schematic diagram and is not necessarily a precise illustration.

[0009] <Configuration of System 1> System 1 in the present disclosure is an information processing system that evaluates the quality of a learning model. The learning model includes any artificial intelligence model such as a machine learning, artificial intelligence, or deep learning model. The quality of the learning model includes any indicator indicating the quality of the artificial intelligence model, such as versatility, accuracy, robustness, speed and efficiency, and reliability. System 1 includes information processing devices, a server 10 and a user terminal 20, connected via a network N. FIG. 1 is a block diagram showing the functional configuration of system 1. FIG. 2 is a block diagram showing the functional configuration of the server 10. FIG. 3 is a block diagram showing the functional configuration of the user terminal 20. FIG. 12 is a block diagram showing the functional configuration of a combined model.

[0010] Each information processing device is configured by a computer including an arithmetic unit and a storage device. The basic hardware configuration of the computer and the basic functional configuration of the computer realized by this hardware configuration will be described later. For the server 10 and the user terminal 20, descriptions that overlap with the basic hardware configuration and basic functional configuration of the computer described later will be omitted. In this disclosure, the user terminal 20 and the server 10 have different device configurations as an example, but this is not limited to this. Specifically, the configuration of the user terminal 20 may include all of the configuration of the server 10. In this case, the user terminal 20 can execute the information processing according to the present disclosure using a standalone configuration. Additionally, the hardware configuration for implementing the information processing system may be any system configuration as long as it is capable of executing the information processing according to the present disclosure.

[0011] <Configuration of Server 10> The server 10 is an information processing device that provides an information processing service for evaluating the quality of a learning model. The server 10 includes a storage unit 101 and a control unit 104.

[0012] <Configuration of Storage Unit 101 of Server 10 > The storage unit 101 of the server 10 includes an application program 1011 , a user table 1012 , a data table 1013 , an evaluation table 1014 , a model table 1021 , a group table 1022 , and a sub-model table 1023 .

[0013] The application programs 1011 are programs for causing the control unit 104 of the server 10 to function as each functional unit. The application programs 1011 include applications such as a web browser application.

[0014] The user table 1012 is a table that stores and manages information about member users (hereinafter referred to as users) who use the service. When a user registers to use the service, the user's information is stored in a new record in the user table 1012. This allows the user to use the service according to the present disclosure. The user table 1012 is a table that has columns for user ID and user name, with the user ID as the primary key. FIG. 4 is a diagram showing the data structure of the user table 1012.

[0015] The user ID is an item that stores user identification information for identifying a user. The user identification information is an item in which a unique value is set for each user. The user name is an item that stores the user's name. The user name may not be a name, but may be any character string such as a nickname.

[0016] The data table 1013 is a table for storing and managing a data set (test data) used to evaluate the quality (versatility) of a learning model (main model). The data table 1013 may also be configured to store any data set used for input to the learning model, such as training data or validation data. The data table 1013 is a table having columns for data ID, user ID, data, and metadata, with the data ID as the primary key. FIG. 5 is a diagram showing the data structure of the data table 1013.

[0017] The data ID is an item that stores data identification information for identifying data. The data identification information is an item that has a unique value assigned to each piece of data information. The user ID is an item that stores user identification information for identifying a user. The data includes any structured or unstructured data, such as image data, video data, text data, and audio data. When evaluating the quality of an artificial intelligence model, such as an object detection model used in autonomous driving, the data includes the following information: Image Data The image data includes still image data from an in-vehicle camera. The still image data includes road markings, other vehicles, pedestrians, obstacles, etc. In the present disclosure, the image data may also include video data. Video Data The video data includes time-series video data output from an in-vehicle camera or other sensors. The video data includes video footage of traffic conditions, such as the movement of other vehicles, pedestrians, and bicycles, intersections, crosswalks, and traffic light changes, and footage of busy roads and roads with less traffic. The video data includes footage of various weather conditions, such as sunny, rainy, snowy, and foggy, as well as footage of daytime, nighttime, and crepuscular conditions. Video data includes footage of roads, such as highways, urban areas, rural roads, and mountain roads, which are related to road types. Video data also includes footage of traffic signs, traffic signals, road markings, and other traffic rules. Text data: Text data includes log data from vehicle sensors and data obtained from external sources (e.g., traffic information and weather forecasts). Audio data: Audio data includes audio input from an in-vehicle microphone and information on ambient sounds input from an external microphone. Metadata includes annotation information stored in association with the data, correct labels in object detection models, classification models, etc. Image data used in artificial intelligence models such as object detection models used in autonomous driving, etc., includes log data from vehicle sensors and any additional information obtained from external sources when the data was created. The following information is included.- Data measured by various sensors such as cameras, sunlight sensors, and acceleration sensors, separate from the data input into learning models such as the main model - Log data such as records of event occurrences and surrounding traffic conditions, collected separate from the data input into learning models such as the main model - Information that supplements the data input into learning models such as the main model or the environment in which a learning model such as a vehicle is operated, such as information on the measurement date and time, location, situation, and events that occurred before and after - Information that represents the electronic attributes of the data, such as file size - Data obtained by performing arbitrary processing on one or more of the above data - Any other information that is related or presumed to be related to one or more of the above data Furthermore, the metadata in this disclosure may include internal features output from the intermediate layer of a learning model such as the main model, output data output from a learning model such as the main model, input data input into a learning model such as the main model, data obtained by performing arbitrary processing on the input data, and other combinations of multiple metadata. The metadata in the present disclosure includes first metadata including internal features output from an intermediate layer of a learning model such as a main model, second metadata not including internal features output from an intermediate layer of a learning model such as a main model, third metadata which is output data output from a learning model such as a main model, fourth metadata which is input data input to a learning model such as the main model, fifth metadata which is data obtained by performing optional processing on the input data, and any combination of one or more of the first to fifth metadata. Furthermore, the metadata in the present disclosure may include data quantitatively measuring the difference between the model prediction calculated during learning evaluation and the correct data, values ​​of an evaluation function and a loss function, values ​​indicating error patterns (Example 1: incorrectly predicting a dog as a cat, Example 2: incorrectly predicting a cat as a dog), and a value quantifying the reliability of the prediction.

[0018] The internal features include, for example, the following information. In a convolutional neural network (CNN) used to classify an input image, the internal features include output data from each layer, from the input layer that receives the image to the fully connected layer that outputs input data to be input to an activation function used in an output layer (classification layer), such as a Softmax function, that uses convolutional data as input data and outputs a class classification. In a CNN, final class classification is performed based on the internal features output from the fully connected layer. In addition to the output data output from each layer of the network, the internal features include information used or generated in the calculation process, such as weight parameters and gradients used in the calculation of each layer. In tasks other than classification, the internal features include output data that generates the final prediction for each task and information used or generated in the calculation process. For example, in object detection, the internal features include the coordinate position of the object (bounding box) and the class classification of the object. In the present disclosure, when evaluating the performance and quality of a main model (including quality evaluation processing), one or more appropriate internal features may be selected from the above for each task. Specifically, the quality evaluation processing may include a step of selecting one or more internal features effective for quality evaluation from among multiple internal feature candidates (feature candidates) based on the evaluation results of the main model. For example, a first quality evaluation processing is performed based on a combination of multiple internal features (referred to as a first internal feature set and a second internal feature set), and the evaluation results for the first internal feature set are compared with the evaluation results for the second internal feature set. If the evaluation results for the first internal feature set are superior to the evaluation results for the second internal feature set, the method may include a step of selecting the first internal feature set as the internal features to be used in the second quality evaluation processing.

[0019] The evaluation table 1014 is a table for storing and managing the evaluation results (evaluation information) of the main model for each data. The evaluation table 1014 is a table having columns for the main model ID, data ID, inference data, and error data. Fig. 6 is a diagram showing the data structure of the evaluation table 1014.

[0020] The main model ID is an item that stores main model identification information for identifying the main model. The data ID is an item that stores data identification information for identifying data. The inference data is stored as output data output (inferred) from the main model when data (test data) identified by the data ID is applied as input data to the main model (training model) identified by the main model ID. The error data is an item that stores information indicating the content, type, etc. of an error determined as a result of comparing the inference data with the metadata of the data identified by the data ID. Specifically, the error data includes information indicating the type of error (error code) and the content of the error (string information indicating the content of the error) as follows: Label Error: This refers to an error in labeling within a dataset. A label error can adversely affect model training due to inaccurate training data. For example, if an image of a cat is labeled "dog," the model will learn this incorrect information. Label errors often occur due to human error during data collection and annotation or errors in the automated labeling process. - EdgeCase: Indicates a case in which the model makes a mistake between similar categories or categories with only minor differences, such as "passenger car" and "pickup truck" in classification tasks. - False Positive (FP): Indicates a case in which a model is incorrectly predicted as a certain class, but is not actually that class. - False Negative (FN): Indicates a case in which a model should be predicted as a certain class, but is not. - Overfitting: Indicates a state in which the model is over-fitted to the training data and is unable to generalize well to new, unknown data. - Bias: Indicates a case in which predictions are biased towards a particular class or characteristic. Variance: Indicates a case in which there is a large variation in predictions across different datasets or environments.

[0021] The model table 1021 is a table for storing and managing main models whose quality is to be evaluated. In the present disclosure, an example in which the learning model is permanently stored in the model table 1021 is used as an example, but the present disclosure is not limited to this. For example, the learning model may be stored in a volatile storage medium that can be read and written at high speed in a memory (such as RAM) provided in the server 10. Specifically, a learning model received from outside the server 10 may be temporarily deployed to the volatile storage medium and processed when performing the quality evaluation process and combined model inference process according to the present disclosure. In this case, information processing can be performed at high speed even for large-scale learning models. The model table 1021 is a table having columns for the main model ID, user ID, main model, and model quality, with the main model ID as the primary key. FIG. 7 is a diagram showing the data structure of the model table 1021.

[0022] The main model ID is a field that stores main model identification information for identifying the main model. The main model identification information is a field that has a unique value set for each main model. The user ID is a field that stores user identification information for identifying a user. The main model is a field that stores learning model data related to the main model whose quality is to be evaluated. The learning model is an inference model that outputs (infers) output data in response to input data. The input data may include information about image data, video data, text data, and audio data. The output data may include information about metadata. The learning process of the learning model will be described later. The learning model is, for example, a type of machine learning, artificial intelligence, or deep learning model. The learning model does not need to be a single learning model, and may be realized by switching between multiple independent learning models. As an example of a learning model, a deep learning model using a deep neural network in deep learning will be described. The learning model does not necessarily need to be a deep learning model, and may be any machine learning or artificial intelligence model. The model quality is a field that stores information indicating the evaluation result of the quality of the main model. Specifically, model quality is an item that stores an index indicating how well a learning model related to a main model was able to perform inference on a data set related to the data. In the present disclosure, model quality includes coverage, which is the number of data sets that were properly inferred (no errors occurred) relative to the number of data sets. Model quality includes evaluation indexes of inference quality such as accuracy rate, precision rate, recall rate, F1 score, mean absolute error, and mean squared error.

[0023] The group table 1022 is a table for storing and managing information about groups (group information). The group table 1022 is a table having the group ID as a primary key and columns for group ID, main model ID, data IDs, number of data, metadata, error data, and group quality. FIG. 8 is a diagram showing the data structure of the group table 1022.

[0024] The group ID is an item that stores group identification information for identifying a group. The group identification information is an item that has a unique value set for each group information. The main model ID is an item that stores main model identification information for identifying a main model. The data IDs are an item that stores data identification information for data belonging to a group. One or more pieces of data identification information are stored in the data IDs. The number of data is an item that stores the number of pieces of data belonging to a group. The metadata is an item that stores metadata that represents (characterizes) a group. The metadata does not necessarily have to be stored in association with a group. In the present disclosure, it is sufficient that at least one of the metadata and the error data is stored in association with each group. The error data is an item that stores information that indicates the content and type of an error that represents (characterizes) a group. The group quality is an item that stores information that indicates the evaluation result of the inference quality of the main model for data belonging to a group. The group quality includes information that comprehensively understands the execution results of the model, its importance, and its effect. Specifically, the group quality includes the following information: evaluation indicators of inference quality such as accuracy rate, precision, recall, F1 score, mean absolute error, and mean squared error. counts_total: Indicates the total number of data belonging to the group. This allows the scale of the data to be evaluated to be understood. model_effect_score: Indicates the impact of specific data on the model output results. Specifically, it is expressed as the sum of the loss functions for each data belonging to the group, and is an index showing the extent of the negative impact on the main model. Instead of a loss function, any index that quantifies the quality of data prediction or the degree of negative impact on operation may be used. priority_score: Indicates an index weighted based on other metrics such as the ease of improvement or the necessity / urgency of improvement for model_effect_score. For example, if specific data is very easy to improve, the priority_score would be a score obtained by weighting model_effect_score based on a predetermined priority.

[0025] The sub-model table 1023 is a table for storing and managing sub-models. The sub-model table 1023 has columns for the sub-model ID, main model ID, sub-model, and application condition, with the sub-model ID as the primary key. FIG. 9 shows the data structure of the sub-model table 1023.

[0026] The submodel ID is an item that stores submodel identification information for identifying a submodel. The submodel identification information is an item that is set to a unique value for each piece of submodel information. The main model ID is an item that stores main model identification information for identifying a main model. The submodel is an item that stores learning model data related to the submodel. The application condition is an item that specifies the range of application conditions for input data to which the submodel is applied instead of the main model. Specifically, the application condition stores information indicating the range of metadata for the input data. The application condition may also include one or more group IDs (group IDs) that specify a group to which the submodel is applied. For example, based on the group IDs, the group ID item in the group table 1022 is referenced to specify data IDs. Based on the identified data IDs, the data ID item in the data table 1013 is referenced to specify metadata. A configuration may also be adopted in which the range of application conditions for input data to which the submodel is applied instead of the main model is specified based on the metadata.

[0027] <Configuration of the control unit 104 of the server 10> The control unit 104 of the server 10 includes a user registration control unit 1041. The control unit 104 executes the application program 1011 stored in the storage unit 101, thereby realizing each functional unit.

[0028] The user registration control unit 1041 performs a process of storing information about users who wish to use the service disclosed herein in the user table 1012. Information stored in the user table 1012 is generated when a user opens a webpage operated by the service provider from an information processing terminal, enters information into a predetermined input form, and transmits the information to the server 10. The user registration control unit 1041 stores the received information in a new record in the user table 1012, completing user registration. This allows the user stored in the user table 1012 to use the service. Prior to the user registration control unit 1041 registering user information in the user table 1012, the service provider may conduct a predetermined screening process to restrict whether or not the user can use the service. The user ID may be any string or number that can identify the user. It may be any string or number desired by the user, or the user registration control unit 1041 may automatically set any string or number.

[0029] <Configuration of User Terminal 20> The user terminal 20 is an information processing device operated by a user who uses a service. The user terminal 20 may be, for example, a mobile terminal such as a smartphone or tablet, or a stationary personal computer (PC) or laptop PC. It may also be a wearable terminal such as a head mounted display (HMD) or a wristwatch terminal. The user terminal 20 includes a storage unit 201, a control unit 204, an input device 206, and an output device 208.

[0030] <Configuration of Storage Unit 201 of User Terminal 20 > The storage unit 201 of the user terminal 20 includes a user ID 2011 and an application program 2012 .

[0031] The user ID 2011 is the user's account ID. The user transmits the user ID 2011 from the user terminal 20 to the server 10. The server 10 identifies the user based on the user ID 2011 and provides the user with the services according to the present disclosure. The user ID 2011 includes information such as a session ID temporarily assigned by the server 10 to identify the user using the user terminal 20.

[0032] The application program 2012 may be stored in advance in the storage unit 201, or may be downloaded from a web server operated by a service provider via a communication IF. The application program 2012 includes an application such as a web browser application. The application program 2012 includes an interpreter-type programming language such as JavaScript (registered trademark) that is executed on a web browser application stored in the user terminal 20.

[0033] <Configuration of control unit 204 of user terminal 20> The control unit 204 of the user terminal 20 includes an input control unit 2041 and an output control unit 2042. The control unit 204 executes an application program 2012 stored in the storage unit 201, thereby realizing each functional unit.

[0034] <Configuration of Input Device 206 of User Terminal 20> The input device 206 of the user terminal 20 includes a camera 2061 , a microphone 2062 , a position information sensor 2063 , a motion sensor 2064 , and a touch device 2065 .

[0035] <Configuration of Output Device 208 of User Terminal 20 > The output device 208 of the user terminal 20 includes a display 2081 and a speaker 2082 .

[0036] <Operation of System 1> Each process of System 1 will be described below. Fig. 10 is a flowchart showing the operation of the quality evaluation process. Fig. 11 is a flowchart showing the operation of the combined model inference process. Fig. 13 is a first example of a screen showing the operation of the quality evaluation process. Fig. 14 is a second example of a screen showing the operation of the quality evaluation process.

[0037] <Quality Evaluation Processing> The quality evaluation processing is processing for evaluating the quality of a learning model (main model) and presenting the evaluation results. The quality evaluation processing may include processing for creating a sub-model with better quality for part of the main model based on the evaluation results.

[0038] <Overview of quality evaluation process> The quality evaluation process is a series of processes that evaluates a dataset consisting of multiple test data by applying it to a main model, groups (clusters) the dataset based on the evaluation results or metadata, evaluates the quality of the main model, visualizes the evaluation results, accepts selection of a specified group, and creates a sub-model that is an improvement over the main model within a range of application determined based on the selected specified group.

[0039] <Details of Quality Evaluation Processing> Details of the quality evaluation processing will be described below.

[0040] In step S101, the control unit 104 of the server 10 executes a model storage step of storing a learning model. The user operates the input device 206 of the user terminal 20 to input the URL of a page for executing the quality evaluation process (quality evaluation processing page) into a web browser or the like, and opens the quality evaluation processing page. The control unit 204 of the user terminal 20 sends a request to open the quality evaluation processing page to the server 10. The control unit 104 of the server 10 generates a quality evaluation processing page based on the received request and sends it to the user terminal 20. The control unit 204 of the user terminal 20 displays the received quality evaluation processing page on the display 2081 of the user terminal 20. The user operates the input device 206 of the user terminal 20 to select a file upload button or the like provided on the quality evaluation processing page, thereby selecting a learning model to be quality evaluated in the quality evaluation process, which is stored in the storage unit 201 of the user terminal 20 or any other location, such as a predetermined cloud service. The control unit 204 of the user terminal 20 transmits the user ID 2011 and the selected learning model to the server 10. The control unit 104 of the server 10 stores the received user ID and learning model in the user ID and main model fields of a new record in the model table 1021, respectively. The main model ID is newly assigned with main model identification information. In the present disclosure, the learning model stored in this step is referred to as the main model. Note that the user may operate the input device 206 of the user terminal 20 to execute a command line or any program, thereby storing the learning model to be quality evaluated in the quality evaluation process in the model table 1021.

[0041] In step S102, the control unit 104 of the server 10 executes a data acquisition step of acquiring a dataset consisting of multiple data pieces, each piece of which is associated with metadata. The user operates the input device 206 of the user terminal 20 to select a file upload button or the like provided on the quality evaluation processing page, thereby selecting a dataset consisting of multiple test data pieces to be used for evaluating the quality of the learning model in the quality evaluation processing. The multiple test data pieces are stored in association with metadata and are stored in any location, such as the storage unit 201 of the user terminal 20 or a predetermined cloud service. The control unit 204 of the user terminal 20 transmits the user ID 2011 and the selected dataset to the server 10. The control unit 104 of the server 10 stores the received user ID, the multiple test data pieces included in the dataset, and the metadata in the user ID, data, and metadata fields of a new record in the data table 1013, respectively. The multiple test data pieces included in the dataset are stored in association with metadata in the data and metadata fields of the data table 1013. The data ID is assigned a new data identification number. In addition, the user may operate the input device 206 of the user terminal 20 to execute a command line or any program, thereby storing multiple test data and metadata used to evaluate the quality of the learning model in the quality evaluation process in the data table 1013.

[0042] The control unit 104 of the server 10 applies each of the multiple data (test data) stored in the data table 1013 as input data to the main model, and obtains internal features and multiple output data as inference results. The control unit 104 of the server 10 compares each of the multiple output data with metadata (correct answer data) stored in association with each of the internal features and the input data (test data), thereby evaluating the inference quality of the main model regarding whether the output data is correct or incorrect. Specifically, the control unit 104 of the server 10 determines that the output data is correct if it matches the correct answer data, and incorrect if it does not match. For example, if the main model is a classification model, the main model outputs an output label (classification label) in response to input data. The control unit 104 of the server 10 compares the output label with the correct answer label included in the metadata stored in association with each of the input data. For the multiple input data, the control unit 104 of the server 10 determines that the output label matches the correct answer label as correct, and incorrect if it does not match. In addition, when there is a combination of output data and correct answer data where it is not possible to determine whether the data is correct or incorrect, or when there is no correct answer data, any indicator may be used to evaluate the inference quality of the main model.

[0043] If the output data is incorrect (if the inference result is incorrect), the control unit 104 of the server 10 identifies information (error information) indicating the content of the error related to the reason for the incorrect answer. For example, the information indicating the content of the error includes an error code such as Label Error or EdgeCase (error type information regarding the type of error in the inference result). For example, if the error code is Label Error, the information indicating the content of the error includes information (such as character string information) regarding the content of the error, such as "a dog image is labeled as a cat" or "the position or size of the bounding box is inaccurate." For example, if the error code is EdgeCase, the information indicating the content of the error includes information such as "detection failed under certain lighting conditions" or "misdetection of an object with an unusual posture or viewpoint." Note that the information indicating the content of the error may be identified by the control unit 104 of the server 10 using any machine learning model, deep learning model, artificial intelligence model, etc. based on the output data and metadata, or may be identified manually by a user. This allows us to obtain evaluation results of inference quality, including indicators measuring the accuracy and error of the main model's predictions for the dataset.

[0044] The control unit 104 of the server 10 stores the main model ID, the data ID of the input data, the output data, and information indicating the content of the error in the main model ID, data ID, inference data, and error data fields of a new record in the evaluation table 1014. Note that if the inference result is a "correct answer" or an inference result that can be considered to have no errors defined by an arbitrary index, null or blanks may be stored in the error data field (nothing may be stored). As a result, the evaluation result of the inference quality for the test data is stored as evaluation information in the evaluation table 1014.

[0045] Note that, in the present disclosure, an example in which a quality evaluation process is executed based on error information has been disclosed as an example. However, if the output data is correct (if the inference result is correct), information indicating the content of the correct answer related to the reason for the correct answer (correct answer information) may be identified. Specifically, a clustering process (first embodiment) and a clustering process (second embodiment) described below may be executed using correct answer information instead of error information. Furthermore, a quality evaluation process and a combined model inference process may be executed using correct answer information instead of error information. Furthermore, a clustering process (first embodiment) and a clustering process (second embodiment) described below may be executed using both error information and correct answer information (correct / incorrect information). Furthermore, a quality evaluation process and a combined model inference process may be executed using correct / incorrect information.

[0046] <Clustering Process (First Embodiment)> In step S103, the control unit 104 of the server 10 performs clustering based on the similarity of metadata on the data set acquired in the data acquisition step, and executes a clustering step in which the clusters formed by the clustering are identified as groups. Specifically, the control unit 104 of the server 10 searches the user ID item in the data table 1013 based on the user ID 2011, and acquires a data set consisting of multiple data IDs, data, and metadata. The control unit 104 of the server 10 performs clustering process on multiple data (test data) included in the data set, based on the similarity of metadata, to classify the multiple data into multiple different groups.

[0047] Specifically, the similarity of metadata is calculated as the distance between multiple data in each vector space, such as internal feature values, shooting date and time, time period (morning, noon, night, etc.), vehicle speed, vehicle location information, and weather information. Any distance measure can be selected for the distance between data, such as cosine similarity, Manhattan distance, or Euclidean distance. The control unit 104 of the server 10 classifies each of the multiple data included in the dataset into multiple groups (clusters) based on the distance between the multiple data using any clustering method, such as hierarchical clustering, K-means, or DBSCAN.

[0048] The control unit 104 of the server 10 stores, for each of a plurality of groups, the main model ID, the data IDs of the data classified into the group, and the number of data items classified into the group in the main model ID, data IDs, and number of data items fields of a new record in the group table 1022. Group identification information is newly assigned as the group ID. Group metadata characterizing each group (e.g., a representative vector, a centroid vector, etc. in a cluster) may be stored in the metadata field of the group table 1022. The group metadata may be identified using the method described in the clustering process (second embodiment).

[0049] The control unit 104 of the server 10 executes a group quality evaluation step in which one or more data items included in the group identified in the clustering step are applied as input data to the learning model stored in the model storage step to acquire an inference result for each of the one or more data items. The group quality evaluation step acquires error information regarding errors in the inference result for each of the one or more data items. Specifically, the control unit 104 of the server 10 references the evaluation table 1014 for a specific group to acquire evaluation information calculated for the data items included in the group. For specific group information, the control unit 104 of the server 10 searches the data ID item in the evaluation table 1014 based on the multiple data IDs stored in the data IDs of the group information, and acquires multiple inference data and error data items for each of the multiple data items. Note that evaluation information may be calculated in this step.

[0050] The control unit 104 of the server 10 executes a group quality determination step of determining a group inference quality that characterizes at least some of the multiple groups based on the inference results acquired in the group quality evaluation step.

[0051] The control unit 104 of the server 10 calculates, for a predetermined group, an evaluation index (group inference quality) of the inference quality of the main model in the predetermined group, such as accuracy rate, precision rate, recall rate, F1 score, mean absolute error, or mean squared error, depending on the content (number of correct answers and incorrect answers) of the plurality of error data included in the predetermined group. The control unit 104 of the server 10 may also calculate evaluation indexes such as counts_total, model_effect_score, and priority_score described in the group quality section of the group table 1022. Additionally, the control unit 104 of the server 10 may output the group inference quality by applying the acquired plurality of inference data and error data for the predetermined group as input data to any machine learning model, deep learning model, artificial intelligence model, or the like.

[0052] The group quality identification step identifies group error information and group error type information that characterize at least some of the multiple groups. Specifically, the control unit 104 of the server 10 identifies the most common error data among the multiple error data acquired for a specific group as the error data (group error data) that characterizes the specific group. For example, if the most common error code for a specific group is a Label Error, Label Error is identified as the error code (group error information, group error type information) that characterizes the specific group. Note that if the inference quality for a specific group is sufficiently good, such as if the inference quality evaluation index is greater than a predetermined value, error data that characterizes the specific group may not be identified. In other words, error data that characterizes the group may not be identified for all groups, but may be identified for some groups. Additionally, the control unit 104 of the server 10 may output group error data, group error information, and group error type information by applying the acquired multiple pieces of inference data and error data for a specified group as input data to any machine learning model, deep learning model, artificial intelligence model, etc.

[0053] The group quality evaluation step acquires group error type information regarding the type of error in the inference result for each of one or more data.

[0054] The control unit 104 of the server 10 executes a group quality storage step of storing the group inference quality identified in the group quality identification step in association with at least some of the multiple groups. The group quality storage step stores the group error information and group error type information identified in the group quality identification step in association with at least some of the multiple groups. Specifically, the control unit 104 of the server 10 stores the evaluation index of the inference quality of the main model calculated for the group and the error code characterizing the group in the group quality and error data fields of the record identified by the group ID of the group in the group table 1022. As a result, the error data and group quality are stored in association with each group classified by clustering in the group table 1022.

[0055] <Clustering Process (Second Embodiment)> In step S103, the control unit 104 of the server 10 executes a data inference step in which each of the plurality of data included in the dataset acquired in the data acquisition step is applied as input data for the learning model stored in the model storage step to acquire an inference result for each of the plurality of data. The data inference step acquires error information regarding errors in the inference result and error type information regarding the type of error in the inference result for at least some of the plurality of data. Specifically, the control unit 104 of the server 10 searches the main model ID item in the evaluation table 1014 based on the main model ID and acquires evaluation information including the data ID, inference data, and error data. This allows the control unit 104 of the server 10 to acquire an inference result based on the main model of the dataset stored in the data table 1013.

[0056] In step S103, the control unit 104 of the server 10 executes clustering according to the inference results acquired in the data inference step, and executes a clustering step in which the clusters formed by the clustering are identified as groups. The clustering step executes clustering on the data set acquired in the data acquisition step based on the similarity of the error information and error type information acquired in the data inference step, and identifies the clusters formed by the clustering as groups. Specifically, the control unit 104 of the server 10 executes a clustering process on multiple data (test data) included in the data set, based on the similarity of error information included in the error data, to classify the multiple data into different groups.

[0057] Specifically, the error information includes information about the type of error, such as an error code. The similarity of the error information is calculated as the distance between multiple pieces of data in a vector space defined based on the type of error. Any distance measure can be selected for the distance between the pieces of data, such as cosine similarity, Manhattan distance, or Euclidean distance. The control unit 104 of the server 10 classifies each of the multiple pieces of data included in the dataset into multiple groups (clusters) based on the distance between the multiple pieces of data using any clustering method, such as hierarchical clustering, K-means, or DBSCAN.

[0058] The error information or the error type itself may be used as a classification label and classified into groups. For example, the data set may be classified by error code, such as into a LabelError group and an EdgeCase group. Data with no error data stored (data for which the inference result is "correct") may be classified into a group indicating no error.

[0059] The control unit 104 of the server 10 stores, for each of a plurality of groups, the main model ID, the data IDs of the data classified into the group, and the number of data items classified into the group in the main model ID, data IDs, and number of data items fields of a new record in the group table 1022. New group identification information is assigned to the group ID. Note that group error data, group error information, and group error type information (e.g., representative vectors and centroid vectors of error data in a cluster) that characterize each group may be stored in the error data field of the group table 1022. The group error data, group error information, and group error type information may be identified by the method described in the clustering process (first embodiment).

[0060] The control unit 104 of the server 10 executes a group metadata identification step in which, based on metadata associated with one or more data items included in the group identified in the clustering step, the control unit 104 of the server 10 identifies group metadata characterizing at least some of the groups. Specifically, the control unit 104 of the server 10 references the data table 1013 for a specific group and acquires metadata stored in association with the data items included in the group. For specific group information, the control unit 104 of the server 10 searches the data table 1013 for the data IDs stored in the data IDs of the specific group information, and acquires multiple metadata items stored in association with each of the multiple data items. The control unit 104 of the server 10 identifies the most numerous metadata item among the acquired multiple metadata items for the specific group as metadata characterizing the specific group (group metadata). For example, metadata items (such as the time being nighttime, the weather being cloudy, and the vehicle speed being 50 to 60 km / h) that are common among multiple data items included in the group (a predetermined percentage or more of the data) may be defined as group metadata. Alternatively, metadata characterizing each group (e.g., a representative vector in a cluster, a centroid vector, etc.) may be used as group metadata. Alternatively, the control unit 104 of the server 10 may output group metadata by applying the acquired metadata for a specific group as input data to any machine learning model, deep learning model, artificial intelligence model, etc.

[0061] The control unit 104 of the server 10 executes a group metadata storage step of storing the group metadata identified in the group metadata identification step in association with at least some of the groups. Specifically, the control unit 104 of the server 10 stores the identified group metadata for a specific group in the metadata field of a record identified by the group ID of the group in the group table 1022.

[0062] <Main Model Quality Evaluation Process> In step S104, the control unit 104 of the server 10 executes a model evaluation step for evaluating the inference quality of the learning model for one or more pieces of data included in each of multiple groups containing one or more pieces of data for the dataset acquired in the data acquisition step. The model evaluation step evaluates the inference quality of the learning model for one or more pieces of data included in the group identified in the clustering step. The model evaluation step calculates the inference accuracy of the learning model. Specifically, the control unit 104 of the server 10 searches the group table 1022 for the main model ID based on the main model ID to acquire the group ID, the number of pieces of data, error data, and group quality. The control unit 104 of the server 10 can evaluate the inference quality of the learning model for each group by referring to the error data (group error data) and group quality for each group. For example, the control unit 104 of the server 10 compares the evaluation index of the inference quality for each group with a predetermined value to calculate evaluation results related to inference quality, such as the number of groups with sufficiently good inference quality, the number of groups with insufficient inference quality, the level of inference quality for each group (determined based on the evaluation index), and the number of pieces of data included in each group. Furthermore, the control unit 104 of the server 10 retrieves inference data and error data by searching for the main model ID in the evaluation table 1014 based on the main model ID. Based on the retrieved inference data and error data, the control unit 104 of the server 10 calculates a coverage, which is the number of datasets that were properly inferred (without errors) relative to the total number of datasets, as an evaluation result of inference quality. Additionally, the control unit 104 of the server 10 may output an evaluation result of inference quality by applying the inference data and error data to any machine learning model, deep learning model, artificial intelligence model, or the like. The control unit 104 of the server 10 stores the evaluation result of inference quality in the model quality field of the record identified based on the main model ID in the model table 1021. As a result, the model quality is stored in association with the main model.

[0063] <Evaluation Result Presentation Process> In step S105, the control unit 104 of the server 10 executes a quality presentation step in which multiple groups are associated with the inference quality evaluated in the model evaluation step and presented. The control unit 104 of the server 10 searches the group table 1022 based on the main model ID and acquires group information. The control unit 104 of the server 10 transmits the acquired group information to the user terminal 20. The control unit 204 of the user terminal 20 generates an evaluation result presentation screen based on the received group information, displays it on the display 2081 of the user terminal 20, and presents it to the user. This allows the quality of the entire learning model to be globally interpreted for each group range of the dataset. For example, it is possible to visually and intuitively confirm what proportion of the dataset has good inference quality and what proportion has poor inference quality.

[0064] FIG. 13 shows an example of the first screen of the evaluation result presentation screen. The evaluation result presentation screen D1 shows a list of 204 group information records in a table format having columns for group quality indicators D11, D12, and D13 included in the group information and the number of data D14. In the present disclosure, only groups (referred to as hotspots) whose group quality is below a predetermined value (poor) are listed, and groups whose group quality is above a predetermined value (good) are excluded. It is also possible to list both groups (referred to as hotspots) whose group quality is below a predetermined value (poor) and groups whose group quality is above a predetermined value (good). The user can rearrange the presented multiple group information items according to the order of the group quality indicators D11, D12, and D13, the number of data D14, etc. Furthermore, in addition to the table format, a hierarchical structure may be defined for each group based on the similarity between group metadata or group quality, and the group information may be visualized in any format, such as a treemap.

[0065] The quality presentation step presents the multiple groups in association with their influence levels on the learning model, which are calculated based on the inference quality evaluated in the model evaluation step. The evaluation result presentation screen D1 includes a model_effect_score. This allows the quality of the entire learning model to be broadly interpreted according to the level of influence each group has on the learning model for each group range of the dataset.

[0066] FIG. 14 shows an example of a second screen of the evaluation result presentation screen. The evaluation result presentation screen D3 includes a heat map D30. A heat map is a mapping of a multidimensional space based on the metadata of a dataset into a two-dimensional space using an arbitrary subspace or multidimensional scaling. The heat map D30 includes points D31, D32, etc. The points D31, D32, etc. are depicted in different colors according to the inference quality evaluation index included in the group quality of the corresponding group. Specifically, the better the inference quality, the greener the color, and the worse the inference quality, the redder the color. The heat map D30 shows the extent of the metadata space, and the position of each group in the space is depicted as points D31, D32, etc. By viewing the heat map D30 from a bird's-eye view, the user can visually and intuitively confirm where in the metadata space groups with good and poor inference quality are positioned (in which metadata areas the main model is weak or weak).

[0067] <Group Selection> In step S106, the control unit 204 of the user terminal 20 executes a group selection step in which the user selects a specific group from among multiple groups. The user can select a group (each row) included in the evaluation result presentation screen D1 by operating the input device 206 of the user terminal 20. The user can select points D31, D32, etc. included in the evaluation result presentation screen D3 by operating the input device 206 of the user terminal 20. The control unit 204 of the user terminal 20 acquires and accepts the group ID associated with the selected group. Note that step S106 may be omitted. For example, the control unit 104 of the server 10 may be configured to automatically select a group that satisfies a predetermined condition, such as group quality being below a predetermined value (e.g., inference quality being poor), from the group table 1022 without accepting a selection operation from the user. Alternatively, the user may execute a command line or an arbitrary program by operating the input device 206 of the user terminal 20 to execute the group selection step in which the user selects a specific group from among multiple groups.

[0068] <Sub-model creation> In step S107, the control unit 104 of the server 10 executes a model correction step of creating a correction function by correcting the learning model stored in the model storage step, based on one or more data included in the group. The model correction step creates a correction function by correcting the learning model stored in the model storage step, based on one or more data included in the predetermined group selected in the group selection step. The model correction step creates a correction function by correcting the learning model, based on the inference quality of the group identified by applying one or more data included in the group as input data for the learning model.

[0069] Specifically, the control unit 104 of the server 10 searches the group table 1022 for the group ID of the group selected in step S106, and obtains group information including data IDs, metadata, error data, and group quality. The control unit 104 of the server 10 transmits the group information to the user terminal 20. The control unit 204 of the user terminal 20 displays and presents the received group information on the display 2081 of the user terminal 20. The user can check the groups with poor inference quality along with the metadata, error data, and group quality. By operating the input device 206 of the user terminal 20, the user creates a training model (corrected model, sub-model) by correcting the main model while referring to group information such as metadata, error data, and group quality. Specifically, for groups with poor inference quality, the user creates a sub-model with better inference quality than the main model for the dataset included in the group. Note that the following methods are possible for creating a sub-model, but any method can be applied. The main model is re-trained based on the test data included in the group to create a sub-model. A sub-model is a model modified through feature engineering, such as deleting unimportant features of the main model, creating new features, or converting existing features. Feature engineering may be performed automatically or by a user. A sub-model is a model obtained by adjusting hyperparameters of the main model, such as the learning rate, regularization parameter, and model depth. The adjustments may be performed automatically or by a user. A sub-model may be a model obtained by fine-tuning the main model based on predetermined data. Note that, although the present disclosure has used an example of creating a modified model as an example, this is not limiting. For example, any function (that maps an input value to a predetermined output value), a constant function that outputs a predetermined constant regardless of the input value, or other modified functions (pre-processing or post-processing functions) including the modified model already described may be created.In the embodiments of the present disclosure, a case where a correction model is used as a correction function will be described as an example, but the scope of application of the present invention is not limited to this.

[0070] In step S107, the control unit 104 of the server 10 executes a condition storage step of storing application conditions for applying the correction function created in the model correction step in association with the correction function. In the condition storage step, the group metadata stored in association with the group of the correction function created in the model correction step is stored as application conditions in association with the correction function. Specifically, the control unit 104 of the server 10 stores the main model ID, the created sub-model, and metadata (group metadata) included in the group information of the selected group in the main model ID, sub-model, and application condition fields of a new record in the sub-model table 1023, respectively. Note that, although the scope of metadata has been described as an example of the application conditions in this disclosure, this is not limiting. For example, the application conditions may be any one or a combination of two or more of the metadata included in the group information, error data (group error data), and group quality. Furthermore, the application conditions may include conditions for information regarding internal features output from the intermediate layer of the main model. That is, the application conditions may include conditions for input data once to the main model and output data from the main model. Alternatively, the user may be able to set any conditions as application conditions for the sub-model.

[0071] In the present disclosure, the created correction function is used to create a binding model in step S108, which will be described later, but is not limited thereto. For example, the correction function may be used to verify or correct a dataset, and may be configured to output output data for this purpose. For example, when the quality of the dataset created by the user itself is low (e.g., when an image of a "dog" is annotated as a "cat"), the correction function may output a seemingly incorrect inference result (e.g., "dog" for the annotation "cat"), but the inference result may actually be correct (the "dog" output by the correction function is actually correct). In such cases, the correction function may output information for issuing a notification, such as an alert, to the user. The control unit 104 of the server 10 notifies the user of the output content according to the information output by the correction function. The correction function may also be configured to correct the dataset.

[0072] <Combined Model Creation> In step S108, the control unit 104 of the server 10 executes a combined model creation step of creating a combined model based on the learning model and the correction function. Specifically, the control unit 104 of the server 10 creates the combined model by combining a main model and one or more sub-models. FIG. 12 is a block diagram showing the functional configuration of the combined model. The operation of the combined model will be described later in the combined model inference process. Note that the combined model may be configured as a single model separate from the main model and one or more sub-models, or may be configured as a model combining the main model and one or more sub-models. In the present disclosure, the combined model includes a combined model in which input data is input to one or more sub-models for a specific region of input data of the main model. The combined model includes a combined model in which one or more sub-models output output data for a specific region of output data of the main model.

[0073] The combined model created in the combined model creation step outputs output data by a modified model associated with the application conditions when the input data is included in the application conditions stored in the condition storage step, and outputs output data by a learning model when the input data is not included in the application conditions stored in the condition storage step. The combined model created in the combined model creation step outputs output data by a modified model associated with the group metadata when metadata associated with the input data is included in the group metadata, and outputs output data by a learning model when metadata associated with the input data is not included in the group metadata.

[0074] The inference process using the combined model will be explained in detail in the combined model inference process.

[0075] <Combined Model Inference Processing> The combined model inference processing is processing for inferring output data for input data based on a combined model that combines two types of learning models, a main model and a sub-model.

[0076] <Overview of combined model inference processing> The combined model inference processing is a series of processes that accepts input data, selects a learning model to which the input data is to be applied from among a main model and one or more sub-models based on the accepted input data, and applies the input data to the learning model to output output data as an inference result.

[0077] <Details of the Combined Model Inference Process> The combined model inference process will be described in detail below.

[0078] In step S301, the control unit 104 of the server 10 accepts input data. The accepted data may be test data included in the data set accepted in step S101 of the quality evaluation process. The input data may be accepted from a user, or real-world data captured by an in-vehicle camera or the like in a production environment of an arbitrary information processing service or the like may be accepted as input data. The method for accepting input data is not limited.

[0079] In step S302, the control unit 104 of the server 10 determines whether or not the input data is included in the application conditions of the sub-model table 1023 (whether or not it matches), based on the metadata included in the input data.

[0080] If the control unit 104 of the server 10 can identify sub-model information related to matching application conditions (if one or more records can be extracted from the sub-model table 1023), it identifies the sub-model included in the extracted record in the sub-model table 1023. Note that if multiple records are extracted, only one sub-model may be selected based on an arbitrary algorithm. For example, a priority order may be assigned to each sub-model during extraction, or the sub-models may be extracted randomly. If the control unit 104 of the server 10 cannot identify sub-model information related to matching application conditions (if one or more records cannot be extracted from the sub-model table 1023), it identifies the main model. This allows for the creation of a combined model in which the applied model is switched between a correction model and a learning model for each range of metadata in the input data. The combined model can output output data with superior inference quality for input data compared to the main model.

[0081] In this disclosure, the range of metadata has been described as an example of the application condition, but the application condition is not limited to this. The control unit 104 of the server 10 may identify a sub-model if the input data received in step S301 falls within the application condition, and may identify a main model if the input data does not fall within the application condition.

[0082] Furthermore, the selection of the applicable model does not necessarily need to be performed before inputting input data into either the main model or the sub-model. For example, if the application conditions for the sub-model include conditions related to input data to the main model and internal features output from the intermediate layer of the main model, the application of the sub-model may be selected based on information about the output data output from the input data. That is, the application conditions may include the output data output from the main model. In this case, the input data input to the sub-model does not necessarily have to be the input data received in step S301; the output data of the main model may be input to the sub-model. As such, the combined model in the present disclosure is not limited to a case where the input data received in step S301 is selectively input to either the main model or the sub-model, and the sub-model may be selected based on the content of the output data output from the main model. Additionally, if the quality of the output data of the main model corresponding to the input data is insufficient (e.g., if the reliability, accuracy, etc. are lower than a predetermined value), the output data from the sub-model may be used as the output data of the combined model.

[0083] In step S303, the control unit 104 of the server 10 inputs the input data received in step S301 as input data to the sub-model or main model identified in step S302. The control unit 104 of the server 10 acquires output data output from the sub-model or main model. If a sub-model is identified in step S302, the control unit 104 of the server 10 acquires the inference result of the sub-model for the input data as output data. If a main model is identified in step S302, the control unit 104 of the server 10 acquires the inference result of the main model for the input data as output data. This makes it possible to create a combined model by switching the model to be applied between the correction model and the learning model for each range of input data (range of group). The combined model can output output data with excellent inference quality for the input data input to the combined model.

[0084] In step S302, if the application conditions for the sub-model include conditions for inputting input data to the main model and for internal features output from the intermediate layer of the main model, it is preferable to use the output data from the sub-model as the output data of the combined data, thereby enabling output data with excellent inference quality to be output for the input data input to the combined model.

[0085] FIG. 12 is a block diagram showing the functional configuration of the binding model. The binding model inference process will be described based on the functional block diagram of the binding model. The binding model M1 outputs output data M18 in response to input of input data M11. The input data M11 input to the binding model M1 is subjected to input data validation M12. The input data M11 is input to either the main model M13 or the sub-model M14 depending on the validation results for the input data M11. The main model M13 outputs output data in response to the input of the input data M11, and validation M15 is performed on the output data. Similarly, the sub-model M14 outputs output data in response to the input of the input data M11, and validation M16 is performed on the output data. Note that if the application conditions include conditions related to internal features output from the intermediate layer of the main model, depending on the results of validation M15, if the application conditions are met, the input data M11 or the output data of the main model is input to the sub-model M14. In this case, the output data of the sub-model M14 is output as the output data M18 of the combined model M1, and the output data of the main model M13 is not output as the output data M18 of the combined model M1. The output data from the main model M13 or the sub-model M14 is output as the output data M18 of the combined model M1. The output data from the sub-model M14 is manually verified M17 depending on the results of the validation M16 for the output data, and the output data after the manual verification is also reflected in the output data M18. The output data from the sub-model M14 does not necessarily need to be verified M17. For example, when the quality of the dataset itself created by the user is low (e.g., when an image of a "dog" is annotated as a "cat"), the correction function may output a seemingly incorrect inference result (e.g., "dog" for the annotation "cat"), but the inference result may actually be correct (the "dog" output by the correction function is actually correct). In such a case, validation M16 of the output data of the submodel involves processing for issuing a notification such as an alert to the user, and manual verification M17 is then performed.The main model M13 and sub-model M14 may be configured by combining a plurality of models, correction functions, etc., or by passing through a sub-model after validation M16.

[0086] 15 is a block diagram showing the basic hardware configuration of a computer 90. The computer 90 includes at least a processor 901, a main storage device 902, an auxiliary storage device 903, and a communication IF 991 (interface), which are electrically connected to one another by a communication bus 921.

[0087] The processor 901 is hardware for executing an instruction set written in a program, and is composed of an arithmetic unit, a register, a peripheral circuit, and the like.

[0088] The main storage device 902 is used to temporarily store programs and data to be processed by the programs, etc. For example, it is a volatile memory such as a DRAM (Dynamic Random Access Memory).

[0089] The auxiliary storage device 903 is a storage device for saving data and programs, such as a flash memory, a hard disk drive (HDD), a magneto-optical disk, a CD-ROM, a DVD-ROM, or a semiconductor memory.

[0090] The communication IF 991 is an interface for inputting and outputting signals for communicating with other computers via a network using a wired or wireless communication standard. The network is composed of the Internet, a LAN, various mobile communication systems constructed using wireless base stations, etc. For example, the network includes 3G, 4G, and 5G mobile communication systems, LTE (Long Term Evolution), and wireless networks (e.g., Wi-Fi (registered trademark)) that can connect to the Internet via a predetermined access point. In the case of a wireless connection, communication protocols include, for example, Z-Wave (registered trademark), ZigBee (registered trademark), and Bluetooth (registered trademark). In the case of a wired connection, the network also includes a network that is directly connected using a USB (Universal Serial Bus) cable, etc.

[0091] It should be noted that the computer 90 can be virtually realized by distributing all or part of each hardware configuration across multiple computers 90 and interconnecting them via a network. In this way, the concept of the computer 90 includes not only a computer 90 housed in a single housing or case, but also a virtualized computer system.

[0092] <Basic Functional Configuration of Computer 90> A description will be given of the functional configuration of the computer realized by the basic hardware configuration (FIG. 15) of the computer 90. The computer includes at least the functional units of a control unit, a storage unit, and a communication unit.

[0093] The functional units of the computer 90 can also be realized by distributing all or part of the functional units among multiple computers 90 interconnected via a network. The computer 90 is a concept that includes not only a single computer 90 but also a virtualized computer system.

[0094] The control unit is realized by the processor 901 reading various programs stored in the auxiliary storage device 903, loading them into the main storage device 902, and executing processing in accordance with the programs. The control unit can realize functional units that perform various types of information processing depending on the type of program. In this way, the computer is realized as an information processing device that performs information processing.

[0095] The storage unit is realized by a main storage device 902 and an auxiliary storage device 903. The storage unit stores data, various programs, and various databases. The processor 901 can allocate a storage area corresponding to the storage unit in the main storage device 902 or the auxiliary storage device 903 in accordance with the programs. The control unit can cause the processor 901 to add, update, and delete data stored in the storage unit in accordance with the various programs.

[0096] The term "database" refers to a relational database, which manages data sets called tables and masters in a tabular format structurally defined by rows and columns, by associating them with each other. In a database, a table is called a table, a master, a column in a table is called a column, and a row in a table is called a record. In a relational database, relationships between tables and masters can be established and associated. Typically, each table and each master is assigned a column that serves as a primary key to uniquely identify a record, but assigning a primary key to a column is not required. The control unit can cause the processor 901 to add, delete, or update records in specific tables and masters stored in the storage unit according to various programs. Furthermore, by storing data, various programs, and various databases in the storage unit, the information processing device and information processing system according to the present disclosure can be considered to have been manufactured.

[0097] Note that the databases and masters in this disclosure may include any data structure in which information is structurally defined (such as a list, dictionary, associative array, or object). The data structure also includes data that can be considered as a data structure by combining data with functions, classes, methods, etc. written in any programming language.

[0098] The communication unit is realized by the communication IF 991. The communication unit realizes the function of communicating with other computers 90 via a network. The communication unit can receive information transmitted from other computers 90 and input the information to the control unit. The control unit can cause the processor 901 to execute information processing on the received information in accordance with various programs. Furthermore, the communication unit can transmit information output from the control unit to other computers 90.

[0099] <Additional Notes> The matters described in the above embodiments will be added below.

[0100] (Supplementary Note 1) A program to be executed by a computer including a processor and a storage unit, the program comprising: a model storage step (S101) in which a learning model is stored; a data acquisition step (S102) in which a dataset consisting of a plurality of pieces of data, each piece of data being associated with metadata; and a model evaluation step (S104) in which, for each of a plurality of groups containing one or more pieces of data, the inference quality of the learning model for one or more pieces of data included in the dataset acquired in the data acquisition step is evaluated. This makes it possible to evaluate the inference quality of the learning model for each group. For example, it is possible to identify areas in which the learning model is weak as a group.

[0101] (Supplementary Note 2) The model evaluation step (S104) is a step of calculating the inference accuracy of the learning model in the program according to Supplementary Note 1. This makes it possible to evaluate the inference accuracy of the learning model for each group.

[0102] (Supplementary Note 3) The program according to Supplementary Note 1 or 2, wherein the processor executes a data inference step (S103) of acquiring an inference result for each of the plurality of data by applying each of the plurality of data included in the dataset acquired in the data acquisition step as input data for the learning model stored in the model storage step, and a clustering step (S103) of performing clustering according to the inference result acquired in the data inference step and identifying clusters formed by the clustering as groups, and the model evaluation step (S104) is a step of evaluating the inference quality of the learning model for one or more data included in the group identified in the clustering step. This makes it possible to evaluate the inference quality of the learning model for each range of cluster identified according to the inference result.

[0103] (Supplementary Note 4) The program according to Supplementary Note 3, wherein the data inference step (S103) is a step of acquiring error information regarding errors in inference results for at least a portion of the plurality of data, and the clustering step (S103) is a step of performing clustering on the data set acquired in the data acquisition step based on the similarity of the error information acquired in the data inference step, and identifying clusters formed by the clustering as groups. This makes it possible to evaluate the inference quality of the learning model for each range of clusters identified according to the similarity of errors in the inference results.

[0104] (Supplementary Note 5) The program according to Supplementary Note 4, wherein the data inference step (S103) is a step of acquiring error type information regarding the types of errors in the inference results for at least a portion of the plurality of data, and the clustering step (S103) is a step of performing clustering on the data set acquired in the data acquisition step based on the similarity of the error type information acquired in the data inference step, and identifying clusters formed by the clustering as groups. This makes it possible to evaluate the inference quality of the learning model for each range of clusters identified according to the similarity of the types of errors in the inference results.

[0105] (Supplementary Note 6) The program according to any one of Supplementary Notes 3 to 5, wherein a processor executes a group metadata identification step (S103) of identifying group metadata characterizing at least some of the groups based on metadata associated with one or more data items included in the groups identified in the clustering step, and a group metadata storage step (S103) of storing the group metadata identified in the group metadata identification step in association with at least some of the groups. This allows the contents of each cluster identified based on the inference result to be interpreted based on the group metadata. For example, for a cluster with poor inference quality, the cause of the poor inference quality can be interpreted based on the metadata.

[0106] (Supplementary Note 7) The program according to Supplementary Note 1 or 2, wherein the processor executes a clustering step (S103) of performing clustering based on metadata similarity on the data set acquired in the data acquisition step and identifying clusters formed by the clustering as groups, and the model evaluation step (S104) is a step of evaluating the inference quality of the learning model for one or more pieces of data included in the groups identified in the clustering step. This makes it possible to evaluate the inference quality of the learning model for each range of clusters identified according to the metadata similarity.

[0107] (Supplementary Note 8) The program according to Supplementary Note 7, wherein the processor executes a group quality evaluation step (S103) of acquiring an inference result for each of the one or more data sets by applying each of the one or more data sets included in the group identified in the clustering step as input data for the learning model stored in the model storage step, a group quality identification step (S103) of identifying, for at least some of the multiple groups, a group inference quality that characterizes the group based on the inference result acquired in the group quality evaluation step, and a group quality storage step (S103) of storing the group inference quality identified in the group quality identification step in association with at least some of the multiple groups. This makes it possible to evaluate the inference quality of the learning model for each range of clusters identified according to the similarity of metadata.

[0108] (Supplementary Note 9) The program according to Supplementary Note 8, wherein the group quality evaluation step (S103) is a step of acquiring error information regarding errors in inference results for each of one or more data, the group quality identification step (S103) is a step of identifying group error information that characterizes at least some of the multiple groups, and the group quality storage step (S103) is a step of storing the group error information identified in the group quality identification step in association with at least some of the multiple groups. This makes it possible to evaluate the inference quality regarding errors in the inference results of the learning model for each range of clusters identified according to the similarity of metadata.

[0109] (Supplementary Note 10) The program according to Supplementary Note 9, wherein the group quality evaluation step (S103) is a step of acquiring group error type information regarding the type of error in the inference result for each of one or more data, the group quality identification step (S103) is a step of identifying group error type information characterizing at least some of the multiple groups, and the group quality storage step (S103) is a step of storing the group error type information identified in the group quality identification step in association with at least some of the multiple groups. This makes it possible to evaluate the inference quality regarding the type of error in the inference result of the learning model for each range of clusters identified according to the similarity of metadata.

[0110] (Supplementary Note 11) The program according to any one of Supplementary Notes 1 to 10, wherein a processor executes a quality presentation step (S105) of presenting the multiple groups in association with the inference quality evaluated in the model evaluation step. This allows the quality of the entire learning model to be globally interpreted for each group range of the dataset. For example, it is possible to visually and intuitively confirm what proportion of the dataset has good inference quality and what proportion has poor inference quality.

[0111] (Supplementary Note 12) The program according to Supplementary Note 11, wherein the quality presentation step (S105) is a step of presenting a plurality of groups in association with an influence degree regarding the degree of influence on the learning model calculated based on the inference quality evaluated in the model evaluation step. This allows the quality of the entire learning model to be globally interpreted according to the degree of influence on the learning model for each group range of the dataset.

[0112] (Supplementary Note 13) The program according to any one of Supplementary Notes 1 to 12, wherein the processor executes a model correction step (S107) of creating a correction function by correcting the learning model stored in the model storage step based on one or more data included in the group. This makes it possible to create a correction function by correcting the learning model for each group.

[0113] (Supplementary Note 14) The program according to Supplementary Note 13, wherein the model correction step (S107) is a step of creating a correction function by correcting the learning model based on the inference quality of the group identified by applying one or more data included in the group as input data for the learning model. This makes it possible to create a correction function by correcting the learning model for a specific group whose inference quality is low, for example.

[0114] (Supplementary Note 15) The program according to Supplementary Note 13, wherein the processor executes a group selection step (S106) of accepting a selection of a predetermined group from a plurality of groups from a user, and a model correction step (S107) of creating a correction function by correcting the learning model stored in the model storage step based on one or more data included in the predetermined group selected in the group selection step. This makes it possible to create a correction function by correcting the learning model for the selected group in accordance with the group selection by the user.

[0115] (Supplementary Note 16) The program according to any one of Supplementary Notes 13 to 15, wherein the processor executes a condition storage step (S107) of storing application conditions for applying the modification function created in the model modification step in association with the modification function, and a combined model creation step (S108) of creating a combined model based on the learning model and the modification function, wherein the combined model created in the combined model creation step outputs output data using the modification function associated with the application condition if the input data is included in the application condition stored in the condition storage step, and outputs output data using the learning model if the input data is not included in the application condition stored in the condition storage step. This allows the combined model to be created by switching the model to be applied between the modification function and the learning model for each range of input data (range of group). The combined model can output output data with excellent inference quality for the input data.

[0116] (Supplementary Note 17) The program according to Supplementary Note 16, wherein a processor executes a group metadata identification step (S103) of identifying group metadata characterizing at least some of a plurality of groups based on metadata associated with one or more data included in the groups, and a group metadata storage step (S103) of storing the group metadata identified in the group metadata identification step in association with at least some of the plurality of groups. The condition storage step (S107) is a step of storing the group metadata stored in association with the group of the correction function created in the model correction step in association with the correction function as an application condition. The combined model created in the combined model creation step outputs output data based on the correction function associated with the group metadata if metadata associated with the input data is included in the group metadata, and outputs output data based on a learning model if metadata associated with the input data is not included in the group metadata. This allows the combined model to be created by switching the model to be applied between the correction function and the learning model for each range of metadata in the input data. The combined model can output output data with excellent inference quality for the input data.

[0117] (Supplementary Note 18) A method executed by a computer having a processor and a memory, wherein the processor executes all of the steps executed in any of the inventions according to Supplementary Note 1 to Supplementary Note 17. This allows the inference quality of a learning model to be evaluated for each group. For example, areas in which the learning model is weak can be identified as groups.

[0118] (Supplementary Note 19) An information processing device including a control unit and a storage unit, wherein the control unit executes all steps executed in any of the inventions according to Supplementary Note 1 to Supplementary Note 17. This allows the inference quality of a learning model to be evaluated for each group. For example, areas in which the learning model is weak can be identified as groups.

[0119] (Supplementary Note 20) A system comprising means for executing all steps performed in any of the inventions according to Supplementary Note 1 to Supplementary Note 17. This allows the inference quality of a learning model to be evaluated for each group. For example, areas in which the learning model is weak can be identified as groups.

[0120] 1 System, 10 Server, 101 Storage unit, 104 Control unit, 106 Input device, 108 Output device, 20 User terminal, 201 Storage unit, 204 Control unit, 206 Input device, 208 Output device

Claims

1. A program to be executed by a computer having a processor and a memory unit, wherein the processor executes the following steps: a model storage step of storing a learning model; a data acquisition step of acquiring a dataset consisting of a plurality of data, each of which is associated with metadata; and a model evaluation step of evaluating, for the dataset acquired in the data acquisition step, for each of a plurality of groups containing one or more of the data, the inference quality of the learning model for the one or more data included in the group.

2. The program according to claim 1, wherein the model evaluation step is a step of calculating the inference accuracy of the learning model.

3. The program of claim 1, wherein the processor executes: a data inference step of acquiring an inference result for each of the plurality of data included in the dataset acquired in the data acquisition step by applying each of the plurality of data included in the data set acquired in the data acquisition step as input data for the learning model stored in the model storage step; and a clustering step of performing clustering according to the inference result acquired in the data inference step and identifying clusters formed by the clustering as groups; and wherein the model evaluation step is a step of evaluating the inference quality of the learning model for the one or more data included in the group identified in the clustering step.

4. The program of claim 3, wherein the data inference step is a step of acquiring error information regarding errors in inference results for at least a portion of the plurality of data, and the clustering step is a step of performing clustering on the data set acquired in the data acquisition step based on the similarity of the error information acquired in the data inference step, and identifying clusters formed by the clustering as groups.

5. The program according to claim 4, wherein the data inference step is a step of acquiring error type information regarding the type of error in the inference result for at least a portion of the plurality of data, and the clustering step is a step of performing clustering on the data set acquired in the data acquisition step based on the similarity of the error type information acquired in the data inference step, and identifying clusters formed by the clustering as groups.

6. The program of claim 3, wherein the processor executes: a group metadata identification step of identifying group metadata characterizing at least some of the groups based on metadata associated with one or more data included in the groups identified in the clustering step; and a group metadata storage step of storing the group metadata identified in the group metadata identification step in association with at least some of the groups.

7. The program of claim 1, wherein the processor executes a clustering step in which clustering is performed on the data set acquired in the data acquisition step based on the similarity of the metadata and clusters formed by the clustering are identified as groups; and the model evaluation step is a step of evaluating the inference quality of the learning model for the one or more data included in the group identified in the clustering step.

8. The program of claim 7, wherein the processor executes the following steps: a group quality evaluation step of obtaining an inference result for each of the one or more data by applying each of the one or more data included in the group identified in the clustering step as input data for the learning model stored in the model storage step; a group quality identification step of identifying a group inference quality that characterizes at least a portion of the plurality of groups based on the inference result obtained in the group quality evaluation step; and a group quality storage step of storing the group inference quality identified in the group quality identification step in association with at least a portion of the plurality of groups.

9. The program of claim 8, wherein the group quality assessment step is a step of acquiring error information regarding errors in the inference results for each of the one or more data, the group quality identification step is a step of identifying group error information characterizing at least a portion of the multiple groups, and the group quality storage step is a step of storing the group error information identified in the group quality identification step in association with at least a portion of the multiple groups.

10. The program described in claim 9, wherein the group quality evaluation step is a step of obtaining group error type information regarding the type of error in the inference result for each of the one or more data, the group quality identification step is a step of identifying the group error type information characterizing at least a portion of the multiple groups, and the group quality storage step is a step of storing the group error type information identified in the group quality identification step in association with at least a portion of the multiple groups.

11. The program according to any one of claims 1 to 10, wherein the processor executes a quality presentation step of presenting the plurality of groups in association with the inference quality evaluated in the model evaluation step.

12. The program of claim 11, further comprising: a quality presentation step for presenting the plurality of groups in association with a degree of influence regarding the degree of influence on the learning model calculated based on the inference quality evaluated in the model evaluation step.

13. The program according to any one of claims 1 to 10, wherein the processor executes a model modification step of creating a modification function that modifies the learning model stored in the model storage step based on one or more data included in the group.

14. The program according to claim 13, wherein the model correction step is a step of creating the correction function that corrects the learning model based on the inference quality of the group identified by applying one or more data included in the group as input data to the learning model.

15. The program according to claim 13, wherein the processor executes: a group selection step of accepting, from a user, a selection of a specific group from among the plurality of groups; and the model modification step is a step of creating the modification function that modifies the learning model stored in the model storage step based on one or more data items included in the specific group selected in the group selection step.

16. The program of claim 13, wherein the processor executes a condition storage step of storing application conditions for applying the modification function created in the model modification step in association with the modification function, and a combined model creation step of creating a combined model based on the learning model and the modification function, wherein the combined model created in the combined model creation step outputs output data using the modification function associated with the application condition when input data is included in the application condition stored in the condition storage step, and outputs output data using the learning model when input data is not included in the application condition stored in the condition storage step.

17. The program of claim 16, wherein the processor executes a group metadata identification step of identifying group metadata characterizing at least some of the groups based on metadata associated with one or more data included in the groups, and a group metadata storage step of storing the group metadata identified in the group metadata identification step in association with at least some of the groups, wherein the condition storage step is a step of storing the group metadata stored in association with the group of the modification function created in the model modification step in association with the modification function as the application condition, and wherein the combined model created in the combined model creation step outputs output data by the modification function associated with the group metadata if metadata associated with input data is included in the group metadata, and outputs output data by the learning model if metadata associated with input data is not included in the group metadata.

18. A method implemented on a computer having a processor and a memory, the processor performing all of the steps performed in any one of claims 1 to 10.

19. An information processing device comprising a control unit and a storage unit, the control unit executing all of the steps executed in any one of the inventions according to claims 1 to 10.

20. A system comprising means for executing all the steps performed in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method of creating learned model, method of classifying data, computer and program

    JP2020052935A

  • Data processing method and apparatus for training depth information estimation model

    JP2023147276A

  • Active learning to reduce noise in labels

    US20190354810A1

  • Providing performance views associated with performance of a machine learning system

    US20200349466A1

  • Cluster targeting for use in machine learning

    US20230297886A1